Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Math Book Downloads / PDF originals

math21-2011 Kazdan Part I

PDF · 425 pages · 1.9 MB
Open PDF file

Lecture notes by J. Kazdan for a second-year Harvard course combining linear algebra with intermediate calculus, with a 1964 preface and a 1966 afterword. Chapters cover review of sets and reals, infinite series and power series, vector spaces, norms and inner products, Fourier series, and linear operators. This is a downloaded copy of someone else's text kept in Phil's math book folder.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
INTERMEDIATE CALCULUS AND LINEAR ALGEBRA Part I J. KAZDAN Harvard University Lecture Notes ii Preface These notes will contain most of the material covered in class, and be distributed before each lecture (hopefully). Since the course is an experimental one and the notes written before the lectures are delivered, there will inevitably be some sloppiness, disorganization, and even egregious blunders—not to mention the issue of clarity in exposition. But we will try. Part of your task is, in fact, to catch and point out these rough spots. In mathematics, proofs are not dogma given by authority; rather a proof is a way of convincing one of the validity of a statement. If, after a reasonable attempt, you are not convinced, complain loudly. Our subject matter is intermediate calculus and linear algebra. We shall develop the material of linear algebra and use it as setting for the relevant material of intermediate calculus. The first portion of our work—Chapter 1 on infinite series—more properly belongs in the first year, but is relegated to the second year by circumstance. Presumably this topic will eventually take its more proper place in the first year. Our course will have a tendency to swallow whole two other more advanced courses, and consequently, like the duck in Peter and the Wolf, remain undigested until regurgitated alive and kicking. To mitigate—if not avoid—this problem, we shall often take pains to state a theorem clearly and then either prove only some special case, or offer no proof at all. This will be true especially if the proof involves technical details which do not help illuminate the landscape. More often than not, when we only prove a special case, the proof in the general case is essentially identical—the equations only becoming larger. September 1964 iii Afterward I have now taught from these notes for two years. No attempt has been made to revise them, although a major revision would be needed to bring them even vaguely in line with what I now believe is the “right” way to do things. And too, the last several chapters remain unwritten. Because the notes were written as a first draft under panic pressure, they contain many incompletely thought-out ideas and expose the whimsy of my passing moods. It is with this—and the novelty of the material at the sophomore level—in mind, that the following suggestions and students’ reactions are listed. There are three categories, A), Material that turned out to be too difficult (they found rigor hard, but not many of the abstractions), B), changes in the order of covering the stuff, and C), material—mainly supplementary at this level—which is not too hard, but should be omitted if one ever hopes to complete the ”standard” topics within the confines of a year course. (A)It was too hard (unless one took vast chunks of time). (1) Completeness of reals. Only “monotone sequences converge” is needed for infinite series. (2) Term-by-term differentiation and integration of power series. The statement of the main theorem should be fully intelligible—but the proof is too complicated. (3) Cosets. This is apparently too abstract. It might be possible to do after finding general solutions of linear inhomogeneous O.D.E.’s. (4)L2and uniform convergence of Fourier series. Again, all I ended up doing was to try to state what the issues were, and not to attempt the proof. The ambitious student should be warned that my proof of the Weierstrass theorem is opaque (one should explicitly introduce the idea of an approximate identity). (5) Fundamental Theorem of Algebra. The students simply don’t believe inequalities in such profusion. (6) I you want to see rank confusion, try to teach the class how to compute higher order partial derivatives using the chain rule. That computation should be one of the headaches of advanced calculus. (7) Existence of a determinant function. I don’t know a simple proof except for the one involving permutations—and I hate that one. (8) Dual spaces. As lovely as the ideas are, this topic is too abstract, and to my knowledge, unneeded at this level where almost all of the spaces are either finite dimensional or Hilbert spaces. One should, however, mention the words “vector” and “covector” to distinguish column from row vectors. I forgot to do so in these notes and it did cause some confusion. (B)Changes in Order and Timing . The structure of the notes is to investigate bare linear spaces, then linear mappings between them, and finally non-linear mappings between them. It is with this in mind that linear O.D.E.’s came before nonlinear maps from Rn→R. The course ended by treating the simplest problem in the calculus of variations as an example of a nonlinear map from an infinite dimensional space iv to the reals. My current feeling is to consider linear andnon-linear maps between finite dimensional spaces before doing the infinite dimensional example of differential equations. The first semester should get up to the generalities on solving LX=Y, p. 319 [incidentally, the material on inverses (p. 355 ff) belongs around p. 319]. Most students find the material on linear dependence difficult—probably for two reasons: 1) they are not used to formal definitions, and ii) they think they have learned a technique for doing something, not just a naked definition, and can’t quite figure out just what they can do with it. In other words, they should feel these definitions about the anatomy of linear spaces are similar to those describing a football field and of little value until the game begins—i.e., until the operators between spaces make their grand entrance. Because of time shortages, the sections on linear maps from R1→RnandRn→R1, pp. 320-41 were regrettably omitted both years I taught the course. The notes were written so that these sections can be skipped. (C)Supplementary Material . A remarkable number of fascinating and important topics could have been included—if there were only enough time. For example: (1) Change of bases for linear transformations (including the spectral theorem). (2) Elementary differential geometry of curves and surfaces. (3) Inverse and implicit function theorems. These should be stated as natural gener- alizations of the problems of a) inverting a linear map, b) finding the null space of a linear map, and c) generalizing dim D(L) = dimR(L) + dimN(L) all to local properties of nonlinear maps via the tangent map. (4) Change of variable in multiple integration. Determinants were deliberately in- troduced as oriented volume to make the result obvious for linear maps and plausible for nonlinear maps. (5) Constrained extrema using Lagrange multipliers. (6) Line and surface integrals along with the theorems of Gauss, Green, and Stokes. The formal development of differential forms takes too much time to do here. Perhaps a satisfactory solution is to restrict oneself to line integrals and these theorems in the plane, where the topological difficulties are minimal. (7) Elementary Morse Theory. One can prove the Morse inequalities easily for the real line, the circle, the plane, and S2merely by gradually flooding these sets and observing the number of lakes and shore line changes only at the critical points. (8) Sturm-Liouville theory. An elegant fusion of the geometry of Hilbert spaces to differential equations. (9) Translation-invariant operators with applications to constant coefficient differ- ence and differential equations. The Laplace and Fourier transforms enter natu- rally here. (10) The Calculus of Variations. The formalism of nonlinear functionals on R/multicloseleft, i.e., mapsf:R/multicloseleft→R, generalizes immediately to nonlinear functionals defined on infinite dimensional spaces. v (11) The deleted rigor. (12) Linear operators with finite dimensional (perhaps even compact) range. One parting warning. When covering intermediate calculus from this viewpoint, it is all too natural to forget the innocence of the class, to enchant with glitter, and to numb with purity and formalism. Emphasis should be placed on developing insight and intuition along with routine computational facility. My classes found frequent reviews of the mathematical edifice, backward glances at the previous months’ work, not only helpful but mandatory if they were to have any conception of the vast canvas which was being etched in their minds over the course of the year. The question, “What are we doing now and how does it fit into the larger plan?” must constantly be raised and at least partially resolved. May, 1966 Contents 0 Remembrance of Things Past. 1 0.1 Sets and Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 0.2 Relations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 0.3 Mathematical Induction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 0.4 Reals: Algebraic and Order Properties . . . . . . . . . . . . . . . . . . . . . 7 0.5 Reals: Completeness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 0.6 Appendix: Continuous Functions and the Mean Value Theorem . . . . . . . 15 0.7 Complex Numbers: Algebraic Properties . . . . . . . . . . . . . . . . . . . . 22 0.8 Complex numbers: Completeness and Functions . . . . . . . . . . . . . . . 28 1 Infinite Series 33 1.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 1.2 Tests for Convergence of Positive Series . . . . . . . . . . . . . . . . . . . . 36 1.3 Absolute and Conditional Convergence . . . . . . . . . . . . . . . . . . . . . 41 1.4 Power Series, Infinite Series of Functions . . . . . . . . . . . . . . . . . . . . 43 1.5 Properties of Functions Represented by Power Series . . . . . . . . . . . . . 48 1.6 Complex-Valued Functions, ez,cosz,sinz. . . . . . . . . . . . . . . . . . . 65 1.7 Appendix to Chapter 1, Section 7. . . . . . . . . . . . . . . . . . . . . . . . 70 2 Linear Vector Spaces: Algebraic Structure 75 2.1 Examples and Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 75 a) The Space R2. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 75 b) The Space Rn. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 c) The Space C[a,b] . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77 d) D. The Space Ck[a,b] . . . . . . . . . . . . . . . . . . . . . . . . . . 77 e) E. The Space l1.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 78 f) F. The Space L1[a,b] . . . . . . . . . . . . . . . . . . . . . . . . . . . 78 g) G. The Space fn. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79 h) Appendix. Free Vectors . . . . . . . . . . . . . . . . . . . . . . . . . 80 2.2 Subspaces. Cosets. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 84 2.3 Linear Dependence and Independence. Span. . . . . . . . . . . . . . . . . . 88 2.4 Bases and Dimension . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93 3 Linear Spaces: Norms and Inner Products 101 3.1 Metric and Normed Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . 101 3.2 The Scalar Product in E2. . . . . . . . . . . . . . . . . . . . . . . . . . . . 107 3.3 Abstract Scalar Product Spaces . . . . . . . . . . . . . . . . . . . . . . . . . 113 vii viii CONTENTS 3.4 Fourier Series. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 132 3.5 Appendix. The Weierstrass Approximation Theorem . . . . . . . . . . . . . 140 3.6 The Vector Product in R3. . . . . . . . . . . . . . . . . . . . . . . . . . . . 146 4 Linear Operators: Generalities. V1→Vn,Vn→V1147 4.1 Introduction. Algebra of Operators . . . . . . . . . . . . . . . . . . . . . . . 147 4.2 A Digression to Consider au/prime/prime+bu/prime+cu=f. . . . . . . . . . . . . . . . . . 161 4.3 Generalities on LX=Y. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 170 4.4L:R1→Rn. Parametrized Straight Lines. . . . . . . . . . . . . . . . . . . 177 4.5L:Rn→R1. Hyperplanes. . . . . . . . . . . . . . . . . . . . . . . . . . . . 182 5 Matrix Representation 187 5.1L:Rm→Rn. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 187 5.2 Supplement on Quadratic Forms . . . . . . . . . . . . . . . . . . . . . . . . 210 5.3 Volume, Determinants, and Linear Algebraic Equations. . . . . . . . . . . . 217 a) Application to Linear Equations . . . . . . . . . . . . . . . . . . . . 234 5.4 An Application to Genetics . . . . . . . . . . . . . . . . . . . . . . . . . . . 243 5.5 A pause to find out where we are . . . . . . . . . . . . . . . . . . . . . . . . 246 6 Linear Ordinary Differential Equations 249 6.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 249 6.2 First Order Linear . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 252 6.3 Linear Equations of Second Order . . . . . . . . . . . . . . . . . . . . . . . 258 a) A Review of the Constant Coefficient Case. . . . . . . . . . . . . . . 258 b) Power Series Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . 259 c) General Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 266 6.4 First Order Linear Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . 278 6.5 Translation Invariant Linear Operators . . . . . . . . . . . . . . . . . . . . . 283 6.6 A Linear Triatomic Molecule . . . . . . . . . . . . . . . . . . . . . . . . . . 286 7 Nonlinear Operators: Introduction 293 7.1 Mappings from R1toR1, a Review . . . . . . . . . . . . . . . . . . . . . . 293 7.2 Generalities on Mappings from RntoRm. . . . . . . . . . . . . . . . . . . 295 7.3 Mapping from E1toEn. . . . . . . . . . . . . . . . . . . . . . . . . . . . 300 8 Mappings from EntoE: The Differential Calculus 309 8.1 The Directional and Total Derivatives . . . . . . . . . . . . . . . . . . . . . 309 8.2 The Mean Value Theorem. Local Extrema. . . . . . . . . . . . . . . . . . . 321 8.3 The Vibrating String. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 332 a) The Mathematical Model . . . . . . . . . . . . . . . . . . . . . . . . 333 b) Uniqueness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 334 c) Existence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 336 8.4 Multiple Integrals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 347 9 Differential Calculus of Maps from EntoEm, s. 361 9.1 The Derivative . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 361 9.2 The Derivative of Composite Maps (“The Chain Rule”). . . . . . . . . . . . 373 10 Miscellaneous Supplementary Problems 383 Chapter 0 Remembrance of Things Past. We shall treat a hodge-podge of topics in a hasty and incomplete fashion. While most of these topics should have been learned earlier, section 5 on the completeness of the real numbers has its more rightful place in advanced calculus. Do nottake time to read this chapter unless the particular topic is needed; then read only the relevant portions. The chapter is included for reference. 0.1 Sets and Functions Asetis any collection of objects, called the elements of the set, together with a criterion for deciding if an object is in the set. For example, I) the set of all girls with blue eyes and blond hair, and ii) the less picturesque set of all positive even integers. We can also define a set by bluntly listing all of its elements. Thus, the set of all students in this class is defined by the list in the roll book. Sets are often specified by a notation which is best described by examples. i)S={x:xis an integer }is the set of all integers. ii)T={(x,y):x2+y2= 1}is the set of all points ( x,y) on the unit circle x2+y2= 1 . iii)A={1,2,7,−3}is the set of integers 1 ,2,7 and −3 . Our attitude toward set theory will be extremely casual; we shall mainly use it as a language and notation. Without further ado, let us introduce some notation. x∈S, x is an element of the set S, or justxis inS. x/negationslash∈S, x is not an element of the set S. Z, the set of all integers, positive, zero, and negative. Z+, the set of all positive integers, excluding 0. R the set of all real numbers (to be defined more precisely later). C, The set of all complex numbers (also to be defined more precisely later). ∅, the set with no elements, the empty ornullset. It is extremely uninteresting. Definition: Given the two sets SandT, i) the set S∪T, “SunionT”, is the set of elements which are in eitherSorT, or both. ii) The set S∩T, “Sintersection T”, is the set of elements in bothSandT. If we represent Sby one blob and Tby another, S∪Tis the shaded region while S∩Tis the cross-hatched region. Note that all elements in S∩Tare also in S∪T. Two sets are disjoint ifS∩T=∅, that is, if their intersection is empty. 1 2 CHAPTER 0. REMEMBRANCE OF THINGS PAST. Asubset of a set is another way of referring to a portion of a given set. Formally, Ais the subset of S, writtenA⊂S, if every element in Ais also an element of S. The set Ais a subset of the set Sif and only if either A∪S=S,or, equivalently, A∩S=A. It is possible that A=S, or thatA=∅. If these degenerate cases are excluded, we say thatAis aproper subset ofS. Given the two sets SandT, it is natural to form a new set S×T, “ScrossT”, which consists of all pairs of elements, one from Sand the other from T. For example, if Sis the set of all men in this class, and Tthe set of all women in this class, then S×T is the set of all couples, a natural set to contemplate. Ifx∈Sandy∈T, the standard notation for the induced element in S×Tis (x,y) . Note that the order in ( x,y) is important. The element on the left is from S, while that on the right is from T. For this reason the pair of elements ( x,y) is usually called an ordered pair. The whole set S×Tis called the product ,direct product , orCartesian product ofS andT, all three names being used interchangeably. You have met this idea in graphing points in the plane. Since these points, ( x,y) , are determined by an ordered pair of real numbers, they are just the elements of R×R. From this example it is clear that even though this set R×Ris the product of a set with itself, theorder of the pair ( x,y) is still important. For example the point (1 ,2)∈R×Ris certainly not the same as (2 ,1)∈R×R. Having defined the direct product of two sets SandTas ordered pairs, it is reasonable to define the direct product of three sets S, T, andUas the set of ordered triplets ( x,y,z ) , wherex∈S, y∈T, andz∈U. The extension to nsets,S1×S2× ··· ×Sn, is done in the same way. Let us now recall the ideas behind the notion of a function. Afunctionffrom the set Xinto the set Bis a rule which assigns to every x∈X one and only one element y=f(x)∈B. We shall also say that fmapsXintoB, and write either f:X→B,orXf→B. This alternative notation is useful when XandBare more important than the specific nature off. The setXis the domain off, while the range offis the subset Y⊂Bof all elements y∈Bwhich are the image of (at least) one point x∈X, soy=f(x) , or in suggestive notation, Y=f(X). Automobile license plates supply a nice example, for they assign to every license plate sold a unique car. The domain is the set of all license plates sold, while the range is not all cars, but rather the subset of all cars which are driven. Wrecks and museum pieces neither need nor have license plates since they are not on the roads. Some other examples are i) the function f(n)≡1 n, n= 1,2,3,... which assigns to every n∈Z+the rational number 1 n, and ii) the function f(n,m) =m n, n, m = 1,2,2,..., which assigns to every element of Z+×Z+the rational numberm n. Quite often we shall use functions which map part of some set into part of some other set. In other words the function may be defined on only a subset of a given set and take on values in a subset of some other set. The function f(n,m)≡m nof the previous paragraph is of this nature for we defined it on a subset of Z×Zand takes its values on the positive subset of the set of all rational numbers. 0.1. SETS AND FUNCTIONS 3 There is some standard nomenclature (or $10 words if you like) associated with map- ping. Say X⊂Aand the function f:X→B. Note that we know the definition of f only onX. It may not be defined for the remainder of A. Definition: i) ifevery element of Bis the image of (at least) one point in X, the map fis called surjective oronto. In other words f:X→Bis a surjection if the range of f is all ofB. Thusfis always surjective onto its range. ii) If the map fhas the property that for every x1,x2∈X, we havef(x1) =f(x2) when and only when x1=x2, the map is called injective orone to one (1-1). This is the case if no two different elements in Xare mapped into the same element in B. iii) If the map fis both surjective and injective, that is, if it is both onto and 1-1, then fis called bijective . Examples: For these, we have f:X→BwhereX=B=Z. (1) The map f(n) = 2nis injective but not surjective since the range does not contain the odd integers in B. (2) The map f(n) =/braceleftbiggn 2ifnis even n+1 2ifnis oddis surjective but not injective since every ele- ment inBis the image of two distinct elements of X. (3) The map f(n) =n+ 7 is bijective. Notational Remark : For functions whose domain is ZorZ+it is customary to indicate the element of the range by a notation like aninstead off(n) . Thusf(n) =1 n, where n∈Z+, is written as an=1 n. Such a function is usually called a sequence . The concepts we have just defined are useful if we try to define what we mean by the inverse of a function. Definition: A function f:X→Bisinvertible if to every b∈Bthere is one and only onex∈Xsuch thatb=f(x) . Thusfis invertible if and only if it is bijective. If fis invertible, we denote the inverse function by f−1, sox=f−1(b) . Iff:A→B, andg:B→C, then when composed (put together) these two functions induce a mapping, g◦f, ofAintoC. Slightly more generally, if B⊂R, andf:A→B whileg:R→C, th eng◦f:A→C. You should be able to see why the composed map g◦fis only defined on A, and then understand that our stipulation that B⊂Ris a convenient requirement. Ifx∈Aandz∈C, theng◦fmapsxontoz= (g◦f)(x) , or in more familiar notation,z=g(f(x)) . Now an example. Say the distance syou have walked at time tis specified by the function s=f(t) , and the amount zof shoe leather worn out by walking the distance sis given by the function z=g(s) . Then the amount of shoe leather you have worn out at time tis given by the composed function z=g(f(t)) . Heret∈A, s∈B, andz∈C. Hopefully you have by now recognized that the “chain rule” for derivatives is just the procedure for finding the derivative of composed functions from their constituent parts. In our example the chain rule would be used to finddz dtfromdg dsandd f dt-if these functions were differentiable. We conclude this section with more symbols—if you have not yet had enough. These are borrowed from logic. Although we shall use them only infrequently as a shorthand, they might have greater use to you in class notes. ∀ “for every” 4 CHAPTER 0. REMEMBRANCE OF THINGS PAST. ∃ “there is”, or “there exists” /owner “such that” A⇒B“the truth of statement Aimplies that of statement B”. A⇔B“statement Ais equivalent to statement B, that is, both A⇒Band B⇒A. Exercises (1) IfR={1,4}, S={1,2,3,4,}, andT={2,3,7}, find the six other sets R∪S, R∩ S, R∪T, R∩T, S∪T,andS∩T. Which of these nine sets are proper subsets of which other sets? (2) IfS={x:|x−1| ≤2}andT={x:|x| ≤2}, findS∪TandS∩T. A sketch is adequate. (3) IfA, B , andCare any subsets of a set S, prove (a) (A∪B)∪C=A∪(B∪C) —so that the parenthesis can be omitted without creating ambiguity. (b) (A∩B)∩C=A∩(B∩C) —so that again the parentheses are superfluous. (c) (A∪B)∩C= (A∩C)∪(B∩C). (d) (A∩B)∪C= (A∪C)∩(B∪C). Remark: two setsXandYare proved equal by showing that both X⊂Yand Y⊂X. (4) If the function fhas domain S, and both A⊂CandB⊂S, prove that (i)A⊂B⇒f(A)⊂f(B) . (ii)f(A∩B)⊂f(A)∩f(B) [We cannot hope to prove equality because of coun- terexamples like: let A={−2,−1,0,1,2,3}andB={−4,−3,−2,−1}. Then with f(n) =n2, we havef(A) ={0,1,4,9}, f(B) ={1,4,9,16}, and f(A∪B) ={1,4} /negationslash=f(A)∩f(B) ]. (iii)f(A∪B) =f(A)∪f(B) . (5) For the following functions f:X→B, classify as to injection, surjection, or bijection, or none of these. (i)f(n) =n2withX=Z+andB=Z. (ii) LetX={all rational numbers },B={all rational numbers }, andf(x) =1 m, wherex=n m∈X[Heren mis assumed to be reduced to lowest terms.] (iii)f(x) =1 x, wherex∈XandX=B={all positive rational numbers }. (iv)X={all women born in May }, B={the thirty days in the month of June }, and letfbe the function assigning “her birthday” to each woman born in June. (v)f(n) =|n|, withX=B=Z. 0.2. RELATIONS 5 0.2 Relations A relationship often exists between elements of sets. Some common examples are i) a≥b, ii)a⊥b(perpendicular to), iii) alovesb, and iv)a/negationslash=b. LetSbe a given set, a, b∈S, and let Rbe a relation defined on S(that is, ∀a, b∈S, eitheraRboraRbwith no third alternative possible). Most relations have at least one of the following properties. (i)reflexiveaRa ∀a∈S (ii)symmetric aRb⇒bRa (iii) transitive (aRbandbRc)⇒aRc. Examples: (1) perpendicular ( ⊥) is only symmetric. (2) “loves” enjoys none of these (well, maybe it is reflexive). (3) equality ( = ) has all three properties. (4) geometric congruence ( ∼=) and geometric similarity ( /similarequal) both have all three. (5) parallel ( /bardbl) has all three—if we are willing to agree that a line is parallel to itself. (6) “is less than five miles from” is only reflexive and symmetric. (7) fora, b∈Z+, the relation “ ais divisible by b” is only reflexive and transitive but not symmetric. (8) “less than” ( <) is only transitive. A relation which is reflexive, symmetric and transitive is called an equivalence relation . The standard examples are those of algebraic equality and of geometric congruence. An equivalence relation on a set Spartitions the set into subsets of equivalent elements . Those terms are illustrated in the following. Examples: (1) In the set Sof all triangles, the equivalence relation of geometric congruence parti- tionsSinto subsets of congruent triangles, any two triangles of Sbeing in the same subset (or equivalence class as it is called) if and only i f they are congruent. (2) In the set Pof all people, consider the equivalence relation ”has the same birthday,” disregarding the year. This relation partitions Pinto 366 equivalence classes. Two people are in the same equivalence class if their birthdays fall on the same day of the year. Notice that any two equivalence classes are either identical or disjoint, that is, they have either no elements in common or they coincide. This is particularly clear from the examples with birthdays. By the fundamental theorem of calculus, we know that the indefinite integral of an integrable function fcan be represented by any function Fwhose derivative is f. The 6 CHAPTER 0. REMEMBRANCE OF THINGS PAST. mean value theorem told us that every other indefinite integral of fdiffers from Fby only a constant. Thus, the indefinite integrals of a given function are an equivalence class of functions, differing from each other by constants. The equivalence relation is “equal up to an additive constant”. Exercises (1) Ifa, b, c, d ∈Z+, let us define the following equivalence relation between the elements ofZ+×Z+: (a,b)R(c,d) if and only if ad=bc. Verify that Ris an equivalence relation. [In real life, the pair ( a,b) of this example is written asa b, so all we have said isa b=c dif and only if ad=bc. This equivalence relation partitions the set of rational numbers into very familiar equivalence classes. For example the equivalent rational numbers1 2,2 4,3 6,...are in the same equivalence class, to no one’s surprise]. (2) Explain the fallacy in the following argument by observing that equality “ = ” here is not the usual algebraic equality , but rather some other equivalence relation. “Let =/integraltextdx x. Integration by parts ( p= 1/x,dq=dx), gives A=x(1 x)−/integraldisplay x(−1 x2)dx= 1 +A. Hence 0 = 1 .” 0.3 Mathematical Induction You are familiar with a variety of proofs, viz. direct proofs and proofs by contradiction. There is, however, another type of proof which is not encountered very often in elementary mathematics: proof by induction . Abstractly, you have a sequence of statements P1, P2, P3,..., and a guess for the nature of the general statement Pn. A proof by mathematical induction provides a method for showing the general statement Pnis correct. Here is how it is carried out. First verify that the statement is true in some special case, say for n= 1 , so you check the validity of P1.Second you show that ifit is true in some particular case n=k, then it is true for the next case n=k+ 1 , that is, Pk⇒Pk+1. Now since P1is true, so is P1+1=P2, and consequently so is P2+1=P3, and so on up. Observe that the procedure does not tell you how in the world to guess the general statement Pn, but only shows how to verify it. Let us carry out the procedure for an example. We guess the formula 1 + 2 + ···+n=n(n+ 1) 2(0-1) step 1. Is the formula true for n= 1 ? Yes, since both sides then equal 1. step 2. Assuming the formula is true for n=k, we must show this implies the formula is true forn=k+ 1 . 1 + 2 + ···+k+ (k+ 1) =(k+ 1)(k+ 2) 2. 0.4. REALS: ALGEBRAIC AND ORDER PROPERTIES 7 The formula, assumed to be true, for n=kis 1 + 2 + ···+k=k(k+ 1) 2. Adding (k+ 1) to both sides we find that 1 + 2 + ···+k+ (k+ 1) =k(k+ 1) 2+ (k+ 1) =(k+ 1)(k+ 2) 2 which is exactly the statement we wanted. This proves that formula (0.3) is true for all n≥1 . Exercises Use mathematical induction to prove the given statements. (1) 12+ 22+···+n2=n(n+1)(2 n+1) 6 (2)d dx(xn) =nxn−1(use the formula for the derivative of a product). (3) LetI(n) =/integraltextπ 2 0sinnxdx (a) Prove the following formula is correct when nis an odd integer ≥3 , I(n) =2·4·6· · · ·(n−1) 1·3·5· · · ·n (b) Guess and prove the formula when nis an even integer ≥2 . (4) Let Γ(s) =/integraltext∞ 0e−tts−1dt, wheres>0 (this is the famous gamma function ). (a) Show Γ( s+ 1) =sΓ(s) (Hint: integrate by parts) (b) Ifn∈Z+, guess and prove the formula for Γ( n+ 1) . 0.4 The Real Numbers: Algebraic and Order Properties. The set of all real numbers can be characterized by a set of axioms. These properties are of three different types, i) algebraic properties, ii) order properties, and iii) the completeness property. Of these, the last is by far the most difficult to grasp. But that is getting ahead of our story. Let Sbe a set with the following properties. I. Algebraic Properties A. Addition. To every pair of elements a,b∈S, is associated another element, denoted bya+b, with the properties A - 0. (a+b)∈S A - 1. Associative: for every a,b,c∈S,a+ (b+c) = (a+b) +c. A - 2. Commutative: a+b=b+a A - 3. There is an additive identity , that is, an element ”0” ∈Ssuch that 0 + a=a for alla∈S. A - 4. For every a∈S, there is also a b∈Ssuch thata+b= 0 .bis the additive inverse ofa, usually written −a. 8 CHAPTER 0. REMEMBRANCE OF THINGS PAST. M. Multiplication. To every pair a,b∈S, there is associated another element, denoted byab, with the properties M - 0.ab∈S M - 1. Associative. For every a,b,c∈S, a(bc) = (ab)c. M - 2. Commutative. ab=ba. M - 3. There is a multiplicative identity , that is, an element “ l”∈Ssuch thatla=a for alla∈S. Moreover 1 /negationslash= 0 . M - 4. For every a∈S, a/negationslash= 0 , there is also a b∈Ssuch thatab= 1 .bis the multiplicative inverse ofa, usually written1 aora−1. D. Connection between Addition and Multiplication . D - 1. Distributive . For every a,b,c∈S, a(b+c) =ab+ac. Some sample - and simple—consequences of these nine axioms are i) a+ 0 =a, ii) a·1 =a, and iii)a+b=a+c⇒b=c. Any set whose elements satisfy the axioms A-0 to A-4 is called a commutative (or abelian )group . The group operation here is addition. In this language, we see that the multiplication axioms just state that the elements of S—with the additive identity 0 excluded—also form a commutative group, with the group operation being multiplication. These additive and multiplicative structures are connected by the distributive axiom. Most of high school algebra takes place in this setting; however, the possibility of non-integer exponents is not yet specifically included; in particular the square root of an element of S is not necessarily also in S. Our axioms, or some part of them, are satisfied by sets other than the real numbers. The set of even integers form a commutative group with the group operation being addition, while numbers of the form 2n, n∈Z, form a commutative group under multiplication. The set of rational numbers satisfies all nine axioms. Any such set which satisfies all nine axioms is called a field. Both the real numbers and the rational numbers (a subset of the real numbers) are fields. A more thorough investigation of groups and fields is carried out in courses in modern algebra. II. Order Axioms Besides the above algebraic rules, we shall introduce an order relation , intuitively, the notion of ’greater than”. To do this we need to use an undefined concept of positivity for elements of Sand use it to state our axioms. O -1. Ifa∈Sandb∈Sare positive, so are a+bandab. O -2. The additive identity 0 is not positive. O - 3. For every a∈S, a/negationslash= 0 , either aor−ais positive, but not both. If −ais positive, we shall say that ais negative. Trichotomy Theorem . For any two numbers a,b∈S, exactly one of the following three statements is true, i) a−bis positive, ii) b−ais positive, or iii) b−ais zero. If the notationa < b is used to mean “ b−ais positive,” and a > b meansb < a , then this theorem reads, either a>b, a<b, ora=b. The proof—which you should do—is a simple consequence of our axioms. Some other consequences are a<b andb<c⇒a<c (transitivity of “ <”) a<b andc>0⇒ac<bc a/negationslash= 0⇒a2>0 . (Since 1 = 12, this implies 1 >0 ). The set of rational numbers as well as the set Rof real numbers satisfy all twelve axioms. Any set which satisfies these twelve axioms is called an ordered field . 0.5. REALS: COMPLETENESS 9 Exercises (1) LetTbe a set whose elements are of the form a+b√ 2 , whereaandbare rational numbers (and so are elements of a field). Show that Tis also a field. (2) Consider the set of all integers Zwith the following equivalence relation: m∈Z andn∈Zare equivalent if they have the same remainder when divided by 2. The notation for this equivalence is m≡n(mod2) This equivalence relation partitions Zinto two equivalence classes which we may denote respectively by 0 if the number is even, and 1 if the number is odd . Thus 8≡ −22 (mod 2) and 7 ≡13 (mod 2). Prove that the set Zwith ordinary addition and multiplication but with this equivalence relation forms a field. (3) Prove the trichotomy theorem. (4) Prove that if a/negationslash= 0 , thena2>0 . Use it to prove that 1 >0 and then to conclude that all of the ’positive integers” are, in fact, positive. 0.5 The Real Numbers: Completeness Property. III. Completeness Axiom. So far our axioms do not insure that we can take fractional powers like the square root, of an element of an ordered field Sand still obtain an element of the same field. The issue here is not merely that of fractional powers or other algebraic operations, but a more serious one. Imagine the (as yet undefined) real number line. Although the rational numbers are an infinite number of points on the line, there are many “holes” between the rationals. We already know of one “hole” at√ 2 , there is another at√ 3 , atπ, and ate. In fact, in a sense which can be made precise, almost all of the points on the real number line represent irrational numbers. The completeness axiom is designed to eliminate the possibility of ”holes” in the real number line. It does so by more or less bluntly stating that there are no holes. This is the “Dedekind cut” form of the completeness axiom. we have chosen it over other equivalent axioms because it is easy to visualize—even though the “Cauchy sequence” form is perhaps preferable for more advanced analysis courses. A definition is needed before the axiom can be stated. Definition: LetS1andS2be subsets of an ordered field S. Then the set S1precedes S2if for every a∈S1andb∈S2, we havea≤b. If you imagine the real number line, “ S1precedesS2” should be thought of as meaning that all of S1is to the left of all of S2.S1andS2of course might touch, or might just miss touching. Completeness Axiom. LetS1andS2be nonempty subsets of an ordered field S. If S1precedesS2, then there is at least one number c∈Ssuch thatcprecedesS2and is preceded by S1. In other words, there is (at least) one element of SbetweenS1andS2. Definition: The set of real numbers ,R, is a set which satisfies the above axioms of algebra, order, and completeness. Thus, the real numbers is a complete ordered field. This type of definition of Ramounts to saying “we don’t know or care what the real numbers are, but in any event they have the required properties.” If we had used the 10 CHAPTER 0. REMEMBRANCE OF THINGS PAST. Cauchy sequence version of the completeness axiom, we would have begun the rational numbers—which we do know—and then defined the real numbers as the set of limits of rational numbers. This would have been somewhat more concrete, but would have involved the difficult concept of limit before we even get off the ground. From the picture associated with the completeness axiom, we see that it exactly states that the real number line has no holes, for - emotionally speaking—if there were a hole, let S1be the set of real numbers to the left of the hole, and S2the s et to the right of the hole. Then there would be no real number between S1andS2, since the hole is there, contradicting the completeness axiom. Let us use the idea of the last paragraph to show that the rational numbers, an ordered field, are notcomplete by exhibiting two sets, one preceding the other, which have no rational number between them. Just let S1={x:x>0, x2<2}andS2={x:x>0, x2>2}. The only possible number between S1andS2is√ 2 —which is irrational. This construc- tion is just what we need to prove the following sample. Theorem 0.1 Every non-negative real number a∈Rhas a unique non-negative square root. Proof: Ifa= 0 , then 0 is the square root. If a > 0 , letS1={x:x > 0, x2< a} andS2={x:x > 0, x2> a}. We first show that neither S1norS2is empty. Since (1 +a 2)2= 1 +a+a2 4>a, we know that (1 +a 2)∈S2, soS2/negationslash=∅. Also (a 1+a 2)2<a(check this) so thata 1+a 2∈S1and hence S1/negationslash=∅. BecauseS1precedesS2, by the completeness axiom there is a c∈RbetweenS1andS2. Notice that c >0 , sincecis preceded by S1. It remains to show that c2=a. By the trichotomy theorem, either c2> a, c2< a, orc2=a. The first two possibilities will be shown to give contradictions. If c2>a, since a < (c2+a 2c)2< c2, we see thatc2+a 2c∈S2an d precedes c2, contradicting the property specified in the completeness axiom that c2precedes every element of S2. Similarly the assumption c2< a, with the inequality c2<(2ac c2+a)2< a, leads to a contradiction. The only remaining possibility is c2=a, which shows that cis the desired positive square root ofa. Let us now prove that the positive square root cofais unique. Assume that there are two positive numbers c1andc2such that both c2 1=aandc2 2=a. Then 0 =c2 1−c2 2= (c1−c2)(c1+c2) Sincec1+c2>0 , we conclude that c1−c2= 0 , soc1=c2, completing the proof of the theorem. Definition: The real number Mis an upper bound for the set A⊂Rif for every a∈A, we havea≤M. The number µ⊂Ris aleast upper bound (l.u.b) forAifµis an upper bound for Aand no smaller number is also an upper bound for A.Lower bound and greatest lower bound (g.l.b) are defined similarly. A set A⊂Risbounded if it has both upper and lower bounds. Theorem 0.2 Every non-empty bounded set A⊂Rhas both a greatest lower bound and a least upper bound. 0.5. REALS: COMPLETENESS 11 Proof: Observe first that this theorem utilizes the completeness property in that without it, there might have been a ”hole” just where the g.l.b. and l.u.b. should be. Since the proofs for the g.l.b. and l.u.b. are almost identical we only prove there is a g.l.b. Let S1={x:xprecedesA},andS2=A. By hypothesis S2/negationslash=∅. SinceAis bounded, it has a lower bound m, m ∈S1soS1/negationslash=∅. By the completeness axiom, there is a c∈RbetweenS1andS2. It should be obvious thatcis both greater than or equal to every element of S1, and less than or equal to every element of S2- so it is the required g.l.b. Definition: Theclosed interval [a,b] is the set {x∈R:/Gmir≤/archleftdown≤ }. Theopen interval (a,b) is the set {x∈R:/Gmir</archleftdown<}. All we can do is apologize for the multiple use of the parentheses in notation. Please note that sets are not like doors. Some sets, like ( a,b) ={x∈R:/Gmir≤/archleftdown<}are neither open nor closed. Theorem 0.3 (Nested set property). Let I1,I2,... be a sequence of non-empty closed bounded intervals, In={x:an≤x≤bn}, which are nested in the sense I1⊃I2⊃I3..., so each covers all that follow it. Then there is at least one point c∈Rwhich lies in all of the intervals, that is, cis in their intersection c∈ ∩∞ k=1Ik. Proof: LetS1={x:xprecedes some In,and so allIk, k≥n} S2={x:xpreceded by some In,and so allIk, k≥n}. First, neither S1norS2are empty since a1∈S1andb1∈S2. Thus by the completeness axiom, there is at least one c∈RbetweenS1andS2. Thiscis the required number (complete the reasoning). If the intervals Ikdo not get smaller after, say INbecauseaN=aN+1= . . . and bN=bN+1= . . . , then the whole interval aN≤x≤bNis caught by the preceding argument. The more common case is there the ak’s strictly increase and the bk’s strictly decrease. This is what happens when approximating a real number to successively greater accuracy by the decimal expansion. In the case of√ 2 for example, I1={x: 1≤x≤2}, I2={x: 1.4≤x≤1.5}, I3={x: 1.41≤x≤1.42}, I4={x: 1.414≤x≤1.415}, and so on, gradually squeezing down on√ 2 to any desired accuracy. Definition: The sequence an∈R,/multicloseleft=/notforces,/notsatisfies,.... of real numbers converges to the real numbercif, given any /epsilon1>0 , there is an integer Nsuch that |an−c|</epsilon1for alln>N . We will then write an→c. [In practice no confusion arises for the use of →to denote both convergence and mappings (cf. 1)]. Again ordinary decimals supply an example, for they allow us to get arbitrarily close to any real number. We could have defined the real numbers as all decimals; however there would be a mess avoiding the built-in ambiguity illustrated by 1 .9999....= 2.0000.... Theorem 0.4 Under the hypotheses of the previous theorem, if in addition the length of Intends to zero, (bn−an)→0, then the number c∈Rfound is unique. Furthermore, if uk∈Ikfor allk, that is if ak≤uk≤bk, thenuk→ctoo. Proof: Suppose there were two real numbers cand ˜cin all of the intervals, ak≤c≤bkandak≤˜c≤bkfor allk. 12 CHAPTER 0. REMEMBRANCE OF THINGS PAST. Rewriting the second inequality as −bk≤ −˜c≤ −ak, and adding this to the first inequality, we find that ak−bk≤c−˜c≤bk−ak. Since both sides of this inequality tend to zero, if c−˜c/negationslash= 0 , we would have a contradiction. To proveuk→c, repeat the above reasoning with ˜ creplaced by uk. We find that ak−bk≤c−uk≤bk−ak. Again both sides of this inequality tend to zero. Now let us fiddle with the /epsilon1,N definition of limit t o complete the proof. Since bn−an→0 , given any/epsilon1>0 , there is an Nsuch that |an−bn|</epsilon1for alln>N . Thus for any /epsilon1>0 and the sameN,|un−c|</epsilon1forn>N , which is the definition of un→c. Theorem 0.5 Bolzano-Weierstrass . Every infinite sequence of real numbers {uk}in a bounded interval Ihas at least one subsequence which converges to a number c∈R. Proof: This one is very clever and picturesque. Watch. Bisect Iinto two intervals I1 and ˜I1of equal length. At least one of I1or˜I1must contain an infinite number of the {uk}’s. Continuing in this way we obtain a set of nested intervals I⊃I1⊃I2⊃. . . each of which have an infinite number of the {uk}’s, and the length of Intending to zero. From Theorem 3 we conclude that there must be a c∈Rcommon to all of the intervals. We must now select the subsequence {ukn}of the {uk}’s which converge to c. Since eachIncontains an infinite number of points of the sequence, we can certainly pick one, sayukn∈In. This sequence {ukn}satisfies the hypotheses of Theorem 4. Thus ukn→c. Remarks: 1. If we also assume Iis closed, then we can further assert that c∈I. IfIis not closed, cmay be an end point /ownerI. 2. If a sequence ukconverges to a c∈R, then every infinite subsequence uknalso converges, and to the same number c. Theorem 0.6 . If the sequence {uk}converges, it is bounded. Proof: Sayuk→α, and let/epsilon1= 1 in the definition of convergence. Then there is an N such that |un−α|<1 for alln>N . Thus, when n>N , |un|=|un−α+α| ≤ |un−α|+|α|<1 +|a| Therefore for any kthe number |uk|is bounded by the largest of the N+1 numbers |u1|, |u2|, . . . , |uN|and (1 + |a|) . The following theorem shows how to handle algebraic combinations of convergent se- quences. Theorem 0.7 Ifan→αandbn→β, then i)an+bn→α+β ii)anbn→αβ iii)an bn→α βif bothbn/negationslash= 0, for alln, and ifβ/negationslash= 0. Proof: Since the proofs are all similar, we only prove ii). Observe that |anbn−αβ|=|(anbn−αbn) + (αbn−αβ)| ≤ |an−α||bn|+|alpha||bn−β| By Theorem 6, the |bn|’s are bounded, say by B. Sincean→α, given any ε>0 , there is anN1such that |an−α|<ε 2Bfor alln>N 1, and since bn→β, for the same εthere 0.5. REALS: COMPLETENESS 13 is anN2such that |bn−β|<ε 2|α|for alln>N 2. Thus, ifnis greater than the larger of N1andN2, n> max(N1,N2) , we find that |anbn−αβ|<ε 2+ε 2=ε which does the job. Definition: The sequence a1, a2,...of real numbers is said to be monotone increasing if a1≤a2≤a3≤..., and monotone decreasing a1≥a2≥a3≥.... Both kinds are called monotone sequences. Theorem 0.8 Every bounded monotone sequence a1, a2,... of real numbers converges. In other words, there is an αεRsuch thatan→α. Proof: We assume the sequence is increasing. The proof for decreasing sequence is iden- tical. Since the sequence is bounded, by Theorem 2 it has a least upper bound αεR. We maintainan→α. Given any ε>0 , we know that for all n, a n<a+εbecauseαis an upper bound. Since α−ε<α , andαis the l.u.b. of the sequence, we can find an Nsuch thatα−ε<a N. But then, because the sequence is increasing α−ε<a nforalln≥N. Thus for all n≥N, a−ε<a n<a+ε; that is, |an−α|<εfor alln≥N, proving the convergence to α. We shall close this difficult section with a wonderful procedure for computing the square root of a positive real number. I use it all of the time. It is much easier to understand than the hair-raising method taught in public school. Theorem 0.9 For any positive real numbers Aanda0the infinite sequence defined by an+1=1 2(an+A an), n= 0,1,2,..., (0-2) is monotone decreasing and converges to√ A. Moreover, if we let bn=A an, then thebn’s are monotone increasing and also converge to√ A: ba≤b2≤...≤√ A≤...≤a2≤a1 Proof: We first show that a2 k≥Aand thatak+1≤ak, a2 k−A=1 4(ak−1+A ak−1)2−A=1 4(ak−1+A ak−1)2≥0,soa2 k≥A. From this, it is easy to see that ak+1≤ak, for ak−ak+1=ak−1 2(ak+A ak) =a2 k−A 2ak≥0. Thusa1≥a2≥ ···.≥√ A. That thea2 kconverge is an immediate consequence of Theorem 8, since the sequence {a2 k}is a bounded (by A) monotone decreasing sequence. Denoting the limit by α, a2 k→α, the proof that α=Ais identical to th e reasoning which gave a unique limit in Theorem 4. Sincebn=A an, and thean’s decrease and are ≥√ A, then thebn’s increase and are ≤√ A. This also shows that bn≤an. Sincean→√ A, we havebn=A an→√ Atoo. 14 CHAPTER 0. REMEMBRANCE OF THINGS PAST. Application: We compute√ 8 . Takea0= 3 . Then a1=1 2(3 +8 3) =17 6, and b1= 8·6 17=48 17. Similarly, a2=577 204, b2=1632 577. This gives1632 577≤√ 8≤577 204, or in decimal form 2.82842<√ 8<2.82843, astounding accuracy after only two steps. I carried the computations one step further and found 2.828427124....≤√ 8≤2.828427124...., where the dots indicate I gave up on the arithmetic, having obtained the exact value as far as the approximation went. Digital computers use this method and related ones for similar computations. It is particularly well adapted to them (and me) since only simple arithmetic operations are involved. This Theorem 9 gives another proof that every positive real number has a unique positive square root. It is valuable to compare this proof with that of Theorem 1. The main distinction is that the second proof just given is constructive it actually shows a way to compute successive approximations to the square root of any number. However, you are justified in asking how we ever found the procedure of equation (0.9) in the first place. The secret is that this formula is a statement of Newton’s method for finding roots off(x) = 0 , applied to the particular function f(x) =x2−A. See most calculus books for more information about this method. Hopefully, we will have time to discuss this topic later, for it is a constructive way o f proving the existence of a sought after object. The standard existence theorem for ordinary differential equations is a close relative of Newton’s method. Exercises (1) For the sequences defined below, find which converge, which do not converge but do have at least one convergent subsequence, and which have neither. In all cases n∈Z+. (a)an=1 n+ 1 (b)an=(−1)n n (c)bn=en (d)an=e−2n+1 (e)an= 1 +n (f)an= 2 + ( −1)n (g)an=√n+ 1−√n (h)an=2−3n 5n+1 (i)an=7n n!(tough, isn’t it?) (j)sn= 1 +1 2+1 4+1 8+···+1 2n. (2) Prove that if an→αandbn→β, then (an+bn)→α+β, where all the letters represent real numbers. 0.6. APPENDIX: CONTINUOUS FUNCTIONS AND THE MEAN VALUE THEOREM 15 (3) a). Prove Bernoulli’s inequality (1 +h)n>1 +nh, h /negationslash= 0, h>−1, n≥2. Hereh∈Randn∈Z. I suggest proof by induction. b). Ifs∈R, use part a) to prove that an≡sn→/braceleftbigg0 if |s|<1 ∞if|s|>1. [Hint: If |s|<1 , write |s|=1 1+h, h> 0 , while if |s|>1 , write |s|= 1 +h, h> 0] . 0.6 Appendix: Continuous Functions and the Mean Value Theorem Definition: : The function f(x) iscontinuous at the point x0if, given any /epsilon1>0 , there is aδ(/epsilon1)>0 such that |f(x)−f(x0)|<εwhen 0<|x−x0|<δ(/epsilon1). Remark: This may be rephrased as lim x→x0f(x) =f(x0). Note that either statement requires (1)fbe defined at x0. (2) limf(x) exists. x→x0 x/negationslash=x0 (3) the limiting value of fatx0is equal to the defined value of fatx0. If a function is discontinuous at x0, it has at least one of the four troubles (1) Jump discontinuity (2) Infinite discontinuity (3) Infinite oscillations (4) Removable discontinuity. Here are examples of each trouble at the point x= 0 . (1)f(x) =/braceleftbigg1,0≤x −1, x < 0 (2)f(x) =/braceleftbigg1 xx/negationslash= 0 anything, say 1, x = 0 16 CHAPTER 0. REMEMBRANCE OF THINGS PAST. (3)f(x) =/braceleftbiggsin1 xx/negationslash= 0 anything, say 0, x = 0 (4)f(x) =/braceleftbiggx, x /negationslash= 0 1, x= 0 Note that a function may oscillate infinitely about a point and still be continuous there. This is illustrated by the everywhere continuous function f(x) =/braceleftbiggxsin1 x, x/negationslash= 0 0, x= 0 Theorem 0.10 I. Iff(x)is continuous at x=c, andf(c) =A/negationslash= 0, thenf(x)will keep the same sign as f(c)in a suitably small neighborhood of x=c. Proof: : We construct the desired neighborhood. Assume Ais positive. The proof if A<0 is essentially the same. In the definition of continuity, take /epsilon1=A. Then there is a δ>0 such that |f(x)−A|<A when |x−c|<δ, that is, 0<f(x)<2A,when |x−c|<δ. In other words, f(x) is positive in the interval |x−c|<δ. Theorem 0.11 II. Iff(x)is continuous at every point of a closed and bounded inter- val, then there is a constant Msuch that |f(x)| ≤Mthroughout the interval. Thus a continuous function in a closed and bounded interval is bounded. Proof: : By contradiction. If fis not bounded, there is a sequence of points xnsuch that|f(xn)|>n. From that sequence by Theorem 5 (Bolzano-Weierstrass) we can select a subsequence xnkwhich converges to some point x0in t he interval, xnk→x0. Thus |f(xnk)| → ∞. But we know from the continuity of fthat|f(xnk)| → |f(x0)|. A contradiction. Moreover, with the same hypotheses , we can conclude more. Theorem 0.12 III. Iffis continuous at every point of a closed and bounded interval, then there are points x=αandx=βin the interval where fassumes its greatest and least values, respectively. Proof: : We show that fassumes its greatest value. The proof for the least value is essentially identical. Let Sbe the set of all upper bounds for f. By Theorem II Sis not empty. Therefore by Theorem 2, Shas a g.l.b., call it M0. SinceM0is the greatest lower bound of upper bounds for f, there is a sequence xnsuch that lim n→∞f(xn)→M0. Use Bolzano-Weierstrass to pick a subsequence xnkof thexnsuch that the xnkconverges, say toc. By continuity of f,limnk→∞f(xnk) =f(x) . Thusf(c) =M0, sofdoes assume its greatest value at x=c. Remark: This theorem refers to the absolute maximum andabsolute minimum values. 0.6. APPENDIX: CONTINUOUS FUNCTIONS AND THE MEAN VALUE THEOREM 17 Examples: The following show that the theorem is not necessarily true if any of the hypotheses are omitted. (1)f(x) =x,0<x≤1 . No min. (interval not closed). (2)f(x) =x, x≤0 , andf(x) =1 1+x2,allx, both have no min. (the interval is unbounded.) (3)f(x) =/braceleftbiggx, 0≤x<3.No max. (function is discontinuous.) x−2,3≤x≤4 Theorem 0.13 Iff(x)is continuous at every point of a closed and bounded interval [a,b], and iff(a)andf(b)have opposite sign, then there is at least one point c∈(a,b)such thatf(c) = 0 . Proof: : Sayf(a)<0, f(b)>0 . We find one point c, “the largest xsuch that f(x) = 0 ”. Let S={x∈[a,b]:f(x)≤0}. Sincef(a)<0, Sis not empty. It thus has a l.u.b., c. We prove that f(c) = 0 . Eitherf(c)>0, f(c)<0 , orf(c) = 0 . The first two possibilities cannot happen, since by Theorem I, if they did, fwould be positive (or negative) in a whole neighborhood of c-violating the fact that cis the l.u.b. of S. Corollary 0.14 (intermediate value theorem ). Letf(x)be continuous at every point of a closed and bounded interval [a,b], withf(a) =A,andf(b) =B. Then ifCis any number between AandB, there is at least one point c, a≤c<b , such that f(c) =C. Thus,fassumes every value between AandBat least once. Proof: : Apply Theorem IV to the function ϕ(x) =C−f(x) . Remark: The function may assume values other than just those between AandB. An example is the function f(x) =x2,−1≤x≤3 . The theorem requires that it assume all values between f(−1) = 1 and f(3) = 9 . Besides those values , this function also happens to assume all values between 0 and 1. We can offer another proof of Corollary 0.15 Every positive number khas a unique positive square root. Proof: : Consider f(x) =x2−k, which is clearly continuous everywhere. Since f(0)<0 , andf(1+k 2) = (1+k 2)2−k= 1+k2 4>0 , Theorem IV shows that fmust vanish somewhere in the interval 0 < x < 1 +k 2. This is the root. It is the unique positive square root, for say there were two positive numbers xandysuch thatx2−k= 0 andy2−k= 0 . then x2−y2= 0 . Thus, 0 = x2−y2= (x−1)(x+y) . Sincex+y>0 , we conclude x−y= 0 , orx=y. Remark : It appears that if a function has the property of Corollary 1, the intermediate value property, then it must be continuous. This is false . An example is given by the discontinuous (trouble 3) function f(x) =/braceleftbiggsin1 x, x/negationslash= 0 0, x= 0 18 CHAPTER 0. REMEMBRANCE OF THINGS PAST. about the point x= 0 . Ifais any number <0 , andbany number >0 , thenf(x) assumes every value between f(a) andf(b) , butf(x) is not continuous throughout the interval since it is not continuous at x= 0 . Definition: The function f(x) has a relative maximum (minimum) at the point x0, if, for allxin a sufficiently small interval containing x0as an interior point, we have f(x)≤f(x0) (f(x)≥f(x0)). Remark: By convention, we shall agree notto call the possible max (or min) at the end point of an interval a relative max (or min). This does lead to the possibility of an absolute max (or min) not being a relative max (or min). However, if the absolute max (or min) does occur at an interior point of an interval, it is also a relative max (or min). Definition: The function f(x) isdifferentiable at the point x0if the following limit lim x→x0f(x)−f(x0) x−x0 exists. There are the usual notations: f/prime(x0) ,d f dx/vextendsingle/vextendsingle/vextendsingle x=x0,Df(x0) . Theorem 0.16 Iff(x)is differentiable at x0, then it is continuous there. Proof: : Now if the limit lim x→x0f(x)−f(x0) x−x0 exists, as we have assumed, then the numerator must approach zero as xtends tox0. Thusfis continuous at x0. Theorem 0.17 Iff(x)is differentiable at x0and has a relative maximum or minimum atx0, thenf/prime(x0) = 0 . Proof: : Assumefhas a relative min at x0. Then for all xnearx0, f(x)≥f(x0) . (i) ifx<x 0f(x)−f(x0) x−x0≤0 so lim x→x0x<x 0f(x)−f(x0) x−x0≤0 (ii) ifx>x 0f(x)−f(x0) x−x0≥0 so lim x→x0x>x 0f(x)−f(x0) x−x0≥0 Because the function is differentiable at x0, the two limiting values are f/prime(x0) . Thus f/prime(x))≤0 andf/prime(x0)≥0 . Both statements can be true only if f/prime(x0) = 0 . The trick here was, the slope must be negative to the left, and positive to the right of x0. Since there is a unique slope (the derivative) at x0, the slope must be zero there. At a relative max., the same proof holds with obvious modifications. Examples: 1. Although the function f(x) =|x|has a relative minimum at x= 0 , the conclusion of the theorem does not hold since fis not differentiable there. Note that both (i) and (ii) of the proof still do hold. 2. The differentiable function (for all x) f(x) =/braceleftbiggx4sin1 x, x/negationslash= 0 0, x= 0 has an infinite number of relative max and min in any interval including the origin. 0.6. APPENDIX: CONTINUOUS FUNCTIONS AND THE MEAN VALUE THEOREM 19 Theorem 0.18 (Rolle ). If (i)f(x)is continuous at every point of the closed and bounded interval [a,b] (ii)f(x)is differentiable at every point of the open interval (a,b)and (iii)f(a) =f(b), then there is at least one point c,a<c<b , wheref/prime(c) = 0 . Proof: : Iff(x)≡constant throughout [ a,b] , takecto be any point in ( a,b) . Otherwise f(x) must go either above or below (or both) the value f(a) . Assume it goes above. Then by Theorem III there is a point x=cwherefhas its absolute maximum. Since we assumedf(x) goes above f(a) , the point x=cis an interior point. Thus there is a relative maximum. Since fis differentiable in ( a,b) , we may apply Theorem VI to conclude that f/prime(c) = 0 . If we had assumed fwent below f(a) , then there would have been an absolute (and relative) min. etc. Remarks: 1. From the proof of the theorem, we see that if fhas values both greater and less thanf(a) , then there would be at least two points in ( a,b) wheref/prime= 0 . 2. You should be able to construct examples showing the theorem is not true if any of the hypotheses are dropped. Corollary 0.19 (mean value theorem ) If (i)f(x)is continuous at every point of the closed and bounded interval [a,b]and (ii)f(x)is differentiable at every point of the open interval (a,b), then there is at least one point cin(a,b)where f/prime(c) =f(b)−f(a) b−a. Proof: : “Shift and apply Rolle’s Theorem”. In more detail, consider F(x) =f(x)−f(a)−x−a b−a(f(b)−f(a)). F(x) satisfies all of the assumption of Rolle’s Theorem. Therefore there is a point cwhere F/prime(c) = 0 . Since F/prime(x) =f/prime(x)−f(b)−f(a) b−a, atx=c, we have f/prime(c) =f(b)−f(a) b−a. Remarks: 1. The function f(x) =|x|in the interval [ a,b], a < 0, b > 0 , shows what happens if the function fails to be differentiable at even one point of the open interval ( a,b) . 2. An alternative form of the conclusion is: there is a number θ,0<θ< 1 , such that f(b)−f(a) =f/prime(a+θ(b−a))(b−a). This is because every point in the interval ( a,b) is of the form a+θ(b−a) , for some θ,0<θ< 1 . We shall now give some applications of the Mean Value Theorem. The first one is a specific example, while the others have great significance in themselves. 20 CHAPTER 0. REMEMBRANCE OF THINGS PAST. Example: The function f(x) =a1sinx+a2sin 2x+bcosx+b2cos 2xhas at least one zero in the interval [0 ,2π] , no matter what the coefficients a1,a2,b1andb2are. To show this, we shall show fis the derivative of a function g(x) which satisfies the hypotheses of Rolle’s theorem. This function gis just an anti-derivative of f:g/prime(x) =f(x) g(x) =−acosx−a2 2cos 2x+b1sinx+b2 2sin 2x. Sincegis clearly continuous and differentiable everywhere, we must only see if g(0) = g(2π) , which as also easy. Theorem 0.20 Iff(x)is continuous and differentiable throughout [a,b], and |f/prime|< N there too, then the δ(/epsilon1)in the definition of continuity can be chosen as δ(c) =/epsilon1 N. Thisδ works for everyxin[a,b]. Proof: : Use the form of the mean value theorem in Remark 2. Then for any points x, x 0 in (a,b) , f(x)−f(x0) =f/prime(˜x)(x−x0), where ˜xis somewhere between xandx0. Thus |f(x)−f(x0)| ≤N|x−x0|. We see now that if δ(/epsilon1) =/epsilon1 N, then for any /epsilon1>0 , |f(x)−f(x0)|</epsilon1if|x−x0|<δ. Theorem 0.21 Iffsatisfies the hypotheses of the mean value theorem and if in addition f/prime(x)≡0throughout (a,b), thenf(x)≡const. Proof: : Letx1andx2be any points on ( a,b) . Then by the form of the mean value theorem in Remark 2 f(x2)−f(x1) = 0·(x2−x1) = 0. Thusf(x2) =f(x1) for any two points in ( a,b) , that is,fis identically constant. Corollary 0.22 Iff(x)andg(x)both satisfy the hypotheses of the mean value theorem, and if in addition f/prime(x)≡g/prime(x)for allxin(a,b), thenf(x) =g(x)+c, wherecis some constant. Proof: : consider the function F(x) =f(x)−g(x) . It satisfies the hypothesis of Theorem VII, soF(x)≡c, cconstant. Thus f(x)−g(x) =c. Remark: Theorem IX is the converse of the theorem: “the derivative of a constant function is zero.” a figure goes here 0.6. APPENDIX: CONTINUOUS FUNCTIONS AND THE MEAN VALUE THEOREM 21 Exercises (1) Look over all the theorems (and corollaries) here and be sure you can find examples showing that the theorems are not true if any of the hypotheses are relaxed. (2) Letf(x) =/braceleftbigg1,ifxis a rational number 0,ifxis an irrational number. Isfcontinuous anywhere? (3) Letf(x) be an everywhere differentiable function which is zero at x=aj, j= 1,2,...,n. Find a function which vanishes at least once between each of the zeros of f. (4) Use Theorem VIII to find a δ(/epsilon1) for the given functions. (a)f(x) =x4−7,−2≤x≤3. (b)f(x) =x2sinx,−4≤x≤3 (c)f(x) =1 1+x2,−2≤x≤1 (d)f(x) =x4 3+ 7,−2≤x≤8 (e)f(x) =x√ x2+ 1,−2≤x≤2 (5) (a) The function f(x) satisfies the following condition |f(x)−f(x0)| ≤2|x−x0|3 for every pair of points x, x 0in the interval [ a,b] . Provef(x)≡constant in this interval. (b) Generalize your proof to the case when fsatisfies |f(x)−f(x0)| ≤c|x−x0|α, wherec>0 is some constant and αis any number >1 . (6) Consider the function f(x) =x2 3, in the interval [ −8,8] . Sketch a graph. Note that f(−8) =f(8) = 4 but there is no point where f/prime= 0 ; which hypothesis of Rolle’s theorem is violated? (7) In a trip, the average speed of a car is 180 miles per hour. Prove that at some time during the trip, the speedometer must have registered precisely 180 miles per hour. (8) LetP1:= (x1,y1) andP2:= (x2,y2) be any two points on the parabola y= ax2+bx+c, and letP3:= (x3,y3) be the point on the arc P1P2where the tangent is parallel to the chord P1P2. Show that x3=x1+x2 2. (9) Prove that every polynomial of odddegree P(x) =x2n+1+a2nx2n+···+a1x+a0 has at least one real root. (10) Iffis a nice function and f/prime<0 everywhere, prove that fis strictly decreasing. 22 CHAPTER 0. REMEMBRANCE OF THINGS PAST. 0.7 Complex Numbers: Algebraic Properties . In high school, to be able to find the roots of all quadratic equations ax2+2bx+c= 0 , we were forced to introduce the symbol i≡√−1 , in other words, introduce a special symbol for a root of x2+ 1 = 0 . Before going any further, we should prove that no real number c can satisfy c2+ 1 = 0 . By contradiction, assume that there is such a c. Then necessarily eitherc>0, c< 0,orc= 0 . Ifc= 0 , we have the immediate contradiction that 1 = 0 . If c>0 , orc<0,0<c2. Consequently 0 <c2+ 1 too, which again contradicts 0 = c2+ 1 , and proves our contention that no real number can satisfy x2+ 1 = 0 . Observe that our proof also shows that if we introduce a new symbol for a root of x2+ 1 = 0 , that symbol cannot be an element of an ordered field, for only the ordered field properties of the real numbers were used in the above proof. we shall see that “ i” is an element of a field, but not an ordered field. It is difficult to overestimate the importance of complex numbers for all of mathematics, both from an esthetic as well as from a practical viewpoint. With them we can prove that every quadratic polynomial has exactly two roots (which may coincide). What is more surprising is that every polynomial of order n anxn+an−1xn−1+···+a1x+a0= 0, an/negationslash= 0, has exactly ncomplex roots. This result, the fundamental theorem of algebra , was first proved by Gauss in his doctoral dissertation (1799). It is one of the crown jewels of math- ematics. The difficult part is proving that every polynomial has at least one complex root, from which the general result follows using only the “factor theorem” of high school alge- bra. Later on in the semester we shall discuss this more fully and offer a proof. It is not simpleminded, for the proof is non-constructive pure existence proof, giving absolutely no method of finding the roots. Perhaps we shall even prove some more exotic results. Having gotten carried away, let us retreat and obtain the algebraic rules governing the setCof complex numbers. In order to reveal the algebraic structure most clearly, we shall denote a complex number zby an ordered pair of real numbers: z= (x,y), x, y ∈R. Thus CisR×Rwith the following additional algebraic structure. Definition: Ifz1= (x1,y1) andz2= (x2,y2) are any two complex numbers, then we define Addition: z1+z2= (x1+x2, y1+y2) , and Multiplication: z1·z2= (x1x2−y1y2, x1y2+y1x2) . Equality:z1=z2if and only if both x1=x2andy1=y2. Thus, the complex number zero—the additive identity—is (0 ,0) , while the complex number one—the multiplicative identity—is (1 ,0) . Using the fact that the real numbers Rform a field, we can now prove the Theorem 0.23 The complex numbers Cform a field. Proof: Since the verification of the field axioms are entirely straightforward we give only a smattering. Note that we shall rely heavily on the field properties of R. Addition is commutative: z1+z2= (x1,y1) + (x2,y2) = (x1+x2, y1+y2) = (x2+x1, y2+y1) = (x2,y2) + (x1,y1) =z2+z1.(0-3) 0.7. COMPLEX NUMBERS: ALGEBRAIC PROPERTIES 23 Additive identity: 0 +z= (0,0) + (x,y) = (0 +x,0 +y) = (x,y) =z. Multiplicative inverse: For any z∈C, z/negationslash= (0,0) , we must find a ˆ z= (ˆx,ˆy)∈Csuch thatzˆz= 1 , that is, find real numbers ˆ xand ˆysuch that ( x,y)(ˆx,ˆy) = (1,0) . Using the definition o f complex multiplication, this means we must solve the two linear algebraic equations xˆx−yˆy= 1 yˆx+xˆy= 0/bracerightbigg x, y∈R, for ˆxand ˆy∈R. The result is ˆz= (ˆx,ˆy) = (x x2+y2,−y x2+y2). We will denote this multiplicative inverse, which we have just proved does exist, by1 zor z−1. It is interesting to notice that complex numbers of the form ( x,0) have the same arithmetic definitions as the real numbers, viz. (x1,0) + (x2,0) = (x1+x2,0) (x1,0)(x2,0) = (x1x2,0). We can easily verify that all complex numbers of this form ( x,0) also form a field, asubfield of the field C. On the basis of these last two equations, we can identify a real numberxwith the complex number ( x,0) in the sense th at if we perform any computation with these complex numbers of this form, the result will be the same as if the computation had been performed with the real numbers alone. Thus, numbers of the form ( x,0)∈Care algebraically equivalent to the numbers x∈R. The technical term for such an algebraic equivalence is isomorphic , much as a term for geometric equivalence is congruent. After identifying the real numbers with complex numbers of the form ( x,0) , we can say that the field of real numbers Risembedded as a subfield in the field of complex numbers, R⊂C. After all this chatter, let us at least convince ourselves that every quadratic equation is solvable if we use complex numbers. First we solve z2+ 1 = 0 , which may be written as (x,y)(x,y) + (1,0) = (0,0) , or as the two real equations x2−y2=−1,2xy= 0 . The last equation says that either x= 0 ory= 0 . Now if y= 0 , we are left to solve x2+ 1 = 0, x∈R, which we know is impossible. Therefore x= 0 and then y2= 1 . Thus the two complex numbers (0 ,1) and (0,−1) both satisfy z2+ 1 = 0 . The general case, az2+bz+c= 0 is easily reduced to the special one by completing the square. One by-product of the above demonstration is that we see it is foolhardy to try to define an order relation on Cto obtain an ordered field. This is because the equation x2+ 1 = 0 cannot be solved in any ordered field, as was shown earlier, whereas we have just solved it inC. Observe that every ( x,y)∈Ccan be written as (x,y) = (x,0)(1,0) + (y,0)(0,1), where the complex number (0 ,1) is called the imaginary unit and is denoted by i. If we now utilize the isomorphism between the real number aand complex numbers ( a,0) , the 24 CHAPTER 0. REMEMBRANCE OF THINGS PAST. last equation shows that ( x,y) may be thought of as x+iy. Thus, we have obtained the usual notation for complex numbers. From our development, the algebraic role of ias the symbol for the imaginary unit (0 ,1) is hopefully clarified. The number xis called the real part, andytheimaginary part of the complex number z=x+iy. In symbols, x=Re{z} andy=Im{z}. Our introduction of complex numbers suggests a geometric interpretation. We have defined complex numbers Cas ordered pairs of real numbers, elements of R×R, with an additional algebraic structure. Since the points in the plane are also elements of R×R, it is clear that there is a one to one correspondence between the complex numbers and the points in the plane. If we plot the point z= (x,y) , the real number |z|, the “ absolute value ormodulus ofz” is the distance of the point zfrom the origin. Its value is computed by the Pythagorean theorem |z|=/radicalbig x2+y2. Here are several formulas which are easily verified: |z1z2|=|z1|+|z2| |x| ≤ |z|,|y| ≤ |z| |z1+z2| ≤ |z1|+|z2|(triangle inequality)  (0-4) If the line joining the point zto the origin is drawn, the angle θbetween that line and the positive real (= x) axis is called the argument oramplitude orz. The absolute valuerand argument θof a complex number determine it uniquely, since we have z=r(cosθ+isinθ) (0-5) This is the polar coordinate form of the complex number z. Note that conversely, z determines its argument only to within an additive multiple of 2 π. This observation will prove of value to us shortly. Associated with every complex number, z=x+iythere is another complex number z=x−iy, the complex conjugate ofz. It is the reflection of zin the real axis. Probably the main reason for introducing zis that we can solve for xandyin terms of zandz: x=z+z 2, y =z−z 2i. Again some simple formulas: |¯z|=|z|,|z|2=|¯z|2=zz. (z1+z2) =z1+z2,(z1z2) =z1z2./bracerightbigg (0-6) To illustrate the value of this notation, let us leave the main road to prove the interesting Theorem 0.24 . If the complex number γis a root of the polynomial P(t) =antn+an−1tn−1+···+a1t+a0, where the coefficients a0,a1,...,a narerealnumbers, then γis also a root of P(t). In other words, the roots of realequations occur in conjugate pairs. 0.7. COMPLEX NUMBERS: ALGEBRAIC PROPERTIES 25 Proof: Sinceγis a root, the complex number P(γ) =anγn+···+a1γ+a0 is zero,P(γ) = 0 . This implies that its conjugate is also 0, P(γ) = 0 . By using equations (0.7), we have that P(γ) =anγn+···+a1γ+a0, since the coefficients ajare real,aj=aj. Thus 0 =P(γ) =anγn+···a1γ+a0=P(γ), that is, the complex number γis a root of the same polynomial. Now if the proof looks like it was done with mirrors, go over each step carefully. This type of reasoning is somewhat typical of modern mathematics in that it yields information about an object (the roots of a polynomial in this case) without first obtaining an explicit formula for the object. After this digression let us return and find a geometric interpretation for the arithmetic operations on complex numbers. First, addition. The three points z1,z2andz1+z2 together with the origin determine a parallelogram. (check this). Thus addition of complex numbers is sometimes called the parallelogram rule for additions. Given the points z1and z2, the point z1+z2can be constructed using compass and straight-edge. Subtraction is justz1+ (−z2) . Multiplication is much more difficult to interpret geometrically. We shall use equation (0.7) and write zj=|zj|(cosθj+isinθj),j= 1,2.Then z1z2=|z1|(cosθ1+isinθ1)|z2|(cosθ2+isinθ2) z1z2=|z1z2|[cosθ1+θ2) +isin(θ1+θ2)]. (0-7) Thus the product of z1andz2has modulus |z1z2|and argument θ1+θ2: multiply the moduli and add the arguments. This too may be carried out using compass and straight- edge. Since1 z2=1 |z2|(cosθ2−isinθ2) , division reads z1 z2=/vextendsingle/vextendsingle/vextendsingle/vextendsinglez1 z2/vextendsingle/vextendsingle/vextendsingle/vextendsingle[cos(θ1−θ2) +isin(θ1−θ2)], so the moduli are divided while the arguments are subtracted. We will exploit the multiplication formula (0.7) to find all ncomplex roots of the specific polynomial zn=A, for anyA∈C. This equation is one of the few whose roots can always be found explicitly. The trick is to write Ain its polar coordinate form A=|A|[cos(α+ 2kπ) +isin(α+ 2kπ)], whereαis the argument of Aandkis any integer. Although we get the same Ano matter whatkis used, as was observed following equation (0.7), we shall retain the arbitrary k since it is the heart of the process we have in mind. From equation (0.7) we see that A1 n=|A|1 n[cosα+ 2kπ n+isinα+ 2kπ n] 26 CHAPTER 0. REMEMBRANCE OF THINGS PAST. in the sense that for any value of the integer k,(A1 n)n=A. Askruns through the integers, we get only ndifferent angles of the formα+2kπ n, since the other angles differ from these nangles by multiples of 2 π. For each of these ndifferent angles we obtain a different complex number A1 n. Thesennumbers for A1 nare the desired nroots ofzn=A. It is usually convenient to obtain the angles by letting k= 0,1,2,...,n −1 , although any n integers which do not differ by multiples of nwill do. An example should help clear the air. We shall find the three cube roots of −2 , that is, solvez3=−2 . First, −2 = 2[cos(π+ 2kπ) +isin(π+ 2kπ)], since the argument of −2 isπwhile its modulus is 2. Thus, the roots are z= 21 3[cosπ+ 2kπ 3+isinπ+ 2kπ 3], k= 0,t1,t2... There are only three values of zpossible, no matter what k’s are used. These three cube roots of −2 are k= 0,3,6,...z 1= 21 3[cos(π 3) +isin(π 3)] = 21 3(1 2+i√ 3 2)k= 1,4,7,...z 2= 21 3[cos(π) + isin(π)] =−21 3 k= 2,5,8,...z 3= 21 3[cos(5π 3) +isin(5π 3)] = 21 3(1 2+i√ 3 2). It is time-saving to observe that the nroots of unity, that is, of zn= 1 , can be written down immediately by utilizing the geometric interpretation of multiplication. All of the roots have modulus 1, and so must lie on the unit circle |z|= 1 . Bisecting the circle into nequal sectors by the radii, the first beginning on the positive x-axis, we find the roots of unity,wj, at thensuccessive intersections of these radii with the unit circle. The roots wj, j= 1,2,3,ofz3= 1 are illustrated in the figure as the intersections of θ= 0, θ=2π 3, andθ=4π 3with |z|= 1 . Thus w1= cos 0 +isin 0 = 1, w2= cos2π 3+isin2π 3= −1 2+i√ 3 2,w3= cos4π 3+isin4π 3=−1 2−i√ 3 2. Exercises (1) Express the following complex numbers in the form a+bi. (a) (1 −i)2 (b) (2 +i)(3−i) (c)1 i (d)1+i 2−i (e)1+i 1+2i (f)i3+i4+i271 (2) Compute the absolute values of the complex numbers in Ex. 1. 0.7. COMPLEX NUMBERS: ALGEBRAIC PROPERTIES 27 (3) a) Add (1 + i) and (1 + 2 i) using compass and straight-edge. b) Multiply (1 + i) and (1 + 2 i) using compass and straight-edge. (4) Express in the form r(cosθ+isinθ),with 0 ≤θ<2π: (a)i (b) 2i (c)−2i (d) 4 (e)−1 (f)−1 +i (g) (1 −i)3 (h)1 (1+i)2 (i)1 2(√ 3 +i) (5) Determine the (a) three cube roots of i,−i, and of 1 + i, (b) four fourth roots of −1 and +2 (c) six roots of z6= 1 . (6) LetAbe any complex number, A=|A|[cosα+isinα] , and letw1,...,w nbe the nroots ofzn= 1 . Prove that the nroots ofzn=Aare z1=A1 nw1, z2=A1 nw2,...z n=A1 nwn, where A1 n=|A|1 n(cosα n+isinα n) is the principal nth root ofA. This shows that the problem of finding the roots of a complex number is essentially reduced to the simpler problem of finding the roots of unity. (7) Draw a sketch of the following sets of points in the complex plane. (a){z∈C:|z−2| ≤1} (b){z∈C:|z−1 +i| ≤2} (c){z∈C:|z−2|>3} (d){z∈C: 1≤ |z−2| ≤3} (e){z∈C: 1≤ |z+i|<2} 28 CHAPTER 0. REMEMBRANCE OF THINGS PAST. 0.8 Complex numbers: Completeness Properties, Complex Functions. We have just considered the algebraic properties of complex numbers. Now we look at infinite sequences of complex numbers. To develop the desired properties of C, we shall utilize those of R. Definition: The sequence znof complex numbers converges to the complex number zif, given any/epsilon1>0 , there is an Nsuch that |zn−z|</epsilon1for alln>N . We shall again write zn→z. In order to apply the theorem known for real sequences to complex sequences, the following is vital. Theorem 0.25 Letzn=xn+iyn, andz=x+iy. Thenznconverges to zif and only if both the real and imaginary parts converge to their respective limits. In symbols, zn→z⇐⇒xn→xandyn→y. Proof: Sincezn→z, given any /epsilon1 >0 , we can find an Netc. for the zn’s. Now by equation (0.7) |xn−x| ≤ |zn−z|</epsilon1and|yn−y| ≤ |zn−z|</epsilon1 so bothxn→xandyn→y. Conversely, given any /epsilon1>0 , we can find an N1for thexn’s and anN2for theyn’s. LetNbe the larger of N1andN2,N= max (N1,N2) . ThisNworks for both the xn andyn. But |zn−z|=|xn+iyn−x−iy| ≤ |xn−x|+|yn−y|<2/epsilon1. Thereforezn→z, completing the proof. This theorem states that a definition is equivalent to some other property. We could thus have used either property as a definition. Recall that the real numbers were defined so that there would be no“hole” in the real line. This was the completeness property. It guaranteed that if a sequence of real numbers an“looked like” they were approaching a limiting value, then indeed th ere is some a∈R such thatan→a. The issue here was to avoid the problem of a sequence of rational numbers approaching an irrational number—which is a “hole” if our set just consisted of the rationals. One consequence of the las t theorem is that the set of complex numbers C is also complete. Theorem 0.26 . Every bounded infinite sequence of complex numbers {zk}has at least one subsequence which converges to a number z∈C. (By bounded, we mean that there is somer∈Rsuch that |zk|<r for allk). Proof: Since the {zk}are bounded, we know {xk}and{yk}are also bounded se- quences of real numbers. The conclusion is now a consequence of the Bolzano-Weierstrass theorem 5 applied to {xk}and{yk}, and of theorem 12 just proved. There is a fine point though: how to get a subsequence of the zkwhose real and imaginary parts both converge. The trick is first to select a subsequence {xkj}={Rez kj}of the {xk}which converge to some x∈R. Then, from the related subsequence {ykj}={Imz kj}, select 0.8. COMPLEX NUMBERS: COMPLETENESS AND FUNCTIONS 29 a subsequence {ykjn}which converges to some y∈R. Then {xkjn}also converges to x∈Rsozkjn→z, and we a re done. With these technical results under our belts, sequences in Cbecome no more difficult than those in R. Let us briefly examine the elements of functions of a complex variable. A complex- valued function f(z) of the complex variable zis a mapping of some subset z U ⊂C into the complex numbers C, f:U→C. Two examples are f(z) =z2, andf(z) =1 z. Both the domain and range of f(z) =z2are all of C, while the domain and range of f(z) =1 zare all of Cwith the exception of 0. Iffmaps R→R, likef(x) = 1 +xorf(x) =ex, since R⊂C, one asks how the domain of definition of fcan be extended from RtoC. Of course there are many possible ways to do this, but most of them are entirely artificial. For f(x) = 1 +x, the natural extension is f(z) = 1+z, z∈C. Similarly, if P(x) =/summationtextN k=0akxkis any polynomial defined for x∈R, the natural extension to z∈CisP(z) =/summationtextN k=0akzk. We are thus led to extendf(x) =exforx∈R, toz∈Cby defining f(z) =e2. The only problem is that we have absolutely no idea what it means to raise a real number, e, to a complex power. Taylor (power) series are needed to resolve this issue. This will be carried out at the end of Chapter 1. Continuity of complex functions is defined in a natural way. Let z0be an interior point of the setU⊂C(that is,z0is not on the boundary of U). Definition: The function f:U→Cis continuous at the interior point a0/epsilon1Uif, given any/epsilon1>0 there is a δ>0 such that |f(z)−f(z0)|</epsilon1for allzin 0<|z−z0|<δ. Reasonable theorems like, if fandgare continuous at the interior point z0/epsilon1U, so is the function f+g, are true too - with the same proof as was given for real-valued functions of a real variable. Although we could go on and define the derivative and integral for complex-valued functionsf(z) of a complex variable, the development would take too much work. For our future purposes, it will be sufficient to define the derivative and integral of a complex- valued function f(x) of the realvariablex. The first step is to split f(x) into its real and imaginary parts, that is, find real valued functions u(x) andv(x) such that f(x) = u(x) +iv(x) . This decomposition c an always be done by taking u(x) =f(x) +f(x) 2, v(x) =f(x) +f(x) 2i. Sinceu(x) =u(x) andv(x) =v(x) , bothu(x) andv(x) are real-valued functions. It is clear thatf(x) =u(x) +iv(x) . Example: For the functions f(x) = 1 + 2ix, we havef(x) = 1−2ix, so u(x) =(1 + 2ix) + (1 −2ix) 2= 1, v(x) =(1 + 2ix)−(1−2ix) 2i= 2x. as expected. Becausef(x) is a complex number for every xin the domain where fis defined, we |f(x)|=/radicalbig u2(x) +v2(x). With this notion of absolute value, the definitions of continuity and differentiability read just as iffwere itself real-valued. For example 30 CHAPTER 0. REMEMBRANCE OF THINGS PAST. Definition: : The complex-valued function f(x) of the real variable xisdifferentiable at the point x0if lim x→x0f(x)−f(x0) x−x0 exists. A more convenient way of dealing with the derivative is supplied by the following Theorem 0.27 . The function f(x) =u(x) +iv(x)is differentiable at a point x0if and only if both u(x)andv(x)are differentiable there, and df dx=du dx+idv dx. Proof: We shall use Theorem 12. Let {xn}be any sequence whose limit is x0. Define the sequences {an},{αn},and{βn}by an=f(xn)−f(x0) xn−x0, αn=u(xn)−u(x0) xn−x0,andβn=v(xn)−v(x0) xn−x0. We must show that lim n→∞anexists if and only if both limits lim n→∞αnand lim n→∞βn exist, for the existence of these limits is equivalent to the existence of the respective deriva- tives. But notice that an=αn+iβn, since an=f(xn)−f(x0) xn−x0=u(xn) +iv(xn)−(u(x0) +iv(x0)) xn−x0=αn+iβn. Thus we can appeal to Theorem 12 to conclude that lim anexists if and only if both lim αn and limβnexist. The formula f/prime=u/prime+iv/primeis an immediate consequence since an→f/prime(x0), αn→u/prime(x0),andβn→v/prime(x0) Examples: a) Iff(x) = 1 + 2ix,d f dx=d dx1 +id dx2x= 2i b) Iff(θ) = cos 7θ+isin 7θ+ 2θ−iθ2 d f dθ=d dθ[2θ+ cos 7θ] +id dθ[−θ2+ sin 7θ] = 2−7 sin 7θ+i[−2θ+ 7 cos 7θ] A related result which is even easier to prove is Theorem 0.28 . The complex-valued function f(x) =u(x) +iv(x), x/epsilon1Ris continuous at x0/epsilon1Rif and only if both u(x)andv(x)are continuous at x0. Proof: An exercise. Integration is defined more directly. Definition: Letf(x) =u(x) +iv(x), x/epsilon1R. If the real-valued functions u(x) , andv(x) are integrable for x/epsilon1[a,b] , we define the definite integral off(x) by /integraldisplayb af(x)dx=/integraldisplayb au(x)dx+i/integraldisplayb av(x)dx. 0.8. COMPLEX NUMBERS: COMPLETENESS AND FUNCTIONS 31 The standard theorems, like if cis any complex constant, then /integraldisplayb acf(x)dx=c/integraldisplayb af(x)dx,and, ifa≤b,/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayb af(x)dx/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤/integraldisplayb a|f(x)|dx are proved by using the definition above and the corresponding theorems for real functions. We shall, however, need the more difficult Theorem 0.29 . If the complex-valued function f(t) =u(t)+iv(t), t/epsilon1R, is continuous for allt/epsilon1[a,b], then there is a constant Ksuch that |f(t)| ≤Kfor allt/epsilon1[a,b]. Furthermore ifx, x 0/epsilon1[a,b], the n/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx x0f(t)dt/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤K|x−x0|. (0-8) Notice that the left-hand side absolute value is in the sense of complex numbers. Proof: Sincef(t) is continuous in [ a,b] , by Theorem 15 so are both u(t) andv(t) . But a real-valued function which is continuous in a closed and bounded interval is bounded. Thus there are constants K1andK2such that |u(t)| ≤K1,|v(t)| ≤K2for allt/epsilon1[a,b] . then |f(t)|=/radicalbig u2(t) +v2(t)≤/radicalBig K2 1+K2 2≡K. To prove the inequality (0.29), we use the inequality mentioned before the theorem to see that ifx0≤x /vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx x0f(t)dt/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤/integraldisplayx x0|f(t)|dt. Since |f(t)| ≤K, we find that /integraldisplayx x0|f(t)|dt≤K|x−x0|. Combining these last two inequalities, we obtain the desired inequality (0.29) if x0≤x. The other case, x≤x0, can be reduced to that already proved by observing that /vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx x0f(t)dt/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle−/integraldisplayx x0f(t)dt/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx x0f(t)dt/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤K|x0−x|=K|x−x0|. Exercises (1) In the complex sequences below, which ones converge, which do not converge but have at least one convergent subsequence, and which do neither? In all cases n= 1,2,3,.... (a)zn=i n+ 3i−4 (b)zn= 2i+ (−1)n (c)zn=n−i (d)zn=in (e)zn= 1 +i√ 3−(−1)n 7n 32 CHAPTER 0. REMEMBRANCE OF THINGS PAST. (f)zn=(4+6 i)n−5 1−2ni. (2) Write the following complex-valued functions f(x) of the real variable xasf(x) = u(x) +iv(x) , whereuandvand real-valued. (a)f(x) =i+ 2(3 −2i)x2, (b)f(x) = (1 + 2ix)2 (c)f(x) = cos 3x2−(3 +i) sinx (d)f(x) =1 1+2i−x (3) (a) Use the definition of the derivative to computed f dxfor the function in Exercise 2a above. (b) Findd f dxfor all the functions in Exercise 2 above. (4) Evaluate (a)/integraltext3 −1(1 + 2ix)dx (b)/integraltext4 1[x+ (1−i) cos 2x]dx Chapter 1 Infinite Series 1.1 Introduction In elementary calculus you have met the notion of the limit of a sequence of numbers (see also Chapter 0, sections 5 and 7). This concept of limit is just what essentially distinguishes calculus from algebra. It was crucial in the definition of the derivative as the limit of a difference quotient and the integral as the limit of a Riemann sum. We now propose to discuss another limiting process, infinite series, in detail. An infinite series is a sum of the form ∞/summationdisplay k=1ak=a1+a2+a3+···, (1-1) where the ak’s are real or complex numbers. Since there is no added difficulty we shall suppose the ak’s are complex numbers. One immediate trouble is that it would take us an infinite amount of time to add an infinite sum. For example, what is (a)/summationtext∞ k=11 = 1 + 1 + 1 + 1 + ··· = ? (b)/summationtext∞ k=11 = 1−1 + 1−1 + 1−1··· = ? (c)/summationtext∞ k=11 2k−1= 1 +1 2+1 4+1 8+1 16+···=? Thus, we are faced with the realization that be sum (1) is not really well defined, even in cases where we feel it might make sense. Our first task is to give a more adequate definition. Let Snbe the sum of the first n terms: Sn:=a1+a2+···+an=n/summationdisplay k=1ak. Then for each n, we have a complex number Sn, called the nthpartial sum of the series (1). Definition: If lim n→∞Sn=S, whereSis a (finite) complex number, we say that the infinite series converges toS. If the sequence S1,S2,S3,... has no limit, we say that the infinite series diverges . For the examples given just above, we have (a)Sn=/summationtextn 11 =n→ ∞ so the infinite series diverges to ∞. (b)Sn=/summationtextn 1(−1)n+1=/braceleftbigg1nodd, 0neven,/bracerightbigg .which does not have a limiting value since it oscillates between 1 and 0. 33 34 CHAPTER 1. INFINITE SERIES (c)Sn=/summationtextn 11 2k−1= 2(1 −1 2n)→2 , so the infinite series converges to the number 2 (we found the sum of the series by realizing it is a simple geometric series: 1 +r+r2+···+rN=1−rN+1 1−r) for (r/negationslash= 1). With an adequate definition of convergence of infinite series, it is clear that we should develop some tests for determining if a given series converges. That will be done in the next section. In preparation, let us examine some simple types of series which occur often and prove a few useful theorems. There are two types of series whose sums can always be found, and for which the question of convergence is exceedingly elementary. Definition: An infinite geometric series is a series of the form ∞/summationdisplay k=0ark=a+ar+ar2+···. The partial sums are Sn=a+ar+···+arn=a1−rn+1 1−rfor (r/negationslash= 1). Theorem 1.1 The infinite geometric series/summationtext∞ k=0ark, a/negationslash= 0, converges if and only if |r|<1. Then the sum isa 1−r. Proof: limn→∞rn+1exists only if |r|<1 . Then the limit is zero so lim n→∞Sn=a 1−r (the non-convergence when |r|= 1 follow from Theorem 6, p. ?) Examples: (a)/summationtext∞ k=0(1 +i)kdiverges since |1 +i|=√ 2≥1 . (b)/summationtext∞ k=0(1+i 2)kconverges since/vextendsingle/vextendsingle1+i 2/vextendsingle/vextendsingle=√ 2 2<1.The sum of this series is 1 + i. (c)/summationtext∞ k=11 diverges since |1|= 1 . (d)/summationtext∞ k=1(−1)kdiverges since |−1|= 1 . Definition: An infinite telescopic series is one of the form ∞/summationdisplay k=1(αk−αk+1) = (α1−α2) + (α2−α3) + (α3−α4) +···. It is clear that most of the terms cancel each other. Theorem 1.2 Ifαk→α, then/summationtext∞ k=1(αk−αk+1) =α1−α. Proof:Sn= (α1−α2) + (α2−α3) +···+ (αn−αn+1) =α1−αn+1→α1−α. Examples: (a)1 1·2+1 2·3+1 3·4+···=/summationtext∞ k=11 k(k+1)=/summationtext∞ k=1(1 k−1 k+1) = 1 . (b)1 4·12−1+1 4·22−1+1 4·32−1+···=1 2/summationtext∞ k=1(1 2k−1−1 2k+1) =1 2 1.1. INTRODUCTION 35 We close this section with some reasonable (and desirable) theorems. The proofs are immediate consequences of the definition of convergence of infinite series and the related theorems about limits of sequences of numbers. Theorem 1.3 . If/summationtext∞ k=1ak→a, andcis any number then/summationtext∞ k=1cak→ca. Theorem 1.4 If/summationtextn k=1ak→aand/summationtextn k=1bk→b, then/summationtextn k=1(ak+bk)→a+b. Theorem 1.5 Letak=αk+iβk, whereαkandβkare real. The infinite series/summationtextak converges if and only if the two real series/summationtextαkand/summationtextβkboth converge. That is, an infinite complex series converges if and only if both its real and imaginary parts converge. Proof: We must look at the partial sums. Let σn=/summationtextn k=1αk, andτn=/summationtextn k=1βk. Then Sn=n/summationdisplay k=1αk=n/summationdisplay k=1(αk+iβk) =n/summationdisplay k=1αk+in/summationdisplay k=1βn=σn+iτn. But we know from Theorem 12 of Chapter 0 that the complex sequence Snconverges if and only if both its real part, σn, and imaginary part, τn, both converge—in other words, if the series/summationtextakand/summationtextβkboth converge. Two remarks should be made in an attempt to mitigate some confusion. First, the index kof the series/summationtext∞ k=1akcould have been any other letter. Thus/summationtext∞ k=1ak=/summationtext∞ j=1aj. This is perhaps indicated most clearly if we left an empty box instead of using any letter at all: . The connecting line means that the same letter must be used in both boxes. Now you can fill in any letter that makes you happy. No matter w hat you write, it still means a1+a2+a3+···. In a similar way, the index need not begin with 1. Thus, for example,/summationtext∞ k=1ak=/summationtext∞ k=17ak−16=a1+a2+···. Although this manipulation looks like unwanted silliness here, it is sometimes quite useful. Later on this year you will need it. The related transformation for integrals is illustrated by /integraldisplay3 21 tdt=/integraldisplay2 11 t+ 1dt. Exercises (1) Find a closed form expression for the nthpartial sum of the following infinite series and determine if they converge. (a)2 3+2 9+2 27+···+2 3n+···=/summationtext∞ k=12 3k. (b) 1 +i+i2+i3+···+in+··· (c)1 2!+2 3!+3 4!+···=/summationtext∞ k=2k−1 k!=/summationtext∞ k=2(1 (k−1)!−1 k) (d) ln1 2+ ln2 3+ ln3 4+···+ ln(n n+1) +··· (e)/summationtext∞ m=0(3−4i 7)m (f)/summationtext∞ n=12−3i n(n+1) 36 CHAPTER 1. INFINITE SERIES (2) The repeating decimal 1 .565656 ···can be written as 1 +56 102+56 104+56 106+···= 1 + 56∞/summationdisplay k=1(1 102)k. Sum the geometric series and find what rational number the repeating decimal repre- sents. In a similar way, every decimal which begins to repeat eventually is a rational number. What rational number is represented by 1.4723? (3) A ball is dropped from a height of 20 feet. Every time it bounces, it rebounds to3 4 of its height on the previous bounce. What is the total distance traveled by the ball? (4) If/summationtext∞ k=1ak→aand/summationtext∞ k=1bk→b, and ifαandβare any numbers, prove that/summationtext∞ k=1(αak+βbk)→αa+βb. (5) Ifan>0 and/summationtextanconverges, prove that/summationtext1 andiverges. (6) Does the convergence of/summationtext∞ n=1animply the convergence of/summationtext∞ n=1(an+an+1) ? (7) (a) If the partial sums of/summationtextanare bounded, and {bn}is a strictly decreasing sequence with limit 0, bn/arrowsoutheast0 , prove that/summationtextanbnconverges. (b) Use (a) to prove that if/summationtext∞ n=1nanconverges then so does the series/summationtext∞ n=1an. (c) Use (a) to discuss the convergence of/summationtext∞ n=1sinnx n. 1.2 Tests for Convergence of Positive Series Tests to determine convergence are of several types, i) those that give sufficient conditions, ii) those that give necessary conditions, and iii) those that give both necessary and sufficient conditions. Theorem 1 of the last section governing geometric series was of the last type; however it is more common to find convergence tests of the first two types since they are usually easier to come by. You should be careful to observe the nature of a test . A simple theorem should make the point clear. Theorem 1.6 . If the series/summationtext∞ k=1ak—whereakmay be complex—converges, then lim k→∞|ak|= 0. Proof: LetSn=a1+a2+···+an. Then |an|=|Sn−Sn−1|. Asn→ ∞ bothSnand Sn−1tend to the same limit, so |an| →0 . Returning to the point made before, this theorem states a necessary but not sufficient (as we shall see) condition for an infinite series to converge. We can apply it to see that/summationtextk k+1diverges—sincek k+1→1/negationslash= 0 . Thus this theorem is useful as a quick crude test to weed out series which diverge badly. But all it tells us about the series/summationtext∞ k=11 k—for which 1 k→0 so the criterion of the theorem is satisfied—is that it might converge. In fact, this series diverges too, as we shall now prove. ∞/summationdisplay k=11 k= 1 +1 2+1 3+1 4+1 5+···+1 8+1 9+···+1 16+1 17+···+1 32+··· 1.2. TESTS FOR CONVERGENCE OF POSITIVE SERIES 37 1 +1 2+1 4+1 4+1 8+···+1 8+1 16+···+1 16+1 32+···+1 32+··· = 1 +1 2+1 2+1 2+1 2+1 2+···.. ThusS1= 1, S2= 1+1 2, S4>1+1 2+1 2= 1+2 ·1 2, S8>1+3·1 2, S16>1+4·1 2,...,S 2n> 1 +n·1 2.We can easily see that as n→ ∞, S2n→ ∞ , so the series/summationtext1 k, called the harmonic series , diverges. For the many series which slip through the test of Theorem 6, more refined criteria are needed. The criteria we shall present in the remainder of this section are for series with positive terms,an≥0 . Application of these criteria to series with complex terms will be made in the next section. Theorem 1.7 . Ifak≥0for eachk, then the series/summationtext∞ k=1akconverges if and only if the sequence of partial sums is bounded from above. Proof: Since all the ak’s are non-negative, Sn+1≥Sn. Thus the Sn’s are a monotone increasing sequence of real numbers. By Theorems 6 and 8 of Chapter 0, this sequence Sn converge if and only if it is bounded. Example: The series/summationtext∞ k=11 k!of positive terms converges, since 1 k!=1 1·2·3···.k≤1 1·2·2·2···.2=1 2k−1 so Sn=n/summationdisplay k=11 k!≤n/summationdisplay k=11 2k−1≤∞/summationdisplay k=01 2k= 2. The convergence now follows since Snis bounded from above. We can extract an exceedingly useful idea from these examples: check the convergence of a given series by comparing it with another series which we know to converge or diverge. Theorem 1.8 . (comparison test ) Let/summationtextakand/summationtextbkbe two positive series for which ak≤bkforn>N . Then i) if/summationtextbkconverges, so does/summationtextak. ii) if/summationtextakdiverges, so does/summationtextbk. Proof: Letsn=/summationtextn k+N+1akandtn=/summationtextn k+N+1bk. Thensn≤tnfor alln>N , so i) if tn→t, thensnis bounded ( sn≤t) , ii) ifsn→ ∞ , thentn→ ∞ too. Remark: The “n>N ” part of the hypothesis reflects the fact that it is only the infinite tail of an infinite series that we need to worry about. Any finite number of terms can always be added later on. Examples: (a)/summationtext∞ k=11 2k+1converges since1 2k+1<1 2kand/summationtext1 2kconverges. (b)/summationtext∞ k=11√ kdiverges since1√ k≥1 k(fork≥1 ) and/summationtext1 kdiverges. Our next test is based upon comparison with a geometric series/summationtextrn. 38 CHAPTER 1. INFINITE SERIES Theorem 1.9 . (ratio test ) Let/summationtextanbe a series with positive terms such that the following limit exists lim n→∞an+1 an=L. Then i) ifL<1, the series converges ii) ifL>1, the series diverges iii) ifL= 1, the test is inconclusive. Remark: If the assumed limit does not exist, a variant of the theorem is still true but we shall not discuss it. Proof: i) IfL < 1 , pick any r, L < r < 1 . Then there is an Nsuch that for all n≥N,an+1 an< r. Therefore an< ra n−1< r2an−2< ... < rn−NaN, so thatan< Krn, n≥N, whereK >aN rN. The series/summationtext∞ n=1an=/summationtextN−1 n=1an+/summationtext∞ n=Nanconsists of a finite sum plus an infinite tail which is dominated by the geometric series/summationtextKrn. Since r<1 , the geometric series converges and by the comparison test, so does/summationtextan. ii) IfL>1 , thenan+1>anfor alln>N ; thus lim n→∞an/negationslash= 0 . By Theorem 6, the series/summationtextancannot converge. iii) This is seen from the two examples. (a)/summationtext1 n,with lim n→∞an+1 an= lim n→∞n n+1= 1 , which we know diverges. (b)/summationtext1 n(n+1),with lim n→∞an+1 an= lim n→∞(n+1)(n+2) n(n+)= 1 , which we know (Theorem 2, Example a) converges. In both these cases L= 1 . You should notice that the criterion uses the limiting value ofan+1/an. The divergent harmonic series/summationtext1 n, whose ratio n/n+ 1 is less than one for finiten, but 1 in the limit shows the mistake you will make if you use the ratio before passing to the limit. Examples: (a)/summationtext1 n!: Since lim n→∞(an+1 an) = lim n→∞(n! (n+1)!) = lim n→∞1 n+1= 0<1 , the ratio is less than one so the series converges. (b)/summationtext10n n!:Since limn→∞(an+1 an) = lim n→∞(10 n+1) = 0<1 , the series converges. (c)/summationtextn! 2n:Since limn→∞(an+1 an) = lim n→∞(n+1 2) =∞, the series diverges. Our last test for series with positive terms is associated with a picture. The crux of the matter is very simple and clever. We associate an area with the infinite series/summationtext∞ n=1an. For the term anwe use a rectangle between n≤x≤n+ 1 of height anand base one. Then the sum of the infinite series is represented by total area under the rectangles. Now by Theorem 7, if all the an’s are positive we know the series converges if the total area is finite. Thus, if we can find a function f(x) whose graph lies above the rectangles, and whose total area is finite, then we know the area contained in the rectangles is finite and so the series converges. Theorem 1.10 . (integral test ) Let/summationtext∞ n=ianbe a series of positive decreasing terms: 0< a n+1≤an, andf(x)a continuous decreasing function with f(n) =an. Then the sequence SN=N/summationdisplay n=1anandTN=/integraldisplayN 1f(x)dx 1.2. TESTS FOR CONVERGENCE OF POSITIVE SERIES 39 either both converge or both diverge, in fact, SN−a1≤TN≤SN−1. Proof: First of all, /integraldisplayN 1f(x)dx=/integraldisplay2 1+/integraldisplay3 2+···+/integraldisplayN N−1=N−1/summationdisplay n=1/integraldisplayn+1 nf(x)dx. Since in the interval n≤x≤n+ 1 we know that an=f(n)≥f(x)≥f(n+ 1) =an+1, we see that an=/integraldisplayn+1 nf(n)dx≥/integraldisplayn+1 nf(x)dx≥/integraldisplayn+1 nf(n+ 1)dx=an+1. Adding these up, we find N−1/summationdisplay n=1an≥N−1/summationdisplay n=1/integraldisplayn+1 nf(x)dx≥N−1/summationdisplay n=1an+1 or N−1/summationdisplay n=1an≥/integraldisplayN 1f(x)dx≥N/summationdisplay n=2an. Thus SN−1≥TN≥SN−a1. From this last inequality, we see that lim n→ ∞TNis finite if and only if lim n→∞SN is finite. Since the sequences SNandTNare both monotone increasing sequences, by Theorem 7 the sequences converge or diverge together. And we are done. Examples: (a)./summationtext∞ n=11 npconverges if p>1 , diverges if p≤1 . We use the function f(x) =1 xp, which satisfies the hypothesis of the theorem, and examine the integral TN=/integraldisplayN 11 xpdx=/braceleftBigg N1−p−1 1−p, p/negationslash= 1. lnN , p = 1/bracerightBigg . AsN→ ∞,lnN→ ∞ , and so does N1−pifp<1 , whileN1−p→0 ifp>1 . Therefore TNconverges if and only if p >1 , so by our theorem/summationtext∞ n=11 npconverges if and only if p > 1 . In the special case p= 1 we have again proven that the harmonic series/summationtext1 n diverges. Another often seen special case is p= 2,/summationtext1 n2, which converges. Sometime later we shall prove the amazing/summationtext∞ n=11 n2=π2 6. (b)/summationtext∞ n=21 nlnndiverges since as N→ ∞,/integraltextN 2dx xln 2= ln(lnN)−ln(ln 2) → ∞ Exercises (1) Determine if the following series converge or diverge. (a)/summationtext∞ n=11 n2+1 40 CHAPTER 1. INFINITE SERIES (b)/summationtext∞ n=11 2n−1 (c)/summationtext∞ n=11 n(lnn)2 (d)/summationtext∞ n=11 10n2 (e)/summationtext∞ n=1n n2+1 (f)/summationtext∞ n=11 2n+3 (g)/summationtext∞ n=1cos2n 2n (h)/summationtext∞ n=1√n n3+1 (i)/summationtext∞ n=1n2 2n (j)/summationtext∞ n=1ne−n2 (k)/summationtext∞ n=11√ n(n+1)(n+2) (l)/summationtext∞ n=1n! 22n (m)/summationtext∞ n=1|an| 10n,|an|<10 (n)/summationtext∞ n=1npe−n, p∈R (2) Ifan≥0 andbn≥0 for alln≥1 , and if there is a constant csuch thatan≤cbn, prove that the convergence of/summationtextbnimplies the convergence of/summationtextan. (3) Use the geometric idea of the integral test to show lim n→∞[1 +1 2+···+1 n−lnn] converges to a constant γ, and show that1 2<γ < 1 .γis called Euler’s constant . (4) If/summationtextanconverges, where an≥0 , prove that/summationtextan 1+analso converges. (5) (a). If/summationtextanconverges, where an≥0 , andcnhave the property 0 ≤cn≤K, the sameKfor alln, then prove that/summationtextcnanconverges. (b). Deduce the result of Exercise 4 from Exercise 5a. (6) Use the geometric idea behind the integral test to prove that (a). lnn! = ln 1 + ln 2 + ln 3 + ···+ lnn >/integraltextn 1lnxdx=nlnn−n+ 1 whenn≥2 . From this deduce that (b).n!>e(n e)n, whenn≥2 . (c). As an application of (b), prove that lim n→∞xn n!= 0 . (7) (a). Use the idea in the proof of the divergence of the harmonic series,/summationtext1 n, to prove the following test for convergence: Let {an}be a positive monotonically decreasing sequence. Then/summationtextanconverges or diverges respectively if and only if the “condensed” series/summationtext2na2nconverges or diverges. (b). Apply the test of part (a) to again prove that/summationtext1 npconverges if p >1 , and diverges if p≤1 . (c). Apply the test of part (a) to determine the values of pfor which the series/summationtext∞ n=21 n(lnn)pconverges and diverges. 1.3. ABSOLUTE AND CONDITIONAL CONVERGENCE 41 1.3 Absolute and Conditional Convergence The tests just given for series with positive terms can be applied to many series with complex termsanby utilizing the concept of absolute convergence. Definition: The series/summationtext∞ k=1ak, where the akmay be complex numbers, converges absolutely if the series of positive numbers/summationtext∞ k=1|ak|converges. It is called conditionally convergent if/summationtext∞ k=1akconverges but/summationtext∞ k=1|ak|diverges. Absolute convergence is stronger than ordinary convergence because Theorem 1.11 . If/summationtext∞ n=1|an|converges, then/summationtext∞ n=1anconverges. Proof: LetaN=αn+iβn. We shall show that the real series/summationtextαnand/summationtextβnboth converge. Then by Theorem 5/summationtextanconverges too. To show that/summationtextαnconverges, let cn=αn+|an|. Since |αn| ≤/radicalbig (α2n+β2n) =|an|, we know that 0 ≤cn≤2|an|. Thus the positive series/summationtextcnis bounded,/summationtextcn≤2/summationtext|an|<infty, and so converges by the comparison test (Theorem 8). Since/summationtextαn=/summationtext(cn− |an|) , and both/summationtextcnand/summationtext|an| converge, then/summationtextαnalso converges by Theorem 4. Similarly, by taking dn=βn+|an|, the series/summationtextdnconverges, from which we can conclude that/summationtextβnconverges. Examples: (a) The complex series1 12+i 22+i2 32+i3 42+···=/summationtext∞ n=1in n2converges absolutely since/vextendsingle/vextendsingle/vextendsinglein−1 n2/vextendsingle/vextendsingle/vextendsingle=1 n2and the positive series/summationtext∞ n=11 n2converges. (b) 1 +1 22−1 23−1 24+1 25+1 26−1 27−1 28+···, which is the geometric series/summationtext1 2nwith negative signs thrown in, converges absolutely since/summationtext1 2nconverges. (c)/summationtextrn, rcomplex, converges absolutely if/summationtext|r|nconverges, that is, if |r|<1 . (d) 1−1 2+1 3−1 4+1 5...=/summationtext∞ n=1(−1)n+1 n, the alternating harmonic series does not converge absolutely because/summationtext1 ndiverges. It does converge though, as we shall see shortly. Thus the alternating harmonic series is conditionally convergent. On the basis of this last theorem, many complex series can be proved to converge by proving they converge absolutely. Since absolute convergence concerns itself with series having only positive terms, all the tests for convergence developed in the previous section may be used. This is the most common way of proving a complex series converges. If it does not converge absolutely, the proof of convergence will usually be more difficult and use special ingenuity based on the particular series at hand. There is one case of conditional convergence which is easy to treat, that of alternating series. Definition: A series of real numbers is called alternating if the positive and negative terms occur alternately. They have the form ∞/summationdisplay n=1(−1)n−1an=a1−a2+a3−a4+···, where thean’s are all positive. 42 CHAPTER 1. INFINITE SERIES Theorem 1.12 . The alternating series/summationtext∞ n=1(−1)n−1an, an>0, converges if i) the an are monotone decreasing ( an/arrowsoutheast), and ii) limn→∞an= 0. IfSis the sum of the series, the inequality 0<|S−SN|<aN+1 (1-2) shows how much the Nthpartial sum differs from the limit S. In words inequality (2) says that the error which results by using the first Nterms is less than the first neglected termaN+1. Proof: The idea is quite simple. Observe that since an/arrowsoutheast, S2n−S2n−2=a2n−1−a2n>0 , so theS2n’s increase. Similarly the S2n+1’s decrease. Also both sequences are bounded— from below by S2and from above by S1(you should check this). Therefore by Theorem 8 Chapter 0, the bounded monotonic sequences S2nandS2n+1converge to real numbers Sand ˆSrespectively. Let us show that S=ˆS. ˆS−S= lim n→∞S2n+1−lim n→∞S2n= lim n→∞(S2n+1−S2n) = lim n→∞a2n+1= 0 Thus the alternating series converges to the unique limit S. All that is left to verify is inequality (2). Because S2nis increasing and S2n+1is decreasing, we know that S2n<S andS <S 2n+1 Therefore 0<S−S2n<S 2n+1−S2n=a2n+1 and 0<S 2n−1−S <S 2n−1−S2n=a2n. These two inequalities are the cases Neven andNodd in (2). Examples: (a)/summationtext∞ n=1(−1)n−1 nconverges since it is an alternating sequence and1 ndecreases mono- tonically to zero. Later we shall show that its sum is ln 2 . (b)/summationtext∞ n=2(−1)n lnnconverges since1 lnndecreases monotonically to zero. (c)/summationtext∞ n=1(−1)n−1n n+1diverges by Theorem 6 since lim n→∞(−1)n−1n n+1is not zero. Exercises (1) Determine which of the following series converge absolutely, converge conditionally, or diverge. (a)/summationtext∞ n=1(−1)n+1 √n (b)/summationtext∞ n=1(2−3i)n n! (c)/summationtext∞ k=2(2k+i)2 ek (d)/summationtext∞ n=1(−1)n−1lnn n (e) 1 −1 2+1 3−1 22+1 5−1 23+1 7−1 24+1 9− ···. (f)/summationtext∞ n=11 n2+2i 1.4. POWER SERIES, INFINITE SERIES OF FUNCTIONS 43 (g)/summationtext∞ n=11 n+2i (h)/summationtext∞ n=1(−1)n−1 n+2i (i)/summationtext∞ n=1(−1)n−1 np, p> 0, (j)/summationtext∞ n=1(−1)n(1+i)n2 2n2+1 (k)/summationtext∞ n=1cosnθ n2, θarbitrary. (2) If/summationtextanand/summationtextbnare absolutely convergent, and αandβare any complex numbers, prove that/summationtext(αan+βbn) also converges absolutely. (3) Show that/summationtext∞ n=1nznconverges absolutely if |z|<1 . (4) Show that for any θ∈R, then/summationtext∞ n=0cosnθdiverges, and that if θ/negationslash= 0,±π,±2π,... , then/summationtext∞ n=0sinnθalso diverges. 1.4 Power Series, Infinite Series of Functions As you will all agree, the simplest functions are polynomials. With infinite series at hand, it is reasonable to consider an “infinite polynomial” a0+a1z+a2z2+a3a3+···=∞/summationdisplay n=0anzn. Because of the appearance of the powers of z, this is called a power series . The question of convergence of a power series is trivial at z= 0 , for then we have only the one term a0. Does this series converge for any other values of z, and if so, for which ones? The answer depends on the coefficients an, but in any case, the set of complex numbers, z∈C, for which the series converges is always a disc |z|<ρ—with possibly some additional points on the boundary |z|=ρ—in the complex pane bC. This number ρis called the radius of convergence of the power series. We shall first prove that the set z∈Cfor which a power series converges is always a disc. Then we shall give a way of computing the radius ρof that disc. Theorem 1.13 . The setz∈Cfor which the power series/summationtextanznconverges is always a disc |z|< ρ, inside of which it even converges absolutely. We do not exclude the two extreme possibilities that the radius of this disc is zero or infinity. The series might converge at some, none, or all of the points on the boundary of the disk|z|=ρ. Proof: We shall show that if the series converges for any ζ∈C, then it converges absolutely for all complex zwith|z|<|ζ|. Ifζ= 0 , there is nothing to prove, so assume ζ/negationslash= 0 . Because/summationtextanζnconverges, lim n→∞|anζn| →0 . Thus all the terms are bounded in absolute value, that is, there is an Msuch that |anζn|<M for alln. Then, since |anzn|=/vextendsingle/vextendsingle/vextendsingle/vextendsingleanζnzn ζn/vextendsingle/vextendsingle/vextendsingle/vextendsingle<M/vextendsingle/vextendsingle/vextendsingle/vextendsinglez ζ/vextendsingle/vextendsingle/vextendsingle/vextendsinglen for alln, 44 CHAPTER 1. INFINITE SERIES the series/summationtext|anzn|is dominated by M/summationtext/vextendsingle/vextendsingle/vextendsinglez ζ/vextendsingle/vextendsingle/vextendsinglen . But this last series is a geometric series which does converge since |z|<|ζ|, so/vextendsingle/vextendsingle/vextendsinglez ζ/vextendsingle/vextendsingle/vextendsingle<1 . Thus by the comparison test/summationtextanzn converges absolutely for all z∈Cwith|z|<|ζ|. Therefore, if the power series/summationtextanznconverges for some complex number ζ, then it converges in the whole disc |z|<|ζ|. The radius of convergence ρis then the radius of the largest disc |z|<ρfor which the series converges. See Exercise 3 for examples concerning convergence on the boundary of the disk. Let us now give a method of computing ρwhich covers most cases arising in practices. Theorem 1.14 . If limn→∞/vextendsingle/vextendsingle/vextendsinglean+1 an/vextendsingle/vextendsingle/vextendsingle=Lexists, the power series/summationtextanznhas radius of convergence ρ=1 LifL/negationslash= 0,∞ifL= 0. In other words, if L/negationslash= 0 the series converges in the disc |z|<1 Land diverges if |z|>1 L. On the circumference |z|= 1/L, anything may happen (see Exercise 3 at the end of this section). If L= 0, the series converges in the whole complex plane. Proof: This is a simple application of the ratio test. The series converges if the limit of the ratio of successive terms lim n→∞/vextendsingle/vextendsingle/vextendsinglean+1zn+1 anzn/vextendsingle/vextendsingle/vextendsingleis less than one and diverges if it is greater than one. Thus we have convergence if lim n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglean+1z an/vextendsingle/vextendsingle/vextendsingle/vextendsingle=|z|L<1,i.e. if |z|<1 L, and divergence if lim n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglean+1z an/vextendsingle/vextendsingle/vextendsingle/vextendsingle=|z|L>1,i.e. if |z|>1 L. Remark: In the one additional case/vextendsingle/vextendsingle/vextendsinglean+1 an/vextendsingle/vextendsingle/vextendsingle→ ∞ asn→ ∞ , the series diverges for every |z| /negationslash= 0 , as can easily be seen again by the ratio test. Examples: (a)/summationtext∞ n=0znconverges where lim n→∞/vextendsingle/vextendsinglezn+1/zn/vextendsingle/vextendsingle<1 that is; for |z|<1 . (b)/summationtext∞ n=0nzn 2nconverges where lim n→∞/vextendsingle/vextendsingle/vextendsingle(n+1)zn+1 2n+1/nzn 2n/vextendsingle/vextendsingle/vextendsingle<1 . Since lim n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsingle(n+ 1)zn+1 2n+1/nzn 2n/vextendsingle/vextendsingle/vextendsingle/vextendsingle= lim n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsingle(n+ 1)z 2n/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsinglez 2/vextendsingle/vextendsingle/vextendsingle, the series converges for all |z|<2 . (c)/summationtext∞ n=0zn n!converges where lim n→∞/vextendsingle/vextendsingle/vextendsinglezn+1 (n+1)!/zn n!/vextendsingle/vextendsingle/vextendsingle<1 . Since lim n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglezn+1 (n+ 1)!/zn n!/vextendsingle/vextendsingle/vextendsingle/vextendsingle= lim n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglez n+ 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle= 0, the series converges for all z∈C, that is, in the whole complex plane. 1.4. POWER SERIES, INFINITE SERIES OF FUNCTIONS 45 (d)/summationtext∞ n=0n!znconverges where lim n→∞/vextendsingle/vextendsingle/vextendsingle(n+1)!zn+1 n!zn/vextendsingle/vextendsingle/vextendsingle<1 . But lim n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsingle(n+ 1)!zn+1 n!zn/vextendsingle/vextendsingle/vextendsingle/vextendsingle= lim n→∞|(n+ 1)z|=∞ unlessz= 0 . Thus the ratio is less than one only at z= 0 , so the series converges only at the origin. Only minor modifications are needed for the more general power series a0+a1(z−z0) +a2(z−z0)2+···=∞/summationdisplay n=0an(z−z0)n, wherea0∈C. Again the series converges in a disc in the complex plane, only now the disc has its center at z0instead of the origin, so if the radius of convergence is ρ, the series converges for |z−z0|<ρ. An example should make this clear. Example:/summationtext∞ n=1(z−2i)n n. By the ratio test, this converges when lim n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsingle(z−2i)n+1 n+ 1/(z−2i)n n/vextendsingle/vextendsingle/vextendsingle/vextendsingle<1, that is, when |z−2i|<1 . This is a disc with center at 2 iand radius 1. A few words should be said about real power series/summationtextan(x−x0)nwhere both xand x0are real (some people only use this phrase if the anare also real). This is a special case of/summationtextan(z−z0)nwherez0is on the real axis and we only ask for what realzthe series converges. However we know that/summationtextan(z−z0)nconverges only for those zin the disc of convergence |z−z0|<ρ—and possibly some boundary points. Thus the realvalues of zfor which the series/summationtextan(z−z0)nconverges are exactly those points on the real axis which are also inside the disc of convergence of the complex power series. In particular the series/summationtextan(x−x0)n, with both xandx0real converges for |x−x0|< ρ, i.e., in the intervalx0−ρ≤x≤x0+ρ. Example: For whatx∈Rdoes/summationtext∞ n=01 2n(x−1)nconverge? The related complex series/summationtext∞ n=01 2n(z−1)nconverges in the disc |z−1|<2 . The points on the real axis which are in this disc are |x−1|<2 , which is −1<x< 3 . A direct check shows the series diverges at both end points x=−1 andx= 3 . If/summationtextanand/summationtextbnboth converge, can we define their product in a meaningful way (∞/summationdisplay n=0an)(∞/summationdisplay n=0bn) =∞/summationdisplay n=0cn? and if so, does the resulting series converge? The most simple-minded approach is to insert powers ofz(a bookkeeping device), giving (/summationtextanzn)(/summationtextbnzn) , try long multiplication and see what happens. A computation shows that (a0+a1z+axz2+···)(b0+b1z+b2z2+···) =a0b0+ (a0b1+a1b0)z +(a0b2+a1b1+a2b0)z2+···+ (a0bn+a1bn−1+···+anb0)zn+···. 46 CHAPTER 1. INFINITE SERIES Motivated by this, we make the following Definition: The formal product , called the Cauchy product , of the series/summationtextanand/summationtextbn is defined to be (∞/summationdisplay n=0an)(∞/summationdisplay n=0bn)≡∞/summationdisplay n=0cn, where cn=a0bn+azbn−1+···+anb0=n/summationdisplay k=0akbn−k. With this definition we shall answer the question we raised about multiplication of power series. Theorem 1.15 . If/summationtext∞ n=0an=Aand/summationtext∞ n=0bn=Bboth converge absolutely, then the Cauchy product series (∞/summationdisplay n=0an)(∞/summationdisplay n=0bn)≡(∞/summationdisplay n=0cn), where cn=∞/summationdisplay k=0akbn−k, also converges absolutely, and to C=AB. Proof: LetAN=/summationtextN n=0an, BN=/summationtextN n=0bn,andCN=/summationtextN n=0cn. We shall show that by pickingNlarge enough, |ANBN−CN|can be made arbitrarily small. Since ANBN→ AB, this will complete the proof. Observe that CN=a0b0+ (a0b1+a1b0) +···+ (a0bN+···+aNb0) =/summationdisplay/summationdisplay ajbk, while ANBN= (a0+···+aN)(b0+···+bN) =N/summationdisplay j=0N/summationdisplay k=0ajbk. Therefore |ANBN−CN|=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingleN/summationdisplay j=0N/summationdisplay k=0ajbk/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤N/summationdisplay j=0N/summationdisplay k=0|aj| |bk|. Sincej+k>N , eitherj >N/ 2 ork>N/ 2 , so |ANBN−CN| ≤N/summationdisplay j>N 2N/summationdisplay k=0|aj| |bk|+N/summationdisplay j=0N/summationdisplay k>N 2|aj| |bk|. Because the original series both converge absolutely, they are bounded, ∞/summationdisplay j=0|aj|<M and∞/summationdisplay k=0|bk|<M. Consequently, |ANBN−CN| ≤M(∞/summationdisplay j>N 2|aj|+∞/summationdisplay k>N 2|bk|). 1.4. POWER SERIES, INFINITE SERIES OF FUNCTIONS 47 Again using the absolute convergence of the original series, we see that for Nlarge, the right side can be made arbitrarily small. Since we shall need the ideas later on, let us digress briefly and examine the convergence of infinite series of functions,/summationtextun(z) . In the special case where un(z) =an(z−z0)n, this is a power series. Generally, there is little one can say about the convergence of such series except to apply our general tests and hope for the best. We shall only illustrate the situation with two Examples: (a)/summationtext∞ n=1cosnθ n2, whereθis any real number. This converges for all θsince it converges absolutely, that is/summationtext+/vextendsingle/vextendsinglecosnθ n2/vextendsingle/vextendsingleconverges. We can see this last statement is true by comparing/summationtext+/vextendsingle/vextendsinglecosnθ n2/vextendsingle/vextendsinglewith the larger convergent series (since |cosnθ| ≤ 1 )/summationtext∞ n=11 n2. (b)/summationtext∞ n=1nenx. By the ratio test, converges if lim n→∞/vextendsingle/vextendsingle(n+ 1)e(n+1)x/nenx/vextendsingle/vextendsingle<1 . Since lim n→∞/vextendsingle/vextendsingle/vextendsingle(u+ 1)e(n+1)x/nenx/vextendsingle/vextendsingle/vextendsingle= lim n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglen+ 1 n/vextendsingle/vextendsingle/vextendsingle/vextendsingleex=ex, the series converges if ex<1 , which happens only when x<0 . Exercises (1) Find the disc of convergence of the following power series by finding the center and radius of the disc. (a)/summationtext∞ n=0zn n+1 (b)/summationtext∞ n=0(z−2)n n (c)/summationtext∞ n=0in 2n−1zn−1 (d)/summationtext∞ n=0(n+ 1)[z−2 + 3i]n (e)/summationtext∞ n=0(2z+3)n n2+2i (f)/summationtext∞ n=01 lnnzn−2 (g)/summationtext∞ n=0(2n−i) 3nzn (h)/summationtext∞ n=02nzn n!(0!≡1) (i)/summationtext∞ n=0(1 2n+i 3nzn (j)/summationtext∞ n=0(z+i)n 22n (k)/summationtext∞ n=0z2n (2n)! (l)/summationtext∞ n=0nn(z−1)n (m)/summationtext∞ n=0z2n 4n (n)/summationtext∞ n=0(1 n+i n2+1)(z−√ 2i)n (2) Find the set x∈Rfor which the following series converge. 48 CHAPTER 1. INFINITE SERIES (a)/summationtext∞ n=0(x−1)n n2n (b)/summationtext∞ n=0cosnx 2n (c)/summationtext∞ n=01 n(x−1 x)n (d)/summationtext∞ n=0e−n(x+1) (e)/summationtext∞ n=02n(sinx)n n (f)/summationtext∞ n=0(1 +ex)n (g)/summationtext∞ n=0(1−ex)n (3) The point of this exercise is to show that a power series might converge at some, none, or all of the points on the boundary of the disk of convergence. (a) Show that/summationtext∞ n=0zndiverges at every point on the boundary of its disc of con- vergence. (b) Show that/summationtext∞ n=0zn n+1diverges for z= 1 but converges for z=−1 (in fact, it converges everywhere on |z|= 1 except at z= 1 ). (c) Show that/summationtext∞ n=0xn (n+1)2converges at every point on the boundary of its disc of convergence. (4) If/summationtextanzndiverges for z=ζ∈C, prove that it diverges for all z∈Cwith|z|>|ζ|. (5) For what z∈Cdoes/summationtext∞ n=0z2 (1+z2)nconverge? Find a formula for the nth partial sumSn(z) . Evaluate lim n→∞Sn(z) . Is the limit function continuous? (6) Let/summationtext∞ n=0P(n)anznhave radius of convergence rho, and letP(n) be any polyno- mial. Prove that/summationtext∞ n=0P(n)anznconverges and also has ρas its radius of conver- gence. (By P(n) w e mean P(n) =Aknk+Ak−1nk−1+···+A1n+A0). 1.5 Properties of Functions Represented by Power Series Having found that a power series/summationtextan(z−z0)nconverges in some disc, |z−z0|< ρ, it is interesting to study the function f(z) defined by the power series for zin the disc of convergence f(z) =∞/summationdisplay n=0an(z−z0)n, |z−z0|<ρ. It turns out that functions f(z) defined by a convergent power series are delightful, as nicely behaved as functions can be. In particular, they are not only continuous, but also automatically have an infinite number of continuous derivatives and many other amazing properties. This section will be devoted to proving the more elementary properties of functions represented by power series, while in the next section we will begin with given functions, like sinx, and see if there is a convergent power series associated with the m, as well as showing a way of obtaining the coefficients anof that power series. The profound theory of functions represented by convergent power series is called analytic functions of a complex variable . 1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 49 Definition: A function f(z) of the complex variable zis said to be analytic in the disc |z−z0|<ρiff(z) can be represented by a convergent power series in that disc: f(z) =∞/summationdisplay n=0an(z−z0)n, |z−z0|<ρ. Since we have not developed the notion of the derivative,d f dz, of a complex valued functionf(z) of the complex variable z, nor have we considered the corresponding theory of integration,/integraltext f(z)dz, the scope of our treatment will regrettably have to be narrowed. However our proofs will have the property that as soon as an adequate theory of differen- tiation and integration is given, the theorems and proofs remain unchanged. Instead of considering power series in the complex variable z, we shall restrict our attention to series in the real variable x f(x) =∞/summationdisplay n=0an(x−x0)n,|x−x0|<ρ, (1-3) still allowing the coefficients anto be complex. Thus, f(x) is a complex-valued function of the real variable x. The definitions of derivative and integral for such functions were given in Section 7 of Chapter 0. We shall use that material here . Our aim is the following: Theorem 1.16 . Suppose that/summationtext∞ n=0anxnhas radius of convergence ρ>0(possibly ∞). Then (a)the function f(x)defined by f(x) =∞/summationdisplay n=0anxn,|x|<ρ, has an infinite number of derivatives; (b)the series/summationtext∞ n=0nanxn−1has the same radius of convergence ρand f/prime(x) =∞/summationdisplay n=0nanxn−1,|x|<ρ, and (c)the series/summationtext∞ n=0an n+1xn+1has the same radius of convergence ρ, and /integraldisplayx 0f(t)dt=∞/summationdisplay n=0an n+ 1xn+1,|x|<ρ. Remark: If we omit f(x) from the picture and write (b) and (c) directly in terms of the infinite sum, we find (b)/primed dx[∞/summationdisplay n=0anxn] =∞/summationdisplay n=0nanxn−1 and (c)/prime/integraldisplayx 0[∞/summationdisplay n=0antn]dt=∞/summationdisplay n=0an n+ 1xn−1. 50 CHAPTER 1. INFINITE SERIES These two statements are usually abbreviated “a power series may be differentiated term by term” and “a power series may be integrated term by term” within their domain of convergence (these statements are notgenerally true for an arbitrary infinite series of functions/summationtextun(x) , see Exercise 4 below). The generalization to/summationtext∞ n=0an(x−x0)nis obvious. Our proof will be given in several parts. We begin with the Lemma 1 . Under the hypothesis of the theorem, f(x) is continuous for all ˜ xwith|˜x|<ρ.Proof: (This is a little dull). Given any /epsilon1>0 , we must find a δ>0 such that |f(x)−f(˜x)|</epsilon1when |x−˜x|<δ. Let us write fN(x) =/summationtextN n=0anxnandRN(x) =/summationtext∞ N+1anxn, so thatf(x) =fN(x) + RN(x) . Observe that |f(x)−f(˜x)|=|fN(x)−fN(˜x) +RN(x)−RN(˜x)| ≤ |fN(x)−fN(˜x)|+ |RN(x)|+|RN(˜x)|. We shall show that each of these three terms can be made </epsilon1 3by picking xclose enough to ˜xandN-which is entirely at our disposal- large enough. First work with RN(x) andRN(˜x) . Choose rsuch that |˜x|< r < ρ . This is to insure that we stay away from the boundary |x|=ρwhere the series may diverge. Then/summationtext∞ n=0|anrn|converges absolutely, say to the number S. If we let SN=/summationtextN 0|anrn|, we know that by picking Nlarge enough,/summationtextN N+1|anrn|=S−SN</epsilon1 3. But |RN(x)|=/vextendsingle/vextendsingle/summationtext∞ N+1anxn≤/summationtext∞ N+1|anxn|/vextendsingle/vextendsingle, so that if|x|+≤r,by using the same Nfound above, we have |RN(x)| ≤∞/summationdisplay N+1|anrn|=S−SN</epsilon1 3. Since by the definition of rwe know |˜x| ≤r, this also proves that for this same N|RN(˜x)|</epsilon1 3. Thus by restricting |x| ≤r, we have seen that both |RN(x)|and|RN(˜x)| can be made less than/epsilon1 3. Having fixed N, f N(x) is a polynomial -which we know is continuous. Thus there is a δ, >0 such that |fN(x)−fN(˜x)|</epsilon1 3when |x−˜x|<δ 1. This shows that |f(x)−f(˜x)|</epsilon1ifxis in the intersection of the intervals |x| ≤ |˜x|< r<ρ ) and |x−˜x|<δ 1. That there is some interval contained in both of these intervals is easy to see since both contain all points sufficiently close to ˜ x. And the proof is completed. As you have observed, the proof involves no new ideas but is rather technical. With this lemma proved, we know that f(x) is continuous -and hence integrable. Thus we can work with/integraltextx 0f(t)dt. Our next task is to prove a portion of Part (c) of Theorem 16. Lemma 1.17 If/summationtext∞ n=0anxnhas radius of convergence ρ>0, then ∞/summationdisplay n=0an n+ 1xn+1=/integraldisplayx 0f(t)dtfor all |x|<ρ. Proof: We shall show that /vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx 0f(t)dt−N/summationdisplay n=0an n+ 1xn+1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle(1-4) 1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 51 can be made arbitrarily small by choosing Nlarge enough. Write f(t) =N/summationdisplay n=0antn+∞/summationdisplay n=N+1antn. Then since we can integrate any finite sum term by term, we have /integraldisplayx 0f(t)dt=N/summationdisplay n=0an/integraldisplayx 0tndt+/integraldisplayx 0[∞/summationdisplay n=N+1antn]dt=N/summationdisplay n=1an n+ 1xn+1=/integraldisplayx 0[∞/summationdisplay n=N+1antn]dt, so that (4) reduces to showing that /vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx 0∞/summationdisplay n=N+1antndt/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle can be made small by choosing Nlarge. The idea here is to apply Theorem 16 of Chapter 0. This means we need to estimate the size of the above integrand. By now you should recognize the method. Because |x|< ρ, we can choose an rsuch that |x|< r < ρ . Then/summationtextanrnis convergent so its terms are bounded, say M≥ |anrn|for alln, that is, |an| ≤M rn. Therefore, since |t|<|x|, we find the inequality /vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle∞/summationdisplay N+1antn/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤∞/summationdisplay N+1|an||t|n≤∞/summationdisplay N+1M rn|x|n. But the last series is a geometric series whose sum is/vextendsingle/vextendsinglex r/vextendsingle/vextendsingleNM|x| r−x. Thus /vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle∞/summationdisplay N+1antn/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤/vextendsingle/vextendsingle/vextendsinglex r/vextendsingle/vextendsingle/vextendsingleNM|x| r−x. Applying Theorem 16 of Chapter 0, we find that /vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx 0(∞/summationdisplay N+1antn)dt/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤/vextendsingle/vextendsingle/vextendsinglex r/vextendsingle/vextendsingle/vextendsingleNM|x|2 r−x. that is,/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx 0f(t)dt−N/summationdisplay 0an n+ 1xn+1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤/vextendsingle/vextendsingle/vextendsinglex 4/vextendsingle/vextendsingle/vextendsingleNM|x|2 r−x. Since/vextendsingle/vextendsinglex r/vextendsingle/vextendsingle<1 , we know that/vextendsingle/vextendsinglex r/vextendsingle/vextendsingleN→0 asN→ ∞ , which completes the proof of the lemma. Incidentally, all we have left to prove of part c of the theorem is that the radius of convergence of the integrated series is no larger than ρ(since the lemma shows it is at least ρ). But this will have to wait until after Lemma 1.18 If/summationtextanxnhas radius of convergence ρ, the series obtained by formally differentiating term by term,/summationtextnanxn−1, has the same radius of convergence. 52 CHAPTER 1. INFINITE SERIES Remark: This lemma does notsay that the derived series is equal to the derivative of the function defined by the original series. It only discusses the radius of convergence, not the relationship of the functions represented b y the two series. Proof: Letρ1be the radius of convergence of/summationtextnanxn−1. First we show that ρ1≤ρ. If/summationtextnanxn−1converges for some fixed x, then so does/summationtextnanxn. But the terms of this last sequence are larger than those of/summationtextanxnsince|nanxn| ≥ |anxn|. Thus by the comparison test/summationtextanxnalso converges for that x, which shows ρ1≤ρ. To show that ρ≤ρ1, assume/summationtextanxnconverges for some xand choose rbetween |x| andρ,|x|<r<ρ . As in the proof of Lemma 2 we find that |an|<Mr−n. Then the terms in the series/summationtext/vextendsingle/vextendsinglenanxn−1/vextendsingle/vextendsingleare smaller than the corresponding terms in/summationtextnM r|x| rn−1. By the ratio test this last series converges, since |x|<r. Thus the derived series/summationtextnanxn−1 also converges, showing that ρ≤ρ1and completing the proof of the lemma. Now we can complete the proof of part c of Theorem 16. Corollary 1.19 If/summationtextanxnhas radius of convergence ρ, then the series obtained by for- mally integrating term by term,/summationtextan n+1xn+1also has radius of convergence ρ. Proof: The series/summationtextanxnis the formal derivative of the series/summationtextan n+1xn+1, and we have just seen that these two series have the same radius of convergence. We shall next prove part (b) of Theorem 16 as Lemma 1.20 f(x)≡/summationtext∞ n=0anxnhas radius of convergence ρ>0then df dx=d dx[∞/summationdisplay n=0anxn] =∞/summationdisplay n=0nanxn−1, and this series also has radius of convergence ρ. Proof: In Lemma 3 we proved that the radii of convergence are the same. What we must prove here is that the derivative of the function is given by the derivative of the series. This is a more or less immediate consequence of Lemma 2, for let us apply this integration lemma to the function g(x) defined by g(x)≡∞/summationdisplay n=1nanxn−1,|x|<ρ. Then we find that/integraldisplayx 0g(t)dt=∞/summationdisplay n=1anxn=f(x)−a0,|x|<ρ. By the fundamental theorem of calculus, we can take the derivative of the left side, and it isg(x) . Thus g(x) =f/prime(x), that is, ∞/summationdisplay n=1nanxn−1=d dxf(x). This incidentally also proves the otherwise not obvious fact that f(x) , only known to be continuous so far (Lemma 1) is also differentiable. 1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 53 To complete the proof of Theorem 16, we must prove Lemma 5. If the power series/summationtextanxnconverges for |x|< ρ, then the function f(x) defined by f(x)≡/summationtext∞ n=0anxn has an infinite number of derivatives. The derivatives are represented by the formal series obtained by term-by-term differentiation. Proof: By induction, Lemma 4 shows us that f(x) has one derivative. Assume f(x) has kderivatives. We shall show that is has k+ 1 . Letf(k)(x) =/summationtextbnxnbe the series for the kthderivative of f. Applying Lemma 4 to this series we find that f(k)(x) is differentiable. This proves that fhask+ 1 derivatives and completes the induction proof. Examples: (a) We know that 1 1 +t=∞/summationdisplay n=0(−t)n= 1−t+t2−t3+···. where the geometric series converges for |t|<1 . Applying the theorem, we integrate term by term to find that ln(1 +x) =/integraldisplayx 01 1 +tdt=∞/summationdisplay n=0(−1)n·xn+1 n+ 1,|x|<1, or ln(1 +x) =x−x2 2+x3 3−x4 4+x5 5+···. Thus the function ln(1 + x) is equal to the power series on the right. With a little more work we can prove that the series, which converges at x= 1 , converges to ln(1 + 1) and obtain the following interesting formula. ln 2 = 1 −1 2+1 3−1 4+1 5− ···. The power series for ln(1 + x) can be used to illustrate the possibilities of computing with infinite series. If 0 <x< 1 the series for ln(1 + x) is a strictly alternating series to which we can apply inequality (2) of Theorem 12. For this series it reads 0</vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingleln(1 +x)−k/summationdisplay n=0(−1)nxn+1 n+ 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle<xk+2 k+ 2, x> 0. This inequality states that if only the first kterms of the infinite series are used to compute ln(1 +x) , the error will be less thanxk+2 k+2. Say we want to compute ln(1 +1 4) = ln5 4to 5 decimal places. Then we want t o choose kso that 1 4k+2 k+ 2<1 1,000,000= 10−6 Cross-multiplying, writing 4 = 22, we wantksuch that 106<(k+ 2)22k+4, sincek+ 2≥ 2,22k+5≤(k+ 2)22k+4. Thus, we are done if we can find ksuch that 106≤22k+5. 54 CHAPTER 1. INFINITE SERIES But since 210= 1024>103, we know 220>106. Thus if 2 k+ 5≥20 , ork= 8 we will have the desired accuracy. This means that ln5 4=1 4−1 2(1 4)2+···+1 9(1 4)8+1+ error where the error is less than 10−6. From the form of the error estimate, it is clear that the series converges faster if xis smaller. This power series, valid only if |x|<1 can be used to compute ln(1+ x) if|x|>1 by utilizing the observation illustrated by ln 6 = 3 ln(3 2) + 2 ln(4 3) = 3 ln(1 +1 2) + 2 ln(1 +1 3), where both ln(1 +1 2) and ln(1 +1 3) can be computed using the power series. We should confess that this series converges too slowly to be of much value for that purpose in real life. (b) Since1 1+t2is also the sum of a geometric series 1 1 +t2= 1−t2+t4−t6+t8+···=∞/summationdisplay 0(−1)nt2n,|t|<1, if we integrate term by term, we find tan−1x=/integraldisplay2 0dt 1 +t2+∞/summationdisplay n=0(−1)nx2n+1 2n+ 1=x−x3 3+x5 5+···, which converges if |x|<1 . Further investigation shows that the series also converges atx= 1 and represents the function at that point. This yields the wonderful formula (obtained by letting x= 1 ) π 4= 1−1 3+1 5−1 7+··· from which we can compute πto any desired accuracy. Exercises (1) Write down an infinite series whose sum is1 1−tand integrate the series term by term to obtain a power series for ln(1 −x) . For what xdoes the series converge? (2) Find a power series which converges about x= 0 for the functionx (1−x)2by recog- nizing1 (1−x)2as the derivative of a function whose power series in known. For what xdoes the series converge? (3) Compute ln9 8to 4 decimal places, proving the error in your approximation is correct. (4) Show that/summationtext∞ n=1sinn2x n2converges for all xbut the series obtained by differentiating term-by-term does not converge, say at x= 0 . (5) Exercise your ingenuity and apply the theorems of this section to find the function whose power series is (a)a+ 2x2+ 4x4+ 6x6+ 8x8+···+ (2n)x2n+···. (b) 2 + 3 ·2x+ 4·3x2+ 5·4x3+···+ (k+ 2)(k+ 1)xk+··· 1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 55 6. Taylor’s Theorem. Representation of a Given Function in a Power Series. The Binomial Theorem. In this section we prove Taylor’s Theorem, an important generalization of the mean value theorem, and use it to investigate the questions i) when does a given function f(x) have a power series? and ii) if f(x) has a power series about x0, f(x) =/summationtext∞ n=0an(x−x0)n, how can we find the coefficients an? As a partial answer to i) we know from Theorem 16 of the last section that if f(x) has a power series about x0, it must necessarily have an infinite number of derivatives at x0. It turns out that this is not enough. Perhaps it is easiest to begin with question ii). Assumef(x) has a power series about x0, f(x) =∞/summationdisplay n=0an(x−x0)n, which converges for |x−x0|< ρ. How can we find the coefficients an? By Theorem 16 we know that fhas an infinite number of derivatives at x0. Moreover these derivatives can be calculated by differentiating the power series term-by-term. F or convenience we let x0= 0 . f(x) =a0+a1x+a2x2+a3x3+···+anxn+···, f/prime(x) =a1+ 2a2x+ 3a3x2+···+nanxn−1+···, f/prime/primex(x) = 2a2+ 2·3a3x+ 3·4·a4x2+···+n(n−1)anxn−2+···, f(3)(x) = 2·3a3+ 2·3·4a4x+ 3·4·5a5x2+···+, f(n)(x) =n!an+ (n+ 1)!an+1x+(n+ 2) x!an+2x2+···. By letting x= 0 in each line, we find a0=f(0), a1=f/prime(0), a2=f/prime/prime(0) 2,...,a n=f(n)(0) n!. This proves Theorem 1.21 Iff(x) =/summationtextan(x−x0)nhas a convergent power series representation aboutx0, then the coefficients anare equal to f(n)(x0)/n!, so in fact f(x) =∞/summationdisplay n=0f(n)(x0) n!(x−x0)n. (1-5) This formula (1.21) completely solves the problem of finding the coefficients anof a function if that function has a power series. A simple consequence is the Corollary 1.22 A function f(x)has at most one convergent Taylor series about a point x0. Proof: By the above theorem, if f(x) =/summationtextan(x−x0)nandf(x) =/summationtextbn(x−x0)n, then an=f(n)(x0) n!=bn, so the power series are identical. Remark: Whenfhas a power series expansion about x0, the series is usually called the Taylor series offatx0. In the special case x0= 0 , the series is sometimes called the Maclaurin series forf. 56 CHAPTER 1. INFINITE SERIES Examples: (a) Iff(x) =exhas a power series about x= 0 , what is it? Since f(n)(0) =dn dxnex/vextendsingle/vextendsingle x=0= e0= 1 , we know that an−1/n! so that the power series is/summationtext∞ n=01 n!xn. We cannot yet writeex=/summationtext∞ n=01 n!xnsince we have not proved that exdoes have a power series. (b) Iff(x) = cosxhas a power series about x= 0 , what is it? f(0) = 1, f/prime(0) = −sin 0 = 0, f/prime/prime(0) = −cos 0 = −1, f/prime/prime/prime(0) = sin 0 = 0 , f(4)(0) = cos 0 = 1 ,.... All the odd derivatives at 0 are zero while the even derivatives alternate between +1 and −1 . Therefore the series is a−1 2!x2+1 4!x4−1 6!x6+···∞/summationdisplay n=0(−1)nx2n (2n)!. Again we cannot yet claim that this is cos x. (c) Iff(x) =/braceleftBigg e−1 x2, x/negationslash= 0 0, x= 0/bracerightBigg .has a power series about x= 0 what is it? The computation is somewhat more difficult here. f/prime(x) =2 x3e−1 x2, f/prime/prime(x) = (−6 x4− 4 x6)e−1 x2, and generally f(n)(x) = (α3n x3n+···+α2n−2 xn+2)e−1 x2where the αkare real numbers we don’t need to find. If we let x= 0 inf(n)(x) , the resulting expression has the indeterminate form ∞ ·0 . Thus l’Hˆ ospital’s rule must be invoked. Now f(n)(x) is the sum of terms of the forme−1/x2 xk, k > 0 . What is lim x→0e−1/x2 xk? Let t=1 x2, and we must evaluate lim t→0tk/2e−t= lim t→∞tk/2 et. Ifkis an even integer, applying l’Hˆ ospital’s rule k/2 times leaves a constant in the numerator and etin the denominator, so the limit is lim t→∞const et= 0 . Ifkis odd, applying l’Hˆ ospital’s rule (k+ 1)/2 times leaves a function of the formconst√ tet, which also tends to 0 as t→ ∞ . What we have just shown is that f(n)(0) = 0 . The power series associated with e−1/x2 is 0 + 0·x+0 2!x2+···0 n!xn+··· ≡ 0. This function e−1/x2, whose power series about x= 0 is zero, is an example of a function which is clearly not equal to the power series, 0, associated with it. To find if a given function has a power series expansion about x0we turn to Taylor’s Theorem (also known as the extended mean value theorem). Now if a function fdefined in a neighborhood of x0has a power series expansion there, we know the series is given by (5). Thus we should investigate RN(x)≡f(x)−N/summationdisplay n=0f(n)(x0) n!(x−x0)n. To say that fis equal to its series expansion is the same as saying that the remainder, RN(x) , becomes arbitrarily small as N→ ∞ . We must now seek an estimate of this remainder RN(x) . Taylor’s theorem is one way of finding an estimate. 1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 57 Theorem 1.23 . (Taylor’s Theorem). Let fbe a real-valued function with N+ 1 con- tinuous derivatives defined on an interval containing x0andx. There exists a number ζ betweenx0andxsuch that f(x) =f(x) +f/prime(x0)(x−x0) +f/prime/prime(x0) 2!(x−x0)2+f/prime/prime/prime(x0) 3!(x−x0)3+··· +f(N)(x0) N!(x−x0)N+f(N+ 1)(ζ) (N+ 1)!(x−x0)N+1. (1-6) In other words, RN(x) =f(N+1)(ζ) (N+ 1)!(x−x0)N+1. (1-7) Remark: 1 The proof will only tell us that such a ζexists but will give us no way to find it. In practice we often try to find some upper bound Mforf(N+ 1)(ζ) , so/vextendsingle/vextendsinglef(N+ 1)(ζ)/vextendsingle/vextendsingle≤M, for allN, and only use the crude resulting estimate |RN(x)| ≤M (N+ 1)!|x−x0|N+1. (1-8) An example of this is the series for cos x. Assuming the proof of the theorem, we know that (see Example b above) about x0= 0 , cosx= 1−x2 2!+x4 4!+···+(−1)N (2N)!x2N+RN(x), where RN(x) =1 (2N+ 2)![d2N+2 dx2N+2cosx]x=ζx2N+2, ζ/epsilon1(0,x). Since /vextendsingle/vextendsingle/vextendsingle/vextendsingled2N+2 dx2N+2cosx/vextendsingle/vextendsingle/vextendsingle/vextendsingle x=ζ≤1, we find that |RN(x)| ≤1 (2N+ 2)!|x|2N+2 Because, for fixed x, this remainder tends to 0 as n→ ∞ , we have proved that the power series for cos xatx0= 0 does converge to cos x, so in the limit cosx=∞/summationdisplay n=0(−1)n (2n)!x2n. We can apply Theorem 16 and differentiate both sides of this to find the series for sin x. Remark: 2 Observe that Taylor’s Theorem is only proved for real-valued functionsf. It is not true if fis complex-valued. However using it we will be able to prove the inequality (7) for complex-valued f. Proof: (Taylor’s Theorem). Our proof is short—perhaps a little too slick. The trick is to appeal to the mean value theorem (really only Rolle’s theorem is used). 58 CHAPTER 1. INFINITE SERIES Fixxand define the real number Aby f(x) =N/summationdisplay n=0f(n)(x0) n!(x−x0)n+A(x−x0)N+1 (N+ 1)!. (1-9) Now let H(t) :=f(x)−/bracketleftBig f(t) +f/prime(t)(x−t) +f/prime/prime(t)(x−t)2 2!+···+f(N)(t) N!(x−t)N/bracketrightBig −A(x−t)N+1 (N+ 1)!. Thus we are letting x0vary, notx. Observe that H(x) = 0 (obviously) and H(x0) = 0 (by definition of A). Since H(t) satisfies the hypotheses of the mean value theorem, we conclude that there is some ζbetweenx0andxsuch thatH/prime(ζ) = 0 . But H/prime(t) =−f/prime(t)−/bracketleftbig f/prime/prime(t)(x−t)−f/prime(t)/bracketrightbig − ··· −/bracketleftBigf(N+1)(t) N!(x−t)N−f(N)(t) (N−1)!(x−t)N−1/bracketrightBig −A(x−t)N N!=(x−t)N N!/bracketleftbig A−f(N+1)(t)/bracketrightbig . Amazingly, almost all the terms canceled. Since H/prime(ζ) = 0 and ζ/negationslash=x, we now know thatA=f(N+1)(ζ) . Substitution of this value of Ainto (8) gives us exactly (6), which is just what we wanted to prove. As an application let us prove the Binomial Theorem. That is the name given to the Maclaurin series for (1 + x)α, whereα∈R. The derivatives are easy to compute. f(x) = (1 +x)α f/prime(x) =α(1 +x)α−1 f/prime/prime/prime(x) =α(α−1)(1 +x)α−2 ... f(n)=α(α−1)···.(α−n+ 1)(1 +x)α−n. Thus the power series about 0 associated formally with (1 + x)αis ∞/summationdisplay n=0α(α−1)···(α−n+ 1) n!xn. By the ratio test this series converges for |x|<1 . Does it converge to (1 + x)αwhen |x|<1 ? Ifαis a positive integer, α=N, the terms in the power series from n=N+ 1 on all are zero since they contain the factor ( N−N) . In this case we have only a finite series so convergence is trivial. The resulting polynomial is the familiar Binomial Theorem of high school algebra. Let us therefore assume αis not a positive integer (or 0). Then we have an honest infinite series. In order to prove that (1 + x)αis equal to the infinite series, we must show that the remainder RN(x)≡(a+x)α−N/summationdisplay n=0α(α−1)···(α−n+ 1) n!xn 1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 59 tends to zero as N→ ∞ . By Taylor’s Theorem RN(x) =α(α−1)···(α−N) (N+ 1)!(1 +ζ)α−N−1xN+1, whereζis between 0 and x. We shall prove that this tends to 0 as N→ ∞ only when 0≤x<1 . It is also true for −1<x≤0 , but the proof is much longer so we will not give it [however a different attack yields the proof easily]. Now if 0 ≤x<1 , since 0<ζ <x , then 1<1 +ζ. Therefore for N≥α, we have (z+ζ)α−N−1<1 . Thus |RN(x)|</vextendsingle/vextendsingle/vextendsingle/vextendsingleα(α−1)···(α−N) (N+ 1)!xN+1/vextendsingle/vextendsingle/vextendsingle/vextendsingle which does tend to zero as N→ ∞ (since it is the N+ 1 st term of the convergent series/summationtext∞α(α−1)···(α−n+1) n!xn,|x|<1) . Although we have proved it only if 0 ≤x<1 , we shall state the complete Theorem 1.24 (Binomial Theorem). The function (1 +x)αis equal to a power series which converges for |x|<1. It is (1 +x)α=∞/summationdisplay n=0α(α−1)···(α−n+ 1) n!. (1-10) In practice it is silly to memorize this formula since it is easier to expand (1 + x)α directly in a Maclaurin series, which we have just shown (partly anyway) is equal to the function. We close this section with the generalization of Taylor’s Theorem to complex-valued functionf(x) . Theorem 1.25 . Letf(x) =u(x) +iv(x)be a complex-valued function with N+ 1 continuous derivatives defined on an interval containing x0andx. There exists a real numberMNdepending on Nsuch that /vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglef(x)−N/summationdisplay n=0f(n)(x0) n!(x−x0)n/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤MN (N+ 1)!|x−x0|N+1(1-11) Proof: SincefhasN+ 1 continuous derivatives, so do the real-valued functions u(x) andv(x) . Applying Taylor’s Theorem to uandv, we find numbers ζ1andζ2, both betweenx0andx, such that u(x)−N/summationdisplay n=0u(n)(x0) n!(x−x0)n=u(N+1)(ζ1) (N+ 1)!(x−x0)N+1, and v(x)−N/summationdisplay n=0v(n)(x0) n!(x−x0)n=v(N+1)(ζ2) (N+ 1)!(x−x0)N+1. 60 CHAPTER 1. INFINITE SERIES Thus, by addition, since f(n)=u(n)=iv(n), we find f(x)−N/summationdisplay n=0f(n)(x0) n!(x−x0)n=u(N+1)(ζ1) +iv(N+1)(ζ2) (N+ 1)!(x−x0)N+1. However since u(N+1)andv(N+1)are assumed continuous in an interval containing x0andx, they are bounded there, say by ˆMNand ˜MN. Taking absolute values of the last equation, we obtain equation (10) where MN=/radicalBig ˆM2 N+˜M2 N. Exercises (1) Find the Taylor series about the specified point x0and determine the interval of convergence for the following functions. You need not prove that the series do converge to the functions. (a) sinx, x 0= 0, (b) lnx, x 0= 1, (c)1 x, x0=−1, (d)√x, x 0= 6, (e)1 2(ex+e−x), x0= 0 (f)x+i 1+x, x0= 0, (g) cosx, x 0=π 4, (h)1 i+x, x0= 0 (i)e−x2, x0= 0, (j) (1 +x+x2)−1, x0= 0, (k) cosx+isinx, x 0= 0, (l)1√1+2x, x0= 0. (2) Prove that in their interval of convergence about 0 the following power series associated with the given functions converge to the functions. Do this by proving that the remainder |RN(x)| →0 asN→ ∞ . (a) sinx, (b)1 1+x4, (c)e−x (d) coshx[Recall the definition: cosh x=ex+e−x 2]. (3) One often approximates1√ 1+x2by 1−x2 2when |x|is small. Give some estimate of the error if a) |x|+<10−1, b)|x|<10−2, c)|x|<10−4. (4) Use the Taylor series e−x2= 1−x2+x4 2!−x6 3!+···+(−1)nx2n n!+···. to evaluate/integraltext1 0e−x2dxto three decimal places. I suggest using Theorem 16 and the error estimate of Theorem 12. 1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 61 (5) Assume the ordinary differential equation y/prime−y= 0 , with y(0) = 1 has a power series solution y(x) =/summationtext∞ n=0anxnaboutx= 0 . a). Substitute this series directly into the differential equation and solve for the coefficients an. b). Find when the series converges; c). justify (a posteriori) the fact that the function defined by the convergent series does satisfy the differential equation. [We do not yet know that this is the only solution. All we know is that it is the only solution which has a power series]. (6) In this exercise you will prove that eis irrational. It all hinges on the series for 3 . e= 1 + 1 +1 2+1 3!+···+1 n!+···. (a) Prove that 2 <e< 3 , soeis not an integer (cf. page 58, bottom). (b) Assume eis rational, e=p q, wherepandqare integers with no common factor and q≥2 . Then use the Taylor series with qterms and the remainder Rqto show that e·q! =N+eζ q+1, where 0<ζ < 1 , andNis an integer. (c) From this deduce thateζ q+1must be an integer, and show that this contradicts eζ<e/prime<3 , andq+ 1≥3 . (7) This exercise generalizes the form of the remainder (6’) in Taylor’s Theorem. Fix x and define the number Bby f(x) =N/summationdisplay n=0f(n)(x0) n!(x−x0)n+B(x−x0)α, α≥1. Then consider the function H(t) defined by H(t)≡f(x)−N/summationdisplay n=0f(n)(t) n!(x−t)n−B(x−t)α. Show that there is a ζbetweenx0andxsuch that B=f(N+1)(ζ) αN!(x−ζ)N+1−α, so that RN=f(N+1)(ζ) αN!(x−x0)α(x−ζ)N+1−α. This is Schlomilch’s form of the remainder. In the special case α=N+ 1 , we obtain Lagrange’s form of the remainder, (6) found previously, while for α= 1 we obtain Cauchy’s form of the remainder RN=f(N+1)(ζ) N!(x−x0)(x−ζ)N. Here are two applications of Taylor’s Theorem to problems other than infinite series. The first one deals with max-min. Let f(x) be a sufficiently smooth function (by which we meanfhas plenty of derivatives—we’ll specify the number later). Now we know that 62 CHAPTER 1. INFINITE SERIES iffhas a local maximum or minimum at x0, thenf/prime(x0) = 0 , and it is a maximum if f/prime/prime(x0)<0 , a minimum if f/prime/prime(x0)>0 . But what if f/prime/prime(x0) = 0 ? Consider the examples f1(x) =x4, f2(x) =−x4, f3(x) =x3, the first of which has a minimum at x= 0 , the second a maximum at x= 0 , while the third has neither. These three examples suggest the criterion will depend upon the lowest non-zero derivative being an even or odd derivative, and on its sign. a figure goes here By the definition of local maximum and minimum, the issue is the behavior of f(x) in a neighborhood of x0, that is, the nature of f(x0+h) for|h|small. We remind you that fhas a local max at x0iff(x0+h)−f(x0)≤0 for all |h|sufficiently small, and a local min atx0iff(x0+h)−f(x0)≥0 for all |h|sufficiently small. Since the behavior of f(x) nearx0is determined by the Taylor polynomial f(x0+h) =f(x0) +f/prime(x0)h+f/prime/prime(x0)h2 2!+···+f(n) n!(x0)hn+f(n+1)(ζ)hn+1 (n+ 1)! whereζis between x0andx0+h, it is natural to look at this polynomial to answer our question. Theorem 1.26 Assumefhas (at least) n+ 1 continuous derivatives in some interval containing x0. Sayf/prime(x0) =f/prime/prime(x0) =...=fn(x0) = 0 butf(n+1)(x0)/negationslash= 0, then (a) ifnis even, then fhas neither a max nor min at x0. (b) ifnis odd, then i)fhas a max at x0iff(n+1)(x0)<0. ii)fhas a min at x0iff(n+1)(x0)>0. Proof: We shall use Taylor’s polynomial with n+ 1 terms. Since the first nderivatives vanish at x0, we havef(x0+h)−f(x0) =f(n+1)(ζ) (n+1)!hn+1, ζ betweenx0andx0+h. Becausef(n+1)(x) is assumed continuous at x0, f(n+1)(ζ) must have the same sign as f(n+1)(x0) in some neighborhood of x0. Restrict your attention to the neighborhood. Ifnis even ,n+ 1 is odd, so that hn+1is positive if h>0 , negative ifh<0 . Thusf(x0+h)−f(x0) changes sign in any neighborhood of x0. However ifn is odd ,hn+1is positive no matter if h>0 orh<0 . Therefore f(x0+h)−f(x0) has the same sign as f(n+1)(x0) throughout some neighborhood about x0. The precise conditions are easy to verify now. Examples: 1.f(x) =x5+ 1 has neither a max nor min at x= 0 , since f/prime(0) =...=f(4)(0) = 0 , butf(5)(0) = 5! /negationslash= 0 . 2.f(x) = (x−1)6−7 has a min at x= 1 since f/prime(1) =...=f(5)(1) = 0 , but f(6)(1) = 6!>0 . Our second application is a geometrical interpretation of the Taylor polynomial. Given the function f(x) , consider the polynomial Pn(x) =f(x0) +f/prime(x0)(x−x0) +f/prime/prime(x0) 2!(x−x0) +···+f(n)(x0) n!(x−x0)n, 1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 63 whose first nderivatives agree with those of fatx=x0.P1(x) =f(x0) +f/prime(x0)(x−x0) is the equation of the tangent to the curve y=f(x) atx0. It is the straight line which most closely approximates the curve at x0. Similarly P2(x) is the parabola which most closely approximates the curve at x0. Generally, Pn(x) is the polynomial of degree n which most closely approximates the curve y=f(x) at the point x0. Using this Taylor polynomial, we can define the order of contact of two curves at a point. Definition: The two curves y=f(x) andy=g(x) have order of contact nat the point x0if their Taylor polynomials of degree natx0are identical, but their n+ 1 st Taylor polynomials differ. An equivalent definition is that f(x0) =g(x0) ,f/prime(x0) =g/prime(x0) , . . . ,f(n)(x0) = g(n)(x0) , butf(n+1)(x0)/negationslash=g(n+1)(x0) . We have assumed that fandghaven+ 1 continuous derivatives. If fandghave contact natx0, then f(x0+h)−g(x0+h) =f(n+1)(ζ1)−g(n+1)(ζ2) (n+ 1)!hn+1. One interesting consequence of this formula is that if fandghave contact of even order, then the curves will cross at x0, while if the contact is of odd order, the curves will notcross in some neighborhood of x0. We can define the curvature of a curve in the plane by using the concept of contact. First we define the curvature of a circle (whose curvature had better be constant). Definition: Thecurvaturekof acircle of radiusRis defined to be1 R, k=1 R. Thus the smaller the circle, the larger the curvature—a natural outcome. Furthermore, a straight line—which may be thought of as a circle with infinite radius—has curvature zero. How can we define the curvature of a given curve? For all non-circles, the curvature will clearly vary from point to point of the curve. Thus, the concept we want is the curvature of a given curve y=f(x)at a pointx0. Our definition should appear reasonable. Definition: Thecurvaturekof a plane curve y=f(x) at the point x0is the curvature of the circle which has contact of order two at x0. This circle which has contact of order two is called the osculating circle to the curve atx0(osculate: Latin, to kiss). Let us convince ourselves that there is only one osculating circle (for if there were two, the curvature would not b e well defined.) Consider all circles of contact one to f(x) atx0. These are all circles tangent to f(x) atx0. Their centers lie on the line lnormal to the curve at x0(“normal” means perpendicular to the tangent line). It is geometrically clear that of these circles with contact 1, there will be exactly one with contact 2. Example: Find the curvature of y=exatx= 0 . The slope of the curve at (0 ,1) is 1. Therefore the equation of the normal is y−1 =−x. Since the center ( x0,y0) of the osculating circle must lie on this line, and the circle contains the point (0 ,1) , subject to y0= 1−x0, the value of x0must be determined from the fact that the second derivative of the circle (0 ,1) must equal the second derivative of y=exatx= 0 , that is, it must equal 1. But for any circle, ( y−y0)y/prime/prime+y/prime2+ 1 = 0 . In our case y/prime= 1 at (0,1) (recall the circle is tangent to exat (0,1) ), so that (1 −y0)·1 + 1 + 1 = 0 , or y0= 3 . The equationy0= 1−x0implies that x0=−2 . Thus the equation of the osculating circle is (y−3)2+ (x+ 2)2= 8 , and the curvature of y=exatx= 0 isk=1√ 8. Later on we will give another definition of curvature which is applicable not only to plane curves, but also to curves in space. 64 CHAPTER 1. INFINITE SERIES Exercises (1) What is the order of contact of the curves y=e−xandy=1 1+x+1 2sin2xatx= 0 ? (2) Find the osculating circle and curvature for the curve y=x2atx= 1 . (3) Show that at x=a, the curve y=f(x) has curvature k=f/prime/prime(a) [1+f/prime(a)2]3 2and the center of the osculating circle is at the point ( a−f/prime(a) f/prime/prime(a)[1 +f/prime(a)2], f(a) +1+f/prime(a)2 f/prime/prime(a)).What is the messy equation of the osculating circle? (4) At the given points, the following curves have slope zero. Determine if the curve has a max, min, or neither there. (a).y= (x+ 1)4, x=−1, (b).y=x2sinx, x= 0. (5) LetP1,P, andP2be three distinct points on the curve y=f(x) , and consider the circle passing through those three points. Show that in the limit as both P1andP2 approachP, this circle becomes the osculating circle. (Hint: Taylor’s Theorem will be needed here). (6) In this problem we outline another derivation of Taylor’s Theorem. Whereas the one in the notes did not use the fact the f(n+1)was continuous, this proof relies upon that fact. (a) Show that /integraldisplayx x0(x−t)k−1 (k−1)!f(k)(t)dt=f(k)x0(x−x0)k k!+/integraldisplayx x0(x−t)k k!f(k+1)(t)dt. (b) Prove by induction that f(x) =f(x0)+f/prime(x0)(x−x0)+···+f(n)(x0) n!(x−x0)n+/integraldisplayx x0(x−t)n n!f(n+1)(t)dt. The remainder is expressed as an integral here. It is because f(n+1)is to be integrated that we require its continuity. (7) (a) Let g(x) have contact of order nwith the function 0 at the point x=a, and assume that f(x) has contact of order at least nwith the function 0 at x=a. Use Taylor’s Theorem to prove that lim x→af(x) g(x)=f(n+1)(a) g(n+1)(a) This is l’Hˆ ospital’s Rule . (b) Apply l’Hˆ ospital’s rule to evaluate i) lim x→0x−sinx x3,ii) lim θ→π 41−tanθ θ−π v 1.6. COMPLEX-VALUED FUNCTIONS, EZ,COSZ,SINZ. 65 (8) Assume fhas two derivatives in the interval [ a,b] , and assume that f/prime/prime≥0 through- out the interval. Prove that if ζis any point in [ a,b] , then the curve y=f(x) never falls below its tangent at the point x=ζ, y=f(ζ) . [hint: Use Taylor’s Theorem with three terms]. (9) Use Cauchy’s form of the remainder (p. 103-4, no. 7) for Taylor’s Theorem to prove that the binomial series converges to (1 + x)αfor−1< x≤0 . This will complete the proof of the binomial theorem. (10) ThenthLegendre polynomial Pn(x) is defined by Pn(x) =1 2nn!dn dxn[(x2−1)n] . Prove thatPn(x) is a polynomial of degree nand hasndistinct real zeros in the interval (−1,1) . (11) Verify that eaxis a solution of y/prime=ay. Prove that every solution has the form Aeax, whereAis a constant. (12) Assume that f(x) has plenty of derivatives in the interval [ a,b] , and that fhas n+ 1 distinct zeros in the interval. Prove that there is at least one c∈(a,b) such thatf(n)(c) = 0 . 1.6 Complex-Valued Functions, ez,cosz,sinz. The task of this section is to answer the following question. Say f(x) is a real or complex valued function of the realvariablex. How can we define f(z) whereziscomplex ? For example, if P(x) =a0+a1x+···+anxnis a polynomial, the answer is easily given: just define P(z) =a0+a1z+···+anzn. Since this function only involves addition and multiplication of complex numbers, for any complex zthe number P(z) can be computed. Similarly any rational function,P(x) Q(x), whereP(x) andQ(x) are both polynomials, can be defined for complex zasP(z) Q(z)since both P(z) andQ(z) are defined separately and we can then take their quotient. But how do we define ez, or cosz, or (1 +z)α, whereα/epsilon1Ris not a positive integer? As might have been suspected, the trick is to use infinite series. Definition: Iff(x), x/epsilon1R, has a convergent Taylor series, f(x) =∞/summationdisplay n=0anxn, |x|<ρ, then we define f(z), z/epsilon1C, by the infinite series f(z) =∞/summationdisplay n=0anzn, and the infinite series converges throughout the disc |z|<ρ. The assertion that the complex series converges throughout the disc |z|< ρ is an immediate consequence of Theorem 13 on page ?. Thus, for example, we define . E(z) =∞/summationdisplay n=01 n!zn, 66 CHAPTER 1. INFINITE SERIES C(z) =∞/summationdisplay n=0(−1)nz2n (2n)!, S(z) =∞/summationdisplay n=0(−1)nz2n+1 (2n+ 1)!, and (1 +z)α=∞/summationdisplay n=0α(α−1)···(α−n+ 1) n!zn, α/epsilon1 R where the first three series converge for all z/epsilon1C, while the last converge for |z|<1 . We have temporarily used the notation E(z) in place of ez, C(z) for cosz, andS(z) for sinz so that you do not jump to hasty conclusion s about these functions by merely extrapolating your knowledge of exetc. For example it is nottrue that |sinz| ≤1 for allz/epsilon1C, even though |sinx| ≤1 for allx/epsilon1R. All properties of these function s for z/epsilon1Cmust be proved again beginning with the power series definitions. Known properties of ex, x/epsilon1Rand wishful thinking don’t prove properties of ez, z/epsilon1C. Let us begin by proving Theorem 1.27 . (a)E(iz) =C(z) +iS(z),for allz∈C. (b)E(−iz) =C(z)−iS(z), for allz∈C. (c)C(z) =1 2[E(iz) +E(−iz)], for allz∈C. (d)S(z) =1 2i[E(iz)−E(−iz)], for allz∈C. Proof: a). b). Just substitute and rearrange the series. For example C(z) = 1−z2 2!+x4 4!−x6 6!+···. iS(z) =i[z−z3 3!+z5 5!−z7 7!+···] so C(z) +iS(z) = 1 +iz−z2 2!−iz3 3!+z4 4!+iz5 5!− ···, where the adding of the two series is justified by Theorem 5(page ?). We must compare the last series with that for E(iz) : E(iz) = 1 +iz+(iz)2 2!+(iz)3 3!+(iz)4 4!+···= 1 +iz−z2 2!−iz3 3!+z4 4!+···, which is identical to the series for C(z) +iS(z) . c)-d). These follow by elementary algebra from a) and b). The formulas a)-d) of Theorem 21 show there is a close connection between the four functionsE(iz), E(−iz), C(z),andS(z) . Our next theorem shows that the formula exey=ex+y, x, y∈R, extends to the function E(z) . Theorem 1.28 .E(z)E(w) =E(z+w), for allz, w∈C. 1.6. COMPLEX-VALUED FUNCTIONS, EZ,COSZ,SINZ. 67 Proof: We must show that (∞/summationdisplay n=0zn n!)(∞/summationdisplay n=0wn n!) =∞/summationdisplay n=0(z+w)n n!, The product of the two series is defined in Theorem 15. Using that definition, we find that (∞/summationdisplay n=0zn n!)(∞/summationdisplay n=0wn n!) =∞/summationdisplay n=0(n/summationdisplay k=0zk k!wn−k (n−k)!). However, the binomial theorem for positive integer exponents (which only uses the algebraic rules for complex numbers) states that (z+w)n=n/summationdisplay k=0n! k!(n−k)!zkwn−k. Upon substituting this into the last equation, we obtain the desired formula. The formula of this theorem is the key to many results, like the following generalization of sin2x+ cos2x= 1 . Corollary 1.29 C(z)2+S(z)2= 1 for allz∈C. Proof: We use equations a) and b) of Theorem 21 to reduce the question to one of exponentials. E(iz)E(−iz) = [C(z) +iS(z)][C(z)−iS(z)] =C2(z) +S2(z). But by Theorem 22, E(iz)E(−iz) =E(iz−iz) =E(0) . Directly from the power series we see thatE(0) = 1 . This proves the formula. Our next corollary states that the addition formulas for sin xand cosxare still valid forC(z) andS(z) . Corollary 1.30 C(z+w) =C(z)C(w)−S(z)S(w)andS(z+w) =S(z)C(w)−S(w)C(z) for allz, w∈C Proof: A direct algebraic computation does the job. C(z+w) +iS(z+w) =E(iz+iw) =E(iz)E(iw) = [C(z) +iS(z)][C(w) +iS(w)] = [C(z)C(w)−S(z)S(w)] +i[S(z)C(w) +S(w)C(z)]. Similarly we find that C(z+w)−iS(z+w) = [C(z)C(w)−S(z)S(w)]−i[S(z)C(w) +S(w)C(z)]. Addition of these two equations gives the formula for C(z+w) , while subtraction gives the formula for S(z+w) . Had we but world enough, and time, we would linger a while. A lovely result we have not proved is that E(z+ 2πi) =E(z) , the periodicity of E(z) , which is a consequence of 68 CHAPTER 1. INFINITE SERIES the formulas C(z+ 2π) =C(z) , andS(z+ 2π) =S(z) , the periodicity of C(z) andS(z) , by using Theorem 21 (but see pp. ??). We shall close this chapter by restating the results proved above in the usual language ofezetc. instead of the temporary notation E(z) etc. we have been using. eiz= cosz+isinz (1-12) e−iz= cosz−isinz (1-13) cosz=1 2(eiz+e−iz) (1-14) sinz=1 2i(eiz−e−iz) (1-15) ezew=ez+w(1-16) sin2z+ cos2z= 1 (1-17) cos(z+w) = coszcosw−sinzsinw (1-18) sin(z+w) = sinzcosw+ sinwcosz (1-19) Generally, all algebraic formulas for sin x,cosx, andexremain valid for sin z,cosz, and ez. In fact any algebraic relationship between any combination of analytic functions remains valid as we change the in dependent variable from a real xto the complex z. Inequalities almost always fall apart in the transition from x∈Rtoz∈C. Exercise 2e below illustrates this. One formula which we will use frequently later on is a specialization of (1-12) to the case when zis real. Then writing the real zasθwe have the famous formula eiθ= cosθ+isinθ, θ∈R. (1-20) We cannot resist stating this formula down again for θ=π: eiπ=−1, an almost mystical identity connecting the four numbers e, iπ , and −1 . Notice that (1.6) also implies/vextendsingle/vextendsingleeiθ/vextendsingle/vextendsingle= 1 . If we write z=x+iy, then using (1.6) we find ez=ex+iy=exeiy=ex(cosy+isiny). (1-21) A consequence of this is |ez|=ex(1-22) Exercises (1) Observe that (directly from the power series) cos(−z) = cosz,and sin( −z) =−sinz. Use this and the addition formula for cos( z+w) to prove that sin2z+ cos2z= 1 . (2) If we define sin hx=1 2(ex−e−x) and coshx=1 2(ex+e−x), x∈R, we prove that 1.6. COMPLEX-VALUED FUNCTIONS, EZ,COSZ,SINZ. 69 (a) cosix= coshx,sinix=isinhx (b) cosz= coshy−isinxsinhy,(z=x+iy) sinz= sinxcoshy+icosxsinhy (c)|cosz|2= cos2x+ sinh2y |cosz|2= cosh2y−sin2x |sinz|2= sin2x+ sinh2y |sinz|2= cosh2y−cos2x (d) Use the identities of part c) to deduce that |sinhy| ≤ |cosz| ≤coshy |sinhy| ≤ |sinz| ≤coshy (e) Prove that there is some z∈Csuch that |sinz|>1,and|cosz|>1. (3) Define the derivative of f(z) atz0, wherez, z 0∈C, as lim z→z0f(z)−f(z0) z−z0, if the limit exists. (a) By working directly with the power series, show that ezis differentiable for all z, and that d dzeaz=aeaz, a, z∈C, (b) Apply this to (1-12) and (1-13) to deduce that d dzcosz=−sinz,d dzsinz= cosz (We cannot appeal to Theorem 16 and differentiate term-by-term since that theorem assumed the independent variable, x, was real). (4) Use the results of Exercise 2c to show that the only complex roots z=x+iyof sinz and coszare at the points on the real axis y= 0 where sin x= 0 and cos x= 0 , respectively. (5) Use the results of this section to prove DeMoirve’s Theorem (cosθ+isinθ)n= cosnθ+isinnθ, θ ∈R, wherenis a positive integer. (6) (a) Show that the sum of the finite geometric series/summationtexteinxis N/summationdisplay n=1einx=ei(N+1/2)x−eix/2 eix/2−e−ix/2. 70 CHAPTER 1. INFINITE SERIES (b) Take the real and imaginary parts of the above formula and prove that for all x/negationslash= 0, x∈(0,2π) , N/summationdisplay n=1cosx=sin(N+ 1/2)x−sin 1/2x 2 sin1 2x N/summationdisplay n=1sinnx=cos1 2x−cos(N+1 2)x 2 sin1 2. 1.7 Appendix to Chapter 1, Section 7. As a special dessert let us take some time out and prove some interesting results you would probably never see otherwise. We have in mind to define a specific number α∈Ras the smallest positive zero of cos x, x∈R—soαhad better turn out as π/2 . Then we prove that 1) sin( x+ 4a) = sinxetc., 2) the ratio of the circumference to diameter of a circle is 2αso that 2αdoes equal the πof public school fame. Furthermore, we also present a way of computing α. In this section we take sin zand cosz, z∈Cto be defined by their power series, and use only the properties of these functions which were obtained from the power series definition. Lemma 1.31 The setA={x∈R: cosx= 0,0< x < 2}is not empty, that is, the equation cosx= 0 has at least one real root for x∈(0,2). Proof: Since cosxis defined by a convergent power series, it is continuous (even infinitely differentiable); furthermore because x∈Rand the power series has real coefficients, we know that cos x, x∈Ris real-valued. Observe that cos 0 = 1 >0 , and the following crude inequality cos 2 = 1 −22 1·2+∞/summationdisplay n=2(−1)n22n (2n)!<−1 +∞/summationdisplay n=222n (2n)! <−1 +24 4!∞/summationdisplay k=0(2 5)2k=−1 +50 63<0.(1-23) Thus cos 0 >0 and cos 2 <0 , so there is at least one point in (0 ,2) where the real-valued continuous function cos xvanishes. This proves the lemma. Denote the g.l.b of A(which does exist since Ais bounded—say by 0 and 2) by α. We shall show that α∈A. Sinceαis the g.l.b. of A, there exists a sequence of points αk∈A(theαkmay just be the same point repeated over and over) such that αk→α and cosαk= 0 . But since cos xis continuous, 0 = lim k→∞cosαk= cosα, so in fact cos α= 0 too ⇒α∈A. Now cosxmust be positive throughout the interval [0 ,α) , since it is positive at x= 0 andαis the first place it vanishes. Therefore the formulad dxsinx= cosx—obtained by differentiating the realpower series for sin xterm by term—shows that sin xis increasing 1.7. APPENDIX TO CHAPTER 1, SECTION 7. 71 forx∈[0,α) . Since sin 0 = 0 , we see that sin x≥0forx∈[0,α) . Thus the formula d dxcosx=−sinxtells us that cos xis decreasing in the interval [0,α] . From the formula 1 = sin2α+ cos2α= sin2α, and the fact that sin α > 0 , we find that sin α= 1 . We can thus conclude from the addition formulas for sin xand cosxthe: Theorem 1.32 Letαdenote the smallest zero of cosxforx>0. Then cosα= 0,cos 2α=−1,cos 3α= 0,cos 4α= 1 sinα= 1,sin 2α= 0,sin 3α=−1,sin 4α= 0, or more generally cos(z+α) =−sinz,sin(z+α) = cosz cos(z+ 4α) = cosz,sin(z+ 4α) = sinz This proves that the sinzand coszare periodic with period 4α. As you have guessed, αis another name for π/2 —and serves as our definition of π. This is based upon power series and is independent of circles or triangles—or even the entire concept of angle. A simple consequence is the Corollary 1.33 The function ezis periodic with period 4αi, ez+4αi=eze4αi=ez. Proof:ez+4α=eze4iα=ez(cos 4α+isin 4α) +ez(1 +i0) =ez. Two issues remain to be settled before closing up. We should 1) prove that the ratio of the circumference Cof a circle to its diameter Disπ, i.e.,C= 2αD, and 2) find some way of approximating αnumerically (for all we know of alpha so far is that it is the smallest element in a set and 0 <α< 2 ). The two problems are closely related. The circle of radius Rhas the equation x2+y2=R2. Consider the portion in the first quadrant. Then using the familiar formulas for arc length, we find that C 4=R/integraldisplayR 0dx√ R2−x2=R/integraldisplay1 0dt√ 1−t2, where the change of variable x=Rthas been used to obtain the last integral [this is legal since the mapping “multiply by R” is a bijection and hence an invertible function]. Thus, the desired result, C= 2αD= 4αRwill be proved if we can p rove Theorem 1.34/integraltext1 0dt√ 1−t2=α(=π 2) Corollary 1.35 IfCdenotes the arc length of the circumference of a circle of radius R, thenC= 4αR. 72 CHAPTER 1. INFINITE SERIES Proof: of Theorem. We want to make the change of variable t= sinζ, wheret∈[0,1] . In order to do this we must only check that the function sin ζis differentiable and invertible function there. We know it is differentiable . Since sin xis continuous and monotone increasing for x∈[0,α] , and since the end points are mapped into 0 and 1 respectively ( sin 0 = 0,sinα= 1 ), the function f(ζ) = sinζis invertible for x∈[0,α]⇐⇒t∈[0,1] . The usual formulas are applicable and yield /integraldisplay1 01√ 1−t2dt=/integraldisplayα 0dζ=α Q.E.D. To compute π= 2α, it is convenient to introduce tan z= sinz/cosz, for allzwhere cosz/negationslash= 0 . In particular tan xis defined for all real xin the interval 0 ≤x<α/ 2 . From the behavior of sin xand cosxin the interval x∈[0,α/2) , it is easy to show that tan x has infinitely many derivatives and is increasing for x∈[0,α/2) , assuming the values from 0 = tan 0 to 1 = tanα 2. The function tan xis therefore invertible in that interval, so we can make the natural change of variable t= tanxand obtain /integraldisplay1 0dt 1 +t2=/integraldisplayα/2 01 1 + tan2x(d dxtanx)dx=/integraldisplayα/2 0dx=α 2. But the integral on the left can be approximated readily because of the algebraic identity 1 1 +t2=N/summationdisplay 0(−1)nt2n+(−1)N+1t2N+2 1 +t2,allt/negationslash=i. Thus π 4=α 2=/integraldisplay1 0dt 1 +t2=N/summationdisplay 0(−1)n/integraldisplay1 0t2ndt+ (−1)N+1/integraldisplay1 0t2N+2 1 +t2dt, or π 4= 1−1 3+1 5−1 7+···+(−1)N 2N+ 1+RN, where since 2 t≤1 +t2the remainder RNcan be estimated by |RN|=/integraldisplay1 0t2N+2 1 +t2dt</integraldisplay1 0t2N+2 2tdt=1 4N+ 4 If the first 250 terms in the series are used, N= 250 , we find π 4= 1−1 3+1 5− ··· +1 251+R250, where |R250|<1 1004<1 1000, so three decimal accuracy is obtained. This is quite slow—but it does work. For practical computations, a series which converges much faster is needed. See exercise 2 below; it is neat. SinceRN→0 asN→ ∞ , the following formula is a consequence of our effort: π 4= 1−1 3+1 5−1 7+1 9− ···.. Exercises 1.7. APPENDIX TO CHAPTER 1, SECTION 7. 73 (1) Use the method illustrated here to slow that ln 2 =/integraldisplay1 01 1 +xdx= 1−1 2+1 3−1 4+1 5− ··· +(−1)N+1 N+RN, where lim N→∞RN= 0 . Find an Nsuch that |RN|<10−3. [Hint: Write1 1+x=/summationtextN 0(−1)nxn+(−1)N+1xN+1 1+x, x/negationslash=−1]. (2) To approximateπ 4with fewer terms, the following clever device works. Write 1 1 +t2=N−1/summationdisplay 0(−1)nt2n+(−1)Nt2N 2+ ((−1)Nt2N 2+(−1)N+1t2N+2 1 +t2) and show that π 4= 1−1 3+1 5+···+(−1)N−1 2N−1+(−1)N 2(2N−1)+˜RN, where ˜RN+(−1)N 2/integraltext1 0t2N−t2N+2 1+t2dt. (a) Prove that/vextendsingle/vextendsingle/vextendsingle˜RN/vextendsingle/vextendsingle/vextendsingle<1 8N2+8N. (b) What should Nbe to make/vextendsingle/vextendsingle/vextendsingle˜RN/vextendsingle/vextendsingle/vextendsingle<10−3? Amazing saving, isn’t it? The technique does generalize to other series and can be refined to yield even better results. (c) Apply the method given here to problem 1 above to show that ln 2 = 1 −1 2+ 1 3+···+(−1)N N−1+1 2(−1)N+1 N+˜RN, where/vextendsingle/vextendsingle/vextendsingle˜RN/vextendsingle/vextendsingle/vextendsingle<1 (2N+1)(2 N+3). PickNso that/vextendsingle/vextendsingle/vextendsingle˜RN/vextendsingle/vextendsingle/vextendsingle<10−3. 74 CHAPTER 1. INFINITE SERIES Chapter 2 Linear Vector Spaces: Algebraic Structure 2.1 Examples and Definition In order to develop intuition for linear vector spaces, a slew of standard examples are needed. From them we shall abstract the needed properties which will then be stated as a set of axioms. a) The Space R2. We begin by informally examining a space of two dimensions (whatever that means). It is constructed by taking the Cartesian Product of Rwith itself. We are thus looking at R×R, which is denoted by R2. A pointXin this space is an ordered pair, X= (x1,x2) , where x1∈R, x2∈R.x1andx2are called the coordinates orcomponents of the point x. Let us propose a reasonable algebraic structure on R×R. IfX= (x1, x2) , andY= (y1, y2) are any two points, and αis any real number, we define addition:X+Y= (x1+y1,x2+y2) . multiplication by scalars: α·X= (αx1,αx 2), α∈R. equality:X=Y⇐⇒x1=y1, x2=y2 The addition formula states that the parallelogram rule is used to add points, whereas the second formula states that a point Xis “stretched” by αby stretching each coordinate byα. Some immediate consequences of the above definitions are, for all X, Y, Z inR×R, (1) addition is associative (X+Y) +Z=X+ (Y+Z) (2) addition is commutative X+Y=Y+X (3) There is an additive identity , 0=(0,0) with the property that X+ 0 =Xfor anyX. (4) Every X= (x1,x2)∈R×Rhas an additive inverse (−x1,−x2) , which we denote by−X. ThusX+ (−X) = 0 . Thus the set of points in R×Rforms an additive abelian group. The following additional properties are also obvious, where αandβare arbitrary real numbers. 75 76 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE (5)α(βX) = (αβ)X (6) 1 ·X=X. and the two distributive laws. (7) (α+β)X=αX+βX (8)α(X+Y) =αX+αY. To insure that you too feel these properties are obvious, let us prove, one, say 7. (α+β)·X= (α+β)·(x1,x2) = ((α+β)x1,(α+β)x2) = (αx1+βx1,αx 2+βx2) = (αx1,αx 2) + (βx1,βx 2) =α·(x1,x2) +β·(x1,x2) =α·X+β·X(2-1) Example: IfX= (2,1) , then 3X= (6,3) and −2X= (−4,−2) . Instead of thinking of the elements ( x1,x2) inR2as points, it is sometimes useful to think of them as directed line segments, from the origin (0,0) directed to the point ( x1,x2) . The figure at the right illustrates this. Note that the axes need not be perpendicular to each other in the space R2. They could just as well veer off at some outrageous angle, as in the diagram. This is because we have yet to place a metric (distance) structure on R2or introduce any concept of angle measurement. When we do that, we will have Euclidean 2-space E2. But right now all we have is R2, which might be thought of as a floppy Euclidean space. b) The Space Rn . This is a simple-minded generalization of R2. A pointXinRn=R×...×Ris an orderedntuple,X= (x1,x2,...,x n) of real numbers, xk∈R. IfX= (x1,...,x n) and Y= (y1,...,y n) are any two points in Rn, andαis any real number, we define addition :λ+Y= (x1+y1,x2+y2,...,x n+yn) multiplication by scalars :α·X= (αx1,αx 2,...,αx n), α∈R. equality :X=Y⇐⇒xj=yjfor allj. Example: The pointX= (1,2,3) , and1 2X= (1 2,1,3 2) inR3are indicated in the figure. Again the coordinate axes need not be mutually perpendicular. Properties 1-8 listed earlier remain valid - and with the proofs essentially unchanged (just add dots inside the parentheses). Remark: . At this stage, you probably are anxiously waiting for us to define multiplication inRn, that is, the product of two points in Rn,X·Y=Z∈Rn, possibly using the multiplication of complex numbers (points in R2) as a guide. Well, we would if we could. It turns out that it is possible to define such a multiplication only inR1,R2,R4, and in R8–but in no others . This is a famous theorem. In R2ordinary complex multiplication does the job. To do it in R4, we have to abandon the commutative law for multiplication. The result is called quaternions. InR8, the multiplication is neither commutative nor associative. The result there is the Cayley numbers . Here we shall not have time to treat this issue. All we shall do (later) is introduce a “pseudo multiplication” in R3—the so called cross product - obtained from the quaternion 2.1. EXAMPLES AND DEFINITION 77 algebra in R4. The major importance of this pseudo multiplication which holds only in R3 is the fact of life that our world has three space dimensions. This multiplication is extremely valuable in physics. c) The Space C[a, b]. Our next example is of an entirely different nature, it is a space of functions, a function space . The space C[a,b] is the set of all real-valued functions of a real variable xwhich are continuous for x∈[a,b] . Iffandgare continuous for x∈[a,b] , that is if fand g∈C[a,b] , and ifαis any real number, we define, in the usual way, addition: ( f+g)(x) =f(x) +g(x) , multiplication by scalars: ( αf)(x) =α[f(x)]. α∈R equality:f=g⇐⇒f(x) =g(x) for allx∈[a,b]. Notice that the sum of two functions in C[a,b] is again in C[a,b] , and the product of a continuous function - in C[a,b] —by a constant αis also an element of C[a,b] . We shall ignore the fact that the product of two continuous functions is also a continuous function. Properties 1-8 listed earlier are also valid here, that is, if f,g, andhare any elements inC[a,b] , then (1)f+ (g+h) = (f+g) +h (2)f+g=g+f (3)f+ 0 =f (4)f+ (−1)f= 0 (5)α(βf) = (αβ)f (6) (1)f=f 1∈R (7) (α+β)f=αf+βf (8)α(f+g) =αf+αg. Again, 1-4 state that the elements of C[a,b] form an abelian group with the group operation being addition. When we define the dimension of a vector space, it will turn out that the space C[a,b] isinfinite dimensional, but don’t let that bother you. This nice space,C[a,b] , and Rnare the two most useful examples of a vector space. d) D. The Space Ck[a, b]. The space Ck[a,b] consists of all real-valued functions f(x) which have kcontinuous derivatives for xin the interval [ a,b]⊂R. Whenk= 0 , this reduces to the space C[a,b] . Addition and scalar multiplication are defined just as in C[a,b] . The key property is that the sum of two functions with kcontinuous derivatives of x∈[a,b] is also a function with kcontinuous derivatives. All of properties 1-8 are valid in Ck[a,b] . Every function f(x) which has one continuous derivative is necessarily continuous. This is a basic result from elementary calculus; it may be written as C1[a,b]⊂C[a,b] . Since the function |x|, x∈[−1,1] is inC[−1,1] but not in C1[−1,1] , we see that C1and Care not the same, that is C1is a proper subset of C. Similarly, Ck+1[a,b]⊂Ck[a,b] (see Exercise 7). 78 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE The space C∞[a,b] consists of all functions with an infinite number of continuous derivatives for x∈[a,b] . All functions which have a convergent Taylor series for x∈[a,b] are inC∞[a,b] . In addition, C∞[a,b] contains functions like f(x) =e−1/x2, x/negationslash= 0, f(0) = 0 , which have an infinite number of continuous derivatives (see p. ??) but do not have convergent Taylor series. Another example of a function space is the set of analytic functions A(z0,R) , functions which have a convergent Taylor series in the disc with center at z0∈Cand radius at least R. e) E. The Space l1. The spacel1(tired yet?) consists of all infinite sequences X= (x1,x2,x3,...) which satisfy the condition∞/summationdisplay n=1|xn|<∞. Addition and multiplication by scalars are defined in a natural way. IfXandYare inl1, then X+Y+ (x1+y1,x2+y2,x3+y3,...) and, ifαis any complex number α·X= (αx1,αx2,...). Equality is defined by X+Y⇐⇒xj=yjfor allj. We should show that if XandYare inl1, then so is X+Yandx·X. To prove that X+Y∈l1, we must show that/summationtext|xn+yn|<∞. But since |xn+yn| ≤ |xn|+|yn|, we have for any N∈Z+ N/summationdisplay n=1|xn+yn| ≤N/summationdisplay n=1|xn|+N/summationdisplay n=1|yn| ≤∞/summationdisplay n=1|xn|+∞/summationdisplay n=1|yn|<∞. Now letting N→ ∞ on the left, we see that∞/summationdisplay n=1|xn+yn|<∞. IfX∈l1, it is obvious thatα·Xis also inl1since ∞/summationdisplay n=1|αxn|=∞/summationdisplay n=1|α||xn|=|α|∞/summationdisplay n=1|xn|<∞. f) F. The Space L1[a, b]. Yes, the space L1[a,b] does consist of all functions f(x) (possibly complex-valued) with the property that/integraltextb a|f(x)|dx<∞. It is the integral analogue of l1. Addition and scalar multiplication are defined as in C[a,b] , that is, as usual. If fandgare inL1[a,b] , then so aref+gandαf, whereα∈C, since /integraldisplayb a|f(x) +g(x)|dx≤/integraldisplayb a|f(x)|dx+/integraldisplayb a|g(x)|dx<∞, 2.1. EXAMPLES AND DEFINITION 79 and /integraldisplayb a|αf(x)|dx=|α|/integraldisplayb a|f(x)|dx<∞. For example, f(x) =xis inL1[0,1] butf(x) =1 x2isnotinL1[0,1] . It is simple to check that properties 1-8 are satisfied in L1[a,b] . g) G. The Space fn. IfP(x) =a0+a1x+...+anxnis any polynomial of degree nwith real coefficients and Q(x) =b0+b1x+...+bnxnis another one, then with ordinary addition, multiplication by real scalars and equality the set fnof all polynomials of degree nsatisfy conditions 1-8. Since a0+a1x+...+an−1xn−1=a0+a1x+...+an−1xn−1+ 0xn, it is clear that fn−1⊂fn. Enough examples for now. You must have gotten the point. We shall meet more later on. Let us give the abstract definition of a linear vector space. Definition: . LetSbe a set with elements X, Y, Z,... andFbe a field with elements α,β... . The setSis alinear vector space (linear space, vector space )over the field Fif the following conditions are satisfied. For any two elements X, Y∈S, there is a unique third element X+Y∈S, such that (1) (X+Y) +Z=X+ (Y+Z); (2)X+Y=Y+X; (3) There exists an element 0 ∈Shaving the property that 0 + X=Xfor allX∈S; (4) for every X∈S, there is an element −X∈S; such that X+ (−X) = 0 . Furthermore, if αis any element of the field F, there is a unique element αX∈S such that, for any α,β∈F, (5)α(βX) = (αβ)X; (6) 1 ·X=X. The additive and field multiplicative structures are related by the following distribu- tive rules (7) (α+β)X=αX+βX (8)α(X+Y) =αX+αY. Elements of the field Fare called scalars , whereas elements of Sare called vectors . We shall usually take the real numbers Rfor our field F, although the complex numbers Cwill be used at times. Exercise 4 shows the need for Axiom 6 (in case you thought it was superfluous). All of the examples of this section are linear spaces. For most purposes the simple example R2will serve you well as a guide to further expectations. The pictures there are simple. In fact, with a certain degree of cleverness, the “right” proof for R2immediately generalizes to all other linear spaces - even “infinite dimensional” ones. 80 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE Since you probably think that everything is a linear space, here is an example to dispel the delusion. Let Sbe the subset of all functions f(x) inC[0,1] which have the property f(0) = 1 . Then if fandgare inS, we are immediately stuck since f(0) +g(0) = 2 , so thatf+gisnotinS. Also, 0 /negationslash∈S. Both here, and before (p.?) when defining a field, axioms “0” have been used. They all express roughly the same concept. We have some set Sand an operation * defined on the set. These axioms all stated that for any x,y∈S, we also have x∗y∈S. In other words, the setSisclosed under the operation * in the sense that performing that operation does not take us out of the set. We shall find this concept useful. h) Appendix. Free Vectors One more example is needed, an exceedingly important example. There are “physicists’ vectors” or free vectors . I always thought they were easy to define - until today. Twelve hours and fifty pages later, I begin again on the fifth attempt. The essential idea is easy to imagine but difficult to convey in a clear and precise exposition. Say you are given two elements XandYofRn, which we represent by directed line segments from the origin. Somehow we want to find a directed line segment Vfrom the tip ofXtothe tip ofY. NowV“looks” like a vector. The problem is that all of the vectors we have met so far have been directed line segments in Rnbeginning at the origin. In order to find a way out, it is best to examine the problem for the most simple case −R1, the ordinary line. Watch closely since we will be so shrewd that all the formalism will be adequate without change for the general case of Rn. We are given two points, XandYofR1which we shall represent by directed line segments from the origin. To make the picture clear, we will draw them slightly above the line. a figure goes here We want a directed line segment Vfrom the tip of Xto the tip of Y. Of course you recognize this as the problem of solving X+V=Y The solution, V=Y−X, is the difference of the two real numbers YandX. But where should we draw V? If we are stubborn and demand that all real numbers must be represented by line segments beginning at the origin, we have the picture a figure goes here but what we really want to do is place the tail of Vat the tip of Xand add the line segments. Why not relent and allow ourselves this added flexibility. a figure goes here There! Now we have solved our problem. But we have made an important generalization in doing so. You see, this Vhas been released from its bondage to the origin and is now free to move along the whole of R. Although we were led to this Vfrom the pair XandY, the sameVcould have been generated by a different pair ˜Xand ˜Y, as the diagram below indicates, 2.1. EXAMPLES AND DEFINITION 81 a figure goes here for we still have ˜X+V=˜Y. In the first case we might have had X= 2 andY= 3 , so that V= 1 , while in the second, we might have had X=−4 andY=−3 , and again V= 1 . Even though we have let this Vgo free, sliding from place to place along R, we still want to say that this is only one V, and in fact, we want to identify this Vwith theVtied to the origin in (2). In other words, we would like to say that all three V’s used above are equivalent to each other. More formally, the element Visgenerated by an ordered pair, V= [X,Y] , which we read as the vector fromXtoY, forX, Y∈R. If some ˜Vis generated by another ordered pair, ˜V= [˜X,˜Y],˜X,˜Y∈R, then we want equality V=˜Vto mean that ˜Y−˜X=Y−X. Moreover, we want to representV= [X,Y] , the vector from XtoY, by the vector from the origin 0 to Y−X, V = [0,Y−X] .This representation of Vis unique , since if any other pair also generates V, V = [˜X,˜Y] , the representative V= [0,˜Y−˜X] = [0,Y−X] sinceV=Vimplies that ˜Y−˜X=Y−X. Therefore much as each rational number is an equivalence class, represented by a single rational number - as1 2represents the equivalence class1 2,2 4,3 6,..., eachVis an equivalence class of ordered pairs V= [X,Y] , whereX,Y∈R. It is uniquely represented by an element of R, viz.V=Y−X, the representation being independent of the particular ordered pair [ X,Y] which generates V. It is possible to think of Veither as an ordered pair with an equivalence relation, or just as the representative V= [0,Y−X] of the whole equivalence class, the representation being written more simply as an element of R:V=Y−X, where here equality is between elements of R. The generalization is now easily made Definition: . (Free vectors). Let XandYbe any elements of Rn. An element V∈Vn, “physicists’ n-space”, is defined as an equivalence class of ordered pairs of elements in Rn, V= [X,Y], X, Y ∈Rn, with the following equivalence relation: If V= [X,Y] and ˜V= [˜X,˜Y] , then V=˜V⇐⇒ ˜Y−˜X=Y−X, where the second equality is that of elements in Rn. If we are given XandYinRn, we speak ofV= [X,Y] as the free vector going fromXtoY. Previous reasoning also shows that eachV∈Vnis uniquely represented by the ordered pairV= [0,Y−X] . This representation is independent of the elements [ X,Y] which generatedV. We were led to this definition of Vnby examining the situation in the special case of V1. Since our formal reasoning there was quite algebraic and general, we know that the definition works algebraically. The geometry works too. An example in V2should make the general case clear. LetX= (1,3) andY= (2,1) . These two points in R2generate the ordered pair V= [(1,3),(2,1)] in V2.Vis the vector going fromX= (1,3)toY= (2,1) . Of all equivalentV’s, the unique representative which begins at the origin is V= [(0,0),(1,−2)] , which we simply write as V= (1,−2) and represent as an ordinary element of R2. On the same diagram we exhibit the vector from ˜X= (−2,2) to ˜Y= (−1,0) , which is ˜V= [(−2,2),(−1,0)] . The unique representative (of all ˜V’s equivalent of ˜V) which begins 82 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE from (0,0) is ˜V= [(0,0),(1,−2)] , which we write simply as ˜V= (1,−2) . Comparison of Vand ˜Vreveals that they are equal, V=˜V. Thus, from the diagram, we see that a free vector is an equivalence class of directed line segments, with two directed line segments V,˜Vbeing equivalent as vectors in V2if they are equivalent to the same directed line segment which begins at the origin. In more geometrical language, V=˜Vif by sliding them “parallel to themselves”, they can be made to coincide with their representer which begins at the origin. (We shall not define “parallel” here. It is not needed because we already have a satisfactory algebraic definition of equivalence.) Notice that X= (1,3) andY= (2,1) also generates a second ordered pair ˆV= [(2,1),(1,3)] , the vector fromY= (2,1)toX= (1,3) . Its unique representation which begins at the origin is ˆV= [(0,0),(−1,2)] , or more simply ˆV= (−1,2) . Comparison with the previous example shows that ˆV=−V:the vector from YtoXis the negative of the vector from XtoY. We need the little arrow on our picture of V= [X,Y] to distinguish it from −V= [Y,X] which is also between the same points but headed in the opposite direction. From now on we shall denote a vector V∈VnfromXtoYby its representative Y−XinRn, soV=Y−X. Hence the vector from (1,3) to (2,1) will be immediately written as V= (1,−2) . As we have said many times, the representation V=Y−Xas an element on Rnis independent of which particular pair [ X,Y] happened to generate V. The following diagram shows a whole bunch of equivalent vectors Vj∈V2, a figure goes here Vj=Vk, and their particular representative Vchained to the origin. In order to justify calling the elements of Vnvectors, we should prove that the elements ofVndo form a vector space . Addition and scalar multiplication must first be defined, an easy task. Since every V∈Vnis uniquely represented as an element of Rn, V=Y−X∈ Rn, we use addition and scalar multiplication for elements of Rn—which has already been defined. Because Rnis known to be a vector space, it is a tedious triviality to prove. Theorem 2.1 .Vnis a linear vector space. Proof: . Only a smattering. (1)Vnis closed under addition. Say V1andV2are in Vn. Then they are represented as the difference of two elements of Rn, sayV1=Y1−X1andV2=Y2−X2. Thus V1+V2= (Y1−X1) + (Y2−X2) = (Y1+Y2)−(X1+X2), so that their sum is generated by [ X1+X2, Y1+Y2] . In other words, there is at least one pair of elements, [ X3,Y3], X 3=X1+X2andY3=Y1+Y2, inRnwhich generateV1+V2, so thatV3=V1+V2∈Vn. Of course [0 ,Y3−X3] and many other pairs also generate V3. (2)Commutativity. V1+V2= (Y1−X1) + (Y2−X2) = (Y2−X2) + (Y1−X1) =V2+V1. (3) (α+β)V1= (α+β)(Y1−X1) =α(Y1−X1) +β(Y1−X1) =αV1+βV1 2.1. EXAMPLES AND DEFINITION 83 Example: IfA= (4,2,−3), B= (0,1,−2), C= (−1,0,1 2) andD= (4,−1 2,1) , find the vectorV1fromAtoBand the vector from CtoD. Then compute V1+ 2V2and V1−V2. solution: V1=B−A= (0,1,−2)−(4,2,−3) = ( −4,−1,1) V2=D−C= (4,−1 2,1)−(−1,0,1 2) = (5,−1 2,1 2) V1+ 2V2= (−4,−1,1) + 2(5,−1 2,1 2) = (−4,−1,1) + (10,−1,1) = (6,−2,2) V1−V2= (−4,−1,1)−(5,−1 2,1 2) = (−4,−1,1) + (−5,1 2,−1 2) = (−9,−1 2,1 2) Exercises (1) (a) Find the vector representing the free vectors from the given A∈RntoB∈Rn. (i)A= (3,1), B= (2,2). (ii)A= (−3,3), B= (0,4). (iii)A= (2,2,3), B= (5,2,17) (iv)A= (0,0,0)B= (9,8,−3) (v)A= (1,2,3), B= (0,0,−1) (vi)A= (0,0,−1), B= (1,2,3) (b) LetV1andV2be the respective vectors of iii) and v) above. Compute V1+ V2, V1−V2, and 2V1−3V2. (c) Draw a diagram on which you indicate the vector going from A= (3,1) to B= (2,2) , and indicate the representer of that vector which begins at the origin. Do the same with the vector from BtoA. (2) Which of the following subsets of C[−1,1] are linear spaces: (a) The set of all even functions in C[−1,1] , that is, functions f(x) with the addi- tional property f(−x) =f(x) , likex2and cosx. (b) The set of all functions finC[−1,1] with the additional property that |f(x)| ≤ 1 . (c) The set of all functions finC[−1,1] with the property that f(0) = 0 . (3) In R3, letX= (1,−1,2) andY= (0,4,−3) . FindX+ 2Y, Y−X, and 7X−4Y. (4) (a) Show that for every X∈R3you can find scalars αj∈Rsuch thatXcan be written as X=α1e1+α2e2+α3e3, wheree1= (1,0,0), e2= (0,1,0), e3= (0,0,1) . (b) IfX∈R3, can you find scalars αj∈Rsuch that X=α1θ1+α2θ2+α3θ3, whereθ1= (1,−1,0), θ2= (−1,1,0), θ3= (0,0,1) , andαj∈R? Proof or counter-example. 84 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE (c) Find two polynomials P1(x) andP2(x) in ? 1such that for every polynomial P(x)∈?1you can find scalars αj∈Rsuch thatPcan be written in the form P(x) =α1P1(x) +α2P2(x). (5) LetV=R×Rwith the following definition of addition and scalar multiplication X+Y= (x1+x2, y1+y2), αX = (αx1,0), 0 = (0,0),−X= (−x1,−x2). IsVa vector space? Why? (6) Show that any field can be considered to be a vector space over itself. (7) Consider the set S={u∈C2[0,1]:a2u/prime/prime+a1u/prime+a0u= 0}, where theaj(x)∈C[0,1] . IsSa linear space? Note that we do not yet know that Shas any elements at all. The proof that Sis not empty is the existence theorem for ordinary differential equations. (8) By integrating |x|the “right” number of times, find a function which is in Ck[−1,1] but is not in Ck+1[−1,1] . 2.2 Subspaces. Cosets. With this section we begin the process of assigning names to the various concepts sur- rounding the idea of a linear vector space. This name calling will take us the balance of the chapter. Although the ideas are elementary and theorems simple, do not deceive yourselves into thinking this must be some grotesque joke that mathematicians have perpetrated. You see, we are in the process of building a machine. Most of its constituent parts are very easy to grasp. But when combined, the machine will be equipped successfully to assault a diversity of problems which appear off hand to be unrelated. The value of this abstract formalism is that many seemingly distinct complicated specific problems are just one single problem in a variety of fancy dresses. By ignoring the extraneous paraphernalia we can concentrate on the essential issues. a figure goes here We begin by defining what is meant by a subspace of a vector space W. While reading the definition, think of a plane through the origin, which is a subspace of ordinary three dimensional space. Definition: . A setAis alinear subspace (linear variety, linear manifold) of the linear spaceWif i)Ais a subset of W, and ii)Ais also a linear space under the operations of vector addition and multiplication by scalars already defined on V. 2.2. SUBSPACES. COSETS. 85 Examples: (1) LetA={X∈R3:X= (x1,x2,0)}, that is, the points in R3whose last coordinate is zero. Since A⊂R3, and a simple check shows that Ais also a linear space, we see that Ais a linear subspace of R3. Intuitively, this set Acertainly “looks like” R2. You are right, and recall that the fancy word for this equivalence - of R2= (x1,x2) and the points in R3of the form ( x1,x2,0) —is isomorphic . Similarly, the setB={X∈R3:X= (x1,0,x3)}is also a subspace of R3.Bis also isomorphic to R2. (2) LetA={X∈Rn:X= (x1,x2,...,x k,0,0,...,0)}, that is, the points in Aare those points in Rnwhose last n−kcoordinates are zero. It is easy to see that Ais a linear subspace of Rn, and that Ais isomorphic to Rk. (3) LetA={f∈C[0,1]:f(0) = 0 }.Ais a subset of the linear space C[0,1] , and is also a linear space (check this). Thus Ais a linear subspace of C[0,1] . (4) LetA={f∈C[0,1]:f(0) = 1 }.Ais a subset of C[0,1] , but it is not a linear subspace since - as we saw in the last section (p. ?)— Ais itself not a linear space. The following lemma supplies a convenient criterion for checking if a given subset Aof a linear space Wis a subspace. Theorem 2.2 . IfAis a non-empty subset of the linear space W, thenAis a linear subspace of W⇐⇒Ais closed under addition of vectors in Aand multiplication by all scalars. Proof: .⇒. SinceAis a subspace, it is itself a linear space. But all linear spaces are, by definition, closed under addition and multiplication by scalars. ⇐. BecauseAis a subset of W, and properties 1,2,5,6,7, and 8 hold in W, they also hold for the particular elements in Wwhich happened to be in A. Notice that here we use the fact that Ais closed under addition. Therefore only the existential axioms 3 and 4 need be checked. Since Ais not empty, it contains at least one element, say X∈A. Because Ais closed under multiplication by scalars we see that 0 = 0 ·X∈A. Furthermore, for everyX∈A, also −X= (−1)·X∈A. Example: LetA={f∈C1[0,1]:f/prime(0) = 0 }. SinceAis a subset of the linear space C1[0,1] , all we need show is that Ais closed under addition and multiplication by scalars in order to prove Aa linear subspace of C1[0,1] . Iff, g∈A, then (f+g)/prime(0) = (f/prime+g/prime)(0) = f/prime(0) +g/prime(0) = 0 , so f+g∈A. Also, for any α∈R,(αf)/prime(0) =α(f/prime)(0) =α·0 = 0 , so αf∈A. Theorem 2.3 . The intersection of two subspaces is also a subspace, but the union of two subspaces is not necessarily a subspace. More generally, the intersection of any collection of subspaces is also a subspace. Proof: . LetA, B be subspaces of W. We show that A∩Bis a subspace. Since A∩B⊂W, all we need show is the closure properties of A∩B. IfX, Y∈A∩B, then XandYare both in AandB, soX+Y∈AandX+Y∈B⇒X+Y∈A∩B too. Similarly for scalar multiples. The proof that A∩B∩C∩...is a subspace is identical except for a notational mess. 86 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE For the second part of the theorem we merely exhibit an example of two subspaces A,B for whichA∪Bis not a subspace. In R2letAbe the linear subspace “horizontal axis”, that is, A={X∈R2:X= (x1,0)}, whileBis “the vertical axis”, B={X∈ R2:X= (0,x2)}. ThenA∪Bis the “cross” of all points on either the horizontal axis or the vertical axis. This is not a linear space because points like (1 ,0)∈A,(0,1)∈Bdo not have their sum (1 ,0) + (0,1) = (1,1) inA∪B. Precisely for this reason R2=R1×R1 was constructed as the Cartesian product of R1with itself; for if it had been constructed asR1×R1, then only the points situated on the axes themselves would get caught. More generally - and for the same reason - the Cartesian product is the process always used to “glue” together a larger space from several linear spaces. Only when A⊂B(orB⊂A) isA∪Balso a subspace (Exercise 4). Your image of a linear space should be R3, and a subspace Sis a plane or line in R3. Note that since every subspace must contain 0, these planes or lines must pass through the origin . Example: LetSc={X∈R2:x1+ 2x2=c, creal}. Thus, the set Scis all points S= (s1,s2)∈R2on the straight line s1+ 2s2=c. For what value(s) of cisSca subspace? If Scis a subspace, then we must have aS∈Scfor all scalars a, that is aS= (as1,as 2)∈Sc⇒as1+ 2as2=c. But fora= 0 this states that c= 0 . Therefore the only possible subspace is S0={X∈R2:x1+ 2x2= 0}. It is easy to check that if S1 andS2are in S0, then so are S1+S2andaS1. Thus S0is a subspace. Similarly, every straight line through the origin is a subspace. Our question now is, how can we talk about the other straight lines or planes which do not happen to pass through the origin? First we answer the question for our example above. There we have the linear space R2and the subspace S0which will be simply written as S.Sis a line through the origin. Let X1be any element in R2(think ofX1as a point). Then the set of all elements of R2which can be written in the form S+X1, whereS∈S, is the line “parallel” to Swhich passes through X1. This line is written as S+X1. More explicitly, say X1= (1,3 2) . The set S+X1is the set of all points X= (x1,x2)∈R2of the form X=S+X1,whichS∈S, or (x1,x2) = (s1,s2) + (1,3 2),wheres1+ 2s2= 0. Consequently x1=s1+ 1 , andx2=s2+3 2. Using the relation s1+ 2s2= 0 , we find that x1+2x2= 4 — exactly the equation of the straight line through X1= (1,3 2) and “parallel” to the subspace S. This subset, S+X1={X∈R2:X=S+X1, whereS∈S}, is called theX1coset of S. Thus, cosets are the names given to “linear objects” which are not subspaces. They are subspaces translated to pass through X1. You might prefer to call them affine subspaces instead of cosets. Please observe that the cosets S+X1andS+X2, whereX1, X 2∈W, are not necessarily distinct. In our example, these cosets coincide if and only if X2is on the line S+X1, that is, if X2∈S+X1. The easiest way to test this is to see if X2−X1∈S. SayX1= (1,3 2) as before, and that X2= (2,1) . Then the cosets S+X1andS+X2are the same since the point X2−X1= (1,−1 2) is in S. It should be geometrically clear that 2.2. SUBSPACES. COSETS. 87 the relation of equality among these cosets is an equivalence relation (and so deserving of the title “equality”). We shall state these ideas formally as we turn from this special - but characteristic - example to the general situation. The general problem of describing lines or planes or “higher dimensional linear objects” which do not pass through the origin - so are not subspaces - is solved similarly. Definition: . LetWbe a linear space, Sa subspace of V, andX1any element of W. All elements in Wwhich can be written in the form S+X1, whereS∈S, is called the X1coset ofS, and written as S+X1. Our first theorem states that if X2is in theX1coset of S, thenX1is in theX2 coset of S: Theorem 2.4 .X2∈S+X1⇐⇒X1∈S+X2. Proof: SinceX2∈S+X1, there is an S∈Ssuch thatX2=S+X1. Therefore X1= (−S) +X2. Because Sis a linear space, ( −S)∈S. ThusX1has been written as the sum of X2and an element of S, which means that X1∈S+X2. By the same argument, one sees that any two cosets S+X1andS+X2are either identical or are disjoint (have no element in common). Thus the cosets of SpartitionW in the sense that every element of Wis in exactly one coset, just as for our example, every point in the plane R2was in exactly one straight line parallel to the subspace determined byx1+ 2x2= 0 . Although we were motivated by geometrical considerations, the ideas apply without alteration to any linear space. This is illustrated by again examining the set A={f∈C[−1,1]:f(0) = 1 }, which is not a subspace. It is a coset of a subspace SofC[−1,1] which is constructed as follows. Consider the subspace Swhich is “naturally” associated with A, viz. S={g∈C[−1,1]:g(0) = 0 }. ThenAis the coset S+1, A=S+1 . This is true since clearly A⊃S+1 . AlsoA⊂S+1 because for every f∈A, f(x) = [f(x)−1] + 1 =g(x) + 1,whereg∈S. ThereforeA=S+1 . Similarly, we could have written AasS+ˆf, where ˆfisanyfunction inA, for example A=S+ cosx. Exercises (1) Find which of the following subsets of Rnare subspaces. (a){X∈Rn:x1= 0}, (b){X∈Rn:x1≥0}, (c){X∈Rn:x1−x2= 0}, (d){X∈Rn:x1−x2= 1}, (e){X∈Rn:x2 1−x2= 0}, 88 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE (2) In P3, the linear space of all polynomials of degree ≤3 , letA={p(x)∈P3:p(0) = 0}, and letB={p(x)∈P3:p(1) = 0 }. (a). Show that AandBare subspaces of P3. (b). FindA∩BandA∪B. Give an example which shows that A∪Bis not a subspace of P3. (3) (a) If X1andX2are given fixed vectors in R2then is A={X∈R2:X=a1X1+a2X2, a1anda2any scalars } a subspace of R2? (b) Same as (a) but replace R2by an arbitrary linear space W. (c) IfX1,X2,...,X k∈W, then is A={X∈W:X=k/summationdisplay 1ajXj,for any scalars aj}, a subspace of W? (4) LetAandBbe subspaces of a linear space W. Prove that A∪Bis also a subspace if and only if either A⊂BorB⊂A, that is, if one of the subspaces contains the other. (5) Let SandTbe subspaces of a linear space W, and suppose that Ais a coset ofSandBis a coset of T. Prove that (a). A⊂B⇒S⊂T, and also (b). A=B⇒S=T. (6) (a) Write the plane 2 x1−3x2+x3= 7 as a coset of some suitable subspace S⊂R36. (b) Write the set A={f∈C[0,4]:f(0) = 1, f(1) = 3 }, as a coset of some suitable subspace S⊂C[0,4] . (c) Write the set A={f∈C1[0,4]:f(1) = 1, f/prime(1) = 2 }as a coset of some suitable subspace S⊂C1[0,4] . 2.3 Linear Dependence and Independence. Span. IfWis a linear space and X1,X2,...,X k∈W, then we know that, for any scalars aj, Y=k/summationdisplay j=1ajXj=a1X1+a2X2+...+akXk is also in V.Yis alinear combination of theXj’s. Now if 0 can be expressed as a linear combination of the Xj’s, where at least one of the aj’s is not zero we expect that there is something degenerate around. In fact, if 0 = a1X1+...+akXkwhere saya1/negationslash= 0 , then we can solve for X1as a linear combination of X2,X3,...,X k, X1=−1 a1(a2X2+...+a,Xk). This leads us to make a definition and state a theorem. 2.3. LINEAR DEPENDENCE AND INDEPENDENCE. SPAN. 89 Definition: . A finite set of elements Xj∈W, j = 1,...,k is called linearly dependent if there exists a set of scalars aj, j= 1,...,k ,notall zero such that 0 =k/summationdisplay 1ajXj. If theXj are not linearly dependent, we say they are linearly independent . Theorem 2.5 . A set of vectors Xj∈W, j = 1,...,k is linearly dependent if and only if at least one of the Xj’s can be written as a linear combination of the other Xj’s. To test if a given set of vectors is linearly independent, an equivalent form of Theorem 5 is useful. Corollary 2.6 A set of vectors Xj∈W, j = 1,...,k is linearly independent if and only if k/summationdisplay j=1ajXj= 0 implies that a1=a2=...=ak= 0. Examples: (1) The vectors X1= (2,0), X 2= (0,1), X 3= (1,1) in Rare linearly dependent since 0 =X1+ 2X2−2X3. Equivalently, we could have applied the theorem since X3can be written as a linear combination of X1andX2 X3=1 2X1+X2. (2) The functions f1(x) =ex, f2(x) =e−x, f3(x) =ex+e−x 2inC[0,1] are linearly depen- dent since 0 =f1+f2−2f3 (3) The vectors X1= (2,0,1), X 2= (−1,0,0) in R3are linearly independent, since if for somea1, a2, 0 =a1X1+a+ 2X2= (2a1,0,a1) + (−a2,0,0), then 0 = (0,0,0) = (2a1−a2,0,a1), which implies that 2 a1−a2= 0 , anda1= 0 =⇒a1=a2= 0 . a figure goes here A simple consequence of these ideas is the following Theorem 2.7 . IfAandBare any subsets of the linear space Wand ifA⊂B, then i)Ais linearly dependent ⇒Bis linearly dependent; and the contrapositive: ii) Bis linearly independent ⇒Ais linearly independent. We now prove the transitivity of linear dependence. 90 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE Theorem 2.8 . IfZis linearly dependent on the set {Yj}, j= 1,...,n and eachYj is linearly dependent on the set {Xl}, l= 1,...,m thenZis linearly dependent on the {Xl}. Proof: . This is trivial arithmetic. We know that Z=a1Y1+...+anYn, and that Yj=c1jX1+c2jX2+...+cmjXm By substitution then Z=al(cllXl+···+cmlXm) +a2(c12X1+···+cm2Xm) +···+an(c1mX1+···+cmnXm) = (a1c11+a2c12+···+ancln)X1+ (a1c21+···+anc2n)X2 +···+ (a1cml+···+cmn)Xm =γ1X1+···+γmXm,whereγl=n/summationdisplay j=1ajclj. More concisely: Z=n/summationdisplay j=1ajYj=n/summationdisplay j=1aj/parenleftBiggm/summationdisplay l=1cljXl/parenrightBigg =m/summationdisplay l=1 n/summationdisplay j=1ajclj Xl=m/summationdisplay l=1γlXl. LetX1andX2be any elements of a linear space W. Is there a smallest subspace AofWwhich contains X1andX2? There are two possible ways of answering this, constructively and non-constructively. First, constructively. We observe that the desired subspace must contain X1and X2, and all linear combinations of X1andX2, that is,Amust contain all X∈W of the form X=a1X1+a2X2for all scalars a1anda2. But observe that the set B={X∈V:X=a1X1+a2X2}is a linear space, since if XandY∈B, thenaX∈B for any scalar a, and alsoX+Y∈B. Thus the desired subspace Ais justBitself. The constructive proof goes as follows: just let Abe the intersection of all subspaces containing X1andX2. By Theorem 3 the intersection of these subspaces is also a subspace. It is clearly the smallest one. Do you feel cheated? This type of reasoning is often used in modern mathematics. Although it reveals little more than the existence of the sought-after object, it is an extremely valuable procedure when you really don’t want anything more than to know the object exists. More important, procedures like this are vital when there is no constructive proof available. More generally, if S={Xj}, j= 1,...,k , is any finite subset of a linear space W, we ask for the smallest subspace AofWwhich contains S. There are two proofs - exactly as in the simple case above (where k= 2 ). From the constructive proof we find that A={X∈W:X=k/summationdisplay 1ajXj, ajscalars }, soAis the set of all linear combinations of the Xj’s. This set Ais called the span ofS, and denoted by A= span(S) . We also say that SspansA, or thatAisgenerated byS. 2.3. LINEAR DEPENDENCE AND INDEPENDENCE. SPAN. 91 Examples: (1) In R3letX1= (1,0,0) andS2= (0,1,0) . Then the span of S={Xj, j= 1,2}is allX∈R3of the form X=a1X1+a2X2= (a1,a2,0) . If we imagine R3as ordinary 3-space, then the span of X1andX2is the entire x1,x2plane. (2) In R3, letX1= (1,0,0), X 2= (0,1,0) , andX3= (0,0,1) . Then the span of T={Xj, j= 1,2,3}is allX∈R3of the form X=a1X1+a2X2+a3X3(a1,a2,a3) . Since all of R3can be so represented, we have span( T) =R3, that is, the set Bspans R3. Comparing these two examples, we see that S⊂Tand span(S)⊂span(T) . (3) In R3, letX1= (1,0) andX2= (0,1) . Then the span of S={X1,X2}is all of R2, since every X∈R2can be written as X=a1X1+a2X2, wherea1anda2are scalars. Many other sets also span R2. In fact almost every set of two vectors X1and X2inR2span R2. This can be seen from the diagram, where we have drawn a net parallel to X1andX2. ThenX=a1X1+a2X2. Any vectors X1andX2would do equally well, as long as they do notpoint in the same (or opposite) direction. We collect some properties of the span Theorem 2.9 . LetR, S , andTbe subsets of a linear space W. Then (a)R⊂span(R). (b)R⊂S=⇒span(R)⊂span(S). (c)R⊂span(S)andS⊂span(T) =⇒R⊂span(T). (d)S⊂span(T) =⇒span(S)⊂span(T). (e) span(span( T)) = span(T). (f)A vectorXj∈Sis linearly dependent on the other elements of S⇐⇒span(S) = span(S−{Xj}). (HereS−{Xj}means the set Awith the one vector Xjdeleted). Proof: These all depend on the representation of span( S) as a linear combination of the elements of S. (a) and (b)—Obvious. They really should be if you understand the definitions. (c). A direct translation of Theorem 7. (d). This is the special case R= span(S) of part c. (e). By part (a) span(span( T))⊃span(T) . The opposite inclusion span(span( T))⊂ span(T) is the special case S= span(T) of part (d). (f).Xjlinearly dependent on S−{Xj}=⇒S⊂span(S−{Xj}) . Thus by part (d), span(S)⊂span(S−{Xj}) . Inclusion in the opposite direction span( S−{Xj})⊂span(S) follows from part (b). Therefore span( S) = span(S− {Xj}) means that Xj∈span(S) can be expressed as a linear combination S− {Xj}, i.e., the other Xk’s. Now most likely this proof was your first taste of abstract juggling and you find it difficult. Relax and don’t be impressed with how formidable it appears. Except for parts a and b, the whole business hinges on the explicit construction of Theorem ?. Since (d) is a special case of (c), a good exercise is to write out the proof of (d) without relying on (c). InR2, letX1= (1,0) ,X2= (0,1) , andX3be any vector in R2. Observe that X1 andX2together span R2. ThusX3can be expressed as a linear combination of X1and X2, so thatX1, X 2, andX3are linearly dependent. The next theorem is a generalization of this idea. 92 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE Theorem 2.10 . If a finite set A={Xj, j= 1,...,n }spans a linear space W, then ev- ery set ˜S={Yj∈V, j= 1,...,m>n }with more than nelements is linearly dependent. In other words, every linearly independent set has at most nelements. Proof: Pick anyn+ 1 elements Y1,...Y n+1from ˜Sand throw the rest away. Call the new setS. We shall show that these n+ 1 elements are linearly dependent. Then, since S⊂˜S, Theorem ? tells us that ˜Sis also linearly dependent. The only problem is how to carry out the proof without getting involved in a mess of algebra. By the principle of conservation of effort, this means that there will be some fancy footwork. Reasoning by contradiction, assume Sis linearly independent. If we can show that span(A) = span(S−{Yn+1}) , then span( S)⊂span(S−{Yn+1}) because span( S)⊂V= span(A) = span(S− {Yn+1})) . Since span( S− {Yn+1})⊂span(S) , we can apply part f of Theorem ? to conclude that Sis linearly dependent - the desired contradiction. Thus, assuming S={Y1,...,Y n+1}is linearly independent, we are done if we prove that span(A) = span(S− {Yn+1}) . Consider the set Bk={Y1,...,Y k, Xk+1,...,X n}. We know that B0=A, so that span( B0) = span(A) =W. Then by induction we shall prove that span( Bk) =W=⇒span(Bk+1) =W. Since span( Bk) spansW, thenYk+1 is a linear combination of the elements of Bk. Because the Y’s are assumed linearly independent, this linear combination must involve at least one of Xk+1,...,X n. Say it involvesXk+1(if not, relabel the X’s to make it so). Then we can solve for Xk+1as a linear combination of span( Bk+1) . Therefore W= span(Bk) = span(Bk+1) . Putting this part together, we find that span( A) =W= span(B0) = span(B1) =...= span(Bn) . But Bn=S− {Yn+1}. Thus span( A) = span(S− {Yn+1}) , and the proof is completed. Example . InR2, any three (or more) non-zero vectors are linearly dependent since the two vectorsX1= (1,0) andX2= (0,1) span R2. Exercises (1) (a) In P2p1(x) = 1, p2(x) = 1 +x, p 3(x) =x−x2 (b) In R3, X 1= (0,1,1), X 2= (0,0,−1), X 3= (0,2,3) . (c) InC[0,π], f(x) = sinx, g(x) = cosx. (d) In Rn, e1= (1,0,0,..., 0), e2= (0,1,0,0),...,e n= (0,0,..., 0,1) . (2) Use the result of (d) to show that any set of n+ 1 vectors in Rnmust be linearly dependent. (3) (a) Find a set which spans i)P3, ii)R4 (b) Show that no finite set spans l1. (4) LetX1,...,X kbe any elements of a linear space V. (a) Prove that span( {X1,...,X k}) = span( {X1+aXj,X2,...,X k}) , whereais any scalar and Xjis any of the X2,X3,...,X k,. (b) Prove that span( {X1,...,X k}) = span( {aX1,X2,...,X k}),a/negationslash= 0 . 2.4. BASES AND DIMENSION 93 (c) In Rn, consider the ordered set of vectors {X1,X2,...,X k}, whereXj= (x1j,x2j,...,x nj) . They are said to be in echelon form if i) noXjis zero, and ii) theindex of the first non-zero entry in Xjis less than the index of the first non- zero entry in Xj+1, for eachj= 1,...,k −1 . ThusX1= (0,1,0), X 2= (0,0,1) are in echelon form while X1= (0,1,0), X 2= (1,0,1) are not in echelon form. Prove that any set of vectors in echelon form is always linearly independent. (I suggest a proof by induction). (5) For what real value(s) of the scalar αare the vectors ( α,1,0),(1,α,1) and (0,1,α) inR3linearly dependent? (6) (a) In R3, letX1= (3,−1,2) . Express ( −6,2,−4) linearly in terms of X1. Show that (3,4,−7) cannot be expressed linearly in terms of X1. Can (1,2,1) be expressed linearly in terms of X1? (b) In R3, letA={X1,X2}, whereX1= (1,3,−2) andX2= (2,1,1) . Express (3,−1,4) linearly in terms of A. Show that (0 ,0,2) cannot be expressed linearly in terms of A. Can (0,5,−5) be expressed linearly in terms of A? (7) (a) In C[0,10] , letf1,...,f 8be defined by f1(x) =x2−x+ 2, f 5(x) =x3 f2(x) = (x+ 1)2f6(x) = sinx f3(x) =x+ 3 f7(x) = cosx f4(x) = 1 f8(x) = sin(x+π/4). LetA={f1,f2,f3}. Expressf4linearly in terms of A. Show that f5cannot be expressed linearly in terms of A. Isf6∈span(A) ? Isf8∈span(f6,f7) ? Is f6∈span(f5,f7,f8) ? (b) If we let f9(x) = (x−1)3, f10(x) = 2x−1 , determine which of the following sets are linearly dependent: (i){f1,f3,f10}, (ii){f1,f5,f9}, (iii){f3,f4,f10}, (iv){f1,f4,f5,f9}, 2.4 Bases and Dimension If the set {X1,...,X m}spans the linear space W, is there any set with less than m vectors which also spans W? There certainly is if the {X1,...X m}are linearly depen- dent, for if say Xmdepends linearly upon the {X1,...,X m−1}, then by Theorem ??, span({X1,...,X m}) = span( {X1,...,X m}) =W, so then {X1,...,X m−1}spanW. We can continue and eliminate the extra linearly dependent elements until we obtain a set {X1,...,X n}of linearly independent vectors which still span W. Definition . A set of vectors Xj∈W,j= 1,...,n which is i) linearly independent, and ii) spansWis called a basis forW. Examples . 94 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE (1) In R2, the vectors X1= (1,0) andX2= (0,1) are linearly independent and span R2. Therefore X1andX2form a basis for R2. The vectors X3= (3,−1) and X4= (−2,2) in R2are also linearly independent and span R2. They thus constitute another basis for R2. Almost any two vectors in R2span R2, as long as they do not point on the same or opposite direction. (2) In P2, the polynomials p1(x) = 1 , and p2(x) =x−x2donotform a basis. They are linearly independent but do not span the space - since for example you can never obtain the polynomial p(x) =xwhich is in P2. If we add the third polynomial, say p3(x) =x−2x2, thenp1, p2andp3do form a basis for P2. Bases have an important property. Theorem 2.11 . If{X1,...,X n}form a basis for the linear space W, then every X∈ Wcan be expressed uniquely as a linear combination of the Xj’s. Remark : Every set which spans Whas, by definition, the property that every X∈Wcan be expressed as a linear combination of the Xj’s. The point here is that for a basis, this linear combination is uniquely determined. Proof: Suppose that X=n/summationdisplay 1akXkand alsoX=n/summationdisplay 1bkXk. We must show that ak=bk for allk. Subtracting the two equations we find that 0 =n/summationdisplay 1ckXk, whereck=ak−bk. But since the Xk’s are linearly independent, by the Corollary to Theorem 5, the only way a linear combination can be zero is if ck= 0, k= 1,...,n , that is,ak=bkfor allk. We have observed that a linear space may have several different bases. Is it possible that different bases contain a different number of elements? Our next theorem states that the answer is NO. Theorem 2.12 . If a linear space Whas one basis with a finite number of elements, say n, then all other bases are finite and also have exactly nelements. Proof: We invoke Theorem ?. Let Abe a basis with nelements and Bbe a basis with melements. Now AspansWand the elements of Bare linearly independent, so the Theorem ?, m≤n. Reversing the roles of AandBwe find that n≤m. Therefore n=m. With this result behind us, we can now define the dimension of a linear space. Definition . If a linear space Whas a basis with nelements, then we say that the dimension ofWisn. If a linear space Whas the property that no finite set of elements spans it, we say it is infinite dimensional . Remarks . Theorem ? states that the dimension of Wis independent of which basis we happened to pick. If we want to emphasize the dimension of a finite dimensional space, we will writeWn. Announcement . The dimension of Rnisn, for thenelementse1= (1,0,0,..., 0), e2= (0,1,0,..., 0),..., e n= (0,..., 0,1) are linearly independent and span Rn. A picture. We have seen that e1= (1,0,0), e2= (0,1,0), e3= (0,0,1) form a basis inR3. Thus every X∈R3can be expressed uniquely as a linear combination of the ej’s, X=a1e1+a2e2+a3e3. If we represent e1as a directed line segment from the origin to (1,0,0) , and similarly for e2ande3, thenXis the geometrical sum of a1e1+a2e2+a3e3, 2.4. BASES AND DIMENSION 95 and is represented as a directed line segment from the origin to ( a1,a2,a3) . In R3, e1is usually written as i,e2asjande3ask, so that a vector X∈R3is written as X=a1i+a2j+a3k. The points in the plane x3= 0 , which is isomorphic to R2, are then represented as X=a1ˆi+a2ˆj+ 0ˆk=a1ˆi+a2ˆj. We would retain this notation except that one runs out of letters when considering spaces of higher dimension. For that reason the subscript notation e1,e2,... is better suited to our purposes. It behooves us to show that the linear space C[0,1] of functions continuous in the interval [0,1] is infinite dimensional. This will be done by proving that the functions f0(x) = 1, f1(x) =ex, f2(x) =e2x,...,f n(x) =enx,... are linearly independent. Assume that 0 =N/summationdisplay k=0akekx, whereNis any non-negative integer. We must show that all the ak’s are zero. The trick is to use induction. For N= 0 , we know that 0 = a0only ifa0= 0 . Suppose 1,ex,e2x,...,e(N−1)xare linearly independent. ThenN−1/summationdisplay k=0akekx= 0 if and only if all of the ak’s are zero. Let us show that this implies thatN/summationdisplay k=0akekx= 0 if and only if all theak’s vanish. Take the derivative. The constant term drops out and we are left with 0 =a1ex+ 2a2e2x+···+NaNeNx. Factor out ex 0 =ex(a1+ 2a2ex+···+NaNe(N−1)x). Sinceexis never zero, we know that 0 =a1+ 2a2ex+···+NaNe(N−1)x. By our induction hypothesis, this linear combination of ! ,ex,...,e(N−1)xcan be zero if and only if a1=a2=a3=···=aN= 0 . It remains to show that a0= 0 . This is an immediate consequence ofN/summationdisplay k=0akekx= 0 and the vanishing of ak, fork≥1 . Since the functions 1 ,ex,e2x,..., are inCk[a,b] for anykwe have shown that these spaces are infinite dimensional too. Moreover, the exact same proof also shows that the set{eα1x,eα2x,...,eαNx}, whereα1,...,a Nare arbitrary distinct complex numbers, is linearly independent. This fact will be needed later. Perhaps we shall present a different proof - or several different ones - at that time. All of the other proofs still involve some calculus - but that should be no surprise since we used calculus to define the exponential function in the first place. Not all spaces of functions are infinite dimensional. For example, the function space A={f∈C[−1,1]:f(x) =a+bex, a, b∈R}has dimension 2. The functions f1(x) = 1 andf2(x) =exconstitute a basis for Abecause every f∈Acan be written in the form f=a1f1+a2f2, wherea1anda2are real numbers. Another basis for Aisf3(x) = 1+ex andf4= 2−ex. There are many ways to see this. One is to observe that f3+f4= 3 and 2f3−f4= 3ex. Thus iff(x) =a+bex∈A, thenf=a 3(f3+f4) +b 3(2f3−f4) = (a 3+2b 3)f3+ (a 3−b 3)f4. 96 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE The function space B={f∈C[−1,1]:f(x) =asin(x+α), α,a ∈R}, also has dimension two, since f(x) = (acosα) sinx+ (asinα) cosx=a1sinx+a2cosx. Thus f1(x) = sinxandf2(x) = cosxform a basis. Actually, we have only shown that f1 andf2spanB, but not that they are linearly independent. You can settle that point yourselves. A few more remarks should be added. If Ais a subspace of an ndimensional space Wn, we would like to enlarge a basis {e1,...,e k}forAto a larger basis {e1,...,e n}for all ofW. SinceA⊂Wn, it is clear that k= dimA≤n. IfA=Wn, we are done since {e1,...,e k}already span Wn. Otherwise there is some element ek+1inWnwhich is not inA. LetA1= span {e1,...,e k+1} ⊂Wn. IfA1=Wn, then {e1,...,e k+1}form a basis for Wn. Otherwise there is some element ek+2inWnwhich is not in A1. Form A2=sp{e1,...,e m+2}. Repeat this process until you finally get a basis for all Wn. Only a finite number of steps are needed since the dimension of Wnis finite. This proves Theorem 2.13 . IfAis a subspace of (finite dimensional) space W, then any basis for Acan be extended to a basis that spans all of W. Consider a subspace Aof a linear space W. Somehow we would like to discuss - and give a name to - the part A/primeofVwhich is not in A. We would like A/primeto be a subspace ofVsuch that the only element of VwhichAandA/primeshare is 0, and such that every element in Vcan be written as the sum of an element in Aand an element in A/prime. Definition: . LetAbe a subspace of the linear space V. Acomplementary subspace A/prime ofAis a subset of Vwith the properties 1.A/primeis a subspace of V, 2. IfX∈V, thenX=X1+X2, whereX1∈AandX2∈A/prime. 3.A∩A/prime= 0 . (The zero vector, notthe empty set). Our first task is to prove Theorem 2.14 . Every subspace A⊂Vhas at least one complement A/prime. . Proof: Let{e1,...,e m}be a basis for A, and {e1,...e m,em+1,...,e n}an extension to a basis for V. We shall verify that A/prime=sp{em+1,...,e n}satisfies both criteria. Now ifX∈AandX∈A/prime, then we can write X=a1e1+...+amem∈A, and X=am+1ee+1+...+anen∈A/prime. Subtracting these equations, we find 0 =a1e1+...+amem−am+1em+1−...−anen. But since {e1,...,e n}is a basis for V, the elements are linearly independent. Thus a1=a2=...=am=am+1=...=an= 0 , soX= 0 . Therefore A∩A/prime= 0 . Furthermore, if X∈Vsince {e1,...,e n}is a basis for V, then X=n/summationdisplay j=1cjej=m/summationdisplay j=1cjej+n/summationdisplay j=m+1cjej. Thus we just let X1=c1e1+...+cmem∈AandX2andX2=cm+1em+1+...+cnen. It is easy to see that the above construction of A/primeisindependent of the basis chosen for A. This is because the construction of em+1,...,e n(Theorem ??) did not depend on the particular basis for A. That construction only utilized the fact that we can pick elements 2.4. BASES AND DIMENSION 97 notinA. However, the construction of A/primedoes depend on which elements em+1,...e n (not inA) we pick. For example, let V=R2, andAbe some one dimensional subspace. Then we pick e, as any vector in A, ande2as any vector not in A. The resulting complement A/primeis then the span of e2. But {e1}could have been extended to a basis forVby choosing another vector ˜ e2/ownerA. This determines a different complement ˜A/primeof A. A subspace has many possible complements. This ambiguity will not bother us since we shall only use the properties of a particular complement which do not depend on which particular complement is chosen. The dimension of the complement is such a property. It only depends on the dimension of the subspace Aand the larger space V, and has the reasonable formula dim A/prime= dimV−dimA, which we now prove. Theorem 2.15 . IfAis a subspace of a linear space Vand ifA/primeis any complement of A, then dimA+ dimA/prime= dimV. Thus, the dimension of A/primeis determined by AandValone. Proof: The dimAand dimVare given data. We shall compute dim A/prime. Since the union of a basis for Awith a basis for any A/primespansV(property 2), it is clear that dimA+ dimA/prime≥dimV. However Aand anyA/primeintersect only at the origin (property 3) and are subspaces of V. Thus the union of their bases can span at most V, that is, dimA+ dimA/prime≤dimV. These two inequalities prove the theorem. REMARK. Some people refer to dim A/primeas the codimension ofA(complementary dimension). In this way they avoid mentioning A/primeat all. The last theorem can be written as dimA+ codimA= dimV. A simple result closes the chapter. Theorem 2.16 . IfAis a subspace of VandA/primeis a complement of A, then forX∈V the decomposition X=X1+X2, X 1∈A, X 2∈A/primeis unique. Proof: Assume there are two decompositions, X=X1+X2andX=˜X1+˜X2. Then ˜X1+˜X2=X1+X2or˜X1−X1=X2−˜X2. However the left side of this equation is in Awhile the right is in A/prime. The only element in both AandA/primeis 0. Thus ˜X1=X1and ˜X2=X2. a figure goes here EXERCISES (1) (a) Let A={X∈R2:x1= 0}. Find a basis for Aand extend it to a basis for all ofR2. Use this to define a complement A/primeofA. SketchAandA/prime. Extend the same basis for Ain a different way to a basis for all of R2. Use this to define another complement ˜A/primeofA. Sketch ˜A/prime. (b) Find a basis for the subspace A={X∈R2:x1+x2+x3= 0}. Extend this basis to one for all of R3. Define a complement A/primeofAinduced by this extension. Write X= (−1,0,7) asX=Y1+Y2whereY1∈AandY2∈A/prime. (2) (a) Let A={p∈P2:p(0) = 0 }. Find a basis for Aand extend it to a basis for all of P2. DefineA/primeinduced by this extension. Is the particular polynomial p(x) = 1 +x2inA? inA/prime? Writep(x) asp(x) =q1(x) +q2(x) where q1(x)∈A, q 2(x)∈A/prime. 98 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE (b) LetA={p∈P2:p(1) = 0 }. Find a basis for Aand extend it to a basis for all of P2. (3) LetAbe a subspace of a linear space V. Show by an example that a basis for V need not contain a basis for A. (4) If dimV=nandV=sp{X1,...,X n}, prove that X1,...,X nare linearly inde- pendent. (5) LetV=P4andAthe subspace spanned by 1 ,x2andx4. Find three different subspaces complementary to A(you may specify a subspace by giving a basis for it). After all this about bases, it is probably best to notify you that properties of linear spaces are best defined and proved without introducing a particular basis . As soon as you define a property of a linear space in terms of a basis, you must then prove that the property is intrinsic to the space itself and does not depend upon the basis you choose. We met this problem in defining the dimension in terms of a basis - and were consequently forced to prove Theorem ? which stated that the property really only depended on the space itself, not on the basis chosen. This, in fact, corresponds to one of the major requisites for laws of physics: they should not depend upon the particular coordinate system you choose (picking a coordinate system is equivalent to picking a basis). Moreover, the laws should not depend on the units you choose for each axis of the coordinate system. But these are long, involved questions which must be investigated deeply to make our remarks precise. One should, however, distinguish theoretical issues from computational ones. In theo- retical questions , the rule is never pick a specific basis unless there is no way out . On the other hand, for computational questions you must always pick a basis . Just as in physics, on order to perform any measurements, you must pick some specific coordinate system and specific units. If the theoretical foundations are firm, then you can feel confident that no matter what choice of basis you make, the essential nature of the results will remain unchanged. As an example, let us consider a point Pand two different fixed coordinate systems in the plane of this paper. You should feel that any motion of the point Pcan be described adequately in either coordinate system - and that when the observers in the two coordinate systems get together and discuss the motion of P, they will agree as to what happened. A common example is the meeting of two people from countries using different units of money. Exercises (1) Prove that any n+ 1 elements in a linear space of dimension nmust be linearly dependent. (2) Prove that Pnhas dimension n+ 1 . (3) Since a basis for a linear space of dimension nmust contain exactly nelements, all one must test is that the nelements which are candidates for a basis are lin- early independent - or equivalently that they span the space. Show that the vectors {X1,...,X n}form a basis for Rnif and only if e1,e2,...,e ncan all be expressed as a linear combination of the {X1,...,X n}. 2.4. BASES AND DIMENSION 99 (4) Use Exercise 3 to determine which of the following sets from bases for R3. (a)X1= (1,1,0), X 2= (1,0,1), X 3= (0,1,1). (b)X1= (1,0,1), X 2= (1,1,1). (c)X1= (1,0,1), X 2= (1,1,0), X 3= (0,−1,1). (d)X1= (1,1,1), X 2= (1,2,3), X 3= (17,3,9), X 4= (−2,7,−1). (e)X1= (−1,0,2), X 2= (1,1,1), X 3= (1 2,1 3,−1). (5) Prove that the subspace of functions in C[0,π] which vanish at x= 0 and at x=πis infinite dimensional by showing that the functions f1(x) = sinx, f 2(x) = sin 2x,...,f k(x) = sinkx,... are all linearly independent. [Hint: Assume that 0 = N/summationdisplay k=1aksinkx, for arbitrary Nand show that all the ak’s must be zero by multiplying both sides by sin nxand utilizing the important formula /integraldisplayπ 0sinnxsinkxdx =/braceleftbigg0, k/negationslash=n π 2, k=n./bracerightbigg .] (6) LetC∗[a,b] denote the set of all complex-valued functions f(x) =u(x) +iv(x) which are continuous for x∈[a,b] . The complex number field Cis the field of scalars for C∗. What is the dimension of the subspace A={f∈C∗[−π,π]:f(x) = aeix+be−ix, a, b∈C}? Show that f1(x) = cosxandf2(x) = sinxconstitute a basis forA. [Hint: Use (?) on p. ?]. (7) Which of the following sets of vectors form a basis for R4? (a)X1= (1,0,0,5), X 2= (0,3,2,6), X 3= (0,0,1,2), X 4= (0,0,0,1). (b)X1= (1,6,7,0), X 2= (−2,2,5,0), X 3= (4,5,6,0), X 4= (7,8,3,0). (c)X1= (1,2,5,7), X 2= (4,9,11,8), X 3= (6,3,12,2), X 4= (3,−4,7,6), X5= (0,0,0,1). (d)X1= (1,2,3,4), X 2= (0,2,3,4), X 3= (0,0,3,4), X 4= (0,0,0,4). (8) Find a basis for the following subspaces. (a)A={X∈R2:x1+x2= 0} (b)B={X∈R3:x1+x2+x3= 0} (c)C={p∈P3:p(0) = 0 } (d)D={p∈P3:p(1) = 0 } (e)E={u∈C1[−1,1]:u/prime−u= 0} (f)F={u∈C1[−1,1]:u/prime+ 2u= 0}. 100 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE Chapter 3 Linear Spaces: Norms and Inner Products 3.1 Metric and Normed Spaces Until now we have been contented with being able to add two elements X1andX2of a linear space, and to multiply them by scalars, aX. Since only these algebraic operations have been defined, only algebraic questions could have been raised and answered. Notably absent were any mention of convergence, because the idea of one element of a linear space being “close” to another was not defined. In this chapter we shall introduce a distance ormetric structure into linear spaces. Instead of lingering in the realm of generalities, we shall define metric and norm in this first section and devote the balance of the chapter to a particular kind of metric which generalizes the “Pythagorean distance” of ordinary Euclidean space. Fourier series supply a wonderful and valuable application. Our first notion of distance, that of a metric , makes sense for elements X, Y, Z of an arbitrary set S. The idea is to define the distance d(X,Y) between any two elements of S. This distance is a function which assigns to every pair of points ( X,Y) apositive real numberd(X,Y) called the “distance between XandY”. Definition . LetSbe a non-empty set. A metric onSis a real-valued function d: S×S→R, whereX, Y∈S, which has the three properties: i)d(X,Y)≥0. d (X,Y) = 0⇐⇒X=Y ii) (symmetry) d(X,Y) =d(Y,X) , iii) (triangle inequality) d(X,Z)≤d(X,Y) +d(Y,Z) . Well, they certainly are reasonable requirements for any function we intend to think of as measuring distance. Examples . (1) This first example is trivial but acts as an important check on intuition. With it, you see that every non-empty set can be regarded as a metric space with the following metric d(X,Y) =/braceleftbigg0,ifX=Y 1,ifX/negationslash=Y. A moments reflection will show that this is a metric—but not too useful since it is so coarse. 101 102 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS (2) For the real line, R, with the usual definition of absolute value we define d(X,Y) = |X−Y|, which is clearly a metric. (3) Another less common metric may be given to R. We define d(X,Y) =|X−Y| 1+|X−Y|. Only the triangle inequality is not evident—and that involves some algebra. This metric has the property that the distance between any two points is always less than one,d(X,Y)<1 for allX,Y∈R. (4)Rncan be endowed with many metrics. Let X= (x1,x2,...,x n) ,Y= (y1,...,y n) andZ= (z1,...z n) be arbitrary points in Rn. The metric you most expect is the Euclidean distance d(X,Y) = [(x1−y1)2+...+ (xn−yn)2]1/2= [n/summationdisplay k=1(xk−yk)2]1/2 Again, only the triangle inequality is not obvious. It is a consequence of the Cauchy- Schwarz inequality/parenleftBiggn/summationdisplay k=1xkyk/parenrightBigg2 ≤n/summationdisplay k=1x2 kn/summationdisplay k=1y2 k, (3-1) which in turn is an immediate consequence of the algebraic identity /parenleftBiggn/summationdisplay k=1xkyk/parenrightBigg2 =n/summationdisplay k=1x2 kn/summationdisplay k=1y2 k−1 2n/summationdisplay i=1n/summationdisplay j=1(xiyj−xjyi)2. And now the triangle inequality. Let ak=xk−yk, andbk=yk−zk. Then xk−zk=ak+bk. Thus, using Cauchy- Schwarz in the second line below, we find that [d(X,Z)]2=n/summationdisplay k=1(ak+bk)2=n/summationdisplay k=1a2 k+ 2n/summationdisplay k=1akbk+n/summationdisplay k=1b2 k ≤n/summationdisplay k=1a2 k+ 2/bracketleftBiggn/summationdisplay k=1a2 kn/summationdisplay k=1b2 k/bracketrightBigg1/2 +n/summationdisplay k=1b2 k = /parenleftBiggn/summationdisplay k=1a2 k/parenrightBigg1/2 +/parenleftBiggn/summationdisplay k=1b2 k/parenrightBigg1/2 2 = [d(X,Y) +d(Y,Z)]2,(3-2) so d(X,Z)≤d(X,Y) +d(Y,Z). Another proof of the Schwarz and triangle inequalities for this metric will be given later in the chapter. (5) A second metric for R/multicloseleftis d(X,Y) =n/summationdisplay k=1|xk−yk| The axioms for a metric are easily verified. 3.1. METRIC AND NORMED SPACES 103 (6) A third metric for Rnis d(X,Y) =/bracketleftBiggn/summationdisplay k=1|xk−yk|p/bracketrightBigg1/p , 1≤p<∞. Example 4 is the special case p= 2 , while example 5 is the special case p= 1 . And again, all but the triangle inequality are obvious. However the triangle inequality, called Minkowski’s inequality in this general case, is not simple. We shall not prove it here. Perhaps it will appear as an exercise later. (7) The usual metric for C[a,b] is the uniform metric d(f,g) = max a≤x≤b|f(x)−g(x)|. Geometrically, this distance is the largest vertical distance between the graphs of f andgfor allx∈[a,b] . (8) The space L1[a,b] of functions whose absolute value is integrable has the “natural” metric d(f,g) =/integraldisplayb a|f(x)−g(x)|dx, which can be interpreted as the total area between the two curves. Since every function which is continuous for x∈[a,b] is integrable there, i.e., C[a,b]⊂L1[a,b] , this metric is another metric for C[a,b] . (9) For the function space C1[a,b] , the standard metric is d(f,g) = max a≤x≤b|f(x)−g(x)|+ max a≤x≤b/vextendsingle/vextendsinglef/prime(x)−g/prime(x)/vextendsingle/vextendsingle The metric for Ck[a,b] is defined similarly. There are many theorems one can prove about metric spaces (a metric space is a set Son which a metric is defined). Look in any book on general topology (or point set topology, as it is often called) and you will find more than enough to satisfy you. For most of our purposes metric spaces are too general. Normed linear spaces will suffice. The norm /bardblX/bardblof an element Xin a linear space Vis the “distance” of Xfrom the origin—the 0 element of V. Definition. LetVbe a linear space over the real or complex field. If to every element X∈Vthere is associated a real number /bardblX/bardbl, the norm ofX, which has the three properties i)/bardblX/bardbl ≥0. /bardblX/bardbl= 0⇐⇒X= 0 ii)/bardblaX/bardbl=|a| /bardblX/bardbl(homogeneity), ais a scalar, iii)/bardblX+Y/bardbl ≤ /bardblX/bardbl+/bardblY/bardbl, (triangle inequality), then we say that Vis anormed linear space . How does a norm differ from a metric? First of all, a norm is only defined on a linear space (sinceaXandX+Yappear in the definition) whereas a metric may be defined on any set (cf. example 1 above). But if we 104 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS restrict our attention to linear spaces, how do the concepts of norm and metric differ? Every normed linear space can be made into a metric space in such a way that /bardblX/bardblis indeed the distance of Xfrom the origin ,d(X,0) =/bardblX/bardbl. The explicit formula for d(X,Y) should surprise no one d(X,Y) =/bardblX−Y/bardbl. It is easy to check that d(X,Y) is a metric. Thus every normed linear space has a “natural” metric induced upon it. However, a linear space which has a metric need not be a normed linear space . For example in R, the linear space of the real numbers, the metric of example 3 d(X,Y) =|X−Y| 1 +|X−Y| is not associated with a norm because axiom ii) for a norm is not satisfied. Of the examples considered earlier, all but the first and third metrics arise from norms, in the sense that d(X,Y) =d(X−Y,0) =/bardblX−Y/bardbl. By far the most common norm in Rnis that given by the Pythagorean theorem (ex- ample 4). Then /bardblX/bardbl=/radicalBig x2 1+x2 2+···+k2n=/parenleftBiggn/summationdisplay k=1x2 x/parenrightBigg1/2 and the induced metric is d(X,Y) =/bardblX−Y/bardbl=/parenleftBiggn/summationdisplay k=1(xk−yk)2/parenrightBigg1/2 For obvious historical reasons, we shall refer to R2with this Pythagorean norm as Eu- clideann-space , and denote it by En. Note that Enis a linear space with a particular way of measuring length specified. A metric removes the floppiness from Rn, giving the addi- tional structure needed to investigate those geometrical concepts which utilize the notion of distance. Once we have a norm (or metric) it becomes possible to discuss convergence of a se- quence of elements. Definition : IfVis a normed linear space, the sequenceXn∈Vconverges toX∈Vif, given any/epsilon1>0 , there in an Nsuch that /bardblXn−X/bardbl</epsilon1 for alln>N. As an example, we shall prove the sample Theorem 3.1 . A sequence of points Xn= (x(n) 1,x(n) 2,...x(n) k)inEkconverges to the pointX= (x1,...,x k)inEkif and only if each component x(n) jconverges to its respective limit, lim n→∞x(n) j=xj,j= 1,...,k . Proof: i)Xn→X⇒x(n) j→xj. This is a consequence of the trivial inequality /vextendsingle/vextendsingle/vextendsinglex(n) j−xj/vextendsingle/vextendsingle/vextendsingle≤/radicalBig (x(n) 1−x1)2+···+ (x(n) k−xk)2=/bardblXn−X/bardbl; 3.1. METRIC AND NORMED SPACES 105 for if /bardblXn−X/bardbl< /epsilon1forn > N , then/vextendsingle/vextendsingle/vextendsinglex(n) j−xj/vextendsingle/vextendsingle/vextendsingle< /epsilon1forn > N too. Thus x(n) j→xj. If the subscripts are cluttering up the proof, go through it again in a special case, say x(n) 2→x2. ii)x(n) j→xj⇒Xn→X. By hypothesis, given any /epsilon1 > 0 , there are numbers N1,N2,...N ksuch that/vextendsingle/vextendsingle/vextendsinglex(n) 1−x1/vextendsingle/vextendsingle/vextendsingle< /epsilon1, for alln > N 1,/vextendsingle/vextendsingle/vextendsinglex(n) 2−x2/vextendsingle/vextendsingle/vextendsingle< /epsilon1 for alln > N2,...,/vextendsingle/vextendsingle/vextendsinglex(n) k−xk/vextendsingle/vextendsingle/vextendsingle< /epsilon1 for alln > N k. PickN= max(N1,N2,...,N k) . ThisNwill work for all the x(n) j’s, that is, for every j, /vextendsingle/vextendsingle/vextendsinglex(n) j−xj/vextendsingle/vextendsingle/vextendsingle</epsilon1 for alln>N. Thus /bardblXn−X/bardbl=/radicalBig (x(n) 1−x1)2+...+ (x(n) k−xk)2 </radicalbig /epsilon12+...+/epsilon12=/epsilon1√ k,for alln>N.(3-3) Sincekis a fixed finite number, this shows that /bardblXn−X/bardblmay be made arbitrarily small by picking nbig enough, so Xndoes converge to X. Example . In E4, the sequence Xn= (n n+1,2,−1 n,0) converges to X= (1,2,0,0) since n n+1→1,2→2,−1 n→0 , and 0 →0 . A useful elementary result is Theorem 3.2 . IfVis a normed linear space, and if Xn→X, Y n→YinV, then for any scalars aandb, aX n+bYn→aX+bY. Proof: There are essentially no changes from the case of R1. We must show that /bardblaXn+ bYn−aX−bY/bardblcan be made arbitrarily small by picking nlarge enough. One application of the triangle inequality /bardblaXn+bYn−aX−bY/bardbl ≤ /bardblaXn−aX/bardbl+/bardblbYn−bY/bardbl, and the homogeneity of a norm, yields ≤ |a| /bardblXn−X/bardbl+|b| /bardblYn−Y/bardbl. BecauseXn→XandYn→Y, ifn>N 1, then /bardblXn−X/bardbl</epsilon1. Also, ifn>N 2, then /bardblYn−Y/bardbl</epsilon1. PickN= max(N1,N2) . Thus /bardblaXn+bYn−aX−bY/bardbl<|a|/epsilon1+|b|/epsilon1= (|a|+|b|)/epsilon1, n>N, and the desired convergence is proved. For a given linear space V, there may be two (or even more) norms defined, say /bardbl /bardbl and/bardbl /bardbl 1to distinguish them. Why carry them both around? First of all, a sequence may converge in one norm and not in the other. Second, even if both norms yield the same convergent sequences, one norm may be more convenient in some particular computation. Example . Consider the linear space C[−1,1] of functions f(x) continuous for x∈[−1,1] , with the two norms (Examples 7 and 8) /bardblf/bardbl∞= max −1≤x≤1|f(x)|;/bardblf/bardbl1=/integraldisplay1 −1|f(x)|dx, 106 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS that is, the uniform norm and the L1norm. We shall exhibit a sequence of functions which converge in the second norm but not in the first. Let fn(x) be fn(x) =  0, x ∈[−1,−1 n2] n3(x+1 n2)x∈[−1 n2,0] −n3(x−1 n2)x∈[0,1 n2] 0 x∈[1 n2,1] Then by inspection from the graph ( /bardbl /bardbl 1is the area), we see that /bardblfn/bardbl∞=n, and /bardblfn/bardbl1=1 n. Asn→ ∞,/bardblfn−0/bardbl1→0 so thatfn→0 in theL1norm. On the other hand, /bardblfn/bardbl∞→ ∞ so the limit does not exist in the uniform norm. If you look at the graph,fnis zero except for a spike in the interval [ −1 n2,1 n2] . Asn→ ∞ , the function is zero in essentially the whole interval, except for the bit around the origin where it blows up—but it blows up slowly enough that the area under the curve tends to zero. However, we can prove the Theorem 3.3 . Letfnandfbe continuous functions, n= 1,2,.... Iffn→fin the uniform norm, then also fn→fin theL1norm. Remark . We have just seen that the converse is false. Proof: An immediate consequence of the Lemma 3.4 /bardblfn−f/bardbl1≤(b−a)/bardblfn−f/bardbl∞ Proof: /bardblfn−f/bardbl1=/integraldisplayb n|fn(x)−f(x)|dx≤/integraldisplayb a/bardblfn−f/bardbl∞dx =/bardblfn−f/bardbl/integraldisplayb adx= (b−a)/bardblfn−f/bardbl∞(3-4) Exercises (1) In the set Z, defined(m,n) =|m−n|where |x|is ordinary absolute value. Prove thatd(m,n) is a metric. (2) Suppose that d1(X,Y) andd2(X,Y) are both metrics for a set S, whereX,Y∈S. a). Show that [ d1(X,Y)]2isnot, in general, a metric. b). Prove that d1+d2and/radicalbig d2 1+d2 2are also metrics for S. (3) Prove that the function d(X,Y) =|X−Y| 1+|X−Y|, X, Y ∈R, is a metric, but that it is not a norm on R. (4) LetX= (x1,...,x k)∈Rk. Define /bardblX/bardbl∞= max 1≤l≤k|xl|. (a) Prove that /bardblX/bardbl∞is a norm for Rk, and write down the induced metric. (b) Let /bardblX/bardbl1=k/summationdisplay l=1|xl|, and /bardblX/bardbl2=/parenleftBiggk/summationdisplay l=1|xl|2/parenrightBigg1/2 . Prove /bardblX/bardbl∞≤ /bardblX/bardbl2≤ /bardblX/bardbl1≤k/bardblX/bardbl∞. 3.2. THE SCALAR PRODUCT IN E2107 (c) Consider the sequence Xn= (1−1 n,−7,1 n2) inR3. In which of the norms /bardbl /bardbl ∞,/bardbl /bardbl 2,/bardbl /bardbl 1does it converge, and to what? (5) LetXnbe a sequence of elements in a normed linear space V(not necessarily finite dimensional). Prove that if Xn→X, then the sequence Xnis bounded in norm (a sequenceXnin a normed linear space is bounded if there is an M∈Rsuch that /bardblXn/bardbl ≤Mfor alln). [Hint: Compare with Theorem 6, page ??]. (6) Compute the /bardbl /bardbl 1,/bardbl /bardbl 2,and/bardbl /bardbl ∞(cf. Ex. 4) norms of the following vectors in R3. a)X= (1,2,2) , b)Y= (2,−2,1) , c)Z= (0,3,−4) , d)W= (0,−1,0) . (7) Compute the /bardbl /bardbl 1,/bardbl /bardbl 2,and/bardbl /bardbl ∞norms of the following functions for the interval [−1,1] . a)f(x) =−2x+ 3 b) g(x) = sinπx, c) hn(x) =xn d) Asn→ ∞ , does the sequence hnconverge in any of these three norms? (8) Which of the following define norms for the given linear spaces? a) For R3,[X] =x2 1+x2 x+x2 3 b) For P3,[p] = max 0≤x≤1p(x) c) For P3,[p] = max 0≤x≤1|p(x)| d) For R3,[X] =|x1|+|x2| e) For R4,[X] =/radicalbig 1 +x2 1+x2 2+x2 3+x2 4. (9) Prove that [ X] =/radicalbig x2 1+x2 2defines a norm for R2(some algebraic fortitude will be needed to prove the triangle inequality). 3.2 The Scalar Product in E2 In Euclidean space E2—which we remind you is R2with the Euclidean norm /bardblX/bardbl=/radicalbig x2 1+x2 2—one can introduce many geometric concepts and develop a corresponding geo- metric theory. Most important of these concepts is that of angle—especially orthogonality (perpendicularity). It turns out that these ideas generalize almost immediately to all En, and even to some exceedingly important infinite dimensional spaces. This section is devoted to the most simple situation: E2. Please look at the pictures. To begin, we introduce the scalar product (also called the dot product , orinner product ) of two vectors XandY. Definition: If XandYare two vectors in E2, their scalar product /angbracketleftX, Y/angbracketright(sometimes writtenX·Y) is defined by /angbracketleftX, Y/angbracketright=/bardblX/bardbl /bardblY/bardblcosθ, whereθis the angle between XandY. Notice that the scalar product of two vectors is a real number, a scalar, notanother vector. We need not specify the direction in which θis measured, counterclockwise or clockwise, since cos θ= cos( −θ) . Further, we can use either the acute or obtuse angle 108 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS betweenXandYsince cos(2 π−θ) = cosθ. It is important to observe that the scalar product of two vectors is defined independent of any coordinate system. We are immediately led to some simple consequences. Lemma 3.5 . Two vectors XandYare orthogonal if and only if /angbracketleftX, Y/angbracketright= 0 Proof: IfXandYare orthogonal, the angle θbetween them isπ 2, so/angbracketleftX, Y/angbracketright= /bardblX/bardbl /bardblY/bardblcosπ 2= 0 . In the other direction, if /angbracketleftX, Y/angbracketright= 0 , then /bardblX/bardbl /bardblY/bardblcosθ= 0 . If neither /bardblX/bardblnor/bardblY/bardbl= 0 , then cos θ= 0 , that is θ=π 2or3π 2. ThusXis orthogonal toY. If/bardblX/bardblor/bardblY/bardbl= 0 , then one of them is just the point at the origin, the zero vector. We agree to say that the zero vector is orthogonal to every other vector. With this agreement, /angbracketleftX, Y/angbracketright= 0⇒X⊥Y, and the second half of the theorem is proved too. There is a nice geometric interpretation of the scalar product. A hint of it appeared in our last lemma. Let ebe a unit vector ,/bardble/bardbl= 1 . Consider /angbracketleftX, e/angbracketright=/bardblX/bardblcosθ(see figure). This is the length of the projection ofXin the direction of e, or in other words, the length of the projection of Xinto the subspace spanned by the single vector e. Strictly /angbracketleftX, e/angbracketrightis not really a “length”, since “length” carries the implication of being positive, whereas the real number /angbracketleftX, e/angbracketrightwill be negative if the projection “points” in the direction opposite to e. We shall, however, allow ourselves this abuse of language. The vector U1which is the projection of Xinto the subspace spanned by eisU1=/angbracketleftX, e/angbracketrighte. IfYis a (non-zero) vector in E2which is not a unit vector, the above geometric idea goes through by making the simple observation that given any vector Y/negationslash= 0 , the vector e=Y//bardblY/bardblis a unit vector in the direction of Y. Now you are certainly wondering how in the world we compute this scalar product. You could take out your ruler, protractor and table of cosines—but we will present a more convenient method. In order to compute this as is always the case, a particular basis must be chosen. Then the vectors XandYcan be given explicitly in terms of the basis. Since we want to show that the concepts are independent of any particular basis , you must relax and be patient. Only after the theory has been exposed will we reveal how to compute in terms of a given basis. Theorem 3.6 (a)/angbracketleftX, X/angbracketright=/bardblX/bardbl2 (b)/angbracketleftX, Y/angbracketright=/angbracketleftY, X/angbracketright (c)/angbracketleftaX, Y /angbracketright=a/angbracketleftX, Y/angbracketrightwherea∈R. (d)/angbracketleftX, aY /angbracketright=a/angbracketleftX, Y/angbracketright, wherea∈R. (e)/angbracketleftX+Y, Z/angbracketright=/angbracketleftX, Z/angbracketright+/angbracketleftY, Z/angbracketright (f)/angbracketleftX, Y +Z/angbracketright=/angbracketleftX, Y/angbracketright+/angbracketleftX, Z/angbracketright (g)|/angbracketleftX, Y/angbracketright| ≤ /bardblX/bardbl /bardblY/bardbl(Cauchy- Schwarz inequality) Proof: (a) Obvious since θ= 0 and cos 0 = 1 . (b) Obvious since cos( −θ) = cosθ. 3.2. THE SCALAR PRODUCT IN E2109 (c) The vectors XandaXlie along the same line through the origin. There are two cases,a>0 anda<0 (a= 0 is trivial). Ifa>0 , the angle θbetweenXandY is identical to that between aXandY. Since /bardblaX/bardbl=a/bardblX/bardblfora>0 , this case is proved, for /angbracketleftaX, Y /angbracketright=/bardblaX/bardbl /bardblY/bardblcosθ=a/angbracketleftX, Y/angbracketright. Ifa<0 , thenaXpoints in the direction opposite to X. Thus the angle θ1between aXandYequalsπ−θ, whereθis the angle between XandY. The following computation completes the proof: /angbracketleftaX, Y /angbracketright=/bardblaX/bardbl /bardblY/bardblcosθ1=|a|/bardblX/bardbl/bardblY/bardblcos(π−θ) =−|a| /bardblX/bardbl /bardblY/bardblcosθ=a/bardblX/bardbl /bardblY/bardblcosθ=a/angbracketleftX, Y/angbracketright(3-5) (d) By (b) and (c) and (b) again we are done /angbracketleftX, aY /angbracketright=/angbracketleftaY, X /angbracketright=a/angbracketleftY, X/angbracketright=a/angbracketleftX, Y/angbracketright. (e) This is the most subtle part. We shall rely on the interpretation of the scalar product /angbracketleftU, e/angbracketrightas the length of the projection of Uin the subspace spanned by e. First, let e=Z//bardblZ/bardblbe the unit vector in the direction of Z. We shall show that /angbracketleftX+Y, e/angbracketright= /angbracketleftX, e/angbracketright+/angbracketleftY, e/angbracketright. A picture is all that is needed now. Two situations are illustrated, where both XandYare on the same side of the line perpendicular to eand a figure goes here whereXandYare on opposite sides of that line. The vector X+Yis found from XandYby the parallelogram rule for addition. Interpreting the scalar product of a vector with eas the length of the projection into the subspace (line) spanned by e, we see (look) that we must prove → OP=→ OQ+→ OM . But since→ OAand→ BC are on opposite sides of the same parallelogram, know that → OM=→ QPboth in magnitude and direction. The natural substitution yields → OP=→ OQ+→ QP, which is indeed all we desired. Thus /angbracketleftX+Y, e/angbracketright=/angbracketleftX, e/angbracketright+/angbracketleftY, e/angbracketright To prove the general result for Z=/bardblZ/bardble, multiply the last equation by /bardblZ/bardbl, which is a scalar. Then by part a we find /angbracketleftX+Y,/bardblZ/bardble/angbracketright=/angbracketleftX,/bardblZ/bardble/angbracketright+/angbracketleftY,/bardblZ/bardble/angbracketright, or /angbracketleftX+Y, Z/angbracketright=/angbracketleftX, Z/angbracketright+/angbracketleftY, Z/angbracketright, We are done. 110 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS (f) By parts (b), (e) and (b) again we obtain the result. /angbracketleftX, Y +Z/angbracketright=/angbracketleftY+Z, X/angbracketright=/angbracketleftY, X/angbracketright+/angbracketleftZ, X/angbracketright=/angbracketleftX, Y/angbracketright+/angbracketleftX, Z/angbracketright (g) Obvious since |cosθ| ≤1 . It is evident that equality occurs when and only when cosθ±1 , that is, when XandYlie along the same line (possibly pointing in opposite directions). Ifeis a unit vector, we know how to find the projection U1of a given vector Xinto the subspace spanned by e, it isU1=/angbracketleftX, e/angbracketrighte. Similarly, if Yis any vector—not necessarily of length one, then since Y//bardblY/bardblis a unit vector in the direction of Y, the projection of X into the subspace spanned by Yis/angbracketleftX, Y/ /bardblY/bardbl/angbracketrightY//bardblY/bardbl=/angbracketleftX, Y/angbracketrightY//bardblY/bardbl2. We can also find the projection U2ofXinto the subspace orthogonal to the unit vector e. Since the sum ofU1andU2must add up to X, X =U1+U2, we find that U2=X−U1=X−/angbracketleftX, e/angbracketrighte. Thus, we have proved Theorem 3.7 . IfXandYare any two vectors, /bardblY/bardbl /negationslash= 0, thenXcan be decomposed into two vectors U1andU2, X=U1+U2such thatU1is in the subspace spanned by Y andU2is in the orthogonal subspace. The decomposition is given by U1=/angbracketleftX, Y/angbracketrightY//bardblY/bardbl2 andU2=X− /angbracketleftX, Y/angbracketrightY//bardblY/bardbl2, so that X=/angbracketleftX, Y/angbracketrightY /bardblY/bardbl2+ (X− /angbracketleftX, Y/angbracketrightY /bardblY/bardbl2). Without further delay, we shall show how to compute the scalar product of two vectors. In order to carry this out we must introduce a basis. Let X1andX2be any two vectors in E2which span E2. Then every vector X∈E2can be written in the form X=a1X1+a2X2, where the scalars a1anda2are determined uniquely by the vector X. Now it is most convenient to have a basis whose vectors are i) orthogonal to each other and ii) have unit length. Such a basis is called an orthonormal basis (orthogonal and normalized to have unit length). In other words e1ande2are an orthonormal basis for E2if/bardblej/bardbl= 1 and /angbracketlefte1, e2/angbracketright= 0 . This requirement is most conveniently stated by introducing the Kronecker symbolδjk δjk=/braceleftbigg0j/negationslash=k 1j=k. Then the orthonormality property reads /angbracketleftej, ek/angbracketright=δjk, j,k = 1,2 . The notation is perhaps excessive for this simple case, but will really be useful in our generalizations. Therefore, let e1ande2be an orthonormal basis for E2, so that if X∈E2, X= x1e1+x2e2.Fixthe basis throughout the ensuing discussion. Observe that x1andx2 can be computed in terms of X, and the basis vectors e1ande2, viz/angbracketleftX, e 1/angbracketright=/angbracketleftx1e1+ x2e2, e1/angbracketright=x1/angbracketlefte1, e1/angbracketright+x2/angbracketlefte2, e1/angbracketright=x1, since /angbracketlefte1, e1/angbracketright= 1 and /angbracketlefte1, e2/angbracketright= 0 . Similarly, /angbracketleftX, e x/angbracketright=x2. Thus we have proved Theorem 3.8 . If{ej}, j = 1,2, form an orthonormal basis for E2, then every vector X∈E2can be written as X=2/summationdisplay j=1xjej, wherexjis the length of the projection of X into the subspace spanned by ej, x j=/angbracketleftX, e j/angbracketright. 3.2. THE SCALAR PRODUCT IN E2111 IfX=x1e1+x2e2andY=y1e1+y2e2are any two vectors in E2, then /angbracketleftX, Y/angbracketright=/angbracketleftx1e1+x2e2, y1e1+y2e2/angbracketright =/angbracketleftx1e1+x2e2, y1e1/angbracketright+/angbracketleftx1e1+x2e2, y2e2/angbracketright =/angbracketleftx1e1, y1e1/angbracketright+/angbracketleftx2e2, y1e1/angbracketright+/angbracketleftx1e1, y2e2/angbracketright+/angbracketleftx2e2,, y 2/angbracketright =x1y1/angbracketlefte1, e1/angbracketright+x2yx/angbracketlefte2, e1/angbracketright+x1y2/angbracketlefte1, e2/angbracketright+x2y2/angbracketlefte2, e2/angbracketright =x1y1+ 0 + 0 +x2y2=x1y1+x2y2. Now you see how easy it is to compute the scalar product of XandYin terms of the representation from an orthonormal basis. Let us rewrite our result formally. Theorem 3.9 . Let {ej}, j= 1,2, form an orthonormal basis for E2. IfX=2/summationdisplay j=1xjej andY=2/summationdisplay j=1vjej, then /angbracketleftX, Y/angbracketright=2/summationdisplay j=1xjyj=x1y1+x2y2. Some numerical examples should reassure you of the basic simplicity of the computation. As our orthonormal basis in E2, we choose the vectors e1= (1,0) ande2= (0,1) . These both have unit length, and are perpendicular (one is on the horizontal axis, the other on the vertical axis). Let X= (−2,3) . ThenX=−2e1+ 3e2. Notice that −2e1and 3e2 are exactly the projections of Xinto the subspaces spanned by e1ande2respectively. If Y= (1,−2) , then our theorem shows that /angbracketleftX, Y/angbracketright= (−2)(−1) + (3)( −2) =−2−6 =−8. From this computation we can reverse the geometric procedure and find the angle θbetween XandY, for we know the formula cosθ=/angbracketleftX, Y/angbracketright /bardblX/bardbl /bardblY/bardbl. In this example, /angbracketleftX, Y/angbracketright=−8,/bardblX/bardbl=√4 + 9 =√ 13 and /bardblY/bardbl=√1 + 4 =√ 5 . Thus θ= cos−1(−8√ 65) which can be evaluated by consulting your favorite numerical tables. It is equally simple to check if two vectors are orthogonal. Let X= (2,−3) and Y= (6,4) . Then /angbracketleftX, Y/angbracketright= (2)(6) + ( −3)(4) = 0 ; consequently XandYare orthogonal. Another consequence is the law of cosines. Let X= (x1,x2) andY= (y1,y2) . Then from the parallelogram construction, the length of the segment joining the tip of Xto the tip ofYhas length /bardblY−X/bardbl. But /bardblY−X/bardbl2=/angbracketleftY−X, Y−X/angbracketright =/bardblX/bardbl2+/bardblY/bardbl2−2/angbracketleftX, Y/angbracketright =/bardblX/bardbl2+/bardblY/bardbl2−2/bardblX/bardbl /bardblY/bardbl??θ. One more example. We shall find the distance of the point P= (−3,2) from the coset A={X= (x1,x2)∈E2:x1−2x2= 2}. Pick some point in X0inA, sayX0= (3,1 2) . The distance dfromPtoAis then the length of the projection of the segment X0P onto a line lorthogonal to A. First of all, we can replace the segment X0Pby a vector from the origin 0 to the point Q=P−X0= (−6,3 2) , for the length of the projection of ¯0Qonto a line lorthogonal to Ais equal to the length of the projection of ¯X0Ponto 112 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS l(see figure). Now we have the vector Q= (−6,3 2) ; all we need to compute the desired projection is another vector Northogonal to A, for thend=|/angbracketleftY, N//bardblN/bardbl/angbracketright|. To find a vector Northogonal to A, we realize that Nwill also be orthogonal to the subspace Sparallel to the coset AsoA=S+X0, where S={X= (x1,x2)∈ E2:x1−2x2= 0}. IfN= (n1,n2) andXis any element of S, sinceN⊥S, we must have 0 = /angbracketleftX, N/angbracketright=x1n1+x2n2. However X∈Ssox1−2x2= 0 . We want the equation x1n1+x2n2= 0 to hold for allpoints onx1−2x2= 0 , that is for all X∈S. This is only possible if n1= 1·candn2=−2·c, wherecis any constant. Thus N=c(1,−2) and /bardblN/bardbl=|c|√ 5 . The distance dbetween the point Pand the coset Ais then d=|/angbracketleftY, N//bardblN/bardbl/angbracketright|=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/angbracketleft(−6,3 2),c |c|√ 5(1,−2)/angbracketright/vextendsingle/vextendsingle/vextendsingle/vextendsingle =/vextendsingle/vextendsingle/vextendsingle/vextendsingle(−6)(c |c|√ 5) + (3 2)(−2c |c|√ 5)/vextendsingle/vextendsingle/vextendsingle/vextendsingle=9√ 5.(3-6) This example contained a plethora of ideas. It would be wise to go through it again and list the constructions and concepts used. The exercises will develop many of them in greater generality. Now you should try some problems on your own. Exercises (1) IfX= (3,4) andY= (5,−12) are two points in E2, find the angle between→ OX and→ OY, where 0 is the origin. (2) IfX= (3,−4) andY= (5,12) are two vectors in E2, find vectors U1∈span(Y) andU2orthogonal to span( Y) such that X=U1+U2. (3) Show that the vector N= (a1,a2) is perpendicular to the straight line whose equation isa1x1+a2x2=c(you will have to supply the natural definition of what it means for a vector to be perpendicular to a straight line). (4) (a) Find the distance of the point P= (2,−1) from the coset A={X∈E2:x1+ x2=−2}. (b) Find the distance between the two “parallel” cosets Adefined above and B= {X∈E2:x1+x2= 1}. (Hint: Draw a figure and observe that P∈B) . (5) (a) Prove that the distance dof the point P= (y1,y2) from the coset A={X∈ E2:a1x1+a2x2=c}is given by d=|a1y1+a2y2−c|/radicalbig a2 1+a2 2. (b) Prove that the distance dbetween the two “parallel” cosets A={X∈E2:a1x1+ a2x2=c1}, andB={X∈E2:a1x1+a2x2=c2}is given by d=|c1−c2|/radicalbig a2 1+a2 2. (Hint: If you use part (a) and are cunning, the derivation takes but one line). 3.3. ABSTRACT SCALAR PRODUCT SPACES 113 (6) (a) If it is known that /angbracketleftX, Y 1/angbracketright=/angbracketleftX, Y 2/angbracketright, and that /bardblX/bardbl /negationslash= 0 for a fixedX, can you “cancel” Xfrom both sides and conclude that Y1=Y2? Reason? (b) If it is known that /angbracketleftX, Y/angbracketright= 0 for everyX, can you conclude that Y= 0 ? Reason? (c) If it is known that /angbracketleftX, Y 1/angbracketright=/angbracketleftX, Y 2/angbracketrightforeveryX, can you conclude that Y1=Y2? Reason? (7) (a) Show that the vector Z=/bardblX/bardblY+/bardblY/bardblX /bardblX/bardbl+/bardblY/bardbl. bisects the angle between the vectors XandY. (b) Show that the vector /bardblX/bardblY+/bardblY/bardblXis perpendicular to the vector /bardblY/bardblX− /bardblX/bardblY. (8) Express the angle between an edge and a diagonal of a rectangle in terms of the scalar product. (9) Let two of the sides of a parallelogram be given by the vectors XandY. The parallelogram theorem states that the sum of the squares of the sides is equal to the sum of the squares of the diagonals, that is, /bardblX+Y/bardbl2+/bardblX−Y/bardbl2= 2/bardblX/bardbl2+ 2/bardblY/bardbl2. Prove this in two ways: i) using elementary geometry, and ii) using only the fact that XandYare elements of a linear space, and the properties of the scalar product contained in Theorem 4 (using 4a to define /bardbl /bardbl ). (10) LetXbe any vector in E2, and letebe a unit vector. Define the vector U=ae, wherea=/angbracketleftX, e/angbracketrightis the length of the projection of Xinto the subspace spanned by e, andV=αe, whereαis any scalar. Prove that /bardblX−V/bardbl2≥ /bardblX−U/bardbl2=/bardblX/bardbl2− /bardblU/bardbl2=/bardblX/bardbl2−a2. This shows that in the subspace spanned by e, the vector closest to Xis the projec- tionUofXinto that subspace. (11) IfXis orthogonal to Y, prove the Pythagorean theorem /bardblX+Y/bardbl2=/bardblX/bardbl2+/bardblY/bardbl2 using only /bardblV/bardbl2=/angbracketleftV, V/angbracketrightand the properties of a scalar product in Theorem 4. (12) LetXandYbe orthogonal elements of E2, with neither /bardblX/bardblnor/bardblY/bardblzero. Prove thatXandYare linearly independent. Do notintroduce a basis. 3.3 Abstract Scalar Product Spaces We shall turn the tables around. Whereas in the last section we defined the scalar product geometrically and deduced its properties, in this section we define a scalar product space as a linear space upon which a scalar product is defined, and the scalar product is stipulated to have the properties deduced earlier. After presenting our abstract definition, we shall give examples—other than E2—of scalar product spaces. 114 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS Definition. A linear space His called a real scalar product space if to every pair of elementsX,Y∈His associated a real number /angbracketleftX, Y/angbracketright,the scalar product of XandY, which has the properties 1./angbracketleftX, X/angbracketright ≥0 with equality if and only if X= 0 . 2./angbracketleftX, Y/angbracketright=/angbracketleftY, X/angbracketright 3./angbracketleftaX, Y /angbracketright=a/angbracketleftX, Y/angbracketright, a∈R 4./angbracketleftX+Y, Z/angbracketright=/angbracketleftX, Z/angbracketright+/angbracketleftY, Z/angbracketright You should observe that the scalar product in E2does have these properties (Theorem 4). Using E2as our model, it is natural to define /bardblX/bardbl=/radicalbig /angbracketleftX, X/angbracketrightand suspect that /bardbl /bardbl is indeed a norm on the linear space H. This is true, but proving the triangle inequality for this norm using only properties 1-4 will take some work. We shall do just that after presenting Examples (1) LetX= (x1,...,x n) andY= (y1,...,y n) be points in the linear space R2. We define /angbracketleftX, Y/angbracketright=x1y1+x2y2+...+xnyn. Only easy algebra is needed to verify that the real number /angbracketleftX, Y/angbracketrightsatisfies all of the properties of a scalar product. It turns out (after we prove the triangle inequality) that the natural norm /bardblX/bardbl=/radicalbig /angbracketleftX, X/angbracketrightis the Euclidean norm, so this is E2. (2) This example is the first hint that our abstractions are fruitful. Let the functions f(x) andg(x) be points in the linear function space C[a,b] of real-valued functions continuous for a≤x≤b. We define /angbracketleftf, g/angbracketright=/integraldisplayb af(x)g(x)dx. You might be surprised; in any event let us verify that the real number /angbracketleftf, g/angbracketrightasso- ciated with the pair of functions fandgdoes satisfy the four properties of a scalar product. (i)/angbracketleftf, f/angbracketright=/integraltextb af2(x)dx. This is clearly non-negative and f= 0 implies that /angbracketleftf, f/angbracketright= 0 . All we must show is that if /angbracketleftf, f/angbracketright=/integraltextb af2(x)dx= 0 , then f= 0 . By contradiction, assume f(x)/negationslash= 0 . Then there is some point x0∈[a,b] such thatf(x0) =c/negationslash= 0 . Thus f2(x0) =c2>0 . Sincef—and hence f2—is continuous, this means that f2is positive in some interval about x0(p. 29b, Theorem I), so that/integraltextb af2(x)dx> 0 , the desired contradiction. (ii)/angbracketleftf, g/angbracketright=/integraltextb af(x)g(x)dx=/integraltextb ag(x)f(x)dx=/angbracketleftg, f/angbracketright. (iii)/angbracketleftαf, g/angbracketright=/integraltextb aαf(x)g(x)dx=α/integraltextb af(x)g(x)dx=α/angbracketleftf, g/angbracketright, whereα∈R. (iv)/angbracketleftf+g, h/angbracketright=/integraltextb a(f(x) +g(x))h(x)dx =/integraltextb af(x)h(x)dx+/integraltextb ag(x)h(x)dx =/angbracketleftf, h/angbracketright+/angbracketleftg, h/angbracketright. There. We did it. After we prove the triangle inequality for an abstract scalar product space, the natural candidate for a norm /bardblf/bardblis a norm: /bardblf/bardbl=/radicalBigg/integraldisplayb af2(x)dx. 3.3. ABSTRACT SCALAR PRODUCT SPACES 115 I like this space very much. You will be meeting it often, becoming much more intimate with its finer features. We shall—somewhat improperly—refer to this linear space with the given scalar product as L2[a,b] . The name is improper sinceL2[a,b] is customarily used for our space but with more general functions and an extended notion of integration. (3) Letf(x) andg(x) be inC[0,∞] . This time define /angbracketleftf, g/angbracketright=/integraldisplay∞ 0f(x)g(x)e−xdx. Sincee−xis continuous and positive for all x, we are assured that /angbracketleftf, f/angbracketright ≥0 , with equality if and only if f= 0 . The other properties of an inner product follow from simple manipulations. Do them. Remark : Complex scalar product spaces are defined similarly. For them, /angbracketleftX, Y/angbracketright may be a complex number, and complex scalars are admitted. The only change in the axioms is that property 2 is dropped in favor of ¯2./angbracketleftY, X/angbracketright=¯/angbracketleftX, Y/angbracketright, where the bar means take the complex conjugate of the complex number /angbracketleftX, Y/angbracketright. Since we shall not develop the theory far enough, our attention henceforth will be restricted to real scalar product spaces. The first order of business is to prove that the natural candidate for a norm /bardblX/bardbl=/radicalbig /angbracketleftX, X/angbracketrightis in fact a norm for the linear space V. Only properties 1-4 may be used. (1)/bardblX/bardbl ≥0 , with equality if and only if X= 0 . This follows immediately from the corresponding property of /angbracketleftX, X/angbracketright. (2)/bardblaX/bardbl=|a| /bardblX/bardbl. For /bardblaX/bardbl=/radicalbig /angbracketleftaX, aX /angbracketright=/radicalbig a2/angbracketleftX, X/angbracketright=|a|/radicalbig /angbracketleftX, X/angbracketright= |a| /bardblX/bardbl. The proof of the triangle inequality (3)/bardblX+Y/bardbl ≤ /bardblX/bardbl+/bardblY/bardblinvolves more labor. We shall first need to prove the Cauchy- Schwarz inequality (cf. Theorem 4,g). Theorem 3.10 (Cauchy-Schwarz inequality). |/angbracketleftX, Y/angbracketright| ≤ /bardblX/bardbl/bardblY/bardbl. Proof: If either /bardblX/bardblor/bardblY/bardblis zero, this is immediate. Thus, assume that neither /bardblX/bardbl nor/bardblY/bardblis zero and define U=X /bardblX/bardbl, V =Y /bardblY/bardbl, so that both UandVare unit vectors, /bardblU/bardbl=/bardblV/bardbl= 1 . Then 0≤ /bardblU±V/bardbl2=/angbracketleftU±V, U±V/angbracketright =/angbracketleftU, U/angbracketright ± /angbracketleftU, V/angbracketright ± /angbracketleftV, U/angbracketright+/angbracketleftV, V/angbracketright =/bardblU/bardbl2±2/angbracketleftU, V/angbracketright+/bardblV/bardbl2,. Since /bardblU/bardbl= 1 and /bardblV/bardbl= 1 , this shows ±/angbracketleftU, V/angbracketright ≤1 . Substituting for UandV, we obtain the inequality sought: |/angbracketleftX, Y/angbracketright| ≤ /bardblX/bardbl/bardblY/bardbl. 116 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS Theorem 3.11 (Triangle inequality) /bardblX+Y/bardbl ≤ /bardblX/bardbl+/bardblY/bardbl. Proof: This is identical to that given in section 1. /bardblX+Y/bardbl2=/angbracketleftX+Y, X +Y/angbracketright= /bardblX/bardbl2+ 2/angbracketleftX, Y/angbracketright+/bardblY/bardbl2. By Cauchy-Schwarz, /angbracketleftX, Y/angbracketright ≤ /bardblX/bardbl /bardblY/bardbl, so /bardblX+Y/bardbl2≤ /bardblX/bardbl2+ 2/bardblX/bardbl /bardblY/bardbl+/bardblY/bardbl2= (/bardblX/bardbl+/bardblY/bardbl)2. Now take square root of both sides to find /bardblX+Y/bardbl ≤ /bardblX/bardbl+/bardblY/bardbl. Nice, eh? See how clean everything is. We have proved Theorem 3.12 . IfHis a scalar product space and we define /bardblX/bardbl=/radicalbig /angbracketleftX, X/angbracketrightin terms of the scalar product, then /bardbl /bardbl is a norm and His a normed linear space with that norm. This special case where the norm is induced by a scalar product is called a pre-Hilbert space (an honest Hilbert space has the additional property of being “complete”). Let us state two easy algebraic consequences of our axioms for a scalar product. The proofs are identical to those of Theorem 4 in the previous section. Theorem 3.13 /angbracketleftX, aY /angbracketright=a/angbracketleftX, Y/angbracketright, a∈R (3-7) /angbracketleftX, Y +Z/angbracketright=/angbracketleftX, Y/angbracketright+/angbracketleftX, Z/angbracketright, (3-8) Needless to say, we hope you are still thinking in the geometric terms presented earlier. In particular, the next definition should be reasonable. Definition Two vectors X,Y are said to be orthogonal if/angbracketleftX, Y/angbracketright= 0 . The Pythagorean theorem suggests Theorem 3.14 . IfXandYare orthogonal, then /bardblX±Y/bardbl2=/bardblX/bardbl2+/bardblY/bardbl2, and conversely. Proof: Both parts are an immediate consequence of the identity /bardblX±Y/bardbl2=/angbracketleftX+Y, X +Y/angbracketright=/bardblX/bardbl2±2/angbracketleftX, Y/angbracketright+/bardblY/bardbl2. Examples . (1) LetX= (2,3,−1) andY= (1,−1,−1) be points in E3, where we use the scalar product of example 1 in this section. Then /angbracketleftX, Y/angbracketright= 2·1 + 3( −1) + (−1)(−1) = 0 soXandYare orthogonal. Similarly X= (2,3,1,−1) andY= (3,−3,3,0) in E4are orthogonal. A useful example is supplied by the vectors e1= (1,0,0,..., 0) , e2= (0,1,0,0,..., 0),...,e n= (0,0,..., 0,1) in En. These are orthonormal since /angbracketleftek, ek/angbracketright= 1 , but /angbracketleftek, el/angbracketright= 0, k/negationslash=l, that is, /angbracketleftek, el/angbracketright=δkl. 3.3. ABSTRACT SCALAR PRODUCT SPACES 117 (2) Consider the functions Φ k(x) = sinkxinL2[−π,π] , wherek= 1,2,3,.... Then, since sinθsin Ψ =1 2[cos(θ−Ψ)−cos(θ+ Ψ)] , we find that /angbracketleftΦk,Φk/angbracketright=/integraldisplayπ −πsin2kxdx =π and fork/negationslash=l /angbracketleftΦk,Φl/angbracketright=/integraldisplayp i−πsinkxsinlxdx = 0. as a computation reveals. Thus in L2[−π,π] the function sin kxis orthogonal to the function sin lxwhenk/negationslash=l. The whole computation may be summarized by /angbracketleftΦk,Φl/angbracketright=/angbracketleftsinkx,sinlx/angbracketright=πδkl. It is only the factor πwhich does not allow us to say that the Φ kareorthonormal— but that is easily patched up. Let ek(x) =sinkx√π. Then /angbracketleftek, el/angbracketright=/angbracketleftsinkx√π,sinlx√π/angbracketright =1 π/angbracketleftsinkx,sinlx/angbracketright, or /angbracketleftek, el/angbracketright=δkl. Therefore the functions ek(x) =sinkx√πare orthonormal. Don’t attempt to imagine it. Just keep on thinking of a big E2and all will be well. So far we have discussed the notion of two vectors XandYbeing orthogonal. This can be restated as one vector Xbeing orthogonal to the subspace Aspanned by Y, for all vectors in Aare of the form aYwhereais a scalar, and /angbracketleftX, aY /angbracketright= 0⇐⇒ /angbracketleftX, Y/angbracketright= 0 since /angbracketleftX, aY /angbracketright=a/angbracketleftX, Y/angbracketright. One can also introduce the concept of a vector Xbeing orthogonal to an arbitrary subspace A. Think of Aas being a plane (through the origin of course). Definition The vector Xisorthogonal to the subspace AifXis orthogonal to every vector in the subspace A. In practice, the usual way to check if Xis orthogonal to the subspace Ais as follows. Pick some basis {Y1,Y2,...}forA. Then every Y∈Ais of the form Y=/summationdisplay akYk (if the basis has an infinite number of elements—that is, if Ais infinite dimensional— one should worry about convergence; however we shall ignore that issue for now). By the algebraic rules for the scalar product, we find that /angbracketleftX, Y/angbracketright=/angbracketleftX,/summationdisplay akYk/angbracketright=/summationdisplay ak/angbracketleftX, Y k/angbracketright. Thus, X is orthogonal to the subspace A if X is orthogonal to every element in some basis forA/angbracketleftX, Y k/angbracketright= 0 . For example, if Ais thex1x2plane in E3, andXis the vector (0 ,0,1) , then we can show that X= (0,0,1) is orthogonal to Aby showing it is orthogonal to both 118 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS the vector e1= (1,0,0) and toe2= (0,1,0) , sincee1ande2form a basis for A. The computation /angbracketleftX, e 1/angbracketright= 0 and /angbracketleftX, e 2/angbracketright= 0 is immediate. Because Y1= (1,2,0) and Y2= (1,−1,0) also form a basis for A, we could prove that Xis orthogonal to Aby showing that /angbracketleftX, Y 1/angbracketright= 0 and /angbracketleftX, Y 2/angbracketright= 0 —which is equally simple. A less obvious example is supplied by the function Ψ( x) = cosxwhich is orthogonal to the subspace Aspanned by Φ 1(x) = sinx,Φ2(x) = sin 2x,..., Φn(x) = sinnxin Lx(−π,π) . The proof is a consequence of the integration formula /angbracketleftΨ,Φk/angbracketright=/integraldisplayπ −πcosxsinkxdx = 0 for all k. Even more general than a vector being orthogonal to a subspace is the idea that two subspaces A and B are orthogonal , by which we mean that every vector in Ais orthogonal to every vector in B. IfAis a subspace of a scalar product space H, then it is natural to define the orthogonal complement A⊥ofAas the set A⊥={X∈H:/angbracketleftX, Y/angbracketright= 0 for all Y∈A} of vectorsXorthogonal to A, that is, orthogonal to every vector Y∈A.The setA⊥is a subspace since it is closed under vector addition and multiplication by scalars (Theorem 2, p. 142). Without fear of evoking surprise, we define the angle θbetween two vectors Xand Yby the formula cosθ=/angbracketleftX, Y/angbracketright /bardblX/bardbl /bardblY/bardbl. No matter what XandYare, this defines a real angle since the right side of the equation is a real number between −1 and +1 (by the Cauchy-Schwarz inequality). To be honest, there is little use for the concept of angles other than right angles. In E3the formula has some use, but is totally unused for more general scalar product spaces. If we are given a set of linearly independent vectors {X1,X2,...}which span a linear scalar product space H, how can we construct an orthonormal set {e1,e2,...}which also spans the space? The process is carried out inductively. Let e1=X1 /bardblX1/bardbl. Now we want a unit vector e2orthogonal to e1. A reasonable candidate is ˜e2=X2− /angbracketleftX2, e1/angbracketrighte1, which isX2with the projection of X2ontoe1subtracted off (see fig.) This vector ˜ e2is orthogonal to e1since /angbracketleft˜e2, e1/angbracketright= 0 . We divide by its length to obtain the unit vector e2, e2=X2− /angbracketleftX2, e1/angbracketrighte1 /bardblX2− /angbracketleftX2, e1/angbracketrighte1/bardbl. Next we take X2and subtract off both its projection into the subspace spanned by e1and e2 ˜e3=X3−[/angbracketleftX3, e1/angbracketrighte1+/angbracketleftX3, e2/angbracketrighte2]. This vector ˜ e3is orthogonal to both e1ande2. Normalize it to get e3= ˜e3//bardbl˜e3/bardbl. 3.3. ABSTRACT SCALAR PRODUCT SPACES 119 More generally, say we have used the vectors X1, X 2,...,X kto obtain the orthonormal sete1,e2,...,e k. Thenek+1is given by ek+1=Xk+1−k/summationdisplay l=1/angbracketleftXk+1, el/angbracketrightel /bardblXk+1−k/summationdisplay l=1/angbracketleftXk+1,el/bardbl This procedure is called the Gram-Schmidt orthogonalization process . With it we can assert that if some set of linearly independent vectors spans a linear space A, we might as well suppose that those vectors constitute an orthonormal set, for if they don’t just use Gram- Schmidt to construct a set that is orthonormal. The next result is a useful observation. Theorem 3.15 . A set {X1,X2,...,X n}of orthogonal vectors, none of which is the zero vector, is necessarily linearly independent. Proof: The hypothesis states that /angbracketleftXj, Xk/angbracketright= 0, j/negationslash=kand that /angbracketleftXj, Xj/angbracketright /negationslash= 0 . Assume there are scalars a1,a2,...a nsuch that 0 =a1X1+a2X2+...+anXn. We shall show that a1=a2=...=an= 0 . Take the scalar product of both sides with the vector X1. Then /angbracketleft0, X 1/angbracketright=a1/angbracketleftX1, X 1/angbracketright+a2/angbracketleftX2, X 1/angbracketright+···+an/angbracketleftXn, X 1/angbracketright. so that 0 =a1/angbracketleftX1, X 1/angbracketright. Since /angbracketleftX1, X 1/angbracketright /negationslash= 0 , we conclude that a1= 0 . Similarly, by taking the scalar product with X2we find that a2= 0 , and so on. An easy consequence of this theorem is the fact that the functions fn(x) = sinnx, n = 1,2,...,N wherex∈[−π,π] are linearly independent, for they are orthogonal (cf. Exercise 5, p. ???). Say we are given an orthonormal set of nvectors, {ej}, j = 1,...,n, /angbracketleftej, ek/angbracketright= δjk, andXan element of the linear space Aspanned by the {ej}. Then X=n/summationdisplay j=1xjej, where thexjare uniquely determined just from the general theory of linear spaces (p. 160, Theorem 10). In the special case of a scalar product space we can conclude even more. Theorem 3.16 . Let {ej, j= 1,...,n }be an orthonormal set of vectors which span A. Then every vector X∈Acan be uniquely written as X=n/summationdisplay j=1xjej, wherexjis the length of the projection of Xinto the subspace spanned by ej, that is,xj=/angbracketleftX, e j/angbracketright. Thexjare the Fourier coefficients of X with respect to the orthonormal basis {ej}. 120 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS Proof: This is identical to Theorem 6 of the last section. Take the inner product of both sides ofX=n/summationdisplay n=1xjejwithek. Then /angbracketleftX, e k/angbracketright=/angbracketleftn/summationdisplay j=1xjej, ek/angbracketright =n/summationdisplay j=1xj/angbracketleftej, ek/angbracketright=n/summationdisplay j=1xjδjk,(3-9) so that /angbracketleftX, e k/angbracketright=xk. Furthermore, Theorem 3.17 . Let {ej}, j= 1,...,n be an orthonormal set of vectors which span A. IfX=n/summationdisplay j=1xjejandY=n/summationdisplay j=1yjejare vectors in A, then /angbracketleftX, Y/angbracketright=n/summationdisplay j=1xjyj=x1y1+x2y2+···+xnyn. Proof: Identical to Theorem 7 of the last section. /angbracketleftX, Y/angbracketright=/angbracketleftn/summationdisplay j=1xjej,n/summationdisplay k=1ykek/angbracketright =n/summationdisplay j=1xj/angbracketleftej,n/summationdisplay k=1ykek/angbracketright =n/summationdisplay j=1xj(n/summationdisplay k=1yk/angbracketleftej, ek/angbracketright) =n/summationdisplay j=1xj(n/summationdisplay k=1ykδjk),(3-10) so that /angbracketleftX, Y/angbracketright=n/summationdisplay j=1xjyj. Remark : We shall see that these two theorems extend to the case n=∞. Examples . (1) The vectors e1= (1,0,0), e2= (0,1,0) , ande3(0,0,1) clearly form an orthonormal basis for E3. LetX= (2,−1,4) . We shall compute the xjin X=3/summationdisplay j=1xjej. 3.3. ABSTRACT SCALAR PRODUCT SPACES 121 Sincexj=/angbracketleftX, e j/angbracketright, we findz1=/angbracketleftX, e 1/angbracketright=/angbracketleft(2,−1,4),(1,0,0)/angbracketright= 2·1+(−1)·0+4·0 = 2 , and similarly, x2=−1, x3= 4 as expected. Thus (2,−1,4) = 2e1−e2+ 4e3. In the same way, if Y= (7,1,−3) , then Y= 7e1+e2−3e3. Also, /angbracketleftX, Y/angbracketright= (2)(7) + ( −1)(1) + (4)( −3) = 1. The projection of Xinto the subspace spanned by Yis /angbracketleftX, Y/ /bardblY/bardbl/angbracketrightY /bardblY/bardbl=1 59(7,1,−3) =7 59e1+1 59e2−3 59e3.(3-11) Another orthonormal basis for E3is ˜e1= (1√ 2,1√ 2,0),˜e2= (−1√ 2,1√ 2,0) , and ˜e3= (0,0,1) , since /angbracketleft˜ej,˜ek/angbracketright=δjk. The expansion for Xin this basis is X=3/summationdisplay j=1˜xj˜ej, where ˜x1=/angbracketleftX,˜e1/angbracketright=/angbracketleft(2,−1,4),(1√ 2,1√ 2,0)/angbracketright=1√ 2, (3-12) ˜x2=−3√ 2,and ˜x3= 4. (3-13) Thus X=1√ 2˜e1−3√ 2˜e2+ 4˜e3. Similarly, Y=8√ 2˜e1−6√ 2˜e2−3˜e3. Therefore /angbracketleftX, Y/angbracketright= (1√ 2)(8√ 2) + (−3√ 2)(−6√ 2) + (4)( −3) = 1. Notice that the number /angbracketleftX, Y/angbracketrightis the same no matter which basis is used. This is nota coincidence. Recall that the scalar product /angbracketleftX, Y/angbracketrightwas defined independently of any basis. Hence its value should not be dependent upon which basis we happen to choose. If you think of /angbracketleftX, Y/angbracketrightgeometrically in terms of the projection, it should be clear that the number should not depend upon which particular basis is used to describe the vectors. 122 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS (2) For our second example, we consider the set of orthonormal functions e1(x) =sinx√π, e2(x) =sinx√πand letAbe the set in L2(−π,π) which they span. We would like to expand some function f(x) +2/summationdisplay j=1fjej(x). The only trouble is that Theorems 14 and 15 only allow us to expand functions f which are in the subspace A, that is, are a linear combination of the basis elements e1 ande2. Since we secretly know that f(x) = sinxcosx(=1 2sin 2x) is such a function, let us find its expansion. By elementary integration, f1=/angbracketleftf, e 1/angbracketright=/integraldisplayπ −π(sinxcosx)sinx√xdx= 0, and f2=/angbracketleftf, e 2/angbracketright=/integraldisplayπ −π(sinxcosx)sin 2x√πdx=√π 2. Therefore f= 0·e1+√π 2e2=√π 2e2 or sinxcosx=√π 2(sin 2x√π) =sin 2x 2, which we knew was the case from trigonometry. If the orthonormal set {ej}, j= 1,...,m spans a subspaceAof a linear scalar product spaceH, and ifX∈H, can any sense be made of the expansion X?=m/summationdisplay j=1xjej? One way to seek an answer is to examine a special case. Again geometry will supply the key. LetH=E3and letAbe the subspace spanned by the orthonormal vectors e1= (1,0,0) ande2= (0,1,0) . Then if X∈E3, how can we interpret X?=2/summationdisplay j=1xjej=x1e1+x2e2? Plowing blindly ahead, we take the scalar product of both sides with e1and then with e2. This gives us xj=/angbracketleftX, e j/angbracketright. Thus the right side, x1e1+x2e2, is the projection ofXinto the subspace Aspanned by {ej}. It is now clear how our original quandary is resolved. Definition If the orthonormal set {ej}, j= 1,...,m spans a subspace Aof a linear scalar product space H, and ifX∈H, then the vectorm/summationdisplay j=1xjej, wherexj=/angbracketleftX, e j/angbracketright,is the projection of X into the subspace A. Remark . It is customary to denote the projection of XintoAbyPAX. Think of PA as an operator (function) which maps the vector Xinto its projection in A. With this notation the above definition reads PAX=m/summationdisplay j=1xjej, 3.3. ABSTRACT SCALAR PRODUCT SPACES 123 wherexj=/angbracketleftX, e j/angbracketrightand the orthonormal set {ej}spansA. Since the projection PAXis defined in terms of a particular basis for A, we should show that this geometrical object is independent of the basis you choose for A. But we shall not take the time right now. In reality, Theorem 17 below leads us to make a better definition of projection. Theorem 3.18 . If the orthonormal set {ej}, j= 1,...,m spans a subspace A⊂H, and ifXandYare inH, then a)/angbracketleftPAX, P AY/angbracketright=m/summationdisplay j=1xjyj, wherexj=/angbracketleftX, e j/angbracketrightandyj=/angbracketleftY, e j/angbracketright. In particular b)/bardblPAX/bardbl=/radicaltp/radicalvertex/radicalvertex/radicalbtm/summationdisplay j=1x2 j. Furthermore, X−PAX∈A⊥, that is, for every Y∈A c)/angbracketleftX−PAX, Y/angbracketright= 0 EveryX∈Hcan be written as d)X=PAX+PA⊥X,wherePA⊥X≡X−PAXis inA⊥. Proof: Since both vectors PAX=m/summationdisplay j=1xjejandPAY=m/summationdisplay j=1yjejare inAitself, a) and b) are immediate consequences of Theorem 15. Although the equation c) is geometrically clear, we shall compute it too. Since theejspanA, this is equivalent to showing it is orthogonal to all the ej. Now /angbracketleftX−PAX, e j/angbracketright=/angbracketleftX, e j/angbracketright−/angbracketleftPAX, e j/angbracketright=xj−xj= 0 . Since trivially X=PAX+(X−PAX) , the only content of part d) is that ( X−PAX)∈A⊥, which is just what part c) proved. Corollary 3.19 a)/bardblX/bardbl2=/bardblPAX/bardbl2+/bardblX−PAX/bardbl2 /bardblX/bardbl2=/bardblPAX/bardbl2+/bardblPA⊥X/bardbl2/bracerightbigg (Pythagorean Theorem) b)/bardblX/bardbl2≥ /bardblPAX/bardbl2=m/summationdisplay j=1x2 j (Bessel’s Inequality) Proof: a) is a result of the fact that PAX∈Ais orthogonal to X−PAX∈A⊥and Theorem 12. The inequality b), Bessel’s inequality, is simply a weaker form of a)—since /bardblX−PAX/bardbl ≥0 . There is equality if and only if X∈A, for only then does /bardblX−PAX/bardbl= 0 . Examples : 124 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS (1) LetAbe the subspace of E3spanned by e1= (1,0,0) ande2= (0,1,0) . The projection of X= (3,−1,7) intoAis represented by PAX=/angbracketleftX, e 1/angbracketrighte1+/angbracketleftX, e 2/angbracketrighte2= 3e1−e2∈A Also PA⊥X=X−PAX= 3e1−e2+ 7e3−(3e1−e2) = 7e3∈A⊥. Since ˜e1= (1√ 2,1√ 2,0) and ˜e2= (−1√ 2,1√ 2,0) also form an orthonormal basis for A, we can equally well write PAX=/angbracketleftX,˜e1/angbracketright˜e1+/angbracketleftX,˜e2/angbracketright˜e2=2√ 2˜e1−4√ 2˜e2. (2) LetAbe the subspace of L2[−π,π] spanned by the orthonormal functions e1(x) = sinx√π, e2(x) =sin 2x√π. The projection of the function f(x)≡xintoAis represented by PAf=/angbracketleftf, e 1/angbracketrighte1+/angbracketleftf, e 2/angbracketrighte2. Since an integration by parts shows that /integraldisplayπ −πxsinkxdx =−xcoskx k/vextendsingle/vextendsingleπ −πcoskxdx =−xcoskx k/vextendsingle/vextendsingleπ −π=−2π kcoskπ=/braceleftbigg2π k, kodd −2π k, keven/bracerightbigg = (−1)k+12π k, we find /angbracketleftf, e 1/angbracketright=/angbracketleftx, e 1/angbracketright=/integraldisplayπ −πxsinx√πdx= 2√π and /angbracketleftf, e 2/angbracketright=/angbracketleftx, e 2/angbracketright=/integraldisplayπ −πxsin 2x√πdx=−√π. Thus PAx= 2√πsinx√π−√πsin 2x√π, or PAx= 2 sinx= sin 2x. Also, /bardblPAX/bardbl2=/angbracketleftf, e 1/angbracketright2+/angbracketleftf, e 2/angbracketright2= 5π. More generally, we can let ˜Abe the subspace of L2[−π,π] spanned by {ek}, k= 1,2,...,N , whereek(x) =sinkx√π. Then the projection of xonto ˜Ais given by P˜Ax=N/summationdisplay k=1/angbracketleftx, e k/angbracketrightek(x). Since /angbracketleftx, e k/angbracketright=/integraldisplayπ −πxsinkx√πdx= (−1)k+12√π k, 3.3. ABSTRACT SCALAR PRODUCT SPACES 125 we have P˜Ax=N/summationdisplay k=1(−1)k+12√π ksinkx√π = 2N/summationdisplay k=1(−1)k+1 ksinkx = 2(sinx−sin 2x 2+sin 3x 3−...+ (−1)N+1sinNx N).(3-14) Furthermore, /bardblP˜Ax/bardbl2=N/summationdisplay k=1/angbracketleftx, e k/angbracketright2=N/summationdisplay k=14π k2= 4πN/summationdisplay k=11 k2. It is from this formula that we eventually intend to obtain the famous formula ∞/summationdisplay k=11 k2= 1 +1 22+1 32+1 42+···=π2 6. We will observe that /bardblf/bardbl2=/bardblX/bardbl2=/integraldisplayπ −πx·xdx=2π3 3 and prove lim N→∞/bardblP˜A⊥x/bardbl= lim N→∞/bardblx−P˜Ax/bardbl= 0 Then from the Corollary to Theorem 16, /bardblX/bardbl2= lim N→∞/bardblP˜Ax/bardbl2, or 2π3 3= 4π?/summationdisplay k=11 k2⇒π2 6=∞/summationdisplay k=11 k2. Geometry leads us to the next theorem—and the proof too. Let Xbe a given vector andPAXits projection into the subspace A. Since distance is measured by dropping a perpendicular, we expect that PAXis the vector in Awhich is closest to X, that is, most closely approximates X. Theorem 3.20 . LetXbe a vector in a scalar product space HandAa subspace of H. Then ifVis any vector in A, /bardblX−PAX/bardbl ≤ /bardblX−V/bardbl. Proof: We shall prove the stronger statement (cf. fig. above) /bardblX−PAX/bardbl2+/bardblV−PAX/bardbl2=/bardblX−V/bardbl2. Observe that ( V−PAX)∈A, since both terms are in AandAis a subspace. Moreover X−PAX∈A⊥(Theorem 16c). Therefore X−PAXis orthogonal to PAX−V, so the identity is a consequence of Theorem 12. 126 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS Remark . With this theorem in mind, we could define the projection PAXinto a subspace Aas the element in Awhich is closest to X. This definition is independent of any basis, whereas our original definition was not. One must, however, be somewhat careful when defining the projection into an infinite dimensional subspace. Although it is clear that the number /bardblX−V/bardblhas a g.l.b. as Vwanders throughout A, it is notclear that it has an actual min, that is, there really is a vector U∈Asuch that /bardblX−U/bardbltakes on its g.l.b. as a min. If there is such a U, we call it PAX. Otherwise there is no projection. When projecting into a finite dimensional space this difficulty does not arise (but we will stop without further explanation of this detail). Some discussion of these results is needed to place the material in its proper perspective. If you are given an orthonormal set of vectors {ej}which span some subspace Aof a scalar product space H, then for any XinHyou can find a representation for PAXin terms of that basis, PAX=/summationtextxjej. If the vector Xhappened to already lie in A, thenPAX=X soX=/summationtextxjejand/bardblX/bardbl=/radicalBig/summationtextx2 j. This last equation for the length of Xis the Pythagorean Theorem. If Xdid not lie entirely in A, but “stuck out” of it into the rest ofH, thenPAX=/summationtextxjejonly represents a piece of X, its projection into A. Since part ofXhas been omitted, we expect that /bardblX/bardbl>/bardblPAX/bardbl=/radicalBig/summationtextx2 j. This inequality was the content of the Corollary to Theorem 16. Informally, if no vector X∈Hsticks out of the linear space spanned by the {ej}, then the set {ej}is said to be complete (do not confuse this with the complete of Chapter 0; they are entirely different concepts, an unfortunate coincidence). More precisely, Definition An orthonormal set is complete for the scalar product space Hif that or- thonormal set is not properly contained in a larger orthonormal set. There are many ways to check if a given orthonormal set is complete for H. Geometry suggests them all. Theorem 3.21 . Let {ej}be an orthonormal set which spans the subspace Aof the scalar product space H. The following statements are equivalent (a)The set {ej}is complete for H. (b)If/angbracketleftX, e j/angbracketright= 0 for allj, thenX= 0. (c)A=H. (d)IfX∈H, thenX=/summationtextxjej, wherexj=/angbracketleftX, e j/angbracketright. (e)IfXandY∈H, then /angbracketleftX, Y/angbracketright=/summationdisplay xjyj, wherexj=/angbracketleftX, e j/angbracketrightandyj=/angbracketleftY, e j/angbracketright (f)IfX∈H, then (Pythagorean Theorem) /bardblX/bardbl2=/summationdisplay x2 j,wherexj=/angbracketleftX, e j/angbracketright Proof: We shall use the chain of reasoning a⇒b⇒c...⇒f⇒a. a⇒b. If/angbracketleftX, e j/angbracketright= 0 butX/negationslash= 0 , thenX//bardblX/bardblis a unit vector orthogonal to all the ej. This means that {X /bardblX/bardbl, e1,e2,...}is an orthonormal set which contains {e1,e2,...} as a proper subset. b⇒c. If there is an X∈HbutX/ownerA, thenPA⊥X=X−PAX∈A⊥and is not zero. Since all the ej∈A, we have /angbracketleftPA⊥X, e j/angbracketright= 0 for all jbutPA⊥X/negationslash= 0 , contradicting b). ThusH⊂A. SinceA⊂Hby hypothesis, this proves that H=A. 3.3. ABSTRACT SCALAR PRODUCT SPACES 127 c⇒d. Since every X∈Ahas the form X=/summationtextxjej(by Theorem 14) and since H=A, the conclusion is immediate. d⇒e⇒f. A restatement of Theorem 16 since for every X∈H, we know that PAX=PHX=X. f⇒a. If{ej}is not complete, it is contained in a larger orthonormal set. Let e be a vector in that larger set which is not one of the ej. Then by f), and the fact that /angbracketlefte, ej/angbracketright= 0 , /bardble/bardbl2=/summationdisplay /angbracketlefte, ej/angbracketright2= 0. Thereforee= 0 . Remarks . 1. Because each of the six conditions a-f are equivalent, any one of them could have been used as the definition of a complete orthonormal set. 2. If the orthonormal set {ej}has a (countably) infinite number of elements, the theorem is still valid but some convergence questions for d-f arise because of the then infinite seriesX=∞/summationdisplay 1xjej. The appropriate sense of convergence is that the remainder afterNterms,∞/summationdisplay N+1xjej=X−N/summationdisplay 1xjejtends to zero in the norm of the scalar product space , that is, if lim N→∞/bardblX−N/summationdisplay 1xjej/bardbl= 0. We shall meet this in the next section for the space L2[−π,π] . Condition f) gives us no convergence problems since the series is an infinite series of positive terms which is always bounded by /bardblX/bardbl2(Bessel’s Inequality—Corollary b to Theorem 16), and so always converges. This criterion just asks if the sum of the series actually equals /bardblX/bardbl2(we know it is no larger). Examples (1) The set of orthonormal vectors e1= (1√ 2,1√ 2,0) ande2= (1√ 2,−1√ 2,0) are not complete for E3since any basis for E3must have three elements because its dimension is 3. This could also be seen geometrically from the fact that, for example X= (1,2,3) sticks out of the space spanned by e1ande2, or from the fact that e3= (0,0,2) is a non-zero vector orthogonal to both e1ande2, or in many other ways. The dimension argument is the easiest to apply if H is finite dimension , for then the number of elements in a complete orthonormal set {ek}must equal the dimension of H. (2) The set {˜en}where ˜en(x) =sinnx√πis an orthonormal set of functions in the scalar product space L2[−π,π] , but it is nota complete orthonormal set for that space since the function cos xis a non-zero function in L2[−π,π] which is orthogonal to all the ˜en, /angbracketleftcosx,˜en/angbracketright=/integraldisplayπ −πcossinnx√πdx= 0. Thus, although the set {˜en}has an infinite number of elements, it is still not big enough to span all of L2[−π,π] . The next section will be devoted to proving that the 128 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS larger orthonormal set, e0,e1,˜e1,e2,˜e2,..., where e0=1√ 2π, en(x) =cosnx√π,˜en(x) =sinnx√π is a complete orthonormal set for the scalar product space L2[−π,π] . This is a difficult theorem. Specific applications of the ideas in this section are contained in the exercises. For many of them you would be wise if you referred to their corresponding special cases which appeared in Section 2. Exercises (1) LetXandYbe points in Rn. Determine which of the following make Rninto a scalar product space, and why—or why not. (a)/angbracketleftX, Y/angbracketright=n/summationdisplay k=11 kxkyk. (b)/angbracketleftX, Y/angbracketright=n/summationdisplay k=1(−1)kxkyk. (c)/angbracketleftX, Y/angbracketright=/radicaltp/radicalvertex/radicalvertex/radicalbtn/summationdisplay k=1x2 ky2 k. (d)/angbracketleftX, Y/angbracketright=n/summationdisplay k=1akxkyk, whereak>0 for allk. (2) Letfandgbe continuous real-valued functions in the interval [0 ,1] , sof, g∈ C[0,1] . Determine which of the following make C[0,1] into a scalar product space, and why—or why not. (a)/angbracketleftf, g/angbracketright=/integraldisplay1 0f(x)g(x)1 1 +x2dx. (b)/angbracketleftf, g/angbracketright=/integraldisplay1 0f(x)g(x) sin 2πxdx . (c)/angbracketleftf, g/angbracketright=/integraldisplay1 0f(x)g2(x)dx (d)/angbracketleftf, g/angbracketright=/integraldisplay1 0f(x)g(x)ρ(x)dx, whereρ(x) is a fixed continuous function with the propertyρ(x)>0 . (e)/angbracketleftf, g/angbracketright=f(0)g(0) . (3) This is the analogue of L2for sequences. Let l2be the set of all sequences X= (x1,x2,xe,...) with the property that /bardblX/bardbl=/radicaltp/radicalvertex/radicalvertex/radicalbt∞/summationdisplay j=1x2 j<∞. Prove that l2is a normed linear space (cf. the example for l1in Section 1). 3.3. ABSTRACT SCALAR PRODUCT SPACES 129 (4) Use the Cauchy-Schwarz inequality to prove that if∞/summationdisplay n=1n2a2 n<∞, then∞/summationdisplay n=1|an|<∞. (Hint: |an|=1 n|nan|). (5) Consider the following linearly independent vectors in E3: X1= (1,0,−1), X 2= (0,3,1), X 3= (2,−1,0). (a) Use the Gram-Schmidt orthogonalization process to find an orthonormal set of vectors,e1,e2ande3such thate1is in the subspace spanned by X1. (b) WriteX= (1,2,3) asX=3/summationdisplay j=1xjej, where the ejare those of part a). Also, compute /bardblX/bardbland/bardblPX/bardbl. (6) Consider the following linearly independent set of functions in L2[−1,1] f1(x) = 1, f 2(x) =x, f 3(x) =x2. (a) Use the Gram-Schmidt orthogonalization process to find an orthonormal set of functionse1(x),e2(x) ande3(x) such that e1is in the subspace spanned by f1. (b) Find the projection of the function f(x) = (1+x)3into the subspace of L2[−1,1] spanned by e1(x),e2(x) , ande3(x) . Also, compute /bardblf/bardbland/bardblPf/bardbl. (7) LetPn(x) =1 2nn!dn dxn(1−x2)n,n= 0,1,2,.... These are the Legendre Polynomials . (a) Prove that /angbracketleftPn, Pm/angbracketright=/integraldisplay1 −1Pn(x)Pm(x)dx= 0, n/negationslash=m, that is, the Pnare orthogonal in L2[−1,1] by first proving that /integraldisplay1 −1Pn(x)xmdx= 0, m<n. (b) Show that /bardblPn/bardbl2=2 2n+1. Thus the functions en(x) =/radicalbigg 2n+ 1 2Pn(x) are an orthonormal set of functions for L2[−1,1] . Compute e0(x),e1(x) , and e2(x) and compare with Exercise 6a. (8) (a) Show that the vector N= (a1,a2,a3) is orthogonal to the coset (a plane in E3) A={X∈E3:a1x1+a2x2+a3x3=c}. (b) Show that the vector N= (a1,...,a n) is orthogonal to the coset (a hyperplane inEn)A={X∈En:a1x1+...a nxn=c}. (c) Find the coset A⊂E3which passes through the point X0= (1,−1,2) and is orthogonal to N= (1,3,2) . In ordinary language, Ais the plane containing the pointX0which is orthogonal to N. 130 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS (d) Show that the coset A⊂Enwhich passes through the point X0= (˜x1,..., ˜xn) and is orthogonal to N= (a1,...,a n) is A={X∈En:/angbracketleftX, N/angbracketright=/angbracketleftX0, N/angbracketright}. (9) (a) Use Problem 8a to show that the distance dfrom the point P= (y1,y2,y3)∈E3 to the coset a1x1+a2x2+a3x3=cinE3is d=|a1y1+a2y2+a3y3−c|/radicalbig a2a+a2 2+a2 3=|/angbracketleftN, P/angbracketright −c| /bardblN/bardbl (b) Show that the distance dfrom the point P= (y1,...,y n)∈Ento the coset a1x1+...+anxn=cinEnis d=|a1y1+a2y2+···+anyn−c|/radicalbig a2 1+a2 2+···+a2n=|/angbracketleftN, P/angbracketright −c| /bardblN/bardbl. (c) Show that the distance dbetween the “parallel” cosets a1x1+···+anxn=c1 anda1x1+···+anxn=c2inEnis d=|c1−c2|/radicalbig a2 1+···+a2n=|c1−c2| /bardblN/bardbl. (Hint: Pick a point Pin one of the cosets and apply part b). (10) Find the angle between the diagonal of a cube and one of its edges. (11) LetY1andY2be fixed vectors in a scalar product space H. a). If /angbracketleftX, Y 1/angbracketright= 0 for all X∈H, prove that Y1= 0 . b). If /angbracketleftX, Y 1/angbracketright=/angbracketleftX, Y 2/angbracketrightfor allX∈H, prove that Y1=Y2. (12) LetY0be a fixed vector in a scalar product space H. LetA={X∈H:/angbracketleftY, Y 0/angbracketright= 0Rightarrow /angbracketleftX, Y/angbracketright= 0}. Prove that Ais the span of Y0:{X∈ARightarrowX = cY0}for some scalar c. Make sure to see the geometrical situation for the case H=E3. [Hint: LetBbe the set of all vectors orthogonal to Y0, soY∈B. SinceHis composed of two parts, Y0andB, everyX∈Hcan be written as X=cY0+Z, wherecY0is the projection of Xinto the subspace spanned by Y0 (soc=/angbracketleftX, Y 0/angbracketright//bardblY0/bardbl2) andZ= (X−cY0)∈B. Now show that X∈A⇒Z= 0) . (13) (a) Let X= (1,3,−1) andY= (2,1,1) . Find a vector Nwhich is orthogonal to the subspace spanned by XandY. (b) LetX= (x1,x2,x3) andY= (y1,y2,y3) . Find a vector Nwhich is orthogonal to the subspace spanned by XandY. [Answer. N=c(x2y3−y2x3,y1x3− x1y3,x1y2−y1x2) , wherecis any non-zero scalar]. (14) LetAbe the subspace of L2[−π,π] spanned by the orthonormal set {en(x)}, n= 1,2,...,N , whereen(x) =sinnx√π. (a) Find the projection of f(x) =x2, intoA. (The answer should surprise you). Compute /bardblf/bardbland/bardblPAf/bardbltoo. 3.3. ABSTRACT SCALAR PRODUCT SPACES 131 (b) Find the projection of f(x) = 1 + sin3xintoA. Compute /bardblf/bardbland/bardblPAf/bardbl. (c) Iff(x) is an even function, f(x) =f(−x) , show that its projection into Ais zero. Now look at part (a) again. (15) (a) If f∈C[a,b] , show that (/integraldisplayb af(x)dx)2≤(b−a)/integraldisplayb af2(x)dx. [Hint: Write f(x) = 1·f(x) and use the Cauchy-Schwarz inequality for L2[a,b] ]. (b) Iff∈C1[a,b] , prove that |f(x)−f(a)|2≤(x−a)/integraldisplayb af/prime(x)2dx, x∈(a,b). [Hint: Write f(x)−f(a) =/integraltextx af/prime(t)dtand apply part a)] (c) Iff∈C1[a,b] andf(a) = 0 , use part b to prove that /integraldisplayb af2(x)dx≤(b−a)2 2/integraldisplayb af/prime(x)2dx. (16) (a) Let A={h∈C1[a,b]:h(a) =h(b)}and letB={h∈C1[a,b]:/angbracketleft1, h/prime/angbracketright= 0}, whereh/prime=dh dx. Show that the subspaces AandBare identical h∈A⇐⇒ h∈B. (b) Letf(x) be any continuous function such that/integraldisplayb af(x)h/prime(x)dx= 0 for all h(x)∈C1[a,b] withh(a) =h(b) . Show that f≡constant. [Hint: Use part (a) and the result of Exercise 12]. (17) Iff(x)∈C[a,b] and satisfies the condition/integraldisplayb af(x)h(x)dx= 0 for all h(x)∈C[a,b] which satisfy the conditions /integraldisplayb ah(x)dx= 0,/integraldisplayb axh(x)dx= 0,...,/integraldisplayb axnh(x)dx= 0, prove that f∈Pn, that is,fis of the form f(x) =a0+a1x+···+anxn, where theajare constants. [Hint: Use Exercise 12]. (18) Determine which of the following orthonormal sets are complete for their respective spaces. (a) In E3, e1= (0,1,0), e2= (3 5,0,4 5), e3= (−4 5,0,3 5) (b) In E4, e1= (1,0,0,0), ex= (0,1,0,0), e3= (0,0,1√ 2,1√ 2). (c) In E4, e1,e2,e3as in (b), and e4= (0,0,1√ 2,−1√ 2). 132 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS (19) Lete1,e2, ande3be an orthonormal basis for E3, and letAbe the subspace spanned by X1= 3e1−4e3. Find an orthonormal basis for A⊥. (20) LetAbe a subspace of a scalar product space H. IfX∈H, prove that PA(PAX) = PAXand interpret this geometrically. This result can be written as P2 A=PA. (21) LetAbe any operator (not necessarily linear) on a scalar product space. Prove the polarization identity 2/angbracketleftAX, AY /angbracketright=/bardblAX+AY/bardbl2− /bardblAX/bardbl2− /bardblAY/bardbl2. 3.4 Fourier Series. Throughout this section we shall only use the scalar product of L2[a,b] , /angbracketleftf, g/angbracketright=/integraldisplayb af(x)g(x)dx. We begin with the observation that in the interval [ −π,π] /angbracketleftsinnx,sinmx/angbracketright=/integraldisplayπ −πsinnxsinmxdx =πδnm, (3-15) /angbracketleftsinnx,cosmx/angbracketright=/integraldisplayπ −πsinnxcosmxdx = 0, (3-16) and /angbracketleftcosnx,cosmx/angbracketright=/integraldisplayπ −πcosnxcosmxdx =πδnm, (3-17) wheren, m = 0,1,2,3,.... Thus the functions e0(x) =1√ 2π, en(x) =cosnx√π,˜cn(x) =sinnx√π form an orthonormal set: /angbracketleften,˜em/angbracketright=δnm/angbracketleften,˜em/angbracketright= 0,/angbracketleft˜en,˜em/angbracketright=δnm. Thus, iff∈Lx[−π,π] , we can find the projection PNfoffinto the subspace spanned bye0,e1,˜e,...,e N,˜eN. (PNf) =a0e0+N/summationdisplay n=1anen+bn˜en, (3-18) where ak=/angbracketleftf, ek/angbracketrightandbk=/angbracketleftf,˜ek/angbracketright (3-19) More explicitly, (PNf)(x) =a01√ 2π+N/summationdisplay n=1ancosnx√π+bnsinnx√π, (3-20) 3.4. FOURIER SERIES. 133 where a0=/integraldisplayπ −πf(x)·1√ 2πdx, and an=/integraldisplayπ −πf(x)cosnx√πdx, b n=/integraldisplayπ −πf(x)sinnx√πdx (3-21) A natural question arises: as N→ ∞ , does the series converge: PNf→f, in the sense that/bardblf−PNf/bardbl →0 ? In other words, is the set {ej(x),˜ej(x)}, j= 0,1,2,... acomplete orthonormal set of functions for L2[−π,π] ? The answer is yes, as we shall prove. Thus for anyf∈L2[−π,π] , f(x) =a01√ 2π+∞/summationdisplay n=1ancosnx√π+bnsinnx√π, (3-22) where the Fourier coefficients ,an,bnare determined by the formulas (2). The expansion (3) is called the Fourier series for f. Historically, Fourier series did not arise from the geometrical considerations we have developed. Mathematical physics—in particular the vibrations of strings and the flow of heat in a bar—take the credit for these ideas. Only in recent years has the geometrical viewpoint been investigated. Later on we shall discuss some of the fascinating problems in mathematical physics to which Fourier series can be applied. Beware . The equality which appears in (3) is equality in the L2[−π,π] norm, viz. /bardblf−PNf/bardbl=/radicalBigg/integraldisplayb a[f(x)−(PNf)(x)]2dx→0 This is quite different than the convergence of infinite series to which you’re accustomed, which is the uniform norm /bardblf−PNf/bardbl∞= max −π≤x≤π|f(x)−(PNf)(x)|. In Section 1 (p. 176) you saw one instance of where a sequence of functions converged in some norm (the L1norm there) but did not converge in the uniform norm. Such is also the case here. In fact, contrasting the situation in the L2norm, there do exist continuous functionsfwhose Fourier series (3) does not converge to fin the uniform norm. However if the function fhas one derivative, then its Fourier series does converge to fin the uniform norm. In addition, there are some discontinuous functions whose Fourier series converge. These ideas will become clearer later on. You should be warned that our definition (1), (3) of a Fourier series is not the standard one. Most books do notwork with the orthonormal sete0=1√ 2π, en=cosnx√π,˜en=sinnx√π, but rather use just an orthogonal set which is notnormalized θ0=1 2, θn= cosnx,˜θn= sinnx. For these people, f(x) =A0 2+∞/summationdisplay n=1Ancosnx+Bnsinnx, where An=/integraldisplayπ −πf(x)cosnx πdx, B n=/integraldisplayπ −πf(x)sinnx πdx, 134 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS n= 0,1,2.... As you can see, these differ from our formulas only by factors of√π. Needless to say, the resulting Fourier series for a given function fdoes not depend which intermediate formulas you use. We prefer the less standard ones because they are more intimately tied to geometry (so there is less to remember). Before discussing the difficult issues of convergence in detail, we will find the Fourier series associated with some specific functions. Examples . (1) Find the Fourier series associated with the functions f(x) =x,−π≤x≤π. We actually found this in the previous section. A computation (involving integration by parts) shows that a0=/angbracketleftf, e 0/angbracketright=/integraldisplayπ −πx·1√ 2πdx= 0 an=/angbracketleftf, en/angbracketright=/integraldisplayπ −πxcosnx√πdx= 0, n = 1,2,... bn=/angbracketleftf,˜en/angbracketright=/integraldisplayπ −πxsinnx√πdx=2(−1)n+1√π n, n = 1,2,... Thus, upon substituting into (3) we find that x=∞/summationdisplay n=12(−1)n+1 n√πsinnx√π= 2∞/summationdisplay n=1(−1)n+1 nsinnx or x= 2[sinx−sin 2x 2+sin 3x 3−sin 4x 4+···]. Again we remind you that the equality here is in the sense of convergence in L2. For this particular function, there is also equality in the usual sense of convergence for infinite series for all x∈(−π,π) . Direct substitution reveals that it does notconverge in the usual sense at x=±π. These remarks are based upon convergence theorems we have yet to prove. At x=π 2, this yields π 4= 1−1 3+1 5−1 7+··· (2) Since the formulas (2)’ make sense even if the function f(x) has a finite number of discontinuities, we are tempted to find the Fourier series for discontinuous functions (in contrast, recall that the coefficients of an infinite power series are only defined if the function had an infinite number of derivatives). We shall find the Fourier series associated with the discontinuous function f(x) =/braceleftbigg0,−π≤x≤0 π,0<x<π 3.4. FOURIER SERIES. 135 The computations are particularly simple. a0=/angbracketleftf, e 0/angbracketright=/integraldisplay0 −π0·1√ 2πdx+/integraldisplayπ 0π·1√ 2πdx =π2 √ 2π,(3-23) an=/angbracketleftf, en/angbracketright=/integraldisplay0 −π0·cosnx√πdx+/integraldisplayπ 0π·cosnx√πdx= 0, n> 0, (3-24) bn=/angbracketleftf,˜en/angbracketright=/integraldisplay0 −π0·sinnx√πdx+/integraldisplayπ 0π·sinnx√πdx (3-25) =√π n(1−cosnπ) =/braceleftBigg 2√π n, n odd 0, n even(3-26) Therefore the Fourier series associated with this function is f(x) =π2 √ 2π·1√ 2π+ 2√π/parenleftbiggsinx√π+sin 3x 3√π+sin 5x 5√π+···/parenrightbigg , or f(x) =π 2+ 2(sinx+sin 3x 3+sin 5x 5+···). As usual, the equality is meant in the sense of convergence in the L2norm. The series also converges to the function fin the uniform norm in the whole interval except for a neighborhood of x= 0 . At 0 it hasn’t got a chance because of the discontinuity of f there. A glance at the series reveals that at x= 0 , the right side is π/2 —the arithmetic mean between the values of fjust to the left and right of 0. This is the usual case at a discontinuity: a Fourier series converges to the average of the function values to the right and left of the point where fis discontinuous . We still offer no proof for these statements. Observe that the Fourier series (3) for any function f(x) depends only upon the values ofxin the interval −π≤x≤π. However the series itself is periodic with period 2 π. If the function f(x) , which we considered only for x∈[−π,π] is defined for all other x by the formula f(x+ 2π) =f(x) (making fperiodic too), then both sides of the Fourier series (3) are periodic with period 2 π. Therefore whatever they do in the interval [ −π,π] is repeated every 2 π. For example, the function f(x) =x, x∈[−π,π] when continued outside the interval [−π,π] as a function periodic with period 2 πbecomes a figure goes here Since the Fourier series for this particular function converges uniformly for all x∈(−π,π) , it also converges uniformly to the periodically continued function for all x∈(kπ,kπ + 2π), k= 1,±1,±2,... . This also makes it clear why the Fourier series for f(x) =x converges to zero at x=±π, for the series is just converging to the arithmetic mean of its neighboring values at the discontinuity. It is pleasant to look at a picture. Let us see how the first four terms of its Fourier series approximates the function x x= 2(sinx−sin 2x 2+sin 3x 3−sin 4x 4+···) a figure goes here 136 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS Notice that as more terms are used, the projection PNx P Nx= 2(sinx−sin 2x 2+···+ (−1)N+1sinNx N) more and more closely approximate x. This reflects the convergence of the Fourier series, PNf→f. One popular interpretation of a Fourier series is as a sum of “waves” which approximate a given function. Thus the function xis the sum of 2 times the wave sin xplus ( −1) times the wave sin 2 xand so on. In other words, the Fourier series for the function f(x) =x represents that function as the superposition of sine waves. The term 2 sin xis spoken of as the first harmonic , the term – sin 2 xas the second harmonic, the term2 3sin 3xas the third harmonic, etc. Although it is difficult to believe, the ear hears by taking the sound wave f(x) which impinges on the ear drum and splitting it up into its Fourier components (3). It then analyzes each component anen+bn˜en—only considering the coefficients anandbn. These Fourier coefficients measure the intensity of the nth harmonic. Particular sounds are then heard in terms of the intensity of their various harmonics. We recognize familiar sounds by recognizing that the sound waves have similar Fourier coefficients. Amazing. It is time to consider the convergence of Fourier series. The question is: does the partial Fourier series PNf=a0e0+N/summationdisplay n=0anen+bn˜en converge to the function fasN→ ∞ . Since there are several norms, in particular the L2 norm /bardbl /bardbl and the uniform norm /bardbl /bardbl ∞, we must investigate convergence in each norm. Even though our proofs are reasonably slick, they are neither short nor particularly simple. A great deal of analytical technique will be needed. The proofs to be presented have been chosen because each of the devices invoked are important devices in their own right. We begin with some useful facts which have nothing especially to do with Fourier series. Theorem 3.22 (Weierstrass Approximation Theorem). If f(x)is continuous in the inter- val[−π,π]andf(−π) =f(π), then given any /epsilon1>0there is a trigonometric polynomial TN(x) =α0+N/summationdisplay n=1αncosnx+βnsinnx = ˆα0e0+N/summationdisplay n=1ˆαnen+ˆβn˜en,(3-27) (whereα0= ˆα0√ 2π, α n= ˆαn√π, β n=ˆβn√π), such that /bardblf−TN/bardbl∞= max −π≤x≤π|f(x)−TN(x)|</epsilon1. Note that the numbers αnandβnarenotnecessarily the Fourier coefficients of f. The proof, which is placed as an appendix at the end of this section, will indicate how they can be found. The following theorem states that convergence in the uniform norm implies convergence in theL2norm. Theorem 3.23 . Ifθ(x)is any bounded integrable function, then (if b>a ) /bardblθ/bardbl ≤√ b−a/bardblθ/bardbl∞. 3.4. FOURIER SERIES. 137 Proof: Since /bardblθ/bardbl∞= max x∈[a,b]|θ(x)|we find immediately that /integraldisplayb aθ(x)2dx≤/integraldisplayb a/bardblθ/bardbl2 ∞dx=/bardblθ/bardbl2 ∞/integraldisplayb adx= (b−a)/bardblθ/bardbl2 ∞ from which the conclusion is obvious. On geometrical grounds the theorem is even easier, since /bardblθ/bardbl∞is the greatest height of the curve θ(x) . Although convergence in the L2norm does notimply convergence in the uniform norm (the example in Section 1 comparing L1convergence and uniform convergence also works forL2), a useful weaker statement is true. Theorem 3.24 . (cf. Ex. 15 Section 3). If θ∈C1[a,b]andθ(x0) = 0 , wherex0∈[a,b], then for every x∈[a,b] |θ(x)| ≤√ b−a/radicalBigg/integraldisplayb aθ/primedt=√ b−a/bardblθ/prime/bardbl. Since the right side is independent of x, this implies that /bardblθ/bardbl∞=max x∈[a,b]|θ(x)| ≤√ b−a/bardblθ/prime/bardbl=√ b−a/bardblDθ/bardbl Proof: By the fundamental theorem of calculus, θ(x) =θ(x)−θ(x0) =/integraldisplayx x0θ/prime(t)dt. Thus the Cauchy-Schwarz inequality yields |θ(x)|2=/parenleftbigg/integraldisplayx x01·θ/prime(t)dt/parenrightbigg2 ≤/integraldisplayx x012dt/integraldisplayx x0θ/prime(t)2dt = (x−x0)/integraldisplayx x0θ/prime(t)2dt ≤(b−a)/integraldisplayb aθ/prime(t)2dt.(3-28) Therefore |θ(x)|2= (b−a)/bardblθ/prime/bardbl2. With these preliminaries behind us we turn to the convergence of Fourier series. First up is convergence in the L2norm. Theorem 3.25 . Assume fis continuous in the interval [−π,π]andf(−π) =f(π). Denote the sum of the first Nterms of its Fourier series by PNf. Then lim N→∞/bardblf−PNf/bardbl= lim N→∞/radicalBigg/integraldisplayπ −π[f(x)−(PNf)(x)]2dx= 0 138 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS Proof: Given any/epsilon1>0 , letTN(x) be the trigonometric polynomial given by Weierstrass Approximation Theorem. The trick is to apply Theorem 17. Using the NofTN, we know that PNf=a0e0+N/summationdisplay n=1anen+bn˜en and TN= ˆa0e0+N/summationdisplay n=1ˆαnen+ˆβn˜en. LetAbe the subspace of H=L2[−π,π] spanned by e0,e1,˜e1,...e N,˜eNThen both PNf andTNare inA. Thus by Theorem 17 of the last section (where slightly different notation was used), /bardblf−PNf/bardbl ≤ /bardblf−TN/bardbl, and by Theorem 20 ≤√ b−a/bardblf−TN/bardbl∞<√ b−a/epsilon1. Thus lim N→∞/bardblf−PNf/bardbl= 0, proving the theorem. Corollary 3.26 (Parseval’s Theorem). If f(x)is continuous in the interval [−π,π]and f(−π) =f(π), then /bardblf/bardbl2= lim N→∞/bardblPNf/bardbl2, that is,/integraldisplayπ −πf2(x)dx=a2 0+∞/summationdisplay n=1(a2 n+b2 n), where the Fourier coefficients ajandbjare determined by equations (2) or (2)’. Proof: The Corollary to Theorem 16 states that /bardblf/bardbl2=/bardblPNf/bardbl2+/bardblf−PNf/bardbl2. If we now let N→ ∞ , the second term on the right vanishes by the theorem just proved. Remark : The theorem and corollary state that the orthonormal set of functions e0= 1√ 2π, en(x) =cosnx√π, and ˜en(x) =sinnx√πis a complete orthonormal set for the scalar prod- uct spaceL2[−π,π] . The formula contained in the corollary is a generalization of the Pythagorean Theorem to L2[−π,π] . The proof of convergence in the uniform norm if the function has one continuous deriva- tive is only slightly more difficult. We shall need a preliminary Lemma 3.27 . Assumef∈C1[−π,π]. Extend it as a periodic function with period 2π byf(x+ 2π) =f(x). Let (PNf)be the sum of the first Nterms of its Fourier series. Then the sum of the first Nterms in the Fourier series for Df=d f dxisPN(Df), that is PN(Df) =D(PNf). This in not necessarily true for other bases in L2[−π,π]. 3.4. FOURIER SERIES. 139 Proof: We know that (PNf)(x) =a01√ 2π+N/summationdisplay n=1ancosnx√π+bnsinnx√π. Since we can differentiate a finite sum term by term, we find that D(PNf)(x) =N/summationdisplay n=1−nansinnx√π+nbncosnx√π, where theanandbnare found by using formulas (2)’. If PN(Df) =A01√ 2π+N/summationdisplay n=1Ancosnx√π+Bnsinnx√π, where theAnandBnare also found by using (2)’, we must show that A0= 0, A n=nbn,andBn=−nan But A0=/integraldisplayπ −π(Df(x))1√ 2πdx=1√ 2π[f(π)−f(−π)] = 0 since fis periodic. Integrating by parts, we further find An=/integraldisplayπ −π(Df(x))cosnx√πdx=n/integraldisplayπ −πf(x)sinnx√πdx=nbn and Bn=/integraldisplayπ −π(Df(x))sinnx√πdx=−n/integraldisplayπ −πf(x)cosnx√πdx=−nan. Our result is now only a few steps away. Theorem 3.28 . Iff∈C1[−π,π]and if both fandf/primeare periodic with period 2π, then the Fourier series PNfconverges to fin the uniform norm lim N→∞/bardblf−PNf/bardbl∞= 0. Proof: The key observation is that f/primeis a continuous function, so that Theorem 22 can be applied to its Fourier series. This shows that lim N→∞/bardblDf−PN(Df)/bardbl= 0. By the above lemma, D(f−Pnf) =Df−D(PNf) =Df−PN(Df). Thus lim N→∞/bardblD(f−PNf)/bardbl= 0. (3-29) 140 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS We would like to apply Theorem 21 to the function θN=f−PNf. In order to do so, we must only verify that θNvanishes somewhere in [ −π,π] . But the area under θN=f−PNf is /integraldisplayπ −πθN(x)dx=√ 2π/angbracketleftθN, e0/angbracketright= 0 sincef−PNfis orthogonal to the space spanned by e0,e1,˜e1,...,e N,˜eN(Theorem 16c). BecauseθN(x) is a continuous function (the difference of the C1) function fand the infinitely differentiable trigonometric polynomial PNf), the area under it can be zero only ifθNvanishes somewhere. Thus Theorem 21 is applicable and yields the inequality /bardblf−PNf/bardbl∞≤√ b−a/bardblD(f−PNf)/bardbl. We now pass to the limit N→ ∞ and use equation (4) to complete the proof of the theorem: lim N→∞/bardblf−PNf/bardbl∞≤lim N→∞√ b−a/bardblD(f−PNf)/bardbl= 0. Remarks . The hypothesis that f∈C1[−π,π] and is periodic with period 2 πhas been proved a sufficient condition for the Fourier series to converge to the function in the uniform norm. Much weaker hypotheses also suffice to prove the same result—but mere continuity is not enough. Convergence of Fourier series or generalizations thereof is a vast and deep subject, one still the object of intense study. On the basis of the theorems we have proved, many other problems are reasonably accessible—like the convergence of the Fourier series for a function which is nice except for a finite number of jump discontinuities. But there is not time for this pleasant excursion. a figure goes here 3.5 Appendix. The Weierstrass Approximation Theorem . The proof—which is difficult—will be given as a series of lemmas. Lemma 3.29 . Iff(x)is continuous and periodic with period 2π, then for any a∈R, the following equality holds /integraldisplaya+2π af(x)dx=/integraldisplay2π 0f(x)dx. Proof: This is clear from a graph of f, since the area under one period of fdoes not depend upon where you begin measuring. We also offer a computational proof. Write /integraldisplaya+2π af(x)dx=/integraldisplay0 af(x)dx+/integraldisplay2π 0f(x)dx+/integraldisplaya+2π 2πf(x)dx. Letx=t+2πin the last integral and use the fact that f(t+2π) =f(t) . The last integral is then −/integraldisplay0 af(t)dt, which cancels the unwanted term in the last equation and proves the lemma. 3.5. APPENDIX. THE WEIERSTRASS APPROXIMATION THEOREM 141 Lemma 3.30 ./integraldisplayπ/2 0cos2nt dt=1 2cn, wherecn=v1 π2·4·6···(2n) 1·3·5···(2n−1). Proof: A computation. Integrate by parts to show that I2n=/integraldisplayπ/2 0cos2ntdt= (2n−1)(I2n−2−I2n). ThusI2n=2n−1 2nI2n−2. Now induction can be used to do the rest, since by observation I0=π/2 . Lemma 3.31 . Assumef(x)is continuous and periodic with period 2π. Let TN(x) =cN 2/integraldisplayπ −πf(t) cos2N(t−x 2)dt (3-30) Then given any /epsilon1>0, there is an Nsuch that /bardblf−TN/bardbl∞= max −π≤x≤π|f(x)−TN(x)|</epsilon1. Proof: How did we guess the formula (4)? We observed that cos2Nxis one atx= 0 , and strictly less than one for all other x∈[−π,π] . Thus, for large N,cos2Nxis one at x= 0 , and decreases sharply thereafter so cos2N(t−x 2) has the same property at x−t= 0 , wherex=t. Then essentially the only values of f(t) which will count are those about t=x, so what comes out will be f(x) . Let us proceed with the details. Takes=t−x 2. Then TN(x) =cN/integraldisplayπ/2 −π/2f(x+ 2s) cos2Nsds. Split the integral into two pieces, from −π 2to 0 and from 0 toπ 2, and then replace sby −sin the first one. This gives TN(x) =cN/integraldisplayπ/2 0[f(x) + 2s) +f(x−2s)] cos2Nsds. From Lemma 2 we know that f(x) =cN/integraldisplayπ/2 02f(x) cos2Nsds, sincef(x) is a constant in the integration with respect to s. Therefore TN(x)−f(x) =cN/integraldisplayπ/2 0[f(x+ 2s)−2f(x) +f(x−2s)] cos2Nsds. Now given any /epsilon1 >0 , from the continuity of fwe can pick a δ >0 independent of x such that |f(x1)−f(x2)|</epsilon1 2when |x1−x2|<δ. This/epsilon1will be the /epsilon1of our conclusion. Break the integral into two parts, one from 0 to δ and the other from δtoπ/2 , whereδis theδwe just found. Then in the [0 ,δ] interval, |f(x+ 2s)−2f(x) +f(x−2s)| ≤ |f(x+ 2s)−f(x)|+|f(x)−f(x−2s)|</epsilon1, 142 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS while in the [ δ,π 2] interval, |f(x+ 2s)−2f(x) +f(x−2s)| ≤ |f(x+ 2s)|+ 2|f(x)|+|f(x−2s)| ≤4M, whereM= max x∈[−π,π]|f(x)|. Hence |f(x)−TN(x)|<cN[/epsilon1/integraldisplayδ 0cos2Nsds+ 4M/integraldisplayπ/2 0cos2Nsds]. Now we observe that /integraldisplayδ 0cos2Nsds</integraldisplayπ/2 0cos2Nsds=1 2cN, and that, since cos sdecreases as sgoes toπ/2 , /integraldisplayπ δcos2Nsds</integraldisplayπ/2 δcos2Nδds<π 2γN, whereγ= cos2δ<1 . Thus |f(x)−TN(x)|</epsilon1 2+ 2πMc NγN. NowπcN=/parenleftBig 2 3·4 5···2N−2 2N−1/parenrightBig ·2N < 2N, so that 2πMc NγN<4MNγN. Becauseγ <1 , we know that lim N→∞NγN= 0 . Thus, pick Nso large that NγN</epsilon1 8M, where this is the same/epsilon1as before. Consequently, for this N, |f(x)−TN(x)|</epsilon1. Since/epsilon1is independent of x, /bardblf(x)−TN(x)/bardbl= max x∈[−π,π]|f(x)−TN(x)|</epsilon1 too. A difficult lemma is thereby proved. The whole proof is completed in the following simple Lemma 3.32 . The function TN(x)defined by (3) TN(x) =cN 2/integraldisplayπ −πf(t) cos2N(t−x 2)dt is a trigonometric polynomial. Proof: This can be horribly messy unless one is shrewd. We shall use the formula eiθ= cosθ+isinθand the binomial theorem (top p. 108). First notice that cos2Nθ=/parenleftbiggeiθ+e−iθ 2/parenrightbigg2N =1 22N2N/summationdisplay k=0(2N)! (2N−k)k!eikθe−i(2N−k)θ. 3.5. APPENDIX. THE WEIERSTRASS APPROXIMATION THEOREM 143 Letdk= (2N)!/22N(2N−k)!k! Then cos2Nθ=2N/summationdisplay k=0dke−i(2N−2k)θ =2N/summationdisplay k=0dk[cos(2N−2k)θ−isin(2N−2k)θ].(3-31) Since cos2Nθis real, the sum of the imaginary terms on the right must be zero. Thus, replacing 2 θbyt−x, we find that cos2N(t−x )2 =2N/summationdisplay k=0dkcos(N−k)(t−x) =2N/summationdisplay k=0dk[cos(N−k)tcos(N−k)x+ sin(N−k)tsin(N−k)x].(3-32) Split the sum into two parts, one from 0 to N, the other from N+ 1 to 2N, and let n=N−kin the first, n=k−Nin the second. This gives cos2N(t−x 2) =N/summationdisplay n=0dN−n[cosntcosnx+ sinntsinnx] +N/summationdisplay n=1dN+n[cosntcosnx+ sinntsinnx],(3-33) so cos2N(t−x 2) =dN+N/summationdisplay n=1(dN+n+dN−n)[cosntcosnx+ sinntsinnx], which is much more simple than one might have anticipated. Substituting this into (4) and realizing that the tintegrations just yield constants, we find that TN(x) is indeed a trigonometric polynomial. Coupled with Lemma 3, the proof of Weierstrass’ Approximation Theorem is completely proved. Exercises (1) Find the Fourier series with period 2 πfor the given functions. (a)f(x) =/braceleftbigg0,−π≤x≤0 2,0<x<π (b)f(x) =/braceleftbigg−2,−π≤x<0 2,0≤x<π (c)f(x) = sin 17x+ cos 2s,−π≤x<π (d)f(x) = sin2x,−π≤x≤π; (e)f(x) =x2,−π≤x≤π 144 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS (f)f(x) =/braceleftbiggx+π,−π≤x≤0 −x+π,0≤x≤π(Also, compute /bardblf/bardbl2anda2 0+∞/summationdisplay n=1(a2 n+b2 n) for (a)-(f)). (2) (a) Apply Parseval’s Theorem (Corollary to Theorem 22) to the function f(x) =x and its Fourier series to deduce that π2 6= 1 +1 22+1 32+1 42+··· (cf. the example before Theorem 17 of Section 3). (b) Do the same for the function f(x) =x2(Ex. 1, e above) to evaluate 1 +1 24+1 34+1 44+···=? (3) A function f(x) iseven iff(−x) =f(x) ,oddiff(−x) =−f(x) . Thus 2 + x2is an even function, x3−sinxis an odd function, while 1 + xis neither even nor odd. Letanandbnbe the Fourier coefficients of the piecewise continuous function f(x) . Prove the following statements. (a) Iffis an oddfunction, an= 0, b n= 2/integraldisplayπ 0f(x)sinnx√πdx (b) Iffis an even function an= 2/integraldisplayπ 0f(x)cosnx√πdx, b n= 0 (c) A function fdefined in [0 ,π] may be extended to [ −π,π] as either an even or odd function by the formulas even extension :f(−x) =f(x), x≥0, or odd extension :f(−x) =−f(x), x≥0. The even extension of f(x) =x, x∈[0,π] isf(x) =|x|, x∈[−π,π] , while its odd extension is f(x) =x, x∈[−π,π] . The odd extension of f(x) =x2, x∈ [0,π] isf(x) =/braceleftbiggx2, x∈[0,π] −x2, x∈[−π,0]. Extend the function f(x) = 1, x∈[0,π] to the interval [ −π,π] as an odd function and sketch its graph. Find its Fourier series using part (a). (4) (a) Let f(x) be a given function. Find a solution of the O.D. E. u/prime/prime+λ2u=f, where λis a real number and u(x) satisfies the boundary condition u(−π) =u(π) = 0 , by the following procedure: Expand fin its Fourier series and assume uhas a Fourier series whose coefficients are to be found. Find a formula for the Fourier coefficients of uin terms of those for fin the case where λis not an integer. 3.5. APPENDIX. THE WEIERSTRASS APPROXIMATION THEOREM 145 (b) Ifλ=nis an integer, show that there is a solution if and only if 0 = /angbracketleftf,˜en/angbracketright=/integraldisplayπ −πf(x)sinnx√πdx. (5) (a) State Parseval’s Theorem for the special cases i) fis a continuous even function in [−π,π] , and ii)fis a continuous odd function in [ −π,π] . (b) Iffis a continuous even function in [ −π,π] and /integraldisplayπ 0f(x) cosnxdx = 0, n = 0,1,2,3,..., show thatf= 0 in [ −π,π] . (c) State and prove a theorem similar to (b) in the case of a continuous odd function. (6) In this exercise you show how a function f∈L2[−A,A] can be expanded in a modified Fourier series (so far we know only L2[−π,π] ). Lety=πx A—this maps the interval [−A,A] onto [ −π,π] —and define g(y) by f(x) =f(Ay π) =g(y) =g(πx A). Sinceg(y)∈L2[−π,π] , it can be expanded in a Fourier series g(y) =a01√ 2π+∞/summationdisplay n=1ancosny√π+bnsinny√π, where theanandbnare given by the usual formulas (2)’. (a) Prove that f(x)∈L2[−A,A] has the modified Fourier series f(x) =a01√ 2A+∞/summationdisplay n=1cosnx Ax+bn√ Asinnπ Ax, where a0=1√ 2A/integraldisplayA −Af(x)dx an=1√ A/integraldisplayA −Af(x) cosnπx Adx, b n=1√ A/integraldisplayA −Af(x) sinnπx Adx. (b) Find the modified Fourier series for f(x) =|x|, in the interval [ −1,1] . The following exercises all concern the Weierstrass Approximation Theorem. (7) Prove the following version of the Weierstrass Approximation Theorem. Let f∈ C[a,b] . Then given any /epsilon1>0 , there is a polynomial Q(x) such that /bardblf−Q/bardbl∞= max x∈[a,b]|f(x)−Q(x)|</epsilon1. (Hint: Let y=−π+2(x−a) b−aπ. This maps [ a,b] into [ −π,π] . Defineg(y), y∈[−π,π] by f(x) =f(a+(b−a) 2π(y+π) =g(y) =g(−π+ 2(x−a) b−aπ). 146 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS Use the version of the theorem proved to approximate g(y), y∈[−π,π] by a trigono- metric polynomial TN(y) to within /epsilon1/2 . Then approximate sin nyand cosnyto withinc/epsilon1(you pickc) by a finite piece of their Taylor series—which are polynomials. Put both parts together to obtain the complete proof for g(y) . The transition back tof(x) is trivial.] (8) (Riemann-Lebesgue Lemma). Let f∈C[a,b] . Prove that lim λ→∞/integraldisplayb af(x) sinλxdx = 0. [Hint: Integrate by parts to prove it first for all f∈C1[a,b] . For arbitrary f, approximate fby a polynomial—Ex. 7 above—to within /epsilon1/2 and realize that every polynomial is in C1[a,b] ]. (9) Iff∈C[0,1] , prove that lim n→∞n/integraldisplay1 0f(x)xndx=f(1). [Hint: Use the hint in Ex. 8]. (10) Iff∈C[a,b] , and if /integraldisplayb af(x)xndx= 0, n = 0,1,2,3,..., show that f= 0 . [ Hint: This implies that/integraldisplayb af(x)Q(x)dx= 0 , where Qis anypolynomial. fcan be approximated by some polynomial ˜Q. Now show that/integraldisplayb af2(x)dx= 0 .] 3.6 The Vector Product in R3 . As you grasped many years ago, the world we live in has three space dimensions. For this reason the material in this section is important in many applications. What we intend to do is define a way to multiply two vectors XandYinR3. Whereas the scalar product /angbracketleftX, Y/angbracketrightis ascalar , this product X×Y, the vector product , orcross product as it is often called, is a vector . For several reasons [i) we shall not cover this in class, and ii) I can probably not do as good a job as appears in many books] we shall let you read about this topic elsewhere. But make sure to read about it even though you’ll never be examined on it. Chapter 4 Linear Operators: Generalities. V1→Vn, Vn→V1 4.1 Introduction. Algebra of Operators . LetVby a linear space. So far we have considered the algebraic structure of such a space; however most significant reason for studying linear spaces is so that one can study operators defined on them. Operator is another, more organic, name for function. Thus an operator T:A→B Tmaps elements in its domain Ainto elements of B, whereBcontains the range of T. IfX∈A, thenT(X) =Y∈B. Think of feeding Xinto the operator T, andYbeing a figure goes here whatTsends out in return. It is useful to think of Tas some type of machine or factory, the input (raw material) is X, and the output is Y. Some examples should illustrate the situation and its potential power. Examples: (1) Let V=R2. IfX= (x1,x2)∈R2, andY= (y1,y2,y3) , we define T(X) =Yby T(X) =  x1+ 2x2=y1 x1+x2=y2 3x1+x2=y3  , or T(X) =T(x1,x2) = (x1+ 2x2, x1+x2,3x1+x2) = (y1,y2,y3) =Y. This operator Thas the property that to every X∈R2it assigns a Y∈R3. In other words Tmaps the two dimensional space R2into the three dimensional space R3 T:R2→R3. 147 148 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 R2is the domain ofT, denoted by D(T) , while the range ofT,R(T) is contained inR3, D(T) =R2,R(T)⊂R3. Sincey1=y2= 0 implies that x1=x2= 0 , which in turn implies that y3= 0 , we see that the point (0 ,0,1)∈R3isnotin the range of T. Thus,Tisnot surjective onto R3. It is injective (one-to-one) since every point Y∈R(T) is the image of exactly one X∈D(T) . This an be seen by observing that y1andy2suffice to determineX= (x1,x2) uniquely by solving the first two equations −y1+ 2y2=x1 y1−y2=x2. Hence ifY=T(X1) and also Y=T(X2) , thenX1=X2. Since the operator Tis completely determined by the coefficients in the equations, it is reasonable to represent this Tby the matrix T= 1 2 1 1 3 1  If you care to think of Xas the input into a paint-making machine, then x1might represent the quantity of yellow and x2the quantity of blue used. In this case y1,y2 andy3represent the quantities of three different shades of green the machine yields. For this machine, as soon as you specify the desired quantities of any two of the greens, say y1andy2, the quantities x1andx2of the input colors are completely determined, as is the quantity y3of the remaining shade of green. (2) Let VbeR2again. With X= (x1,x2)∈R2, andY= (y1)∈R1, defineTby x2 1+x2 2=y1, or T(X) =x2 1+x2 2. This operator Tmaps R2intoR1 T:R2→R1. It is not surjective onto R1since the negative half of R1is completely omitted from R(T) . Furthermore, it is not injective either since each point y1∈R(T) other than zero is the image of infinitely many points—all of those on the circle x2 1+x2 2=y1. (3) Let VbeC[−1,1] . Iff∈C[−1,1] , we define Tby T(f) =f(0). Thus, iff(x) = 2 + cos x, thenTf= 3 . This operator Tis usually denoted by δand called the Dirac delta functional . It was first used by Dirac in his work on quantum mechanics and is extremely valuable in modern mathematics and physics. Tassigns to each continuous function fits value at x= 0 , a real number. Therefore T:C[−1,1]→R1. 4.1. INTRODUCTION. ALGEBRA OF OPERATORS 149 The operator Tis not injective, since for example the element 2 ∈R1is the image of bothf(x) = 1 +exandf(x) = 2 . It is surjective since every element a∈R1is the image of at least one element in C[−1,1] (iff(x)≡a, then clearly T(f) =a). (4) Let VbeC[−1,1] . Iff∈C1[−1,1] then the differentiation operator Dis defined by (Df)(x) =df dx(x). It maps each function into its derivative. If f(x) =x2, then (Df)(x) = 2x. Since the derivative of a continuously differentiable function (a function in C1) is necessarily continuous, we see that D:C1[−1,1]→C[−1,1]. Dis not injective since, for example, the function g(x) = 1 is the image of both f1(x) =xandf2(x) = 2 +x.Dis surjective onto C[−1,1] . R(D) =C[−1,1], since ifg(x) is any element of C[−1,1] , thengis the image of the particular function f∈C1[−1,1] defined by f(x) =/integraldisplayx 0g(s)ds, becauseDf=gby the fundamental theorem of calculus. Throughout this and the next chapter we will study some of the elementary aspects of linear operators. It is reasonable to denote a linear operator by L. Definition Let V1andV2both be linear spaces over the same field of scalars. An operatorLmapping V1intoV2is called a linear operator if for every Xand ˜X inV1and any scalar a,Lsatisfies the two conditions 1.L(X+˜X) =L(X) +L(˜X) 2.L(aX) =aL(X). Whenever ambiguity does not arise, we will omit the parentheses and write LX instead ofL(X) . An equivalent form of the definition is Theorem 4.1 .Lis a linear operator ⇐⇒ L(aX+b˜X) =aL(X) +bL(˜X), whereX,˜X∈V1andaandbare any scalars. Proof: ⇒ L(aX+b˜X) =L(aX) +L(b˜X) (property 1) =aLX +bL˜X(property 2) .(4-1) ⇐Property 1 is the special case a=b= 1 . Property 2 is the special case b= 0 . Remark: It is useful to observe that always L(0) =L(0·X) = 0L(X) = 0 . This identity is often the easiest way to test if an operator is notlinear. 150 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 Examples: (1) The operator Ldefined by example 1 where L:R2→R3is LX= (x1+ 2x2,x1+x2,3x1+x2), is linear. Let X= (x1,x2) and ˜X= (˜x1,˜x2) . Then L(X+˜X) = (x1+ ˜x1+ 2x2+ 2˜x2,x1+ ˜x1+x2+ ˜x2,3x1+ 3˜x1+x2+ ˜x2) = (x1+ 2x2,x1+x2,3x1+x2) + (˜x1+ 2˜x2,˜x1+ ˜x2,3˜x1+ ˜x2) =LX+L˜X and L(aX) = (ax1+ 2ax2,ax 1+ax2,3ax1+ax2) =a(x1+ 2x2,x1+x2,3xz+x2) =aLX. (4-2) (2) The operator TX=x2 1+x2 2with domain R2and range R1isnotlinear, since T(aX) = (ax1)2+ (ax2)2=a2[x2 1+x2 2]/negationslash=aTX except for the particular scalars a= 0,1 . (3) The operator Df=d f dxwith domain C1[−1,1] and range C[−1,1] is linear since if f1andf2are inC1[−1,1] andaandbare any real numbers, then by elementary calculus D(af1+bf2) =d dx(af1+bf2) =adf1 dx+bdf2 dx =aDf 1+bDf 2.(4-3) (4) The operator Ldefined as Lu=a2(x)u/prime/prime+a1(x)u/prime+a0(x)u,(/prime=d dx), whereu(x)∈D(L) =C2, and where a0(x) ,a1(x) , anda2(x) are continuous functions, is a linear operator, L:C2→C. IfAandBare any constants (scalars for C2), then for any u1andu2∈C2, L(Au1+Bu2) =ax[Au1+Bu2]/prime/prime+a1[Au1+Bu2]/prime+a0[Au1+Bu2] =a2Au/prime/prime 1+a2B/prime/prime 2+a1Au/prime 1+a1Bu/prime 2+a0Au1+a0Bu2 =A[a2u/prime/prime 1+a1u/prime 1+a0u1] +B[a2u/prime/prime 2+a1u/prime 2+a0u2] =ALu 1+BLu 2.(4-4) 4.1. INTRODUCTION. ALGEBRA OF OPERATORS 151 (5) The identity operator Iis the operator which leaves everything unchanged. Because it is so simple, it can be defined on an arbitrary set Sand mapsSintoitselfS→S in a trivial way. If X∈S, then we define IX=X. What could be more simple? If Sis a linear space V(soaXandX1+X2are defined), then Iis trivially a linear operator, since I(aX1+bX2) =aX1+bX2=aIX 1+bIX 2 Why are linear operators important? There are several reasons. First, they are much easier to work with than nonlinear operators. Second, most of the operators which arise in applications are linear. The feature possessed by linear operators which is central to applications is that of superposition . IfLu1=fandLu2=g, thenL(u1+u2) =f+g. In other words, if u1is the response to some external influence fandu2the response to g, then the response to f+gis found by adding the separate responses. The special case of a linear operator whose range is the real number line R1arises often enough to receive a name of its own. Definition: A linear operator whose range is R1is called a linear functional ,/lscriptV→R1. The Dirac delta functional is such an operator. So is the operator l(f) =/integraldisplay1 0f(x)dx, which assigns to every continuous function f∈C[0,1] the real number equal to the area between the graph of fand thex-axis. Check that /lscriptis linear. If the linear operator L:V1→V2the range of L—a subset of the linear space V2—has a particularly nice structure. In fact, R(L) is not just any clump of points in V2but Theorem 4.2 . The range of a linear operator L:V1→V2is a linear subspace of V2. Remark: Even more is true. We shall prove (p. 312-3) that dim R(L)≤dimD(L) so that no matter how large V2is, the range has at most the same dimension as the domain. Proof: The range of Lconsists of all elements Y∈V2of the form Y=LX where X∈V1. We know that R(L) is a subset of the linear space V2. The only task is to prove that it is actually a subspace. Since V2is a linear space, it is sufficient to show that the setR(L) is closed under multiplication by scalars, and under addition of vectors. i) R(L) is closed under multiplication by scalars. If Y∈R(L) , there is an X∈V1=D(L) such thatY+LX. We must find some ˜XinV1such thataY=L˜X, whereais any scalar. SinceaY=aLX =L(aX) , we take ˜X=aX. ii)R(L) is closed under addition of vectors. If Y1andY2are in R(L) , there are elements X1andX2inV1=D(L) such that Y1=LX 1andY2=LX 2. We must show that Y1+Y2∈D(L) , that is, find some ˜X∈V1such thatY1+Y2=L˜X. ButY1+Y2= LX 1+LX 2=L(X1+X2) . Thus we can take ˜X=X1+X2. Before moving further on into the realm of special linear operators, we shall take this opportunity to define algebraic operations (addition and multiplication) for linear operators. But first we define equality, L1=L2, in a straightforward way. 152 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 Definition: (equality ) IfL1andL2both mapV1intoV2, whereV1andV2are linear spaces, and if L1X=L2Xfor allXinV1, thenL1equalsL2. Thus, two operators are equal if they have the same effect on any vector. Addition is equally simple. Definition: (addition ). IfL1:V1→V2andL2:V1→V2then their sum, L1+L2, is defined by the rule (L1+L2)X=L1X+L2X, X ∈V1 Examples: (1) LetL1:R2→R3be defined by L1(X) = (x1+x2,x1+ 2x2,−x2), X = (x1,x2)∈R2 andL2:R2→R3be defined by L2X= (−3x1+x2,x1−x2,x1), X = (x1,x2)∈R2. ThenL1+L2is defined, and is (L1+L2)X+L1X+L2X= (x1+x2,x1+ 2x2,−x2) + (−3x1+x2,x1−x2,x1) = (−2x1+ 2x2,2x1+x2,x1−x2) (4-5) (2) LetD:C1→Cbe defined by Du=du dxu∈C1, andL:C1→Cbe defined by Lu=/integraldisplay1 0ex−tu(t)dt u ∈C1 =ex/integraldisplay1 0e−tu(t)dt.(4-6) (In reality, Lmay be defined on a much larger class of functions— u∈Cis plenty, while its image is the smaller space, constant ex⊂C. We have decided on the smaller domain and larger image space so that the sum D+Lis defined). Then for any u∈C1. (D+L)u=Du+Lu=du dx+/integraldisplay1 0ex−tu(t)dt. The following theorem is a statement of some simple facts about the sum of two linear operators. Theorem 4.3 . LetL1,L2,L3,... be any linear operators which map V1→V2, so that their sums are defined. Then 0.L=L1+L2is a linear operator (1)L1+ (L2+L3) + (L1+L2) +L3, 4.1. INTRODUCTION. ALGEBRA OF OPERATORS 153 (2)L1+L2=L2+L1 (3)Let 0 be the operator which maps every element of V1into 0∈V2, so 0X= 0. Then L1+ 0 =L1. (4)L1+ (−L1) = 0 . Here −L1is the operator which maps every element X∈V1 into−(L1X). Proof: These are just computations. Let X1,X2∈V1. 0. L(aX1+bX2) = (L1+L2)(aX1+bX2) =L1(aX1+bX2) +L2(aX1+bX2) =aL1X1+bL1X2+aL2X1+bL2X2 =a(L1X1+L2X1) +b(L1X2+L2X2) =a(L1+L2)X1+b(L1+L2)X2 =aLX 1+bLX 2.(4-7) (1) (L1+(L2+L3))X=L1X+(L2+L3)X=L1X+L2X+L3X= (L1+L2)X+L3X= ((L1+L2) +L3)X. (2) (L1+L2)X=L1X+L2X=L2X+L1X= (L2+L1)X. The step L1X+L2X= L2X+L1Xis justified on the grounds that the vectors Y1:=L1XandY2:=L2X are elements of V2—which is a linear space—so that Y1+Y2=Y2+Y1. (3) (L1+ 0)X=L1X+ 0X=L1X+ 0 =L1X Note that the 0 in 0 Xis an operator, while the 0 in the next step is an element of V2. This ambiguity causes no trouble once you understand it. (4) (L1+ (−L1))X=L1X+ (−L1)X=L1X−L1X= 0 The crucial step ( −L1)X=−L1Xis the definition of the operator ( −L1) . Remark: This theorem states that the set of all linear operators mapping one linear space V1into another V2form an abelian group under addition. Multiplication of operators is not much more difficult. If L1andL2are linear opera- tors, then their product L2L1in that order is defined by the rule L2L1X=L2(L1X) . In other words, firstoperate on XwithL1giving a vector Y=L1X.Then operate on this new vector YwithL2, givingL2Y=L2(L1X) . It is clear that in order for this to make sense, for every X∈D(L1) , the new vector Y=L1Xmust be in the domain of L2. Thus to form the product L2L1, we require that R(L1)⊂D(L2) . Look at our machine again. a figure goes here 154 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 The multiplication L2L1means sending the output from L1as input into L2. In order to join the machines in this way, surely one necessary requirement is that L2is equipped to act on the output from R(L1) , that is, R(L1)⊂D(L2) . Of course the L2machine might be able to digest input other than what L1sends out. But all we care is that L2 can digest at least whatL1sends it. Definition: (multiplication ). LetL1:V1→V2andL2:V3→V4. If the range of L1 is contained in the domain of L2,R(L1)⊂D(L2) , then the productL2L1is definable by the composition rule L2L1X=L2(L1X),whereX∈V1=D(L1). The product L2L1maps the input V1forL1into the output V4forL2, L2L1:V1→ V3→V4. We exhibit a little diagram (cf. p. ???). a figure goes here The way to get from V1toV4usingL2L1is to first use L1to reachV2. Then use L2 to get toV4. Remarks: IfL2L1is defined, it is notnecessarily true that L1L2is defined (Example 1 below). Furthermore, even if L1L2is also defined, it is only a rare coincidence that multiplication is commutative. Usually L2L1/negationslash=L1L2when both products are defined. Thus the orderL2L1isimportant . Examples: (1) LetL1:R2→R3be defined as L1X= (x1−x2,x2,−x1−2x2),whereX= (x1,x2)∈R2, and letL2:R3→R1be defined as L2Y= (y1+ 2y2−y3),whereY= (y1,y2,y3)∈R3. Then R(L1)⊂R3=D(L2) so that the product L2L1is definable and L2L1:R2L1→ R3L2→R1. Consider what L2L1does to the particular vector X0= (−1,2)∈R2. L2L1X0=L2(L1X0) =L2(−3,2,−3) = ( −3 + 4 + 3 = 4) ThusL2L1maps ( −1,2)∈R2into 4 ∈R1. More generally, if Xis any vector in R2, L2L1X=L2(L1X) =L2(x1−x2,x2,−x1−2x2) = (x1−x2+ 2x2+x1+ 2x2) = 2x1+ 3x2∈R1.(4-8) ThusL2L1maps (x1,x2)∈R2into 2x1+ 3x2∈R1. Since R(L2) =R1andD(L1) =R2,R(L2) not ⊂D(L1) so that the product L1L2 isnotdefined. You might be thinking that R1is part of R2. What you mean is that R2has one dimensional subspaces. It certainly does—an infinite number of them, all of the straight lines through the origin. Because there are so many subspaces of R2 4.1. INTRODUCTION. ALGEBRA OF OPERATORS 155 which are one dimensional, there is no natural way of regarding R1as being contained inR2. [On the other hand, there is a natural way in which C1can be regarded as contained in C. We used this above in our second example for addition of linear operators]. (2) Define L1:R2→R2by the rule L1X= (2x1−3x2,−x1+x2) andL2:R2→R2by the ruleL2X= (2x2,x1+x2) Then R(L1) =R2=D(L2) so thatL2L1is defined. It is given by L2L1X=L2(2x1−3x2,−x1+x2) = (−2x1+ 2x2,x1−2x2) In particular, L2L1mapsX0= (1,2) into (2,−3) . Now R(L2) =R2=D(L1) , so thatL1L2is also definable. It is given by L1L2X=L1(2x2,x1+x2) = (2·2x2−3·(x1+x2),−2x2+ (x1+x2)) = (−3x1+x2,x1−x2).(4-9) In particular, L1L2mapsX0= (1,2) into ( −1,−1) . SinceL1L2andL2L1map the pointX0= (1,2) into two different points, it is clear that L1L2/negationslash=L2L1, the operators do not commute. (3) LetAbe the subspace of R2spanned by some unit vector e1andBbe the subspace spanned by another unit vector e2. Consider the projection operators PAandPB. They are linear since, for example, PA(aX1+bX2) =/angbracketleftaX1+bX2, e1/angbracketrighte1 =a/angbracketleftX1, e1/angbracketrighte1+b/angbracketleftX2, e1/angbracketrighte1 =aPAX1+bPAX2.(4-10) BecausePA:R2→R2andPB:R2→R2, both products PAPBandPBPAare defined. We have PAPBX=PA(PBX) =PA(/angbracketleftX, e 2/angbracketrighte2) =/angbracketleftX, e 2/angbracketrightPAe2=/angbracketleftX, e 2/angbracketright/angbracketlefte2, e1/angbracketrighte1.(4-11) Also, PBPAX=PB(PAX) =PB(/angbracketleftX, e 1/angbracketrighte1) =/angbracketleftX, e 1/angbracketrightPBe1=/angbracketleftX, e 1/angbracketright/angbracketlefte1, e2/angbracketrighte2.(4-12) SincePAPBX∈A⊂R2, whilePBPAX∈B⊂R2, it is clear that usually PAPB/negationslash= PBPA. They will happen to be equal if A=B, or ifA⊥B(for thenPAPB= PBPA= 0 ). See the figure at the beginning of this example—and draw some more special cases for yourself. (4) LetL:C∞→C∞(C∞is the space of infinitely differentiable functions) be defined by (Lu)(x) =xu(x), u∈C∞, 156 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 andD:C∞→C∞be defined by (Du)(x) =du dx(x), u∈C∞. Then R(L) =D(D) so that the product DLis definable by DLu =D(Lu) =D(xu) =d dx(xu(x)) =xu/prime+u. Also, R(D) =D(L) soLDis definable by LDu =L(Du) =L(u/prime) =xu/prime. Notice that LD/negationslash=DLunlessu= 0 . We collect some properties of multiplication. Theorem 4.4 . IfL1:V1→V2, L2:V3→V4, andL3:V5→V6, whereV1⊂V3and V4⊂V5, then 0. The operator L=L2L1is a linear operator. 1.L3(L2L1) = (L3L2)L1—Associative law. Proof: 0. L(aX1+bX2) =L2/parenleftbig L1(aX1+bX2)/parenrightbig =L2/parenleftbig aL1X1+bL1X2/parenrightbig =L2(aL1X1) +L2(bL1X2) =aL2L1X1+bL2L1X2 =aLX 1+bLX 2.(4-13) (1) By definition of the product, [L3(L2L1)]X=L3[(L2L1)X] =L3[L2(L1X)] and [(L3L2)L1]X= (L3L2)(L1X) =L3[L2(L1X)]. Now match the ends. Notice that the commonly occurring special case V1=V2=V3=V4=V5=V6is included in this theorem. In this special case, even more can be proved. For then the identity operator I, defined by IX=Xfor allX∈Vcan be used to multiply any other operator. Moreover, addition, L1+L2also makes sense. Theorem 4.5 . If the linear operators L1,L2,L3all mapVintoV, then representing any one of these by L, (1)LI=IL=L. (2)For any positive integer n, we define Lninductively by the rule Ln+1=LLn, andL0=I. Then for any non-negative integers mandn, Lm+1=LmLn. 4.1. INTRODUCTION. ALGEBRA OF OPERATORS 157 (3) (L1+L2)L3=L1L3+L2L3. (4)L3(L1+L2) =L3L1+L3L2(This is needed in addition to 3 because of the non- commutativity). Proof: (1) IfX∈V, (LI)X=L(IX) =LX (IL)X=I(LX) =LX (2) We shall prove Lm+n=LmLnby induction on m. The statement is true, by definition, for m= 1 . Assume it is true for m=k, soLk+n=LkLn. Our job is to prove the statement for m=k+ 1 . By the definition and the induction hypothesis, we have Lk+n+1=LLk+n=L(LkLn). Since multiplication is associative, we find that L(LkLn) = (LLk)Ln. But, by definition, LLk=Lk+1. Thus, Lk+n+1=Lk+1Ln. This completes the induction proof. (3) IfX∈V, [(L1+L2)L3]X= (L1+L2)(L3X) LetL3X=Y∈V. Then (L1+L2)Y=L1Y+L2Y.Thus [(L1+L2)L3]X=L1(L3X) +L2(L3X) = (L1,L3)X+ (L2L3)X. (4) Same proof as 3. Remark: IfV1andV2are two linear spaces, the set of all linear operators which mapV1intoV2is usually denoted by Hom( V1,V2) —Hom rhymes with Mom and Tom. In this notation, the last theorem concerned Hom( V,V) . The abbreviation Hom is for the impressive word “homomorphism”. Tell your friends. Examples: ConsiderD:C∞→C∞defined by ( Du)(x) =du dx(x) . ThenDn= dn dxn. Exercises (1) Determine which of the following are linear operators. 158 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 (a)T:R2→R2 TX= (x1+x2,x1−x2), whereX= (x1,x2)∈R2. (b)T:R2→R2, T(X) = (x1+x2+ 1,x1−x2) (c)T:R3→R2 T(X) = (x1+x1x2,x2) (d)T:R3→R1 T(X) = (x1+x2−x3) (e)T:R3→R1 T(X) + (x1+x2−x3+ 2) (f)D:P2→P1. IfP(x) =a2x2+a1x+a0∈P2then D(P) = 2a2+a1∈P1. (g)T:C1[−1,1]→R1. Ifu(x)∈C1[−1,1] , then T(u) =u(0) +u/prime(0). (h)T:C[2,3]→C[2,3] . Ifu∈C[2,3] , (Tu)(x) =/integraldisplay3 2ex−tu(t)dt (i)T:C[2,3]→C[2,3] , (Tu)(x) = 1 +/integraldisplay3 2ex−tu(t)dt (j)T:C[2,3]→C[2,3], (Tu)(x) =/integraldisplay3 2ex−tu2(t)dt (k)S1:C[0,∞]→C[0,∞] (S1u)(x) =u(x+ 1)−u(x) (l)L:A→C[0,∞] , whereA={u∈C[0,∞]:/integraltext∞ 0|u(t)|dt<∞}, (Lu)(x) =/integraldisplay∞ 0e−xtu(t)dt, [Our restriction on Ais just to insure that the integral exists. Luis usually called the Laplace transform ofu]. 4.1. INTRODUCTION. ALGEBRA OF OPERATORS 159 (m)T:C[0,∞]→C[0,∞] (Tu)(x) =a2u(x2) +a1u(x+ 1) +a0u(x), where theak(x) are continuous functions. (n)T:C[0,1]→C[0,1] . (Tu)(x) = 2xu(x). (o)T:R2→R1 TX=|x1+x2|,whereX= (x1,x2)∈R2. [Answers : a,d,f,g,h,k,l,m,n are linear]. (2) (a) If l(x) is a linear functional mapping R1→R1, prove that l(x) =αx, where α=l(1) . (b) Ifl(X) is a linear functional mapping Rn→R1, prove that l(X) =n/summationdisplay k=1αkxk, whereX= (x1,...,x n) . (3) LetL1:R1→R2be defined by L1X= (x1,3x1),whereX= (x1)∈R1, andL2:R2→R2be defined by L2Y= (y1+y2,y1+ 2y2),whereY= (y1,y2)∈R2. ComputeL2L1X0, whereX0= 2∈R1. IsL1L2defined? (4) LetA:R2→R2be defined by AX= (x1+ 3x2,−x1−x2), X = (x1,x2)∈R2 andB:R2→R2by BX= (−x1+x2,2x1+x2). a). Compute ABX,BAX,B2X,A2BX, and (A+B)X. b). Find an operator Csuch thatCA=I. [hint: LetCX= (c11x1+c12x2,c21x1+ c22x2) and solve for c11,c12, etc.] (5) Consider the operators D:C∞→C∞,(Du) =u/primeandL:C∞→C∞,(Lu)(x) =/integraltextx 0u(t)dt. (a) Show that DL=I, LD =I−δ, whereδis the delta functional. (L2u)(x) =/integraldisplayx 0/parenleftbigg/integraldisplay2 0u(t)dt/parenrightbigg ds. Integrate by parts to conclude that (L2u)(x) +/integraldisplayx 0(x−t)u(t)dt. 160 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 (b) Observe that D2L2=D(DL)L=DIL =DL=I. Use this observation to find a solution of the differential equation D2u=fforu, wheref∈C∞. Solve the particular equation ( D2u)(x) =1 1+x2 (6) LetA:R2→R2be defined by AX= (a11x1+a12x2,a21x1+a22x2), andB:R2→R2be defined by BX= (b11x1+b12x2,b21x1+b22x2). (a) Compute AB. (b) Find a matrix Bsuch thatAB=I, that is, determine b11,b12,... in terms of a11,a12,...such thatAB=I. [In the course of your computation, I suggest in- troducing a symbol, say ∆ , for a11a12−a12a21when that algebraic combination crops up.] (7) In the plane E2, consider the operator Rwhich rotates a vector by 90oand the operatorPprojecting onto the subspace spanned by e(see fig). (a) Prove that R is linear. (b). Let X= (x1,x2) be any point on E2. Compute PRX andRPX . Draw a sketch for the special case X= (1,1) . (8) In R3, letAdenote the operator of rotation through 90oabout the x1-axis (so A: (0,1,0)→(0,0,1) ),Bthe operator of rotation through 90oabout thex2-axis andCthe operator of rotation through 90oabout thex3-axis (see fig.) Prove these operators are linear (just do it for A). Show that A4=B4=C4=I, AB /negationslash=BA, and thatA2B2=B2A2. Is it true that ABAB =A2B2? (9) Let Pdenote the linear space of all polynomials in x. Forp∈P, consider the operatorsDp=dp dxandLp=xp. Show that DL−LD=I. (10) (a) If L1L2=L2L1, prove that (L1+L2)2=L2 1+ 2L1L2+L2 2. (b) IfL1L2/negationslash=L2L1, then (L1+L2)2=? (11) IfL1andL2are operators such that L1L2−L2L1=I, prove the formula Ln 1L2− L2Ln 1=nLn−1 1, wheren= 1,2,3,.... (12) IfL1is a linear operator, L1:V1→V2[orL1∈Hom(V1,V2) ], andais any scalar, define the operator L=aL1by the rule LX= (aL1)X=a(L1X) , whereX∈V1. Prove (0).L=aL1is a linear operator, L:V1→V1. (5).a(bL1) = (ab)L1, wherea,bare any scalars. (6). 1 ·L1=L1. (7). (a+b)L1=aL1+bL1. (8).a(L1+L2) =aL1+aL2, whereL2∈Hom(V1,V2) . 4.2. A DIGRESSION TO CONSIDER AU/prime/prime+BU/prime+CU=F. 161 Coupled with Theorem 3, this exercise proves that the set of all linear operators map- ping one linear space in to another linear is itself a linear space , that is, Hom ( V1,V2) is a linear space . (13) (a). In E2, letLdenote the operator which rotates a vector by 90o. ThenL:E2→ E2. IfX= (x1,x2) =x1e1+x2e2, wheree1= (1,0) ande2= (0,1) , writeLas LX= (a11x1+a12x2,a21x1+a22x2), That is, find the coefficients a11,a12,.... This gives two ways to represent L, as a rotation (geometrically), and by linear equations in terms of a particular basis (algebraically). (b). In E2, consider the operator Lof rotation through an angle α. Show that Le1= (cosα,sinα), Le 2= (−sinα,cosα), and then deduce that if X= (x1,x2) =x1e1+x2e2, LX= (x1cosα−x2sinα, x 1sinα+x2cosα). (14) Consider the space Pnof all polynomial of degree n. DefineL:Pn→Pnas the translation operator ( Lp)(x) =p(x+ 1) , and D:Pn→Pnas the differentiation operator, ( Dp)(x) =dp dx(x) . Show that L=I+D+D2 2!+···+Dn−1 (n−1)!+Dn n! (15) Consider the linear operators L1=a1D2+b1D+c1I, andL2=a2D2+b2D+e2I. BothL1andL2map the linear space of infinitely differentiable function into itself, Lj:C∞→C∞. If the coefficients a1,a2,b1,... areconstants , prove that L1L2= L2L1. 4.2 A Digression to Consider au/prime/prime+bu/prime+cu=f. Essentially the only linear equation youcan solve explicitly are linear algebraic equations, like two equations in two unknowns. Since our theory applies to much more general situa- tions, we shall develop a different example for you to keep in the back of your minds along with that of linear algebraic equations. The example we have chosen has the additional virtue that it contains most of the solvable differential equations which arise anywhere. Watch closely because we shall be brief and with a high density of valuable ideas. Problems concerning vibration or oscillatory phenomena are among the most important and significant ones which arise in applications. The simplest case is that of a simple harmonic oscillator . We have a figure goes here 162 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 a massmattached to a spring. Pull the mass back a little and watch it move back and forth, back and forth. These are oscillations. To make the situation simple, we assume that the spring has no mass and that the surface upon which the mass rests is frictionless. Letu(t) denote the displacement of the center of gravity of the mass from the equilibrium position. Two experimental results are needed from physics. 1.Newton’s Second Law :m...u=/summationtextF, where/summationtextFmeans the resultant of all the forces on the center of gravity of the mass (we assume all forces are acting horizontally). 2.Hooke’s Law : If a spring is not stretched too far, then the force it exerts is propor- tional to the displacement, F=−ku, k> 0. We chose the minus sign since if a spring is displaced, the force it exerts is in the direction opposite to the displacement. [Under larger displacements, actually F(u) =a1u+a2u2+a3u3+... -wherea0=F(0) = 0 . If the displacement uis small, the lowest term in the Taylor series forF(u) gives an adequate approximation. This is a more precise statement of Hooke’s Law]. Putting these two results together, we find that m¨u=−ku+F1, (notation : ¨ u=d2u dt2) whereF1represents all of the remaining forces on the mass. One possible force (so far incorporated into F1) is a so-called viscous damping force . It is of the form Fv=−µ˙uwhere µ>0 ; at low velocities, this force is experimentally found to account for air resistance. It is directed opposite to the velocity, and increases as the speed does (speed = /bardblvelocity /bardbl). [Again,Fv=b1¨u+b2˙u2+..., that isFv( ˙u) is given by a Taylor series with Fv(0) = 0 . At low speeds, the higher order terms can be neglected to yield a reasonable approximation.] Thus, to our approximation, m¨u=−ku−µ˙u+F2, whereF2represents the forces yet unaccounted for. Let us assume that these remaining forces do not depend on the motion and are applied by the outside world. Then the force F2depends only on time, F2=f(t) . It is called the applied orexternal force . Newton’s law gives m¨u=−ku−µ˙u+f(t), or Lu: =a¨u=b˙u+cu=f(t), wherea=m,b=µ, andc=k. For the purposes of our discussion, we shall assume that kandµdo not depend on time. Then a,bandcare non negative constants. In order to determine the motion of the mass, we must solve the ordinary differential equationLu=fforu. Have we given enough information to determine the solution? In other words, is the solution unique? For any physically reasonable problem, we expect the mathematical model has a unique solution since (neglecting quantum mechanical effects) once we let the mass go, it will certainly move in one particular way, the same way every time we perform the same experiment. It is clear that the motion will depend on the initial 4.2. A DIGRESSION TO CONSIDER AU/prime/prime+BU/prime+CU=F. 163 positionu(t0) . But if two masses have the same initial position, the resulting motion will still be different if their initial velocities ˙ u(t0) are different. Thus we must also specify the initial velocity ˙ u(t0) as well as the initial position u(t0) . Are these sufficient to determine the motion? Yes, however that requires proof. What must be proved is that if we have two solutionsu1(t) andu2(t) of the same ordinary differential equation (1) and if their initial positions and velocities coincide, then the solutions coincide, u1=u2for all later time, t≥t0. Theorem 4.6 (Uniqueness). Let u1(t)andu2(t)be two solutions of the ordinary differ- ential equation Lu: =a¨u+b˙u+cu=f(t), wherea, b, andcare constants, a >0, b≥0, c≥0. Ifu1(t0) =u2(t0), and ˙u1(t0) = ˙u2(t0), thenu1(t) =u2(t)for allt≥0, in other words, the solution is uniquely determined by the initial position and velocity. Remark: The theorem is true under much more general conditions - as we shall prove in Chapter 6. Proof: Letw(t) =u2(t)−u1(t) . We shall show that w(t)≡0 for allt≥t0. Now Lw=L(u2−u1) =Lu2−Lu1=f−f= 0, that is, a¨w+b˙w+cw= 0 (4-14) Furthermore w(t0) = 0 and ˙w(t0) = 0, (4-15) sincew(t0) =u2(t0)−u1(t0) = 0 , and ˙ w(t0) = ˙u2(t0)−˙u1(t0) = 0 . This reduces the question to showing that if Lw= 0 , and if whas zero initial position and velocity, then in factw≡0 . The trick is to introduce a new function, E(t) , associated with (2) (which happens to be the total energy of the system) E(t) =1 2a˙w2+1 2cw2. How does this function change with time? We compute its derivative. ˙E(t) =a˙w¨w+cw˙w= ˙w(a¨w+cw). Using (2) we know that a¨w+cw=−b˙w. Therefore ˙E(t) =−b˙w2≤0 (since b≥0) [Thus energy is dissipated ( b>0 ) - or conserved ˙E= 0 in the special case of no damping (b= 0 ).] Consequently E(t)≤E(t0) for allt≥t0 (4-16) Now observe that for the mechanical system associated with w, we haveE(t0) =a w˙w2(t0)+ c 2w2(t0) = 0 . Furthermore, it is obvious from the definition of E(t) (sinceaandcare positive) that 0 ≤E(t) . Substitution of this information into (4) reveals 0≤E(t)≤0 for all t≥t0. 164 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 This proves E(t)≡0 for allt≥t0, which in turn implies w(t)≡0 —again from the definition of E(t) . Our proof is completed. We have taken some care since all of our uniqueness proofs will use essentially no additional ideas. A more general case ( a, bandc still constants but not necessarily positive) will be treated in Exercise 9. Having proved that there is at most one solution of the initial value problem Lu: =a¨u+b˙u+cu=f(t) (differential equation) u(t0) =αand ˙u(t0) =β (initial conditions) we must now prove there is at least one solution. This is the question of existence. For the special equation (5), the solution is shown to exist by explicitly exhibiting it. In the case of more complicated equations we are not as fortunate and must content ourselves with just showing that a unique solution exists but cannot exhibit it in closed form. It is easiest to begin with the homogeneous equation Lu= 0 , that is, find a solution of a¨u+b˙u+cu= 0 withu(t0) =α,and ˙u(t0) =β. Without motivation, let us see what the substitution u(t) =eλtyields. Here λis a constant. We must compute Leλt. Leλt= (aλ2+bλ+c)eλt. Canλbe chosen so that eλtis a solution of Lu= 0 ? Since eλt/negationslash= 0 for any t, this means, is it possible to pick λso thataλ2+bλ+c= 0 ? Yes. In fact that “quadratic equation formula” yields two roots λ1=−b+√ b2−4ac 2a, λ 2=−b−√ b2−4ac 2a of the characteristic polynomial p(λ) =aλ2+bλ+c. Notice that we have assumed a/negationslash= 0 . Thus, two solutions of the homogeneous equation are u1(t) =eλ1tandu2(t) =eλ2t. Since the operator Lis linear, every linear combination of solutions is also a solution, L(Au1+Bu2) =ALu 1+BLu 2= 0 . Therefore u(t) =Au1(t) +Bu2(t) is a solution of the homogeneous equation Lu= 0 for any choice of the scalars AandB. What about the initial conditions u(t0) =α,˙u(t0) =β; can they be satisfied by picking the constants AandBsuitably? Let us try. We want to pick AandBso that Aeλ1t0+Beλ2t0=α (u(t0) =α) Aλ1eλ1t0+Bλ2eλ2t0=β( ˙u(t0) =β). These equations can be solved as long as 0/negationslash=λ2e(λ1+λ2)t0−λ1e(λ2+λ1)t0= (λ2−λ1)e(λ1+λ2)t0, which means λ1/negationslash=λ2orb2−4ac/negationslash= 0 . [The linear equations Ar1+Bs1=α, Ar 2+Bs2=β can be solved for AandBifr1s2−r2s1/negationslash= 0 ]. Before dealing with the degenerate case b2−4ac= 0 , let us consider an 4.2. A DIGRESSION TO CONSIDER AU/prime/prime+BU/prime+CU=F. 165 Example: Solve ¨u+ 3 ˙u+ 2u= 0 with the initial conditions u(0) = 1 and ˙ u(0) = 0 . If we seek a solution of the form u(t) =eλt, the characteristic polynomial is λ2+ 3λ+ 2 = 0 , and has roots λ1=−1, λ2=−2 . Therefore u(t) =Ae−1+Be−2tis a solution. Since λ1/negationslash=λ2, we can solve for AandBby using the initial conditions. We find A+B= 1 (u(0) = 1), −A−2B= 0 ( ˙u(0) = 0). These two equations yield A= 1, B=−1 . Thus u(t) = 2e−t−e−2t is the unique solution of our initial value problem. The degenerate case b2−4ac= 0 must be discussed separately. In this case λ1=λ2= −b 2a, so the two solutions eλ1tandeλ2tare really the same solution. Without motivation (but see Exercise 12) we claim that teλ1tis also a solution. This is easy to verify by a calculation. L(teλ1t) =a(tλ2 1eλ1t+ 2λ1eλ1t) +b(eλ1t+λ1teλ1t) +cteλt = (aλ2 1+bλ1+c)teλt+ (2aλ1+b)eλ1t(4-17) Since (aλ2 1+bλ1+c) = 0 by definition of λ1, andλ1=−b 2ain our special case, both terms on the right vanish. Hence both u1(t) =eλ1tandu2(t) =teλ1tare solutions of Lu= 0 (ifb2−4ac= 0) , sou(t) =Aeλ1t+Bteλ1tis a solution for any choice of AandB. It is possible to pick AandBto satisfy arbitrary initial conditions u(t0) =α,˙u(t0) =β. Aeλ1t0+Bt0eλ1t0=α, (u(t0) =α) Aλ1eλ1t0+B(1 +λ1t0)eλ1t0=β ( ˙u(t0) =β). These can be solved for AandBsince 0/negationslash= (1 +λ1t0)e2λ1t0−λ1t0e2λ1t0=e2λ1t0. Example: Solve ¨u+6 ˙u+9u= 0 with the initial conditions u(1) = 2 , ˙u(1) = −1 . Seeking a solution in the form eλt, we are led to the characteristic equation λ2+ 6λ+ 9 = 0 , which hasλ1=−3, λ2=−3 , as roots. Therefore u1(t) =e−3tis a solution of Lu= 0 . Since λ1=λ2, another solution is u2(t) =te−3t. Thusu(t) =Ae−35+Bte−3tis a solution for anyAandB. To solve for AandBin terms of the initial conditions, we must solve the algebraic equations Ae−3+B·1·e−3= 2, (u(1) = 2), −3Ae−3+B(1−3)e−3=−1, ( ˙u(1) = −1). We find that A=−3e3andB= 5e3. Thus u(t) =−3e3e−3t+ 5e3te−3t, or, equivalently, u(t) =−3e−3(t−1)+ 5te−3(t−1). Our results will now be collected as 166 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 Theorem 4.7 . The initial value problem a¨u+b˙u+cu= 0, a/negationslash= 0,withu(t0) =α,˙u(t0) =β, wherea,b, andcare constants, has a unique solution. i) Ifb2−4ac/negationslash= 0, it is of the form u(t) =Aeλ1t+Beλ2t. ii) Ifb2−4ac= 0, soλ1=λ2, it is of the form u(t) =Aeλ1t+Bteλ1t. Hereλ1andλ2are the roots of the characteristic equation aλ2+bλ+c= 0 , and the constantsAandBare determined from the initial conditions. Remark: We have omitted the condition a>0 ,b≥0 ,c≥0 from our theorem since the construction presented to find a solution did not depend on this. Uniqueness for that case is treated as exercise 9, as we mentioned earlier. Another Example: Solve ¨u−2 ˙u+ 2u= 0 , with the initial conditions u(0) = 1,˙u(0) = 1 . The characteristic polynomial is λ2−2λ+2 = 0 . Its roots are λ1= 1+i, andλ2= 1−i. Since λ1/negationslash=λ2, the solution is of the form u(t) =Ae(1i)t+Be(1−i)t. From the initial conditions, we find that A+B= 1, (u(0) = 1), (1 +i)A+ (1−i)B= 1, ( ˙u(0) = 1). ThusA=1 2, B=1 2, so u(t) =1 2e(1+i)t+1 2e(1−i)t. Recalling that ex+iy=ex(cost+isiny) , this solution may be written in a more familiar form: u(t) =1 2et(cost+isint) +1 2et(cost−isint), that is, u(t) =etcost. What has been done can be summarized elegantly in the language of linear spaces. We have sought a solution of a second order linear O.D.E., which we write as Lu= 0 . It was found that every solution of this equation could be expressed as a linear combination of two specific solutions u1andu2, u(t) =Au1(t) +Bu2(t) , where the constants AandB are uniquely determined from u(t0) and ˙u(t0) . Thus, the set of functions u which satisfy Lu= 0 form a two dimensional subspace of D(L) =C2. The functions u1andu2span that subspace. If we call the set of all solutions of Lu= 0 the nullspace of L,N(L) , then our result simply reads “dim N(L) = 2 ”. A particular solution of Lu= 0 is found by specifyingu(t0) and ˙u(t0) . The inhomogeneous equation Lu=fis treated by finding a coset of the nullspace ofL. For ifu0is a particular solution of the inhomogeneous equation Lu0=f, then 4.2. A DIGRESSION TO CONSIDER AU/prime/prime+BU/prime+CU=F. 167 u= ˜u+u0; where ˜u∈N(L) , is also a solution since Lu=L(˜u+u0) =L˜u+Lu0= 0+f=f. Therefore, if one solution u0of the inhomogeneous equation Lu=fis found, the general solution is u= ˜u+u0where ˜u∈N(L) . In particular, the solution ˜ u∈N(L) can be chosen so that arbitrary initial conditions for u,u(t0) =α,˙u(t0) =β, can be met. We shall defer (until our systematic treatment of linear O.D.E.’s) presenting a general method for finding a solution u0of the inhomogeneous equation. In our example, the particular solution will be found by guessing. Example: SolveLu: = ¨u−u= 2t, with the initial conditions u(0) = −1,˙u(0) = 3 . The homogeneous equation Lu= 0 has the general solution ˜ u(t) =Aet+Be−t. We observe that the function u0(t) =−2tis a particular solution of the inhomogeneous equation, Lu= 2t. Thusu(t) =Aet+Be−t. The initial conditions lead us to solve the following equations for AandB, A+B=−1 (4-18) A−B−2 = 3. (4-19) A computation gives A= 2 ,B=−3 . Thus the solution of our problem is u(t) = 2et−3e−t−2t It is routine to verify that this function u(t) does satisfy the O.D.E. and initial conditions (you should verify the solutions to check for algebraic mistakes). Exercises (1) Solve the following homogeneous initial value problems, (a). ¨u−u= 0, u (0) = 0,˙u(0) = 1 . (b). ¨u+u= 0, u (0) = 1,˙u(0) = 0. (c). ¨u−4 ˙u+ 5u= 0, u(0) = −1,˙u(0) = 2. (d). ¨u+ 2 ˙u−8u= 0, u(2) = 3,˙u(2) = 0. (e). ¨u= 0, u (0) = 7,˙u(0) = 3. (2) Solve the following inhomogeneous initial value problems by guessing a particular solution of the inhomogeneous equation. Check your answers. (a) ¨u−u=t2, u (0) = 0,˙u(0) = 0 hint: Tryu0(t) =a1t2+a2t+a3and solve for a1,a2,a3.] (b) ¨u−4 ˙u+5u= sint, u (0) = 1,˙u(0) = 0 [ hint: Tryu0(t) =a1sint+a2cost.] (3) Consider an undamped harmonic oscillator with a sinusoidal forcing term, ¨ u+n2u= sinγt. Find the general solution if γ2/negationslash=n2[tryu0(t) =a1sinγt+a2cosγtfor a particular solution]. What happens if γ→+−n? This is called resonance . (4) You shall discuss damping in this problem. Consider the equation ¨ u+ 2µ˙u+ku= 0 , whereµ>0 , andk>0 . We shall let γ=/radicalbig |µ2−k|. 168 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 (a)Light damping (µ2<k) . Show that the solution is u(t) =e−µt(Acosγt+Bsinγt), and sketch a rough graph for the case A= 1, B= 0 . This is the kind of oscillation you want for a pendulum clock, with µsmall. (b)Heavy damping (µ2>k) . Show that the solution is u(t) =e−µt(Aeγt+Be−γt). Show that u(t) vanishes at most once. Sketch a graph for the two cases A= B= 1 andA=−1,B= 3 . The first describes the oscillation of an ideal screen door, while the second describes the ideal oscillation of a slammed car door. (5) It is often useful to study the oscillations described by ¨ u+ 2µ˙u+ku= 0 by sketching the solution in the u,˙uplane - or phase space as it is called. Investigate the curves for heavily and lightly damped oscillators. Show that the curve for a heavily damped oscillator will be a straight line through the origin for special initial conditions. What does the phase space curve look like for an undamped oscillator ( µ= 0, k> 0) ? (6) Consider the linear operator Lu=a¨u+b˙u+cu, wherea,b,c are constants. We have seen thatLert=p(r)ertwherep(r) is the characteristic polynomial. (a) Ifrisnotone of the roots of the characteristic polynomial, observe that you can find a particular solution of Lu=ert. What is it? (b) If neither r1norr2is a root of the characteristic polynomial, find a particular solution of Lu=a1er1t+a2er2t, wherea1anda2are specified constants. (c) Use this procedure to find a particular solution of i)¨u−4u= cosht, ii )¨u+ 4u= sint (7) (a) Imitate our procedure and develop a theory for the first order homogeneous O.D.E.Lu: = ˙u+bu= 0 , where bis a constant. In particular, you should prove that there exists aunique solution satisfying the initial condition u(t0) =α, and give a recipe for finding it. Use your recipe to solve ˙ u+ 2u= 0, u(0) = 3 . (b) And now you will show us how to find a particular solution of the inhomogeneous equationLu=f, wheref(t) is some given continuous function and Lu: = ˙u+bu. [hint: Try to find a function µ(t) such that µ( ˙u+bu) =d dt(µu) . Then integrated dt(µu) =µf, and solve for u]. Use your method to find a particular solution for ˙ u+2u=x, and then a solution of the same equation which satisfies the initial condition u(0) = 1 . (8) Find a solution of u/prime/prime/prime−2u/prime/prime−u/prime+ 2u= 0 which satisfies the initial conditions u(0) =u/prime(0) = 0, u/prime/prime(0) = 1 . [ hint: The cubic equation γ3−2γ2−γ+ 2 has roots +1,−1 and 2]. (9) You will prove the uniqueness theorem for the equation ¨ u+b˙u+cu= 0 , where band care any constants (we have let a= 1 , because if it is not 1, just divide the whole equation by a). The trick is to reduce this to the special case b≥0, c≥0 , already done. 4.2. A DIGRESSION TO CONSIDER AU/prime/prime+BU/prime+CU=F. 169 (a) Show that in order to prove the solution of ¨u+b˙u+cu=f,whereu(t0) =α,˙u(t0) =β is unique, it is sufficient to prove that the only solution of ¨w+b˙w+cw= 0, w(t0) = 0,˙w(t0) = 0 isw(t)≡0 . (b) Define ϕ(t) byw(t) =eγtϕ(t) . Observe: to prove w= 0 , it is sufficient to proveϕ≡0 ( hereγis any constant). Use the differential equation and initial conditions for wto find the differential equation and initial conditions for ϕ. Show that γcan be picked so that the D.E. for ϕis ¨ϕ+˜b˙ϕ+˜0ϕ= 0, where ˜band ˜care positive. Deduce that ϕ≡0 , and from that, that w≡0 , completing the proof. (10) A boundary value problem for the equation u/prime/prime+bu/prime+cu= 0 is to find a solution of the equation with given boundary values, say u(0) =αand u(1) =β. Assumebandcare real numbers. (a) Show that a solution of the boundary value problem always exists if b2−4c≥0 (the caseb2−4c= 0 will have to be done separately). (b) Prove that if b2−4c≥0 , the solution is unique too. [I suggest letting u(t) = eγtv(t) , and then choosing γso that the equation satisfied by vis of the form v/prime/prime+ ˜cv= 0 , where ˜ c≤0 . The case ˜ c= 0 is trivial. If˜(c)<0 , can the solution have a positive maximum or negative minimum?] (11) If a spring is hung vertically and a mass mplaced at its end, an external force of magnitude mgdue to gravity is placed on the system. Assume there are no dissipative forces of any kind. (a) Set up the differential equation of motion. Remember that you must specify which is the positive direction. (b) If the tip of the spring is displaced a distance dby placing the mass on it (no motion yet), so the equilibrium position isdbelow the unstretched end of the spring, show that the spring constant kis given by k=mg/d . (c) Let the body weigh 32 pounds, and dbe 2 feet. Find the subsequent motion if the body is initially displaced from rest one foot below its equilibrium position. [Take |g|= 32 ft/sec2]. (12) * Consider au/prime/prime+bu/prime+cu= 0 . Ifγ1/negationslash=γ2, are the roots of the characteristic equation, observe that the function ˜u(t) =eγ1t−eγ2t γ1−γ2 170 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 is also a solution (it is a linear combination of eγ1tandeγ2t). Now pass to the limitγ2→γ1(leaveγ1fixed and let γ2move) by using the Taylor series for eγt. The function you get is then a “guess” for a second solution in the degenerate case γ1=γ2. This supplies some motivation for the guess made earlier. (13) * Consider Lu: =u/prime/prime+ 2u=f, wherefis given. You know how to solve Lu=Asinnx(Exercise 6). Find a particular solution to the general inhomoge- neous equation in the interval [ −π,π] by expanding fin a Fourier series and then use superposition. Apply this to solve u/prime/prime+ 2u=x. (14) Consider an undamped harmonic oscillator, whose motion is specified by u(t) , where mu/prime/prime+ku= 0,k> 0 . Show that the solution u(t) =A1cos/radicalbigg k mt+B1sin/radicalbigg k mtmay be written in the form u(t) =Asin(wt+θ), whereAis the amplitude of the oscillation, w= 2πv, v is the frequency , andθis thephase . Show that u(t) is periodic, u(t+T) =u(t) , where the periodT= 1/v. Interpret the amplitude and phase and determine A,w, andθin terms of A1, B1, k andm. [I suggest looking at a specific example and its graph first]. 4.3 Generalities on LX =Y. Undoubtedly the fundamental problem in the theory of linear (and nonlinear) operators is to determine the nature of the range of an operator L. One particular aspect of this is the vast problem of solving the equation LX=Y forXwhenYis given to you. The question here is, “is a given Yin the range of L?”, or “can we find some Xsuch thatLX=Y?” If one can solve the problem uniquely for anyY, then the solution is written as X=L−1(Y), whereL−1is the operator inverse to L, in the sense that L−1L=I(so to solve LX=Y, applyL−1, X=L−1LX=L−1Y). Let us give some examples, familiar and unfamiliar, of problems of the form LX=Y, whereYis given. 1.LX= (2x1+ 3x2,x1+ 2x2), X ∈R2, L:R2→R2. The problem of solving LX=YwhereY= (−1,2)∈R2is that of solving the two equations 2x1+ 3x2=−1 x1+ 2x2= 2 for two unknowns ( x1,x2) =X. 4.3. GENERALITIES ON LX=Y. 171 2.Lu=u/prime/prime+ 2u/prime+ 3u,whereu∈C2, L:C2→C. The problem of solving L(u) =xis that of solving the inhomogeneous ordinary differential equation Lu: =u/prime/prime+ 2u/prime+ 3u=x foru(x) . 3.Lu=/integraldisplayπ 0cos(x−t)u(t)dt, u ∈C[0,π]. You should check that Lis a linear operator. The problem of solving L(u) = sinxis that of solving the integral equation Lu: =/integraldisplayπ 0cos(x−t)u(t)dt= sinx for the function u. In this example, it is instructive to examine the range more closely. Since cos(x−t) = cosxcost+ sinxsintand since functions of xare constant with respect totintegration, we see that Lumay be written as Lu: = cosx/integraldisplayπ 0costu(t)dt+ sinx/integraldisplayπ 0sintu(t)dt, or Lu: =α1cosx+α2sinx, where the numbersα1andα2are α1=/integraldisplayπ 0(cost)u(t)dt;α2=/integraldisplayπ 0(sint)u(t)dt. Thus, the range of Lis the linear space spanned by cos xand sinx, which has dimension two. This linear operator Ltherefore maps the infinite dimensional space C[0,π] into a finite (two) dimensional space. In order to even have a chance of solving Lu=ffor this operatorL, we first check to see if feven lies in this two dimensional subspace (for if it doesn’t, it is futile to go further). The particular function sin xdoes, so it is reasonable to look for a solution - which we shall not do right now (however there are infinitely many solutions, among them u(x) =2 πsinx). One particularly important equation which arises frequently is the homogeneous equa- tion LX= 0, which is the special case Y= 0 of the inhomogeneous equation , LX=Y. SinceLis a linear operator, there is no problem of our finding onesolution of LX= 0 forX= 0 is a solution, the so-called trivial solution of the homogeneous equation. The problem is to find a non-trivial solution, or better yet, all solutions. In the previous section, this question was answered fully for the particular operator Lu=au/prime/prime+bu/prime+cu, where a, b, andcare constants. Many of our results there generalize immediately, as we shall see now. 172 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 Definition: The set of all solutions of the homogeneous equation LX= 0 where Lis a linear operator is called the nullspace ofL. This nullspace of L,N(L) , consists of all X in the domain of Lwhich are mapped into zero by L, L:N(L)→0.N(L)⊂D(L). We have called N(L) the null space ofL, not the null setbecause of Theorem 4.8 . The nullspace of a linear operator L:V1→V2is a linear space, a subspace of the domain of L. Proof: Since the domain of L,D(L) =V1, is a linear space and N(L)⊂D(L) , by Theorem 2, p.142 all we need show is that the set N(L) is closed under multiplication by scalars and under addition of vectors. Say X1andX2∈N(L) . ThenLX 1= 0 and LX 2= 0 . We must show that L(aX1) = 0 for any scalar a, and that L(X1+X2) = 0 . ButL(aX1) =aL(X1) =a·0 = 0 , and L(X1+X2) =LX 1+LX 2= 0 + 0 = 0 . Thus N(L) is a subspace of D(L) =V1. One important reason for examining the null space of a linear operator is because ifN(L) is known, and if any onesolution of the inhomogeneous equation is known, say LX 1=Y(whereYwas given and X1is the solution we know), then every solution of the inhomogeneous equation is of the form ˜X+X1, where ˜X∈N(L) . In other words every solution of LX=Yis inN(L) +X1, theX1coset of the subspace N(L) . Theorem 4.9 . LetL:V1→V2be a linear operator. If X1andX2are any two solutions of the inhomogeneous equation LX=Y, whereYis given, then X2−X1∈N(L), or X2=˜X+X1where ˜X∈N(L). Proof: Let ˜X=X2−X1. We shall show that ˜X∈N(L) . L˜X=L(X2−X1) =LX 2−LX 1=Y−Y= 0. By using this theorem, we see that if allsolutions of the homogeneous equation LX= 0 are known - the nullspace of L—and if onesolution of the inhomogeneous equation LX 1=Yis known, then allof the solutions of the inhomogeneous equation are known. This solution set of the inhomogeneous equation is the X1coset of N(L) . Example: 1 LetL:R2→R2be defined by LX= (x1+x2,x1−x2)∈R2 Then N(L) is the set of all points in R2such thatLX= 0 , that is, which satisfy the equations x1+x2= 0 x1−x2= 0 Thus the nullspace of Lconsists of the intersection of the two lines x1+x2= 0, x1−x2= 0 . The only point on both lines is 0. Thus N(L) is just the point 0. To solve the inhomogeneous equationLX=Y, whereY= (1,1) . x1+x2= 1, x1−x2= 1, 4.3. GENERALITIES ON LX=Y. 173 we find one solution of it, X1= (1,0) . Then every solution of the inhomogeneous equation is of the form X=˜X+X1, where ˜X∈N(L) . But since ˜X+0 is the only point in N(L) , every solution is of the form X= 0 +X1=X1. Thus every solution is exactly X1, which is the unique solution of LX=Y. This situation is a general one. Again, we also saw this forLu=au/prime/prime+bu/prime+cu. Theorem 4.10 . If the nullspace N(L)of the linear operator consists only of 0, then the solution of the inhomogeneous equation LX=Y(if a solution exists) is unique. (Thus, if the nullspace contains only 0, then Lis injective). Proof: Say there were two solution X1andX2. ThenLX 1=YandLX 2=Y, which impliesL(X2−X1) =LX 2−LX 1=Y−Y= 0 . Therefore ( X2−X1)∈N(L) . Since the only element of N(L) is 0, X 2−X1= 0 , or,X1=X2. In other words, the two solutions are the same. Example: 2 LetL:C2→Cbe defined on functions u∈C2by Lu: =a(x)u/prime/prime+b(x)u/prime+c(x)u. The nullspace of Lconsists of all solutions of the homogeneous equation Lu= 0 . It turns out (see chapter 6) - as in the constant coefficient case - that every solution of this homogeneous O.D.E. has the form u=Au1+Bu2, whereu1andu2are any two linearly independent solutions of the equation, and where AandBare constants. Thus N(L) is a two dimensional space spanned by u1andu2. Ifu1is a particular solution of the inhomogeneous equation Lu1=f, then allthe solutions of Lu=fare just the elements of theu1coset of N(L) , that is, functions of the form u= ˜u+u1, where ˜u∈N(L) . With every linear operator L:V1→V2, V1=D(L) , we have associated two other linear spaces, the nullspace N(L)⊂D(L) =V1and range R(L)⊂V2. There is a valuable and elegant way to connect D(L),N(L) and R(L) . The result we are aiming at is certainly the most important theorem of this section. We know that R(L)⊂V2. The space V2may be of arbitrarily high dimension. However, since R(L) is the image of D(L) , we suspect that R(L) can take up “no more room” then D(L) . To be more precise, dimR(L)≤dimD(L). Thus, for example, if L:R2→R17, we expect that the range of Lis a subspace of dimension no more than two in R17. Not only is this a justifiable expectation, but even more is true. If Dim R(L) = Dim D(L) , essentially all of D(L) is carried over under the mapping. But if Dim R(L)<DimD(L) , what has happened to the remainder of D(L) ? Let us look at N(L)⊂D(L) . The elements of N(L) are all squashed into the zero element of V2. In other words, a set of dim N(L) inV1=D(L) is mapped into a set of dimension zero inV2. DoesLdecompose D(L) +V1into two parts, N(L) and a complement N(L)/prime such thatLmaps N(L) into zero and the dimension of the remainder, N(L)/prime, is preserved underL(so dim N(L)/prime= dim R(L)) . Of course, a figure goes here 174 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 Theorem 4.11 . Let the linear operator LmapV1=D(L)intoV2. IfD(L)has finite dimension, then dimD(L) = dim R(L) + dim N(L). Proof: LetN(L)/primebe a complement of N(L) (cf. pp. 163a-d). Since dim N(L) + dimN(L)/prime= dim D(L) , it is sufficient to prove that dim N(L)/prime= dim R(L) . ForX∈V1, we can write X=X1+X2, whereX1∈N(L) andX2∈N(L)/prime. Now LX=LX 1+LX 2, so the image of D(L) is the same as the image of N(L)/prime. In addition, if X2∈N(L)/prime, thenLX 2= 0 if and only if X2= 0 , merely because N(L)/primeis a complement of the nullspace. Let {θ1,...,θ k}be a basis for N(L/prime) . IfX2∈N(L)/prime, we can write X2= k/summationdisplay j=1ajθj, andLX 2=k/summationdisplay j=1ajLθj. LetLθ1=Y1, Lθ 2=Y2,...,Lθ k=Yk. Since the image ofN(L)/primeisR(L) , the vectors Y1,...,Y kspanR(L) . Thus, dim R(L)≤k= dim N(L)/prime. To show that there is equality, dim R(L) = dim N(L)/prime, we prove that Y1,...,Y kare linearly independent. If c1Y1+···+ckYk= 0 , then 0 = c1Lθ1+···+ckLθk=L(c1θ1+ ···+ckθk) =L˜Xwhere ˜x=c1θ1+···+ckθk∈N(L)/prime. However for any ˜X∈N(L)/prime, we knowL˜x= 0 implies that ˜X= 0 . The linear independence of θ1,...,θ kfurther shows that c1=c2=···=ck= 0 . The hypothesis c1Y1+···+ckYk= 0 has led us to conclude that the cj’s are all zero, that is, the Yj’s are linearly independent. Therefore dimR(L) = dim N(L)/prime. Coupled with our first relationship, this proves the result. Corollary 4.12 :dimR(L)≤dimD(L). Proof: dimN(L)≥0 . Two examples, one an illustration, the other an application. Example: 1 Consider a projection operator, PA, mapping vectors from Eninto a subspace AofEn, where the dim A=m < n . Let us first show that PAis a linear operator. If e1,...e mis an orthonormal basis for A, then for any XandYinEn, PA(X+Y) =m/summationdisplay k=1/angbracketleftX+Y, e k/angbracketrightek=m/summationdisplay k=1(/angbracketleftX, e k/angbracketright+/angbracketleftY, e k/angbracketright)ek =m/summationdisplay k=1/angbracketleftX, e k/angbracketrightek+m/summationdisplay k=1/angbracketleftY, e k/angbracketrightek=PAX+PAY. Similarly,PA(aX) =aPAXfor every scalar a. Thus the projection operator is a linear operator. Since R(PA) =Aand dimA=m, while dim En=n, we conclude that dimN(PA) =n−m. This could have been arrived at immediately since PAwill certainly map everything perpendicular to A, that isA⊥, into 0 (see fig. illustrating the case E2→A, whereAis a line). Thus N(PA) +A⊥, so dim N(PA) = dimA⊥=n−m. Example: 2 DefineL:Rn→Rkby, LX= (a11x1+a12x2+···+a1nxn, a21x1+···+a2nxn,···,aklx1+ak2x2+···+aknxn) whereX= (x1,x2,...x n)∈Rn. If we let Y= (y1,...,y k)∈Rk, then writing Y=LX, the linear operator Lmay be defined by the kequations (for y1,...,y k) inn“unknowns” 4.3. GENERALITIES ON LX=Y. 175 (x1,...,x n) , a11x1+a12x2+···+a1nxn=y1 a21x1+a22x2+···+a2nxn=y2 ... ak1x1+ak2x2+···+aknxn=yk. The problem of solving LX=Y, whereYis given, is that of solving kequations with n “unknowns”. Consider the special case k <n , when there are less equations than unknowns. Since the range of Lis contained in Rk,R(L)⊂Rk, then dim R(L)≤dimRk=k. Because D(L) =Rn, we also know that dim D(L) = dim Rn=n. Thus dimN(L) = dim D(L)−dimR(L)≥n−k>0. However ifdimN(L)>0 , then N(L) must contain something other than zero. Thus there is at least one non-trivial solution ˜Xof the homogeneous equation ,L˜X= 0 . Since a˜Xis also a solution, where ais any scalar, there are, in fact an infinite number of solutions . Notice that the above was a non-constructive existence theorem. We proved that a solution does exist but never gave a recipe to obtain it. One consequence of this result is that, if dim N(L)>0 , and if a solution of the inhomogeneous equation LX=Yexists, it is not unique ; for ifLX 1=Y, then also L(X1+˜X) =Y, where ˜Xis any solution of the homogeneous equation. In the special case n=k, and dimN(L) = 0 a fascinating (and non-constructive) theorem falls out of Theorem 11: the inhomogeneous equation LX=Yalways has a solution and the solution is unique . Put in more conventional terms, if there are the same number of equations as unknowns, and if the only solution of the homogeneous equation is zero, then the inhomogeneous equation always has a unique solution. Thus, if n=k, uniqueness implies existence . Since dim N(L) = 0 , then dim R(L) = dim D(L) =n. However L:Rn→Rnin this case (n=k) . Since R(L)⊂Rnand dim R(L) =n, we see that R(L) must be all of Rn, that is, every Y∈Rnis in the range of L, which means that the inhomogeneous equation LX=Yis solvable for every Y∈Rn. Theorem 10 gives the uniqueness. We shall obtain a better theorem later. Remark: Some people refer to dim R(L) as the rank of the linear operator L. We shall, however, refer to it as the dimension of the range of L. IfL1:V1→V2andL2:V2→V3, it is easy to make a few statements about dimR(L2L1) . Theorem 4.13 . IfL1:V1→V2andL2:V3→V4, whenV2⊂V3, (soL2L1) is defined), then dimR(L2L1)≤min(dim R(L1),dimR(L2)). Proof: The last corollary states that an operator is like a funnel with respect to dimension: the dimension can only get smaller or remain the same. After passing through two funnels, we obtain no more than the smallest allowed through. One might think that there should be equality in the formula. That this is not the case can be seen from the possibility illustrated in the figure. Only the shaded stuff gets through. 176 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 Exercises (1) LetL:Rn→Rnbe defined by LX= (x1,x2,...,x k,0,..., 0), whereX= (x1,x2,...x n)∈Rn. Describe R(L) and N(L) . Compute dim R(L) and dim N(L) . (2) (a) Describe the range and nullspace of the linear operator L:R3→R3defined by LX= (x1+x2−x3,2x1−x2+x3, x2−x3), X= (x1,x2,x3)∈R3. (b) Compute dim R(L) and dim N(L) . (c) Is (1,2,0)∈R(L) ? Is (1,2,1)∈N(L) ? Is (1,2,2)∈N(L) ? Is (0,−1,−1)∈N(L) ? (3) LetA={u∈C2[0,2]:u(0) =u(1) = 0 }, and define L:A→C[0,1] byLu= u/prime/prime+b(x)u/prime−u, whereb(x) is some continuous function. Prove N(L) = 0 . [ hint: If u∈N(L) , canuhave a positive maximum or negative minimum?] (4) Consider the linear operator L:C[0,1]→C[0,1] defined by (Lu)(x) =u(x) + 2/integraldisplay1 0ex−tu(t)dt (a) Find the nullspace of L. (b) SolveLu= 3ex. Is the solution unique? (c) Show that the unique solution of Lu=f, wheref∈C[0,1] is u(x) =f(x)−cex,wherec=2 3/integraldisplay1 0e−tf(t)dt. (5) LetL:V→V(soLkis defined for k= 0,1,2...). Prove that (a)R(L)⊂N(L) if and only if L2= 0 . (b)N(L)⊂N(L2)⊂N(L3)⊂... (c)N(L)/prime⊃N(L2)/prime⊃N(L3)/prime⊃.... (6) IfL1:V1→V2andL2:V3→V4whereV2⊂V3, Theorem 12 gives an upper bound for dim R(L2L1) . (a) Prove the corresponding lower bound dimR(L2L1)≥dimR(L1) + dim R(L2)−dimV3. [hint: Prove the equivalent inequality dim R(L1)≤dimR(L2L1) + dim N(L2) by letting ˜V=R(L1) and applying Theorem 11 to L2defined on ˜V]. (b) Prove: if dim N(L2) = 0 , then dimR(L2L1) = dim R(L1). 4.4.L:R1→RN. PARAMETRIZED STRAIGHT LINES. 177 (c) If dim N(L1) = 0 , is it then true that dim R(L2L1) = dim R(L2) ? Proof or counterexample. (d) If dimV1= dimV2= dimV3and dim N(L1) = 0 , is it true that dim R(L2L1) = dimR(L2) ? Proof or counterexample. (7) IfL1andL2both mapV1→V2, prove |dimR(L1)−dimR(L2)| ≤dimR(L1+L2). (8) Consider the operator L:C2[0,∞)→C[0,∞) defined by Lu: =u/prime/prime+ 3u/prime+ 2u. (a) Describe N(L) . What is dim N(L) ? Isf(x) = sinx∈R(L) ? (b) Consider the same operator Lbut mapping AintoC[0,∞] , whereA={u∈ C2[0,∞):u(0) +u/prime(0) = 0 }. Answer the same questions as part a). (c) Same as bbutA={u∈C2[0,∞):u(1) +u/prime(1) = 0 }this time. 4.4 L:R1→Rn. Parametrized Straight Lines. Our study of particular linear operators begins with the most simple case: a linear operator which maps a one- dimensional space R1into anndimensional space Rn. Since the dimension of the range of Lis no greater than that of the domain R1and dim R1= 1 , then dimR(L)≤1. This proves Remark: IfL:R1→Rn, then the dimension of the range of Lis either one or zero. The case dim R(L) = 0 is trivial, for then Lmust map all of R1into a single point, and that single point must be the origin since the range of Lis a subspace. Thus, ifdimR(L) = 0, thenLmaps every point into 0 . Without change, the same holds if L:V1→V2(where V1andV2are any linear spaces) and dim R(L) = 0 . Not very profound. If dim R(L) = 1 , then the subspace R(L) inRnis a one dimensional subspace in the ndimensional space Rn, this is, R(L) is a “straight line” through the origin of Rn. This straight line is determined if any one point P/negationslash= 0 on it is known. Then there is a point X1∈R1such thatLX 1=P. Since R1is one dimensional it is spanned by any element other than zero, so every X∈R1can be written as X=sX1. Therefore, if Xis any element of R1, LX=L(sX1) =sLX 1=tP. In other words, this last equation states that the range of Lis a multiple of a particular vectorP, that is, a straight line through the origin. Example: IfL:R1→R2such that the point X1= 2∈R1is mapped into the point P= (1,−2)∈R2, then L:X=s2→(s,−2s), In particular, the point X= 3(s=3 2) is mapped into the point (3 2,−3) . 178 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 a figure goes here In applications, the domain R1usually represents time, while the range represents the position of a particle. Then L:R1→Rnis an operator which specifies the position of a particle at a given time. Since Lis linear and L0 = 0 , the path of the particle must be a straight line which passes through the origin at t= 0 . Later on in this section we shall show how to treat the situation of a straight line not through the origin, while in Chapter 7 we shall examine curved paths (non-linear operators). Example: This is the same example as before. L:R1→R2is such that at time t= 2∈R1 a particle is at the point (1 ,−2) (while at t= 0 it is at the origin). At any time t=s2 , the particle is at ( s,−2s) . In particular, at t= 3(s=3 2) , the particle is at (3 2,−3) . It is also convenient to rewrite the position ( s,−2s) directly in terms of the time. Since t= 2s, the position at time tis (t 2,−t) . Thus we can write L:t→(t 2,−t), which clearly indicates the position at a given time. If a point in the space R2is specified byY= (y1,y2)∈R2, then the operator can be written as y1=1 2t y2=−t. All of these are useful ways to write the operator L. In some situations, one might be more useful than another. This brings us to an issue which perhaps seems a bit pedantic but can serve you well in times of need. How can we represent the operator in a picture? There are three distinct ways. Some clarity can be gained by distinguishing them carefully. The same ideas carry over immediately to nonlinear operators. Our first picture has two parts. If L:R1→Rn, then the first part is a diagram of R1, the second part a diagram of Rn, and between them are arrows to indicate the image of each point in R1. The picture below the first example was of this type. All of the arrows get in the way, so a more convenient picture is needed. That comes next. The second picture is the graph of an operator L. The graphL:R1→Rnis the set of points (X,LX ) in the Cartesian product space R1×Rn. Thus, ifV1is time, and Rn space with Lassigning a position to every time, then the points on the graph are points in time - space ( X,LX ) . For the previous example, these are the points ( t,(t 2,−t)) in R1×R2, a straight line in time-space ( or space-time if you prefer). To each time, there is a unique point in space. In a sense, this second picture, the graph, associated with an operator results from gluing together the two pieces of the first picture. By using the graph of an operator, we avoid the arrow mess of the first picture. The third picture just indicates the range of an operator (when thinking of pictures, the range is often referred to as the path of the operator). In terms of the time- position example, this picture only shows the path of a particle in space and ignores when a particle had a given position. Thus, this picture is the second half of the first picture. From our physical interpretation, it is clear that two different operators might have the same path (for two particles could travel the same path without having the same position at every time). Thus, this picture is an incomplete representation of an operator. 4.4.L:R1→RN. PARAMETRIZED STRAIGHT LINES. 179 Example: If˜L:R1→R2such that the point X1= 1∈R1is mapped into the point P= (1,−2)∈R2, then ˜L:X=s·1→(s,−2s). In particular, the point X= 3(s= 3) is mapped into the point (3 ,−6) . The graph of˜L is the set of points ( s,s,−2s) , which is a straight line in R1×R2. Compare this with the operatorLconsidered previously (we remind you that L:X= 2s→(s,−2s)) . The graph ofLwas the set of point (2 s,s,−2s) . These two sets of points the graphs of ˜LandL, respectively, do not coincide since the operators are the same. On the other hand, the path of˜Lis the set of points ( s,−2s) , which is exactly the same set of points as the path of L. Shortly, we shall ask the question, how can we describe a straight line in Rn? One way is to find an operator whose path is that straight line. Since many operators have the same path, there will be many possible ways to describe the straight line. All we need do is pick one, any one will do. LetL:R1→RnandY0be some fixed point in Rn. Consider the operator MX := LX+Y0. SinceM0 =L0 +Y0=Y0/negationslash= 0 , we see that Mis not a linear operator; it is called an affine operator oraffine mapping . The range of Mis the subspace translated by the vector Y0, a straight line which does not necessarily pass through the origin ( it will if and only if Y0∈R(L) ). In other words, R(M) is theY0coset of the subspace R(L) . Example: TakeLto be the same as before, so L:X= 2s→(s,−2s) orL(2s) = (s,−2s) . LetY0= (−3,2) . ThenMX :=LX+Y0= (s,−2s) + (−3,2) = (s−3,−2s+ 2) , where X= 2s. In particular, Mmaps the point X= 3∈V1(s=e 2) into ( −3 2,−1)∈R2. The figure shows the path of LandM. SinceX= 2s, we can eliminate sfrom the above formula and write MX = (1 2X−3,−X+ 2), X ∈R1. If we denote by Y= (y1,y2) a general point in R2, thenMmay be written in the standard form y−1 =1 2X−3 y−2 =−X+ 2. Of course, one could eliminate Xfrom these too and be left with 2 y1+y2=−4 , which is the equation of the path and could come from any mapping with the same path. It is instructive to investigate the reverse question, given two points PandQinRn, find an equation for the straight line passing through them. Any mapping whose path is the desired line will do. We have learned that MX =LX+Y0is the general equation of a straight line through Y0. There is complete freedom in specifying which points are mapped into PandQ, so we would be foolish not to pick the most simple case. Let M: 0→PandM: 1→Q. ThenP=M(0) =L(0) +Y0=Y0, soY0=P, and Q=M(1) =L(1) +Y0=L(1) +P, soL: 1→P−Q. This completely determines M (sinceLis determined once the image of one point is known, L: 1→P−Q, and the vectorY0is also determined, Y0=P). Example: Find an equation for the straight line passing through the two points P= (1,2,−3,−4) ,Q= (−1,3,2,−2) in R4. SayPis the image of 0 and Qis the image of 1, soM: 0→PandM: 1→Q. Then since MX =LX+Y0⇒P=M(0) =Y0so Y0= (1,2,−3,−4) . AlsoQ=L(1)+Y0⇒L(1) =Q−Y0=Q−P= (−2,1,5,2) . Because 180 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 everyX∈R1can be written as X=s·1⇒LX=L(s·1) =sL(1) =s(−2,1,5,2) , or LX= (−2s,s,5s,2s) , whereX=s·1∈R1. ThusMX =LX+Y0= (−2s,s,5s,2s) + (1,2,−3,−4) , or MX = (−2s+ 1,s+ 2,5s−3,2s−4),whereX=s· ∈R1. If we useY= (y1,y2,y3,y4) to indicate a general point in R4, thenM:R1→R4can be written as four equations. y1=−2s+ 1 y2=s+ 2 y3= 5s−3 y4= 2s−4 whereX=s·1∈R1. For example, the image of X= 2(s= 2) in R1is the point Y= (−3,4,7,4)∈R4. The discussion before the example contained most of the proof of Theorem 13 . LetP andQbe two points in Rn. Then the affine mapping MX =P+s(Q−P), has as its path the straight line passing through PandQ. Remark: 1 The affine mapping ˜MX =P+ks(Q−P) , wherek/negationslash= 0 is some constant, has the same path too. The only change is that while M: 0→PandM: 1→Qthis mapping ˜M: 0→Pand ˜M:ks→Q. In other words for ˜Mwe have chosen to take ks (nots) as the pre-image of Q. This pre-image of Qwas entirely arbitrary anyway. Remark: 2 The equation MX =P+s(Q−P) ofM:R1→Rn, whereX=s·1∈R1 is called a parametric equation of the straight line which passes through PandQin Rn, andsis called the parameter . Other parametric equations of the same line arise if X=ks·1∈R1(cf. Remark 1), where kis some non-zero constant. In order to introduce the slope of a straight line, let us paraphrase the last few para- graphs in terms of particle motion. If PandQare two points in Rn, thenMt= P+t(Q−P)M:R1→Rn, wheret∈R1describes the position of the particle at time t. Att= 0 the particle is at P, while att= 1 the particle is at Q. Another particle movingktimes as fast has the position ˜Mt=P+kt(Q−P) . This other particle is also atPwhent= 0 , but takes time t=1 kto reach the point Q. It still has the same path as the first particle. If we denote by Y= (y1,y2,...,y n) an arbitrary point in Rn, then the position Yat timetis y1=p1+kt(q1−p1) y2=p2+kt(q2−p2)... yn=pn+kt(qn−pn). Now consider the mapping Mt=P+kt(Q−P) . The derivative att=t1is dM dt/vextendsingle/vextendsingle t=t1= lim t2→t1M(t2)−M(t1) t2−t1 4.4.L:R1→RN. PARAMETRIZED STRAIGHT LINES. 181 It represents the velocity at t=t1. To have this make sense, we must introduce a norm inRnso that the limit can be defined. Use the Euclidean norm (although any other one could be used, for it turns out that there is no need for a limit in the case of a straight line). SinceM(t2)−M(t1) =P+kt2(Q−P)−[P+kt1(Q−P)] =k(t2−t1)(Q−P) , we have M(t2)−M(t1) t2−t1=k(Q−P), so dM dt(t) =k(Q−P). Because this is independent of t, it is the derivative at anytimet. Thus, the derivative is a vector,k(Q−P) . The derivative represents the velocity of a particle moving on the line. The speed is the length of the velocity vector, speed = /bardblk(Q−P)/bardbl. What is the slope of the line? Since the line is the path of a mapping, it should not depend on which mapping is used. In terms of mechanics, the slope should not depend on the speed of the particle moving along the line, but only that it moved along the straight line, that is its velocity vector was along the line. Thus we define the slope as a unit vector in the direction of the velocity. In our case, slope = Q−P//bardblQ−P/bardbl. This is a unit vector from PtoQand only depends upon the mapping to specify a positive direction (orientation) for the line. Example: A particle moves on a straight line from P= (1,−2,1) att= 0 toQ= (3,1,−5) att= 2 . Find the position of the particle as a function of time, the velocity and speed of the particle, and slope of the path. The equation of the path is Mt=P+kt(Q−P) , wherekis determined from Q=M(2) =P+ 2k(Q−P) , sok=1 2. ThusM(t) = (1,−2,1) +1 2t(2,3,−6) = (1 +t,2 +3 2t,1−3t) . Velocity =1 2(Q−P) = (1,3 2,−3) . Speed = /bardblvelocity /bardbl=7 2. Slope = velocity /speed = (2 7,3 7,−6 7) . A glance at the formulas which precede the example reveals that the position of a particle which moves along a straight line through Pcan be written in any of the forms 1. M (t) =P+kt(Q−P). whereQis another point on the path and the particle is at Qwhent=1 k, 2. M (t) =P+dM dtt or 3. M (t) =P+Vt, whereVis the velocity. See Exercise 5 too. Exercises (1) (a) If L:R1→R2such that the point X1= 3∈R1is mapped into P= (1,0) , which of the following points are in R(L) i) (2,0) , ii) (1,2) , iii) ( −1,0) ? (b) Sketch two pictures, one of the graph of L, the other of the path of L. (c) Find another operator ˜L:R1→R2whose path is the same as that for L. 182 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 (2) Find a mapping whose path is the straight line passing through the points (2 ,−1,3) and (1,−3,−5) . Find its slope too. (3) If a point is at (1 ,−1,0) att= 0 and at (2 ,3,8) att= 3 , find the position as a function of time if the particle moves along a straight line. What is the velocity and speed of the particle? (4) If a particle is initially at (0 ,1,0,1) and has constant velocity (1 ,−2,3,−1) , find its position as a function of time. Where is it at t= 3 ? (5) A particle moves along a straight line in such a way that at t=t0it is at ˜P, while att=t1it is at ˜Q. (a) Show that its position M(t) as a function of time is M(t) =˜P+ (t−t0)˜Q−˜P t1−t0 (b) What is the velocity? (c) Show that M(t) =M(t0) +dM dt(t−t0). (6) Two straight lines are parallel if they have the same slope. If M(t) =P+t(Q−P) is a parametric equation of one line, find an equation for the parallel line which passes through the point ˜P. 4.5 L:Rn→R1. Hyperplanes. Whereas in the previous section we examined linear mappings from a one-dimensional linear space into an ndimensional space, now we shall look at the opposite extreme, linear mappings from an ndimensional space into a one- dimensional space. LetL:Rn→R1. We would like to find a representation theorem for this linear operator. The most natural way to do this is to work with a basis {e1,...,e n}forRn. Then every X∈Rncan be written as X=n/summationdisplay 1xkek. Consequently, LX=L/parenleftBiggn/summationdisplay 1xkek/parenrightBigg =n/summationdisplay 1L(xkek) =n/summationdisplay 1xkL(ek). It is clear that LXis determined once we know all the numbers Lek. In other words, the linear mapping Lis determined by the effect of the mapping on a basis for the domain of the operator. This proves Theorem 4.14 . LetL:Rn→R1linearly. If {ek}is a basis for the domain of L,Rn, then LX=a1x1+a2x2+...+anxn=n/summationdisplay k=aakxk, 4.5.L:RN→R1. HYPERPLANES. 183 whereX=n/summationdisplay 1xkekandak=Lek. Notice that the akare scalars since they are in the range ofL—and the range of LisR1by hypothesis. Examples: (1) Consider the linear operator L:R3→R1, which maps L:e1= (1,0,0)→1, L:e2= (0,1,0)→0 , andL:e3= (0,0,1)→0 . Since the ekconstitute a basis for R3, the mappingLis completely determined by using Theorem 14. If X= (x1,x2,x3)∈R3, thenX=x1e1+x2e2+x3e3. Thus LX=x1Le1+x2Le2+x3Le3=x1−x2 or LX=x1. For example, L: (2,1,7)→2 . The nullspace of L—those points X∈R3such that LX= 0 —are the points X= (x1,x2,x3)∈R3such thatx1= 0 which is the x2x3 plane. (2) LetL:R4→R1such thatLe1= 1 ,Le2=−2 ,Le3= 5 ,Le4=−3 , where e1= (1,0,0,0) ,e2= etc. Then if X= (x1,x2,x3,x4)∈R4, we have LX=x1−2x2+ 5x3−3x4. The nullspace of Lis again a hyperplane, the hyperplane x1−2x2+ 5x3−3x4= 0 inR4. So far we have not given any attention to the range of L, all of our pictures being in the domain of L. Since the range is R1, its picture is a simple straight line which is not very interesting. However the graph of Lis interesting. Let L:Rn→R1andY= (y)∈R1. Then y=a1x1+...+anxn. The graph of Lis the set of points ( X,LX ) [or (X,Y) whereY=LX] inRn×R1∼= Rn+1. A point (X,Y) = (x1,...,x n,y)∈Rn×R1is on the graph if the coordinates satisfy the equation y=a1x1+...+anxn. This equation can be written as 0 = a1x1+...+ anxn+ (−1)ywhich is a hyperplane in Rn+1. Thus we have found two ways to associate a hyperplane with L:Rn→R1, i) AllXsuch thatLX= 0 , which is the nullspace of L, a linear space of dimension n−1 (since dim N(L) = dim D(L)−dimR(L) =n−1) . ii) The graph of L, that is, all points of the form ( X,LX ) , is a linear space of dimension n+ 1 . Although this is confusing, both ways are used in practice, whichever is most convenient for the problem at hand. For the remainder of this section, we shall confine our attention to hyperplanes defined in the first way. Since linear mappings L:Rn→R1all have the form LX=a1x1+...+anxn, and since it is natural to think of the sum as the scalar product of the vectors N= (a1,...,a n) andX= (x1,...,x n) . Theorem 14 may be rephrased as Theorem 4.15 . IfL:Rn→R1, thenLX=/angbracketleftN, X/angbracketright, whereNis the vector N= (Le1,...,Le n)and{ek}form a basis for Rn. 184 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 Remark: The vector Nis an element of the so-called dual space ofRn. From the above, it is clear that the dual space of Rnalso has dimension n. Theorem 14’ is a “representation theorem”. It states that every linear mapping L:Rn→ R1may be represented in the form LX:=/angbracketleftN, X/angbracketrightfor some vector Nwhich depends on L. You may wish to think of Nas a vector perpendicular to the hyperplane LX= 0 (cf. Ex. 8, p. 225). Example: Consider the operator Lof Example 2 in this section. For it, LX=/angbracketleftN, X/angbracketright whereNis the particular vector N= (1,−2,5,3) . Recall that a linear functional is a linear operator lwhose range is R1. Since the operatorsL:Rn→R1we are considering have range R1, they are all linear functionals. We may again rephrase Theorem 14 in this language. It states that every linear functional defined on Rnmay be represented in the form l(X) =/angbracketleftN, X/angbracketright, whereNdepends on the functional lat hand. This is just a restatement of Theorem 14 with the realization that ourL’s are linear functionals. Don’t let the excess language bewilder you. So far in this section, we have concentrated our attention on the algebraic representation of a linear operator (functional) L:Rn→R1. Let us turn to geometry for a bit. In passing we observed that the nullspace of the operator was a hyperplane in the domain of L(a hyperplane in a linear space Vis a “flat” subset of Vwhose dimension is one less than V, that is, of codimension one). These hyperplanes, {X∈Rn:LX= 0}, all passed through the origin of Rn. A plane parallel to this one which passes through the particular point X0∈Rnhas the form L(X−X0) = 0. It is clear that the point X=X0does satisfy the equation. From the representation theorem, L(X−X0) =a1(x1−x0 1) +a2(x2−x0 2) +...+an(xn−x0 n) = 0, is the equation of this hyperplane, where X= (x1,x2,...,x n) andX0= (x0 1,x0 2,...,x0 n) . If we again write N= (a1,a2,...,a n) , then the equation of the hyperplane is /angbracketleftN, X−X0/angbracketright= 0, all vectors Xsuch thatX−X0is perpendicular to N. Examples: (1) Find the equation of a plane which passes through the point X0= (1,2,−5) and is parallel to the plane −2x1+ 7x2+ 4x3= 0 . Solution : HereN= (−2,7,4) ,X= (x1,x2,x3) , so the plane has the equation 0 =/angbracketleftN, X−X0/angbracketright=−2(x1−1) + 7(x2−2) + 4(x3+ 5), which may be written as −2x1+ 7x2+ 4x3=−8. The equation has been cooked up so that X0= (1,2,−5) does satisfy it. (2) Find the equation of a plane which passes through the point X0= (1,2,−5) and is parallel to the plane −2x1+ 7x2+ 4x3= 37 . Solution : Since this plane is also parallel to the plane −2x1+ 7x2+ 4x3= 0 , the solution is that of Example 1. 4.5.L:RN→R1. HYPERPLANES. 185 (3) Find the equation of the plane in R4which is perpendicular to the vector N= (1,−2,3,1) and passes through the point X0= (1,0,1,−1) . Easy. The plane is all pointsXsuch that /angbracketleftN, X−X0/angbracketright= 0, that is (x1−1)−2(x2−0) + 3(x3−1) + (x4+ 1) = 0, or x1−2x2+ 3x3+x4= 3. (4) Find the equation of the plane in R3which passes through the three points X1= (7,0,0), X2= (1,0,−2), X3= (0,5,1). We shall find this by using the general equation of a plane, a1(x1−x0 1) +a2(x2−x0 2) +a3(x3−x0 3) = 0. HereX0= (x0 1,x0 2,x0 3) is a particular point on the plane. We may use any of X1,X2, orX3for it. Since X1is simplest, we take X0= (7,0,0) . All that remains is to find the coefficients a1,a2, anda3in a1(x1−7) +a2x2+a+ 3x3= 0. SinceX2andX3are in the plane (and so must satisfy its equation), the substitution X=X2andX=X3yields two equations for the coefficients, a1(1−7) +a20 +a3(−2) = 0 a1(0−7) +a2(5) +a3(1) = 0. These two equations in three unknowns may be solved for any two in terms of the third. We find a3=−3a1anda2= 2a1, so the equation is a1(x1−7) + 2a1x2−3a1x3= 0. Factoring out the coefficient a1, we obtain the desired equation x1−7 + 2x2−3x3= 0. (It is clear from the general equation of a plane that the coefficients are determined only to within a constant multiple). Exercises (1) LetL:R2→R1map L: (1,0)→3, L : (0,1)→ −2. WriteLXin the form Lx=a1x1+a2x2.L: (7,3)→? 186 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1 (2) LetL:R2→R1map L: (2,1)→1, L : (0,3)→ −2. WriteLXin the form LX=a1x1+a2x2.L: (7,3)→? (3) Find the equation of a plane in R3which passes through the point (3 ,−1,2) and is parallel to the plane x1−x2−2x3= 7 . (4) Find the equation of a plane in R5which is perpendicular to the vector N= (6,2,−3,1,−1) and contains the point (1 ,1,1,1,4) . (5) Find the equation of a plane in R4which contains the four points X1= (2,0,0,0) , X2= (1,0,2,0) ,X3= (0,−1,0,−1) ,X4= (3,0,1,1) . (6) In this problem, you will have to use the norm induced by the scalar product. a). Show that the distance between the point Y∈Rnand the plane A={X∈ Rn:/angbracketleftN, X−X0/angbracketright= 0}is d(Y,A) =/vextendsingle/vextendsingle/angbracketleftN, Y−X0/angbracketright/vextendsingle/vextendsingle /bardblN/bardbl. b). Prove that the distance between the parallel planes A={X∈Rn:/angbracketleftN, X−X1/angbracketright= 0}andB={X∈Rn:/angbracketleftN, X−X2/angbracketright= 0}is d(A,B) =/vextendsingle/vextendsingle/angbracketleftN, X2−X1/angbracketright/vextendsingle/vextendsingle /bardblN/bardbl. Chapter 5 Matrices and the Matrix Representation of a Linear Operator 5.1 L:Rm→Rn. The simplest example of a linear operator Lwhich maps RmintoRnis supplied by n linear algebraic equations with mvariables. Let X= (x1,...,x m)∈Rm. Then we define LX= a11x1+a12x2+···+a1mxm a21x1+a22x2+···+a2mxm ............................ an1x1+······ +anmxm (5-1) Notice the right side of this equation is a (column) vector with ncomponents. If we let Y= (y1,...,y n) , then the equation LX=Yor m/summationdisplay j=1aijxj=yi, i = 1,2,...,n, determines a vector YinRnfor everyXinRm. Since the operator Lis essentially specified by the coefficients a11,a12,...,a nm, it is convenient to represent it by the notation L= a11a12···alm a21a22···a2m ................... anlan2···anm , and use the notation LX= a11a12···alm a21a22···a2m ................... anl........ a nm  x1 x2 . xm (5-2) The ordered array of m×ncoefficients is called a matrix associated with L, and the numbersaijare called the elements of the matrix. The first index irefers to the rowwhile 187 188 CHAPTER 5. MATRIX REPRESENTATION the second index jrefers to the column . We may also write L= ((aij)) as a shorthand to refer to the whole matrix. Since we shall only use linear operators in this chapter it is convenient to drop the letter Lfor the operator and use A= ((aij)) instead. This will facilitate the notation when referring to other matrices B= ((bin) , etc. since there will be enough subscripts without adding to the confusion by using L1,L2, etc. for linear operators. In this section we shall work out the meaning of operator algebra applied to the special case of operators L:Rm→Rnwhich are represented by matrices. It turns out that every operatorL:Rm→Rncan be represented by a matrix (proved later in this very section). Let us first i) define equality, ii) exhibit the matrices for the zero operator O(X) = 0 (additive identity). If A= ((aij)) andB= ((bij)) both map Rm→Rn, then by definition, A=Bif and only if AX=BX for everyX∈Rm, that is, for all X= (x1,x2,...,x m) , ai1x1+ai2x2+···+aimxm=bilxl+···+bimxm, i= 1,2,...,n orm/summationdisplay j=1aijxj=m/summationdisplay j=1bijxj, i = 1,2,...,n. Subtracting, we find that m/summationdisplay j=1(aij−bij)xj= 0, i = 1,2,...,n must hold for any choice of X= (x1,x2,...,x m) . From the particular choice X= (1,0,0,...0) , we see that ai1−bi1= 0, i = 1,2,...,n, that is, a11=b11,a21=b21,...,a nl=bnl. Similarly, by using other vectors X, we conclude Theorem 5.1 1 (equality ). IfA= ((aij))andB= ((bij))both map Rm→Rn, then A=Bif and only if the corresponding elements of their matrices are equal, aij=bij, i = 1,2,...,n, j = 1,2,...,m. It is clear that the n×mmatrix all of whose elements are zero 0 = 0 0 ··· 0 0 0 ··· 0 ............ 0 0 ··· 0  has the property that it maps every X∈Rminto zero, and thus satisfies the conditions for the zero matrix. That this is the only such matrix follows from Theorem 1, since any other matrix which acts the same way on every vector X∈Rmmust have the same elements - all zeroes. 5.1.L:RM→RN. 189 Theorem 5.2 2. The zero matrix 0:Rm→Rnis uniquely represented by a matrix with nrows andmcolumns, all of whose elements are zero. How is the identity matrix Idefined? Since I:Rn→Rnmaps every vector into itself,IX=X, the linear equations (1) must have the property that given any vector X= (x1,x2,...,x n)∈Rn, then n/summationdisplay j=1δijxj=xi, i = 1,2,...,n Ifaij=δij(the Kronecker delta), so a11=a22=···=ann= 1 while aij= 0, i/negationslash=j, then indeed n/summationdisplay j=1δijxj=xi, i = 1,2,...n is satisfied. Thus, the coefficients of the identity matrix are I= ((δij)) . This is a square (nxm) matrix, I= 1 0 ··· 0 0 1 ··· 0 ............ 0 0 ··· 1  with ones along the main diagonal and zeroes elsewhere. Theorem 5.3 3. The identity matrix I:Rn→Rnis uniquely represented by a square (n×n)matrix whose elements are I= ((δij)). We turn to addition. Let A= ((aij)) andB= ((bij)) be twon×mmatrices, so they both represent operators mapping RmintoRn. Their sum C=A+Bis defined as the operator which acts upon Xaccording to the rule (p. 268) CX=AX+BX, X ∈Rm. The elements cijof the matrix Cconsequently satisfy m/summationdisplay j=1cijxj=m/summationdisplay j=1aijxj+m/summationdisplay j=mbijxj, j = 1,2,...,n. or =m/summationdisplay j=1(aij+bij)xj, j = 1,2,...,n for allX= (x1,x2,...,x m) . Thus, the cijare in fact aij+bij(by Theorem 1) Theorem 5.4 4. IfA= ((aij))andB= ((bij))both map RmintoRn, then their sum C=A+Bhas elements cij=aij+bij. 190 CHAPTER 5. MATRIX REPRESENTATION Remark: From this it follows that the zero matrix is actually the additive identity, for if A= ((aij)) , thenC=A+ 0 has elements cij=aij+ 0 =aij, that is,A+ 0 =A. Example: 1 LetAandBwhich map R3→R4be represented by the matrices A= −3 0 1 7 2 −1 5 4 −3 0 1 1 ;B= 2 2 2 −3 0 0 −4−2 2 0−1−1 . Then A+B= −1 2 3 4 2 −1 1 2 −1 0 0 0 . Example: 2. LetAandBbe the operators on p. 268 (called L1andL2there) which mapR2→R3. Then A= 1 1 1 2 0−1 , B = −3 1 1−1 1 0 , so A+B= −2 2 2 1 1−1  which agrees with the sum obtained there. IfA= ((aij)) , is there a matrix ˜Asuch thatA+˜A= 0 ? Clearly the matrix ˜A defined by ˜A= ((−aij)) does the job since A+˜A= ((aij)) + (( −aij)) = ((0)) by definition of addition. We shall denote the matrix with elements (( −aij)) by “ −A” sinceA+ (−A) = 0 . This matrix “ −A” is the additive inverse to A. Example: If A= 1−1 −π 2 0−1  then −A= −1 1 π−2 0 1 . Since a linear operator which is represented by a matrix is still a linear operator, Theorem 3 (p. 269) certainly holds for matrix addition. We shall rewrite it. Theorem 5.5 LetA,B,C,... be matrices which map RmintoRn(so they are n× mmatrices). The set of all such matrices forms an abelian (commutative) group under addition, that is, 1.A+ (B+C) = (A+B) +C 2.A+B=B+A 3.A+ 0 =A 4. For every A, there is a matrix ( −A) such that A+ (−A) = 0. 5.1.L:RM→RN. 191 Proof: No need to do this again since it was carried out in even greater generality on p. 270. For practice, you might want to write out the proof in the special case of 2 ×3 matrices and see how much more awkward the formulas become when you use the specific elements instead of proceeding more abstractly as we did in the proof on p. 270. Ifαis a scalar and A= ((aij)) is ann×mmatrix which represents a linear operator mapping Rm→Rn, the operator αAisdefined by the rule (αA)X=A(αX) whereXis any vector in Rm. In terms of the elements (( aij)) , this means that the elements ((˜ aij)) ofαAare given by m/summationdisplay j=1˜aijxj=m/summationdisplay j=1aij(αxj), i = 1,2,...,n =m/summationdisplay j=1(αaij)xj, i = 1,2,...,n so ˜aij=αaij. Thus, the matrix αAis found by multiplying each of the elements of Aby α, α a11a12···a1m a21........ a 2m ................... anl... ... a nm = αa11αa12···αa1m αa21αa22···αa2m ....................... αanl...···αanm  Example: −1 7 1 3 −2−1 4 9 6 5 −3 1 −1 = −14−2−6 4 2 −8 −18−12−10 6−2 2 . The following theorem concerns multiplication of matrices by scalars. It is proved either by direct computation - or more simply by realizing that it is a special case of Exercise 12, p. 284. Theorem 5.6 . IfAandBare matrices which map Rm→Rn, and ifα,β are any scalars, then 1.α(βA) = (αβ)A 2.1·A=A 3.(α+β)A=αA+βA 4.α(A+B) =αA+αB. Remark: Theorems 5 and 6 together state that the set of all matrices which map Rminto Rnforms a linear space . It is easy to show that the dimension of this space is m·n(by exhibitingm·nlinearly independent matrices which span the whole space). Now we get more algebraic structure and see how to multiply. Let AmapR?intoRn andB= ((bij)) map RrintoRs. By definition of operator multiplication (p. 271-2), the productABis defined on an element X∈Rr=D(B) by the rule ABX =A(BX). 192 CHAPTER 5. MATRIX REPRESENTATION Since the vector BX∈Rsmust be fed into A, we find that BX∈Rmtoo. Thus, in order for the product AB of ann×mmatrixAwith as×rmatrixBto make sense, we must have ? =m, that is, the range of Bmust be contained in the domain of A, a figure goes here IfC= ((cij)) =AB, then for every X∈Rr CX=A(BX) or r/summationdisplay k=1cijxk=s/summationdisplay j=1aij/parenleftBiggr/summationdisplay k=1bjkxk/parenrightBigg , i= 1,2,...,n so =r/summationdisplay k=1 s/summationdisplay j=1aijbjk xk, i= 1,2,...,n. Therefore, the elements cikof the product ABare given by the formula cik=s/summationdisplay j=1aijbjki= 1,2,...,n k= 1,2,...,r. Since the summation signs have probably overwhelmed you, we repeat it in a special case. LetBbe determined by the linear equations b11x1+b12x2+b13x3=y1 b21x1+b22x2+b23x3=y2. ThenB:R3→R2. Also letA:R2→R2be determined by a11y1+a12y2=z1 a21y1+a22y2=z2. The product AB maps a vector X∈R3first intoY=BX∈R2and then into Z= ABX ∈R2. a figure goes here Ordinary substitution yields Z=ABX as a function of X: a11(b11x1+b12x2+b13x3) +a12(b21x1+b22x2+b21x3) =z1 a21(b11x1+b12x2+b13x3) +a22(b21x1+b22x2+b23x3) =z2, or (a11b11+a12b21)x1+ (a11b12+a12b22)x2+ (a11b13+a12b23)x3=z1 (a21b11+a22b21)x1+ (a21b12+a22b22)x2+ (a21b13+a22b23)x3=z2. 5.1.L:RM→RN. 193 If we write this in the matrix form /parenleftbiggc11c12c13 c21c22c23/parenrightbigg x1 x2 x3 =/parenleftbiggz1 z2/parenrightbigg , we find c11=a11b11+a12b21, c 12=a11b12+a12b22 etc., just as was dictated by the general formula for the multiplication of matrices. Theorem 5.7 . IfA= ((aij))andB= ((bij))are matrices with B:Rr→Rsand A:Rs→Rn, then the product C=AB is defined and the elements of the product C= ((cij))are given by the formula cik=s/summationdisplay j=aaijbjk, i= 1,2,...,n ;k= 1,2,...,r. Remark: Since this formula for matrix multiplication is impossible to remember as it stands, it is fortunate that there is an easy way to remember it. We shall work with the example of matrices A:R2→R2andB:R3→R2discussed earlier. Then AB=/parenleftbigga11a12 a21a22/parenrightbigg /parenleftbiggb11b12b13 b21b22b23/parenrightbigg =/parenleftbiggc11c12c13 c21c22c23/parenrightbigg . To compute the element cik, we merely observe that cik=2/summationdisplay j=1aijbjk=ai1b1k+a12b2k cikis the scalar product of theith row inAwith thekth column in B(see fig.). Thus, the element c21inC=ABis the scalar product of the 2nd row of Awith the 1st column ofB. Do not be embarrassed to use two hands to multiply matrices. Everybody does. Examples: (1) (cf. p. 274 where this was done without matrices). If A=/parenleftbigg2−3 −1 1/parenrightbigg , B =/parenleftbigg0 2 1 1/parenrightbigg , then AB=/parenleftbigg2−3 −1 1/parenrightbigg/parenleftbigg0 2 1 1/parenrightbigg =/parenleftbigg−3 1 1−1/parenrightbigg and BA=/parenleftbigg0 2 1 1/parenrightbigg/parenleftbigg2−3 −1 1/parenrightbigg =/parenleftbigg−2 2 1−2/parenrightbigg . Notice that even though AB andBA are both defined, we have AB/negationslash=BA—the expected noncommutativity in operator multiplication. 194 CHAPTER 5. MATRIX REPRESENTATION (2) (cf. p. 272 bottom where this was done without matrices). If A= 1−1 0 1 −1−2 , B = (1,2,−1), then BA= (1,2,−1) 1−1 0 1 −1−2 = (2,3). However the product ABdoes not make sense. From the general theory of linear operators (Theorem 4, p. 276) we can conclude Theorem 5.8 . Matrix multiplication is associative, that is, if RkA→RlB→RmC→Rn, so the products C(BA)and (CB)Aare defined, then C(BA) = (CB)A. Thus the parenthesis can be omitted without risking chaos. Remark: Returning to linear algebraic equations, you will observe that the matrix notationAX there [eq(2)] can now be viewed as matrix multiplication of the n×m matrixA= ((aij)) with the m×1 matrix (column vector) X. In developing the algebra of matrices - and operators in general - we have been ne- glecting one important issue, that of an inverse operator. If L:V1→V2, can we find an operator ˜L:V2→V1which reverses the effect of L, that is, if LX=Y, where X∈V1andY∈V2, is there an operator ˜Lsuch that ˜LY=X? If so, then ˜LLX =˜LY=X, and we write ˜LL=I. This operator ˜Lis the left(multiplicative) inverse of L. Similarly, an operator ˆL such thatLˆL=Iis the right (multiplicative) inverse ofL. We shall shortly prove that ifan operator Lhas both a left inverse Land a right inverse ˆL, then they are equal, ˆL=˜L, so without ambiguity one can write L−1fortheinverse. a figure goes here 5.1.L:RM→RN. 195 To begin, we compute the inverse of the matrix A=/parenleftbigg5−2 3−1/parenrightbigg associated with the system of linear equations 5x1−2x2=y1 3x1−x2=y2. These equations specify a mapping from R2intoR2. They map a point XintoY. Finding the inverse of Ais equivalent to answering the question, if we are given a point Y, can we find the Xwhence it came? AX=Y, X =A−1Y. Finding the Xin terms of Ymeans solving these two equations, a routine task. The answer is x1=−y1+ 2y2 x2=−3y1+ 5y2.so/parenleftbiggx1 x2/parenrightbigg =/parenleftbigg−1 2 −3 5/parenrightbigg/parenleftbiggy1 y2/parenrightbigg Thus, X=A−1Y, where A−1=/parenleftbigg−1 2 −3 5/parenrightbigg . The matrix A−1is the matrix inverse to A. It is easy to check that AA−1=/parenleftbigg5−2 3−1/parenrightbigg/parenleftbigg−1 2 −3 5/parenrightbigg =/parenleftbigg1 0 0 1/parenrightbigg =I and A−1A=/parenleftbigg−1 2 −3 5/parenrightbigg/parenleftbigg5−2 3−1/parenrightbigg =/parenleftbigg1 0 0 1/parenrightbigg =I. Thus, this matrix A−1is both the right and left inverse of A. Our second example is of a more geometric nature. We shall consider a matrix Rwhich represents rotation of a vector in E2through an angle α. a figure goes here Ris represented by the matrix (cf. Ex. 13b p. 285) R=/parenleftbiggcosα−sinα sinα cosα/parenrightbigg . It is geometrically clear that in inverse of this operator Ris an operator which rotates through an angle −α, unwinding the effect of R. Thus, immediately from the formula for R, we find R−1=/parenleftbiggcos(−α)−sin(−α) sin(−α) cos( −α)/parenrightbigg =/parenleftbiggcosαsinα −sinαcosα/parenrightbigg . 196 CHAPTER 5. MATRIX REPRESENTATION To check that geometry has not deceived us, we should multiply out RR−1andR−1R. Do it. You will find RR−1=R−1R=I. One could also have found R−1by solving linear algebraic equations as was done in the first example. The problem of finding the matrix inverse to any square matrix, A= a11a12···a1n a21a22···a2n ...... an1an2···ann  is equivalent to the dull problem of solving nlinear algebraic equations in nunknowns a11x1+··· +a1nxn=y1 a21x1+··· +a2nxn=y2 ...... an1x1+··· +annxn=yn forXin terms of Y, X =A−1Y. Forn= 2 the computation is not too grotesque, and yields the formulas x1=a22 ∆y1−a12 ∆y2 x2=−a21 ∆y1+a11 ∆y1 where ∆ = a11a22−a12a21(= determinant of A, for those who have seen this before). From this formula we read off that the inverse of the 2 ×2 matrix A=/parenleftbigga11a12 a21a22/parenrightbigg isA−1=1 ∆/parenleftbigga22−a12 −a21a11/parenrightbigg . As a check, one computes that AA−1=A−1A=I. Thus the 2×2matrixAhas an inverse if and only if ∆ :=a11a22−a12a21/negationslash= 0 . a figure goes here Fortunately, one rarely needs the explicit formula for the inverse of a square n×n matrix other than the reasonable cases n= 2 andn= 3 . The inverse of a matrix has greater conceptual use as the inverse of an operator. Having relegated the computation of the inverse of a matrix to the future, let us see what can be said about the inverse without computation. This will necessarily be a bit more abstract. Since the issues involve solving systems of linear algebraic equations, we shall invoke the theory concerning that which was developed in Chapter 4 Section 3. For this discussion, it is convenient to use the following definition (cf. p. 6). Definition: An operator A:V1→V2isinvertible if it has the two properties i) IfX1/negationslash=X2thenAX 1/negationslash=AX 2(injective, 1-1) ii) To every Y∈V2, there is at least one X∈V1such thatAX=Y(surjective, onto). 5.1.L:RM→RN. 197 Thus, an operator is invertible if and only if it is bijective. An invertible matrix is usually called non-singular , while a matrix which is not invertible is called singular . To show that this definition is identical with the previous one, we must show that every invertible linear operator Ahas a right and left inverse. A more pressing matter though, is Theorem 5.9 . If the linear operator A:V1→V2whereV1andV2are finite dimen- sional, is invertible, then dimV1= dimV2, so a matrix must necessarily be square for an inverse to exist (but being square is not sufficient, as was seen in the 2×2case where the additional condition a11a22−a21a12/negationslash= 0 we needed). In other words, you haven’t got a chance to invert a matrix unless it is square, but being square is not enough. Proof: Condition i) states that N(A) = 0 , for if X1/negationslash= 0 , thenAX 1/negationslash= 0 . Therefore dimR(A) = dim D(A)−dimN(A) = dimV1−0 = dimV1. On the other hand, condition ii) states that V2⊂R(A) . SinceA:V1→V2, we know that R(A)⊂V2. Therefore R(A) =V2. Coupled with the first part, we have dimV1= dim R(A) = dimV2. Theorem 5.10 . Given an operator Awhich is invertible, there is a linear operator A−1 such thatAA−1=A−1A=I. Proof: If˜Y∈V2, there is an ˜X∈V1such thatA˜X=˜Y(by property ii), and that ˜X is unique (property i). Therefore without ambiguity we can define A−1˜Y=˜X. A similar process defines the operator A−1for everyY∈V2. From our construction, it is clear (or should be) that AA−1=A−1A=I. All that remains is to show A−1is linear. If A˜X=˜YandAˆX=˜Y, then since Ais linear, A(a˜X+bˆX) =aA˜X+bAˆX=a˜Y+bˆY. ThusA−1(a˜Y+bˆY) =a˜X+bˆX=aA−1˜Y+bA−1ˆY. Remark: Glancing over this proof, it should be observed that finite dimensionality (or even the concept of dimension) never entered - so the result is true for infinite dimensional spaces. Furthermore, linearity was only used to show that A−1was linear. Thus the theorem (except for the claim that A−1is linear) is true for nonlinear operators as well. Needless to say, this construction of A−1one point at a time is useless as a method for findingA−1(since even in the simplest case A:R1→R1it involves an infinite number of points). This theorem shows that if an operator Ais invertible, then there are right and left inverses which are equal AA−1=A−1A=I. We can reverse the theorem and prove Theorem 5.11 . Given the linear operator A:V1→V2, if there are linear operators ˆA (right inverse) and ˜A(left inverse) such that AˆA=˜AA=I, thenAis invertible and A−1=ˆA=˜A. 198 CHAPTER 5. MATRIX REPRESENTATION Proof: Verify condition i: If AX 1=AX 2, then ˜AAX 1=˜AAX 2. Since ˜AA=I, this impliesX1=X2. Verify condition ii. If Yis any element in V2, letX=ˆAY. ThenAX=AˆAY=Y, so thatYis the image of Xunder the mapping. The proof that A−1=˜A=ˆAis delightfully easy. Only the associative property of multiplication is used: ˆA= (A−1A)ˆA=A−1(AˆA) =A−1= (˜AA)A−1=˜A(AA−1) =˜A. Examples: (1) The identity operator Ion every linear space is invertible, for it trivially satisfies both criteria. Not only that, but it is its own inverse for II=I. (2) The zero operator is never invertible, for even though X1/negationslash=X2, we always have 0(X1) = 0 = 0(X2) . (3) The 2 ×2 matrix A=/parenleftbigg1 3 −2−6/parenrightbigg is not invertible since, from the formula AX=/parenleftbigg1 3 −2−6/parenrightbigg/parenleftbiggx1 x2/parenrightbigg =/parenleftbiggx1+ 3x2 −2x1−6x2/parenrightbigg , we see that the vector ( −3,1)/negationslash= 0 is mapped into zero by A(whereas criterion i). states that only 0 can be mapped into 0 by an invertible linear operator). Another way to see that Ais not invertible is to observe that ∆ = a11a22−a12a21= 0 . thus violating the explicit condition for 2 ×2 matrices found earlier. In this last example, we observed that if a linear operator Ais invertible, then by property i) the equation AX= 0 has exactly one solution X= 0 . IfA:V1→V2on a finite dimensional space, and dim V1= dimV2the converse is true also. Theorem 5.12 If the linear operator Amaps the linear space V1intoV2anddimV1= dimV2<∞, then Ais invertible ⇐⇒AX= 0impliesX= 0. Proof: ⇒A restatement of condition i) in the definition. ⇐A restatement of lines 7-10 on page 316. Corollary 5.13 . A square matrix A= ((aij))is invertible if and only if its columns A1= a11 a21 · · · an1 ,A2= a12 a22 · · · an2 ,...,An= a1n a2n · · · ann  are linearly independent vectors. 5.1.L:RM→RN. 199 Proof: To test for linear independence, we examine xaA1+x2A2+···+xnAn= 0, and try to prove that x1=x2=···=xn= 0 . But writing the equation in full, it reads a11x1+a12x2+··· +a1nxn= 0 a21x1+a22x2+··· +a2nxn= 0 · · · · · · · · · an1x1+an2x2+··· +annxn= 0, or AX= 0. By the theorem, Ais invertible if and only if the equation AX= 0 has only the solution X= 0 . Thus Ais invertible if and only if the only solution of x1A1+x2A2+···+xnAn= 0 isx1=x2=···=xn= 0 . We close our discussion of invertible operators with Theorem 5.14 . The set of all invertible linear operators which map a space into itself constitutes a (non- commutative) group under multiplication; that is, if L1,L2,... are invertible operators which map Vinto itself then they satisfy 0. Closed under multiplication ( L1L2is an invertible linear operator which maps V into itself). (1)L1(L2L3) = (L1L2)L3- Associative (2)There is an identity Isuch that IL=LI=L. (3)For every operator Lin the set, there is another operator L−1for which LL−1=L−1L=I. Proof: 0)L1L2is a linear operator which maps Vinto itself by part 0. of Theorem 4 (p. 276). It is invertible since its inverse can be written in the explicit form (an important formula) (L1L2)−1=L−1 2LL−1 1, as we will verify: (L1L2)(L−1 2L−1 1) =L1(L2L−1 2)L−1 1=L1IL−1 1=L1L−1 1=I connect these?? ( L−1 2L−1 1)(L1L2) =L−1 2(L−1 1L1)L2=L−1 2IL2=L−1 2IL2=L−1 2L2=I. (1) Part 1 of Theorem 4 (p. 276). 200 CHAPTER 5. MATRIX REPRESENTATION (2) Part 1 of Theorem 5 (p. 277) (3) A direct restatement of the fact that our set consists only of invertible operators. Closely associated with a matrix A:Rn→Rm A= a11a12···a1n a21··· ··· · · · · · · · am1am2···amn . is another matrix A∗, the transpose oradjoint ofA, which is obtained by interchanging the rows and columns of A, viz. A∗= a11a21···am1 a12a22··· · · · · · · · a1n··· ···amn . For example, ifA= 1 2 4−2 5−2 ,thenA∗=/parenleftbigg1 4 5 2−2−1/parenrightbigg . IfA= ((aij)) , thenA∗= ((aij)) . The adjoint of an m×nmatrix is an n×mmatrix. Thus, ifA:Rn→RmthenA∗:Rm→Rn, and for any Z∈Rm, we have A∗Z= a11a21···am1 a12a22··· · · · · · · · a1n··· ···amn  z1 · · · zm = a11z1+a21z2+··· +am1zm a12z1+··· ··· +am2zm · · · · · · a1nz1+··· ··· +amnzm , so thejth component ( A∗Z)jof the vector A∗Z∈Rnis (A∗Z)j=m/summationdisplay i=1aijzi=a1jz1+a2jz2+···+amjzm. Beware : The classical literature on matrices uses the term “adjoint of a matrix” for an entirely different object. Our nomenclature is now standard in the theory of linear operators. A real square matrix Ais called symmetric orself-adjoint ifA=A∗. For example, A= 7 2 −3 2−1 5 −3 5 4 =A∗. For a symmetric matrix A, we haveaij=aji. 5.1.L:RM→RN. 201 The significance of the adjoint of a matrix (as well as its relation to the more general conception of the adjoint of an arbitrary operator) arises in the following way. If A:En→ Em, then for any XinEnthe vectorY=AX is a vector in Em. We can form the scalar product of this vector Y=AX with any other vector ZinEm(becauseYandZare both in Em /angbracketleftZ, Y/angbracketright=/angbracketleftZ, AX /angbracketright. SinceA∗:Em→En, andZ∈Em, thenA∗Zmakes sense, and is a vector in En, so /angbracketleftA∗Z, X/angbracketrightis a real number for any X∈En.Claim : /angbracketleftZ, AX /angbracketright=/angbracketleftA∗Z, X/angbracketright. This is easy to verify. Let A= ((aij)) . Then (AX)i=n/summationdisplay j=1aijxjand (A∗Z)j=m/summationdisplay i=1aijzi, so that /angbracketleftZ, AX /angbracketright=m/summationdisplay i=1zi(AX)i=m/summationdisplay i=1zi(n/summationdisplay j=1aijxj) =m/summationdisplay i=1n/summationdisplay j=1ziaijxj. In the same way, /angbracketleftA∗Z, X/angbracketright=n/summationdisplay j=1(A∗Z)jxj=n/summationdisplay j=1(m/summationdisplay i=1aijzi)xj =m/summationdisplay i=1n/summationdisplay j=1ziaijxj. Comparison reveals we have proved Theorem 5.15 . IfA:En→Em, then for any X∈Enand anyZ∈Em, /angbracketleftZ, AX /angbracketright=/angbracketleftA∗Z, X/angbracketright, whereA∗is the adjoint of A. Remark: From a more abstract point of view, the operator A∗is usually defined as the operator which has the above property. If this definition is adopted, one must use it to prove the adjoint A∗of a matrix Ais found by merely interchanging the rows and columns (try to do it!). It is remarkably easy to obtain some properties of the adjoint by using Theorem 14. Our attention will be restricted to square matrices (although the results are still true with but minor modifications for a rectangular matrix). 202 CHAPTER 5. MATRIX REPRESENTATION Theorem 5.16 . LetAandBben×nmatrices (so the products AB,BA,B∗A∗, A+Betc. are all defined). Then 0.I∗=I(becauseIis symmetric) 1.(A∗)∗=A 2.(AB)∗=B∗A∗ 3.(A+B)∗=A∗+B∗. 4.(cA)∗=cA∗, cis a real scalar. 5.Ais invertible if and only if A∗is invertible, and (A∗)−1= (A−1)∗. 6.Ais invertible if and only if the rows of Aare linearly independent. Proof: We could use subscripts and the aijstuff - but it is clearer to use the result of Theorem 14. In order to do so, an important preliminary result is needed. Theorem 5.17 . IfC:En→Em, then the equation /angbracketleftCx, Y /angbracketright= 0 for allXinEnandYinEm ⇐⇒Cis the zero operator, C= 0. Thus ifC1andC2mapEninto itself, the equation /angbracketleftC1X, Y/angbracketright=/angbracketleftC2X, Y/angbracketrightfor allX,Y∈En⇐⇒C1=C2. Proof: ⇒By contradiction, if C/negationslash= 0 there is some X0such that 0 /negationslash=CX 0∈En. Now just pickY0=CX 0. Then 0 =/angbracketleftCX 0, Y0/angbracketright=/angbracketleftCX 0, CX 0/angbracketright=/bardblCX 0/bardbl2>0 because by assumption CX 0/negationslash= 0 . A glance at this line reveals the desired contradiction. ⇐Obvious. The last assertion of the theorem follows by subtraction, 0 =/angbracketleftC1X, Y/angbracketright − /angbracketleftC2X, Y/angbracketright=/angbracketleftC1X−C2X, Y/angbracketright=/angbracketleft(C1−C2)X, Y/angbracketright and letting C=C1−C2. Now we return to the Proof of Theorem 15 : The vectors X,Z will be in En. (0) Particularly clear because Iis symmetric. You should try constructing another proof patterned on those below. (1) Two successive interchanges of the rows and columns of a matrix leave it unchanged. Again, try to construct another proof patterned on those below. (2)/angbracketleft(AB)∗Z, X/angbracketright=/angbracketleftZ, ABX /angbracketright=/angbracketleftZ, A(BX)/angbracketright=/angbracketleftA∗Z, BX /angbracketright =/angbracketleftB∗(A∗Z), X/angbracketright=/angbracketleft(B∗A∗)Z, X/angbracketright for allX,Z inEn. Application of Theorem 16 yields the result. 5.1.L:RM→RN. 203 (3)/angbracketleft(A+B)∗Z, X/angbracketright=/angbracketleftZ,(A+B)X/angbracketright=/angbracketleftZ, AX +BX/angbracketright =/angbracketleftZ, AX /angbracketright+/angbracketleftZ,BX, =/angbracketright/angbracketleftA∗Z, X/angbracketright+/angbracketleftB∗Z, X/angbracketright =/angbracketleftA∗Z+B∗Z, X/angbracketright=/angbracketleft(A∗+B∗)Z, X/angbracketright. And apply Theorem 16. (4)/angbracketleft(cA)∗Z, X/angbracketright=/angbracketleftZ, cAX /angbracketright=c/angbracketleftZ, AX /angbracketright =c/angbracketleftA∗Z, X/angbracketright=/angbracketleft(cA∗)Z, X/angbracketright. Apply Theorem 16. (5) IfAis invertible, then AA−1=A−1A=I. An application of parts 0 and 2 shows (A−1)∗A∗= (AA−1)∗=I∗=I. Similarly,A∗(A−1)∗=I. ThusA∗has a left and right inverse, so it is invertible by Theorem 11. The above formulas reveal ( A∗)−1= (A−1)∗. In the other direction, assume A∗is invertible. Since A∗∗=A(part 1) the matrix Ais the adjoint of A∗. But we just saw that if a matrix is invertible then its adjoint is too. Thus the invertibility of A∗implies that of A. (6) By the Corollary to Theorem 12, A∗is invertible if and only if its columns are linearly independent. Since the columns of A∗are the rows of A, we find that A∗is invertible if and only if the rows of Aare linearly independent. Coupled with Part 5, the proof is completed. In our later work we shall need an inequality. Why not insert it here for future reference. Theorem 5.18 . IfA= ((aij))is anm×nmatrix, so A:En→Em, then for any X inEnandYinEm /bardblAX/bardbl ≤k/bardblX/bardbl and |/angbracketleftY, AX /angbracketright| ≤k/bardblX/bardbl /bardblY/bardbl, where k2=m/summationdisplay i=1n/summationdisplay j=1a2 ij. Proof: By definition /bardblAX/bardbl2=m/summationdisplay i=1(AX)2 i=m/summationdisplay i=1(n/summationdisplay j=1aijxj)2, where (AX)iis theith component of the vector AX. The Schwarz inequality shows (n/summationdisplay j=1aijxj)2≤n/summationdisplay j=1a2 ijn/summationdisplay j=1x2 j=/bardblX/bardbl2n/summationdisplay j=1a2 ij. 204 CHAPTER 5. MATRIX REPRESENTATION Thus, /bardblAX/bardbl2≤ /bardblX/bardbl2m/summationdisplay i=1(n/summationdisplay j=1a2 ij) =k2/bardblX/bardbl2, which proves the first part. The second part follows from this and one more application of Schwarz: |/angbracketleftY, AX /angbracketright| ≤ /bardblY/bardbl/bardblAX/bardbl ≤k/bardblX/bardbl/bardblY/bardbl. After all of this detailed discussion of matrices as an example of a linear operator L mapping one finite dimensional space into another, our next theorem will show why matrices are so ubiquitous. You see, we shall prove that every such linear operator L:V1→V2can be represented as a matrix after bases forV1andV2have been selected. Theorem 5.19 . (Representation Theorem) Let Lbe a linear operator which maps one finite dimensional space into another L:V1→V2. Let{e1,e2,...,e n}be a basis for V1, and {θ1,θ2,...θ m}be a basis for V2. Then in terms of these bases Lmay be represented by the matrix θLewhosejth column is the vector (Lej)θ, that is, the vector Lej(which is a vector in V2) written in terms of the θ basis forV2. Pictorially we have θLe= ((Le1)θ···(Len)θ). Proof: Finding the representation of Lin terms of given bases for V1andV2means: given a vector XinV1which is represented in the ebasis forV1(write it as Xe) to find a matrix θLesuch that the image vector θLeXeis the image ( LX)θofXwritten in the θbasis forV2. We have used the cumbersome notation θLeto make explicit the fact that it maps vectors written in the ebasis forV1into vectors written in the θbasisV2. To avoid even further notation, we shall carry out the details only for the particular case where the domain V1is two dimensional with basis {e1,e2}andV2is three dimensional with basis {θ1,θ2,θ3}. The general case is proved in the same way. Since the vectors Le1andLe2are inV2, they can be written in the θbasis, say, Le1=a1θ1+b1θ2+c1θ3, Le 2=a2θ1+b2θ2+c2θ3, so (Le1)θ= a1 b1 c1 and (Le2)θ= a2 b2 c2 . GivenXinV1, it can be written in the ebasis forV1, X=x1e1+x2e2,soXe=/parenleftbiggx1 x2/parenrightbigg . Then LX=L(x1e1+x2e2) =x1Le1+x2Le2 =x1(a1θ1+b1θ2+c1θ3) +x2(a2θ1+b2θ2+c2θ3) = (a1x1+a2x2)θ1+ (b1x1+b2x2)θ2+ (c1x1+c2x2)θ3. 5.1.L:RM→RN. 205 If we write LXas a column vector in the θbasis it is (LX)θ= a1x1+a2x2 b1x1+b2x2 c1x2+c2x3  which is recognized as a product θLe= a1a2 b1b2 c1c2 /parenleftbiggx1 x2/parenrightbigg = a1a2 b1b2 c1c2 Xe therefore, the matrix we want is θLe= a1a2 b1b2 c1c2 ((Le1)θ(Le2)θ), a matrix whose jth column is the vector Lejwritten in the θbasis forV2. Example: Consider the integral operator L: =/integraltextx 0as a map of the two dimensional space P1into the three dimensional space P2. Any bases for P1andP2will do, however we must simply fix our attention to specific bases. Say basis for P1:={e1(x) = 1, e 2(x) =x} basis for P2:={θ1(x) =1+x 2, θ 2(x) =1−x 2, θ 3(x) =x2}. Then Le1=/integraldisplayx 01dt=x=θ1−θ2 and Le2=/integraldisplayx 0tdt=x2 2=1 2θ3. Therefore (Le1)θ= 1 −1 0 ,and (Le2)θ= 0 0 1 2 , so θLe= ((Le1)θ(Le2)θ) = 1 0 −1 0 01 2  is the matrix representing Lin terms of the given ebasis for P1andθbasis for P2. To make you believe this, let us evaluate LP=/integraldisplayx 0P for some polynomial p∈P1by using the matrix. For example, p(x) = 3−x= 3e1−e2, so in theebasis for P1, P 3=/parenleftbigg3 −1/parenrightbigg . Its image under Lin terms of the θbasis for P2is then (Lp)θ=θL(pe) e= 1 0 −1 0 01 2 /parenleftbigg3 −1/parenrightbigg = 3 −3 −1 2 ; 206 CHAPTER 5. MATRIX REPRESENTATION that is, Lp= 3θ1−3θ2−1 2θ3= 3(1 +x 2)−3(1−x 2)−1 2(x2) = 3x−1 2x2 which, of course, agrees with /integraldisplayx 0p(t)dt=/integraldisplayx 0(3−t)dt= 3x−1 2x2. WARNING : If we had used a different basis for either P1orP2, the resulting matrix representing Lwould be different. For example, if the same basis were used for P1but a different basis for P2, ˜θbasis for P2:={˜θ1(x) = 1,˜θ2(x) =x,˜θ3(x) =x2}, then Le1=x=˜θ2 andLe2=x2 2=1 2˜θ3, so (Le1)˜θ= 0 1 0 , (Le2)˜θ= 0 0 1 2 . Therefore the matrix ˜θLewhich represents Lin terms of the ebasis for P1and the ˜θ basis for P2is ˜θLe= 0 0 1 0 01 2 . Again, ifp(x) = 3−x= 3e1−e2, then in the ˜θbasis (Lp)˜θ=˜θLePe= 0 0 1 0 01 2 /parenleftbigg3 −1/parenrightbigg = 0 3 −1 2 ; that is, Lp= 0˜θ1+ 3˜θ2−1 2˜θ3= 3x−1 2x2, to no one’s surprise. Observe that the matrices θLeand ˜θL3both represent L—but with respect to dif- ferent basis. The second matrix ˜θLeis somewhat simpler that the first since it has more zeroes. It is often useful to pick bases in order that the representing matrix be as simple as possible. We shall not discuss that issue right now. There is a simple class of operators (transformations) which are not linear, but enjoy most of the properties which linear ones do. They are affine operators, or affine transfor- mations. To define them, it is best to first define the translation operator. Definition: IfVis any linear space and Y0a particular element of V, then the operator T:V→Vdefined by TY=Y+Y0, Y∈V, is the translation operator . It translates a vector Yinto the vector Y+Y0. 5.1.L:RM→RN. 207 Definition: Anaffine transformation Ais a linear transformation Lfollowed by a trans- lation. ifL:V1→V2andY0∈V2, it has the form AX:=LX+Y0. X ∈V1, Y 0∈V2. Affine transformations can be added and multiplied by the same definition which gov- erned linear transformations. Thus, if AandBare affine transformations mapping V1 intoV2, (A+B)X:=AX+BX. In particular, if AX=L1X+Y0andBX=L2X+Z0, whereY0andZ0are inV2, then (A+B)X=AX+BX=L1X+Y0+L2X+Z0 = (L1+L2)X+ (Y0+Z0). Similarly, if A:V1→V2andB:V3→V4, whereV2⊂V3, then (BA)X:=B(AX) =B(L1X+Y0) =L2(L1X+Y0) +Z0 =L2L1X+L2Y0+Z0, whereY0∈V2andZ0∈V4. You will carry out the (straightforward) proofs of the algebraic properties for affine transformations in Exercise 23. The curtain on this longest of sections will be brought down with a brief discussion of the operators which characterize rigid body motions, or Euclidean motions, as they are often called. Definition: The transformation R:En→Enis an isometric transformation , (orEuclidean transformation or rigid body transformation) if the distance between two points is preserved (invariant) under the transformation. Thus, Ris an isometry if /bardblRX−RY/bardbl=/bardblX−Y/bardbl for allXandYinEn. It is interesting to think for a moment how all these names originated. The phrase rigid body transformation arises from the idea that any motion of a rigid body (such as a translation or rotation) does not alter the distance between any two points in the body. In the framework of Euclidean geometry the whole notion of congruence is defined to be just those properties of a figure which are invariant under isometries. By allowing deformations other than isometries, one obtains geometries, so affine geometry is the study of properties invariant under all affine motions. The study of isometric transformations is mainly contained in that of a special case, orthogonal transformations . These are isometries which leave the origin fixed, R0 = 0 . It should be clear from our next theorem (part 3) that the idea of an orthogonal transformation generalizes the idea of a rotation to higher dimensional space. Reflections (mirror images) are also orthogonal transformations. Theorem 20 states that every isometric transformation is the result of an orthogonal transformation followed by a translation. Example: The matrix R=/parenleftbigg1 0 0−1/parenrightbigg defines an orthogonal transformation since if X=/parenleftbiggx1 x2/parenrightbigg ,thenRX=/parenleftbigg1 0 0−1/parenrightbigg/parenleftbiggx1 x2/parenrightbigg =/parenleftbiggx1 −x2/parenrightbigg , 208 CHAPTER 5. MATRIX REPRESENTATION and if Y=/parenleftbiggy1 y2/parenrightbigg ,thenRY=/parenleftbiggy1 −y2/parenrightbigg . Consequently /bardblRX−RY/bardbl=/bardblX−Y/bardbl=/radicalbig (x1−y1)2+ (x2−y2)2, soR, being isometric and linear is an orthogonal transformation. It represents a reflection across the x1axis. Our definition of an orthogonal transformation does not presume its linearity. This is because the linearity is a consequence of the given properties. A proof is outlined in Ex. 16, p. 390. For convenience, the linearity will be assumed in the following theorem where we collect the standard properties of orthogonal transformations. Theorem 5.20 . LetR:En→Enbe a linear transformation. The following properties of Rare equivalent. (1)Ris an orthogonal transformation, that is /bardblRX−RY/bardbl=/bardblX−Y/bardblandR0 = 0. (2)/bardblRX/bardbl=/bardblX/bardbl (3)/angbracketleftRX, RY /angbracketright=/angbracketleftX, Y/angbracketright(so angles are preserved) (4)R∗R=I (5)Ris invertible and R−1=R∗. (Only in this part do we use the finite dimension- ality of En). Proof: We shall prove the following chain of implications: 1 = ⇒2 =⇒3 =⇒4 =⇒5 =⇒ 4 =⇒1 1 =⇒2 . Trivial, for /bardblRX/bardbl=/bardblRX−R0/bardbl=/bardblX−0/bardbl=/bardblX/bardbl. 2 =⇒3 . By linearity and part 2) applied to the vector X+Y, we have /bardblRX+RY/bardbl=/bardblR(X+Y)/bardbl=/bardblX+Y/bardbl. Now square both sides and express the norm as a scalar product: /angbracketleftRX+RY, RX +RY/angbracketright=/angbracketleftX+Y, X +Y/angbracketright. Upon expanding both sides, we find that /bardblRX/bardbl2+ 2/angbracketleftRX, RY /angbracketright+/bardblRY/bardbl2=/bardblX/bardbl2+ 2/angbracketleftX, Y/angbracketright+/bardblY/bardbl2. Since by part 2) /bardblRX/bardbl=/bardblX/bardbland/bardblRY/bardbl=/bardblY/bardbl, we are done. 3 =⇒4 . By part 3) and Theorem 14 (p. 369), /angbracketleftR∗RX, Y /angbracketright=/angbracketleftRX, RY /angbracketright=/angbracketleftX, Y/angbracketright. Thus, an application of the second part of Theorem 16 (p. 371) gives us R∗R=I. 4 =⇒5 . SinceX=R∗RX, we see that RX= 0 implies X= 0 , consequently, Ris invertible (Theorem 12, p. 364). Moreover R∗R=IsoR∗=R−1. 5 =⇒4 . Clear, since R∗=R−1. 5.1.L:RM→RN. 209 5 =⇒1 . Because Ris linear,R0 = 0 . It remains to show that /bardblRX−RY/bardbl=/bardblX−Y/bardbl, an easy computation. /bardblRX−RY/bardbl2=/bardblR(X−Y)/bardbl2=/angbracketleftR(X−Y), R(X−Y)/angbracketright, so using 4) =/angbracketleftR∗R(X−Y), X−Y/angbracketright=/angbracketleft(X−Y),(X−Y)/angbracketright=/bardblX−Y/bardbl2. Done. Earlier in this section (p. 357-8) we considered a matrix Rwhich represented the operator which rotates a vector in E2through an angle α. This matrix is the simplest (non-trivial) example of a rigid body transformation which leaves the origin fixed, that is, an orthogonal transformation. R=/parenleftbiggcosα−sinα sinα cosα/parenrightbigg . To prove that Ris an orthogonal matrix, by Theorem 19 part 3, it is sufficient to verify /angbracketleftRX, RY /angbracketright=/angbracketleftX, Y/angbracketrightfor allXandYinE2. A calculation is in order here. RX=/parenleftbiggcosα−sinα sinα cosα/parenrightbigg/parenleftbiggx1 x2/parenrightbigg =/parenleftbiggx1cosα−x2sinα x1sinα+x2cosα/parenrightbigg . Similarly for RY, just replace x1andx2byy1andy2respectively. Then /angbracketleftRX, RY /angbracketright= (x1cosα−x2sinα)(y1cosα−y2sinα) + missing? (x1sinα+x2cosα)(y1sinα+y2cosα) =x1y1cos2α−(x1y2+x2y1) sinαcosα+x2y2sin2α +x1y1sin2α+ (x1y2+x2y1) sinαcosα+x2y2cos2α =x1y1+x2y2=/angbracketleftX, Y/angbracketright.Done. We previously found an expression for R−1(p. 358) by geometric reasoning. It is reassuring to notice R−1=R∗, just as part 5 of our theorem states. The most general rotation in E3may be decomposed into a product of these simple two dimensional rotations. For a brief discussion - complete with pictures - open Goldstein, Classical Mechanics to pp. 107-9. Now to the last theorem of this section. Theorem 5.21 . IfR:En→Enis a rigid body transformation, then for every X∈En RX=R0X+X0, whereR0is an orthogonal transformation (rotation) and X0is a fixed vector in En. Thus, every rigid body motion is composed of a rotation (by R0and a translation (through X0). Proof: LetR0X=RX−R0 . Since R00 =R0−R0 = 0, the operator R0has the property R00 = 0 . Furthermore, for any XandYinEn, /bardblR0X−R0Y/bardbl=/bardblRX−R0−RY+R0/bardbl =/bardblRX−RY/bardbl=/bardblX−Y/bardbl. ThereforeR0satisfies the definition of an orthogonal transformation. The proof is com- pleted by defining X0to be the image of the origin under R, X 0=R0 . Then R0X=RX−X0, or RX=R0X+X0. 210 CHAPTER 5. MATRIX REPRESENTATION 5.2 Supplement on Quadratic Forms Quadratic polynomials of the form Q(X) =αx2 1+βx1x2+γx2 2, X = (x1,x2) and the generalization to nvariablesX= (x1,x2,...,x n) Q(X) =n/summationdisplay in/summationdisplay j=1αijxixj often arise in mathematics. They are called quadratic forms and can always be represented in the form /angbracketleftX, SX /angbracketrightwhereSis a self adjoint matrix. For example, the first quadratic form can be written as Q(X) = (x1,x2)/parenleftbiggαβ 2β 2γ/parenrightbigg/parenleftbiggx1 x2/parenrightbigg =/angbracketleftX, SX /angbracketright, whereSis the matrix indicated. The procedure for finding the elements (( aij)) of the matrix Sis simple. First take care of the diagonal terms by letting aiibe the coefficient of x2 1inQ(X) . Realizing that xixj=xixj, collect the terms αijxixjandαjixjxiinQ(X) , getting ( αij+αji)xixj. Then let aij=aji=1 2(αij+αji)i/negationslash=j. Example: Q(X) =x2 1−2x1x3−x2 2+ 6x1x2+ 4x3x1. Rewrite this as Q(X) =x2 1−x2 2+ 6x1x2+ 2x1x3. Then S= 1 3 1 3−1 0 1 0 0  and Q(X) =/angbracketleftX, SX /angbracketright. as you can easily verify. Definition: A quadratic form Q(X) ispositive semi definite ifQ(X)≥0 for allXand positive definite ifQ(X)>0, x/negationslash= 0. Q(X) isnegative semi definite ornegative definite if, respectively, Q(X)≤0 , orQ(X)<0, X/negationslash= 0 . IfSis the self adjoint matrix associated with the quadratic form Q(X) , thenSis positive semi definite, positive definite, etc., if Q(X) has the respective property. We may think of Q(X) as representing a quadratic surface. Thus, if Sis diagonal, for example S= 2 0 0 0 1 0 0 0 3 , with positive diagonal elements, then the equation Q(X) = 1 , where Q(X) =/angbracketleftX, SX /angbracketright= 2x2 1+x2 2+3x2 3, represents an ellipsoid. This matrix Sis positive definite since by inspection Q(X)>0, X /negationslash= 0 . It is easy to see if a diagonal matrix Sis positive semi definite, negative semi definite, positive definite, or negative definite. 5.2. SUPPLEMENT ON QUADRATIC FORMS 211 Example: The diagonal matrix S= γ1...0 ... 0...γ n  is (a) positive semi definite if and only if γ1,...,γ nare all non-negative, (b) positive definite if and only if γ1,...,γ nare all positive (not zero), and the obvious statements for negative semi definite and negative definite. The problem of determining if a non diagonal symmetric matrix is positive etc. is more subtle. We shall find necessary and sufficient conditions for the two variable case, but only necessary conditions for the general case. Consider the 2 ×2 self-adjoint matrix S=/parenleftbigga b b c/parenrightbigg and the associated quadratic form Q(X) =ax2+ 2bxy+cy2. There are several cases. (i)Ifa= 0 , then Q(X) = (2bx+cy)y. Ifb/negationslash= 0 , by choosing xandyappropriately, we can make Q(X) assume bothpositive and negative values. Thus, for a= 0, b/negationslash= 0, Q can be neither a positive nor a negative semi-definite form. On the other hand, if a= 0 , andb= 0 , thenQis positive (negative) semi definite if and only if c≥0 (c≤0) . Ifa= 0, Q can never be positive definite or negative definite since if X= (x,0) wherex/negationslash= 0 , thenQ(X) = 0 butX/negationslash= 0 . (ii)Ifa/negationslash= 0 , thenQcan be written as Q(X) =1 a[(ax+by)2= (ac−b2)y2]. We can immediately read off the conditions from this. Qis positive semi definite (definite) if and only if a>0 andac−b2≥0 (ac−b2>0) , and negative semi definite (definite) if and only if a<0 andac−b2≥0 (ac−b2>0) . In summary, we have proved Theorem 5.22 A. LetQ(X) =ax2+ 2bxy+cy2, andSbe the associated symmetric matrix. Then (a)Qis positive semi definite if and only if a≥0andac−b2≥0(this implies c≥0 too). (b)Qis positive definite if and only if a >0andac−b2>0(this implies c >0 too). 212 CHAPTER 5. MATRIX REPRESENTATION The general case of a quadratic form in nvariables is much more difficult to treat. There are known necessary and sufficient conditions, but they are not too useful in practice, especially for a large number of variables. We shall only prove one necessary condition for a quadratic form to be positive semi-definite (or positive definite), a condition which is both transparent to verify in practice and even easier to prove. THEOREM B. If the self adjoint matrix S= ((aij)) is positive definite, then the diagonal elements must all be positive, a11,a22,...,a nn>0 . Similarly, if Sis negative definite then the diagonal elements must all be negative. Proof:Q(X) =/angbracketleftX, SX /angbracketright=n/summationdisplay i,j=1aijxixj. SinceQis positive definite, Q(X)>0 for all X/negationslash= 0 . In particular, Q(ek)>0, k= 1,...,n , whereekis thekth coordinate vector ek= (0,0,..., 0,1,0,..., 0) . ButQ(ek) =akk. Thusakk>0, k= 1,...,n , just what we wanted to prove. Examples: 1. The quadratic form Q(X) = 3x2+743xy−y2+4z2+xzis positive definite or semi definite since the coefficient of y2is negative. It is not negative definite or semi definite since the coefficient of x2is positive. 2. The quadratic form Q(X) =x2−5xy+y2+ 2z2satisfies the necessary conditions of Theorem B, but the conditions of Theorem D were not sufficient conditions for positive definiteness. Thus, we cannot conclude this Q(X) is positive definite. In fact, this Q(X) is notpositive definite or semi definite since, for example, if X= (1,1,1) , thenQ(X) =−1 . It is clearly not negative definite or semi definite. Exercises (1) Find the self-adjoint matrix Sassociated with the following quadratic forms: (a)Q(X) =x2 1−2x1x2+ 4x2 2. (b)Q(X) =−x2 1+x1x2−x1x3+x2 2−3x2x1−2x3x2+ 3x2 3 (c)Q(X) = 2x1x2−3x3x2+ 4x2x4+x3x4+ 7x2 2 [Answers: (a)/parenleftbigg1−1 −1 4/parenrightbigg , (b) −1−1−1 2 −1 1 −1 −1 2−1 3 , (c) 0 1 0 0 1 7 −3 22 0−3 201 2 0 21 20  (2) Use Theorem AorBto determine which of the following quadratic forms in two variables are positive or negative definite, or semi definite, or none of these. (a)Q(X) =x2 1−2x1x2+ 4x2 2 (b)Q(X) =−x2 1+x1x2−4x2 2 (c)Q(X) =x2 1−6x1x2−4x2 2 (d)Q(X) =x2 1−6x1x2+ 4x2 2 (e)Q(X) =x2 1−6x1x2+ 4x2x3−x2 2+ 4x2 3 (3) If the self-adjoint matrix Sis positive definite, prove it is invertible. Give an example of an invertible self-adjoint matrix which is neither positive nor negative definite. 5.2. SUPPLEMENT ON QUADRATIC FORMS 213 (4) Find all real values for λfor which the quadratic form Q(X) = 2x2+y2+ 3z2+ 2λxy+ 2xz is positive definite. [Hint: Q(X) = (5 3−λ2)x2+ (λx+y)2+ (√ 3z+1√ 3x)2] (5) Let the integer nbe≥3 . If the quadratic form Q(X) =n/summationdisplay i,j=1aijxixj, a ij=aji is the product of two linear forms Q(X) = (n/summationdisplay i=1λixi)(n/summationdisplay j=1µjxj), show that det A= det((aij)) = 0 . (6) If the self-adjoint matrix Sis positive definite or semi-definite, prove the generalized Schwarz inequality : |/angbracketleftY, SX /angbracketright|2≤ /angbracketleftY, SY /angbracketright/angbracketleftX, SX /angbracketright for allXandY. [Hint: Observe [ X,Y] :=/angbracketleftY, SX /angbracketrightsatisfies all the axioms for a scalar product]. (7) If the self-adjoint matrix Sis positive definite (so S−1exists by Exercise 3), prove thatS−1is also positive definite. [Hint: Use the generalized Schwarz inequality, Exercise 6, with Y=S−1Xand the inequality /angbracketleftX, SX /angbracketright ≤k2/bardblX/bardbl2of Theorem 17, p. 373]. (8) Proof or counterexample: (a) If a matrix A= ((aij)) is positive definite, then all of its elements are positive, aij>0 for alli,j. (b) If a matrix Ais such that all of its elements are positive, aij>0 , then the matrix is positive definite. Exercises (1) Write out the matrices associated with the operators AandBin Exercise 4a, p. 281, and carry out the computation there using matrices. (2) Write out the matrices RA,RB, andRCfor the rotation operators A,B , andC in Exercise 8 p. 281 and complete that problem using matrices. [Ans. RA= 1 0 0 0 0 −1 0 1 0 in terms of the basis e1= (1,0,0), e 2= (0,1,0), e 3= (0,0,1) ]. (3) Prove Exercise 2b (p. 281) as a corollary of Theorem 18. 214 CHAPTER 5. MATRIX REPRESENTATION (4) If A=/parenleftbigg0 0 0 1/parenrightbigg , B =/parenleftbigg0 1 0 0/parenrightbigg computeAB, BA , andB2. (5) Compute A−1if (a).A=/parenleftbigg1 2 3 4/parenrightbigg [ans.A−1=/parenleftbigg−2 1 3 2−1 2/parenrightbigg ] (b).A= 4 0 5 0 1 −6 3 0 4  [ans.A−1= 4 0 −5 −18 1 24 −3 0 4 ] (c)A= 1 1 0 0 0 1 1 0 0 0 1 1 0 0 0 1 [ans.A−1= 1−1 1 −1 0 1 −1 1 0 0 1 −1 0 0 0 1 ] (6) IfAis the matrix of 5a) above, from the definition compute directly , (a)−6A−1+1 2A∗[ans./parenleftbigg25 2−9 2 −8 5/parenrightbigg ] . (b) (A∗)−1and (A−1)∗. Compare them. State and prove a general theorem. (c)AA∗andA∗A. (7) If A=/parenleftbigg1 2 3 4/parenrightbigg , B =/parenleftbigg1 1 2−1/parenrightbigg , compute (AB)∗, A∗B∗, andB∗A∗. Compare ( AB)∗andB∗A∗and explain the outcome. (8) Prove that I∗=IandA∗∗=Ausing only Theorems 14 and 16 (cf. Parts 2-4 of Theorem 15). (9) IfA:Rn→RmandB:Rm→Rnwheren>m , prove that BA(ann×nmatrix) is singular. Is ABnecessarily singular? (Proof or counterexample). (10) Given two square matrices AandBsuch thatAB= 0 , which of the following statements are always true. Proofs or counterexamples are called for. [I suggest you confine your search for counterexamples to the case of 2 ×2 matrices.] (a).A= 0 . (b).B= 0. (c).Aand/orBare (is) singular (not invertible). (d).Ais singular. (e).B−1exists. (f). IfA−1exists, then B= 0 . (g). IfBis nonsingular, then A=C. 5.2. SUPPLEMENT ON QUADRATIC FORMS 215 (h).BA= 0 . (i). IfA/negationslash= 0 andB/negationslash= 0 , then neither AnorBare invertible. (11) (a). If Ais a square matrix which satisfies A2−2A−I= 0, findA−1in terms of A. [Hint: Find a matrix Bsuch thatAB=BA=I.] (b). IfAis a square matrix which satisfies An+an−1An−1+an−2An−2+...+a1A+a0I= 0, a 0/negationslash= 0, wherea0,a1,...,a n−1are scalars, prove that Ais invertible and find A−1in terms ofA. (12) (a). If L:En→Em, prove N(L∗) =R(L)⊥ [Hint: Show (in two lines) that X∈N(L∗)⇐⇒ /angbracketleftX, LZ /angbracketright= 0 for all Z∈En—from which the result is immediate.] (b). Use part (a) to show that dim R(L) = dim R(L∗) . (c). Do exercise 19, page 441. (13) (a). If T:En→Enis a translation, TX=X+X0, proveTis invertible by explicitly findingT−1(which is a trivial task). [Answer: T−1X=X−X0.] (b). IfR:En→Enis a rigid body transformation, show that Ris always invertible by exhibiting R−1. [Answer: If RX =R0X+X0, thenRcan be written as Rx= (TR 0)X. R−1=R∗ 0T−1.] (14) IfAis anyn×nmatrix, find matrices A1andA2such thatAis decomposed into the two parts A=A1+A2 whereA1is symmetric and A2isanti-symmetric , i.e.,A∗ 2=−A2. [Hint: Assume there is such a decomposition and use it to find A1andA2in terms of AandA∗. Then verify that these work.] (15) Consider the operator D=d dxonP5. Prove that Dis not invertible (return to the definition p. 360) but exhibit an operator Lwhich is a right inverse, DL=I. (16) This problem proves that Ris orthogonal if and only if Ris linear and isometric. (a) Prove that if Ris linear and isometric, then it is orthogonal. (Trivial!). (b) IfRis orthogonal, prove that i)/bardblRX/bardbl=/bardblX/bardbl ii)/angbracketleftRX, RY /angbracketright=/angbracketleftX, Y/angbracketright(Hint: Use /bardblRX−RY/bardbl2=/bardblX−Y/bardbl2) iii)R(aX) =aRX (Hint: Prove /bardblR(aX)−aRX/bardbl2= 0 ) iv)R(X+Y) =RX+RY(Hint: Prove /bardbl“something” /bardbl2= 0 ) v)Ris linear and isometric [Warning: If you assume linearity in b), you’ll vitiate the whole problem]. 216 CHAPTER 5. MATRIX REPRESENTATION (17) (a). Let Abe a square matrix such that A5= 0 . Verify that ( I+A)−1=I−A+ A2−A3+A4. (b). IfA7= 0 , then ( I−A)−1= ? (18) Consider the matrices (a)./parenleftbiggα1 2 −1 2δ/parenrightbigg , (b)./parenleftBigg1√ 2β γ1√ 2/parenrightBigg , (c)./parenleftbigg0β γ0/parenrightbigg , (d)./parenleftbigg1β 0 2/parenrightbigg . For what value(s) of α, β, γ andδdo these matrices represent orthogonal transfor- mations? (19) IfA= ((aij)) is a square ( n×n) matrix, the trace ofAis defined as the sum of the elements on the main diagonal, tr A: =a11+a22+...+ann. Prove (a). tr(αA) =αA, whereαis a scalar. (b). tr(A+B) = trA+ trB, whereBis also ann×nmatrix. (c). tr(AB) = tr(BA) . (d). tr(I) =? (20) Assume that A:En→Enis anti-symmetric, A∗=−A. (a). Prove A−Iis invertible. [By Theorem 12, it is sufficient to show ( A−I)X= 0⇒X= 0 . Use the property of Ato prove it AX =X, then /angbracketleftX, AX /angbracketright= /bardblX/bardbl2,/angbracketleftA∗X, X/angbracketright=−/bardblX/bardbl2,and/angbracketleftX, AX /angbracketright=/angbracketleftA∗X, X/angbracketright.] (b). IfU= (A+I)(A−I)−1, thenUis an orthogonal transformation. (21) LetAnbe the orthogonal matrix which rotates vectors in E2through an angle of 2π/n. (a). Find a matrix representing An(use the standard basis for E2). (b). LetBdenote the orthogonal matrix of reflection across the x1axis (p. 382). Show that BAb=A−1 nB. [The group of matrices generated by AnandBand all possible products is the dihedral group of order n]. (22) Prove that the set of all orthogonal transformations of EnintoEnforms a (non- commutative) group under multiplication. (23) An affine transformation AX =LX+X0of a linear space into itself is called non-singular if the linear transformation Lis non-singular. Prove that the set of all such non-singular affine transformations form a (non-commutative) group under multiplication. (24) LetA= ((aij)) be a square matrix. Find all such matrices with the property that tr(AA∗) = 0 (see Ex. 19 for the definition of the trace). 5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 217 (25) Consider the linear space S={f(x):f(x) =a+bcosx+csinx}. with the scalar product /angbracketleftf, g/angbracketright=a˜a+1 2(b˜b+c˜c), whereg(x) = ˜a+˜bcosx+ ˜csinx. Define the linear transformation R:S→Sby the rule (Rf)(x) =f(x+α), α real. (a) Show that Ris an orthogonal transformation by proving that /angbracketleftRf, Rg /angbracketright=/angbracketleftf, g/angbracketright for allf,ginS. (b) Choose a basis for Sand exhibit a matrix eRewhich represents Rwith respect to that basis for both the domain and target. (26) LetA:En→Em. Prove:Ais surjective (= onto) if and only if A∗is injective (= one to one). (27) Define A:P3→R3by A[p(x)] = (p(0), p(1), p(−1)) where p∈P3. Find the matrix for this transformation with respect to the basis e1= 1e2= (x+ 1)2, e3= (x−1)2, e4=x3forP3; and the standard basis for R3. (b). Find the matrix representing Ausing the same basis for R3but using the basis ˆe1= 1,ˆe2=x,ˆe3=x2and ˆe4=x3forP3. (28) IfAandBboth map the linear space Vinto itself, and if Bis the only right inverse ofA, AB =I, proveAis invertible. [Hint: Consider BA+B+I]. (29) LetA:En→Embe represented by the matrix (( aij)) , andB:Em→Enby ((bij)) . If /angbracketleftY, AX /angbracketright=/angbracketleftBY, X /angbracketright for allX∈Enand allY∈Em, proveB=A∗. This proves the statement made in the remark following Theorem 14. (30) LetL:R4→R4be defined by LX= (x1,0,x3,0) , whereX= (x1,x2,x3,x4) . Find a matrix representing Lin terms of some basis. You may use the same basis for both the domain and the target. 5.3 Volume, Determinants, and Linear Algebraic Equations. Often we have stated that thus and so is true if and only if a certain set of vectors are linearly independent. But we still have no adequate criteria for determining if a set of vectors is linearly independent. What would be an ideal criterion? One superb criteria would be as follows. Find a function which assigns to a set of nvectorsX1,X2,...,X nin Rna real number, with the property that this number is zero if and only if the vectors are linearly dependent. 218 CHAPTER 5. MATRIX REPRESENTATION There is a geometric way of solving this problem. For clarity we shall work in two dimensions, E2. IfX1andX2are any two vectors in E2, then intuition tells us X1and X2are linearly dependent if and only if the area of the parallelogram (see fig.) is zero. Thus, once we define the analogue of volume for ndimensional parallelepipeds in Rn, the appropriate criterion appears to be that a set of nvectorsX1,...,X ninEmis linearly dependent if and only if the volume of the parallelepiped they span is zero. The major hurdle is constructing a volume function which behaves in the manner dic- tated by two and three dimensional intuition. Our program is to state a few (four to be exact) desirable properties of a volume function Vfor parallelepipeds, then construct a simpler related function - the determinant D, and observe that V=|D|(absolute value ofD) is a volume function. This determinant function will prove useful in the theory of linear algebraic equations. LetX1andX2be any two vectors in R2. We define the parallelogram spanned by X1andX2to be the set of points XinR2which have the form X=t1X1+t2X2, 0≤t1≤1,0≤t2≤1 You can check that these points are precisely those in the parallelogram drawn above. The volume function (really area in this case) V(X1,X2) which assigns to each parallelogram its volume should have the properties 1.V(X1,X2)≥0. 2.V(λX1,X2) =|λ|V(X1,X2), λ scalar. 3.V(X1+X2,X2) =V(X1,X2) =V(X1,X1+X2) . 4.V(e1,e2) = 1. e 1= (1,0), e2= (0,1). The second property states that if one side is multiplied by λ1then the volume is multiplied by |λ|(see fig.). The third property is more subtle. It states that the volume of the parallelogram spanned by X1andX2is the same as the parallelogram spanned by X1andX1+X2. This is clear from the figure since both parallelograms have the same base and height. The last property merely normalizes the volume. It states that the unit square has volume 1. Our first task is to define a parallelepiped in En. Definition: Thendimensional parallelepiped inEnspanned by a linearly independent set of vectors X1,X2,...,X nis the set of all points XinRnof the form X=t1X1+t2X2+···+tnXn, 0≤tj≤1. It is a straightforward matter to write the axioms for the volume V(X1,X2,...,X n) for thendimensional parallelepiped in En. V-1.V(X1,X2,...,X n)≥0. V-2.V(X1,X2,...,X n) is multiplied by |λ|if someXjis replaced by λXjwhereλ is real. V-3.V(X1,X2,...,X n) does not change if some Xjis replaced by Xj+Xk, where j/negationslash=k. V-4.V(e1,e2,...,e n) = 1 , where e1= (1,0,0,..., 0) , etc. These axioms are amazingly simple. It is surprising that the volume function Vin uniquely determined by them; that is, there is only one function which satisfies these axioms. You might wonder why we did not add the reasonable stipulation that volume remains 5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 219 unchanged if the parallelepiped is subjected to a rigid body transformation. The reason is that this axiom would be redundant, for this invariance of volume under rigid body transformation will be one of our theorems. The most simple way to obtain the volume function is to first obtain the determinant functionD(X1,X2,...,X n) . We define thedeterminant functionD(X1,X2,...,X n) ofn vectorsX1,X2,...,X ninRnby the following axioms (selected from those for V). D-1.D(X1,X2,...,X n) is a real number. D-2.D(X1,X2,...,X n) is multiplied by λif someXjis replaced by λXjwhereλ is real. D-3.D(X1,X2,...,X n) does not change if some Xjis replaced by Xj+Xk, where j/negationslash=k. D-4.D(e1,e2,...,e n) = 1 , where e1= (1,0,0,..., 0) etc. Remarks : (1) IfA= ((aij)) is a (square) n×nmatrix, A= a11a12···a1n a21a22··· · · · · · · · an1an2···ann  we can consider it as being composed of ncolumn vectors A1,A2,...,An, and define the determinant of the square matrix Ain terms of the determinant of these vectors detA=D(A1,A2,...,An) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11a12···a1n a21 a2n · · · · · · an1··· ···ann/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle. (2) Although we have written a set of axioms for D, it is not at all obvious that such a function exists. Rest assured that we will prove the existence of such a function. (3) Observe: if we define V(X1,X2,...,X n) :=|D(X1,X2,...,X n)|, X j∈En, thenVdoes satisfy the axioms for volume. Granting existence of D, we derive some algebraic consequences of the axioms. Theorem 5.23 . LetDbe a function which satisfies axiom D-1 to D-3 (not necessarily D-4). (1)IfXjis replaced by Xj=/summationdisplay k/negationslash=jλkXkthenDdoes not change. 220 CHAPTER 5. MATRIX REPRESENTATION (2)If one of the vectors Xjis zero, then D= 0. (3)If the vectors X1,X2,...,X nare linearly dependent then D= 0. In particular D= 0 if two vectors are equal. (4)Dis a linear function of each of its variables, that is D(...,λY +µZ,... ) =λD(...,Y,... ) +µD(...,Z,... ) (soDis a multilinear function). (5)If any two vectors XiandXjare interchanged, then Dis multiplied by −1. D(...,X i,...,X j,...) =−D(...,X j,...,X i,...) Proof: These proofs, like the statements above, are conceptually simple but notationally awkward. Notice that only Axioms 1-3 but not Axiom 4 will be used. We shall need this fact shortly. (1) We prove this only if Xjis replaced by Xj+λXk, j/negationslash=kandλ/negationslash= 0 . The general case is a simple repetition of this until the other Xk’s are used up. It is simplest to work backward. By Axiom 2, D(...,X j+λXk,...,X k...) =1 λD(...,X j+λXk,...,λX k,...) so by axiom 3 (since λXkis now a vector in D) =1 λD(...,X j,...,λX k,...) and axiom 2 again =D(...,X j,...,X k,...). (2) Write the vector Xj= 0 as 0Xjwhere 0 is now a scalar. This scalar may be brought outsideDby axiom 2. Since Dis a real number, 0 ·D= 0 . (3) LetXj=/summationdisplay k/negationslash=jakXk. By part 1, Ddoes not change if Xjis replaced by Xj+/summationdisplay k/negationslash=jλkXk. Chooseλk=−ak. This gives a Dwith one vector zero, Xj−/summationdisplay k/negationslash=jakXk= 0 . Thus Dis zero by part 2. (4) The trickiest part. Axiom 2 immediately reduced this to the special case λ=µ= 1 . For notational convenience, let Y+Zby in the last slot. We have to prove D(X1,X2,...,Y +Z) =D(X1,X2,...,Y ) +D(X1,X2,...,Z ). IfX1,X2,...,X n−1(which appear in all three terms above) are linearly dependent, we are done by part 3. Thus assume they are linearly independent. Since our linear space Rnhas dimension n, thesen−1 vectors can be extended to a basis for Rn 5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 221 by adding one more, ˜Xn. Now we can write YandZas a linear combination of these basis vectors Y=a1X1+···+an−1Xn−1+an˜Xn, Z =b1X1+···+bn−1Xn−1+bn˜Xn. Substituting this into Dwe obtain D(X1,...,Y +Z) =D(X1,...,...,n−1/summationdisplay 1(aj+bj)Xj+ (an+bn)˜Xn). But by part 1, =D(X1,..., (an+bn)˜Xn) and axiom 1 results in = (an+bn)D(X1,..., ˜Xn). However, again by part 1, D(X1,...,Y ) =D(X1,...,n−1/summationdisplay 1ajXj+an˜Xn) =D(X1,...,...,a n˜Xn) =anD(X1,..., ˜Xn). Similarly D(X1,...,Z ) =bnD(X1,..., ˜Xn). Adding these two expressions and comparing them with the above, we obtain the result. (5) To avoid a mess, indicate only the ith andjth vectors. Our task is to prove D(...,X i,...,X j,...) =−D(...,X j,...,X i,...). This is clever. Watch: By the multilinearity (part 4) D(...,X i+Xj,...,X i+Xj,...) =D(...,X i,...,X i,..., ) +···+D(...,X i,...,X j,...) +D(...,X j,...,X i,...) +···+D(...,X j,...,X j,...). However part 2 states that the left side as well as the first and last terms on the right are zero. Thus 0 =D(...,X i,...,X j,...) +D(...,X j,...,X i,...). Transposition of one of the terms to the other side of the equality sign completes the proof. You should also be able to fashion an easy proof of this part which uses only the axioms directly (and uses none of the other parts of this theorem). Instead of moving on immediately, it is instructive to compute D[X1,X2] whereX1 andX2are vectors in R2, X 1= (a,b), X 2= (c,d) . Then we are computing D/bracketleftbigg/parenleftbigga b/parenrightbigg ,/parenleftbiggc d/parenrightbigg/bracketrightbigg , 222 CHAPTER 5. MATRIX REPRESENTATION which is, equivalently, the determinant of the matrix/parenleftbigga c b d/parenrightbigg . D/bracketleftbigg/parenleftbigga b/parenrightbigg ,/parenleftbiggc d/parenrightbigg/bracketrightbigg =aD/bracketleftbigg/parenleftbigg1 b a/parenrightbigg ,/parenleftbiggc d/parenrightbigg/bracketrightbigg (axiom 2) =aD/bracketleftbigg/parenleftbigg1 b a/parenrightbigg ,/parenleftbiggc d/parenrightbigg −c/parenleftbigg1 b a/parenrightbigg/bracketrightbigg (Theorem 21 part 1) =aD/bracketleftbigg/parenleftbigg1 b a/parenrightbigg ,/parenleftbigg0 ad−cb a/parenrightbigg/bracketrightbigg (algebra) = (ad−bc)D/bracketleftbigg/parenleftbigg1 b a/parenrightbigg ,/parenleftbigg0 1/parenrightbigg/bracketrightbigg (axiom 2) = (ad−bc)D/bracketleftbigg/parenleftbigg1 b a/parenrightbigg −b a/parenleftbigg0 1/parenrightbigg ,/parenleftbigg0 1/parenrightbigg/bracketrightbigg (Theorem 21 part 1) = (ad−bc)D/bracketleftbigg/parenleftbigg1 0/parenrightbigg ,/parenleftbigg0 1/parenrightbigg/bracketrightbigg (algebra) = (ad−bc)D[e1,e2] =ad−bc (axiom 4). Thus |Area|=|(a+c)(b+d)−2bc−cd−ab|=|ad−bc| You can indulge in a bit of analytic geometry (or look at my figure) to show that the area of a parallelogram spanned by X1andX2is|ad−bc|. From our explicit calculation, the existence and uniqueness of the determinant of two vectors in R2has been proved. There are several ways to prove the general existence and uniqueness of a determi- nant function. Our procedure is to first prove there is at most one determinant function (uniqueness). Then we shall define a function inductively, and verify it satisfies the axioms. By uniqueness, it must be the only function. Two interesting and important preliminary propositions are needed. The following lemma shows how to evaluate the determinant if all of the elements above the principal diagonal are zeroes (that is, the determinant of a lower triangular matrix). LEMMA : LetX1,···,Xnbe the columns of a lower triangular matrix  a11 0 0... 0 a21a22 0 · · · 0 · · · · · · an1an2··· ···ann  Then D(X1,···,Xn) =a11a22···annD(e1,···,en) =a11a22···ann, that is, the determinant of a triangular matrix is the product of the diagonal elements. Proof: If any one of the principal diagonal elements are zero, then the determinant is zero. For example, if ajj= 0 , then the n−j+ 1 vectors Xj,···,Xnall have their firstjcomponents zero, and hence can span at most an n−jdimensional space. Since n−j+ 1> n−j, these vectors must be linearly dependent. Therefore, by Theorem 21, part 3, the determinant is zero, as the theorem asserts. [If you didn’t follow this, look at a 3×3 or 4 ×4 lower triangular matrix and think for a moment]. 5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 223 If none of the diagonal elements are zero, we can carry out the following simple recipe. The recipe gives a procedure for reducing the problem to evaluating a matrix which is zero everywhere except along the diagonal. First, we get all zeros to the left of a22in the second row by multiplying the second column,X2, by−a21/a22and adding the resulting vector to X1. This gives a new first column with i= 2, j = 1 element zero. Moreover, the new matrix has the same determinant as the old one (Theorem 21, part 1). It looks like  a11 0 0 0 0a22 0· ˜a31a32a33· · · · 0 · · · · · · · · ˜an1an2···ann . Only the first column has changed. Repeat the same process to get all zeros to the left ofa33. Thus, multiply the third column by −˜a31/a33and−a31/a33and add the result to the first and second columns respectively. This gives a new matrix, again with equal determinant, but which looks like  a11 0 0 0 ··· 0 0a22 0 0 0 0 0a33 0 · ˆa41ˆa42a43a44 · · · 0 · · · · · · ˆan1ˆan2an3· ·ann . Moving on, we gradually eliminate all of the terms to the left of the diagonal but keep thesame diagonal ones. The final result is  a11 0···0 0a22 · · · 0 · · · · · · 0 0 0 ann . It has the same determinant as the original matrix, so D(X1,...,X n) =D(a11e1,...,a nnen) =a11···annD(e1,...,e n), where Axiom 2 has been used to pull out the constants. Now Axiom 4, D(e1,...,e n) = 1 , can be used to complete the proof. Observe that Axiom 4 is not used until the very last step. Thus, the formula D= (something) D(e1,...,e n) depends only on Axioms 1-3. We shall need this soon. The above theorem shows how easy it is to evaluate the determinant of a lower trian- gular matrix. It becomes particularly valuable when coupled with the next theorem which 224 CHAPTER 5. MATRIX REPRESENTATION shows how the determinant of an arbitrary matrix can be reduced to that of a lower tri- angular matrix. The reduction procedure given here is the best practical way of evaluating a determinant . There is a peculiar criss-cross method for evaluating 3 ×3 determinants which is taught in many high schools. Forget it. The method is not very practical and does not generalize to 4 ×4 or larger determinants. Theorem 5.24 . The evaluation of the determinant D(X1,...,X n)can be reduced to the evaluation of a lower triangular matrix - and hence has the form D= (something )D(e1,...,e n). The proof gives a way of computing “something” in terms of the original matrix. Remark : In the above formula, we did not utilize the fact that D(e1,···,en) = 1 since this one step in the proof is the only place where Axiom 4 would be used, so we can (and shall) use the fact that this result holds for any function which only satisfies Axioms 1-3. Proof: This is just a recipe for carrying out the reduction. It essentially is a repetition of the last part of the preceding lemma. Instead of waving our hands at the procedure, we shall work out a representative Example: Evaluate D=D(X1,X2,X3,X4) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 2 −1 0 −1−2 3 1 0−1 4 −3 2 5 0 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle by reducing it to a lower triangular determinant. First we get all zeros to the right of the diagonal in the first row, that is, except in the a11slot, by multiplying X1by the constants −2,1 and 0 and adding the resulting vectors toX2,X3, andX4, respectively. We obtain D=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 2 −1 0 −1−2 3 1 0−1 4 −3 2 5 0 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 −1 0 −1 0 3 1 0−1 4 −3 2 1 0 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 0 0 −1 0 2 1 0−1 4 −3 2 1 2 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle Now we get all zeros to the right of the diagonal in the second row. Since the new a22ele- ment above is zero, interchange the second and third columns (one could have interchanged the second and fourth). This introduces a factor of −1 (by Theorem 21, part 5). Then multiply the new second column by the constants 0 and −1 2, respectively, and add to the last two columns, respectively. This gives D=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 0 0 −1 0 2 1 0−1 4 −3 2 1 2 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 0 0 −1 2 0 1 0 4 −1−3 2 2 1 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 0 0 −1 2 0 0 0 4 −1−5 2 2 1 0/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle 5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 225 And on the third row, where we again want all zeros to the right of the diagonal, so multiply the new third column by −5 and add it to the fourth column: D=−/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 0 0 −1 2 0 0 0 4 −1−5 2 2 1 0/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=−/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 0 0 −1 2 0 0 0 4 −1 0 2 2 1 −5/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=−(1)(2)( −1)(−5) =−10, where we have used the lemma about determinants of lower triangular matrices to evaluate the last determinant. Uniqueness is now elementary. Theorem 5.25 . There is at most one function D(X1,···,Xn), X k∈Rn, which satisfies the 4 axioms for a determinant function. Proof: Assume there are two such functions, D(X1,···,Xn) and ˜D(X1,···,Xn). Let ∆(X1,···,Xn) =D(X1,···,Xn)−˜D(X1,···,Xn). We shall show ∆( X1,···,Xn) = 0 for any choice of X1,···,Xn. Since both Dand ˜D satisfy Axioms 1-4, we have 1). ∆ =D−˜Dis real valued. 2). ∆(...,λX j,...) =D(...,λX j,...)−˜D(...,λX j,...) =λD(...,X j,...)−λ˜D(...,X j,...) =λ∆(...,X j,...). 3). ∆(...,X j+Xk,...) =D(...,X j+Xk,...)−˜D(...,X j+Xk,...) =D(...,X j,...−˜D(...,X j,...) = ∆(...,X j,...), j/negationslash=k. 4). ∆(e1,...,e n) =D(e1,...,e n)−˜D(e1,...,e n) = 1−1 = 0 . Thus, ∆ satisfies the same first three axioms but ∆( e1,...,e n) = 0 in place of Axiom 4. Because the proof of Theorem 22 and its predecessors never used Axiom 4, we know that ∆(X1,...,X n) = (something) ∆( e1,...,e n) = 0. Thus ∆(X1,···,Xn) = 0 for any vectors Xj. If it exists, the determinant function is known to be unique. We intend to define the determinant of order n, that is, of nvectors in Rn, in terms of determinants of order n−1 . The key to such an approach is a relationship between a determinant of order nand 226 CHAPTER 5. MATRIX REPRESENTATION determinants of order n−1 . To motivate our definition, we first examine the case n= 3 and utilize the intimate relation between determinant and volume. LetX1,X2andX3be three vectors in R3. To find the determinant D(X1,X2,X3) , we can resolve one of the vectors, say X1, into its components X1=a11e1+a21e2+a31e3. Since the determinant function is linear (Theorem 21, part 4), D=D(X1,X2,X3) =a11D(e1,X2,X3) +a21D(e2,X2,X3) +a31D(e3,X2,X3). How can we interpret D(a11e1,X2,X3) , D(a11e1,X2,X3) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11a12a13 0a22a23 0a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle? By subtracting suitable multiples of the first column from the other two, we have D(a11e1,X2,X3) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11 0 0 0a22a23 0a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle. Consider the related volume function. The vectors in the last matrix span a parallelepiped whose base is the parallelogram spanned by (0 ,a22,a32) and (0,a23,a33) , while the height isa11. Thus, we expect the volume to be a11times the area of the base. Since the area of the base is/vextendsingle/vextendsingle/vextendsingle/vextendsingledet/vextendsingle/vextendsingle/vextendsingle/vextendsinglea22a23 a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle, we hope /vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11 0 0 0a22a23 0a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=a11/vextendsingle/vextendsingle/vextendsingle/vextendsinglea22a23 a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle. except possibly for a factor of ±1 . This last formula is the connection between determinants of order three and those of order two. Notice that the determinant on the right in the last equation is obtained from that of D=D(X1,X2,X3) by deleting both the first row and first column. It is called the 1,1 minor ofD, and written D11. More generally, the i,jminorDijofDis the determinant obtained by deleting the ith row and jth column of D. IfDis of order n, then each Dijis of order n−1 . In this notation, we expect from the expansion of D(X1,X2,X3) that D(X1,X2,X3) =±?a11D11±?a21D21±?31D31, or/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11a21a13 a21a22a23 a31a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=±?a11/vextendsingle/vextendsingle/vextendsingle/vextendsinglea22a23 a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle±?a21/vextendsingle/vextendsingle/vextendsingle/vextendsinglea12a13 a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle±?a31/vextendsingle/vextendsingle/vextendsingle/vextendsinglea12a13 a22a23/vextendsingle/vextendsingle/vextendsingle/vextendsingle. where ? indicates our doubt as to the signs. Explicit evaluation of both sides (using Theorem 22) reveals that the correct sign pattern is + ,−,+ . Having examined this special case (and the 4 ×4 case too), we are tentatively led to Suspicion (Expansion by Minors). If D(X1,...,X n) is a determinant function, that is, if it satisfies the axioms, then D(X1,X2,...,X n) =n/summationdisplay i=1(−1)i+jaijDij, (5-3) 5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 227 whereXj= (a1j,a2j,...,a nj) . For the case n= 3,j= 1 this is the formula we found above. To verify that the formula is correct, we must verify that the function satisfies our axioms for a determinant. The reasoning goes as follows: we know exactly what determinants of order two are by a previous computation, so the formula gives a candidate for the determinant of order three, which in turn gives a candidate for a determinant function of order four, and so on. Thus, by induction, let us assume that determinants of order k−1 are known. We must prove Theorem 5.26 . The previous function D(X1,...,X k)defined by the above formula is a determinant function, that is, it satisfies the axioms. Proof: 1).D(X1,...,X k) is real valued since, by our induction hypothesis, each of the Dij, determinants of order k−1 , is real valued. 2).D(...,λX l,...) =λD(...,X l,...) . There are two cases. Ifl=j, thenλXj means that a1j,a2j,..., is multiplied by λ. Thus D(...,λX j,...) =n/summationdisplay i=1(−1)i+jλaijDij=λD(...,X j,...), so the axiom is satisfied. Ifl/negationslash=j, then some vector Xother than Xjis multiplied by λ, so D(...,λX l,...,X j,...) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle···λa1l···a1j··· ···λa2la2j··· · · · · · · ···λaklakj···/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle SinceDijis formed by deleting the ith row and jth column of D, andl/negationslash=j, one column in minor Dij will have the factor λappearing in it. By the induction hypothesis, the factor can be pulled out of each one, and hence from any linear combination of them. Because the expansion formula for Dis a linear combination of the minors, the axiom is verified in this case too. 3). Omitted. This one is just plain messy. If you don’t care to try the general case for yourself, at least try the case n= 3 and verify it there. 4). To prove D(e1,...,e n) = 1 . Of the coefficients a1j,a2j,...,a nj, onlyajj/negationslash= 0 , andajj= 1 . Thus D(e1,...,e n) = (−1)j+jajjDjj=Djj. But by the induction hypoth- esis,Djj= 1 since it has only ones on its main diagonal and zero elsewhere. Therefore D(e1,...,e n) = 1 , as desired. This theorem completes (except for one segment) the proof that a unique determinant function exists. The uniqueness was proved directly, while the existence was obtained from the known existence of 2 ×2 determinant functions (the simpler case of 1 ×1 determi- nants could also have been used) and proving inductively that a candidate for the n×n determinant function does satisfy the axioms. Emerging from the jungle of the existence proof, we are fully equipped with the powerful determinant function and the associated volume function. It will be relatively simple to prove the remaining theorems involving determinants. The trick in most of them is to make clever use of the fact that the determinant function is unique. We shall expose this trick in its bare form. 228 CHAPTER 5. MATRIX REPRESENTATION Theorem 5.27 . Let ∆(X1,...,X n)be a function of nvectors in Rnwhich satisfies axioms 1-3 for the determinant. Then for every set of vectors X1,...,X n ∆(X1,...,X n) = ∆(e1,...,e n)D(X1,...,X n). Thus, the function ∆differs from Donly by a constant multiplicative factor, which is the number ∆assigns to the unit matrix (geometrically, the unit cube) in Rn. Proof: If ∆(e1,...,e n) = 1 , then ∆ satisfies Axiom 4 also, so by the uniqueness theorem, it must be Ditself. If ∆( e1,...,e n)/negationslash= 1 , consider ˜D(X1,...,X n) :=D(X1,...,X n) = ∆(X1,...,X n) 1−∆(e1,...,e n). Note that the denominator is a fixed scalar which does not depend on X1,...,X n. It is a mental calculation to verify that ˜Dsatisfies all of Axioms 1-4. Therefore ˜D(X1,...,X n) := D(X1,...,X n) by uniqueness. Solving the last equation for ∆( X1,...,X n) yields the formula. ConsiderD(X1,...,X n) . IfB= ((bij)) is a square n×nmatrix representing a linear transformation from RntoRn, how are D(X1,...,X n) andD(BX 1,BX 2,...,BX n) related? The answer to this question is vital if we are to find how volume varies under a linear transformation B. IfA= ((aij)) is the matrix whose columns are X1,...,X n, and C= ((cij)) is the matrix whose columns are BX 1,BX 2,...,BX n, thenC=BA[since, for example, c11—the first element in the vector BX 1—is c11=b11a11+b12a21+b13a31+···+b1nan1.] BecauseD(X1,...,X n) = detAandD(BX 1,...,BX n) = detC, our question becomes one of relating det C= det(BA) to detA. The result is as simple as one could possibly expect. Theorem 5.28 . IfAandBare twon×nmatrices, then det(BA) = (detB)(detA) = (detA)(detB) = det(AB) or, ifX1,...,X nare the column vectors of A, then this is equivalent to D(BX 1,...,BX n) =D(Be1,Be 2,...,Be n)D(X1,...,X n) (since the matrix whose columns are Be1,...,Be nis justB). Proof: Let ∆(X1,...,X n) :=D(BX 1,...,BX n) . This function clearly satisfies Axiom 1. We shall verify Axioms 2 and 3 at the same time. ∆(...,λX j+µXk,...) =D(...,B (λXn+µXk),...) BecauseBis a linear transformation, we have =D(...,λBX j+µBX k,...). By the linearity of D(Theorem 21, part 4) =λD(...,BX j,...) +µD(...,BX k,...). 5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 229 Ifj/negationslash=k, then the vector BXkin the second term on the right also appears as another column in the same determinant. Hence the second term vanishes. Thus if j/negationslash=k, ∆(...,λX j+µXk,...) =λD(...,BX j,...). The special case µ= 0 shows Axiom 2 holds for ∆ , while the case λ=µ= 1 verifies Axiom 3. Therefore ∆ satisfies Axioms 1-3. Applying the preceding Theorem (25), we have ∆(X1,...,X n) = ∆(e1,...,e n)D(X1,...,X n). By definition, ∆( e1,...,e n) :=D(Be1,...,Be n) . Substitution verifies our formula. The commutativity (detB)(detA) = (detA)(detB) follows from the fact that det Aand detBare real numbers - which do commute under multiplications. Corollary 5.29 . IfAis an invertible matrix, then det(A−1) =1 detA. Proof: SinceAA−1=I, and detI= 1 , we find (detA)(detA−1) = det(AA−1) = detI= 1. Ordinary division completes the proof. Our next theorem is also a corollary, but because of its importance, we call it Theorem 5.30 . The vectors X1,...,X ninRnare linearly independent if and only if D(X1,...,X n)/negationslash= 0. Proof: ⇐IfD(X1,...,X n)/negationslash= 0 , then the vectors X1,...,X nare linearly independent, since if they were dependent, then D= 0 by part 3 of Theorem 21. ⇒. IfX1,...,X nare linearly independent vectors in Rn, then the Corollary to Theorem 12 (p. 364) shows that the matrix Awhose columns are the Xjis invertible. LetA−1be its inverse. From the computation in the corollary preceding this theorem, (detA)(detA−1) = 1. Thus the real number det Acannot be zero. The equivalent form of our theorem is also a consequence of the Corollary to Theorem 12. Example: (cf. p. 157, Ex. 1b). Are the vectors X1= (0,1,1), X 2= (0,0,−1), X 3= (0,2,3) linearly dependent? We compute the determinant D(X1,X2,X3) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle0 0 0 1 0 2 1−1 3/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle. 230 CHAPTER 5. MATRIX REPRESENTATION If we knew that “the determinant of a matrix was equal to the determinant of its adjoint” (a true theorem to be proved below), then taking the adjoint we get a matrix with one column zero 0 which gives D= 0 . Since the quoted theorem is not yet proved, we proceed differently and reduce our 3 ×3 determinant to 2 ×2 determinants expanding by minors (p. 411). The simplest column to use is the second. /vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle0 0 0 1 0 2 1−1 3/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle= (−1)1+2 0/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 2 1 3/vextendsingle/vextendsingle/vextendsingle/vextendsingle+ (−1)2+2 0/vextendsingle/vextendsingle/vextendsingle/vextendsingle0 0 1 3/vextendsingle/vextendsingle/vextendsingle/vextendsingle+ (−1)3+2(−1)/vextendsingle/vextendsingle/vextendsingle/vextendsingle0 0 1 2/vextendsingle/vextendsingle/vextendsingle/vextendsingle =/vextendsingle/vextendsingle/vextendsingle/vextendsingle0 0 1 2/vextendsingle/vextendsingle/vextendsingle/vextendsingle= 0·2−1·0 = 0 by the explicit formula for evaluating 2 ×2 determinants. Thus D= 0 so the vectors X1,X2,X3are linearly dependent. That nice theorem we could have used in the above example is our next target. Theorem 5.31 . IfAis ann×nmatrix, then detA∗= detA. Proof: LetA1,...,Anbe the columns of AandB,...,Bnits rows, A= a11a12···a1n a21··· ···a2n · · · · · · an1··· ···ann B1 }B2 · · · }Bn. Consider the function D(B1,...,Bn) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11···an1 a12 · · · · · a1nann/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle}A1 · · · }An= detA∗. since the rows of Aare the columns of A∗. Let us define a new function ˆD(A1,...,An) :=D(B1,...,Bn). Our task is to verify that ˆD(A1,...,An) satisfies all of Axioms 1-4. Then by uniqueness detA∗:=ˆD(A1,...,An) =D(A1,...,An) = detA. (1) ˆD(A1,...,An) is a real number since det A∗, the determinant of the matrix A∗is a real number. 5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 231 (2) We must show ˆD(...,λAj,...) =λˆD(...,Aj,...) , that is, /vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11a2j···an1 · · λa1jλa2j... λa nj · · · a1n... ... a nn/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=λ/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11···anl · · a1j···anj a1n···ann/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle (a fact we only know so far if a column is multiplied by a scalar). Trick: observe that jth row... 1 0 0 0 0 1 0 0···λ1 0 0 0...0 1  a11···anl · · · · a1j···anj · · · · a1n···ann  = a11···anl · · · · a1j···anj · · · · a1n···ann  The matrix on the left is the identity matrix Iexcept for a λin itsjth row and jth column. Its determinant is λ(since you can factor λfrom thejth column and are left with the identity matrix). By Theorem 26, the determinant of the product on the left is λˆD(A1,...,An) while the right is ˆD(A1,...,An) , proving ˆDsatisfies Axiom 2. (3) The proof of Axiom 3 involves a similar trick. We have to show ˆD(...,Aj+Ak,...) = ˆD(...,Aj,...) wherej/negationslash=k, that is, to show /vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11 ···an1 · · · · a1j+a1k···anj+ank · · · · · · a1n ···ann/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11···an1 · · a1j···anj · · · · · · a1n···ann/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle,j/negationslash=k. 232 CHAPTER 5. MATRIX REPRESENTATION Observe that  1 0 ··· 0 0 1 ··· ··· ··· ··· 0 0 ··· 0 0 0 1  a11···an1 · · a1j···anj · · · a1n···ann = a11 ···an1 · · · · a1j+a1k···anj+ank · · · a1n ···ann , where the matrix on the left is the identity matrix with an extra 1 in the jth row,kth column. Since the determinant of this matrix is one (check by a mental computation), the rule for the determinant of a product of matrices shows that Axiom 3 is satisfied. (4) Easy, for ˆD(e1,...,e n) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 ··· ··· 0 0 1 ··· ··· 0 ··· ··· 1 0 0··· ··· 0 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=D(e1,...,e n) = 1. This verification of the four Axioms coupled with the remarks at the beginning of the proof completes the proof. Corollary 5.32 . The column operations of Theorem 21 are also valid as row operations. Proof: Every row operation on a matrix A(like adding two rows) can be split up to : i) takeA∗so the rows become columns, ii) carry out the operation on the column of A∗and iii) take the adjoint again. Since the determinant does not change under these operations, we are done. Corollary 5.33 . IfRis an orthogonal matrix then detR=±1. Proof: IfRis orthogonal, then R∗R=Iby Theorem 19 (p. 383). Thus, a= detI= det(R∗R) = (detR∗)(detR) = (detR)2, where Theorems 25 and 27 were invoked once each. Now take the square root of both sides. The orthogonal matrices R1=/parenleftbigg1 0 0 1/parenrightbigg andR2=/parenleftbigg0 1 1 0/parenrightbigg , for which det R1= 1 and det R2=−1 show that both signs are possible. IfdetR=−1 , then the orthogonal transformation has not only been a rotation but also a reflection . The transformation given by R2is a figure goes here 5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 233 which can be thought of as the composition (product) of a rotation by +900followed by a reflection (mirror image). In fact, R2may be factored into ˆR˜R=r2, where /parenleftbigg1 0 0−1/parenrightbigg/parenleftbigg0 1 −1 0/parenrightbigg =/parenleftbigg0 1 1 0/parenrightbigg =R2. Pictorially a figure goes here Our theorems about determinants also imply the following valuable result about volume. Theorem 5.34 . LetX1,...,X nspan a parallelepiped QinEnand the matrix Amap EnintoEn. Then the volume is magnified by |detA|, that is, V[AX 1,...,AX n] =|detA|V[X1,...,X n]. If we denote the image of QbyA(Q), then this theorem reads Vol[A(Q)] =|detA|Vol[Q]. Proof: [We should first prove that there is at most one volume function Vsatisfying its four axioms. Since V:=|D|is a volume function, assume there is another volume function V∗and define ˜D(X1,...,X n) by ˜D(X1,...,X n) :=/braceleftBigg V∗(X1,...,X n)D(X1,...,X n) |D(X1,...,X n)|ifD/negationslash= 0 0 if D= 0. It is simple to check that ˜Dsatisfies the axioms for a determinant. By uniqueness, ˜D=D. Solving the last equation, we find V∗(X1,...,X n) =|D(X,...,X n)| ≡V(X1,...,X n) , so the volume function is also unique.] The theorem is easily proved. Since V=|D|, an application of Theorem 26 tells us that V[AX 1,...,AX n] =|D(AX 1,...,AX n)| =|D(Ae1,...,Ae n)| |D(X1,...,X n)| =|detA|V[X1,...,X n]. Done. Corollary 5.35 . Volume is invariant under an orthogonal transformation. V(RQ) =V(Q) Proof: IfRis an orthogonal transformation, |detR|= 1 . Remark 1 . Since we eventually want to define the volume of suitable sets by approximating the sets by parallelepipeds, this theorem will allow us to conclude the same results about how the volume of some set changes under a linear transformation in general and an orthogonal transformation in particular. 234 CHAPTER 5. MATRIX REPRESENTATION Remark: 2 We define the determinant of a linear transformation Lwhich maps Rninto Rnas the determinant of a matrix which represents L. This definition makes it mandatory to prove: “the determinant of two different matrices which represent L(different because of a different choice of bases) are equal.” However the theorem is an immediate consequence of the following fact we never proved: “if AandBare matrices which represent the same linear transformation Lwith respect to different bases then there is a nonsingular matrix Csuch thatB=CAC−1.” The matrix Cis the matrix expressing one set of bases vectors in terms of the other bases. Using this theorem, we find detB= det(CAC−1) = (detC)(detA)(detC−1) = detA. How does volume change under a translation T, TX =X+X0? A little thought is needed. Imagine a parallelepiped Qspanned by X1,...,X n. The crux of the matter is to realize that the parallelepiped has the origin as one of its vertices and X1,...,X nat the others. Under the translation T, not only do the Xj’s get translated through X0, but so does the origin , 0→X0, X 1→X1+X0, X 2→X2+X0, etc. a figure goes here In terms of free vectors, the edge from 0 to Xjbecomes the edge from X0toXj+X0 (see figure). Thus the free vector representing this edge is ( Xj+X0)−X0, that is, it is stillXj! This motivates the Definition: The volume of a parallelepiped is defined to be the volume of the parallelepiped after translating one vertex to the origin. Theorem 5.36 . The change in volume of a parallelepiped Qunder an affine transforma- tionAX=LX+X0, Llinear, is given by: Vol[A(Q)] =|detL|Vol[Q]. In particular, volume is invariant under a rigid body transformation (for then Lis an orthogonal transformation). Proof: The affine transformation may be factored into A=TL, a linear transformation followed by a translation (p. 380). Since Lchanges volume by |detL|while translation preserves the volume, the net result is a change by |detL|as claimed. a) Application to Linear Equations What have our geometrically motivated determinants in common with the determinants of high school fame - where they were used to solve systems of linear algebraic equations? Everything, for they are the same. Since determinants are defined only for square matrices, they are applicable to linear algebraic equations only when there are the same number of equations as unknowns. At the end of this section, we shall make some remarks about the case when the number of equations and unknowns are not equal. Consider the system of equations a11x1+···+a1nxn=y1 a21x1+···+a2nxn=y2 ...... an1x1+···+annxn=yn, 5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 235 which we can write as x1A1+···+xnAn=Y, where Ajis thejth column of the matrix A= ((aij)) andYis the obvious column vector. The problem is to find numbers x1,...,x nsuch thatx1A1+···+xnAn=Y, whereYis given. Theorem 5.37 . LetA= ((aij))be a square n×nmatrix and Ya given vector. The system of linear algebraic equations AX=Ycan always be solved for Xif and only if detA/negationslash= 0. This can be rephrased as, Ais invertible if and only if detA/negationslash= 0. Proof: LetAjbe thejth vector of A. Each Ajis a vector in Rn. If detA/negationslash= 0 , then the An’s are linearly independent by Theorem 27, p. 417. But since they are linearly independent and there are nof them, A1,···,An, they must span Rn. Thus, any Y∈Rn can be written as a linear combination of the Aj’s. The numbers x1,···,xnare just the coefficients in this linear combination. Conversely, if the equations AX=Ycan be solved for anyY∈Rn, then the vec- torsA1,···,Anspan Rn. But ifnvectors span Rn, these vectors must be linearly independent, so det A/negationslash= 0 , again by Theorem 27, page 417. Theorem 5.38 . LetAbe a square matrix. The system of homogeneous equations AX= 0has a non-trivial solution if and only if detA= 0. Proof: By Theorem 27, Page 417, det A= 0 if and only if the column vectors A1,...,An are linearly dependent. Now if the column vectors A1,...,Anare linearly dependent, then there are numbers x1,...,x n, not all zero, such that x1A1+...+xnAn= 0 . The vectorX= (x1,...,x n) is then a non-trivial solution of AX= 0 . Conversely, if there is a non-trivial solution of AX= 0 , then x1A1+···+xnAn= 0 , so the Aj’s are linearly dependent. Hence det A= 0 . In contrast to the above theorems which give no hint of a procedure for finding the desired vector X, the next theorem gives an explicit formula for the solution of AX=Y. Theorem 5.39 (Cramer’s Rule). Let A= ((aij))be a square n×nmatrix with columns A1,...,An. Assume detA/negationslash= 0. Then for any vector Y, the solution of AX=Yis x1=D(Y,A2,...,An) D(A1,...,An), x 2=D(A1,Y,A3,...,An) D(A1,...,An) ... xn=D(A1,...,An−1,Y) D(A1,...,An). For example, in detail, the formula for x2is x2=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11y1a13···a1n ............ an1ynan3···ann/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle /vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11a12a13···a1n ............... an1an2an3···ann/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle. 236 CHAPTER 5. MATRIX REPRESENTATION Proof: A snap. Since det A/negationslash= 0 , by Theorem 31 we know a solution X= (x1,...,x n) exists. Thus x1A1+···+xnAn=Y. Let us obtain the formula for x2as a representative case. Observe that D(A1,Y,A3,···,An) =D(A1,x1A1+···+xnAn,A3,···,An). SinceDis multilinear, we can expand the above to =x1D(A1,A1,A3,···,An) +xnD(A1,A2,A3,···,An) +···+xnD(A1,An,A3···An). Now all of these determinants, except the second one, vanishes since each has two identical columns (part 5 of Theorem 21, page 400). Thus D(A1,Y,A3,·,An) =x2D(A1,A2,···,An). Because det A=D(A1,···,An)/negationslash= 0 , we can divide to find the desired formula for x2. Done. Remark: This elegant formula is mainly of theoretical use. It is not the most efficient procedure for solving such equations. That honor belongs to the method of reducing to triangular form which was outlined in the proof of Theorem 22. To be more vivid, if Cramer’s rule were used to solve a system of 26 equations, approximately (23 + 1)! ≈1028 multiplications would be required. Reduction to triangular form, on the other hand, would only require about (1 /3)(23)3≈6000 multiplications. Think about that. For non-square matrices, determinants are not applicable. Given a vector Y, one would still like a criterion to determine if one can solve AX=Y, that is, one would like a criterion to see ifY∈R(A) . Theorem 5.40 . LetL:V1→V2be a linear operator. Then R(L)⊥=N(L∗); or equivalently (for finite dimensional spaces) R(L) =N(L∗)⊥. Proof: IfX∈V, andY∈R(L)⊥, then for all X 0 =/angbracketleftY, LX /angbracketright=/angbracketleftL∗Y, X/angbracketright. This means L∗Yis orthogonal to all X, consequently, L∗Y= 0 , soY∈N(L∗) . The converse is proved by observing that our steps are reversible. Application . For what vectors Y= (y1,y2,y3) can you solve the equations 2x1,+3x2=y1 x1−x2=y2 x1+ 2x2=y3 ? 5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 237 If the equations are written as AX=Y, then by the above theorem Y∈R(A) if and only ifY⊥N(A∗) . Let us find a basis for N(A∗) . This means solving the homogeneous equationsA∗Z= 0 , 2z1+z2+z3= 0 3z1−z2+ 2z3= 0. If we letz1=α, and solve the resulting equations for z2andz3, we find that z3= −5α/3 andz2=−11α/3 . Consequently, all vectors Z∈(A∗) have the form Z= (3α,−11α,−5α) . A basis for N(A∗) ise= (3,−11,−5) . Therefore, Y⊥N(A∗) if and only if 3y1−11y2−5y3= 0 . By the above reasoning, the equation AX=Ycan be solved for only these Y’s. Remark: The use of Theorem 34 as a criterion for finding if Y∈R(L) is much more valuable in infinite dimensional spaces, for it quite often turns out that N(L∗) is still finite dimensional while R(L) is infinite dimensional. For more on these ideas, see page 389, Exercise 12 and page 501 Exercises 27- 29. Exercises (1) Evaluate the following determinants as you see fit: a)./vextendsingle/vextendsingle/vextendsingle/vextendsingle7 3 2−1/vextendsingle/vextendsingle/vextendsingle/vextendsingle, b)./vextendsingle/vextendsingle/vextendsingle/vextendsingle1 25 −3 4/vextendsingle/vextendsingle/vextendsingle/vextendsingle. c)./vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle−10−2 3 −3 2 1 5 0 −1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle, d)./vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle53 17 29 36 12 39 69 23 75/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle, e)./vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 2 0 1 1 3 4 0 0 1 −5 6 1 2 3 4/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle f)./vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle2 1 1 1 1 2 1 1 1 1 2 1 1 1 1 2/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle, g)./vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea1 0 0 0 b1 0 0 0 c0 0 1 −b c0 0 1 −a d e 1f g/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle [Answers: a) −13 , b) 17, c) −14 , d) 6, e) 5, g) −(b−a)2]. (2) IfAandBare the matrices whose respective determinants appear in #1 a) and b), compute det( AB) by first finding AB. Compare with (det A)(detB) . (3) a). Use Cramer’s rule (Theorem 33) to solve the equation AX=Y, where A is given below. Then observe you have computed A−1, so exhibit it. A= 1 1 1 2−3−1 4 9 1 . [A−1=1 30 6 8 2 −6−3 3 30−5−5 ]. b). Use the formula for A−1to solve the equations AX=YwhereY= (1,2,0). 238 CHAPTER 5. MATRIX REPRESENTATION (4) a). Find the volume of the parallelepiped QinE3which is spanned by the vectors X1= (1,1,1), X 2= (2,−1,−3) andX3= (4,1,9) . [Answer: Volume = 30]. b). The matrix A, A= −10−2 3 −3 2 1 5 0 −1  −(cf. #1,c) maps E3into itself. Find the volume of the image of Q, that is, the volume of A(Q) . [Answer: 420]. (5) LetB=A−λIwhereAis a square matrix. The values λfor whichBis singular are called the eigenvalues ofA. Find the eigenvalues for a).A=/parenleftbigg3 2 2−1/parenrightbigg , b). A=/parenleftbigg3 2 1−1/parenrightbigg . c).A=/parenleftbigga b c d/parenrightbigg . [Hint: IfBis singular, then 0 = det B= det(A−λI) . Now observe that det( A−λI) is a polynomial in λ. The answer to c) is λ=1 2(a+d±/radicalbig (a+d)2−4(ad−bc))]. (6) For what value(s) of αare the vectors X1= (1,2,3), X 2= (2,0,1), X 3= (0,α,−1) linearly dependent? (7) IfX1,X2,X3andY1,Y2,Y3are vectors in R3, prove that D[X1,X2,X3]−D[Y1,Y2,Y3] =D[X1−Y1,X2,X3] +D[X1,X2−Y2,X3] +D[X1,X2,X3−Y3]. [Hint: First work out the corresponding formula for the 2 ×2 case.] (8) Here you shall compute the derivative of a determinant if the coefficients of A= ((aij)) depend ont,aij(t) . LetX1(t),...,X n(t) be the vectors which constitute the columns ofA. The problem is to compute dD(t) dt=d dtD[X1,...,X n](t) =d dt/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11(t),···, a 1n(t) · · · an1(t),···, a nn(t)/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle a). Use Exercise 7 (generalized to n×nmatrices) to show D(t+ ∆t)−D(t)≡D[X1(t+ ∆t), X 2(t+ ∆t),...]−D[X1(t), X 2(t),...] =n/summationdisplay j=1D[X1(t),...,X j−1(t), Xj(t+ ∆t)−Xj(t),Xj+1(t+ ∆b),...] 5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 239 [Hint: Do the cases n= 2 andn= 3 first]. b). Use part a to show that dD dt= lim ∆t→0/bracketleftbiggD(t+ ∆t)−D(t) ∆t/bracketrightbigg =n/summationdisplay j=1D[X1,...,X j−1,dXj dt,Xj+1,...,X n], so the derivative of a determinant is found by taking the derivative one column at a time and adding the result. (9) Letu1(t) , andu3(t) be solutions of the differential equation u/prime/prime+a1(t)u/prime+a0(t)u= 0. Consider the Wronski determinant W(u1,u2)(t) :=/vextendsingle/vextendsingle/vextendsingle/vextendsingleu1(t)u2(t) u/prime 1(t)u/prime 2(t)/vextendsingle/vextendsingle/vextendsingle/vextendsingle (a) Use Exercise 8 to prove dW dt=−a1(t)W. (b) Consequently, show W(t) =W(t0) exp/braceleftbigg −/integraldisplayt t0a1(s)ds/bracerightbigg . (c) Apply this to show that if the vectors ( u1(t),u/prime 1(t)) and (u2(t),u/prime 2(t)) are linearly independent at t=t0, then they are always linearly independent. (d) Letu1(t)...,u n(t) be solutions of the differential equation u(n)+an−1(t)u(n−1)+···+a1(t)u/prime+a0(t)u= 0. Consider the Wronski determinant of u1,...,u n W(u1,...,u n) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingleu1u2···un u/prime 1u/prime 2···u/prime n · · · u(n−1) 1a(n−) 2 ···u(n−1) n/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle ProvedW dt=−an−1(t)W, so again W(t) =W(t0) exp/braceleftbigg −/integraldisplayt t0an−1(s)ds/bracerightbigg . 240 CHAPTER 5. MATRIX REPRESENTATION (e) Use part d) to conclude that the nvectors (u1,u/prime 1,...,u(n−1) 1),(u2,u/prime 2,...,u(n−1) 2,),···(un,u/prime n,...,u(n−1) n) (where the ujare solutions of the O.D.E.) are linearly independent for all tif and only if they are so at t=t0. (10) A matrix Aisupper (lower) triangular if all the elements below (above) the main diagonal are zero, A= a11a12···an 0a22··· · 0 0 · 0 0 ···ann . IfAis upper (or lower) triangular, prove again that detA=a11a22...a nn. by expanding by minors. What is the relation of this result to the exercise ( #4 , p. 157) on echelon form? (11) LetX1,...,X nbe vectors in Rnand let ˆD(X1,...,X n) be a real valued function which has properties 1 and 4 of Theorem 21. Thus ˆDis skew-symmetric, and is linear in each of its columns. Prove ˆDnecessarily satisfies Axioms 2 and 3 for the determinant, and conclude that ˆD(X1,...,X n) =kD(X1,...,X n), where the constant k=D(e1,...,e n) . (12) Letu1(t),...,u n(t) be sufficiently differentiable functions ( Cn−1is enough). Define the Wronskian as in Exercise 9 part d. Prove that if the functions u1,...,u nare linearly dependent, then W(t)≡0 . Thus, if W(t0)/negationslash= 0 , the functions are linearly independent in any interval containing t0. [Do nottry to apply the result of Exercise 9 for it is not applicable]. (13) (a) If Iis then×nidentity matrix, evaluate det( λI) whereλis a constant. (b) IfAis ann×nmatrix, prove det(λA) =λndetA. (c) IfAorBaren×nmatrices, is det(A+B)?= detA+ detB? Proof or counterexample. (14) For what value of αdoes the system of equations x+ 2y+z= 0 −2x+αy+ 2z= 0 x+ 2y+ 3z= 0 have more than one solution? 5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 241 (15) A matrix is nilpotent if some power of it is zero, that is, AN= 0 for some positive integerN. Prove that if Ais nilpotent, then det A= 0 . (16) (a) Solve the systems of equations i)x+y= 1 ,x−.9y=−1 and ii)x+y= 1 ,x−1.1y=−1 , and compare your solutions, which should be almost the same. (b) Solve the systems of equations i)x+y= 1 ,x+.9y=−1, and x+y= 1 ,x+ 1.1y=−1. and again compare your solutions. Explain the result in terms of the theory in this section. (c) Consider the solution of the systems of equations x+y= 1 x+αy=−1 as the point where the lines x+y= 1 andx+αy=−1 intersect. Sketch the graph of these lines for αnear−1 and then for αnear +1 . Use these observations to again explain the phenomena in parts a) and b). (17) Let ∆ nbe then×ndeterminant of a matrix with a’s along the main diagonal and b’s on the two “off diagonals” directly above and below the main diagonal. Thus ∆5=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea b 0 0 0 b a b 0 0 0b a b 0 0 0b a b 0 0 0b a/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle. (a) Prove ∆ n=a∆n−1−b2∆n−2. (b) Compute ∆ 1and ∆ 2by hand. Then use the formula to compute ∆ 3and ∆ 4. (c) Ifa2/negationslash= 4b2, can you show ∆n=1√ a2−4b2 /parenleftBigg a+√ a2−4b2 2/parenrightBiggn+1 −/parenleftBigg a−√ a2−4b2 2/parenrightBiggn+1 ? Later, we shall give a method for obtaining this directly from the equation of part a). [p. 522-523]. (18) Prove Part 5 of Theorem 21 using only the axioms and no other part of Theorem 21. 242 CHAPTER 5. MATRIX REPRESENTATION (19) Apply the result of Exercise 12 on page 389. Try to prove the following. Ais asquare matrix. a). dim N(A) = dim N(A∗). Thus, the homogeneous equation AX= 0 has the same number of linearly indepen- dent solutions as does the equation A∗Z= 0 . b). LetZ1,...,Z kspanN(A∗) . Then the inhomogeneous equation AX=Y has a solution, that is, Y∈R(A) , if and only if /angbracketleftZj, Y/angbracketright= 0, j = 1,2,...,k. In other words, the equation AX=Yhas a solution if and only if Yis orthogonal to the solutions of the homogeneous adjoint equation. c). Consider the system of linear equations 2x−3y+z= 1 −3x+ 2y−4z=α x−4y−2z=β. LetAbe the coefficient matrix. Find a basis for N(A∗) . [Answer: dim N(A∗) = 1 andZ1= (2,1,−1) is a basis]. For what value(s) of the constants α,β can you solve the given system of equations? [Answer: There is a solution if and only if β−α= 2 .] Find a solution if α= 1 andβ= 3 . d). Repeat part c) for the system of equations x−y= 1 x−2y=−1 x+ 3y=α. [Answer: dim N(A) = 1 and Z1= (−5,4,1) is a basis. There is a solution if and only ifα=−1 ]. (20) Use the result of Exercise 12 to prove that each of the following sets of functions are linearly independent everywhere. a)u1(x) = sinx, u 2(x) = cosx b)u1(x) = sinnx, u 2(x) = cosmx, wheren/negationslash= 0. c)u1(x) =ex, u2(x) =e2x, u3(x) =e3x. d)u1(x) =eax,u2(x) =ebx, u3(x) =ecx, wherea,b, andcare distinct numbers. e)u1(x) = 1, u2(x) =x, u 3(x) =x2, u4(x) =x3 f)u1(x) =ex, u2(x) =e−x, u3(x) =xex, u4(x) =xe−x. 5.4. AN APPLICATION TO GENETICS 243 5.4 An Application to Genetics A mathematical model is developed and solved. Although this particular model will be motivated by genetics, the resulting mathematical problem also arises in sociometrics and statistical mechanics as well as many other places. In the literature you will find these mathematical ideas listed under the title Markov chains . Part of the value you should glean from our discourse is insight into the process of going from vague qualitative phenomena to setting up a quantitative model. One part of this scientific process we shall not have time to investigate in detail is the very important step of comparing the quantitative results with experimental data. Furthermore, we shall never delve into the fertile realm of generalizing our accumulated knowledge to more complicated - as well as more interesting and realistic - situations. In bisexual mating, the genes of the resulting offspring occur in pairs, one gene in each pair being contributed by each parent. Consider the simplest case of a trait which is determined by a single pair of genes, each of which is one of two types gandG. Thus, the father contributes Gorgto the pair, and the mother does likewise. Since experimental results show that the pair Ggis identical to the pair gG, the offspring has one of the three pairs GG Gg gg. The geneGdominatesgif the resulting offspring with genetic types GGandGg“appear” identical but both are different from gg. In this case, an individual with genetic type GG is called dominant , while the types ggandGgare called recessive andhybrid , respectively. An offspring can have the pair GG(resp.gg) if and only if bothparents contributed a gene of type G(resp. g) while the combination Ggoccurs if either parent contributed Gand the other g. A fundamental assumption we shall make is that a parent with genetic typeabcan only contribute a gene of type aor of typeb. This assumption ignores such things as radioactivity as a genetic force. Thus, a dominant parent, GGcanonlycontribute a dominant gene, G, a recessive parent, gg, can only contribute g, and a hybrid parent Ggcan contribute eitherGorg(with equal probability). Consequently, if two hybrids are mated, the offspring has probability1 2of getting Gorgfrom each parent, so the probability of his having genetic type GGofggis1 4each, while the probability of having genetic type Ggis1 2. We introduce a probability vector V= (v1,v2,v3) , withv1representing the probability of being genetic type GG,v2of being type Gg, andv3of being type gg. Thus for an offspring of two hybrid parents ,V= (1 4,1 2,1 4) . Observe that, by definition of probability, 0≤vj≤1, j= 1,2,3 , andv1+v2+v3= 1 (since with probability one - certainty - the offspring is either GG,Gg, orgg). Consider the issue of mating an individual whose genetic type is unknown with an individual of known genetic type (dominant, hybrid or recessive). To be specific, assume the known person is of dominant type. Then the following matrix of transition probabilities D= 1 1/2 0 0 1/2 1 0 0 0  describes the probability of the offspring’s genetic type in the following sense: if the unknown parent had genetic type V0(soV0= (1,0,0) if unknown was dominant, V0= (0,1,0) if 244 CHAPTER 5. MATRIX REPRESENTATION hybrid, and V0= (0,0,1) if recessive), then V1=DV 0, is the probability vector of the offspring. For example, if the unknown parent was hy- brid, then V1=DV 0= (1 2,1 2,0) . Thus the offspring can, with equal likelihood, be either dominant or hybrid, but cannot be recessive. Notice that the matrix Dembodies the fact that one of the parents is dominant. If the individual of unknown genetic type were crossed with an individual of hybrid type, then the corresponding matrix His H= 1 21 40 1 21 21 2 01 41 2 , while if the person of unknown type were crossed with the individual of recessive type, then R= 0 0 0 11 20 01 21 . It is of interest to investigate the question of genetic stability under various circum- stances. Say we begin with an individual of unknown genetic type and cross it with a dominant individual, then cross that offspring with another dominant individual, and so on, always mating the resulting offspring with a dominant individual. Let Vnrepresent the genetic probability vector for the offspring in the nth generation. Then Vn=DVn−1=D2Vn−2=···=DnV0, whereV0is the unknown vector for the initial parent (of unknown genetic type). Without knowingV0, can we predict the eventual ( n→ ∞ ) genetic types of the offspring? Intu- itively, we expect that no matter what the type of the initial parent, the repeated mating with a dominant individual will produce a dominant strain. The question we are asking is, does lim n→∞Vnexist, and if so, what is it? Assume for the moment that the limit does exist and denote it by V. ThenV=DV since V= lim n→∞Vn= lim n→∞Vn+1= lim n→∞DVn=D( lim n→∞Vn) =DV Armed with the equation DV =V, we can solve linear equations for the vector V= (v1,v2,v3) v1+1 2v2+ 0 =v1 0 +1 2v2+v3=v2 0 + 0 + 0 = v3. Clearlyv1=v2=v3= 0 is a trivial solution. A non-trivial one can be found by transposing thevj’s to the left side and solving. We find v1= 1,v2= 0,v3= 0(v1= 1 since v1+v2+v3= 1 ). Thus, ifthe limitVnexists, the limit must be V= (1,0,0) . In genetic terms, this sustains our feeling that the offspring will eventually become genetically dominant. 5.4. AN APPLICATION TO GENETICS 245 But does the limit exist? To prove it does, we must show for any probability vector V0= (v1,v2,v3) , wherev1+v2+v3= 1 , that the limit lim n→∞Vn= lim n→∞DnV0, exists and equals V= (1,0,0) . By evaluating D,D2,andD3explicitly, we are led to guess Dn= 1 1−1 2n1−1 2n−1 01 2n1 2n−1 0 0 0 , which is then easily verified using mathematical induction. Thus Vn=DnV0= v1+ (1 −1 2n) + (1 −1 2n−1)v3 0 +1 2nv2 +1 2n−1v3 0 + 0 + 0  = v1+v2+v3−1 2n(v2+ 2v3) 1 2nv2+1 2n−1v3 0  Sincev1+v2+v3= 1 , we find Vn=DnV0= 1 0 0 +1 2n(v2+ 2v3) −1 1 0 . It is now clear that the limit as n→ ∞ does exist, and is V= (1,0,0) . Consequently, if we begin with a random individual (you) and mate that individual and the successive offspring with a dominant gene bearer, then the resulting generations will tend to all dominant individuals. Moreover, the process proceeds exponentially because the “damping factor” is essentially1 2for each generation (see above formula). Were there enough time, you would see a second application of matrices to the special theory of relativity. Given your knowledge of linear spaces, it is possible to present an elegant exposition of the theory. The Lorentz transformation would appear as an orthogonal transformation - a rotation - in world space orMinkowski’s space as it is often called. This is a four dimensional space three of whose dimension are those of ordinary space, while the fourth dimension is an imaginary (i=√−1) time dimension. Goldstein’s Classical Mechanics contains the topic. Regrettably, he does not begin with the Michelson - Morley experiment but rather plunges immediately into mathematical technicalities. Exercises 1. If you begin with an individual of unknown genetic type and cross it with a hybrid individual and then cross the successive offspring with hybrids, does the resulting strain approach equilibrium? If so, what is it? 2. Same as 1 but you mate an individual of unknown type with a recessive individual. 3. Beginning with an individual of unknown genetic type, you mate it with a dominant individual, mate the offspring with a hybrid , mate that offspring with a dominant, and continue mating alternate generations with dominants and hybrids respectively. Does the 246 CHAPTER 5. MATRIX REPRESENTATION resulting strain approach equilibrium? If so, what is it? (You will need to define equilibrium to cope with this problem. There are several reasonable definitions.) 4. a). The city Xhas found that each year 5% of the city dwellers move to the suburbs, while only 1% of the suburbanites move to the city. Assuming the total population of the city plus suburb does not change, show that the matrix of transition probabilities is P=/parenleftbigg.95.01 .05.99/parenrightbigg , where a vector V= (v1,v2) = (proportion of people in city, proportion of people in suburb). b). Given any initial population distribution V, does the population approach an equilibrium distribution? If so, find it. 5. A long queue in front of a Moscow market in the Stalin era sees the butcher whisper to the first in line. He tells her “Yes, there is steak today.” She tells the one behind her and so on down the line. However, Moscow housewives are not reliable transmitters. If one is told “yes”, there is only an 80% chance she’ll report “yes” to the person behind her. On the other hand, being optimistic, if one hears “no”, she will report “yes” 40% of the time. If the queue is very long, what fraction of them will hear “there is no steak”? [This problem can be solved without finding a formula for Pn, although you might find it a challenge to find the formula]. 5.5 A pause to find out where we are . We all know the homily about the forest and the trees. The next few pages are about the forest. In the beginning we introduced dead linear spaces with their algebraic structure (Chap- ter II). Then we investigated the geometry induced by defining an inner product on a linear space and saw how easily many of the results in Euclidean geometry generalize (Chapter III). Our next step was to consider mappings, linear mappings, between linear spaces (Chap- ter IV). Not much could be said in general, so we began investigating a particular case, linear maps between finite dimensional spaces. Two important special cases of this L:R1→Rn, and L:Rn→R1, were treated before the general case, L:Rn→Rm. A key theorem which facilitates the theory of linear mappings between finite dimensional spaces is the representation theorem (page 374): every such map can be represented as a matrix. What next? There are two equally reasonable alternatives: 5.5. A PAUSE TO FIND OUT WHERE WE ARE 247 (A) We can continue with linear maps , L:V1→V2, and consider the case where V1orV2, or both are infinite dimensional. The general theory here is in its youth and still undeveloped. Only one of the sources of difficulty is that a generalization of the representation theorem (page 374) remains unknown - except for some special cases. Thus, many special types of mappings have to be investigated individually. We shall consider only one type of linear mapping between infinite dimensional spaces, those defined by linear differential operators (Chapter VI and Chapter VII, Section 3). (B) The second alternative is to continue our study of mappings between finite dimensional spaces, only now switch to non linear mappings. This theory should parallel the transition in elementary calculus from the analytic geometry of straight lines, f(x) =a+bx, that is, affine mappings, to genuine non linear mappings, as f(x) =x2−7√x or f(x) =x3−esinx. You recall, one important idea was to approximate the graph of a function y=f(x) at a pointx0by its tangent line at x0, since forxnearx0, the curve and the tangent line there approximately agree. For example, one easily proves that at a maximum or minimum, the tangent line must be horizontal, f/prime= 0 . In generalizing this to functions of several variables, Y=F(X) =F(x1,···,xn), the role of the derivative at X0is assumed by the affine map, A(X) =Y0+LX, which is tangent to FatX0. Thus, linear algebra appears as the natural extension of analytic geometry to higher dimensional spaces. See Chapters VII - IX for this. 248 CHAPTER 5. MATRIX REPRESENTATION Chapter 6 Linear Ordinary Differential Equations 6.1 Introduction . Adifferential equation is an equation relating the values of a function u(t) with the values of its derivatives at a point, F(t,u(t),du dt,...,dnu dtn) = 0 (6-1) The order of the equation is the order, n, of the highest derivative which appears. For example, the equations/parenleftbiggd2u dt2/parenrightbigg3 −7du dt+t2u2−sint= 0 du dt−tsinu2= 0 are of order two and one respectively. A function u(t) is a solution of the differential equation if it has at least as many derivatives as the order of the equation, and if substitution of it into the equation yields an identity. Thus, the equation /parenleftbiggdu dt/parenrightbigg2 +u2= 1 has the function u(t) = sintas a solution, since for all t /parenleftbiggd dtsint/parenrightbigg2 + (sint)2= 1. A differential equation (1) for the unknown function u(t) islinear if it has the form Lu:=an(t)dnu dtn+an−1(t)dn−1 dtn−1+···+a0(t)u= 0 (6-2) You should verify that this coincides with the notion of a linear operator used earlier. Equa- tion (2) is sometimes called linear homogeneous to distinguish it from the inhomogeneous equation Lu=f(t), (6-3) 249 250 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS that is an(t)dnu dtn+···+a0(t)u=f(t). (6-4) The subject of this chapter is linear ordinary differential equations with variable coeffi- cients (to distinguish them from the special case where the aj’s are constants). This opera- torLdefined by (2) has as its domain the set of all sufficiently differentiable functions— n derivatives is enough. These functions constitute an infinite dimensional linear space. Thus, the differential operator Lacts on an infinite dimensional space, as opposed to a matrix which acts on a finite dimensional space. Differential equations abound throughout applications of mathematics. This is because most phenomena are described by laws which relate the rate of change of a function - the derivative - at a given time (or point) to the values of the function at that same time. For example, we have seen that at any time the acceleration of a harmonic oscillator is determined by its position and velocity at the same time, ¨u=−µ˙u−ku. When confronted by a differential equation, your first reaction should be to attempt to find the solution explicitly. We were able to do this for linear constant coefficient equations (Chapter 4, Section 2). One of the main goals of this chapter is to show you how to solve as many linear ordinary differential equations as possible. However, it is naive to expect to solve an arbitrary equation which crops up in terms of the few functions we know: xα,ex,logx,sinx, and cosx. In fact, to even solve the elementary equation du dx=1 x, appearing in elementary calculus, we were forced to define a new function as the solution of this equation u(x) = logx+c and obtain the properties of this function and its inverse exdirectly from the differential equation. Many many functions arise which cannot be expressed in terms of the few elemen- tary functions we know and love. Most of these functions - like Bessel’s functions, elliptic functions, and hypergeometric functions, arise directly because they are the solutions of differential equations nature has forced us to consider. How do we know these strange sounding functions are solutions of the differential equa- tions? Well, we somehow prove a solution exists and then simply give a name to the solution - much as babies are given names at birth. Furthermore, as is the case with babies, their actual “names” are the least important aspect. To summarize briefly, we shall solve as many equations as we can. For the remaining ones (which include most equations), we shall attempt to describe a few of the main prop- erties so that if one arises in your work, you will have a place to begin the attack. Later on, we shall again return to the more complicated situation of nonlinear equations. Much less can be said there. Only very few general results are known. Lest you get the wrong idea, we shall cover but a fraction of the known theory for just linear ordinary differential equations. In the next chapter, we shall only look at one partial differential equation (the wave equation for a vibrating violin string). The general theory there is too complicated to allow discussion for more than one particular equation. 6.1. INTRODUCTION 251 Exercises 1. Assume there exists aunique functionE(x) which satisfies the following differential equation for all xand satisfies the initial condition du dx=u, u (0) = 1. (a) Use the “chain rule” and uniqueness to prove for any a∈R E(x+a) =E(a)E(x) [Hint: Prove ˜E(x) :=E(x+a) is also a solution of the equation. Then apply the uniqueness to the function ˜E(x)/E(a) ]. (b) Prove E(−x) =1 E(x). (c) Prove for any x E(nx) = [E(x)]n, n∈Z. In particular, show E(n) = [E(1)]n, n∈Z and E(1 m) = [E(1)]1/m, m ∈Z+ (d) Prove E(n m) = [E(1)]n/m, n∈Z, m ∈Z+ [Thus, the function E(x) is defined for all rational x=n mas the number E(1) to the powern/m . SinceE(x) is continuous (even differentiable by definition, we can extend the last formula to irrational xby continuity: if rjis a sequence of rational numbers converging to the real number x(which may or may not be rational) then by continuity E(x) = lim j→∞E(rj) = lim j→∞[E(1)]rj=E(1)x. Consequently, E(x) is the familiar exponential function ex]. 2. Find the general solutions of the following equations by any method you can. (a)du dx−2u= 0 (b)du dx=x2+ sinx (c)/parenleftbigdu dx/parenrightbig2+ 4u2= 1 (d)du dx=x u+1 (e)du dx=x2eu (f)d2u dx2+ 3du dx−4u= 4 252 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS 6.2 First Order Linear . Except for those differential equations which can be solved by inspection, the next most simple equation is one which is linear and first order, the homogeneous equation du dx+a(x)u= 0, (6-5) and the inhomogeneous equation du dx+a(x)u=f(x). (6-6) The homogeneous equation can be solved by first writing it in the form 1 udu dx=−a(x) and then integrating both sides logu(x) =−/integraldisplayx a(s)ds+C1. Thus u(x) =Ce−/integraldisplayx a(s)ds (6-7) is the solution of equation (4) for any constant C. In the very special case a(s)≡constant, the solution does have the form found earlier (Chapter 4, Section 2) for a linear equation with constant coefficients. How can we integrate the inhomogeneous equation (5)? A useful device is needed. Multiply both sides of this equation by an unknown function q(x) q(x)du dx+q(x)a(x)u=q(x)f(x), Ifwe can find q(x) so that the left side is a derivative, q(x)du dx+q(x)a(x)u=d dx(q(x)u), (6-8) then the equation reads d dx(q(x)u) =q(x)f(x), which can be integrated immediately, q(x)u(x) =/integraldisplayx q(s)f(s)ds+c, (6-9) and then solved for u(x) by dividing by q(x) . Thus, the problem is reduced to finding a q(x) which satisfies (7). Evaluating the right side of (7), we find qdu dx+qau=udq dx+qdu dx, 6.2. FIRST ORDER LINEAR 253 soq(x) must satisfy dq dx=q(x)a(x). It is easy to find a function q(x) which satisfies this - for it is a homogeneous equation of the form (4). Therefore q(x) =eRxa(t)dt, the reciprocal of the solution (6) to the homogeneous equation, does satisfy (7). Notice we have ignored the arbitrary constant factor in the solution since all we want is any one functionq(x) for (7). Now we can substitute into (8) to find the solution of the inhomogeneous equation u(x) =1 q(x)/integraldisplayx q(s)f(s)ds+c q(x), (6-10) whereq(x) is given by the formula at the top of the page. If it makes you happier, substitute the expression for q(x) into (9) to obtain the messy formula. We have left some room. a figure goes here Examples: 1.du dx+2 xu= (1 +x3)17, x/negationslash= 0. First, q(x) = exp(/integraldisplayx2 sds) = exp(2 ln x) = exp(lnx2) =x2. Thusd dx(x2u) =x2(1 +x3)17. Integrating both sides we find x2u(x) =(1 +x3)18 54+C. Therefore u(x) =1 54(1 +x3)18 x2+C x2, x/negationslash= 0. 2.du dx+ 2xu=x First, q(x) = exp(/integraldisplayx 2sds) = expx2. Thus, d dx(ex2u) =ex2x. Integrating both sides, we find ex2u(x) =1 2ex2+C, so u(x) =1 2+Ce−x2. This formula could have been guessed much earlier since we know the general solution of the inhomogeneous equation can be expressed as the sum of a particular solution to that 254 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS equation plus the general solution of the homogeneous equation. The particular solution u0(x) =1 2can be obtained by inspection of the D.E. Let us summarize our results. Theorem 6.1 . Consider the first order linear inhomogeneous equation Lu:=du dx+a(x)u=f(x). Ifa(x)andf(x)are continuous functions, the equation has the solutions u(x) = ˜u(x)/integraldisplayxf(s) ˜u(s)ds+C˜u(x) (6-11) where ˜u(x) = exp( −/integraldisplayx a(s)ds) is a non-trivial solution of the homogeneous equation. Moreover, if we specify the initial conditionu(x0) =α, then the solution which satisfies this initial condition is unique. Proof: Theexistence follows from the explicit formula (9) or (10) and from the fact that a continuous function is always integrable. Uniqueness . This will be quite similar to the proof carried out in Chapter 4. If u1(x) andu2(x) are two solutions of the inhomogeneous equation Lu=f, with the same initial conditions, then the function w(x) :=u1(x)−u2(x) satisfies the homogeneous equation Lw:=w/prime+a(x)w= 0, and is zero at x0, w(x0) =u1(x0)−u2(x0) = 0. Our task is to prove w(x)≡0 . Multiply the equation (20) by w(x) . Then ww/prime=−a(x)w2, or1 2d dxw2=−a(x)w2. Sincea(x) is continuous, for any closed and bounded interval [ A,B] there is a constant k (depending on the interval) such that −a(x)≤kfor allx∈[A,B] . Consequently, 1 2d dxw2≤kw2, ord dxw2−2kw2≤0. Now we need an important identity which can be verified by direct computation: for any smooth function g, and any constant α,g/prime+αg=e−αx(eαxg)/prime. We apply this to the above inequality with g=w2andα=−2kto conclude that e2kxd dx[e−2kxw2]≤0. 6.2. FIRST ORDER LINEAR 255 Becausee2kxis always positive, by the mean value theorem this inequality states that e−2kxw2is a decreasing function of x. Thus e−2kxw2(x)≤e−2kx0w2(x0), x≥x0, or w2(x)≤e2k(x−x0)w2(x0), x≥x0. But sincew(x0) = 0 andw2(x)≥0 this means that 0≤w2(x)≤0. Thereforew(x)≡0x≥x0. To provew(x)≡0 forx≤x0, merely observe that the equation (11) has the same form ifxis replaced by −x. Thus the above proof applies and shows w(x)≡0 forx≤x0 too. Remark: Although a formula has been exhibited for the solution, this does not mean that the integrals which occur can be evaluated in terms of elementary functions. These integrals however can be at least evaluated approximately using a computer if a numerical result is needed. Exercises (1) . Find the solution of the following equations with given initial values (a)u/prime+ 7u= 3, u(1) = 2 (b) 5u/prime−2u=e3x, u(0) = 1. (c) 3u/prime+u=x−2x2, u(−1) = 0. (d)xu/prime+u= 4x3+ 2, u(1) = −1. (e)u/prime+ (cotx)u=ecosx+ 1, u(π 2) = 0.[/integraltext cotxdx= ln(sinx)] . (2) . The differential equation Ldu dt+Ru=Esinωt, L,R,E constants arises in circuit theory. Find the solution satisfying u(0) = 0 and show that it can be written in the form u(t) =ωEL R2+ω2L2e−Rt/L+E√ R2+ω2L2sin(ωt−α) where tanα=ωL R. (3) Bernoulli’s equation is u/prime+a(x)u=b(x)uk, k a constant. 256 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS (a) Use the substitution v(x) =u(x)1−kto transform this nonlinear equation to the linear equation v/prime+ (1−k)a(x)v= (1−k)b(x). (b) Apply the above procedure to find the general solution of u/prime−2exu=exu3/2. (4) . Consider the equation u/prime+au=f(x), whereais a constant, fis continuous in the interval [0 ,∞] , and |f(x)|< M for allx. (a) Show that the solution of this equation is u(x) =e−axu(0) +e−ax/integraldisplayx 0eatf(t)dt (b) Prove (if a/negationslash= 0 ) /vextendsingle/vextendsingleu(x)−e−axu(0)/vextendsingle/vextendsingle≤M a[1−e−ax]. (5) (a) Show the uniqueness proof yields the following stronger fact. If u1(x) andu2(x) are both solutions of the same equation u/prime+a(x)u=f(x) but satisfy different initial conditions u1(x0) =α, u 2(x0) =β, then |u1(x)−u2(x)| ≤ek(x−x0)|α−β|, x≥x0 for allx∈[A,B] , where −a(x)< k in the interval. Thus, if the initial values are close, then the solutions cannot get too far apart. (b) Show that if a(x)≤A < 0 , whereAis a constant, then as x→0 any two solutions of the same equation - but with possibly different initial values - tend to the same function. (6) . Show that the differential equation y/prime=a(x)F(y) +b(x)G(y) can be reduced to a linear equation by the substitution u=F(y)/G(y) oru=G(y)/F(y) if (FG/prime−GF/prime)/Gor (FG/prime−GF/prime)/F, respectively, is a constant. Use this substitution to again solve Bernoulli’s equation. 6.2. FIRST ORDER LINEAR 257 (7) . LetS={u∈C/prime:u(0) = 0 }, and define the operator LfromStoCby Lu=u/prime+u. ProveLis injective and R(L) =C. (8) . Set up the differential equation and solve. The rate of growth of a bacteria culture at any time tis proportional to the amount of material present at that time. If there was one ounce of culture in 1940 and 3 ounces in 1950, find the amount present in the year 2000. The doubling time is the interval it takes for a given amount to double. Find the doubling time for this example. (9) . Find the general solution of x2u/prime+ 3xu= sinx. (10) . Assume that a body decreases its temperature u(t) at a rate proportional to the difference between the temperature of the body and the temperature Tof the sur- rounding air. A body originally at a temperature of 1000is placed in air which is kept at a temperature of 500. If at the end of one hour the temperature of the body has fallen 200, how long will it take for the body to reach 600? (11) . Here is one simple mathematical model governing economic behavior. Think of yourself as a widget manufacturer for now. Let i)S(t) be the supply of widgets available at time t. This is the only function you can control directly. ii)P(t) be the market price of a widget at time t. iii)D(t) is the demand for widgets at time t—the number of widgets people want to buy at time t. You cannot control this given function. It has been found that the market price P(t) changes at a rate proportional to the difference between demand and supply, dP dt=k(D(t)−S(t)), wherek>0 is a fixed constant. You decide to vary the supply so that it is a fixed constant S0plus an amount proportional to the market price, S(t) =S0+αP(t), α> 0. (a) Set up the differential equation for S(t) in terms of the given function D(t) and solve it. (b) Analyze the solution and give an argument making it plausible that the market for widgets behaves roughly in this way. What criticisms can you make of the model? (c) How does the market behave if the demand increases for a long time and then levels off at some constant value, D(t) =D(t1) fort≥t1? A qualitative description of S(t) andP(t) is called for here. In particular, say whether price increases without bound (bringing the evils of inflation) or whether it, too, levels off. 258 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS (12) It is found that a juicy rumor spreads at a rate proportional to the number of people who “know”. If one person knows initially, t= 0 , and tells one other person by the next day, t= 1 , approximately how long does it take before 4000 people know? Analyze the mathematical model as t→ ∞ and state why it is, in fact, the wrong model. (The question to ask yourself is, “how long will it take before everyone even remotely concerned knows?”). The same mathematical model applies to the spreading of contagious diseases - and many other similar phenomena. 6.3 Linear Equations of Second Order In this section we will consider a portion of the general theory of second order linear O.D.E.’s, with variable coefficients, Lu:=a2(x)d2u dx2+a1(x)du dx+a0(x)u=f(x). Although allof the results obtained generalize immediately to linear equation of order n, only the special case n= 2 will be treated. This special case has the advantage of clearly illustrating the general situation and supplying proofs which generalize immediately - while avoiding the inevitable computation complexities inherent in the general case. There are three parts: A). a review of the constant coefficient case, B). power series solutions, and C). the general theory. Whereas the first two parts are concerned with obtaining explicit formulas for the solutions, the last resigns itself to some statements which can be made without finding the solution explicitly. a) A Review of the Constant Coefficient Case. Here we have the operator Lu:=a2u/prime/prime+a1u/prime+a0u, (6-12) wherea0,a1, anda2are constants. In order to solve the homogeneous equation Lu= 0, the function eλxis tried. Substitution yields L(eλx) = (a2λ2+a1λ+a0)eλx=p(λ)eλx. (6-13) labeleq:13 The polynomial p(λ) is called the characteristic polynomial forL. Ifλ1is a root of this polynomial, p(λ1) = 0 , then u1(x) =eλ1xis a solution of the homogeneous equationLu= 0 . Ifλ2is another root of this polynomial λ1/negationslash=λ2, u 2(x) =eλ2xis another solution. Then every function of the form u(x) =Au1(x) +Bu2(x) =Aeλ1x+Beλ2x, (6-14) whereAandBare constants, is a solution of the homogeneous equation. The uniqueness theorem showed that every solution of Lu= 0 is of the form (14). 6.3. LINEAR EQUATIONS OF SECOND ORDER 259 Ifthe two roots ofp(λ)coincide , then a second solution is u2(x) =xeλ1x, and every function of the form u(x) =Au1(x) +Bu2(x) =Aeλ1x+Bxeλ1x(6-15) whereAandBare constants, is a solution of the homogeneous equation. Again the uniqueness theorem showed that every solution of Lu= 0 is of the form (15). In both (14) and (15), the constants AandBcan be chosen to find a unique function u(x) which satisfies the homogeneous equation Lu= 0 as well as the initial conditions u(x0) =α, u/prime(x0) =β, whereαandβare specified constants. It turns out that the inhomogeneous O.D.E. Lu=f, wherefis a given continuous function, can always be solved once two linearly independent solutionsu1andu2of the homogeneous equation Luj= 0 are known. Since the procedure for solving the inhomogeneous equation also works if the coefficients in the differential operatorLare not constant, it is described later in this section in the more general situation (p. 487-8, Theorem 8). Somewhat simpler techniques can be used for the constant coefficient equation if the function fis a linear combination of functions of the form xkerx, wherek is a nonnegative integer and ris some real or complex constant (cf. Exercise 6, p. 300). Because both sin nxand cosnxare of this form, Fourier series can be used to supply a solution for any function fwhich has a convergent Fourier series (cf. Exercise 13, p. 303). Section 5 of this chapter contains an interesting generalization of the theory for constant coefficient ordinary differential operators to operators which are “translation invariant”. b) Power Series Solutions . Many ordinary differential equations (linear and nonlinear) can be solved by merely assuming the solution can be expanded in a power series u(x) =/summationtextcnxn, and plugging into the differential equation to find the coefficients cn. A simple example illustrates this. Example: Solveu/prime/prime−2xu/prime= 0 with the initial conditions u(0) = 1,u/prime(0) = 0 . Solution: We try u(x) =c0+c1x+c2x2+···+cnxn+···. Then u/prime(x) =c1+ 2c2x+ 3c3x2+···+ncnxn−1+··· so 2xu/prime(x) = 2c1x+ 4c2x2+···+ 2ncnxn+··· 260 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS Also u/prime/prime(x) = 2c2+ 2·3c3x+ 3·4c4x2+·+ (n−1)ncnxn+··· Addingu/prime/prime−2xu/prime−uand collecting like powers of xwe find that 0 = u/prime/prime−2xu/prime−u= [2c2−c0] + [2·3c3−2c1−c1]x+ [3·4c4−4c2−c2]x2 +···+ [(k+ 1)(k+ 2)ck+2−2kck−ck]xk+··· If the right side, a Taylor series, is to be zero (= the left side), then the coefficient of each power ofxmust vanish because the only convergent Taylor series for zero is zero itself. The coefficient of x0is 2 c2−c0 x1is 6 c3−3c1 x2is 12 c4−5c2 xkis ( k+ 1)(k+ 2)ck+2−(2k+ 1)ck Equating these to zero we find that c2=c0 2, c 3=c1 2, c 4=5c2 12=5 24c0 and, more generally, ck+2=2k+ 1 (k+ 2)(k+ 1)ck. (6-16) Thus, for this example eeven is some multiple of c0whilecoddis some multiple of c1. Sinceu(0) =c0andu/prime(0) =c1, the constants c0andc1are determined by the initial conditions. c0= 1, c 1= 0. Consequently, all of the odd coefficients c3,c5,... vanish, while c2=1 2, c 4=5 24, c 6=3 10c4=1 16, c 8=..., so the first few terms in the series for u(x) are u(x) = 1 +1 2x2+5 24x4+1 16x6+··· (6-17) We should investigate if this formal power series expansion converges. Using (16), the ratio of successive terms in the series for u(x) is /vextendsingle/vextendsingle/vextendsingle/vextendsingleck+2xk+2 ckxk/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle(2k+ 1) (k+ 2)(k+ 1)x2/vextendsingle/vextendsingle/vextendsingle/vextendsingle Therefore the ratio test shows the formal power series actually converges for all x. By Theorem 16, p. 82, the series can be differentiated term by term and does satisfy the equation. Although the computation is lengthy, the series (17) is a solution. Since there is no way of finding the solution in terms of elementary functions, we must be contented with the power series solution. You have seen (Chapter 1, Section 7) how properties of a function can be extracted from a power series definition. This example is typical. 6.3. LINEAR EQUATIONS OF SECOND ORDER 261 Theorem 6.2 . If the differential equation a2(x)u/prime/prime+a1(x)u/prime+a0(x)u= 0 has analytic coefficients about x= 0, that is, if the coefficients all have convergent Taylor series expansions about x= 0, and ifa2(0)/negationslash= 0, then given any initial values u(0) =α, u/prime(0) =β, there is a unique solution u(x)which satisfies the equation and initial conditions. Moreover, the solution is analytic about x= 0 and converges in the largest interval [−r,r]in which the series for a1/a2anda0/a2both converge. Outline of Proof . There are two parts: i) find a formal power series u(x) =/summationtextcnxn, and ii) prove the formal power series converges. Since explicit formulas can be found for the cn’s (cf. Exercise 30a) the first part is true. Proof of the second part is sketched in the exercises too (Exercise 30b). From the explicit formulas mentioned above for the cn’s, it is clear there is at most oneanalytic solution. But because the general uniqueness proof (p. 510, Theorem 9) states there is at most one solution which is twice differentiable and since u(x) is certainly such a function - the uniqueness of u(x) among all twice differentiable functions follows as soon as Theorem 9 is proved. The restriction a2(0)/negationslash= 0 which was made in Theorem 3 is very important. If a2(0) = 0 then the differential equation a2(x)u/prime/prime+a1(x)u/prime+a0(x)u= 0 is degenerate at x= 0 because the coefficient of the highest order derivative vanishes there. Then the point x= 0 is called a singularity of the differential equation. A simple example illustrates the situation. The function u(x) =x5/2satisfies the differential equation 4x2u/prime/prime−15u= 0 and the initial conditions u(0) = 0,u/prime(0) = 0 . However u(x)≡0 is also a solution. Thus it will be impossible to prove any uniqueness theorem at x= 0 for this equation. Perhaps the singular nature of this equation at x= 0 is more vivid if the equation is written as u/prime/prime−15 4x2u= 0. Although the possibility of a uniqueness result is ruled out for equations with singular- ities, it is important to be able to find the non-zero solutions of these equations, important because many of the equations which arise in practice do happen to have singularities (Bessel’s equation, Legendre’s equation, the hypergeometric equation, ...). In all of the commonly occurring cases, the coefficients a0(x), a1(x) , anda2(x) , a2u/prime/prime+a1u/prime+a0u= 0, are analytic functions. Thus the only obstacle to applying Theorem 3 is the condition a2/negationslash= 0 . We persist, however, in the belief that a power series, or some modification of it, 262 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS should work. The modification must allow for such solutions as u(x) =x3/2which do not have Taylor expansions about x= 0 . Undoubtedly the most naive candidate for a solution is to try u(x) =xρ∞/summationdisplay n=0cnxn, (6-18) whereρmay be any real number. The particular choice ρ= 3/2, c0= 1,c1=c2=c3= ...= 0 does yield the function u(x) =x3/2. It turns out that (18) is usually the correct guess. Again, we turn to an example. Bessel’s equation of ordern, x2u/prime/prime+xu/prime+ (x2−n2)u= 0, which arises in the study of waves in a two dimensional circular domain, like those on tympani, in a tea cup, or on your ear drum. Let us find a solution to Bessel’s equation of order one, x2u/prime/prime+xu/prime+ (x2−1)u= 0 (6-19) This equation does have a singularity at the origin, x= 0 . Ifuhas the form (18), then u(x) =∞/summationdisplay n=0cnxn+ρ, u/prime(x) =∞/summationdisplay n=0(n+ρ)cnxn+ρ−1, and u/prime/prime(x) =∞/summationdisplay n=0(n+ρ)(n+ρ−1)cnxn+ρ−2. Substituting this into the differential equation (19), we find ∞/summationdisplay n=0(n+ρ)(n+ρ−1)cnxn+p+∞/summationdisplay n=0(n+ρ)cnxn+ρ +∞/summationdisplay n=0cnxn+ρ+2−∞/summationdisplay n=0cnxn+ρ= 0. (6-20) We must equate the coefficients of successive powers of xto zero. The lowest power of x which appears is xρ, the nextxρ+1, and so on. xρ:ρ(ρ−1)c0+ρc0−c0= 0 xρ+1: (ρ+ 1)ρc1+ (ρ+ 1)c1−c1= 0 xρ+2: (ρ+ 2)(ρ+ 1)c2+ (ρ+ 2)c2+c0−c2= 0 xρ+3: (ρ+ 3)(ρ+ 2)c3+ (ρ+ 3)c3+c1−c3= 0 · · · · xρ+n: (ρ+n)(ρ+n−1)cn+ (ρ+n)cn+cn−2−cn= 0. 6.3. LINEAR EQUATIONS OF SECOND ORDER 263 From the equation for the power xρ, we find (ρ2−1)c0= 0 The polynomial q(ρ) =ρ2−1 which appears in the coefficient of the lowest power of x in (20) is called the indicial polynomial since it will be used to determine the indexρ. If c0/negationslash= 0 , the equation ( ρ2−1)c0= 0 can be satisfied only if ρis a root of the indicial polynomial. Thus ρ1= 1, ρ2=−1 . Consider the largest root ρ1= 1 . Then the equation for the coefficients of xρ+1in (20) is xρ+1=x2: 3c1= 0⇒c1= 0, while the equation for the coefficient of xρ+nin (20) is xρ+n=x1+n: (n+ 1)ncn+ (n+ 1)cn+cn−2−cn= 0, or cn=−cn−2 n(n+ 2), n = 2,3,... Sincec1= 0 , this equation implies codd= 0 and determines the cevenin terms of c0, c2=−c0 2·4, c 4=−c2 4·6=c0 2·42·6, c 6=−c4 6·8=c0 2·42·62·8 c2k=(−1)kc0 2·42·62·82···(2k)2(2k+ 2)=(−1)kc0 22kk!(k+ 1)!. Thus, the formal series we find for the solution, J1(x) , of the Bessel equation of first order corresponding to the largest indicial root, ρ1= 1 is J1(x) =1 2x1(1−x2 2·4+x4 2·42·6− ··· ) or J1(x) =1 2x∞/summationdisplay k=0(−1)kx2k 22kk!(k+ 1)!(6-21) since it is customary to choose the constant c0forJ1(x) asc0=1 2(and the constant c0 forJn(x) as 1/2nn! whennis a positive integer). The other (smaller) root, ρ2=−1 , is much more difficult to treat. If the above steps are imitated (which you should try), division by zero needed to solve for c2fromc0. It turns out that the solution corresponding to the smaller root ρ2=−1 is not of the form (18). We shall not enter into this matter further except to note that the difficulty occurs because the two rootsρ1andρ2differ by an integer . If the two roots ρ1andρ2donot differ by an integer, the above method yields two different solutions of the form (18) for the equation. In any case, this method always gives a solution of the form (18) for the largest root of the indicial equation. It is easy to check that the power series (21) does converge for all xand is therefore a solution to Bessel’s equation of the first order. From the power series, with considerable effort one can obtain a series of identities for Bessel functions which exactly parallels those for the trigonometric functions. The functions Jn(x) behaving in many ways similar to sinnxor cosnx. Here is a graph of J1(x) : 264 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS a figure goes here Forxvery large, J1(x) is asymptotically J1(x)∼/radicalbigg 2 πcos(x−3π/4)√x, which is a cosine curve whose amplitude decreases like 1 /√x. For good reason this curve resembles the height of surface waves on a lake after a pebble has been dropped into the water, or those on the surface of a cup of tea. Having worked out this example in detail, we shall state a definition in preparation for our theorem. Definition: The differential equation a2(x)u/prime/prime+a1(x)u/prime+a0(x)u= 0, where theaj(x) are analytic about x= 0 , it has a regular singularity atx= 0 if it can be written in the form x2u/prime/prime+A(x)xu/prime+B(x)u= 0, where the functions A(x) andB(x) are analytic about x= 0 . Otherwise the singularity isirregular . Examples: (1) .x2(1 +x)u/prime/prime+ 2(sinx)u/prime−exu=−has a regular singularity at x= 0 since the equation may be written as x2u/prime/prime+2(sinx) 1 +xu/prime−−ex 1 +xu= 0, where the coefficients 2 sin x/(1 +x)xandex/1 +xdo have convergent Taylor series aboutx= 0 . (Here we observed thatsinx x= 1−x2 3!+···). (2) .xu/prime/prime−7u/prime+3 cosxu= 0 has a regular singularity at x= 0 since it can be written in the form x2u/prime/prime−7xu/prime+3x cosxu= 0, where the coefficients −7 and 3x/cosxare analytic about x= 0 . (3) .x2u/prime/prime−2u/prime+xu= 0 has an irregular singularity at x= 0 since it cannot be written in the desired form. (4) .x3u/prime/prime−2xu/prime+u= 0 has an irregular singularity at x= 0 . Theorem 6.3 . (Frobenius) Consider the equation with a regular singularity at x= 0 a2(x)u/prime/prime+a1(x)u/prime+a0(x)u= 0, so it can be written in the form x2u/prime/prime+A(x)xu/prime+B(x)u= 0, 6.3. LINEAR EQUATIONS OF SECOND ORDER 265 where the analytic function A(x)andB(x)have convergent power series for |x|<r. Let ρ1andρ2be the roots of the indicial polynomial q(ρ) =ρ(ρ−1) +A(0)ρ+B(0), whereρ1≥ρ2(orReρ 1≥Reρ 2if roots are complex). Then the differential equation has one solution u1(x)of the form u1(x) =xρ1∞/summationdisplay n=0cnxn(c0/negationslash= 0), the series converging for all |x|<r. Moreover, if ρ1−ρ2is not an integer (or zero), there is a second solution u2(x)of the form u2(x) =xρ2∞/summationdisplay n=0˜cnxn(˜c0/negationslash= 0), where this series also converges in the interval |x|< r. In the special case ρ1−ρ2= integer, there may not be a solution of the form (18) - see Exercise 19c. Notice: although the power series do converge at x= 0, the functions u1(x)andu2(x)may not be solutions at that point because the functions xρmay not be twice differentiable (for example, if ρ=1 2 then√xhas no derivatives at x= 0). Outline of Proof . Like Theorem 2, this proof also has two parts; i) finding the coefficients cnfor the formal power series, and ii) proving the formal power series converges. As in Theorem 3, part i) is proved by exhibiting formulas for the cn’s, while part ii) is proved by comparing the series/summationtextcnxnwith another convergent series/summationtextCnxnwhose coefficients are larger, |cn| ≤Cn. To illustrate the procedure of part i), we will obtain the stated formula for the indicial polynomial q(ρ) . LetA(x) =∞/summationdisplay n=0αnxnandB(x) =∞/summationdisplay n=0βnxnbe the power series expan- sions ofA(x) andB(x) . Then assuming u(x) has a solution in the form (18), we find by substituting these formulas into the differential equation that ∞/summationdisplay n=0(ρ+n)(ρ+n−1)cnxρ+n+ (∞/summationdisplay n=0αnxn)(∞/summationdisplay n=0(ρ+n)cnxρ+n) +(∞/summationdisplay n=0βnxn)(∞/summationdisplay n=0cnxρ+n) = 0. The lowest power of xappearing is xρ, then comes xρ+1,.... xρ:ρ(ρ−1)c0+α0ρc0+β0c0= 0 xρ+1: (ρ+ 1)ρc1+ [α1ρc0+α0(ρ+ 1)c1] + [β1c0+β0c1] = 0 · · · xρ+n: (ρ+n)(ρ+n−1)cn+n/summationdisplay k=0αn−k[(ρ+k)ck] +n/summationdisplay k=0βn−kck= 0, 266 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS the last formula arising from the formula for the coefficients in the product of two power series (p. 76). If c0/negationslash= 0 , the first equation states q(ρ) :=ρ(ρ−1) +α0ρ+β0= 0, whereq(ρ) is the indicial polynomial. Since α0=A(0) andβ0=B(0) , this is precisely the formula given in the theorem. c) General Theory We begin immediately by stating Theorem 6.4 (Existence and Uniqueness). Consider the second order linear O.D.E. Lu:=a2(x)u/prime/prime+a1(x)u/prime+a0(x)u=f(x), where the coefficients a0,a1, anda2as well asfare continuous functions, and a2(x)/negationslash= 0. There exists a unique twice differentiable function u(x)which satisfies the equation and the initial conditions u(x0) =α, u/prime(x0) =β, whereαandβare arbitrary constants. If time permits, the existence proof will be carried out in the last chapter as a special case of a more general result. The uniqueness will be proved later too, as a special case of Theorem 9, page 510 - in the next section. We will not be guilty of circular reasoning. Now what? Although this theorem appears to make further study unnecessary, there are several general statements which can be made because the equation is linear . Two other theorems are particularly nice; the first is dim N(L) = 2 , while the second gives a procedure for solving the inhomogeneous equation once two linearly independent solutions of the homogeneous equation are known. A preliminary result on linear dependence and independence of functions is needed. If the differentiable functions u1(x) andu2(x) are linearly dependent, there are constants c1 andc2not both zero such that c1u1(x) +c2u2(x)≡0. Differentiating this equation, we find c1u/prime 1(x) +c2u/prime 2(x)≡0. Since the two homogeneous algebraic equations for c1andc2have a non-trivial solution, by Theorem 32 (page 428), the determinant W(x) :=W(u1,u2)(x) :=/vextendsingle/vextendsingle/vextendsingle/vextendsingleu1(x)u2(x) u/prime 1(x)u/prime 2(x)/vextendsingle/vextendsingle/vextendsingle/vextendsingle= 0 must vanish. This determinant is called the Wronskian ofu1andu2. We have proved 6.3. LINEAR EQUATIONS OF SECOND ORDER 267 Theorem 6.5 . If the differentiable functions u1(x),u2(x)are linearly dependent in the interval [α,β], then necessarily W(x)≡0throughout [α,β]. Thus, if W/negationslash= 0, theuj’s are independent. Remark: The condition W= 0 is necessary for linear dependence but not sufficient in general, as can be seen from the example u1(x) =/braceleftbiggx2, x≥0, 0, x< 0u2(x) =/braceleftbigg0, x≥0 x2, x< 0, for whichW(u1,u2)≡0 for allxbutu1andu2are linearly independent. However it is sufficient if u1andu2are solutions of a second order linear O.D.E., Luj= 0 . An even stronger statement is true in this case. All we need require is that Wvanish at one point x0. Theorem 6.6 . Letu1andu2both be solutions of Lu:=a2u/prime/prime+a1u/prime+a0u= 0, wherea2/negationslash= 0. IfW(x0) = 0 at some point x0, thenu1andu2are linearly dependent - which implies by Theorem 6 that W(x)≡0for allx. In other words, if W(x0)/negationslash= 0, then u1andu2are linearly independent. Proof: SinceW(x0) = 0 , the homogeneous algebraic equations c1u1(x0) +c2u2(x0) = 0 c1u/prime 1(x0) +c2u/prime 2(x0) = 0 have a non-trivial solution c1,c2. Let v(x) =c1u1(x) +c2u2(x). We went to prove v(x)≡0 . Observe Lv= 0 . Moreover v(x0) = 0 andv/prime(x0) = 0 . Thus by uniqueness, v(x)≡0 , establishing the linear dependence of u1andu2. The same type of reasoning proves Theorem 6.7 . LetLu:=a2u/prime/prime+a1u/prime+a0u, wherea2(x)/negationslash= 0. Then dimN(L) = 2. Proof: We exhibit two special solutions φ1andφ2ofLu= 0 and prove they constitute a basis for N(L) . Let φ1(x) satisfy Lφ1= 0 with φ1(x0) = 1, φ/prime 1(x0) = 0 φ2(x) satisfy Lφ2= 0 with φ2(x0) = 0, φ/prime 2(x0) = 1. There are such functions by the existence theorem. 268 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS i) They are linearly independent. W(x0) =W(φ1,φ2)(x0) =/vextendsingle/vextendsingle/vextendsingle/vextendsingleφ1(x0)φ2(x0) φ/prime 1(x0)φ/prime 2(x0)/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 0 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle= 1/negationslash= 0. Thus by Theorem 7, φ1andφ2are linearly independent. ii) They span N(L) . Letu(x) be any element in N(L) and consider the function v(x) =u(x)−[u(x0)φ1(x) +u/prime(x0)φ2(x)]. ThenLv= 0 andv(x0) = 0,v/prime(x0) = 0 . By uniqueness, v(x)≡0 . Thus every u∈N(L) can be written as u(x) =Aφ1(x) +Bφ2(x), where the constants AandBareA=u(x0),B=u/prime(x0) . All of our attention has been on the homogeneous equation Lu= 0 . Let us solve the inhomogeneous equation. This is particularly simple for a linear differential equation once we have a basis for N(L) . Theorem 6.8 (Lagrange). Let u1(x)andu2(x)be a basis for N(L), whereLu:= a2(x)u/prime/prime+a1(x)u/prime+a0(x)u, witha2/negationslash= 0. Then the inhomogeneous equation Lu=f has the particular solution up(x) =u1(x)/integraldisplayxW1(s) W(s)f(s)ds+u2(x)/integraldisplayxW2(s) W(s)f(s)ds, whereW(s) :=W(u1,u2)(s)andWj(s)is obtained from W(s)by replacing the jth column (uj,u/prime j)ofWby the vector (0,1/a2). Remark: If we let G(x;s) =u1(x)W1(s) +u2(x)W2(s) W(s) then the above formula assumes the elegant form up(x) =/integraldisplayx G(x;s)f(s)ds. Proof: A device (due to Lagrange) called variation of parameters is needed. We already used a form of this device to solve the inhomogeneous first order linear equation (5, p. 457). The trick is to let up(x) =v1(x)u1(x) +v2(x)u2(x) where the functions v1(x) andv2(x) are to be found. This attempt to find upis reminiscent of writing the general solution of the homogeneous equation as Au1+Bu2. Differentiate: u/prime p(x) =v1u/prime 1+v2u/prime 2+ [v/prime 1u1+v/prime 2u2]. The functions v1andv2will be chosen to make v/prime 1u1+v/prime 2u2= 0. 6.3. LINEAR EQUATIONS OF SECOND ORDER 269 Using this, we differentiate again u/prime/prime p(x) =v1u/prime/prime 1=v2u/prime/prime 2+ [v/prime 1u/prime 1+v/prime 2u/prime 2] Now multiply u/prime/prime pbya2, u/prime pbya1, upbya0, and add to find Lup=v1Lu1+v2Lu2+a2[v/prime 1u/prime 1+v/prime 2u/prime 2] =a2[v/prime 1u/prime 1+v/prime 2u/prime 2]. If we can choose v1andv2so thata2[ ] =f, then indeed Lup=f, sou0=v1u1+v2u2 is a particular solution. It remains to see if v1andv2can be found which satisfy the two needed conditions v/prime 1u1+v/prime 2u2= 0 v/prime 1u/prime 1+v/prime 2u/prime 2=f a2. These two linear equations for v/prime 1andv/prime 2may be solved by Cramer’s rule (Theorem 33, page 429), v/prime 1=/vextendsingle/vextendsingle/vextendsingle/vextendsingle0u2 f/a 2u/prime 2/vextendsingle/vextendsingle/vextendsingle/vextendsingle W=f/vextendsingle/vextendsingle/vextendsingle/vextendsingle0u2 1/a2u/prime 2/vextendsingle/vextendsingle/vextendsingle/vextendsingle W=W1 Wf v/prime 2=/vextendsingle/vextendsingle/vextendsingle/vextendsingleu1 0 u/prime 1f/a 2/vextendsingle/vextendsingle/vextendsingle/vextendsingle W=f/vextendsingle/vextendsingle/vextendsingle/vextendsingleu10 u/prime 11/a2/vextendsingle/vextendsingle/vextendsingle/vextendsingle W=W2 Wf Integration of these equations yields v1andv2, which, when substituted into up=u1v1+ u2v2, do give the stated result With this theorem, knowing the general solution of the homogeneous equation L˜u= 0 allows us to find a particular solution of the homogeneous equation Lup=f. The general solutionuof the inhomogeneous equation Lu=fis then the upcoset of N(L) , that is, all functions of the form u=up+ ˜u. This puts the burden on finding the general solution of the homogeneous equation. Examples: (1) . The homogeneous equation x2u/prime/prime−3xu/prime+ 3u= 0, x/negationslash= 0 , has the two linearly independent solutions u1(x) =x, u 2(x) =x3—which might have been found by the power series method. Therefore a particular solution of the inhomogeneous equation x2u/prime/prime−3xu/prime+ 3u= 2x4 can be found by the variation of parameters. We try up=v1x3+v2x and are led to the equations v/prime 1=−2x4 x2x3 2x3, v/prime 2=2x4 x2x 2x3 270 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS or v/prime 1=−x2, v/prime 2= 1. Thus v1(x) =−x3 3, v 2(x) =x. Therefore up(x) =x(−x3 3) +x3(x) =2 3x4 The general solution to the inhomogeneous equation is found by adding the general solution of the homogeneous equation to this particular solution, u(x) =Ax+Bx3+2 3x4. (2) The homogeneous equation u/prime/prime+u= 0 has the linearly independent solutions u1(x) = cosx, u 2(x) = sinx. Let us solve u/prime/prime+u=f(x), wherefis an arbitrary continuous function. Trying up(x) =v1cosx+v2sinx, we are led to v/prime 1=−fsinx 1, v/prime 2=fcosx 1. Thus v1(x) =−/integraldisplayx f(s) sinsds, v 2(x) =/integraldisplayx f(s) cossds. Therefore up(x) =−cosx/integraldisplayx f(s) sinsds+ sinx/integraldisplayx f(s) cossds =/integraldisplayx f(s)[−sinscosx+ cosssinx]ds =/integraldisplayx f(s) sin(x−s)ds. Consequently, the handsome formula u(x) =Asinx+Bcosx+/integraldisplayx f(s) sin(x−s)ds is the general solution of the inhomogeneous equation u/prime/prime+u=f. Exercises (1) Solve the following initial value problems any way you can. Check your answers by substituting back into the differential equation. 6.3. LINEAR EQUATIONS OF SECOND ORDER 271 (a)u/prime+ 2u= 0, u(1) = 2 (b)u/prime/prime+ 3u/prime+ 2u= 7, u(0) = 0, u/prime(0) = 0 (c)u/prime/prime+ 3u/prime+ 2u= 2ex, u(0) = 0, u/prime(0) = 1 (d)u/prime/prime+ 3u/prime+ 2u=e−2x, u(0) = 1, u/prime(0) = 0 (e) (tanx)du dx+u−sin2x= 0, u(π 6) = 1 (f)u/prime/prime+u= tanx, u (0) =u/prime(0) = 1, x/epsilon1(−π 2,π 2). (g)u/prime/prime/prime−8u= 0, u(0) = 1, u/prime(0) = 2, u/prime/prime(0) = 3 (h)u/prime/prime/prime/prime−k4u= 0.General solution. (i)u/prime/prime−6u/prime+ 10u=x2+ sinx, u (0) =u/prime(0) = 0. (j)u/prime/prime/prime/prime−7u/prime/prime/prime−8u/prime/prime= 0, u(0) = 3, u/prime(0) = 8, u/prime/prime(0) = 65, u/prime/prime/prime(0) = 511. (k)xu/prime+u=x3, u(1) = 1. (l)u/prime/prime+ 4u= 4x2+ cos 2x, u (0) = 0, u/prime(0) = 1 (m)u/prime/prime/prime−u/prime=ex. General solution. (n)u/prime/prime/prime= 3u/prime/prime+ 3u/prime−u= 0, u(0) = 1, u/prime(0) = 2, u/prime/prime(0) = 3 (o)u(5)−u(4)+ 3u(3)−3u(2)−4u(1)+ 4u= 0 . General solution. [Hint:λ5−λ4+ 3λ3−3λ2−4λ+ 4 = (λ2−1)(λ2+ 4)(λ−1) ]. (2) Find the first four non-zero terms (if there are that many) in the power series solutions aboutx= 0 for the following equations. (a)u/prime/prime−xu/prime−u= 0, u(0) =u/prime(0) = 1 (b)u/prime/prime−2xu/prime+ 2u= 0, u(0) = 0,u/prime(0) = 1. (c)u/prime/prime−2xu/prime−2u= 0, u(0) = 1, u/prime(0) = 0. (d)u/prime/prime+xu= 0, u(0) = 1, u/prime(0) = −1. (e)u/prime/prime/prime−xu= 0, u(0) = 1,u/prime(0) =u/prime/prime(0) = 0. (f)u/prime/prime−x2u=1 1−x2, u(0) = 0, u/prime(0) = 0.[Hint:1 1−x2= 1 +x2+x4+···] (g)u/prime/prime−1 1−xu= 0, u(0) = 0, u/prime(0) = 1.[Hint:1 1−x=?] (3) a) - e) Find where the power series in Ex. 2 a-e converge. (4) Find the first four non-zero terms (if there are that many) in the power series solutions corresponding to the larger root of the indicial polynomial. (a) 2x2u/prime/prime−3xu/prime+ 2u= 0 (b)xu/prime/prime+ 2u/prime−xu= 0.[Answer:u(x) =c0∞/summationdisplay 0x2n (2n+ 1)!] . (c) 4xu/prime/prime+ 2u/prime+u= 0. (d)xu/prime/prime+ (sinx)u/prime+x2u= 0, u(0) = 0, u/prime(0) = 1. (e)xu/prime/prime+u/prime=x2. (5) (a-e). Investigate the convergence of the series solutions found in Exercise 4 above. 272 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS (6) Find the power series solution about x= 0 for the nth order Bessel equation corre- sponding to the highest root of the indicial polynomial. The answer is: Jn(x) = (x 2)n∞/summationdisplay k=0(−1)k k!(k+n)!(x 2)2k, where we have chosen c0= 1/2nn! . (7) Find two linearly independent power series solutions of u/prime/prime+xu/prime+u= 0 and prove they are linearly independent. Find all solutions. (8) The Hermite equation is u/prime/prime−2xu/prime+ 2αu= 0. For which value(s) of the constant αare the solutions polynomials - that is, a solution with a finite Taylor series. These are the Hermite polynomials . (9) Find the first three non-zero terms in the power series about x= 0 for two linearly independent solutions of 2x2u/prime/prime+xu/prime+ (x−1)u= 0. (10) The homogeneous equation Lu:= 2x2u/prime/prime−3xu/prime−2u= 0 has the two linearly inde- pendent solutions u1(x) =x2, u2(x) =√x(see Ex. 20c below). Find the general solution of the inhomogeneous equation Lu= log(x3) . (11) LetLu= (1−x2)u/prime/prime−2xu/prime+n(n+ 1)uwherenis an integer. Show that Lu= 0 has a polynomial solution - the Legendre polynomial. Compute this for n= 3 . (cf. page 104l Ex. 10). (12) LetJ0(x) be a solution of the zerothorder Bessel equation. ProvedJ0 dxis a solution of the first order Bessel equation. [Hint: Work directly with the equation itself, not with power series]. (13) Consider the equation a2(x)u/prime/prime+a1(x)u/prime+a0(x)u= 0. (a) Letu(x) :=u1(x)v(x) . Show that the result arranged as an equation for v(x) is a2u1v/prime/prime+ (2a2u/prime 1+a1u1)v/prime+ (a2u/prime/prime 1+a1u/prime 1+a0u1)v= 0 (b) Ifu1is known to be one solution of the equation, show that the second solution isu2(x) u2(x) =u1(x)/integraldisplay w(x)dx wherew(x) is a solution of the firstorder equation a2u1w/prime+ (2a2u/prime 1+a1u1)w= 0. 6.3. LINEAR EQUATIONS OF SECOND ORDER 273 Thus, if one solution of a second order linear O.D.E. is known, the problem of finding a second solution is reduced to the problem of solving a first order linear O.D.E. - which can always be solved by separation of variables. (14) Apply Exercise 13 to the following: (a) One solution of 2 x2u/prime/prime−3xu/prime+ 2u= 0 isu1(x) =x2. Find another. (b) One solution of x2u/prime/prime−xu/prime+u= 0 isu1(x) =x. Find another. (c) One solution of (1 + x)xu/prime/prime−xu/prime+u= 0 isu1(x) =x. Find another, and then write down the general solution. (d) One solution of the equation x2u/prime/prime+2xu/prime= 0 is clearly u1(x) = 1 . Find another. Prove the solutions are linearly independent for x>0 . Find the general solution ofx2u/prime/prime+ 2xu/prime= 1 . (15) Consider the O.D.E. u/prime/prime+a(x)u/prime+b(x)u= 0 , where aandbare continuous about x0. If the graphs of two solutions are tangent at x=x0, are these two solutions linearly dependent? Explain: Can you make an even stronger deduction? (16) (a) Let Lbe a constant coefficient differential operator with characteristic polyno- mialp(λ) . Ifp(λ) =p(−λ) , prove L(sinkx) =p(ik) sinkx (b) Apply this to find a particular solution of u/prime/prime/prime/prime−u= sin 2x (17) Find a particular solution of the equation u/prime/prime−n2u=f, n /negationslash= 0. [You will need: sin h(α−β) = sinhαcoshβ−sinhβcoshα]. [Answer:u(x) =1 n/integraldisplayx 0f(s) sinhn(x−s)ds.] (18) Use the method of variation of parameters to find a particular solution to u/prime/prime=f. Compare with Exercise 5, p. 282. (19) Consider the differential operator Lu:=x2u/prime/prime+axu/prime+bu, whereaandbare constants. This is called Euler’s equation . It is the simplest equation with a regular singularity at x= 0 . (a) Show that Lxρ=q(ρ)xρ, whereq(ρ) is the indicial polynomial for L. (b) If the roots of q(ρ) = 0 are distinct, find two solutions of Lu= 0, x > 0 , and prove the solutions are linearly independent for x>0 . (c) If the roots ρ1andρ2ofq(ρ) = 0 coincide, take the derivative with respect to ρ of the equation in a) - holding xfixed - to obtain the candidate u2(x) =xρ1lnx for a second solution. Verify by substitution that u2is a solution in this case and prove the two solutions u1(x) =xρ1, u2(x) =xρ1lnx, x> 0 are linearly independent for x/negationslash= 0 . 274 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS (20) Apply the method of Exercise 19 to find two linearly independent solutions for each of the following Euler equations a).x2u/prime/prime+xu/prime= 0. b). 2x2u/prime/prime−3xu/prime+ 2u= 0. c). 2x2u/prime/prime−3xu/prime−2u= 0. d).x2u/prime/prime−xu/prime+u= 0. (21) (a) Use the result of Ex. 19 a) to find a particular solution of the equation Lu=xα, where Lu:=x2u/prime/prime+axu/prime+bu, withaandbconstant, and where αisnota root of the indicial polynomial q(ρ) (cf. Ex. 6, p. 300). (b) If neither αnotβare roots of q(ρ) , find a particular solution to the inhomo- geneous equation Lu=Axα+Bxβ. (c) Apply this procedure to find the general solution of 2x2u/prime/prime−3xu/prime−2u= 3x−4x1/3. (d) How can you solve Lu=xαifαis a root of the indicial polynomial? (22) (a) If uhasnderivatives and λis a constant, prove Dn[eλxu] =eλx(D+λI)nu. Thus (D+λI)nu=e−λxDn[eλxu] . (b) LetL= (D−a)nbe a constant coefficient differential operator with charac- teristic polynomial p(λ) = (λ−a)n. Showu(x) is a solution of the equation Lu= 0 if and only if u(x) has the form u(x) =eaxQ(x), whereQ(x) is a polynomial of degree ≤n−1 . (23) Consider the O.D.E. Lu=f, whereLis a second order constant coefficient operator, and letλ1andλ2be the characteristic roots of L1. Assume i) Reλ 1<0 and Reλ 2<0 , and ii)there is some constant Msuch that |f(x)| ≤Mfor allx∈[0,∞] . (a) Prove every solution of Lu=fis bounded for x∈[0,∞] . (b) If lim x→∞f(x) = 0 , prove that as x→ ∞ , every solution of Lu=ftends to zero. (24) Consider the operator Lu:=a2(x)u/prime/prime+a1(x)u/prime+a0(x)u, where the aj’s are continuous forx∈[α,β] . Letu1,u2andφ1,φ2both be bases for N(L) . Prove there is a constantk/negationslash= 0 such that W(u1,u2)(x) =kW(φ1,φ2)(x) for all x∈[α,β]. 6.3. LINEAR EQUATIONS OF SECOND ORDER 275 (25) (a) Generalize the procedure of Ex. 21b and show how the inhomogeneous Euler equationLu=fcan be solved if fhas a power series expansion. You will have to assume that no root of the indicial polynomial is a positive integer. (b) Apply a) to find a particular solution (as a power series) of 2x2u/prime/prime+ 3xu/prime−u=1 1−x. (26) Given the equation Lu:=u/prime/prime+a(x)u/prime+b(x)u= 0 has solutions u1(x) = sinx, u2(x) = tanx, find the general solution of the inhomogeneous equation Lu=cosx 1 + sin2x. (27) (a) If Lu:=a2u/prime/prime+a1u/prime+a0uandL∗v:= (a2v)/prime/prime−(a1v)/prime+a0v, prove the Lagrange identity vLu−uL∗v=d dx[a2(u/primev−v/primeu) + (a1−a/prime 2)uv], where the functions ajare assumed to be sufficiently differentiable. The operator L∗is the adjoint ofL. (b) Show that Lisself-adjoint ,L=L∗, if and only if a/prime 2=a1. Write the Lagrange identity in this case. (c) Ifc1u1(x) +c2u2(x) is the general solution of the equation Lu= 0 find the general solution of the adjoint equation L∗v= 0 . [Answer: v=c3u1+c4u2 u1u/prime 2−u/prime 1u2] . (d) Letube a twice differentiable function which vanishes at αandβ. Show the adjoint operator L∗has the property that for all such functions uandv, /angbracketleftv, Lu/angbracketright=/angbracketleftL∗v, u/angbracketright where /angbracketleftf, g/angbracketright:=/integraldisplayβ αf(x)g(x)dx. (28) (a) Let Lbe a self-adjoint operator,L=L∗. IfLX 1=λ1X1andLX 2=λ2X2, whereλ1andλ2are real number, λ1/negationslash=λ2, proveX1andX2are orthogonal /angbracketleftX1, X 2/angbracketright= 0. [Hint: Compare /angbracketleftX2, LX 1/angbracketright=λ1/angbracketleftX2, X 1/angbracketrightwith/angbracketleftLX 2, X 1/angbracketright=λ2/angbracketleftX2, X 1/angbracketright]. (b) LetL=d2 dx2. For what values of λcan you find a non-zero solution uof the equationLu=λuwhereusatisfies the boundary conditions u(0) =u(π) = 0 ? (c) Apply parts a) and b) as well a Ex. 27d to prove /angbracketleftsinnx,sinmx/angbracketright=/integraldisplayπ 0sinnxsinmxdx = 0, wherenandmare unequal integers. 276 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS (29) . Consider the boundary value problem Lu:=u/prime/prime+u=f, u (0) = 0, u(π) = 0, wherefis continuous in [0 ,π] . a). Show that if a solution exists, it is not unique. b). Show a solution exists if and only if /integraldisplayπ 0f(x) sinxdx= 0. [Hint: First find the general solution of the homogeneous equation]. Remark: In the notation of Ex. 27, we have L=L∗. Moreover, N(L∗) = span{sinx}. The conclusions of b) states that R(L) =N(L∗)⊥, and illustrates how Theorem 34, p. 431, is used in infinite dimensional spaces. (30) . A proof of Theorem 3. Since a2(x)/negationslash= 0 , the equation can be written as u/prime/prime+a(x)u/prime+b(x)u= 0. If a(x) =∞/summationdisplay n=0αnx2, b (x) =∞/summationdisplay n=0βnxn, let u(x) =∞/summationdisplay n=0cnxn,whereu(0) =c0, u/prime(0) =c1, (a) Imitate the example to prove the remaining cn’s must satisfy cn+2=−n/summationdisplay k=0[αn−k(k+ 1)ck+1+βn−kck] (n+ 2)(n+ 1). Show that if c0andc1are known, then the remaining cn’s are determined inductively by the above formula. (b) Because the series for a(x) andb(x) converge for |x|<r, ifRis any number less thanr, there is a constant Msuch that for all n,|αn| ≤M Rnand|βn| ≤M Rn (cf. p. 72, line 2). Define constants Cnas C0=|c0|, C1=|c1|, and forn≥0 Cn+2=M Rnn/summationdisplay k=0[(k+ 1)Ck+1+Ck]Rk+MC n+1R (n+ 2)(n+ 1). (i) Prove |cn| ≤Cn, n = 0,1,2,3,... 6.3. LINEAR EQUATIONS OF SECOND ORDER 277 (ii) Prove/vextendsingle/vextendsingle/vextendsingle/vextendsingleCn+1xn+1 Cnxn/vextendsingle/vextendsingle/vextendsingle/vextendsingle=n(n−1) +MnR +MR2 R(n+ 1)n|x|. (iii) Prove∞/summationdisplay n=0Cnxnconverges for |x|< R, whereRis any number less than r. (iv) Prove∞/summationdisplay n=0cnxnconverges for |x|< R, whereRis any number less than r. (31) . (a) Letu(x) andv(x) be solutions of the equations L1u:=u/prime/prime+a(x)u= 0 , andL2v:=v/prime/prime+b(x)v= 0 respectively, in some interval, where aandbare continuous. If b(x)≥a(x) throughout the interval, prove there must be a zero ofvbetween any two zeroes of u. This is the Sturm oscillation theorem . [Hint: Supposeαandβare consecutive zeroes of uandu>0 in (α,β) . Prove 0 =/integraldisplayβ α(vL1u−uL2v)dx=vu/prime/vextendsingle/vextendsingleβ α−/integraldisplayβ α(b−a)uvdx, and show, because u/prime(α)>0, u/prime(β)<0 , there is a contradiction if vdoes not vanish somewhere in ( α,β) .] (b) Letu1(x) andu2(x) be two linearly independent solutions of u/prime/prime+a(x)u= 0 . Prove between any two zeroes of u1, there is a zero of u2and vice verse. Thus, the zeroes interlace. (c) Apply b) to the solutions sin γxand cosγxof the equation u/prime/prime+γ2u= 0 to conclude a well-known fact. (d) Ifb(x)≥δ >0 , whereδis a constant, prove every solution of v/prime/prime+b(x)v= 0 must have an infinite number of zeros by comparing vwith a solution of u/prime/prime+γ2u= 0 , where γis an appropriate constant. (e) Apply d) to prove every solution of v/prime/prime+ (1−3 4x2)v= 0, has an infinite number of zeroes for x≥1 . (f) Letu1(x) be a solution of the first order Bessel equation. Take v(x) =u1(x)√x and show that vsatisfies the equation in e). Deduce that J1(x) has infinitely many zeroes. (32) LetL1andL2be linear constant coefficient differential operators with characteristic polynomials p1(λ) andp2(λ) respectively. (a) If there is a function u(x), u(x)/negationslash≡0 , which satisfies both L1u= 0 andL2u= 0 , prove the polynomials p1andp2have a common root. (b) Ifp1andp2have no common roots, prove the solution of L1L2u= 0 are exactly all functions of the form c1u1+c2u2whereu1is a solution of L1u1= 0 , andu2 ofL2u2= 0 . Thus N(L1L2) may be decomposed into the two complementary subspaces N(L1) and N(L2),N(L1L2) =N(L1)⊕N(L2) . 278 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS (33) Imitate Exercise 30 and prove Theorem 3. Make sure to observe the trouble in trying to find the solution corresponding to the lower root of the indicial polynomial if the roots differ by an integer. (34) The purpose of this exercise is to show that an equation with an irregular singular point may have a formal power series at that point which does not converge to the solution. Try to find a solution of the form (18) for the following equation which has an irregular singularity at x= 0 , x6u/prime/prime+ 3x5u/prime−4u= 0. What happened? Two linearly independent solutions for x/negationslash= 0 are u1(x) =e−1/x2andu2(x) =e1/x2. How does this explain the situation (cf. p. 95-6)? (35) Consider the equation 2 x2u/prime/prime+ 3xu/prime+u=/radicalbig (x) . Two linearly independent solutions of the homogeneous equation are x−1/2andx−1. Find the general solution of the homogeneous equation. (36) Consider the equation u/prime/prime+b(x)u/prime+c(x)u= 0 , where bandcare continuous functions and c(x)<0 . Prove that a solution cannot have a positive maximum or negative minimum. 6.4 First Order Linear Systems Quite often in applications you must consider systems of differential equations. We shall consider a linear system of the form du1 dx+a11(x)u1+a12(x)u2+···+a1n(x)un=f1(x) (6-22) du2 dx+a21(x)u1+a22(x)u2+···+a2n(x)un=f2(x) (6-23) ...... (6-24) dun dx+an1(x)u1+an2(x)u2+···+ann(x)un=fn(x), (6-25) where the functions aij(x) andfj(x) are continuous. If we anticipate the next chapter and write the derivative of a vector U= (u1,...,u n) as the derivative of its components, d dxU(x) =/parenleftbiggdu1 dx,du2 dx,···,dun dx/parenrightbigg , then the above system can be written in the clean form dU dx+A(x)U=F(x), (6-26) where, A(x) = ((aij)), F = (f1,f2,...,f n) 6.4. FIRST ORDER LINEAR SYSTEMS 279 and U(x) = (u1,u2,...,u n). The initial value problem for the system of differential equations (22) is to find a vector U(x) which satisfies the equation as well as the initial condition U(x0) =U0, (6-27) whereU0is a vector of constants. It is useful to observe that the initial value problem for a single linear equation of order n u(n)+an−1(x)u(n−1)+···+a0(x)u=f(x) u(x0) =α1, u/prime(x0) =α2,...,u(n−1)(x0) =αn, can be transformed to the conceptually simpler problem (22)-(23). Let u1(x) :=u(x) , u2(x) :=u/prime(x),..., andun(x) =u(n−1)(x) . Then the components of the vector U(x) = (u1,u2,...,u n) must obviously satisfy the relations du1 dx=u2 du2 dx=u3 · · · dun−1 dx=un dun dx=−a0u1−a1u2− ··· −an−1un+f(x), which may be written as U/prime=MU+F, where M(x) = 0 1 0 ··· 0 0 0 1 ··· 0 ...... 0 0 0 ··· 1 −a0−a1−a2··· −an−1 , and F= (0,0,..., 0,f). The initial conditions read U(x0) = (α1,α2,...,α n). Conversely, if Uis any solution of this system of equations with the proper initial conditions, then the first component u1(x) is a solution of the single nth order equation. Thus, the general theory of a single nth order linear O.D.E. is completely subsumed as a portion of the theory of a system of first order linear O.D.E.’s. You should be warned that this generalization is mainly of theoretical value and is of little use if you are seeking an explicit solution. Both the existence and uniqueness theorems are true for systems, and supply an example where the theoretical advantages of systems become clear. To illustrate this, we shall prove the uniqueness theorem. Our proof is patterned directly after the uniqueness proof for a single equation (Theorem 1). 280 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS Theorem 6.9 (Uniqueness). Let A(x)be a matrix whose coefficients aij(x)are bounded |aij(x)| ≤Mforxin some interval, and let F(x)be a continuous function. Then there is at most one solution U(x)of the initial value problem U/prime+AU=F, U (x0) =U0. Remark: The existence theorem states, if Ais nonsingular and each element is integrable there is at least one solution. Thus, there is then exactly one solution. Proof: AssumeU1andU2are both solutions. Let W=U1−U2. ThenWsatisfies the homogeneous equation and is zero at x0, W/prime+AW= 0, W (x0) = 0. Take the scalar product of this with W, /angbracketleftW, W/prime/angbracketright+/angbracketleftW, AW /angbracketright= 0. But /angbracketleftW, W/prime/angbracketright=w1w/prime 1+ww/prime 2+···+wnw/prime n =1 2d dx(w2 1+w2 2+···+w2 n) =1 2d dx/bardblW/bardbl2. Thus, 1 2d dx/bardblW/bardbl2=−/angbracketleftW, AW /angbracketright. By Theorem 17, p. 173 and the hypothesis |aij(x)| ≤M, we know /vextendsingle/vextendsingle/angbracketleftW, AW /angbracketright/vextendsingle/vextendsingle≤/bracketleftBign/summationdisplay i,j=1|aij|2/bracketrightBig1/2 /bardblW/bardbl2≤nM/bardblW/bardbl2. so that 1 2d dx/bardblW/bardbl2≤nM/bardblW/bardbl2. Therefore, as on p. 462-3 d dx(/bardblW/bardbl2)−2nM/bardblW/bardbl2≤0, or e2nMxd dx[e−2nMx/bardblW/bardbl2]≤0. Becausee2nMxis always positive, by the mean value theorem the quantity [ ] is a decreasing function. Its value for x>x 0is then less than at x0, e−2nMx/bardblW(x)/bardbl2≤e−2nMx 0/bardblW(x0)/bardbl2, x ≥x0 6.4. FIRST ORDER LINEAR SYSTEMS 281 Consequently /bardblW(x)/bardbl ≤enM(x−x0)/bardblW(x0)/bardbl, x ≥x0. SinceW(x0) = 0 and the norm is non negative, we have 0≤ /bardblW(x)/bardbl ≤0, x ≥x0, which implies /bardblW(x)/bardbl= 0, x ≥x0. Therefore, W(x)≡0x≥x0. By replacing xwith−xin the original equation, the same statement is true for x≤x0. Thus, throughout the interval where |aij(x)| ≤M, we have proved W(x)≡0 , that is, U1(x)≡U2(x) , so the solution is indeed unique. Because a single linear nth order O.D.E. can be replaced by an equivalent system of equations, this theorem implies the uniqueness theorem for a single O.D.E. of order nif the coefficients aj(x) are bounded in some interval - which is certainly true in every interval if theaj’s are continuous. With this theorem, a short section closes. Further developments in the theory of systems of linear O.D.E.’s make elegant use of linear operators in general and matrices in particular. As you might well accept, the exercises contain a few of the more accessible results. Exercises (1) . Find functions u1(x),u2(x) which satisfy u/prime 1=u1 u/prime 2=u1−u2, with the initial conditions U(0) := (u1(0),u2(0) = (1,0) . Find the general solution too. [Hint: Solve the equation u/prime 1=u1first, then substitute. Answer: General solution is U(x) = (γ1ex,γ1 2ex+γ2e−x) ]. (2) Consider the system u/prime 1= 2u1−u2 u/prime 2= 3u1−2u2, that is, U/prime=AU, whereA=/parenleftbigg2−1 3−2/parenrightbigg . Letφ1(x) =au1+bu2, φ2(x) =cu1+du2, wherea,b,c anddare constants. Thus, Φ =SU, where S=/parenleftbigga b c c/parenrightbigg , Φ = (φ1,φ2). 282 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS (a) By direct substitution, find the differential equations satisfied by the φj’s and show they can be written as Φ/prime=SAS−1Φ. (b) Pick the coefficients of Sso the matrix SAS−1is a diagonal matrix, SAS−1=/parenleftbiggλ10 0λ2/parenrightbigg ≡Λ (c) Solve the resulting equation Φ/prime= ΛΦ . [Solution: φ1=αex, φ 2=βe−x—you might have φ1andφ2interchanged]. (d) Use this solution to solve the original equations for U. [hint: RecallU= S−1Φ ]. (3) By only a slight modification of Exercise 2, solve v/prime/prime 1= 2v1−v2 v/prime/prime 2= 3v1−2v2. [Hint: Everything, even the algebra, is identical. The only difference is in part c) you have to solve Φ/prime/prime= ΛΦ . Then V=S−1Φ as before]. (4) A bathtub initially contains Q1gallons of gin and Q2gallons of vermouth, where Q1+Q2=Q, Q being the capacity of the tub. Pure gin enters from one faucet at a constant rate of R1gallons per minute, while pure vermouth enters from another faucet at a constant rate R2gallons per minute. The well stirred mixture of martinis leaves the drain at a rate R1+R2gallons per minute (so the total amount of fluid in the tub remains constant at Qgallons). Let G(t) be the quantity of gin in the tub at timetandV(t) be the quantity of vermouth. (a) Show dG dt=R1−G Q(R1+R2) dV dt=R2−V Q(R1+R2). (b) Integrate this simple system of equations to find G(t) andV(t) . Also find their ratioP(t) :=G(t)/V(t) which is the strength of the martinis at time t. (c) Prove lim t→∞P(t) =R1 R2. Compare this with your intuitive expectations. (d) IfQ1= 20,Q2= 0,R1=R2= 1 gal/min, how long must I wait to get a perfect martini (for me, perfect is 5 parts gin to 1 part vermouth). [Needless to say, the mathematical model is applicable to many problems in the mixing of chemicals which do not react with each other. If the chemicals do interact, the model must be changed to account for the interaction]. 6.5. TRANSLATION INVARIANT LINEAR OPERATORS 283 (5) Consider the homogeneous equation U/prime=A(x)U, whereAis non-singular (so detA/negationslash= 0 ). Assuming the validity of the existence theorem, prove there exists nlinearly independent vectors U1(x),U2(x),...,U n(x) which are solutions, U/prime k= AUk, k= 1,...,n . [Hint: Construct nsolutions which are linearly independent at x=x0, and then prove a set of nsolutions are linearly independent in an interval if and only if they are linearly independent at x=x0, wherex0is a point in the interval]. (6) LetLU:=U/prime−A(x)Uas in Exercise 5. Prove dim N(L) =n. (7) LetLU:=U/prime−A(x)U. If a basis U1,...,U n, forN(L) is known, prove the inho- mogeneous equation LU=Fcan be solved by variation of parameters. That is, seek a particular solution UpofLU=Fin the form Up=n/summationdisplay i=1Uivi where thevi(x) are scalar-valued functions ( notvectors). (a) Compute U/prime pand substitute into the O.D.E. to conclude Upis a particular solution ifn/summationdisplay i=1Uiv/prime i=F. (b) LetUbe then×nmatrix whose columns are U1,U2,...,U n. ProveUis invertible and show v/prime i(x) = (U−1F)ith component. (c) Show Up(x) =n/summationdisplay i=1Ui(x)/integraldisplayx [U−1(s)F(s)]ids. This may also be written in the form Up(x) =U(x)/integraldisplayx U−1(s)F(s)ds (d) Apply this procedure to find the general solution of u/prime q=u1+e2xcf. Ex 1 u/prime 2=u1−u2+ 1. 6.5 Translation Invariant Linear Operators This section develops various extensions and applications of the procedure used to solve linear ordinary differential equations with constant coefficients. The results will be proved as a series of exercises interspersed by various remarks. Definition: Thetranslation operator Ttacting on functions u(x) is defined by the prop- erty (Ttu)(x) =u(x−t). x,t ∈R. 284 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS A linear operator Listranslation invariant if LTt=TtL for everyt, that is, if L(Ttu) =Tt(Lu) for everytand for every function ufor which the operators are defined. Example: 1 Let (Lu)(x) := 3u(x)−2u(x−1) . Then [Tt(Lu)](x) = 3u(x−t)−2u(x−t−1), and [L(Ttu)](x) =Lu(x−t) = 3u(x−t)−2u(x−t−1). Thus, LTt=TtL, so the operator Lis translation invariant. 2. Let (Lu)(x) := 3xu(x).Then [Tt(Lu)](x) = 3(x−t)u(x−t), and [L(Ttu)](x) =Lu(x−t) = 3xu(x−t). Thus LTt/negationslash=TtL, so this operator is nottranslation invariant. Exercises (1) Which of the following linear operators (verify!) are also translation invariant? (a) (Lu)(x) :=cu(x), c ≡constant (b) (Lu)(x) :=u(x+h)−u(x) h, h≡constant /negationslash= 0 . (c) (Lu)(x) :=/integraldisplayx −∞k(x−s)u(s)ds (d) (Lu)(x) := (x−1)u(x) (e) (Lu)(x) =du dx(x). (f) Any linear ordinary differential operator with constant coefficients, Lu:=anu(n)+an−1u(n−1)+···+a0u, a kconstants. (g) Any linear ordinary differential operator with variable coefficients. (h) (Lu)(x) =n/summationdisplay k=1aku(x−γk), a kandγkconstants. [Answers: All but d) and g) are translation invariant]. 6.5. TRANSLATION INVARIANT LINEAR OPERATORS 285 (2) IfL1andL2are translation invariant operators which map some linear space into itself, then so are a).AL1+BL 2, A,B constants b).L1L2andL2L1 c). If in addition Lis invertible, then L−1is also translation invariant. Theorem 6.10 . IfLis a translation invariant linear operator, then L(eλa) =φ(λ)eλx. Proof: We know so little about Lthat all we can hope to do is compute TtL(eλx) andLTt(eλx) and see what happens. Let Leλx=ψ(λ;x) , whereψis some unknown function whose value depends on both λandx. Then TtL(eλx) =ψ(λ;x−t), while LTteλx−Leλ(x−t)=L(e−λteλx) =e−λtLeλx=e−λtψ(λ;x). SinceTtL=LTt, we find e−λtψ(λ;x) =ψ(λ;x−t), or ψ(λ;x) =ψ(λ;x−t)eλt. Because the left side does not contain t, the right side must not depend on which value oftis chosen. Using this freedom, we let t=xand conclude ψ(λ;x) =ψ(λ; 0)eλx. By setting φ(λ) =ψ(λ,0) , we find Leλx=ψ(λ;x) =φ(λ)eλx as desired. Exercises (3) By direct substitution, find φ(λ) for those operators in Exercise 1 which are trans- lation invariant. [Answers: a) φ(λ) =c, b)φ(λ) = (e−ah−1)/hc)φ(λ) =/integraldisplay0 −∞k(−s)eλsds, d)φ(λ) =cλ, f)φ(λ) =n/summationdisplay k=0akλk(the characteristic polynomial), h)φ(λ) =n/summationdisplay k=1ake−λγk]. 286 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS (4) With the same assumptions and notation as in the theorem, if φ(λ) = 0 is a poly- nomial equation with Ndistinct roots λ1,λ2,...,λ N, soφ(λj) = 0, j= 1,...,N , prove any linear combination of the function eλjxis inN(L) , that is, Lu= 0 where u(x) =N/summationdisplay 1cjeλjx. (5) Apply the theorem to find the solution of Exercise 4 for the equation Lu= 0 , where (a)Lu:=u/prime/prime−u/prime−u. (b) (Lu)(x) =u(x+ 2)−u(x+ 1)−u(x) . (c) Find a special solution of b) which satisfies the “initial conditions” u(0) =u(1) = 1 . Compute u(2),u(3) andu(4) directly from b). The integers u(n), n∈Z+ are called the Fibonacci sequence . [Answer: u(2) = 2,u(3) = 3,u(4) = 5 , and surprisingly , u(n) =1√ 5 /parenleftBigg 1 +√ 5 2/parenrightBiggn+1 −/parenleftBigg 1−√ 5 2/parenrightBiggn+1 ]. (6) Solveu(x)−au(x−1) +b2u(x−2) = 0 with the initial conditions u(1) =a,u(2) = a2−b2. Compare with Exercise 17, p. 440. (7) Extend Exercises 5(b - c) and 6 to develop a theory of second order difference equations with constant coefficients . Thus Lu:=a2u(x+ 2) +a1u(x+ 1) +a0u(x), a 2/negationslash= 0, x ∈Z. In particular, you should, (a) Find two linearly independent solutions of Lu= 0 . Remember the degenerate casea2 1−4a0a2= 0 . (b) Prove there is at most one solution of the initial value problem Lu=f,u(0) = α0,u(1) =α1. (c) Prove dim N(L) = 2 . Remarks: The ideas presented above generalize immediately to the case where X∈Rn instead of just R1, as well as to the case where the u’s are vectors and not scalars. These few concepts lie at the heart of any treatment of many linear operators with constant coefficients, especially ordinary and partial differential operators. This mildly abstract formulation manages to penetrate through the obscuring details of particular cases to observe a rather simple structure unifying many seemingly different problems. 6.6 A Linear Triatomic Molecule A molecule composed of three atoms is called a triatomic . Consider a triatomic molecule whose equilibrium configuration is a straight line with two atoms of equal mass msituated on either side of a central atom of mass M. 6.6. A LINEAR TRIATOMIC MOLECULE 287 a figure goes here To simplify the situation further, we shall only consider the motion along the straight line (axis) of these atoms, and shall assume the inter-atomic forces can be approximated by springs with equal spring constants k.u1(t),u2(t) andu3(t) will denote the displacements of the atoms (see fig.) from their equilibrium position. Newton’s second law, m¨u=/summationtextF, will give the equations of motion. The atom on the left only “feels” the force due to the spring attached to it, the force being equal to the spring constant ktimes the amount that spring is stretched, u2−u1. Thus, m¨u1=k(u2−u1). The central atom “feels” two forces, one from each side, with the resulting equation of motion M¨u2=−k(u2−u1) +k(u3−u2). In the same way, the equation of motion for the remaining atom is m¨u3=−k(u3−u2). Collecting our equations, we have ¨u1=−k mu1+k mu2 ¨u2=k Mu1−2k Mu2+k Mu3 ¨u3=k mu2−k mu3. These are a system of three linear ordinary differential equations with constant coefficients. They cannot be integrated as they stand since each equation involves functions from the other equations, that is, the equations are copied (not surprising since we are considering coupled oscillators . Now we can integrate such a system immediately if they are in the simple form ¨φ1=λ1φ1 ¨φ2=λ2φ2 ¨φ3=λ3φ3 by integrating each equation separately. By using an important method, we will be able to place our system in this special form. Before doing so, it is suggestive to rewrite the system in matrix form  ¨u1 ¨u2 ¨u3 = −k mk m0 k M−2k Mk M 0k m−k m  u1 u2 u3 . LettingAdenote the 3 ×3 matrix, our hope is to somehow change Ainto a diagonal matrix (one with zeroes everywhere except along the principal diagonal), for then the differential equations will be in a form mentioned above which can be immediately integrated. 288 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS The trick is to replace the basis u1,u2,u3by some other basis in which the matrix assumes a diagonal form. The differential equation can be written in the form ¨U=AU, whereU= (u1,u2,u3) , and the derivative of a vector being defined as the derivative of each of its components. Let φ1(t),φ2(t) , andφ3(t) be three other functions - which we plan to use as a new basis. Then the φj’s can be written as a linear combination of the uj’s, φ=s11u1+s12u2+s13u3 φ2=s21u1+s22u2+s23u3 φ3=s31u1+s32u2+s33u3, wheresijareconstants . WritingS= ((sij)) and Φ = ( φ1,φ2,φ3) , this last equation reads Φ =SU. Taking the derivative of both sides (or going back to the equations defining φjin terms of theuk’s), we find ¨Φ =S¨U. Because both u1,u2andu3as well asφ1,φ2, andφ3are bases for the solution, the matrix Smust be non-singular (its inverse expresses the φ/prime jsin terms of the uj’s). Thus ¨Φ =SAS−1Φ. The problem has been reduced to finding a matrix Ssuch that the matrix SAS−1is a diagonal matrix , SAS−1= λ10 0 0λ20 0 0λ3 ≡Λ. Multiply by S−1on the left: AS−1=S−1Λ. Since this equation is equally between matrices, their corresponding columns must be equal. Thus, if we denote by ˆSi, theith column of S−1, the above equation then reads AˆSi=λiˆSi, or (A−λiI)ˆSi= 0. For eachithis is a system of three linear algebraic equations for the three components of ˆSi. If there is to be a non-trivial solution, we know det(A−λiI) = 0. Since det(A−λiI) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle−k m−λik m0 k M−2k M−λik M 0k m−k m−λi/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle, 6.6. A LINEAR TRIATOMIC MOLECULE 289 (algebra later) =−λi(k m+λi)[λi+ (2 M+1 m)k] We see the three possible values of λfor det(A−λiI) = 0 are λ1= 0,λ2=−k m, λ 3=−k(2 M+1 m). These numbers λiare the eigenvalues of A. The non-trivial solution ˆSiof the homo- geneous equations ( A−λiI)ˆSi= 0 corresponding to the ith eigenvalue is called the eigenvalue ofAcorresponding to the eigenvalue λi. For example, ˆS2is the solution of (A−λ2I)S2= 0 corresponding to λ2=−k/m, 0ˆs12+k mˆs22+ 0ˆs32= 0 k Mˆs12−(2k M−k m)ˆs22+k Mˆs32= 0 0ˆs12+k mˆs22+ 0ˆs32= 0. We see ˆs22= 0 while ˆs12=−ˆs32. Thus, one solution is ˆS2= (1,0,−1) Similarly we find one solution for ˆS1is ˆS1= (1,1,1), while one solution for ˆS3is ˆS3= (1,−2m M,1). The computation is over. All that remains is to put the parts together and interpret the solution. If you got lost, presumably this recapitulation will help. We have found a transformation Sto new coordinates ( φ1,φ2,φ3) such that the differential equations for theφj’s are in diagonal form, ¨φm=λjφj, ¨φ1= 0 ¨φ2=−k mφ2 ¨φ3=−k(2 M+1 m)φ3. The solutions are φ1(t) =A1+B1t, φ2(t) =A2cos/radicalbigg k mt+B2sin/radicalbigg k mt φ3=A3cos/radicalbigg k(2 M+1 m)t+B3sin/radicalbigg k(2 M+1 m)t. 290 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS Since Φ =SU, and the ˆSjare the columns of S−1, S−1= 1 1 1 1 0 −2m M 1−1 1 , we haveU=S−1Φ , u1(t) =φ1(t) +φ2(t) +φ3(t) u2(t) =φ1(t) −2m Mφ3(t) u3(t) =φ1(t)−φ2(t) +φ3(t) Although the solutions φ1(t),φ2(t) , andφ3(t) can now be substituted into the first set of equations for the uj’s, it is more instructive to leave that step to your imagination and analyze the nature of the solution. (1) Ifφ1(t)/negationslash= 0 but φ2(t) =φ3(t) = 0,then u1(t) =u2(t) =u3(t) =A1+B1t. Thus all three atoms - the whole molecule - moves with a constant velocity B1. This is the trivial translation motion of the molecule, simply moving without internal oscillations at all. (2) Ifφ2(t)/negationslash= 0 butφ1(t) =φ3(t) = 0 , then u1(t) =φ2(t) =−u3(t),andu2(t) = 0. Thus, the two outside atoms vibrate in opposite directions with frequency/radicalbig k/m while the center atom remains still: a figure goes here (3) Ifφ3(t)/negationslash= 0 butφ1(t) =φ2(t) = 0 u1(t) =u3(t) =φ3(t), u 2(t) =−2m Mφ3(t). A bit more complicated. The two outside atoms move in the same direction with same frequency/radicalBig k(2 M+1 m) , while the center atom moves in a direction opposite to them and with the same frequency but a different amplitude (to conserve linear momentum m˙u1+ M˙u2+m˙u3= 0 ). In the figure we take m=M. a figure goes here These three simple motions are called the normal modes of oscillation of the molecule. They are the oscillations determined by the φ1,φ2, andφ3. Every motion of the system is a linear combination of the normal modes of oscillation, the particular oscillation depending on what initial conditions are given. By an appropriate choice of the initial conditions, one or another of the normal modes will result. Otherwise some less recognizable motion will result. Exercises Consider the simpler model of a diatomic molecule 6.6. A LINEAR TRIATOMIC MOLECULE 291 a figure goes here which we will represent as two masses joined by a spring with spring constant k. (a) Show the equations of motion are m¨u1=k(u2−u1) M¨u2=−k(u2−u1) (b) Introduce new variables, Φ = SU, φ1=s11+s12u2 φ2=s21u1+s22u2, and findSso that the equation ¨Φ =SAS−1Φ is in diagonal form. (c) Solve the resulting equation and find the normal modes of oscillation. Interpret your results with a diagram. 292 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS Chapter 7 Nonlinear Operators: Introduction 7.1 Mappings from R1toR1, a Review . The subject of this section is one you presumably know well. Our intention is to briefly review the more important results, stating them in a form which suggests the generalizations we intend to develop. Consider a function y=f(x), x∈R. This function assigns to each number xanother real number y. Thus we may write f:R→R. fis a scalar-valued function of a scalar. What are the simplest such functions? Linear ones of course, f(x) =ax+b. In keeping with our more sophisticated terminology, this should be called an “affine” func- tion (mapping, operator, ...) since it is linear only if b= 0 . We shall, however, be abusive and refer to such functions as linear mappings. The study of linear functions in one variable, x, is carried out in elementary analytic geometry. At an early age we enlarged our vocabulary of functions from linear ones to a more general class which includes, for example, f1(x) =ax2+bx+c, f 2(x) = sinx, f 3(x) =√x. These functions are all examples of nonlinear functions. They map the reals (only the positive reals in the case of f3) into the reals. The portion of the reals for which they are defined is called their domain of definition ,D(f) . Thus D(f1) =R1,D(f2) =R1,D(f3) ={x∈R1:x>0}. The class of all real valued functions of a real variable is too large to consider. For most purposes it is sufficient to restrict oneself to the class of continuous or sufficiently differentiable functions. Here is an outline of the basic definitions and theorems from elementary calculus. In our prospective generalization from the simplest case of a function (operator) fwhich maps numbers to numbers, f:R1→R1, to the case of a function from vectors to vectors f:Rn→Rm, all of these concepts and results will need to be extended. 293 294 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION Definition: akconverges to a, a kanda∈R1. Definition: Continuity. Theorem 7.1 The set of continuous functions forms a linear space. Definition: The derivative: limit of difference quotient. Theorem 7.2 1.d dx(af+bg) =ad f dxf+bdg dx(linearity) 2.d dx(fg) =fdg dx+ (d f dx)g(Product rule) 3.d dx(f◦g) =d f dgdg dx(Chain rule) Theorem 7.3 The Mean Value Theorem. Definition: The integral. Theorem 7.4 1./integraldisplayb af(x)dx=−/integraldisplaya bf(x)dx 2./integraldisplayb 1f(x)dx+/integraldisplayc bf(x)dx=/integraldisplayc af(x)dx 3./integraldisplayb a[αf(x) +βg(x)]dx=α/integraldisplayb af(x)dx+β/integraldisplayb ag(x)dx(linearity) 4./integraldisplayb a(f◦φ)(x)dφ dxdx=/integraldisplayφ(b) φ(a)f(x)dx(Change of variable in an integral) Theorem 7.5 1./integraldisplayb adf dx(x)dx=f(b)−f(a) 2.d dx/integraldisplayx af(t)dt=f(x) 3./integraldisplayb af(x)dg dxdx=fg/vextendsingle/vextendsingleb a−/integraldisplayb adf dxg(x)dx(Integration by parts). Remark: These theorems contain essentially all of elementary calculus. What are missing are specific formulas for the derivatives and integrals of the basic functions as well as the application of these theorems to compute maxima, area, etc. Exercises (1) Use the definition of the derivative (as the limit of a difference quotient) to compute the derivatives of the following functions at the given point. a). 3x2−x+ 1, x0= 2 b).1 x+1, x 0= 2 c).x 1+x, x 0= 2 d).x 1−x, x =x0/negationslash= 1. 7.2. GENERALITIES ON MAPPINGS FROM RNTORM. 295 (2) Use the definition of the integral to evaluate /integraldisplay2 0x2dx. You should approximate the area by rectangular strips and evaluate the limit as the width of the thickest strip tends to zero. [Hint: 12+22+32+···+n2=n(n+1)(2 n+1) 6]. (3) Prove that .6<log 2<.8 (log 2 = 0 .693) by using the definition of the integral to find upper and lower bounds for log 2 =/integraldisplay2 11 xdx. (4) Find the equation of the straight line which is tangent to the curve f(x) =x7/3+ 1 atx= 1 . Draw a sketch indicating both the curve and tangent line. Use the tangent line to approximately evaluate (1 .01)7/3. Find some estimate for the error in your approximation. 7.2 Generalities on Mappings from RntoRm. A function, or operator, Fwhich maps RntoRm, is a rule which assigns to each vector XinRnanother vector Y=F(X) inRm. It is a function from vectors to vectors, a vector-valued function of a vector. We have already discussed the case when Fis an affine operator, Y=F(X) =b+LX or in coordinates, y1=b1+a11x2+···+a1nx1n y2=b2+a21x2+···+a2nxn · · · ym=bm+am1x2+···+amnxn Linear algebra can be thought of as the study of higher dimensional analytic geometry, the affine transformations taking the role of the straight line y=b+cx. But now it is time to consider more complicated mappings from RntoRm. Here is an Example:/braceleftbiggy1=x1+x2sinπx3 y2=e1−x1−√x2. This transformation maps vectors X= (x1,x2,x3)∈R3to vectorsY= (y1,y2)∈R2. Note the second function is only defined for x2≥0 . Thus the domain of the transformation F is D(F) =/braceleftbig X∈R3:x2≥0/bracerightbig . For example, Fmaps the point (1 ,4,1 6) into the point (3 ,−1) . 296 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION It is usual to write a transformation Fwhich maps a set A⊂Rnto a setB⊂Rmin terms of its components , y1=f1(x1,...,x n) =f1(X) y2=f2(x1,...,x n) =f2(X) · · · ym=fm(x1,...,x n) =fm(X), or more concisely as Y=F(X). To discuss continuity etc. for nonlinear mappings from RntoRm, it is necessary that the distance between points be defined. We shall use the Euclidean norm - although any other norm could also be used. If X= (x1,...,x k) is a point (or vector, if you like) in Rk, then/bardblX/bardbl=/radicalBig x2 1+···+x2 k. To review briefly, a sequence of pointsXjinRkconverges to a pointXinRkif, given any /epsilon1>0 , there is an integer Nsuch that /bardblXj−X/bardbl</epsilon1 textforall j ≥N. Anopen ball inRkof radiusrabout the point X0is the setB(X0;r) ={X∈Rk:/bardblX− X0/bardbl<r}. Aclosed ball inRkis ¯B(X0;r) ={X∈Rk:/bardblX−X0/bardbl ≤r}. The only difference is the open ball does not contain the boundary of the ball. In two dimensions, R2, the names open and closed discare often used. A setD⊂Rkisopen if each point X∈Dis the center of some ball contained entirely withinD. The radius may be very tiny. Every open ball is open, as can be seen in the figure. A closed ball is not open since there is no way of placing a small ball about a point on the boundary in such a way that the small ball is inside the larger one. A set Ais closed if it contains allof its limit points , that is, if the points Xj∈Aconverge to a point X, X j→X, thenXis also inA. An open ball is not closed, for a sequence of points in the ball may converge to a point on the boundary, and the boundary points are not in the ball. For the special case of R1, these notions coincide with those of open and closed intervals. Again, sets - like doors - may be neither open nor closed. A point set Disbounded if it is contained in some ball (of possibly large radius). The pointXisexterior toDifXdoes not belong to Dand if there is some ball about X none of whose point are in D.Xisinterior toDifXbelongs to Dand there is some ball about Xall of whose points are in D.Xis aboundary point ofDif it is neither interior nor exterior to D. Note that a boundary point of Dmay or may not belong to D. For example, the boundaries of the open and closed balls B(0;r),¯B(0;r) are the same. The boundary of a set Dis denoted by ∂D. It is evident that a set is open if and only if every point is an interior point, and a set is closed if and only if it contains all of its boundary points. Definition: LetAbe a set in RnandCa set in Rm. The function F:A→Cis continuous at the interior point X0∈Aif, given any radius /epsilon1>0 , there is a radius δ>0 7.2. GENERALITIES ON MAPPINGS FROM RNTORM. 297 such that /bardblF(X)−F(X0)/bardbl</epsilon1 textforall /bardblX−X0/bardbl<δ. [Observe the norm on the left is in Rmwhile that on the right is in Rn]. It is easy to prove Theorem 7.6 . An affine mapping F(X) =b+LX from RntoRmis continuous at every point X0∈Rm. Proof: First,F(X)−F(X0) =b+LX−b−LX 0=L(X−X0) . Thus, /bardblF(X)−F(X0)/bardbl=/bardblL(X−X0)/bardbl. Let ((aij)) be a matrix representing Lwith respect to some bases for RnandRm. Then by Theorem 17, p. 373 /bardblL(X−X0)/bardbl2=/angbracketleftL(X−X0), L(X−X0)/angbracketright ≤k/bardblL(X−X0)/bardbl/bardblX−X0/bardbl, where k2=m/summationdisplay i=1n/summationdisplay j=1a2 ij. Therefore /bardblL(X−X0)/bardbl ≤k/bardblX−X0/bardbl. It is now clear that if X→X0, thenL(X−X0)→0 . More formally, given any /epsilon1>0 , if δ=/epsilon1 k+1, we have /bardblF(X)−F(X0)/bardbl</epsilon1 textforall /bardblX−X0/bardbl<δ. The following theorems have the same proofs as were given earlier for special cases. (See a first year calculus book and our Chapter 0). Theorem 7.7 . LetF1andF2mapA⊂RnintoC⊂Rm. IfF1andF2are continuous at the interior point X0∈A, then 1.aF1+bF2is continuous at X0. 2./angbracketleftF1, F2/angbracketrightis continuous at X0. Theorem 7.8 . LetF= (f1,...,f m)mapA⊂RnintoC⊂Rm. ThenFis continuous at the interior point X0∈Aif and only if each of the fj, j= 1,...,m , is continuous at X0. Theorem 7.9 . LetF:A→C, whereAis a closed and bounded (= compact) set. If F is continuous at every point of A, then it is bounded; that is, there is a constant Msuch that/bardblF(x)/bardbl ≤Mfor allX∈A. Moreover, if M0is the least upper bound, then there is a pointX0∈Asuch that /bardblF(X0)/bardbl=M0. Similarly, if m0is the greatest lower bound for/bardblF/bardbl, then there is a point X1∈Asuch that /bardblF(X1)/bardbl=m0. There is nothing better than to close this otherwise unauspicious section with one of the crown jewels of mathematics - the Fundamental Theorem of Algebra, all of whose proofs require the non-algebraic notion of continuity. Let p(z) =a0+a1z+···+anzn, (n≥1), where theaj’s are complex numbers and anis not zero. For every complex number z, the value of the function p(z) is a complex number. Thusp:C→C. We want to prove there is at least one z0∈Csuch thatp(z0) = 0 . 298 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION Lemma 7.10 .p(z)is a continuous function for every z∈C. Proof: Identical to the proof that a real polynomial is continuous everywhere. Lemma 7.11 LetDbe a set in the complex plane in which p(z)/negationslash= 0. The minimum modulus of p(z), that is, the minimum value of |p(z)|, cannot occur at an interior point ofD. It must occur on the boundary ∂DofD. Proof: Letz0be any interior point of D. Rewritep(z) in the form p(z) =b0+b1(z−z0) +···+bn(z−z0)n. Sincep(z0)/negationslash= 0 , we know b0/negationslash= 0 . Also, because pis not identically constant, at least one coefficient following b0is not zero. Take bkto be the first such coefficient. We must write b0, bkandz−z0in polar form, b0=ρ0euαbk=ρ1eiβz−z0=ρeiθ, whereρ0=|p(z0)|,ρ1andρare positive real numbers. Here we are restricting zto a point on a circle of radius ρaboutz0, after taking ρsmall enough to insure this circle is interior to D. Then p(z) =ρ0eiα+ρ1eiβρkeikθ+bk+1(z−z0)k+1+···+bn(z−z0)n =ρ0eiα+ρ1ρkei(β+kθ)+ (z−z0)k+1[bk+1+···+bn(z−z0)n−k−1]. Pick the particular point ˆ zon the circle whose argument θis given by β+kθ=α+π. Thenei(β+kθ)=ei(α+π)=−eiα, so p(ˆz) = (ρ0−ρ1ρk)eiα+ (ˆz−z0)k+1[bk+1+···+bn(ˆz−z0)n−k−1]. By the triangle inequality we find |p(ˆz)| ≤/vextendsingle/vextendsingle/vextendsingleρ0−ρ1ρk/vextendsingle/vextendsingle/vextendsingle+ρk+1[|bk+1|+···+|bn|ρn−k−1]. Choose the radius ρso small that ρ0−ρ1ρk≥0 . Then |p(ˆz)| ≤ρ0−ρ1ρk+ρk+1[|bk+1|+···+|bn|ρn−k−1]. By choosing ρsmaller yet, if necessary, we can make the term ρ[|bk+1|+···+|bn|ρn−k−1]< 1 2ρ1. Consequently, |p(ˆz)| ≤ρ0−ρ1ρk+1 2ρ1ρk=ρ0−1 2ρ1ρk <ρ 0=|p(z0)|. Thus, ifz0is any interior point of a domain Din whichpdoes not vanish, then there is a point ˆzalso interior to Dsuch that |p(z)|<|p(z0)|. Therefore, the minimum of |p(z)| must occur on the boundary of any set in which pdoes not vanish. Lemma 7.12 . Given any real number M, there is a circle |z|=Ron which |p(z)|>M for allz,|z|=R. 7.2. GENERALITIES ON MAPPINGS FROM RNTORM. 299 Proof: Forz/negationslash= 0 , we can write the polynomial p(z) as p(z) zn=an+an−1 z+···+a0 zn. From the triangle inequality written in the form |f1+f2| ≥ |f1| − |f2|, we find /vextendsingle/vextendsingle/vextendsingle/vextendsinglep(z) zn/vextendsingle/vextendsingle/vextendsingle/vextendsingle≥ |an| −/vextendsingle/vextendsingle/vextendsinglean−1 z+···+a0 zn/vextendsingle/vextendsingle/vextendsingle. If|z|is taken large enough, |z| ≥R0, it is possible to make the second term on the right less than |an|/2 , /vextendsingle/vextendsingle/vextendsinglean−1 z+···+a0 zn/vextendsingle/vextendsingle/vextendsingle</vextendsingle/vextendsingle/vextendsinglean 2/vextendsingle/vextendsingle/vextendsingle, texton |z|=R>R 0 Therefore, for |z|=R≥R0 /vextendsingle/vextendsingle/vextendsingle/vextendsinglep(z) zn/vextendsingle/vextendsingle/vextendsingle/vextendsingle≥ |an| −/vextendsingle/vextendsingle/vextendsinglean 2/vextendsingle/vextendsingle/vextendsingle=1 2|an|, so |p(z)| ≥1 2|an|Rn, texton |z|=R. It is now clear that by choosing Rsufficiently large, |p(z)|can be made to exceed any constantMon the circle |z|=R. Theorem 7.13 (Fundamental Theorem of Algebra). Let p(z) =a0+a1z+···+anzn, a n/negationslash= 0,n≥1, be any polynomial with possibly complex coefficients, a0,a1,...,a n. Then there is at least one number z0∈Csuch thatp(z0) = 0 . In other words, every polynomial has at least one complex root. Proof: By Lemma 3, we can find a large circle |z|=R, on which |p(z)|>2|a0|for all |z|=R. Sincep(z) is a continuous function, by Theorem 4 there is a point z0in the closed and bounded disc |z| ≤Rfor which |p|attains its minimum value m0,|p(z0)|=m0. If p(z0) = 0 , we are done. However if pdoes not vanish inside the closed disc, by the important Lemma 2 its minimum value is attained only on the boundary, so z0is on the circle |z0|=R. But on the circle we know |p(z0)|>2|a0|= 2|p(0)|, so the minimum is not atz0after all. The assumption that pdoes not vanish in the disc |z| ≤Rhad led us to a contradiction. Notice the proof does not give a procedure for finding the root whose existence has been proved. Exercises 1. Prove Theorem 2, part 1. 2. Use the Fundamental Theorem of Algebra along with the “factor theorem” of high school algebra to prove that a polynomial of degree nhas exactly nroots (some of which may be repeated roots). 300 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION 7.3 Mapping from E1toEn . As a particle moves along a curve γinEnits position F(t) at timetcan be specified by a vector X=F(t) = (f1(t),f2(t),...,f n(t)), wherexj=fj(t) is thejthcoordinate of the position at time t. Thus, the curve is specified by F(t) , a mapping from numbers to vectors, F:A⊂E1→En, whereAis the domain of definition of F. For example, the mapping F:t→(cosπt,sinπt,t), t∈(−∞,∞) which may also be written as F(t) = (cosπt,sinπt,t) can be thought of as describing the motion of a particle along a helix. It is natural to ask about the velocity, which means derivative must be defined. Definition: LetF(t) define a curve γfortin the interval A= [a,b] . Consider the difference quotient F(t+h)−F(t) h, t textand t +hinA, wheretis fixed. If this vector has a limit as htends to zero, then Fis said to have a derivativeF/prime(t) att, F/prime(t) = lim h→0F(t+h)−F(t) h, while the curve has slopeF/prime(t) att. Some other common notations are ˙F(t),dF dt, D tF. The curve γis called smooth if i) the derivative F/prime(t) exists and is continuous for each t in [a,b] , and if ii) /bardblF/prime(t)/bardbl /negationslash= 0 for any point tin [a,b] . Iftrepresents time, then F/prime(t) is the velocity of the particle at time twhile /bardblF/prime(t)/bardbl is the speed . IfF(t) is given in terms of coordinate functions, F:t→(f1(t),...,f n(t)) , how can the derivative of Fbe computed? Theorem 7.14 . IfF(t) = (f1(t),...,f n(t))is a differentiable mapping of A⊂E1into En, then the coordinate functions are differentiable and dF dt=/parenleftbiggdf1 dt,df2 dt,···,dfn dt/parenrightbigg . Conversely, if the coordinate functions are differentiable, then so is F(t)and the derivative is given by the above formula. 7.3. MAPPING FROM E1TOEN301 Proof: Iftandt+hare both in A, then F(t+h)−F(t) h=1 h[(f1(t+h),...,f n(t+h))−(f1(t),...,f n(t))] =/parenleftbiggf1(t+h)−f1(t) h,···,fn(t+h)−fn(t) h/parenrightbigg Since the limit as h→0 of the expression on the left exists if and only if all of the limits lim h→0fj(t+h)−fj(t) h, j = 1,...,n exist, the theorem is proved. Examples: (1) IfF:t→(cosπt,sinπt,t), t∈(−∞,∞), F is differentiable for all tsince each of the coordinate functions are differentiable. Also, F/prime(t) = (−πsinπt,π cosπt,1). In addition, the curve - a helix - which Fdefines is smooth since F/primeis continuous and /bardblF/prime(t)/bardbl −/radicalbig π2sin2πt+π2cos2πt+ 1 =/radicalbig π2+ 1/negationslash= 0, (2) LetF:t→(a1+b1t,a2+b2t,a3+b3t) =P+QtwhereP= (a1,a2,a3) and Q= (b1,b2,b3) are constant vectors. Then the curve Fdefines is a straight line which passes through the point P= (a1,a2,a3) att= 0 .Fis differentiable for all t, since each of the coordinate functions are differentiable. Furthermore, F/prime(t) =Q= (b1,b2,b3), aconstant vector pointing in the direction Q= (b1,b2,b3) , as is anticipated for a straight line. Because /bardblF/prime(t)/bardbl=/bardblQ/bardbl=/radicalBig b2 1+b2 2+b2 3, this curve is smooth except in the degenerate case b1=b2=b3= 0 , that is, Q= 0 , when the curve degenerates to a single point, F(t) = (a1,a2,a3) =P. (3) The curve defined by the mapping F:t→(t,|t|) is differentiable everywhere and /bardblF/prime(t)/bardbl /negationslash= 0 except at t= 0 . It is not differentiable there since the second coordinate function,f2(t) =|t|is not differentiable at t= 0 . Thus, the curve is smooth except att= 0 . (4) The curve defined by the mapping F:t→(t3,t2) is differentiable everywhere, and F/prime(t) = (3t2,2t). However, /bardblF/prime(t)/bardbl=√ 9t4+ 4t2, so the curve is smooth everywhere except at t= 0 , which corresponds to a cusp at the origin in the x1,x2plane. 302 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION It is elementary to compute the derivative of the sum of two vectors. The derivative of a product can be defined for the inner product, and for the product with scalar-valued function. Theorem 7.15 . IfF(t)andG(t)both map an interval A⊂E1intoEn, and are both differentiable there, then for all t∈A, 1.d dt[aF+bG] =adF dt+bdG dt(linearity of the derivative). 2.d dt/angbracketleftF, G/angbracketright=/angbracketleftF/prime, G/angbracketright+/angbracketleftF, G/prime/angbracketright, (in “dot product” notation:d dt(F·G) =F/prime·G+F·G/prime). Proof: Since these are identical to the proofs of the corresponding statements for scalar- valued functions, we prove only the second statement. d dt/angbracketleftF(t), G(t)/angbracketright= lim h→01 h[/angbracketleftF(t+h), G(t+h)/angbracketright − /angbracketleftF(t), G(t)/angbracketright] = lim h→01 h[/angbracketleftF(t+h)−F(t), G(t+h)/angbracketright+/angbracketleftF(t), G(t+h)−G(t)/angbracketright] = lim h→0/bracketleftbigg/angbracketleftF(t+h)−F(t), h/angbracketright G(t+h)+/angbracketleftF(t),G(t+h)−G(t) h/angbracketright/bracketrightbigg =/angbracketleftF/prime(t), G(t)/angbracketright+/angbracketleftF(t), G/prime(t)/angbracketright. An interesting and simple consequence is the fact that if a particle moves on a curve F(t) which remains a fixed distance from the origin, /bardblF(t)/bardbl ≡ constant = c, then the velocity vector F/primeis always orthogonal to the position vector F. This follows from c2=/bardblF(t)/bardbl2=/angbracketleftF(t), F(t)/angbracketright, so taking the derivative of both sides we find 0 =/angbracketleftF/prime, F/angbracketright+/angbracketleftF, F/prime/angbracketright= 2/angbracketleftF, F/prime/angbracketright. Thus /angbracketleftF, F/prime/angbracketright= 0 for all t, an algebraic statement of the orthogonality. As a particular example, the mapping F(t) = (cosπ 1 +t2,sinπ 1 +t2) has the property /bardblF(t)/bardbl= 1 for all t. You can see the path of the particle in the figure. Att= 0 the particle is at ( −1,0) . As time increases, the particle moves along an arc of the unit circle toward (1 ,0) , reaching (0 ,1) att= 1 . The velocity at time tis F/prime(t) =2πt (1 +t2)2(sinπ 1 +t2,−cos−π 1 +t2). From this expression, it is evident the particle slows down as it approaches (1 ,0) . In fact, the particle never does manage to reach (1 ,0) . We would like to define the notion of a straight line which is tangent to a smooth curve at a given point. There is one touchy issue. You see, the curve may intersect itself, thus having two or more tangents at the same point. Once acknowledged, the difficulty is resolved by realizing that for each value of t, there is a unique point F(t) on the curve. X0is a double point if F(t1) =F(t2) =X0. 7.3. MAPPING FROM E1TOEN303 By picking one value of t, there will be a unique tangent line to the curve for this value oft. Thus, we define the tangent line for t=t1to the curve defined by a differentiable functionF(t) as the straight line whose equation is A(t) =F(t1) +F/prime(t1)(t−t1). Att=t1, the curves defined by F(t) andA(t) have the same value F(t1) =X0and the same derivative (slope), F/prime(t) . Example: Consider the curve defined by the mapping F:t→(3 +t3−t,t2−t), t∈ (−∞,∞) . The point (3 ,0) is a double point since F: 0→(3,0) andF: 1→(3,0) . Thus, the line tangent to the point (3 ,0) whent= 1 is defined by A(t) = (3,0) + (2,1)(t−1) = (3,0) + (2(t−1),(t−1)) or A(t) = (1,−1) + (2t,t). Since we are still working with functions F(t) of one real variable t, the mean value theorem and chain rule follow immediately by applying the corresponding theorems for scalar valued functions to each of the components f1(t),...,f n(t) ofF(t) . Theorem 7.16 (Approximation Theorem and Mean Value Theorem). If the vector valued functionF(t)is continuous for t∈[a,b]and differentiable for t∈(a,b)then fort0∈ (a,b), 1.F(t) =F(t0) +dF dt/vextendsingle/vextendsingle t0(t−t0) +R(t,t0)|t−t0|where lim t→t0/bardblR(t,t0)/bardbl= 0. 2. There is a point τbetweentandt0such that /bardblF(t)−F(t0)/bardbl ≤ /bardblF/prime(τ)/bardbl|t−t0|. 3. IfF= (f1,...,f n), there are points τ1,...,τ nbetweentandt0such that F(t) =F(t0) +L(t−t0), whereLis the linear transformation L= (f/prime 1(τ1),f/prime 2(τ2),...,f/prime n(τn)) Remark: Although 1 and 3 follow from the one variable case f(t) —and will be proved again in greater generality later on - the proof of 2 is difficult under our weak hypothesis. If the stronger assumption, Fis continuously differentiable, is made, then 2 becomes easy, and the factor /bardblF/prime(τ)/bardblcan be replaced by a constant M= max τ∈[a,b]/bardblF/prime(τ)/bardbl, since a continuous function /bardblF/prime(τ)/bardbldoes assume its maximum if τis in a closed and bounded set, τ∈[a,b] . Corollary 7.17 . IfFsatisfies the hypotheses of Theorem 8 and if F/prime(t)≡0for all t∈[a,b], thenFis a constant vector. Proof: Just look at 2 or 3 above to see that for any points t,t0in [a,b] , we have F(t) =F(t0) . 304 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION Theorem 7.18 (Chain Rule). Consider the vector-valued function F(t)which is differ- entiable for t∈(a,b), and the scalar valued function φ(s)which is differentiable for s∈(α,β). If the range of φis contained in (a,b),R(φ)⊂(a,b), then the composed func- tionG(s) = (F◦φ)(s) =F(φ(s))is differentiable as a function of sfor allsin(α,β) and G/prime(s) =F/prime(φ(s))φ/prime(s), that is, dG ds(s) =dF dφ(φ)dφ ds(s) =dF dt(t)/vextendsingle/vextendsingle/vextendsingle/vextendsingle t=φ(s)dφ ds(s). IfF(t) = (f1(t),...,f n(t)), then G(s) =F((s)) = (f1(φ(s)),...,f n(φ(s)),and G/prime(s) =)f/prime 1(φ)φ/prime(s),...,f/prime n(φ)φ/prime(s)) = (f/prime 1(φ),...,f/prime n(φ))φ/prime(s). Proof not given here . It is the same as that given in elementary calculus for n= 1 . A more general theorem containing this one is proved later (p. 701). Examples: 1. IfF(t) = (1 −t2,t3−sinπt) andφ(s) =e−s, thenG(s) = (F◦φ)(s) = (1 − e−2s,e−3w−sinπe−2) . We compute G/prime(s) in two distinct ways, using the chain rule, and directly from the formula for G(s) . By the chain rule: G/prime(s) =F/prime(t)/vextendsingle/vextendsingle t=φ(s)φ/prime(s) = (−2t,3t2−πcosπt)/vextendsingle/vextendsingle t=e−s(−)e−s =−(−2e−s,3e−2s−πcosπe−s)e−s, In particular, at s= 0 , since t= 1 whens= 0 , we find G/prime(0) = −(−2,3 +π) = (2,−3,−π) Directly from the formula for G(s) = (1 −e−2s,−3e−3s−sinπe−s),we find G/prime(s) = (2e−2s,−3e−3s+πe−scosπe−2), which agrees with the chain rule computation. Since the derivative F/prime(t) of a function F(t) from numbers to vectors, F:E1→En, is also a function of the same type, the second and higher order derivatives can be defined inductively; d2 dt2F(t) :=d dtF/prime(t),dk+1 dtk+1F(t) :=d dtF(k)(t). 7.3. MAPPING FROM E1TOEN305 Example: IfF:t→(cosπt,sinπt,t) , then F/prime/prime(t) =d dt(−πsinπt,π cosπt,1) = (−π2cosπt,−π2sinπt,0). IfF(t) represents the position of a particle at time t, thenF/prime/prime(t) is the acceleration of the particle at time t. All of these ideas were used in the last two sections in Chapter 6 where linear systems of ordinary differential equations were encountered. Time permitting, a second application to a non-linear system of O.D.E.’s will be treated in Section of Chapter. There another of the crown jewels in the intellectual history of mankind will be discussed: Newton’s incredible solution of “the two body problem”, that is, to determine the motion of the heavenly bodies. Recall that the length of a curve is defined to be the limit of the lengths of inscribed polygons which approximate the curve as the length of the longest subinterval tends to zero - if the limit does exist. Let the curve γ, which we assume is smooth, be determined by the function F(t),t∈[a,b] . Then the length of the straight line joining F(tj) to F(tj+ ∆tj),tj+1=tj+ ∆tj, is /bardblF(tj+ ∆tj)−F(tj)/bardbl=/bardblF(tj+ ∆tj)−F(tj) ∆tj/bardbl∆tj Adding up the lengths of these segments and letting the largest ∆ tjtend to zero, we find the length of γis given by L(γ) =/integraldisplayb 1/bardblF/prime(t)/bardbldt. If the function Fis defined through coordinates, F(t) = (f1(t),...,f n(t)) , this formula reads L(γ) =/integraldisplayb a/radicalBig f/prime2 1+f/prime2 2+···+f/prime2ndt. You will recognize the special case where F(t) = (x(t),y(t)) L(γ) =/integraldisplayb a/radicalbig x2+ ˙y2dt. Example: Find the length of the portion of the helix γdefined by F(t) = (cost,sint,t) , fort∈[0,2π] . This is one “hoop” of the helix. Since F/prime(t) = (−sint,cost,1) , we have /bardblF/prime(t)/bardbl=/radicalbig sin2t+ cos2t+ 1 =√ 2 , so the length is L(γ) =/integraldisplay2π 0√ 2dt= 2π√ 2. For eacht∈[a,b] , we can define an arc length function s(t) , the arc length from ato t, by s(t) =/integraldisplayt a/bardblF/prime(τ)/bardbldτ. Note we are using a dummy variable of integration τ. By the fundamental theorem of calculus, we have ds dt=/bardblF/prime(t)/bardbl 306 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION Sinceds/dt can be thought of as the rate of change of arc length with respect to time, it is the speed of a particle moving along the curve , the tangential speed . The integral used in arc length is the integral of a scalar- valued function /bardblF/prime(t)/bardbl. How can we define the integral of a vector-valued function F(t) = (f1(t),...,f n(t)) ? Just integrate each component, assuming they are all integrable of course, /integraldisplayb aF(t)dt:= (/integraldisplayb af1(t)dt,...,/integraldisplayb afn(t)dt). For example, if F(t) = (t−3t2,1−√ 2t,e3t) , then /integraldisplay2 0F(t)dt= (/integraldisplay2 0(t−3t2)dt,/integraldisplay2 0(1−√ 2t)dt,/integraldisplay2 0e3tdt) = (−4,−2 3,e6−1). We give no physical interpretation of the integral (as an area or the like) except in the case whereF(t) represents the velocity of a particle. Then/integraldisplayb aF(t)dtis the vector pointing from the position at t=ato the position at t=b. Exercises (1) (a) Describe and sketch the images of the curves F:E1→E2defined by (i)F(t) = (2t,3−t) (ii)F(t) = (2t,|3−t|) (iii)F(t) = (t2,1 +t2) (iv)F(t) = (2t,sint) (v)F(t) = (t2,1 +t4) (b) Which of the above mappings are differentiable and for what value(s) of t? Find the derivatives if the functions are differentiable. Which of the curves defined by these mappings are smooth, and where are they not smooth? (2) Use the definition of the derivative to find F/prime(t) att= 2πfor the functions a).F(t) = (2t,3−t)t∈(−∞,∞). b).F(t) = (1 +t2,sin 2t). t∈(−∞,∞). (3) Find the lengths of the curves γdefined by the mappings a).F(t) = (a1+b1t,a2+b2t,...,a n+bnt),=P+Qt, t ∈[0,1]. b).F(t) = (sin 2t,1−3t,cos 2t,2t3/2),t∈[−π,2π] (4) Consider the curve defined by the equation F(t) = (t−t2,t4−t2+ 1), t∈(−∞,∞) a). Sketch the curve. b). Where does the curve intersect itself? c). Find the line tangent to the curve at the image of t= 1 . 7.3. MAPPING FROM E1TOEN307 (5) IfF:A⊂E1→Enis twice continuously differentiable and F/prime/prime(t)≡0 for allt∈A, what can you conclude? Please prove your assertion. [Hint: First consider the special case where F:E1→E1]. (6) LetF(t) be a twice differentiable function which maps a set in E1intoEnand satisfies the ordinary differential equation F/prime/prime+µF/prime+kF= 0 , where kandµare positive constants. Define the energy as E(t) =1 2/bardblF/prime/bardbl2+1 2k/bardblF/bardbl2 (a) Prove E(t) is a non-increasing function of t(energy is dissipated). [Hint: dE/dt =? ]. (b) IfF(0) = 0 and F/prime(0) = 0 , prove E(t)≡0 . (c) Prove there is at most one function which satisfies the given differential equation as well as the initial conditions F(0) =A, F/prime(0) =B, whereAandBare given vectors. (7) IfF(t) = (1 −e2t,t3,1 1+t2) , andφ(x) =1 1+x, x > −1 , computed dx(F◦φ)(x) by using the chain rule. (8) Compute d2F/dt2for the function F(t) in Exercise 7. (9) (a) Show that the equation of a straight line which passes through the point P1at t= 0 andP2att= 1 is F(t) =P1+ (P2−P1)t. (b) Find the equation of a straight line which passes through the point P1= (1,2,3) att= 0 andP2= (1,−5,0) att= 1 . (c) Find the equation of a straight line which passes through the point P1att=t1 andP2att=t2. (d) Apply this to find the equation of a straight line which passes through P1= (−3,1,−2) att=−1 andP2= (0,2,1) att= 2 . What is the slope of this line? (10) Given a smooth curve all of whose tangent lines pass through a given point, prove that the curve is a straight line. (11) LetF:E1→Endefine a smooth curve which does not pass through the origin. Show that the position vector F(t) is orthogonal to the velocity vector at the point of the curve which is closest to the origin. Apply this to prove anew the well known fact that the radius vector to any point on a circle is perpendicular to the tangent vector at that point. [Hint: Why is it sufficient to minimize ϕ(t) =/angbracketleftF(t), F(t)/angbracketright?] 308 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION Chapter 8 Mappings from Ento E: The Differential Calculus 8.1 The Directional and Total Derivatives . Throughout this and the next chapter we shall consider functions which map Enor a portion of it A, into E. By the statement f:A→E, A ⊂En we mean that to every vector XinA, the function (operator, map, transformation) assigns a unique real number w. Thusw=f(X) in this case is a map from vectors to numbers. Two particular examples prove helpful in thinking conceptually about mappings of this type. (1)The temperature function .f:A→E, where the set A⊂E3is the room in which you are sitting. To every point Xin the room, A, this function fassigns a number - the temperature f(X) atX, w =f(X) . (2)The height function .f:A→E, where the set Ais some set in the plane E2. To every pointXinA, this function fassigns a number - the height f(X) of a surface (or manifold)Mabove that point. Thus, the set of all pairs ( X,f(x)), X∈A, defines a portion of a surface, a surface in E2×E∼=E3. From the second example, it is clear that every function f:A⊂En→Emay be regarded as the graph of a surface in En×E∼=En+1, the surface being regarded as all points in En+1of the form ( X,f(X)) , where X∈A. For example, the temperature function can be thought of as the graph of a surface in E4, the height of the surface w=f(X) aboveXbeing the temperature at X. (Compare with the discussion from p. 322 bottom, to p. 324). In concrete situations, the point X∈Enis specified by giving its coordinates with respect to some fixed bases for EnandE. The particular coordinate system used depends on the geometry of the problem at hand. Rectangular symmetry calls for the standard 309 310CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS rectangular coordinates, while polar coordinates are well suited to problems with circular symmetry. We shall meet these issues head-on a bit later. IfX= (x1,...,x n) with respect to some coordinates for En, then we write w= f(X) =f(x1,...,x n) . The points ( X,f(X)) on the graph are (x1,...,x n, f(x1,...,x n)) , which we may also write as ( x1,...,x n,f) or else as ( x1,...,x n,w) . For low dimensional spaces, E2orE3, it is convenient to avoid subscripts. In these situations we shall write w=f(x,y) andw=f(x,y,z ) for mappings with domains in E2andE3, respectively. We now examine some more specific examples. Examples: (1)w=−1 2x+y−1 . This function assigns to every point X= (x,y) inE2a numberw inE. We can represent the function, an affine mapping from E2→E, as the graph of a plane in E3. The linear nature of the plane reflects the fact that the mapping is an affine mapping - a linear mapping except for a translation of the origin. More generally, the function w=α+a1x1+a2x2+···+anxn, an affine mapping from En→E, represents a plane in En+1. In fact, this can be taken as the algebraic definition of a plane in En+1. These affine functions are the simplest functions which mapEnintoE. Although we shall not, it is customary to abuse the nomenclature and refer to affine mappings as being linear. This is because they share most of the algebraic and geometric properties of proper linear mappings, as opposed to the honestly nonlinear mappings we will be treating as in the next examples. (2)w=x2+y2. This function assigns to every point X= (x,y) inE2a real number w∈E. We can represent the function as the graph of a paraboloid of revolution, obtained by rotating the parabola w=x2about the waxis. If this paraboloid is cut by a plane parallel to the x,yplane, say w= 2 , the intersection of these two surfaces is the circle x2+y2= 2 . (3)w=−x2+y2. This function can be represented as the graph of a very fancy surface - ahyperbolic paraboloid . If this surface is cut by a plane parallel to the x,yplane, w=c, the intersection is the curve c=−x2+y2. Forc>0 , this curve is a hyperbola which opens about the yaxis, while if c<0 , the curve is a hyperbola which opens about the xaxis. Forc= 0 we obtain two straight lines, x=+−y(see fig). The intersection of the surface with the plane x=cis a parabola which opens upward in they,w plane. Similarly, the intersection of the surface with the plane y=cis a parabola which opens downward in the xwplane. This curve is rightly called a saddle , and the origin (0 ,0,0) a saddle point (or mountain pass) since a particle can remain at rest at that point, or ii) move on the surface in one direction and go up, or iii) move on the surface in another direction and go down. Letf(X) be a function from vectors to numbers, f:A⊂En→E. How can we define the notion of derivative for such functions? The derivative should measure the rate of change of f(X) asXmoves about. But if you think of f(X) as the temperature function, it is clear that the temperature will change at different rates depending which direction you move. Thus, if you move across the room in 8.1. THE DIRECTIONAL AND TOTAL DERIVATIVES 311 the direction of the door, the temperature may decrease, while if you move up to the ceiling, the temperature will likely increase. Thus, the natural notion of a derivative is the rate of change in a particular direction - a directional derivative . LetX0denote your position and f(X0) the temperature there. Take ηto be a free vector, which we shall think of as pointing from X0toX0+η. We want to define the rate at which the temperature changes as you move from X0in the direction ηtowardX0+η. Since all points on the line joining X0toX0+ηare of the form X0+λη, whereλis a real number, the difference f(X0+λη)−f(X0) is the difference between the temperatures atX0+ληand atX0. Definition: Letf:A⊂En→E. The derivative offat the interior point X0∈Awith respect to the vector ηis f/prime(X0;η) = lim λ→0f(X0+λη)−f(X0) λ, if the limit exists. In the special case when η=eis aunit vector ,/bardble/bardbl= 1 , we see that λ=/bardblλe/bardbl. Then Def(X0) :=f/prime(X0;e) is the instantaneous rate of change of fper unit length asXmoves fromX0towardX0+ 3 . This normalization to using only unit vectors is necessary to have a meaningful definition of a directional derivative. Thus, the directional derivative of fatX0∈Ain the direction of the unit vector eis the derivative with respect to the unit vectore. It measures how fchanges as you move from X0to a point on the unit sphere aboutX0. For theoretical purposes, the derivative of fwith respect to any vector ηis useful, while for practical purposes, the more restrictive notion of the directional derivative is needed. Example: 1 Find the directional derivative of f(X) =x2 1−2x1x2+ 3x2atX0= (1,0) in the direction η= (−1,1) . Note that ηis not a unit vector. The unit vector is e=η /bardblη/bardbl= (−1√ 2,1√ 2) . Then X0+λe= (1,0) +λ(−1√ 2,1√ 2) = (1 −λ√ 2,λ√ 2), so f(X0+λe) = (1 −λ√ 2)2−2(1−λ√ 2)(λ√ 2) + 3(λ√ 2) = 1−λ√ 2+3 2λ2. Thus, f(X0+λe)−f(X0) λ=1−λ√ 2+3 2λ2−1 λ=−1√ 2+3 2λ. Therefore, the directional derivative Defis Def(X0) = lim λ→0f(X0+λe)−f(X0) λ=−1√ 2. In words, the rate of change of fatX0in the direction of the unit vector eis−1√ 2. One qualitative conclusion we arrive at is that f(X) decreases as Xmoves from X0in the directione. 312CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS 2. Compute f/prime(X;η) iff(X) =/angbracketleftX, AX /angbracketright, whereAis a self-adjoint transformation. f(X+λη) =/angbracketleftX+λη, A (X+λη)/angbracketright =/angbracketleftX, AX /angbracketright+λ/angbracketleftη, AX /angbracketright+λ/angbracketleftX, Aη /angbracketright+λ2/angbracketleftη, Aη/angbracketright and sinceAis self-adjoint, =/angbracketleftX, AX /angbracketright+ 2λ/angbracketleftAX, η /angbracketright+λ2/angbracketleftη, Aη/angbracketright. Thus, f/prime(X,η) = lim λ→0f(X+λη)−f(X) λ= 2/angbracketleftAX, η /angbracketright. In particular when A=Iis the identity operator, f(X) =/bardblX/bardbl2, we findf/prime(X;η) = 2/angbracketleftX, η/angbracketright The directional derivatives of fin the particular direction of the coordinate axes e1= (1,0,...), e2= (0,1,0,...) have special names. They are called the partial derivatives of f. For example, the partial derivative of f(X) =f(x1,x2,...,x n) atX0with respect to x2is ∂f ∂x2(X0) :=f/prime(X0;e2) = lim λ→0f(X0+λe2)−f(X0) λ There are many other competing notations, all of them being used. We shall list them shortly, after observing there is a simple way to compute these partial derivatives. Consider f(X) =f(x−1,x2,x3) . Then ∂f ∂x2(X) = lim λ→0f(X+λe1)−f(X) λ SinceX+λe1= (x1,x2,x3) +λ(1,0,0) = (x1+λ,x 2,x3) we have ∂f ∂x2(X) = lim λ→0f(x1+λ,x 2,x3)−f(x1,x2,x3) λ. But this is the ordinary derivative of fwith respect to the single variable x1, while holding the other variables x2andx3fixed. Thus, ∂f/∂x 1can be computed by merely taking the ordinary one variable derivative of fwith respect to x1, pretending the other variables are constants. Example: Iff(X) =x2 1+x1ex1x2, find the rate of change of fat the point X0in the directionse1= (1,0) ande2= (0,1) . Thus, we want to compute∂f ∂x1(X0) and∂f ∂x2(X0). ∂f ∂x1= 2x1+ex1x2+x1x2ex1x2 ∂f ∂x2=x2 1ex1x2 At the pint X0= (2,−1) , we have ∂f ∂x1/vextendsingle/vextendsingle/vextendsingle/vextendsingle 2,−1= 4−e−2,∂f ∂x2/vextendsingle/vextendsingle/vextendsingle/vextendsingle (2,−1)= 4e−2. 8.1. THE DIRECTIONAL AND TOTAL DERIVATIVES 313 Some common notation. If w=f(x1,x2) , then ∂w ∂x1=∂f ∂x1=D1f=f1=fx1=wx1 ∂w ∂x2=∂f ∂x2=D2f=f2=fx2=wx2 Iff:A⊂En→E, then∂f/∂x jis another function of X= (x1,x2,...,x n) . It is then possible to take further partial derivatives. Example: Letw=f(X) =x2 1+x1ex1x2as in the previous example. Then w11=f11=fx1x1=∂2f ∂x2 1=∂ ∂x1(∂f ∂x1) = 2 +x2ex1x2=x2ex1x2+x1x2 2ex1x2 w12=f12=fx1x2=∂2 ∂x1∂x2=∂ ∂x2(∂f ∂x1) =x1ex1x2=x1ex1x2+x2 1x2ex1x2 w21=f21=fx2x1=∂2f ∂x2∂x1=∂ ∂x1(∂f ∂x2) = 2x1ex1x2=x2 1x2ex1x2=f12 w22=f22=fx2x2=∂2f ∂x2 2=∂ ∂x2(∂f ∂x2) =x3 1ex1x2. And even higher derivatives can be computed too, like f221=fx2x2x1=∂3f ∂x2 2∂x1=∂ ∂x1(∂2f ∂x2 2) = 3x2 1ex1x2+x3 1x2ex1x2. Remark: From this one example, it appears possible that we always have f12=f21, that is∂2f ∂x1∂x2=∂2f ∂x2∂x1. This is indeed the case ifthe second partial derivatives of fare continuous, but for lack of time we shall not prove it (see Exercise 6). So far we have defined the directional derivative of a function f:En→Eand called particular attention to those in the direction of the coordinate axes - the partial derivatives off. Although the actual computation of the partial derivatives has been reduced to the formal procedure of computing ordinary derivatives, the computation of the directional derivative in an arbitrary direction must still be done by using the definition: the limit of a difference quotient. We shall now reduce the computation of all directional derivatives to a simple formal procedure. In order to do so, we shall introduce the concept of the total derivative for functions f:A⊂En→E1. This derivative will not be a directional derivative, but rather a more general object. The motivating idea here is the important one of approximating a non-linear function fat a pointX0by a linear function. If we think of the function f(X) as defining a surface MinEn+1with point ( X,f(X)) , then the picture is that of approximating the surface MnearX0by a plane (or hyperplane) tangent to the surface at X0. We want to write f(X)∼f(X0) +L(X−X0), whereLis a linear operator, L:En→E, which may depend on the “base point” X0. Of course, as X→X0we want the accuracy to improve in the sense that the tangent plane should be a better approximation the closer Xis toX0. AtX=X0, the tangent 314CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS planef(X0) +L(X−X0) and surface Mtouch since they both pass through the point (X0,f(X0)) . Notice that the function f(X)) +L(X−X0) is affine, so it does represent a plane surface. Motivated by the above considerations, we can now make a reasonable Definition: Letf:A⊂En→EandX0be an interior point of A.fisdifferentiable atX0if there exists a linear transformation L:En→Esuch that lim /bardblh/bardbl→0/bardblf(X0+h)−f(X0)−Lh/bardbl /bardblh/bardbl= 0, for any vector hin some small ball about X0(sof(X0+h) is defined). The operator L will usually depend on the base point X0. Iffis differentiable at X0, we shall use the notation df dX(X0) =f/prime(X0) =L(X0)=L, and refer to f/prime(X0) as the total derivative offatX0. [The notation ∇f(X0) and grad f(X0) , for gradient , are also used]. If L=f/prime(X0) , a linear operator from EntoE, exists and depends continuously on the base point X0for allX0∈A, thenfis said to be continuously differentiable in A, writtenf∈C1(A) . Remark: The condition that fbe differentiable at X0can also be written in the following useful form: f(X0+h) =f(X0) +Lh+R(X0,h)/bardblh/bardbl, (8-1) where the remainder R(X0,h) has the property lim /bardblh/bardbl→0R(X0,h) = 0. This abstract operator Lhas the delightful property that it can be computed easily. But before telling you how, we should first prove for a given fthere can be at most one linear operator Lwhich is the total derivative. Theorem 8.1 . (Uniqueness of the total derivative). Let f:A→Ebe differentiable at the interior point X0∈A. IfL1andL2are linear operators both of which satisfy the conditions for the total derivative of fatX0, thenL1=L2. Proof: LetL=L1−L2. We shall show Lis the zero operator. Since Lh=L1h−L2h= [f(X0+h)−f(X0)−L2h]−[f(X0+h)−f(X0−L1h], by the triangle inequality we have /bardblLh/bardbl ≤ /bardblf(X0+h)−f(X0)−L2h/bardbl+/bardblf(X0+h)−f(X0)−L1h/bardbl. Consequently, lim /bardblh/bardbl→0/bardblLh/bardbl /bardblh/bardbl= 0. 8.1. THE DIRECTIONAL AND TOTAL DERIVATIVES 315 To complete the proof, a trick is needed. Fix η/negationslash= 0 . Ifλis a constant, λ→0 , then /bardblλη/bardbl →0 so lim /bardblλ/bardbl→0/bardblL(λη)/bardbl /bardblλη/bardbl= 0. But sinceLis linear, /bardblLλη/bardbl=/bardblλLη/bardbl=|λ| /bardblLη/bardbl, so the factor λcan be canceled in numerator and denominator. Thus the last equation is independent of λ, so/bardblLη/bardbl//bardblη/bardbl= 0 . Becauseη/negationslash= 0 , this implies /bardblLη/bardbl= 0 . Therefore Lmust be the zero operator. Next, we give a method for computing L. Not only that, but we also find an easy way to compute the directional derivatives. Theorem 8.2 . Letf:A→Ebe differentiable at the interior point X0∈A. Then a) the directional derivative of fatX0exists for every direction eand is given by the formula Def(X0) =Le. b) Moreover, if fis given in terms of coordinates, f(X) =f(x1,...,x n), thenLis represented by the 1×nmatrix f/prime(X0) =L= (fx1(X0),...,f xn(X0)). c) Consequently, the directional derivative is simply the product of this matrix Lwith the unit vector e, which can also be thought of as the scalar product of the 1×nmatrix, a vector, and the vector e, Def(X0) =/angbracketleftf/prime(X0), e/angbracketright. Proof: This falls out of the definitions. First Def(X0) = lim λ→0f(X0+λe)−f(X0) λ. = lim λ→0f(X0+λe)−f(X0)−L(λe) +L(λe) λ. SinceL(λe) =λLe and/bardblλe/bardbl=λ = lim /bardblλe/bardbl→0f(X0+λe)−f(X0)−L(λe) /bardblλe/bardbl+Le. Becausefis differentiable at X0, the first term tends to zero. Thus proving the first part. To prove the last part, it is sufficient to observe that if e=ejis one of the coordinate vectors, then by definition Dejf(X0) :=fxj. Thus, ifhis any vector h= (h1,...,h n) = h1e1+···+hnen, by the linearity of Lwe have Lh=L(h1e1+···+hnen) =h1Le1+···+hnLen =h1fx1(X0) +···+hnfxn(X0) = (fx1(X0),...,f xn(X0)) h1 · · · hn . 316CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS Sincehis any vector, we have shown Lis represented by the given matrix. Remark: The theorem states that iffis differentiable, then all the partial derivatives exist andf/prime(X0) :=Lis represented by the above matrix. It does notstate that if the partial derivatives exist, then fis differentiable. This is false (see Exercise 16). However, if the partial derivatives of fexist and are continuous, then fis differentiable. The last statement will be proved as Theorem 3. Example: The same one worked before (p. 573). Find the directional derivative of f(X) = x2 1−2x1x2+ 3x2atX0= (1,0) in the direction η= (−1,1) . Sinceηis not a unit vector, we let e=η /bardblη/bardbl= (−1√ 2,1√ 2) . Now at a point X, L=f/prime(X) = (fx1,fx2) = (2x1−2x2,−2x1+ 3). In particular, at X=X0= (1,0) , L= (2,−2 + 3) = (2,1). Therefore Def(X0) =Le= (2,1)/parenleftBigg −1√ 21√ 2/parenrightBigg =−1√ 2, which checks with the answer found previously. Consider the mapping w=f(X), X∈A⊂En, w∈Eas defining a surface M⊂En+1. It is now evident how to define the tangent plane to Mat the point ( X0,f(X0)) , where X0∈A. Definition: LetF:A⊂En→Ebe a differentiable mapping, thus defining a surface M with points ( X,f(X)), X∈A. The tangent plane toMat the point ( X0,f(X0)) , where X0∈A, is the surface defined by the affine mapping Φ(X) =f(X0) +f/prime(X0)(X−X0). or Φ(X) =f(X0) +L(X−X0),whereL=f/prime(X0), Thus, the tangent plane to the surface defined by fis merely the “affine part” of fat X0. Example: Consider the function w=f(X) = 3 −x2 1−x2 2. This function defines a paraboloid (see fig.). Let us find the tangent plane to this surface at ( X0,f(X0)) , where X0= (1,−1) , sof(X0) = 3−12−(−1)2= 1 . Also fx1(X) =−2x1,fx2(X) =−2x2. Thus f/prime(X0) = (fx1(X0),fx2(X0)) = (−2,2). SinceX−X0= (x1,x2)−(1,−1) = (x1−1,x2+ 1) we find the equation of the tangent plane is Φ(X) = 1 + ( −2,2)/parenleftbiggx1−1 x2+ 1/parenrightbigg = 1−2(x1−1) + 2(x2+ 1), 8.1. THE DIRECTIONAL AND TOTAL DERIVATIVES 317 or Φ(X) = 5−2x1+ 2x2. This tangent plane is the unique plane with the property Φ(X0) =f(X0),and Φ/prime(X0) =f/prime(X0). Although we have given necessary conditions that a function be differentiable (all directional derivatives exist, in particular, all partial derivatives exist), we have not given sufficient conditions. The next theorem gives sufficient conditions for a function to be continuously differentiable. Theorem 8.3 . Letf:A⊂En→E, whereAis an open set. Then fis continuously differentiable throughout Aif and only if all the partial derivatives of fexist and are continuous. Proof: ⇒Iffis continuously differentiable, then the partial derivatives exist by Theorem 2. Furthermore, for any XandYinA, fxi(X)−fxi(Y) =/angbracketleftf/prime(X), ei/angbracketright − /angbracketleftf/prime(Y), ei/angbracketright=/angbracketleftf/prime(X)−f/prime(Y), ei/angbracketright. Thus, applying the Schwartz inequality we find |fxi(X)−fxi(Y)| ≤ /bardblf/prime(X)−f/prime(Y)/bardbl. The statement fis continuously differentiable means the vector f/prime(X) is a continuous function of X. Therefore, given any /epsilon1>0 is aδ >0 such that /bardblf/prime(X)−f/prime(Y)/bardbl</epsilon1for all/bardblX−Y/bardbl<δ. For any/epsilon1>0 , the inequality above shows |fxi(X)−fxi(Y)|is also less than/epsilon1for the same δ. Consequently, fxiis continuous. ⇐. A little more difficult. The idea is to use the mean value theorem for functions of one variable. Let XandYbe points on A. To prove continuity at X, it is sufficient to restrictYto being in some ball about Xwhich is entirely in A(some ball does exist sinceAis open). For notational convenience, we take n= 2 . Then f(Y)−f(X) =f(Y)−f(Z) +f(Z)−f(X), whereZis a point in Awhose coordinates, except the first, are the same as Xand whose coordinates, except the second, are the same as Y. By the one variable mean value theorem, there is a point ˜XbetweenXandZand point ˆXbetweenYandZsuch that f(Z)−f(X) =∂f ∂x1(˜X)(y1−x1), f(Y)−f(Z) =∂f ∂x2(ˆX)(y2−x2). Therefore f(Y)−f(X) =∂f ∂x1(˜X)(y1−x1) +∂f ∂x2(ˆX)(y2−x2), so f(Y)−f(X)−[fxi(X)(y1−x1) +fx2(X)(y2−x2)] = [fxi(˜X)−fxi(X)](y1−x1) + [fx2(ˆX)−fx2(X)](y2−x2). Therefore /bardblf(Y)−f(X)−L(Y−X)/bardbl ≤/vextendsingle/vextendsingle/vextendsinglefxi(˜X)−fxi(X)/vextendsingle/vextendsingle/vextendsingle|y1−x1|/vextendsingle/vextendsingle/vextendsinglefx2(ˆX)−fx2(X)/vextendsingle/vextendsingle/vextendsingle|y2−x2| 318CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS where we have written L= (fx1(X),fx2(X)) . Since |yj−xj| ≤ /bardblY−X/bardbl, we see that /bardblf(X)−f(Y)−L(Y−X)/bardbl /bardblY−X/bardbl≤/vextendsingle/vextendsingle/vextendsinglefx1(˜X)−fx1(X)/vextendsingle/vextendsingle/vextendsingle+/vextendsingle/vextendsingle/vextendsinglefx2(ˆX)−fx2(X)/vextendsingle/vextendsingle/vextendsingle. Becausefx1andfx2are continuous and /bardbl˜X−X/bardbl</bardblY−X/bardbl,/bardblˆX−X/bardbl</bardblY−X/bardbl, by making /bardblY−X/bardblsufficiently small the right side of the above inequality can be made arbitrarily small. This proves the limit as /bardblY−X/bardbl →0 of the expression on the left - exists and is zero. Since Lis linear, the proof that fis differentiable is complete. The continuous differentiability is an immediate consequence of the linearity of Land the continuity of its components - the partial derivatives fxi. Exercises (1) i) Use the definition of the directional derivative to compute the given directional derivatives, ii) Check your answer by computing the directional derivative using the procedure of the Corollary to Theorem I. (a)f(x1,x2) = 1−2x1+ 3x2, at (2,−1) in the direction (3 ,4) . [Answer: +6 4]. (b)f(x,y) =ex+2y, at (3,−2) in the direction (1 ,1) . [Answer: 3 e−1/√ 2 ]. (c)f(u,v,w ) = 3uv+uw−v2, at (1,1,1) in the direction (1 ,−2,2) . (d)f(x,y) = 1−3y+xyat (0,6) in the direction (3 5,−4 5) . (2) i) Compute allof the first and second partial derivatives for the following functions. (a)f(x1,x2) =x1+x1sin 2x1 (b)f(x1,x2,x3) =x2 1x2+ 2x1√x3−x3 (c)f(x,y) =xy (d)f(x1,x2,...,x n) =a+a1x1+a2x2+...+anxn. (e)f(x1,x2,...,x n) =n/summationdisplay i,j=1aijxixj=/angbracketleftX, AX /angbracketright, whereaij=aji(first try the cases n= 2 andn= 3 to see what is happening). ii) Find the 1 ×nmatrixf/prime(x) . (3) For the surfaces defined by the functions f(X) listed below, find the equation of the tangent plane to the surface at the point ( X0,f(X0)) . Draw a sketch showing the surface and its tangent plane. (a)f(X) =x2 1+ 3x2 2+ 1, X 0= (0,0). (b)f(X) =ex1x2, X 0= (0,1) (c)f(X) =x2 1sinπx2, X 0= (−1,1 2) (d)f(X) =−1 2x1+x2+ 1, X 0= (2,1) (e)f(X) =x2 1+ 2x2 2−x1x3+x1, X 0= (1,−2,−1). Why can’t you sketch the surface defined by this function? 8.1. THE DIRECTIONAL AND TOTAL DERIVATIVES 319 (4) Letf(X) andg(X) both map A⊂En→E1. Iffandgare differentiable for all X∈A, prove (a)d dX[af(X) +bg(X)] =ad f dX(X) +bdg dX(X) (Linearity), where aandbare con- stants. (b)d dX[f(X)g(X)] =f(X)dg dX(X) +g(X)d f dX(X) (c)d dX/bracketleftBig f(X) g(X)/bracketrightBig =g(X)f/prime(X)−f(X)g/prime(X) g2(X), ifg(X)/negationslash= 0 . (5) Use the rules (a-c) of Exercise 4 to computed dX[2f−3g],d dX[f·g] , andd dX[f g] , where f(X) =f(x1,x2) = 1−x1+x1x2, andg(X) =g(x1,x2) =ex1−x2. (6) Letf(X) =f(x,y) =/braceleftBigg xy(x2−y2) x2+y2, X = (x,y)/negationslash= 0 0 X= 0 Prove (a)f,fx,fyare continuous for all X∈E2. [Hint: Prove and use 2 xy≤x2+y2]. (b)fxyandfyxexist for all X∈E2, and are continuous except at the origin. (c)fxy(0) = 1,fyx(0) = −1 , sofxy(0)/negationslash=fyx(0) (cf. Remark p. 577). (7) Letf:A⊂En→Ebe a differentiable map. Prove it is necessarily continuous. [Hint: This is a simple consequence of the definition in the form (1)]. (8) Letf:A⊂En→Ebe a continuous map. We say fhas a local maximum at the pointX0interior to Aiff(X0)≥f(X) for allXin some sufficiently small ball aboutX0. If we assume fis continuously differentiable, more can be said. (a) Iffas above has a local maximum at the point X0, prove /angbracketleftf/prime(X0, X−X0)/angbracketright+ R(X0,X)/bardblX−X0/bardbl ≤0 for allXis some small ball about X0. (b) Use the property of R(X0,X) to conclude the stronger statement /angbracketleftf/prime(X0),(X−X0)/angbracketright ≤0. for allXin some small ball about X0. (c) Observe the statement must also hold for the vector X0−X, which points in the direction opposite to X−X0, to conclude /angbracketleftf/prime(X0),(X−X0)/angbracketright ≥0, and hence that in fact /angbracketleftf/prime(X0), Z/angbracketright= 0, for all vectors Z=X−X0. (d) Finally, show that at a maximum, f/prime(X0) = 0. (9) (a) Find the equation of the plane which is tangent at the point X0= (2,6,3) to the surface consisting of the points ( X,f(X)) , where f(X) =f(x,y,z ) = (x2+y2+z2)1/2. 320CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS (b) Use the tangent plane found above to find the approximate value of ((2.01)2+ (5.98)2+ (2.99)2)1/2. (10) Assume the continuously differentiable function f(X) has a zero derivative, f/prime(X)≡ 0 , forXin some ball in En. Prove that f(X)≡constant throughout the ball. (11) (a) Show the following functions satisfy the two dimensional Laplace equation ∂2u ∂x2+∂2u ∂y2= 0 i)u(x,y) =x2−y2−3xy+ 5y−6 ii)u(x,y) = log(x2+y2) , except at the origin, ( x,y) = 0. iii)u(x,y) =exsiny (b) Show the following functions satisfy the one (space) dimensional wave equation utt=c2uxx, c≡constant [Heretis time and xis space;cis the velocity of light, sound, etc.] i)u(x,y) =ex−ct−2ex+ct ii)u(x,y) = 2(x+ct)2+ sin 2(x−ct). (12) Letf:A⊂En→Ebe continuously differentiable throughout A. IfX0∈Ais not a critical point of f, sof/prime(X0)/negationslash= 0 , prove the directional derivative at X0is greatest in the direction emax:=f/prime(X0)//bardblf/prime(X0)/bardbl, and least in the opposite direction, emin:=−emax. [Hint: Use the Schwarz inequality.] (13) Consider the function f(X) =f(x,y) =/braceleftbiggxy x2+y2, X = (x,y)/negationslash= 0 0, X = 0 SinceFis the quotient of two continuous functions, it is continuous except possibly at the origin, where the denominator vanishes. Show that f(X) isnotcontinuous at the origin by finding lim f(X) asX→0 along paths 1 and 2, and showing that lim X→0 path1f(X)/negationslash= lim X→0 path2f(X). (14) LetLbe the partial differential operator defined by Lu=∂2u ∂x2−5∂2u ∂x∂y+ 6∂2u ∂y2. Show that L[eαx+βy] =p(α,β)eαx+βy, wherep(α,β) is a polynomial in αandβ. Find a solution of the linear homogeneous partial differential equation Lu= 0 . Find an infinite number of solutions of Lu= 0 , one for each value of α, by choosing αto depend on βin a particular way. [Answer: e2βx+βyande3βx+βyare solutions for any β]. 8.2. THE MEAN VALUE THEOREM. LOCAL EXTREMA. 321 (15) The two equations x=eucosv y=eusinv defineu=f(x,y) andv=g(x,y) . Find the functions fandgforx>0 . Compute f/prime(X) andg/prime(X) and show f/prime(X)⊥g/prime(X) . (16) This exercise gives an example in which the first partial derivatives of a function exist but the function is not continuous, let alone differentiable. Let f(X) =f(x,y) =/braceleftBigg xy2 x2+y4, X = (x,y)/negationslash= 0 0, X = 0. (a) If cosα/negationslash= 0 , prove the directional derivative at the origin in the direction e= (cosα,sinα) exists and is Def(0) =2 sin2α cosα, cosα/negationslash= 0 while if cos α= 0 , Def(0) = 0, cosα= 0. (b) Provefis discontinuous at the origin by showing lim X→0f(X) has two different values along the two paths in the figure. Then appeal to exercise 7 to conclude fis not differentiable. (17) (a) Let P(X),X∈En, be a polynomial of degree N, that is, P(X) =/summationdisplay k1+k2+···+kn≤Nak1,...,k nxk1 1xk2 2···xknn, wherek1,k2,...,k nare all non-negative integers. Prove P(α) is continuously differentiable. [Hint: How do you prove a polynomial in one variable is continu- ously differentiable.] (b) LetR(X),X∈En, be a rational function - that is, the quotient of two polyno- mials. Prove R(X) is continuously differentiable whenever the denominator is not zero. (18) Iff:E1→E1, show that the definition of differentiability on page 578 coincides with the usual one. 8.2 The Mean Value Theorem. Local Extrema. Although the full “chain rule” will not be proved until Chapter 10, we shall need a very special and elementary case to develop the main features of the theory of mappings from En toE. Letf:A⊂En→Ebe a continuously differentiable function at all interior points ofA. TakeXandZto be fixed interior points of A. Letφ(t) =f(X+tZ) . We want to compute d dtφ(t) =d dtf(X+tZ) that is, the rate of change of f(X) at the point X+tZasXvaries along the line joining XtoZ. 322CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS Theorem 8.4 . Letf:A→Ebe a differentiable function throughout A. IfXandZ are two interior points of A, and if the line segment joining them is in A, then d dtf(X+tZ) =f/prime(X+tZ)Z, t ∈(0,1). By the product f/prime(Y)Zwe mean matrix multiplication. Proof: For fixedXandZ, the function φ(t) :=f(X+tZ) an ordinary scalar valued function of the one variable t. Thus d dtφ(t) = lim λ→0φ(tλ)−φ(t) λ = lim λ→0f(X+tZ+λZ)−f(X+tZ) λ = lim λ→0f(X+tZ+λZ)−f(X+tZ)−f/prime(X+tZ)(λZ) +f/prime(X+tZ)(λZ) λ Sincefis differentiable at X+tZ, then asλ→0 the first three terms tend to zero. The factorλin the last term cancels. Therefore d dtf(X+tZ) = lim λ→0f/prime(X+tZ)Z=f/prime(X+tZ)Z, as claimed. An easy consequence is Theorem 8.5 (The Mean Value Theorem). Let f:A→E, whereAis an open convex set in En, that is, if XandYare any points in Z, then the straight line segment joining XandYis inAtoo. Iffis differentiable in A, there is a point Zon the segment joiningXandYsuch that f(Y)−f(X) =f/prime(Z)(Y−X). If, moreover, f/primeis bounded by some constant C,/bardblf/prime(X)/bardbl ≤Cfor allX∈A, then |f(Y)−f(X)| ≤C/bardblY−X/bardbl a figure goes here Proof: Every point on the segment joining XandYis of the form X+t(Y−X) , where t∈[0,1] . Consider the function φ(t) of one variable, φ(t) =f(X+t(Y−X)). Theorem 4 states φis differentiable. Therefore, by the one variable mean value theorem, there is a number t0in the interval (0 ,1) such that φ(1)−φ(0) =φ/prime(t0) . Butφ(1) = f(Y), φ(0) =f(X) and, by Theorem 4, φ/prime(t0) =f/prime(X+t0(Y−X))(Y−X) . Letting Z=X+t0(Y−X) , a point on the segment joining XtoY, we conclude f(Y)−f(X) =f/prime(Z)(Y−X). 8.2. THE MEAN VALUE THEOREM. LOCAL EXTREMA. 323 The second part of the theorem follows by applying the Schwarz inequality to the function f/prime(Z)(Y−X) which can be written as /angbracketleftf/prime(Z),(Y−X)/angbracketright. Then /angbracketleftf/prime(Z), Y−X/angbracketright ≤ /bardblf/prime(Z)/bardbl/bardblY−X/bardbl. Therefore if /bardblf/prime(Z)/bardbl ≤Cfor allZ∈A, we find |f(Y)−f(X)| ≤C/bardblY−X/bardbl. Corollary 8.6 Letf:A→Ebe a differentiable map and Aan open connected set in En (by a connected open set we mean it is possible to join any two points in Aby a polygonal curve contained in A). If f/prime(X)≡0for everyX∈A, that is, if fx1(X) =...=fxn(X) = 0, thenf(X)≡c,ca constant. Proof: IfAis convex, say a ball, this is an immediate consequence of the second part of the mean value theorem, for /bardblf/prime(X)/bardbl= 0 so |f(Y)−f(X)|= 0 . Thus f(Y) =f(X) = constant for any two points XandY. The requirement that Ais connected is to exclude the possibility that Aconsists of two (or more) disjoint sets, in which case, all we can conclude is that fis constant on each connected part, but not necessarily the same constant. However, if Ais connected, then any two points in Acan be joined by a polygonal curve which is contained in A. Consider some straight line segment in this curve. By the mean value theorem, fmust be constant on it. In particular, it has the same value at both end points. Checking the beginning and end of the whole polygonal curve, we find that f(X) =f(Y) . Because XandYwere any points, we are done. It is not at all difficult to generalize the mean value theorem to Taylor’s theorem and then to power series for functions of several variables. The only problem is one of notation, and that is a problem. As a compromise, we will prove the Taylor theorem - but only the first two terms for functions of three variables f(x,y,z ) . Just as in the mean value theorem, the idea is to reduce the problem to a function φ(t) of one real variable, because we do know the result for these functions. Let fbe differentiable in some open set A⊂E3andX0a point inA. IfX0+his also inA, we would like to express f(X0+h) in terms of fand its derivatives at X0. FixX0andh and consider the real valued function φ(t) of one variable defined by φ(t) =f(X0+th), t ∈[0,1]. Then by Theorem 4, φ/prime(t) =f/prime(X0+th)h=fx(X0+th)h1+fy(X0+th)h2+fz(X0+th)h3, whereh= (h1,h2,h3) . Since each of the partial derivatives are maps from AtoE, they can be differentiated in the same way fwas. So can a sum of such functions. Thus φ/prime/prime(t) =d dt[fx(X0+th)h1+···+fz(X0+th)hn] =fxx(X0+th)h1h1+fxy(X0+th)h1h2+fxz(X0+th)h1h3 +fyx(X0+th)h2h1+fyy(X0+th)h2h2+fyz(X0+th)h2h3 +fzx(X0+th)h3h1+fzy(X0+th)h3h2+fzz(X0+th)h3h2. 324CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS If we introduce a matrix H(X) , the Hessian matrix , whose elements are∂2f(X) ∂xi∂xj, φ/prime/prime(t) can be written as φ/prime/prime(t) =/angbracketlefth, H(X0+th)h/angbracketright. We remark that if fis sufficiently differentiable (two continuous derivatives is enough), then the Hessian matrix is self-adjoint since fxixj=fxjxi, as we mentioned - but did not prove - earlier. If φ(t) is twice differentiable, by Taylor’s theorem for functions of one variable, we know that φ(1) =φ(0) +φ/prime(0) +1 2!φ/prime/prime(τ), τ∈(0,1). Substituting into this formula, we find f(X0+h) =f(X0) +f/prime(X0)h+1 2!/angbracketlefth, H(X0+τh)h/angbracketright. Let us summarize. We have proved Theorem 8.7 ( Taylor’s Theorem with two terms) . Letf:A→E, whereAis an open connected set in En. Assumefhas two continuous derivatives - that is, all the second partial derivatives of fexist and are continuous. If X0is inAandX0+his in a ball about X0inA, then f(X0+h) =f(X0) +f/prime(X0)h+1 2!/angbracketlefth, H(X0+τh)h/angbracketright, whereH(X) = ((∂2f ∂x1∂xj))is then×nHessian matrix and τ∈(0,1). LettingX=X0+handZ=X0+τh, Z being a point on the line segment joining X0toX, this reads f(X) =f(X0) +f/prime(X0)(X−X0) +1 2!/angbracketleftX−X0, H(Z)(X−X0)/angbracketright, or, in more detail, f(X) =f(X0) +n/summationdisplay i=1∂f(X0) ∂xi(xi−x0 i) +1 2n/summationdisplay i=jn/summationdisplay j=1∂2f(Z) ∂xi∂xj(xi−x0 i)(xj−x0 j). Example: Find the first two terms in the Taylor expansion for the function f(X) = f(x,y) = 5 + (2x−y)3about the point X0= (1,3) . We compute fx(X) = 6(2x−y)2, f y(X) =−3(2x−y)2 fxx(X) = 24(2x−y),fxy(X) =fyx(X) =−12(2x−y),fyy(X) = 6(2x−y). Thereforef(X0) = 4,fx(X0) = 6,fy(X0) =−3 , so f(X) = 4 + (6,−3)/parenleftbiggx−1 y−3/parenrightbigg +1 2(x−1,y−3)/parenleftbigg2ξ−η −12(2ξ−η) −12(2ξ−η) 6(2ξ−η)/parenrightbigg/parenleftbiggx−1 y−3/parenrightbigg 8.2. THE MEAN VALUE THEOREM. LOCAL EXTREMA. 325 whereZ= (ξ,η) is a point on the segment between X0= (1,3) andX= (x,y) . Written out, the above equation reads, f(x,y) = 4 + 6(x−1)−3(y−3) +1 2[fxx(x−1)2+ 2fxy(x−1)(y−3) +fyy(y−x)2], where the second derivatives are evaluated at Z= (ξ,η) . We are now in a position to examine the extrema of functions of several variables. Finding the maxima and minima of functions is important for several reasons. First of all, there is the vague emotional feeling that all patterns of action should maximize or minimize something. Second, we can investigate a complicated geometrical object by the relatively easy procedure of finding the local maxima and minima. Without further mention, for the balance of this section f(X) will be a twice continuously differentiable function which maps the open set A⊂EnintoE. Definition: A function f:A→Ehas a local maximum at the interior point X0∈Aif, for allXin some open ball about X0 f(X)≤f(X0). fhas a local minimum atX0if for allXin some open ball about X0 f(X)≥f(X0). Iffhas a local maximum or minimum at X0, isf/prime(X0) = 0 ? Certainly. Theorem 8.8 . Iffhas a local maximum or minimum at X0, thenf/prime(X0) = 0 . In coordinates, this means all the partial derivatives vanish at X0, ∂f ∂x11(X0) =∂f ∂x2(X0) =···=∂f ∂xn(X0) = 0. Proof: Letηbe any fixed vector. Then the function φ(t) of one variable φ(t) =f(X0+tη) has a local maximum or minimum at t= 0 . Consequently φ/prime(0) = 0 . But by Theorem 4, φ/prime(0) =f/prime(X0)ηwhich we may write as /angbracketleftf/prime(X0), η/angbracketright. Thus /angbracketleftf/prime(X0), η/angbracketright= 0 , so the vector f/prime(X0) is orthogonal to η. Sinceηwas any vector, we conclude that f/prime(X0) = 0 . The derivative f/prime(X0) may vanish at points other than maxima or minima. An example is the “saddle point” of the hyperbolic paraboloid at the beginning of Section 1. All points wheref/primevanishes are called critical points orstationary points off. Let us give a precise definition of a saddle point. fhas a saddle point atX0ifX0is a critical point of f and if every ball about X0contains points X1andX2such thatf(X1)< f(X0) and f(X2)>f(X0) . Thus, every critical point is either a local maximum, minimum, or saddle point. There is a more intuitive way to prove Theorem 7. If eis a unit vector, then by Theorem 2, the directional derivative at Xin the direction eisDef(X) =/angbracketleftf/prime(X), e/angbracketright. In what way should you move so fincreases fastest? By the Schwartz inequality, we find |Def(X)| ≤ /bardblf/prime(X)/bardbl /bardble/bardbl=/bardblf/prime(X)/bardbl, 326CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS with equality if and only if the vectors eandf/prime(X) are parallel. Thus, the directional derivative is largest when e has the same direction as f/prime(X), and smallest when e has the opposite direction ,emax=f/prime(X)//bardblf/prime(X)/bardbl, emin=−emax, Demaxf(X) =/bardblf/prime(X)/bardbl, Deminf(X) =−/bardblf/prime(X)/bardbl. IfX0is a local maximum of f, thenf/prime(X0) must be zero, for otherwise you could move in the direction of f/prime(X0) and increase the value of f. Similarly, if X0is a local minimum, f/prime(X0) must be zero. Once we know X0is a critical point of f, f/prime(X0) = 0 , an effective criterion is needed to determine if X0is a local maximum, minimum, or saddle point for f. In elementary calculus, the sign of the second derivative was used. Our next theorem generalizes this test. The idea is essentially the same as in the one variable case (p. 104a-c). If fhas a local maxima or minima, the tangent plane to the surface whose points are ( X,f(X)) is horizontal, that is, f/prime(X0) = 0 . Thus, near X0the quadratic terms - the next lowest power in the Taylor expansion of faboutX0—will determine the behavior of fnearX0. Let X0be the origin and take f(X) =f(x,y) to be a function of two variables with f(0) = 0 . Then near X0= 0 , by Taylor’s theorem, we have f(x,y)∼1 2[ax2+ 2bxy+cy2], wherea=fxx(0), b=fxy(0) , andc=fyy(0) . The nature of the quadratic form Q(X) =ax2+ 2bxy+cy2 has already been determined. If Q(X) is positive definite, then Q(X)>0 forX/negationslash= 0 . Sincef(x,y)∼Q(X) , this means f(x,y) is positive near the origin. Because f(0,0) = 0 , this implies the origin is a minimum for f. Instead of completing and rigorously justifying this special case, we shall immediately treat the general situation. Theorem 8.9 . Assume the twice continuously differentiable function f:A→Ehas a critical point at an interior point X0ofA⊂En, f/prime(X0) = 0 . LetH(X0)be the Hessian matrix/parenleftbigg/parenleftbigg∂2f ∂xi∂xj(X0)/parenrightbigg/parenrightbigg evaluated at X0. (a)IfH(X0)is positive definite, then fhas a local minimum at X0. (b)IfH(X0)is negative definite, then fhas a local maximum at X0. (c)If at least two of the diagonal elements of H(X0), fx1x1(X0),...,f xnxn(X0)have different signs, then X0is a saddle point. (d)Otherwise the test fails. Proof: IfX0is a critical point for f, then Taylor’s theorem (Theorem 6) states f(X0+η) =f(X0) +1 2/angbracketleftη, H(Z)η/angbracketright whereZis between X0andX0+η. The linear term has been dropped since f/prime(X0) = 0 . 8.2. THE MEAN VALUE THEOREM. LOCAL EXTREMA. 327 As in the proof of Taylor’s theorem, let φ(t) =f(X0+tη). Then φ/prime/prime(t) =/angbracketleftη, H(X0+tη)η/angbracketright. Since the second derivatives of fare assumed to be continuous, the function φ/prime/prime(t) is a continuous function of t. Consequently, if φ/prime/prime(0) is positive then φ/prime/prime(t) is also positive for alltsufficiently close to zero (Theorem I p. 29b). Because φ/prime/prime(0) = /angbracketleftη, H(X0)η/angbracketrightand φ/prime/prime(τ) =/angbracketleftη, H(Z)η/angbracketright, whereZ=X0+τη, this implies if His positive definite at X0, it is also positive definite at ZwhenZis close toX0. AssumingH(X0) is positive definite, we see that for all ηsufficiently small, H(Z) is positive definite. Therefore, f(X0+η)−f(X0) =1 2/angbracketleftη, H(Z)η/angbracketright>0, η/negationslash= 0, that is f(X0+η)−f(X0)>0 for allηis some small ball about X0. Thusfhas a local minimum at X0. IfH(X0) is negative definite, the same proof with trivial modifications works. Another way to complete the proof is to apply part a) to the function g(X) :=−f(X) . The Hessian forgatX0will be −H(X0) which is positive definite (since H(X0) was negative definite). Thusghas a local minimum at X0sof: =−ghas a local maximum at X0. If any two of the diagonal elements of H(X0) have opposite sign, say fx1x1(X0)>0 andfx2x2(X0)<0 , then for η=λe1= (λ,0,0,..., 0) ,λany real number, we find /angbracketleftη, H(X0)η/angbracketright=λ2fx1x1(X0)>0 , while for η=λe2= (0,λ,0,..., 0)/angbracketleftη, H(X0)η/angbracketright= λ2fx2x2(X0)<0 . Therefore the quadratic form /angbracketleftη, H(X0)η/angbracketrightassumes positive and negative values in any ball about X0, provingX0is a saddle point. Since this theorem reduces the investigation of the nature of a critical point to testing if a matrix is positive or negative definite, it would do well in this context to repeat Theorem A (p. 386d) which tells us when a 2 ×2 matrix is positive definite. Corollary 8.10 . LetX0be a critical point for the function of two variables f(x,y)with Hessian matrix H(X0) =/parenleftbiggfxx(X0)fxy(X0) fxy(X0)fyy(X0)/parenrightbigg . (a)IfdetH(X0)>0andfxx(X0)>0, thenfhas a local minimum at X0. (b)IfdetH(X0)>0andfxx(X0)<0, thenfhas a local maximum at X0. (c)IfdetH(X0)<0, thenfhas a saddle point at X0(this is a stronger statement than part c of Theorem 8). Proof: Since these merely join Theorem A (p. 386d) with Theorem 8, the proof is done. Examples: 328CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS (1) Find and classify the critical points of the function w=f(x,y) := 3 −x2−4y2+ 2x. A sketch of the surface with points ( x,y,f (x,y)) , a paraboloid, is at the right. At a critical point f/prime(X) = 0 , that is, fx= 0, fy= 0 . Since fx=−2x+ 2,fy=−8y, at a critical point −2x+ 2 = 0,−8y= 0. There is therefore only one critical point, X0= (1,0) . We look at the Hessian to determine the nature of the critical point. Because fxx=−2, fxy=fyx= 0, fyy= −8 , H(X0) =/parenleftbigg−2 0 0−8/parenrightbigg . Since detH(X0) = 16>0 andfxx(X0) =−2<0,H(X0) is negative definite so X0= (1,0) is a local maximum for the function, and at that point f(X0) = 4 . (2) Find and classify the critical points of w=f(x,y) =−x2+y2. The surface ( x,y,f (x,y)) is a hyperbolic paraboloid. We expect a saddle point at the origin. At a critical point fx=−2x= 0, f y= 2y= 0. Thus the origin (0 ,0) is the only critical point. Since H(x,y) =/parenleftbigg−2 0 0 2/parenrightbigg , and detH(0,0) =−4<0 , the origin is a saddle point. This also follows from the observation that the diagonal elements have different signs. (3) Find and classify the critical points of w=f(x,y) = [x2+ (y+ 1)2][x2+ (y−1)2]. At a critical point, fx= 2x[x2+ (y−1)2] + 2x[x2+ (y+ 1)2] = 0 and fy= 2(y+ 1)[x2+ (y−1)2] + 2(y−1)[x2+ (y+ 1)2] = 0. The first equation implies x= 0 . Substituting this into the second we find y= 0,y= 1,y=−1 . Thus there are three critical points X1= (0,0), X 2= (0,1), X 3= (0,−1). We must evaluate the Hessian matrix at these points. Since fxx= 12x2+ 4y2+ 4, f xy= 9xy, f yy= 4x2+ 12y2=−4, H(X1) =/parenleftbigg4 0 0−4/parenrightbigg , H (X2) =/parenleftbigg8 0 0 8/parenrightbigg =H(X3). Because det H(X1) =−16<0, X 1= (0,0) is a saddle point. Because det H(X2)> detH(X3) = 64>0 andfxx(X2) =fxx(X3) = 8>0 , bothX2= (0,1) and X3= (0,−1) are local minima. To complete the computation, we find f(X1) = 1,f(X2) = 0,f(X3) = 0 . A sketch of the surface is at the right. 8.2. THE MEAN VALUE THEOREM. LOCAL EXTREMA. 329 (4) Find and classify the critical points of w=f(x,y,z ) = 1−2x+ 3x2−xy+xz−z2+ 4z+y2+ 2yz. At a critical point, fx=−2 + 6x−y+z,fy=−x+ 2y+ 2z,fz=x−2z+ 4 + 2y. Solving these equations, we find only one critical point, X0= (0,−1,1),where f(X0) = 3 . Since fxx= 6, fxy=−1, fxz= 1fyy= 2,fyx= 2,fzz=−2, then H(X) = 6−1 1 −1 2 2 1 2 −2 . Because the diagonal elements 6 ,2,−2 are not all of the same sign, by part c of the theorem, the critical point X0= (0,−1,1) is a saddle point. (5) Find and classify the critical points of w=f(x,y) :=x2y2. At a critical point, fx= 2xy2= 0, f y= 2x2y= 0. Thus the points where either x= 0 ory= 0 are all critical points. Since fxx= 2y2, fxy= 4xy, f yy= 2x2, we find H(X) =/parenleftbigg2y24xy 4xy2x2/parenrightbigg If eitherx= 0 ory= 0 , then det H= 0 so none of our tests apply to determine the nature of the critical point. However, a glance at the function f(x,y) =x2y2reveals that all of the points where either x= 0 ory= 0 are clearly local minima, since f= 0 there, whilef >0 elsewhere. Exercises (1) Find and classify the critical points of the following functions. (a)f(x,y) =x2−3x+ 2y2+ 10 (b)f(x,y) = 3−2x+ 2y+x2y2 (c)f(x,y) = [x2+ (y+ 1)2][4−x2−(y−1)2] (d)f(x,y) =x3−3xy2(figure on next page) (e)f(x,y) =xy−x+y+ 2 (f)f(x,y) =xcosy (g)f(x,y,z ) = 2x2+ 3xz+ 5z2+ 4y−y2+ 7 330CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS (h)f(x,y,z ) = 5x2+ 4xy+ 2y2+z2−4z+ 31 (2) LetX1,...,X NbeNdistinct points in En. Find a point X∈Ensuch that the function f(X) =/bardblX−X1/bardbl2+···+/bardblX−XN/bardbl2 is a minimum. [Answer: X=1 N/summationtextN j=1Xj, the center of gravity.] (3) (a) Find the minimum distance from the origin in E3to the plane 2 x+y−z= 5. (b) Find the minimum distance from the origin in Ento the hyperplane a1x1+ a2x2+···+anxn=c. (c) Find the minimum distance between the fixed point X0= (˜x1,..., ˜xn) and the hyperplane a1x1+...+anxn=c. (d) Find the minimum distance between the two parallel planes a1x1+···+anxn=c1 anda1x1+···+anxn=c2, (4) Iff(x,y) has two continuous derivatives, use Taylor’s Theorem (Theorem 6) to prove f(x+h1,y+h2) =f(x,y) +fx(x,y)h1+fy(x,y)h2 +1 2[fxx(x,y)h2 1+ 2fxy(x,y)h1h2+fyy(x,y)h2 2] + (h2 1+h2 2)R, whereRdepends on x,y,h 1andh2, and lim h1→0 h2→0R= 0 . (5) (a) If u(x,y) has two continuous derivatives, use the result of Exercise 4 to prove uxx(x,y) =u(x+h1,y)−2u(x,y) +u(x−h1,y) h2 1+h1˜R and uyy(x,y) =u(x,y+h2)−2u(x,y) +u(x,y−h2) h2 2+h2ˆR, where lim h1→0˜R= 0 and lim h2→0ˆR= 0 . (b) Use part a) to deduce that if h1=h2=hthen uxx(x,y) +uyy(x,y) =4 h2[u(x,y)−u(x+h,y) +u(x−h,y) +u(x,y+h) +u(x,y−h) 4] +h2R (8-2) where lim h→0R= 0 . (c) Use part b) to deduce that if his small, the solution of the partial differential equationuxx+uyy= 0 , Laplace’s equation, approximately satisfies the difference equation u(x,y) =u(x+h,y) +u(x−h,y) +u(x,y+h) +u(x,y−h) 4 This difference equation states that the value of uat the center of a cross equals the arithmetic mean (“average”) of its values at the four ends of the cross. One could use the difference equation to solve Laplace’s equation numerically. 8.2. THE MEAN VALUE THEOREM. LOCAL EXTREMA. 331 (d) Prove that any function which satisfies the above difference equation in some set cannot have a maxima or minima inside that set. [Do not differentiate! Reason directly from the difference equation. No computation is necessary.] (6) If all the second partial derivatives of a function f(X) vanish identically in some open connected set, prove that fis an affine function. (7) (The Method of Least Squares). Let Z1,...,Z NbeNdistinct points in En, and w1,...,w Na set ofNnumbers. We imagine the points ( Zj,wj)∈En+1to be points on a surface MinEn+1. Find a hyperplane w=φ(X) =c+ξ1x1+···,+ξnxn≡c+/angbracketleftξ, X/angbracketright which most closely approximates the surface Min the sense that the error E(ξ) E(ξ) :=N/summationdisplay j=1|φ(Zn)−wj|2= is minimized. Note that you are to find the coefficients ξ1,...,ξ nin the equation of the hyperplane. (8) (a) Let u(x,y) be a twice continuously differentiable function which satisfies the partial differential equation Lu:=uxx+uyy+aux+buy−cu= 0 in some open set D, where the coefficients a(x,y),b(x,y) , andc(x,y) are con- tinuous functions. If c >0 throughout D, prove that u(x,y) cannot have a positive maximum or negative minimum anywhere in D. (b) Extend the result of part a) to functions u(x1,...,x n) which satisfy Lu:=n/summationdisplay i=1∂2u ∂x2 1+n/summationdisplay j=1aj∂u ∂xj−cu= 0, in some open set D, wherec>0 throughout D. (c) Ifu(x,y) satisfies the equation of part a) and uvanishes on the boundary of D, u ≡0 on∂D, prove that u(x,y)≡0 throughout D. (d) Assume u(x,y) andv(x,y) both satisfy the same equation Lu= 0, Lv = 0 , whereLis the operator of part a). If u(x,y)≡v(x,y) on the whole boundary ofD, prove that u(x,y)≡v(x,y) throughout the interior of D. (9) Letfbe a twice continuously differentiable function throughout the open set A. Prove that (a) iffhas a local minimum at X0∈A, then its Hessian H(X0) is positive definite or semi-definite there. (b) iffhas a local maximum at X0∈A, then its Hessian H(X0) is negative definite or semi-definite there. 332CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS (10) LetAbe a square n×nself-adjoint matrix and Ya fixed vector in En, and let f(X) =/angbracketleftX, AX /angbracketright −2/angbracketleftX, Y/angbracketright. (a) Iff(X) has a critical point at X0, proveX0satisfies the equation AX 0=Y. (b) IfAis positive definite and X0satisfies the equation AX 0=Y, provef(X) defined above has a minimum at X0. [The results of this problem remain valid ifAis any positive definite linear operator - possibly a differential operator. The nonlinear function f(X) defines a variational problem associated with the equationAX=Y.] (11) Iff:A→Ehas three continuous derivatives in the open set A⊂E2containing the origin, state precisely and prove Taylor’s Theorem with three terms about the origin. The resulting expression will be f(X) :=f(x,y) =f(0) +fx(0)x+fy(0)y+1 2![fxx(0)x2+ 2fxy(0)xy+fyy(0)y2] +1 3![fxxx(Z)x3+ 3fxxy(Z)x2y+ 3fxyy(Z)xy2+fyyy(Z)y3] whereZis on the line segment between 0 and X= (x,y) . (12) Ifu(x,y) has the property uxy(x,y) = 0 for (x,y) in some open set, prove u(x,y) = φ(x) +ψ(y) , whereφandψare functions of one variable. (13) Compute the direction(s) at X0in which the following functions f i) increase most rapidly, ii) decrease most rapidly, iii) remain constant. (a)f(x1,x2) = 3−2x1+ 5x2 atX0= (2,1) (b)f(x,y) =e2x+yatX0= (1,−2) (c)f(x,y,z ) = 2x2+ 3xy+ 5z2+ 4y−y2+ 7 at X0(1,0,−1) (d)f(u,v) =uv−u+v+ 2 at ( −1,1) . 8.3 The Vibrating String. Waves. You have been hearing about them your whole life. Waves are the term used to describe the oscillatory behavior of continuous media; water waves and sound waves being the most familiar. We shall give a mathematical description of a very simple type of wave - those in an oscillating violin string. The resulting mathematical model will be a second order linear partial differential equation - the wave equation - with both initial and boundary conditions. 8.3. THE VIBRATING STRING. 333 a) The Mathematical Model Consider a string of length /lscriptstretched along the xaxis. Imagine the string vibrating in the plane of the paper and let u(x,t) denote the vertical displacement of the point x at timet. In order to end up with a tractable mathematical model several reasonable simplifying assumptions will be made. We assume the tension τand density ρof the string are constant throughout the motion, while the string is taken to be perfectly flexible so the tension force in the string acts along the tangential direction. Dissipative effects (air resistance, heating, etc.) are entirely neglected. One more assumption will be made when needed. It essentially states that the oscillations are small in some sense. Newton’s second law, ma=/summationtextF, is where we begin. Draw your attention to a small segment of the string whose length, at rest, is ∆ x=x2−x1. The mass of the segment is ρ∆x. By Newton’s second law the segment moves in such a way that the product of its center of gravity equals the resultant of the forces acting on it. For the vertical component, this means ρ∆x∂2u ∂t2(˜x,t) =Fv, where ˜x∈(x,x+ ∆x) is the horizontal coordinate of the center of gravity of the segment, andFvmeans the vertical component of the resultant force. There are two types of forces. One is the tension acting at both ends of the segment. The other is gravity acting down with a force equal to the weight of the segment, ρg∆x. To evaluate the tension forces, let θ1andθ2be the angles the string makes with the horizontal at either end of the segment (see figure above). Then the vertical component of the tension force is τsinθ2−τsinθ1. The signs indicate one force is up while the other is down. Adding the tension force to the gravitational force and substituting into Newton’s second law, we find ρ∆x∂2u ∂t2(˜x,t) =τ(sinθ2−sinθ1)−ρg∆x. The dependence of θ1andθ2on the displacement can be brought out by using the relation sinθ=ux/radicalbig 1 +u2x, which follows from the relation ux= tanθfor the slope of the string. Using this, we obtain the equation ρ∆x∂2u ∂t2(˜x,t) =τ/bracketleftBigg ux/radicalbig 1 +u2x/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglex=x2−ux/radicalbig 1 +u2x/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle x=x1/bracketrightBigg −ρg∆x. A simplifying assumption is badly needed. If the function ux//radicalbig 1 +u2xis expanded in a Taylor series, ux/radicalbig 1 +u2x=ux−1 2u3 x+···, 334CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS we see that if the slope uxis small, essentially only the linear term in this series counts. Therefore, we do assume the slope uxis small (this is the same assumption made in treating the simple pendulum). With this simplification, the equation of motion is ρ∆x∂2u ∂t2(˜x,t) =τ[ux(x2,t)−ux(x1,t)]−ρg∆x. Divide both sides of this equation by ∆ x=x2−x1and let the length of the interval shrink to zero. Since lim (x2−x1)→0ux(x2,t)−ux(x1,t) x2−x1=∂ ∂xux(x,t) =∂2u ∂x2(x,t), wherexis the limiting value of x1andx2, we find ρ∂2u ∂t2(x,t) =τ∂2u ∂x2(x,t)−ρg Because the length of the interval has been shrunk to one point x, the center of gravity is now atxtoo. It is customary to let τ/ρ=c2. The constant chas units of velocity, and, in fact, is just the speed with which waves travel along the string. Thus Lu:=utt−c2uxx=−g. This is the wave equation , a second order linear inhomogeneous partial differential equation. As was the case with linear ordinary differential equations, it is easier to attempt first to solve the homogeneous equation Lu:=utt−c2uxx= 0. On physical grounds, we expect the motion u(x,t) of the string will be determined if the initial position u(x,0) and initial velocity ut(x,0) are known, along with the motion of both end points u(0,t) andu(/lscript,t) . However the mathematical model must be examined to see if these four facts do determine the subsequent motion (which it should if the model is to be of any use). Thus we must prove that given the initial position u(x,0) =f(x), x ∈[0,/lscript] initial velocity ut(x,0) =g(x), x ∈[0,/lscript] motion of left end u(0,t) =φ(t)t≥0 motion of right end u(/lscript,t) =ψ(t), t≥0, then a solution u(x,t) of the wave equation utt−c2uxx= 0 does exist which has these properties, and there is only one such solution. Existence and uniqueness theorems must therefore be proved. b) Uniqueness . This is almost identical to all uniqueness theorems encountered earlier, especially that for the simple harmonic oscillator in Chapter 4, Section 2. 8.3. THE VIBRATING STRING. 335 Theorem 8.11 (Uniqueness). There exists at most one twice continuously differentiable functionu(x,t)which satisfies the inhomogeneous wave equation Lu:=utt−c2uxx=F(x,t) and the subsidiary initial conditions: u(x,0) =f(x),ut(x,0) =g(x), x∈[0,/lscript] boundary conditions: u(0,t) =φ(t),u(/lscript,t) =ψ(t), t≥0, whereF,f,g,φ , andψare given functions. Proof: Assumeu(x,t) andv(x,t) both satisfy the same equation and the same subsidiary conditions. Let w(x,t) =u(x,t)−v(x,t) . ThenLw=Lu−Lv=F−F= 0 , sowsatisfies the homogeneous equation Lw:=wtt−c2wxx= 0 and has zero subsidiary data initial conditions: w(x,0)≡0, wt(x,0)≡0, x∈[0,/lscript] boundary conditions: w(0,t)≡0, w(/lscript,t)≡0, t≥0 We want to prove w(x,t)≡0 . Notice that wsatisfies the equation for a vibrating string which is initially at rest on the xaxis, and whose ends never move. Therefore our desire to prove the string never moves, w(x,t)≡0 , is certain physically reasonable. For this function w,define the new function E(t) E(t) =1 2/integraldisplay/lscript 0[w2 t+c2w2 x]dx. We have named the function E(t) since it actually happens to be the energy in the string associated with the motion w(x,t) at timet, except for a factor of ρ. Assume it is “legal” to differentiate under the integral sign (it is). Upon doing so, we get dE dt=/integraldisplay/lscript 0[wtwtt+c2wxwxt]dx. But an integration by parts reveals that /integraldisplay/lscript 0wxwxtdx=wxwt/vextendsingle/vextendsingle/lscript 0−/integraldisplay/lscript 0wtwxxdx. Because the end points are held fixed, w(0,t) = 0 and w(/lscript,t) = 0 , the velocity at those points is zero too, wt(0,t) = 0 andwt(/lscript,t) = 0 . This drops out the boundary terms in the integration by parts. Substituting the last expression into that for dE/dt , we find that dE dt=/integraldisplay/lscript 0wt[wtt−c2wxx]dx. Butwsatisfies the homogeneous wave equation wtt−c2wxx= 0 . Therefore dE/dt ≡0 , so E(t)≡constant = E(0), that is, energy is conserved . Now E(0) =1 2/integraldisplay/lscript 0[w2 t(x,0) +c2w2 x(x,0)]dx. 336CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS Since the initial position is zero, w(x,0) = 0 , its slope is also zero, wx(x,0) = 0 . The initial velocity wt(x,0) is also zero, wt(x,0) = 0 . Thus E(t)≡E(0)≡0, that is, 0 =E(t) =1 2/integraldisplay/lscript 0[w2 t(x,t) +c2w2 x(x,t)]dx. Because the integrand is positive, we conclude wt(x,t)≡0 andwx(x,t)≡0 . Consequently w(x,t)≡constant. Since w(0,t) = 0 , that constant is the zero constant, w(x,t)≡0. Therefore u(x,t)−v(x,t)≡w(x,t)≡0, sou(x,t)≡v(x,t) : the solution is unique. c) Existence For the simple one (space) dimension wave equation, there are many ways to prove a solution exists. The one to be given here is not the simplest (see Exercise 6 for the result of that method), but it does generalize immediately to many other problems. It makes no difference how we find a solution, for once found, by the uniqueness theorem it is the only possible solution. To avoid complications, we shall consider only the homogeneous equation and assume the end points are tied down. Thus, we want to solve Wave equations: utt−c2uxx= 0. Initial conditions: u(x,0) =f(x), ut(x,0) =g(x). Boundary conditions: u(0,t) = 0, u(/lscript,t) = 0. The idea is first to find special solutions u1(x,t), u2(x,t),..., which satisfy the bound- ary conditions but do not necessarily satisfy the initial conditions. Then, as was done for linear O.D.E.’s, we build the solution which does satisfy the given initial conditions as a linear combination of these special solutions, u(x,t) =/summationdisplay Ajuj(x,t), that is, by superposition. Let us seek special solutions in the form of a standing wave , u(x,t) =X(x)T(t). HereX(x) andT(t) are functions of one variable. Our procedure is reasonably called separation of variables . Substitution of this into the wave equation gives ¨T(t)X(x)−c2X/prime/prime(x)T(t) = 0, or X/prime/prime(x) X(x)=1 c2¨T(t) T(t). 8.3. THE VIBRATING STRING. 337 Since the left side depends only on x, while the right depends only on t, both sides must be constant (a somewhat tricky remark; think it over). Let that constant be −γ(using −γinstead ofγis the result of hindsight, as you shall see). X/prime/prime X=1 c2¨T T=−γ. This leads us to the two ordinary differential equations X/prime/prime(x) +γX(x) = 0, ¨T(t) +γc2T(t) = 0. Sinceu(0,t) = 0 and u(/lscript,t) = 0 and u(x,t) =X(x)T(t) , the function X(x) must also satisfy the boundary conditions X(0) = 0, X (/lscript) = 0. There are several ways to show γmust be positive. Perhaps the simplest is to observe that ifγ <0 orγ= 0 , the only function X(t) which satisfies the differential equation X/prime/prime+γX= 0 and boundary conditions X(0) =X(/lscript) = 0 is the zero function X(x)≡0 . Since for this function u(x,t) =X(x)T(t)≡0 , it is devoid of further interest. Another way to show γis positive is to multiply the ordinary differential equation X/prime/prime+γX= 0 byX(x) and integrate over the length of the string, /integraldisplay/lscript 0[X(x)X/prime/prime(x) +γX2(x)]dx= 0. Upon integrating by parts, we find that /integraldisplay/lscript 0?(x)X/prime(x)dx=XX/prime/vextendsingle/vextendsingle/vextendsingle/lscript 0−/integraldisplay/lscript 0X/prime2(x)dx. SinceX(0) =X(/lscript) = 0 , the boundary terms drop out. Substituting this into the above equation, we find that/integraldisplay/lscript 0X/prime2(x)dx=γ/integraldisplay/lscript 0X2(x)dx. IfX(x) is not identically zero, this can be solved for γ γ=/integraldisplay/lscript 0X/prime2(x)dx /integraldisplay/lscript 0X2(x)dx, and clearly shows γ >0 . Enough for that. The solution of X/prime/prime+γX= 0, γ > 0 , is X(x) =Acos√γx+Bsin√γx. The boundary condition X(0) = 0 implies A= 0 , while the boundary condition at the other end point X(/lscript) = 0 , implies 0 =Bsin√γ/lscript. 338CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS IfB= 0 too, then X(x)≡0 , sou(x,t)≡0 . This is of no use to us. The only alternative is to restrict γso that sin√γl= 0 . This means√γ/lscriptis a multiple of π,√γ/lscript=nπ, n = 1,2,..., √γ=nπ /lscript, n = 1,2,.... There is then one possible solution X(x) for each integer n, Xn(x) =Bnsinnπ /lscriptx, where the constants Bnare arbitrary. Remark: There is a similarity of deep significance for mathematics and physics between the work in these last few paragraphs and that done for the coupled oscillators in Chapter 6. There (p. 528-9), we had an operator Aand wanted to find nonzero vectors Snand numbersλsuch that ASn=λnSn. The numbers found λnwere called the eigenvalues of A, andSnthe corresponding eigen- vectors. Here, we were given the operator A=−d2 dx2and wanted to find nonzero functions Xn(t)∈ {X∈C2[0,/lscript]:X(0) =X(/lscript) = 0}which satisfy the equation AXn=γnXn The numbers found, γn=n2π2//lscript2, are also called the eigenvalues ofA, and the function Xn(t) = sinnπ /lscriptx, the eigenfunction ofAcorresponding to the eigenvalue γn. Associated with each possible eigenvalue γn, there is a solution of the time equation, ¨T+γc2T= 0 , Tn(t) =Cncosncπ /lscriptt+Dnsinncπ /lscriptt. We therefore have found one special solution, un(x,t)−Xn(t)Tn(t) , for each value of the indexn, un(x,t) = sinnπx /lscript(αncosncπt /lscript+βnsinncπt /lscript). The arbitrary constants have been lumped in this equation. These special solutions are the “natural” vibrations of the string, or normal modes of vibration . A snapshot at t=t0of the string moving in the nth normal mode would reveal the sine curve un(x,t0) =Csinnπx /lscript, the constant Caccounting for the remaining terms, which are constant for tfixed. In music, the integer nrefers to the octave. The fundamental tone is the case n= 1 , while the tone for n= 2 , the second harmonic orfirst overtone , is one octave higher. a figure goes here 8.3. THE VIBRATING STRING. 339 The time frequency Vnof thenth normal mode is Vn=ncπ /lscript, this is the number of oscillations in 2 πunits of time. It is the time frequency which we usually associate with musical pitch. The (time) periodτnof thenth normal mode is 2 π/Vn, that isτn= 2l/nc. Another name you will want to know is the wave length λnof thenth normal mode, λn= 2/lscript/n(see figures above). Notice that Vnλn=c, an important relationship. Having found the special normal mode solutions, un(x,t) , we hope that arbitrary constantsαnandβncan be chosen so a linear combination u(x,t) =∞/summationdisplay n=1un(x,t) =∞/summationdisplay n=1(αncosncπt /lscript+βnsinncπt /lscript) sinnπx /lscript will satisfy the given initial conditions. Every function u(x,t) of this form automatically satisfies the boundary conditions u(0,t) = 0, u(/lscript,t) = 0 since each of the un’s satisfy them. Ifu(x,0) =f(x) andut(x,0) =g(x) , then from the above equation, we must have f(x) =∞/summationdisplay n=1un(x,0) =∞/summationdisplay n=1αnsinnπx /lscript and g(x) =∞/summationdisplay n=1∂un ∂t(x,0) =∞/summationdisplay n=1nπc /lscriptβnsinnπx /lscript. Thus, the coefficients αnare the coefficients in the Fourier sine series for f, while the βn are essentially the coefficients in the Fourier sine series for g. In fact, this is how Fourier was led to the series bearing his name. These formulas for u(x,y), f(x) , andg(x) become easier on the eye if the length of the string is π,/lscript=π. Then u(x,y) =∞/summationdisplay n=1(αncosnct+βnsinnct) sinnx, while f(x) =∞/summationdisplay n=1un(x,0) =∞/summationdisplay n=1αnsinnx, (8-3) and g(x) =∞/summationdisplay n=1∂un ∂t(x,0) =∞/summationdisplay n=1ncβnsinnx. Finding the coefficients αnandβnis particularly simple if fandgcan be represented by finite series. Examples: Find the solution u(x,t) of the wave equation for a string of length π, l=π, which is pinned down at its end points, u(0,t) =u(π,t) = 0 , and satisfies the given initial conditions. (1)u(x,0) =f(x) = 2 sin 3x, u t(x,0) =g(x) =1 2sin 4x. We have to find αnandβnfor the two series 2 sin 3x=∞/summationdisplay n=1αnsinnx 340CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS 1 2sin 4x=∞/summationdisplay n=1ncβnsinnx. For these simple functions, just match coefficients, giving α3= 2, αn= 0, n/negationslash= 3,andβ4=1 8c, βn= 0, n/negationslash= 4. Therefore, the sum of the two waves u(x,t) = 2 cos 3ctsin 3x+1 8csin 4ctsin 4x is the (unique!) solution of this example. (2)u(x,0) =f(x) =1 2sin 3x−sin 17xand ut(x,0) =g(x) =−9 sinx+ 13 sin 973 x. We have to find αnandβnfor the two series 1 2sin 3x−sin 17x=∞/summationdisplay n=1αnsinnx and −9 sinx+ 13 sin 973 x=∞/summationdisplay n=1ncβnsinnx. By matching again, we find α3=1 2, α17=−1 , andαn= 0 forn/negationslash= 3 or 17 . Also, β1=9 c,β973=13 973c, andβn= 0 forn/negationslash= 1 or 973 . The (unique) solution is then a sum of four waves u(x,t) =−9 3sinctsinx+1 2cos 3ctsin 3x −cos 17ctsin 17x+13 973csin 973ctsin 973x. Sincefandgare not usually given in the simple form of these examples, the full Fourier series is needed. Recall that the string is pinned down at both ends. Therefore both the initial position function f(x) and velocity function g(x) have the property f(0) =f(π) = 0 , and g(0) =g(π) = 0 , where we have taken the length of the string to beπ. It is now possible to extend both fandg, assumed continuous in [0 ,π] , to the whole interval [ −π,π] as continuous odd functions, a figure goes here that is, ifx∈[0,π] , we can define f(−x) =−f(x) andg(−x) =−g(x), since the right sides, −f(x) and −g(x) , are known functions for x∈[0,π] . As odd functions now on the whole interval [ −π,π] , the functions fandghave Fourier sine series (cf. p. 252, Exercise 3a). 8.3. THE VIBRATING STRING. 341 f(x) =∞/summationdisplay n=1bnsinnx√π g(x) =∞/summationdisplay n=1˜bnsinnx√π where bn= 2/integraldisplayπ 0f(x)sinnx√πdx, ˜bn= 2/integraldisplayπ 0g(x)sinnx√πdx (8-4) Comparing with the previous formulas (3) for fandg, we find αn=bn/√π,andβn=˜bn/nc√π Consequently u(x,t) =∞/summationdisplay n=1(bncosnct√π+˜bn ncsinnct√π) sinnx (8-5) the coefficients bnand˜bnbeing determined from the initial conditions by equation (4). Thus, we have almost proved Theorem 8.12 . Iff(x)is twice continuously differentiable and g(x)once continuously differentiable for x∈[0,π]and both functions vanish at x= 0 andx=π, then the functionu(x,t)defined by equation (5) is a solution of the homogeneous wave equation utt−c2uxx= 0 and satisfies the initial conditions: u(x,0) =f(x), ut(x,0) =g(x), x∈[0,π], as well as the boundary conditions: u(0,t) = 0, u(π,t) = 0, t≥0, wherebnand ˜bnare determined from fandgthrough equations (4). Moreover, this solution is unique (by Theorem 9). Outline of Proof . If it is possible to differentiate the infinite series (5) term by term u(x,t) would satisfy the wave equation since each special solution un(x,t) does. In any case, the initial condition u(x,0) =f(x) is clearly satisfied. However, checking the other initial conditionut(x,0) =g(x) also involves differentiating the infinite series term by term. Thus, we must only justify the term by term differentiation of an infinite Fourier series. For power series, we found (p. 82-3, Theorem 16) we can always differentiate term by term within its disc of convergence. Such is not the case with Fourier series. For example, the Fourier series∞/summationdisplay n=1sinn2x n2converges for all x, but the series obtain by differentiating formally,∞/summationdisplay n=1cosn2xdiverges at x= 0 . However, if a function is sufficiently smooth, its Fourier series can be differentiated term by term and does converge to the derivative of the 342CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS function. Since the details of a complete proof are but a rehash of the proof carried out for power series (p. 82ff), we omit it. Example: Find the displacement u(x,t) of a violin string of length πwith fixed end points which is plucked at its midpoint to height h. The initial position is then f(x) =/braceleftbiggxh, x ∈[0,π/2] (π−x)h, x∈[π/2,π], and the initial velocity, g(x) , is zero. We must find the coefficients bnand˜bnin the series (5). After mentally continuing f andgto the interval [ −π,π] as odd functions, the formulas (4) give us bnand˜bn, bn= 2/integraldisplayπ 0f(x)sinnx√πdx=2h√π/braceleftBigg/integraldisplayπ/2 0xsinnxdx +/integraldisplayπ π/2(π−x) sinnxdx/bracerightBigg . Integrating and simplifying, we find that bn=4h√πn2sinnπ 2=  0, neven 1, n= 1,5,9,13,... −1, n= 3,7,11,15. Fromg(x)≡0 , it is immediate that βn= 0 for all n. Thus, u(x,t) =4h π∞/summationdisplay n=11 n2sinnπ 2cosnctsinnx =4h π[cos 3ctsinx 1−cosctsin 3x 32+cos 5ctsin 5x 52+···] is the desired solution. Exercises (1) (a) Find a solution u(x,t) of the homogeneous wave equation for a string of length πwhose end points are held fixed if the initial position function is u(x,0) =1 2sin 4x−sin 7x, while the initial velocity is ut(x,0) = sin 3x+ sin 73x. (b) Same problem as a), but u(x,0) = sin 5x+ 12 sin 6x−7 sin 9x ut(x,0) =−sinx+ 91 sin 273 x. 8.3. THE VIBRATING STRING. 343 (2) Find a solution u(x,t) of the homogeneous wave equation for a string of length π whose end points are held fixed if the string is initially plucked at the point x=π/4 to the height h. (3) Consider a vibrating string of length /lscriptwhose end points are on rings which can slide freely on poles at 0 and /lscript. Then the boundary conditions at the end points are ux(0,t) = 0, ux(/lscript,t) = 0 that is, zero slope. (a) Use the method of separation of variables to find the form of special standing wave solutions. [Answer: un(x,t) = cosnπx /lscript(αncosncπt /lscript+βnsinncπt /lscript) ]. (b) Use these to find a solution with the initial conditions u(x,0) = cosx−6 cos 3x (let/lscript=π) ut(x,0) =1 2cos 2x. (4) Letu(x,t) satisfy the homogeneous wave equation. Instead of keeping the end points fixed, we either put them on rings (cf. Exercise 3) or attach them by elastic bands, in which case the boundary conditions become ux(0,t)−c1u(0,t) = 0, ux(π,t) +c2u(π,t) = 0, c1,c2≥0. (a) Define the energy as before, and prove that energy is dissipated with these bound- ary conditions, unless c1andc2vanish. (b) Prove there is at most one function u(x,t) which satisfies the inhomogeneous wave equation utt−c2uxx=F(x,t) with initial conditions as before, but with elastic boundary conditions ux(0,t)−c1u(0,t) =φ(t), ux(π,t) +c2u(π,t) =ψ(t), wherec1andc2are non-negative constants. (5) To account for the effect of air resistance on a vibrating string, one common assump- tion is that the resistance on a segment of length ∆ xis proportional to the velocity of its center of gravity, Fres=−k∆xut(˜x,t), k> 0, wherekis a numerical constant. This is analogous to the standard viscous resistance force on a harmonic oscillator. (a) Find the equation of motion ignoring gravity. [Answer:1 c2utt+kut=uxx] (b) Find the form of the special standing wave solutions, assuming, the end points are held fixed. (c) Write a formula giving the probable form for the general solution u(x,t) . (d) If the end points are pinned down, what do you expect the behavior of the string will be as t→ ∞ ? Does the formula found in part c) verify your belief (it should). 344CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS (e) Define the energy E(t) as before and show that energy is dissipated if the ends are held fixed. (f) Use the result of e) to prove ˙E(t) + 2kE(t)≥0 , and conclude that E(t)≥ E(0)e−2ktfort≥0 . This shows that the energy is not dissipated too rapidly. (6) It is possible to write the solution of the homogeneous wave equation for a string of lengthπwith fixed end points in a simple closed form by using the trigonometric identities 2 sinnxcosnct= sinn(x−ct) + sinn(x+ct). 2 sinnxsinnct= sinn(x−ct)−cosn(x+ct). (a) Do this and obtain d’Alembert’s formula u(x,t) =f(x−ct) +f(x+ct) 2+1 2c/integraldisplayx+ct x−ctg(ξ)dξ. (b) Solve the example of a plucked string (p. 641) again using this formula. Draw two sketches, one indicating the position of the string at time t=π 2cand another att=π c. (7) (a) Prove the wave operator L:=∂2 ∂t2−c2∂2 ∂x2,ca constant, is translation invariant, that is, ifT:u(x,t)→u(x+x0, t+t0),prove (LT)u= (TL)ufor all values of x0andt0, and for all functions ufor which the operators make sense. (b) Find the function φ(a,b) in the formula Leax+bt=φ(a,b)eax+bt. (c) Use part b) to show that if ais any constant, the four functions ea(x+ct),e−a(x+ct),ea(x−ct),e−a(x−ct) are solutions of the homogeneous wave equation Lu= 0 . (d) Use the fact that each of the above functions satisfies the ordinary differential equationv/prime/prime(x) =a2v(x) to conclude that if linear combinations of these func- tions are to satisfy the boundary conditions v(0) =v(/lscript) = 0 , then necessarily a2<0 , so the constant ais pure imaginary and we can write a=iγ, whereγ is real. (e) Letu(x,t) be a linear combination of the four functions part c) with a=iγ. Show that u(x,t) may be written in the form u(x,t) = sinγx [Acosγct+Bsinγct]. (f) Ifu(0,t) =u(/lscript,t) = 0 , show that γn=nπ /lscript. Find an infinite set of special solu- tionsun(x,t) which satisfy the homogeneous wave equation with zero boundary values [From here on, one proceeds as before to find the general solution. This problem has shown how the idea of translation invariance can also be used to lead one to the special solutions un]. 8.3. THE VIBRATING STRING. 345 (8) (a) By inspection , find a particular solution for the solution of the inhomogeneous wave equations Lu:=utt−c2uxx=g, g ≡constant. (b) How can this particular solution be used to find the solution of the equation Lu=gwhich has given initial conditions and zero boundary conditions? (9) Flow of heat in a thin insulated rod on the xaxis is governed by the heat equation ut(x,t) =k2uxx(x,t), whereu(x,t) represents the temperature at the point xat timet, andk2, the diffusivity , is a constant depending on the material. The “energy” in a rod of length /lscript,0≤x≤/lscript, is defined as E(t) =1 2/integraldisplayl 0u2(x,t)dx. (a) If the ends of the rod have zero temperature, u(0,t) =u(/lscript,t) = 0 , prove “energy” is dissipated, ˙E(t)≤0 , by showing dE(t) dt=−k2/integraldisplay/lscript 0u2 x(x,t)dx. (b) Given a rod whose ends have zero temperature and whose initial temperature is zero,u(x,0) = 0 , prove that the temperature remains zero, u(x,t)≡0 . (c) Prove the temperature of a rod is uniquely determined if the following three data are known: initial temperature: u(x,0) =f(x), x∈[0,/lscript]. boundary conditions: u(0,t) =φ(t), u(/lscript,t) =ψ(t), t≥0. (d) Use the method of separation of variables to find an infinite number of special solutions of the heat equation for a thin rod whose end points have zero temper- ature for all t≥0 . [Answer: un(x,t) =cne−n2k2π2 /lscript2tsinnπ /lscriptx, n = 1,2,...] (e) If the ends of a rod have zero temperature for all t≥0 , what do you intuitively expect the temperature u(x,t) will be as t→ ∞ ? Is this borne out by the formulas for the special solutions? (f) Find the temperature distribution in a rod of length πif the ends have zero temperature and if the initial temperature distribution in the rod is u(x,0) = sinx−4 sin 7x, (10) If the temperature at the ends of the bar of length /lscriptis constant but not necessarily zero, say u(0,t) =θ1, u (/lscript,t) =θ2, the temperature distribution can be found be splitting the solution into two parts, u(x,t) = ˜u(x,t) +up(x,t) , whereup(x,t) is a particular solution having the correct temperature at the ends of the bar and u(x,t) is a general solution which has zero temperature at the ends. 346CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS (a) Find a particular solution of the homogeneous heat equation ut=k2uxxwhich satisfiesu(0,t) = 200, u(/lscript,t) = 500, but does not necessarily satisfy any prescribed initial condition. [Answer: Many possible solutions - for example up(x,t) = 20 + 30x /lscript, orup(x,t) = 20 + 30 sinπx 2/lscript] . (b) Find the temperature distribution in a rod of length πif the initial temperature isu(x,0) = 2 sinx−sin 4x, while the boundary conditions are as in part a). (11) If the ends of a bar of length /lscriptare insulated instead of being kept at zero, the boundary conditions are ux(0,t) =ux(/lscript,t) = 0. (a) Use the method of separation of variables to find an infinite number of spe- cial solutions for the homogeneous heat equation with insulated ends. [Answer: un(x,t) =cne−n2k2π2 /lscript2tcosnπx /lscript, n= 0,1,2,...]. (b) What is the temperature distribution in a rod whose ends are insulated if the initial temperature distribution is u(x,t) = 3 cos2πx /lscript−1 5cos5πx /lscript. (12) In this exercise you will find a quantitative estimate for the rate of decrease of energy for the heat in a rod of length /lscriptwith zero temperature at the ends. (a) Use the result of Exercise 9a to prove the differential inequality dE dt≤ −cE(t), wherecis a positive constant. [Hint: Look at p. 227 Exercise 15c]. (b) Conclude that E(t)≤E(0)e−ct, t≥0. This is the desired estimate for the decrease of energy in the rod. (13) The linear partial differential equation uxx−u=ut governs the temperature distribution in a rod of length /lscriptmade up of a material which uses up heat to carry out a chemical process. Define the energy E(t) in the rod as in Exercise 9. (a) Prove that if the ends of the rod have zero temperature, then the energy is dissipated, ˙E(t)≤0 . (b) Given a rod whose ends have zero temperature and whose initial temperature u(x,0) is zero, use a) to prove that the temperature remains zero, u(x,t)≡ 0, t≥0 . (c) Use part b) to prove that the temperature of the rod described above is uniquely determined if the following three data are known u(x,0) forx∈[0,/lscript], u(0,t) andu(/lscript,t) fort≥0. 8.4. MULTIPLE INTEGRALS 347 (14) In setting up the mathematical model for the vibrating string, we never examined the horizontal components of the forces. (a) Show that the net horizontal force is Fh=τcosθ2−τcosθ1 (b) Under our assumption uxis small, show that the net horizontal force is zero - so there is no horizontal motion of the string. This justifies the statement that the motion of the string is entirely vertical. (15) Use the formula Vn=nπc//lscript (page 635) for the frequency and the relationship be- tweenc,T andρ(page 624) to derive a formula for Vnin terms of the physical constants/lscript,T, andρfor a vibrating string. Interpret the effect on the frequency, Vn, if the physical constants are changed. Does this agree with your experience in tuning stringed instruments? 8.4 Multiple Integrals How can we extend the notion of integration from functions of one variable to functions of several variables? That is the problem we shall face in this section. Letw=f(X) =f(x1,...,x n) be a scalar-valued function defined in C⊂En. For the purposes of this section it will be convenient to think of fas either the height function for a surface MinEn+1overD, or as the mass density of D. In the first case./integraldisplay/integraldisplay Df should be the volume of the solid contained between MandD(see fig.), whereas in the second case,/integraldisplay/integraldisplay Dfshould be total mass of the set D. Two problems have to be solved. First, define the integral in En. Second, give a reasonable procedure for explicitly evaluating the integral in sufficiently simple situations. More so than for the single integral, the problem of defining the multiple integral bristles with technical difficulties. However, after this is done the evaluation of integrals in Encan be reduced to the evaluation of repeated integrals, that is, a sequence of nintegrals in E1, which is in turn effected not by using the definition of the integral, but rather by recourse to the fundamental theorem of calculus. Before starting the formalities, it is well advised to see where some difficulties lie. Suppose we are given a density function fdefined on some domain Dand want to find the total mass of D. To make things even simpler, assume for the moment that the density is constant and equal to 1, for all X∈D⊂En. Then the mass coincides with the volume of the domain. For the special case of functions of one variable D⊂E1is an interval so the “volume” of D(really the length of D) is trivial a figure goes here to compute, Vol ( D) =b−a. However if Dhas two or more dimensions, even finding the volume ofD(area ifD⊂E2) is itself difficult. The problem is that a connected set DinE1can only be a line segment, whereas a connected open set in En, n≥Zcan be much more complicated topologically. In E1, the 348CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS closed “cube” and closed “ball” are both intervals [ a,b] , and every other connected set is also an interval. In E2, not only do the cube and ball become distinct, but also a slew of other possibilities arise. Dmay be riddled with holes and its a figure goes here boundary wild (contrasted to the boundary of a connected set in E1which is always just two points, the end points of the interval). It should be clear that the notion of volume of a setDmay only be definable if the boundary of Dis sufficiently smooth. As you should be anticipating, the volume of a set Dwill be defined by filling it up with little cubes of volume ∆ x1∆x2...∆xn= ∆V, and then proving that as the size of the cubes becomes small, the sum of volumes of the cubes approaches a limit (here is where the smoothness of θDenters). In two dimensions, D⊂E2, this roughly reads Area (D) = lim ∆x→0 ∆y→0/summationdisplay/summationdisplay ∆x∆y=/integraldisplay Ddxdy. Only after the volume of a domain is defined can the more general notion of mass of a setDfor a density function fbe defined. The procedure here is straightforward, however it is important that the density fbe “essentially” continuous. Using the same approximating cubes, we assign to each little cube its approximate density, say by using the value of the density fat the center of the little cube. Adding up the masses of these little cubes and passing to the limit again, we find the total mass of the solid Dwith density f. Again, in two dimensions this roughly reads Mass (D) = lim ∆x→0 ∆y→0/summationdisplay/summationdisplay f(xi,yj)∆x∆y=/integraldisplay/integraldisplay Df(x,y)dxdy. Because of the technical complications, we shall only state a series of propositions which give the existence of the integral. The proofs of several crucial - but believable - results will not be carried out, but can be found in many advanced calculus books. For convenience, the geometric language of the plane, E2, will be used. The ideas extend immediately to higher dimensions. Now some terminology. Definition: Ashaved rectangle is a rectangle with its bottom and left sides omitted, that is, a set of the form Q={X= (x1,x2):aj<xj≤bj, j= 1,2}. Arectangular complex is a finite union of shaved rectangles, which can always be assumed disjoint, that is, non-overlapping. This should more accurately be called a shaved rectan- gular complex, but is not for the sake of euphony. IfDis a set, the characteristic function of D,XDis defined by XD(X) =/braceleftbigg1, X ∈D 0, X /ownerD. Astep function s(X) is a finite linear combination of characteristic functions of shaved rectangles. The graph of this function looks like its name implies. 8.4. MULTIPLE INTEGRALS 349 a figure goes here A function fhascompact support if it is identically zero outside some sufficiently large rectangle. The support of a particular function f, written supp f, is the smallest closed set outside of which fis zero. Thus, it is the set of all points Xwheref(X)/negationslash= 0 and the limit points of those points. We take the area of a shaved rectangle Qas a known quantity - the height times base, anddefine the integral as I(XQ) =/integraldisplay/integraldisplay E2XQdA=/integraldisplay DdA≡Area (Q), where the Area ( Q) is defined in the natural way as length ×width. You may wish to think ofdAas representing an “infinitesimal element of area”. We however assign no meaning to the symbol and use it only as a reminder. Some prefer to do without it altogether and write /integraldisplay/integraldisplay E2XQ= Area (Q). Our task is to define I(f)≡/integraldisplay/integraldisplay E2fdA for density functions other than XQ’s. For example, if Dis some set, for the function XD we want to define Area (D) =/integraldisplay/integraldisplay E2XDdA=/integraldisplay/integraldisplay DdA But this will not make sense unless it is shown that the set Ddoes have a number associated with it which has the properties of area. It is easy to define the integral of a step function S. Let S(X) =n/summationdisplay j=1ajXQj(X), where the Qj’s are disjoint. Then/integraldisplay/integraldisplay SdA should represent the total mass of a plate composed of rectangles Q1,...,Q nwith respective densities a1,...,a n. Thus, we define I(S) =/integraldisplay/integraldisplay SdA≡a1Area (Q1) +...+anArea/prime,(Qn) =n/summationdisplay j=1aj/integraldisplay XQjdA. The integrals of step functions clearly satisfy the following Lemma 8.13 . IfS1(X)andS2(X)are step functions, then a).I(aS1+bS2) =aI(S1) +bI(S2). b).S1(X)≤S2(X)impliesI(S1)≤I(S2). c). IfS(X)is bounded by M, S (X)≤M, then I(S)≤cM, wherecis the area of the support of S. 350CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS The integral of any other more complicated function is defined by using step functions. Definition: A function f:E2→EisRiemann integrable if given any /epsilon1>0 , there are step functionssandSwiths(X)≤f(X)≤S(X) for allX∈E2such thatI(S)−I(s)</epsilon1, that is /integraldisplay/integraldisplay E2SdA−/integraldisplay/integraldisplay E2sdA</epsilon1. Intuitively, a function is Riemann integrable if it can be trapped between two step functionsSandsin such a way that the integrals of Sandsdiffer by an arbitrarily small amount. Definition: Iffis Riemann integrable, let Snandsnbe a trapping sequence , forf, that is,sn(X)≤f(X)≤Sn(X) andI(Sn)−I(sn)<1 n. Then the Riemann integral of f, I(f) is defined as (cf. page 21, for the definition of l.u.b. = least upper bound, and of g.l.b.). I(f)≡l.u.b. n→∞I(sn) We could have equivalently defined I(f) asI(f) = g.l.b.n→∞I(Sn) . Since both limits are the same, it is irrelevant. However, it is important to show that I(f) has the same value if any other trapping sequence ˆSn(X),ˆsn(X) is used. This is the content of Lemma 8.14 . Iffis Riemann integrable, then I(f)does not depend on which trapping sequences are used. Proof not given. Now we exhibit a class of functions which are Riemann integrable. The issue boils down to finding functions which can be approximated well by step functions. Lemma 8.15 . Iffis a continuous function and Dis a closed and bounded set, then fcan be approximated arbitrarily closely from above and below by step functions Sands throughout D. Thus, given any /epsilon1>0, there are step functions Sandssuch that 0≤S(X)−f(x)</epsilon1, and 0≤f(X)−s(X)</epsilon1 for allX∈D. Proof not given. Theorem 8.16 . Iffis a continuous function with compact support, then it is Riemann integrable. Proof: LetS(X) ands(X) be as in the lemma where Dis the support of f. Then s(X)≤f(X)≤S(X) and S(X)−s(X) = [S(X)−f(X)] + [f(X)−s(X)]<2/epsilon1. Thus by Lemma 1, I(S)−I(s) =I(S−s)<2c/epsilon1, wherecis the area of the set (supp S)∪(supps). Becausefhas compact support, the constant cis bounded. Therefore the factor 2c/epsilon1can be made arbitrarily small by choosing /epsilon1small. This verifies all the conditions for integrability. 8.4. MULTIPLE INTEGRALS 351 We have disposed of the problem of integrating continuous functions with compact support. Notice that the above procedure is identical to that used for functions of one variable (see figure.) We still do not know how to find the area of a domain D. Although we anticipate that Area (D) =I(XD) , this does not yet make sense (except for rectangular complexes) since thediscontinuous functionXDis not covered by Theorem 1). Let us remedy this now. The problem is to show the boundary ∂Ddoes not have any area. Definition: A set in E2hascontent zero if it can be enclosed in a rectangular complex whose total area is arbitrarily small. Thus, if a set has content zero, given any /epsilon1>0 , there is a rectangular complex Rcontaining ∂Dsuch that Area (R) =I(XR)</epsilon1. It should be clear that any set with a finite number of points has content zero (since each point can be enclosed on a square of side /epsilon1, so the total area of Nsuch squares is N/epsilon12, which can be made arbitrarily small.) One would also expect that curves will have zero content. This is not necessarily true unless the curve is not too badly behaved. Lemma 8.17 . If a curve is composed of a finite number of smooth curves, then it has zero content. In particular, if the boundary ∂Dof a bounded domain Dis such a curve, it has zero content. Proof not given. Theorem 8.18 . If the boundary ∂Dof a domain D⊂E2has content zero, then the functionXDis Riemann integrable. Consequently, the area of Dis definable and given by Area (D) =/integraldisplay/integraldisplay E2XDdA=/integraldisplay/integraldisplay DdA. Proof: Almost identical to that for Theorem 11. Let /epsilon1 > 0 be given and let Rbe the rectangular complex which encloses the boundary ∂D, whereRhas area less than /epsilon1, I(XR)< /epsilon1. Then the part of Dwhich is enclosed by R, D −=D−R∩D, is a rectangular complex as is D+=R∪D−andD+−D−=R. SinceD+⊃D⊃D−, we have XD−(X)≤XD(X)≤XD+(X) for all X. Also, I(XD+)−I(XD−) =I(XR)</epsilon1. ThusXDis trapped by the step functions S=XD+ands=SD−andI(S)−I(s)</epsilon1, proving the theorem. It is now possible to define /integraldisplay/integraldisplay DfdA for continuous functions fwhereDis not necessarily the support of f. Theorem 8.19 . Iffis continuous in a closed and bounded set Dwhose boundary ∂D has content zero, then the function fXDis Riemann integrable and /integraldisplay/integraldisplay DfdA≡I(fXD). 352CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS Proof: LetRbe the rectangular complex which encloses ∂Dand has area less than /epsilon1, I(XR)</epsilon1. TakeD−=D−R∩DandD+=R∪D−as in Theorem 12. Further let S1ands1be step functions which trap fwithin/epsilon1for allX∈D−(this is possible by Lemma 3) 0≤S1(X)−f(X)</epsilon1,0≤f(X)−s1(X)</epsilon1 for allX/epsilon1D −, so 0≤S1(X)−s1(X)<2/epsilon1for allX∈D. LetMbe an upper bound for |f|onD,|f(X)| ≤Mfor allX∈D−. Then define S=S1+MX Rands=s1−MX R. These functions Sandstrapfon all ofD, s(X)≤f(X)≤S(X) for all X∈D, that is, s≤fXD≤Sfor allX. Furthermore I(S−s) =I(S1−s1) + 2MI(XR) <2c/epsilon1+ 2M/epsilon1= (2c+ 2M)/epsilon1, wherecis the area of D−. SinceSandsare step functions which trap f, and since I(S−s) can be made arbitrarily small, the proof that fXDis Riemann integrable is completed. We follow custom and write I(fXD)≡/integraldisplay/integraldisplay DfdA. Except for the three unproved lemmas, this completes the proof of the existence of the integral. The next theorem summarizes some important properties of the integral. Theorem 8.20 . Iffandgare Riemann integrable, then a).I(af+bg) =aI(f) +bI(g), a,b constants b).f≤gimpliesI(f)≤I(g). c).|I(f)| ≤I(|f|) Proof: a) and b) are immediate consequences of the corresponding statements for step functions (Lemma 1) and the definition of the Riemann integral as the limit of step functions. To prove c), we first observe that if fis integrable, so is |f|. Since −|f| ≤f≤ |f|, by parts a and b −I(|f|)≤I(f)≤I(|f|), which is equivalent to the stated property. Although the approximate value of the integral/integraldisplay/integraldisplay DfdA can be evaluated by using the procedures of the above theorems, we have as yet no routine way of evaluating the integral 8.4. MULTIPLE INTEGRALS 353 iffandDare simple. Some notation will suggest the method. Write dA=dxdy and think ofdxdy as the area of an “infinitesimal” rectangle. Then /integraldisplay/integraldisplay DfdA =/integraldisplay/integraldisplay Df(x,y)dxdy. IfDis the domain in the figure, it is reasonable to evaluate the double integral, which we shall think of as the mass of Dwith density f, by first finding the mass of a horizontal strip g(y) =/integraldisplayγ2 γ1f(x,y)dx, and then adding up the horizontal strips to find the total mass /integraldisplay/integraldisplay Df(x,y)dxdy =/integraldisplayγ4 γ3g(y)dy=/integraldisplayγ4 γ3/parenleftbigg/integraldisplayγ2 γ1f(x,y)dx/parenrightbigg dy. The integral on the right is called an iterated orrepeated integral. In a similar way, one could begin with mass of vertical strips h(x) =/integraldisplayγ4 γ3f(x,y)dy and add these up /integraldisplay/integraldisplay Df(x,y)dxdy =/integraldisplayγ2 γ1h(x)dy=/integraldisplayγ2 γ1/parenleftbigg/integraldisplayγ4 γ3f(x,y)dy/parenrightbigg dx. For most purposes, it is sufficient to consider domains which are of the two types pictured a figure goes here that is,Dis bounded on two sides by straight line segments. More complicated domains can be treated by decomposing them into domains of these two types, where one or both of the straight line segments might degenerate to a point. Theorem 8.21 . Iffis continuous on a domain D1(respectively D2) as above, then the iterated integral /integraldisplayb a/parenleftBigg/integraldisplayφ2(x) φ1(x)f(x,y)dy/parenrightBigg dx [resp./integraldisplayβ α/parenleftBigg/integraldisplayφ2(y) φ1(y)f(x,y)dx/parenrightBigg dy] exists and equals /integraldisplay/integraldisplay DfdA. Proof not given. It is rather technical. Remark: If a domain Dhappens to be of both types (as, for example, rectangles and triangles are ) then either iterated integral can be used and yield the same result - since they are both equal/integraldisplay/integraldisplay DfdA . See Examples 1 and 3 below (Example 2 could also have been done both ways). Examples: 354CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS (1) Evaluate/integraldisplay/integraldisplay DfdA wheref(x,y) =x2yandDis the rectangle in the figure. We shall integrate with respect to xfirst. /integraldisplay/integraldisplay DfdA =/integraldisplay2 1/parenleftbigg/integraldisplay3 1(x2+xy)dx/parenrightbigg dy. The inner integral is the mass of a strip. Think of yas being the fixed height of the strip. Then /integraldisplay3 1(x2+xy)dx=x3 3+x2y 2/vextendsingle/vextendsingle/vextendsinglex=3 x=1= 9 +9y 2−1 3−y 2=26 3+ 4y Therefore, adding up all the strips we find /integraldisplay/integraldisplay DfdA =/integraldisplay2 1(26 3+ 4y)dy= (26 3y+ 2y2)/vextendsingle/vextendsingle/vextendsingley=2 y=1=26 3+ 6 =44 3 Let us evaluate this again, now integrating first with respect to y. /integraldisplay/integraldisplay DfdA =/integraldisplay3 1/parenleftbigg/integraldisplay2 1(x2+xy)dy/parenrightbigg dx. First/integraldisplay2 1(x2+xy)dy= (x2y+xy2 2)/vextendsingle/vextendsingle/vextendsingley=2 y=1=x2+3 2x so/integraldisplay/integraldisplay DfdA =/integraldisplay3 1(x2+3 2x)dx= (x2 3+3 4x2)/vextendsingle/vextendsingle/vextendsinglex=3 x=1=44 3, which agrees with the previous computation. Instead of imagining fas the density ofD, one can also take fto be the height function of a surface above D. Then the integral/integraldisplay/integraldisplay DfdA is the volume of the solid whose base is Dand whose “top” is the surfaceMwith points ( x,y,f (x,y)) . In this case, the volume is 44/3. (2) Evaluate/integraldisplay/integraldisplay DfdA wheref(x,y) =x2+xy+ 2 andDis the domain bounded by the curves φ1(x) = 2x2, φ2(x) = 4 +x2, andx= 0 . Integrate first with respect to y. Thenyvaries between 2 x2and 4 +x2, whilex varies between the two straight lines x= 0 andx= 2 . /integraldisplay/integraldisplay DfdA =/integraldisplay2 0/parenleftBigg/integraldisplay4+x2 2x2(x2+xy+ 2)dy/parenrightBigg dx =/integraldisplay2 0(x2y+xy2 2+ 2y)/vextendsingle/vextendsingle/vextendsingley=4+x2 y=2x2dy /integraldisplay2 0(8 + 8x+ 2x2+ 4x3−x4−3 2x5)dy=464 15 8.4. MULTIPLE INTEGRALS 355 (3) Evaluate/integraldisplay/integraldisplay DfdA wheref(x,y) = (x−2y)2andDis the triangle bounded by x= 1,y=−2 , andy+ 2x= 6. We shall integrate first with respect to x. Thenxvaries between x= 1 and x=−1 2y+ 2 , while yvaries between the lines y=−2 andy= 2 . /integraldisplay/integraldisplay DfdA =/integraldisplay2 −2/integraldisplay1 2y+2 1(x−2y)2dxdy Since /integraldisplay−1 2y+2 1(x−2y)2(x−2y)2dx=1 3(x−2y)3/vextendsingle/vextendsingle/vextendsinglex=−1 2y+2 x=1=1 3(2−5 2y)3−1 3(1−2y)3, we find /integraldisplay/integraldisplay DfdA =1 3/integraldisplay2 −2[(2−5 2y)3−(1−2y)3]dy=164 3. One can also integrate first with respect to y. Thenyvaries between y=−2 and y=−2x+ 6 , while xvaries between the lines x= 1 andx= 3 . /integraldisplay/integraldisplay DfdA =/integraldisplay3 1/parenleftbigg/integraldisplay−2x+4 −2(x−2y)2dy/parenrightbigg dx. Since /integraldisplay−2x+4 −2(x−2y)2dy=−1 6(x−2y)3/vextendsingle/vextendsingle/vextendsingley=−2x+4 y=−2=−1 6[(5x−8)3−(x+ 4)3] we again find /integraldisplay/integraldisplay DfdA =−1 6/integraldisplay3 1[(5x−8)3−(x+ 4)3]dx=164 3. (4) Find the volume of the pyramid Pbounded by the four planes x= 0,y= 0,z= 0, andx+y+z= 1 . The easiest way to do this is to let z=f(x,y) = 1−x−ybe the height function of the tilted plane which we shall take as the top of the pyramid which lies above the triangle D(in thexyplane) which is bounded by the three linesx= 0,y= 0 , andx+y= 1 . Then Volume (P) =/integraldisplay/integraldisplay Df(x,y)dxdy One can integrate with respect to either xoryfirst. We shall do the xintegration first. /integraldisplay/integraldisplay DfdA =/integraldisplay1 0/parenleftbigg/integraldisplay1−y 0(1−x−y)dx/parenrightbigg dy. Since /integraldisplay1−y 0(1−x−y)dx=−1 2(1−x−y)2/vextendsingle/vextendsingle/vextendsinglex=1−y x=0=1 2(1−y)2 356CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS we find Volume (P) =/integraldisplay/integraldisplay DfdA =1 2/integraldisplay1 0(1−y)2dy=−1 6(1−y)3/vextendsingle/vextendsingle/vextendsingle1 0=1 6. This agrees with the usual formula for the volume of a pyramid Vol =1 3altitude ×area of base . The identical methods work for triple integrals. All of the theorems and proofs remain unchanged. Again the integral /integraldisplay/integraldisplay/integraldisplay DfdV can either be interpreted as the mass of a solid Dwith density f, or as the “vol- ume” of a four dimensional solid whose base is Dand top in the surface with points (x,y,z,f (x,y,z )) . Because of conceptual difficulties, one usually thinks of fas a density. Calculation of triple integrals is done by evaluating three integrals, as /integraldisplay/integraldisplay/integraldisplay DfdV =/integraldisplay/parenleftbigg/integraldisplay/parenleftbigg/integraldisplay f(x,y,z )dz/parenrightbigg dy/parenrightbigg dx, where the limits in the iterated integral on the right are determined from the domain D. An example should illustrate the idea adequately, Example: Evaluate/integraldisplay/integraldisplay/integraldisplay DfdV wheref(x,y,z )≡candDis the solid bounded by the two planes z≡0, y≡2 , and the surface z≡ −x2+y2. We have to evaluate/integraldisplay/integraldisplay/integraldisplay DcdV which is the mass of the solid Dwith constant density c, that isctimes the volume of D. It is convenient to carry out the zintegration first, then the xintegration /integraldisplay/integraldisplay/integraldisplay DcdV =c/integraldisplay2 0/parenleftBigg/integraldisplayy −y/parenleftBigg/integraldisplay−x2+y2 0dz/parenrightBigg dx/parenrightBigg dy. Thexlimits of integration have been found by looking at the region of integration in the xyplane beneath the surface z=−x2+y2. This region, found by setting z= 0 , consists of the points between the straight lines 0 = −x2+y2, that is between the lines x=yand x=−y. Then/integraldisplay/integraldisplay/integraldisplay DfdV =c/integraldisplay2 0/parenleftbigg/integraldisplayy −y(−x2+y2)dx/parenrightbigg dy =c/integraldisplay2 0(−x3 3+xy2)/vextendsingle/vextendsingle/vextendsinglex=y x=−ydy=c/integraldisplay2 04 3dy=16 3c. By letting c= 1 , the volume of the solid is seen to be 16 /3 . Exercises (1) Evaluate/integraldisplay/integraldisplay Dxydxdy for the following domains Din two ways:/integraldisplay (/integraldisplay xydx )dy and/integraldisplay (/integraldisplay xydy )dx. 8.4. MULTIPLE INTEGRALS 357 (a)Dis the rectangle with vertices at (1 ,1),(1,5),(3,1) and (3,5) . (b)Dis the triangle with vertices at (1 ,1),(3,1) and (3,5) . (c)Dis the region enclosed by the lines x= 1, y= 2 , and the curve y=x3(a curvilinear triangle). (d)Dis the region enclosed by the curves y=x2andy=√x. (2) Evaluate/integraldisplay/integraldisplay Dsinπ(2x+y)dxdy, whereDis the triangle bounded by the lines x= 1, y= 2 andx−y= 5 . (3) Evaluate/integraldisplay/integraldisplay D(xy−y3)dxdy, whereDis the region enclosed by the lines x=−1, x= 1, y=−2 and the curve y= 2−x2. (4) Evaluate/integraldisplay D(xy+z)dxdydz, whereDis the rectangular parallelepiped bounded by the six planes x=−2, y= 1, z= 0, x= 1, y= 2, z= 3 . (5) Evaluate/integraldisplay/integraldisplay/integraldisplay Dxyzdxdydz, whereDis the solid enclosed by the paraboloid z=x2+y2and the plane z= 4 . (6) Find the volume of an octant of the ball x2+y2+z2≤a2in two ways; (a) by evaluating/integraldisplay/integraldisplay Df(x,y)dxdy wherefis a suitable function and Da suitable domain (b) by evaluating/integraldisplay/integraldisplay/integraldisplay Ddxdydz, whereDis the ball. (7) Iff(x,y)>0 is the density function of a plate D, thexandycoordinates of the center of mass (¯x,¯y) are defined by x=/integraltext/integraltext Dxf(x,y)dxdy/integraltext/integraltext Df(x,y)dxdy,y=/integraltext/integraltext Dyf(x,y)dxdy/integraltext/integraltext Df(x,y)dxdy. Find the center of mass of a triangle whose vertices are at the points (0 ,0),(0,4) , and (2,0) , and whose density is f(x,y) =xy+ 1 . 358CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS (8) The moment of inertia with respect to a point p= (ξ,η) of a plate Dwith density f(x,y) is defined by Jp(D) =/integraldisplay/integraldisplay D[(x−ξ)2+ (y−η)2]f(x,y)dxdy. (a) Find the moment of inertia of the plate in Exercise 7, with respect to the point p= (1,0) . (b) IfDis any plate (with sufficiently smooth boundary), prove that the moment of inertia is smallest if the point f= (ξ,η) is taken to be the center of mass of D. [Hint: Consider Jas a function of the two variables ξandηand showJ has a minimum at (¯ x,¯y) .] (9) (a) Show that /integraldisplay/integraldisplay Dfxy(x,y)dxdy =f(p1)−f(p2) +f(p3)−f(p4), whereDis a rectangle with vertices at p1,p2,p3,p4(see fig.). (b) Use the result of part (a) to again evaluate the integral in Ex. 1a. (c) IfU(x,y) satisfies the partial differential equation Uxy= 0 for 0<y <x and U(x,x) = 0 while U(x,0) =xsinx, findU(x,y) for all points ( x,y) in the wedge 0<y<x . [Answer: U(x,y) =xsinx−ysinyfor 0<y<x ]. (10) Letf(x,y) be a bounded function which is continuous except as a set of points of content zero, and suppose fhas compact support. Prove that fis Riemann integrable. This again proves Theorem 13. (11) LetD1andD2be domains whose boundaries have zero content and whose intersec- tionD1∩D2has zero content. (a) Iffis continuous on D1∪D2, prove that the integral/integraldisplay/integraldisplay D1∪D2fdA exists and that /integraldisplay/integraldisplay D1∪D2fdA =/integraldisplay/integraldisplay D1fdA +/integraldisplay/integraldisplay D2fdA. (b) Give an example showing the above equality does not hold if D1∩D2has non- zero content. (12) (a) By an explicit construction, show that the region D={(x,y)/epsilon1E2:|x|+|y| ≤1} has boundary with zero content. (b) By an explicit construction, show that the circle ? = {(x,y)/epsilon1E2:x2+y2= 1} has zero content. (13) (a) By interchanging the order of integration, show that /integraldisplayx 0(/integraldisplays 0f(t)dt)ds=/integraldisplayx 0(x−t)f(t)dt. (b)/integraldisplayx 0(/integraldisplay2 0(/integraldisplayr 0f(t)dt)dr)ds=? 8.4. MULTIPLE INTEGRALS 359 (14) LetDbe a plate in the x,yplane with density fand total mass M. Ifp= (ξ,η) is an arbitrary point in the plane and ¯ p= (¯x,¯y) is the center of mass of D, prove Jp(D) =J¯p(D) +M/bardblp−p0/bardbl2, where the notation of Exercise 8 has been used. This is the parallel axis theorem . It again proves the result of Exercise 8b. 360CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS Chapter 9 Differential Calculus of Maps from Ento Em, s. 9.1 The Derivative . Now we generalize the ideas of Chapters 7 and 8 and consider nonlinear mappings from a setDinEntoEm, F:D⊂En→Em, orY=F(X) , whereX∈DandY∈Em. In coordinates, these functions look like y1=f1(x1,...,x n) · · · ym=fm(x1,...,x n) where the functions fjare scalar-valued. The special case n= 1, marbitrary, was treated in Chapter 7, section 3, while the special case m= 1, narbitrary, was treated in Chapter 8. One interpretation of maps F:D⊂En→Emis as a geometric transformation from some subset DofEninto all or part of Em. EXAMPLES. (1) The affine map Y=F(X) defined by y1= 2 +x1−2x2 y2= 1 +x1+x2 maps E2intoE2. Under this map, the origin goes into (2 ,1) , thex1axis (i.e. the linex2= 0 ) goes into the line y1−y2= 1 , a figure goes here 361 362 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S. while thex2axis goes into the line y1+ 2y2= 4 . The shaded region indicates the image of the indicated square. (2) The map Y=F(X) defined by y1=x1−x2 y2=x2 1+x2 2 maps all of E2onto the upper half y1y2plane (since y2≥0 ). Let us see what happens to a rectangle under this mapping. Consider the rectangle Rin the figure. Thex1axis,x2= 0 , goes into the parabola y2=y2 1, and the line x2= 1 into y2= 1 + (y1+ 1)2. a figure goes here Similarly, the line x1= 1 is mapped into y2= 1+(y1−1)2, whilex1= 2 is mapped intoy2= 4 + (y1−2)2. By following the images of the boundary ∂R, we now see that the interior of Ris mapped into the shaded curvilinear “parallelogram”. This mapping, though injective when restricted to our rectangle, is not injective for all (x1,x2)∈E2, since, for example, the points X1= (1,2) andX2= (−2,−1) are both mapped into the same point ( −1,5) . (3) The function w=x2 1+x2 2whose graph is a paraboloid, is a map from E2into E1. It can also be regarded as a map from E2intoE3by a useful artifice. Let y1=x1, y2=x2, andy3=w=x2 1+x2 2. Then y1=x1 y2=x2 y3=x2 1+x2 2 is a mapFfrom E2intoE3. The image of the unit square (see figure) is then the shaded region in the figure above the image ( y1,y2) of the square R a figure goes here (4) The map F:E2→E3defined by (cf. example 2) y1=x1−x2 y2=x2 1+x2 2 y3=x1+x2 also represents a surface M. In fact, since y2 1+y2 3= 2y2, this surface is a paraboloid opening out on the y2axis. Again, we investigate where the rectangle Rof example 2 is mapped. Since the y1andy2components of the mapping are the same as before, the image ofRwill lie on the surface Mabove the image ( y1,y2) of (x1,x2) . Thus the image of the rectangle Ris a patch of the surface M. 9.1. THE DERIVATIVE 363 From these examples, we see it is natural to regard any map F:D⊂E2→Em as an ordinary surface, or two dimensional manifold, embedded in Em, much as a map F:D⊂E1→Emwas regarded as an ordinary curve. In the case m= 1 , the surface F:D⊂E2→E1was representable as the graph of the function F. Form= 2 and higher, this surface is seen as the range of the map. In the same way, an ndimensional surface, or manifold, embedded in IEmis a mapF:D⊂En→Em. You might want to think ofnas being the number of “degrees of freedom” on the manifold. In a strict sense, the mapF:D⊂En→Emis not annmanifold embedded inEmunless Emis big enough to hold an mmanifold, i.e. m≥n. However by either using the graph of F, a subset of Em+n, or by using the trick of example 3 we can always think of the map F:En→Emas anndimensional surface. For m≥n, this surface can be embedded as a subset of En. There are several valuable physical interpretations of these vector valued functions of a vector,Y=F(X) . Consider a fluid flowing through a domain DinE3. The fluid could be air and Das the outside of an airplane, or the fluid could be an organic fluid, and D as some portion of the body. The velocity Vof a particle of fluid is a three vector which depends upon the space co- ordinate (x1,x2,x3) as well as the time coordinate tof the particle, V=F(x1,x2,x3,t) = F(X,t) . This velocity vector V(X,t) atXpoints in the direction the fluid is moving. Thus, the velocity function is an example of a mapping from space-time E3×E1∼=E4into vectors in E3. In this case, we think of the velocity vector V=F(X,t) as having its foot at the point X∈Dand imagine the mapping as the domain Dalong with a vector V attached to each point of D(see fig. above). One calls this a vector field defined on the domainD, since it assigns a vector to each point of D. A very common vector field is a field of forces. By this we mean that to every point Xof a domain D, we associate a vector F(X) equal to the force an object at X“feels”. If the forces are time dependent, then the force field is written F(X,t), X∈D. You are most familiar with the force field due to gravity. If e3is the direction toward the center of the earth, and say e1points east and e2north along the surface of the earth (other coordinates must be chosen for the north and south poles), then the gravitational force is usually written as F= (0,0,g) , a constant vector pointing down to the center of the earth. For more precise purposes, one must take into account the fact that gdoes vary from place to place of the earth’s surface. Then F(x) = (0,0,g(X)) . In even more accurate experiments - or in outer space - must further account for the effect of the other heavenly bodies. This brings in the other components of force as well as a time dependence due to the motion of the earth, F(X,t) = (f1(X,t), f2(X,t), f3(X,t)) . The force field is imagined as a vector attached to each point Xin space, the vector having the magnitude and direction of the net force Fthere. An entirely different example of a mapping Ffrom EntoEmis a factory - or an even larger economic system. The vector X= (x1,x2,...,x n) might represent the quantities x1,x2,... of different raw materials needed. Y=F(X) could then represent the output from the factory, the number yjbeing the quantity of the jth product produced from the inputX. Turning to the quantitative mathematical aspect of the mappings F:En→Em, we define the derivative. The definition will be formal, patterned directly on the definition of the total derivative given previously (p. 578-9). Definition: LetF:D⊂En→EmandX0be an interior point of D.Fisdifferentiable 364 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S. atX0, if there exists a linear transformation L(X0):En→Em, depending on the base pointX0, such that lim /bardblh/bardbl→0/bardblF(X0+h)−F(X0)−L(X0)h/bardbl /bardblh/bardbl= 0 for any vector hin some sufficiently small ball about X0. IfFis differentiable at X0, we shall use the notationsdF dX(X0) =F/prime(X0) =L(X0) and refer to them as the derivative of FatX0. IfF/prime(X0) depends continuously on the base point X0for allX0inD, thenFis said to be continuously differentiable inD, writtenF∈C1(D) . Many of the results from Chapter 8 Sections 1 and 2 generalize immediately to the present situation. Proposition 9.1 . The function F:D⊂En→Emis differentiable at the interior point X0∈Dif and only if there is a linear operator L(X0):En→Emand a function R(X0,h) such that F(X0+h) =F(X0) +L(X0)h+R(X0,h)/bardblh/bardbl, where the remainder R(X0,h)has the property lim /bardblh/bardbl→0/bardblR(X0,h)/bardbl= 0. Proof: ⇐IfFis differentiable at X0, letL(X0)be the derivative and take R(X0,h) = [F(X0+h)−F(X0)−L(X0)h]//bardblh/bardbl. Then this L(X0)andR(X0,h) do satisfy the above conditions. ⇒IfL(X0)andR(X0,h) are as above, then lim /bardblh/bardbl→0/bardblF(X0+h)−F(X0)−L(X0)h/bardbl /bardblh/bardbl= lim /bardblh/bardbl→0/bardblR(X0,h)/bardbl= 0. SinceL(X0)is linear, this proves Fis differentiable at X0. There is at most one derivative operator L(X0), that is Proposition 9.2 . (Uniqueness of the derivative). Let F:D⊂En→Embe differentiable at the interior point X0∈D. IfˆL(X0)and ˜L(X0)are linear operators both of which satisfy the conditions for the derivative of FandX0, then ˆL(X0)=˜L(X0). Proof: Word for word the same as the proof of Theorem 1, page 579-80. If the map F=F(X) is given in terms of coordinates, y1=f1(x1,...,x n) y2=f2(x1,...,x n) · · · · · · ym=fm(x1,...,x n), how is the derivative computed, and what is its relationship to the derivative of the indi- vidual coordinate functions fj? The answer is contained in 9.1. THE DERIVATIVE 365 Theorem 9.3 . LetFmapD⊂EnintoEmbe given in terms of the coordinate functions fj(X), j = 1,...,m y1=f1(X)f1(x1,...,x n) · · · ym=fm(X) =fm(x1,...,x m). (a) ThenFis differentiable or continuously differentiable at the interior point X0∈D if and only if all of the fj’s are respectively differentiable or continuously differentiable. (b) Moreover, if Fis differentiable at X0, then the derivative in these coordinates is given by the m×nmatrix of partial derivatives L(X0):=F/prime(X0) = f/prime 1(X0) · · · f/prime m(X0) = ∂f1 ∂x1(X0),...,∂f1 ∂xn(X0) · · · ∂fm ∂x1(X0),...,∂fm ∂xn(X0) . The matrix is sometimes called the Jacobian matrix. Proof: (a) Observe that the limit lim /bardblh/bardbl→0/bardblF(X0+h)−F(X0)−L(X0)h/bardbl /bardblh/bardbl= 0 exists if and only if each of its components tend to zero, lim /bardblh/bardbl→0/bardblfjX0+h)−fj(X0)−Lj(X0)h/bardbl /bardblh/bardbl= 0, j = 1,2,...,m. Thus, ifFis differentiable at X0, each of the coordinate functions fjare differentiable and have total derivative Lj(X0). Conversely, if each of the coordinate functions are dif- ferentiable at X0, all of the above limits exist so the vector valued function Fis also differentiable. (b) Since the differentiability of Fimplies that of the coordinate vectors, we have F/prime(X0) = f/prime 1(X0) · · · f/prime m(X0) . The result now follows by writing out each of the derivatives f/prime 1(X0) = (∂f1(X0) ∂x1,...,∂f1(X0) ∂xn) f/prime 2(X0) =...etc. and then inserting these in the expression for F/prime(X0) . Corollary 9.4 . A function F:D⊂En→Emis continuously differentiable in Dif and only if all the partial derivatives of its components ∂fi/∂x jexist and are continuous. 366 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S. Proof: This follows from this theorem and Theorem 3, p. 585. EXAMPLES. 1. LetFbe an affine map from EntoEm F(X) =Y0+BX, whereBis a linear operator from EntoEm(which you may choose to think of as an m×nmatrix with respect to some coordinate system) and Y0=F(0) is a fixed vector in Em. ThenFis differentiable at every point of Enand it given by the eminently reasonable formula F/prime(X0) =B, where the operator Bdoes not depend on X0. For proof, we observe that F(X0+h)−F(X0) =Y0+B(X0+h)−[Y0+BX 0] =Bh. Thus lim /bardblh/bardbl→0/bardblF(X0+h)−F(X0)−Bh/bardbl /bardblh/bardbl= lim /bardblh/bardbl→00 /bardblh/bardbl= 0. SinceBis linear, this shows the derivatives exists and is B. Let us do this again in coordinates. If B= ((bij)) the function Fis f1(X) =y01+b11x1+b12x2+...+b1nxn f2(X) =y02+b21x1+... +b2nxn · · · fm(X) =y0m+bm1x1+... +bmnxn. Therefore each of the functions fjis clearly differentiable and f/prime 1= (∂f1 ∂x1,...,∂f1 ∂xn) = (b11,...,b 1n) · · · · · · · · · f/prime m= (∂fm ∂x1,...,∂fm ∂xn) = (bm1,...,b mn). Consequently, F/prime(X0) = f/prime 1(X0) · · · f/prime m(X0) = b11, ..., b 1m · · · bm1, ..., b mn =B, which agrees with the result obtained without coordinates. 2. LetF:E2→E3be defined by f1(x1,x2) = 2−x1+x2 2 9.1. THE DERIVATIVE 367 f2(x1,x2) =x1x2−x3 2 f3(x1,x2) =x2 1−3x1x2. Since each of the coordinate functions fjare continuously differentiable, so is F. Because f/prime 1(X) = (−1,2x2), f/prime 2(X)−(x2,x1−3x2 2), f/prime 3(X) = (2x1−3x2,−3x1), we find that at X0= (3,1) F/prime(X0) = f/prime 1(X0) f/prime 2(X0) f/prime 3(X0) = −1 2 1 0 3−9 . IfXis nearX0, then by Proposition 1 with h=X−X0 F(X) =F(X0) +f/prime(X0)(X−X0) + remainder = 0 2 3 + −1 2 1 0 3−9 /parenleftbiggx1−3 x2−1/parenrightbigg + remainder , where the remainder term becomes less significant the closer Xis toX0. Motivated by our previous work, it is natural to formally define the tangent map as follows. Definition: LetF:D⊂En→Embe differentiable at the interior point X0∈D. The tangent map atF(X0) to the (hyper) surface defined by Fis defined to be the affine mapping Φ(X) =F(X0) +f/prime(X0)(X−X0). Examples: (1) LetFbe the function of Example 2 above. Then the tangent map at X0= (3,1) is Φ(X) = 0 2 3 + −1 2 1 0 3−9 /parenleftbiggx1−3 x2−1/parenrightbigg . (2) LetFbe the function of Example 4 (page 679). Then F/prime(X) = 1−1 2x12x2 1 1 . Thus the tangent map at (2 ,1) is Φ(X) = 1 5 3 + 1−1 4 2 1 1 /parenleftbiggx1−2 x2−1/parenrightbigg If we letY= Φ(X) , then the target plane in the tangent space is found from y1= 1 + (x1−2)−(x2−1) y2= 5 + 4(x1−2) + 2(x2−1) y3= 3 + (x1−2) + (x2−1) By eliminating x1andx2from these equations, we find y2=−5 +y1+ 3y3. A graph of the surface Mand the tangent plane can now be drawn. 368 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S. a figure goes here The next result is the generalization of the mean value theorem. Theorem 9.5 . (Mean Value Theorem). Let F:D⊂En→Embe differentiable at every point ofD, whereDis an open convex set in En. IfF/prime(X)is bounded in D, that is, if there is a constant γ <∞such that/vextendsingle/vextendsingle/vextendsingle∂fi ∂xj(X)/vextendsingle/vextendsingle/vextendsingle≤γfor allX∈Dand for all i= 1,...,m , andj= 1,...,n , then /bardblF(X2)−F(X1)/bardbl ≤c/bardblX2−X1/bardbl for allX1andX2inD, whereC=√nmγ . Proof: The idea is to use the components of Fand to appeal to the similar theorem (p. 597-8) for the function from En→E1. By that theorem, if X1andX2are inD, then there is a point Z1on the line segment joining X1toX2such that f1(X2) =f1(X1) +f/prime 1(Z1)(X2−X1), and similarly for the other components f2,f3,...,f m. Thus  f1(X2) · · · fm(X2) = f1(X1) · · · fm(X1) = f/prime 1(Z1) · · · f/prime m(Zm) (X2−X1), whereZ1,...,Z mare all on the segment joining a figure goes here X1toX2. Observe that the f/prime j(Zj) ’s are all vectors. Let Lbe the matrix of derivatives in the last term above, that is L= f/prime 1(Z1) · · · f/prime m(Zm) = ∂f1 ∂x2(Z1)···∂f1 ∂xn(Z1) · · · ∂fm ∂x1(Zm)···∂fm ∂xn(Zm) . The above equation then reads F(X2) =F(X1) +L(X2−X1). (9-1) This equation itself is sometimes referred to as the mean value theorem. Note, however, that the partial derivatives in Larenotall evaluated at the same point. Since/vextendsingle/vextendsingle/vextendsingle∂fi ∂xj(X)/vextendsingle/vextendsingle/vextendsingle≤γfor allX, ifηis any vector in En, by Theorem 17, p. 373. we find that /bardblLη/bardbl ≤√nmγ/bardblη/bardbl. Takingη=X2−X1, and using (1), we are led to the inequality /bardblF(X2)−F(X1)/bardbl ≤√nmγ/bardblX2−X1/bardbl, 9.1. THE DERIVATIVE 369 which holds for any points X1andX2inD. WithC=√nmγ , this is the desired inequality. A few heuristic remarks. We have been considering mappings F:En→Em. In the case of linear mappings, L:En→Em, it was possible to prove that the range of Lhad dimension no greater than n, dim R(L)≤n. Although this does not remain true for an arbitrary nonlinear map F, it is still true if Fis differentiable - after a suitable definition of dimension for an arbitrary point set is made (for the range of Fwill not usually be a linear space, the only sets whose dimension we have so far defined). In the case of differentiable mapsF, it is easy to make a reasonable definition of dimension. The idea is to define dimension of the range of Flocally, that is, in the neighborhood of every point in the range. IfF:D⊂En→EmandFis differentiable at X∈D, then for all hsufficiently small, F(X+h) =F(X) +L(X)h+ remainder . Thedimension of the range ofFatF(X) is defined to be the dimension of its affine part, which is the same as dim bR(L(X)) . SinceL(X)is a linear operator, its range has a well defined dimension. Geometrically, we have defined dimension of the range of FatF(X) as the dimension of the tangent plane at F(X) . Our definition makes good physical sense for it is exactly the number an insect on the surface would use for the dimension. The illustration below is for a map F:DE2→E3whose range has dimension 2, a figure goes here Some special remarks should be made about maps from one space into another of the same dimension, F:DEn→En. Let us assume Fis differentiable throughout D. Then the dimension of the range of FatF(X), X∈D, is the dimension of the range of L(X)=F/prime(X) . IfFis to preserve dimension at every point, then we must have dim R(L(X)) =nfor allX∈D. For maps Fgiven in terms of coordinates, this means the determinant of the n×nmatrixL(X) does not vanish, detL(X)= detF/prime(X)/negationslash= 0 for allx∈D. In more conceptual terms, this states that a map F:D⊂En→Enis dimension preserving at X0∈Dif its “affine part” Φ( X0+h) =F(X0) +F/prime(X0)his dimension preserving at X0(there is no trouble with the constant vector F(X0) since it only represents a translation of the origin - which does not affect dimensionality). From the geometric interpretation of determinants as volume, we see that the condition detF/prime(X0)/negationslash= 0 means that if a small set S⊂Dhas non-zero volume, then its image F(X) also has non-zero volume. In fact, we expect that if Sis a small set about X, then Vol (F(S)) =/vextendsingle/vextendsingledetF/prime(X0)/vextendsingle/vextendsingleVol (S). Our expectation is based upon the realization that if the points of Sare all near X0, then Fwill behave like its affine part, ( X0+h) =F(X0) +F/prime(X0)h, on the points X0+h∈S. The above formula is a restatement of the effect of affine maps on volume (Corollary to Theorem 30, page 426). We shall return to this later (Chapter 10, Section 4). 370 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S. Because of its frequent appearance, det F/prime(X) has a name of its own. It is called the Jacobian determinant or just the Jacobian ofF. IfFis given in terms of coordinates, y1=f1(x1,...,x n) · · · yn=fn(x1,...,x n), then another common notation for the Jacobian is detF/prime(X) =∂(f1,f2,...,f n) ∂(x1,x2,...,x n). For these maps Ffrom a space into one of the same dimension, F:D⊂En→En, there is a very special derivative which appears often. It is the sum of the diagonal elements of the derivative matrix F/prime(X) . One writes this expression as ∇.For÷F, the divergence ofF, ∇ ·F(X) = divF(X) =∂f1(X) ∂x1+∂f2(X) ∂x2+···+∂fn(X) ∂xn For example, if Y=F(X) is defined by y1=x1+ 2x1x2 y2=x2 1−3x2, then F/prime(X) =/parenleftbigg1 + 2x22x1 2x1 −3/parenrightbigg and ∇ ·F(X) = divF(X) = (1 + 2x2) + (−3) =−2 + 2x2. The significance of the divergence will become clear later (Chapter 10, Section 2). You will probably find it helpful to think of ∇as the operator ∇= (∂ ∂x1,···,∂ ∂xn). Then ∇ ·Fis the “scalar product” of the operator ∇with the vector F. EXERCISES. (1) (a) Find the derivative matrix at the given point for the following mappings Y= F(X) . (i)y1=x2 1+ sinx1x2 y2=x2 2+ cosx1x2atX0= (0,0) (ii)y1=x2 1+x3ex2−x3 2 y2=x1−3x2+x1logx3 y3=x2+x3 y4= 5x1x2x3atX0= (2,0,1) (b) Find the equation of the tangent plane to the above surfaces at the given point. 9.1. THE DERIVATIVE 371 (2) Consider the following map from E2→E2, /braceleftbiggu=excosy v=exsiny (a) Find the image of the following regions i)x≥0,0≤y≤π 4 ii)x≥0,0≤y≤π iii)x≤0,0≤y≤2π iv) 1<x< 2,π 6≤y≤π 3. (b) Compute the derivative matrix and its determinant. (3) IfF:D⊂En→Emis differentiable at X0∈D, prove it is then also continuous at X0. (4) LetFandGboth mapD⊂En→Em, so the function f(X) =/angbracketleftF(X), G(X)/angbracketrightis defined for all X∈Dandf:D→E1. (a) IfFandGare differentiable in D, provefis also, and that f/prime=F/primeG+G/primeF (b) Apply this result to the function f(X) =/angbracketleftX, AX /angbracketright −2/angbracketleftX, Y/angbracketright, whereAis a constant linear operator from En→EnandYis a constant vector inEn. How does the result simplify if Ais self adjoint? (5) Ifϕ:D⊂En→E1andF:D⊂En→Em, then the function G(X) :=ϕ(X)F(X) is defined for all x∈DandG:En→Em. (a) Letϕ(x2,x2) =ax1+bx2andF(x1,x2) = (αx1+βx2,γx 1+δx2) . LetG=ϕF and compute G/prime(X) . (b) More generally, prove that if ϕandFare differentiable in D, thenG:=ϕF is also differentiable and find a formula for G/prime. IfFis expressed in terms of coordinate functions, F= (f1,f2,...,f m) , how does your formula read? Check the result with that of part (a). (6) (a) If F:D⊂En→Emis differentiable in the open connected set D, and if F/prime(X)≡0 for allx∈D, prove that Fis a constant vector. (b) IfFandGmapD⊂En→Emare differentiable in the open connected set D, and ifF/prime(X)≡G/prime(X) for allx∈D, what can you conclude? (7) Consider the map F:Q→R3defined by F:x= (a+bcosϕ) cosθ y= (a+bcosϕ) sinθ z=bsinϕ 372 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S. a figure goes here (a) Compute F/prime. (b) Find the equation of the tangent map at (0 ,0) and at ( π/2,π/2) . (c) Determine the range of the tangent map at the above two points and indicate your findings in a sketch. 9.2. THE DERIVATIVE OF COMPOSITE MAPS (“THE CHAIN RULE”). 373 9.2 The Derivative of Composite Maps (“The Chain Rule”). Consider the two mappings F:A⊂En→EmandG:B⊂Em→Er. Then the composite map H:=G◦F:A⊂En→Eris defined if Bcontains the image of all the points from A, F (A)⊂B. a figure goes here The mapH=G◦Ftakes points from A⊂Enand sends them into Er. From knowledge of the derivatives of FandG, it is possible to compute the derivative of the composite map G◦F. Theorem 9.6 . LetF:A⊂En→EmandG:B⊂Em→Erbe differentiable maps defined in the open sets AandB, respectively, with F(A)⊂B(so the composite map H(X) := (G◦F)(X)is defined for all X∈A). IfX0∈A, letY0=F(X0)∈B. Then the composite map His differentiable at X0and H/prime(X0) =G/prime(Y0)◦F/prime(X0). Remark: The multiplication G/prime◦F/primeis the multiplication of the linear operators G/primeand F/prime. IfFandGare given in terms of coordinates, then the formula is just the product of two matrices G/primeandF/prime. Before proving this theorem, we shall illustrate its meaning. Example: LetF:E2→E2andG:E2→E3be defined by Y=F(X) andZ=G(Y) as follows/braceleftbiggy1=x1−x2 2 y2=x2sinπx1  z1=y1y2 z2= 1 +y2 1+y2 z3= 5−y3 2. Then F/prime(X) =/parenleftbigg1 −2x2 πx2cosπx1sinπx1/parenrightbigg , G/prime(X) = y2y1 2y1 1 0−3y2 2 . AtX0= (3,2) , we find Y0=F(X0) = (−1,0) . Thus F/prime(X0) =/parenleftbigg1−4 −2π 0/parenrightbigg , G/prime(Y0) = 0−1 −2 1 0 0 . IfH(X) = (G◦F)(X) =G(F(X)) , then the derivative of HatX0is H/prime(X0) =G/prime(Y0)◦F/prime(X0) = 0−1 −2 1 0 0 /parenleftbigg1−4 −2π 0/parenrightbigg = 2π 0 −2−2π8 0 0 . 374 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S. The derivative could also have been found in a longer way by explicitly finding Z=H(X) from the formulas for FandG z1=y1y2= (x1−x2 2)(x2sinπx1) z2= 1 +y2 1+y2= 1 + (x1−x2 2)2+x2sinπx1 z3= 5−y3 2= 5−(x2sinπx1)3 and now directly computing H/prime(X0) . Proof of Theorem . SinceFis differentiable at X0∈A⊂EnandGis differentiable atY0∈B⊂Er, for all sufficiently small vectors h∈Enandk∈Em, we can write F(X0+h) =F(X0) +F/prime(X0)h+R1(X0,h)/bardblh/bardbl G(Y0+k) =G(Y0) +G/prime(Y0)k+R2(Y0,k)/bardblk/bardbl where lim /bardblh/bardbl→0/bardblR1(X0;h)/bardbl= 0 and lim /bardblk/bardbl→0/bardblR2(Y0,k)/bardbl= 0. Consequently, since H(X) := (G◦F)(X) =G(F(X)) , H(X0+h) =G(F(X0+h)) =G(F(X0) +F/prime(X0)h+R1(X0;h)/bardblh/bardbl =G(F(X0)) +G/prime(Y0)F/prime(X0)h+R3(X0,h)/bardblh/bardbl, where R3(X0;h) =G/prime(Y0)R1(X0;h) +R2(Y0,k)/bardblk/bardbl /bardblh/bardbl, and k=F/prime(X0)h+R1(X0;h)/bardblh/bardbl. Thus, for all sufficiently small h, H(X0+h) =H(X0) +G/prime(Y0)F/prime(X0)h+R3(X/prime 0,h)/bardblh/bardbl. BecauseG/prime(Y0) andF/prime(X0) are linear maps, so is their product. Therefore we are done if we prove lim /bardblh/bardbl→0/bardblR3(X0;h)/bardbl= 0 . By the triangle inequality /bardblR3(X0;h)/bardbl ≤ /bardblG/prime(Y0)R1(X0;h)/bardbl+/bardblR2(Y0,k)/bardbl/bardblk/bardbl /bardblh/bardbl. Since for fixed X0, the operators F/prime(X0) andG/prime(Y0) are constant operators, by Theorem 17, p. 373, there exist constants αandβsuch that for any vectors ξ∈Enandη∈Em, /bardblF/prime(X0)ξ/bardbl ≤α/bardblξ/bardbland /bardblG/prime(Y0)η/bardbl ≤β/bardblη/bardbl. This means /bardblk/bardbl ≤ /bardblF/prime(X0)h/bardbl+/bardblR1(X0;h)/bardbl/bardblh/bardbl ≤(α+/bardblR1(X0;h)/bardbl)/bardblh/bardbl 9.2. THE DERIVATIVE OF COMPOSITE MAPS (“THE CHAIN RULE”). 375 and /bardblG/prime(Y0)R1(X0;h)/bardbl ≤β/bardblR1(X0;h)/bardbl. Thus, /bardblR3(X0;h)/bardbl ≤β/bardblR1(X0;h)/bardbl+ (α+/bardblR1(X0;h)/bardbl)/bardblR2(Y0,k)/bardbl Now, as /bardblh/bardbl → 0 , so does /bardblk/bardbl ≤(α+/bardblR1(X0;h)/bardbl)/bardblh/bardbl. From the definition of R1and R2, this implies /bardblR3(X0;h)/bardbl →0 as/bardblh/bardbl →0 and completes the proof. Incidentally, if one writes Y=F(X) andZ=G(Y) , then the chain rule can be written in the form d dx(G◦F) =dG dY◦dY dX, which could hardly be more simple to remember. For the balance of this section, we shall work out a few more illustrations showing how the chain rule is applied in different concrete situations. We isolate the next example as an important Corollary 9.7 . LetF:D⊂En→Emand the scalar valued function g:Em→E1 both satisfy the hypotheses of Theorem 1. If we write Y=F(X)in coordinates F= (f1,f2,...,f m), and leth=g◦F, then ∂h ∂x1=∂g ∂y1∂f1 ∂x1+∂g ∂y2∂f2 ∂x1+···+∂g ∂ym∂fm ∂x1 · · · ∂h ∂xn=∂g ∂y1∂f1 ∂xn+∂g ∂y2∂f2 ∂xn+···+∂g ∂ym∂fm ∂xn Remark: This is the chain rule for scalar-valued functions. Proof: By Theorem 3, dh dX=dq dYdF dX Since dq dY= (∂g ∂y1,···,∂g ∂ym) and dF dX= ∂f1 ∂x1···∂f1 ∂xn · · · ∂fm ∂x1···∂fm ∂xn , we find upon multiplying the matrices that dh dX= (m/summationdisplay j=1∂g ∂yj∂fj ∂x1,m/summationdisplay j=1∂g ∂yj∂fj ∂x2,···,m/summationdisplay j=1∂g ∂yj∂fj ∂xn). But we also know dh dX= (∂h ∂x1,∂h ∂x2,···∂h ∂xn). 376 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S. Comparison of the last two formulas gives the stated result. EXAMPLE. Let F:E2→E2andg:E2→E1be defined by /braceleftbiggf1(x1,x2) =x1−ex2, g(y1,y2) =y2 1+y1y2. f2(x1,x2) =ex1+x2 Then F/prime(X) =/parenleftbigg1−ex2 ex1 1/parenrightbigg , g/prime(Y) = (2y1+y2,y1). Ifh=g◦F=g(F(x1,x2)) , then dh dX= (2y1+y2,y1)/parenleftbigg1−ex2 ex1 1/parenrightbigg = (2y1+y2+y1ex1,−(2y1+y2)ex2+y1). In particular, we find ∂h ∂x1= 2y1+y2+y1ex1 and ∂h ∂x2=−(2y1+y2)ex2+y1. These formulas could also have been found by directly applying the corollary, viz. ∂h ∂x1=∂g ∂y1∂f1 ∂x1+∂g ∂y2∂f2 ∂x1= (2y1+y2)1 +y1(ex1), and similarly for ∂h/∂x 2. Many applications of the chain rule are more complicated. Consider a real valued functiong(x1,x2,x3,t) , which depends on the point ˜X= (x1,x2,x3) as well as t. The functiongcould be an expression of the temperature at a point ˜Xat timet. If the point ˜Xrepresents your position in the room, then since you move around the room, ˜Xis itself a function of t. Thus, if your position is specified by ˜X=˜F(t) , x1=f1(t), x 2=f2(t), x 3=f3(t), the temperature where you stand is h(t) =g(f1(t),f2(t),f3(t),t) . Since ˜F:E1→E3while g:E4→E1, the chain rule is not directly applicable because gis defined on E4, while the image of ˜Fis inE3. A simple - if artificial - device clears up the difficulty. Introduce another variable x4 and letX= (x1,x2,x3,x4) . Then write g(x1,x2,x3,x4) , as well as X=F(t) , with x1=f1(t), x 2=f2(t), x 3=f3(t), x 4=f4(t)≡t. Now, as before, h(t) =g(f1(t),f2(t),f3(t),t) , butF:E1→E4andg:E4→E1. The chain rule is thus applicable and gives dh dt=dg dXdF dt 9.2. THE DERIVATIVE OF COMPOSITE MAPS (“THE CHAIN RULE”). 377 = (∂g ∂x1,∂g ∂x2,∂g ∂x3,∂g ∂x4 d f1 dtd f2 dtd f3 dt 1 , so thatdh dt=∂g ∂x1∂f1 ∂t+∂g ∂x2∂f2 ∂t+∂g ∂x3∂f3 ∂t+∂g ∂x4. Sincex4≡t, the last equation can also be written as dh dt=∂g ∂x1df1 dt+∂g ∂x2∂f2 ∂t+∂g ∂x3df3 dt+∂g dt. From a less formal viewpoint, this could have been obtained directly from the equation h(t) =g(f1(t),f2(t),f3(t),t) without dragging in the artificial auxiliary variable x4. The variablex4has been introduced to show how the chain rule applies. Once the process is understood, the variable x4can (and should) be omitted. EXAMPLE. Let g(x1,x2,x3,t) =x1t+ 3x2 2−x1x3+4 1+t2, and letx1= 3t−1, x2= et−1, x3=t2−1 . Ifh(t) =g(x1(t), x2(t), x3(t),t) , we find dh dt=∂g ∂x1∂x1 dt+∂g ∂x2dx2 ∂t+∂g ∂x3dx3 dt+∂g ∂t. = (t−x3)3 + (6x2)et−1−(x1)2t+x1−8t (1 +t2)2. In particular, at t= 1 , we have x1= 2, x2= 1, x3= 0 so that dh(1) dt= (1−0)3 + (6)1 −(2)2 + 2 −8 4= 5. It is straightforward to compute the second derivative d2h/dt2from the formula for the first derivative. d2h dt2=∂ ∂x1(dg dt)dx1 dt+∂ ∂x2(dg dt)dx2 dt+∂ ∂x3(dg dt)dx3 dt+∂ ∂t(dg dt). For this example, this gives d2h dt2= (−2t+ 1)3 + (6et−1)et−1+ (−3)2t+ (3 + 6x2et−1−2x1−81−3t2 (1 +t2)3). Att= 1 , we have ∂2h ∂t2(1) = ( −2 + 1)3 + 6 −6 + (3 + 6 −4−8−2 8) = 4. The next example brings to the surface an ambiguity in the notation∂ ∂xfor partial derivatives. This ambiguity is often a source of great confusion. Consider a scalar valued functiong(x1,x2,t,s) . Ifx1=f1(t) andx2=f2(t) , then h(t,s) =g(f1(t), f2(t),t,s) 378 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S. depends on the two variables tands. In order to see how hchanges with respect to t, we regardsas being held fixed and use the previous example to find ∂h dt=∂g ∂x1∂f1 dt+∂g ∂x2∂f2 ∂t+∂g ∂t. We were careful and realized that the functions g(x1,x2,t,s) , a function with four independent variables, and h(t,s) :=g(f1(t), f2(t), t,s) , a function with only two indepen- dent variables, were different functions. The usual (occasionally confusing) approach is to be less careful and write∂g dt=∂g ∂x1∂f1 dt+∂g ∂x2∂f2 ∂t+∂g ∂t. In the above equation, the term ∂g/∂t on the right is the partial derivative of g(x1,x2,t,s) with respect to twhile thinking of all four variables x1,x2,tandsas being independent. On the other hand, the term ∂g/∂t on the left is the partial derivative of g(f1(t),f2(t),t,s) as a function of two variables. After being spelled out like this, the formula does have a clear meaning - but this is not at all obvious from a glance. One might even be mistakenly tempted to cancel the terms ∂g/∂t from both sides of the equation. It is often awkward to introduce a new name, as h(t,s) , forg(f1(t),f2(t),t,s) . Another unambiguous procedure is available: use the numerical subscript notation for the partial derivatives. Then g,1always refers to the partial derivative of gwith respect to its first variable,g,2with respect to the second variable, etc. Thus, for the above example of g(x1,x2,t,s) wherex1=f1(t) andx2=f2(t) , we have ∂g ∂t=g,1df1 dt+g,2df2 dt+g,3. This clearly distinguishes the two time derivatives g,3and∂g/∂t . The seemingly unnecessary comma in the notation is to take care of the possibility of vector valued functions G(x1,x2,t,s) whose coordinate functions are indicated by sub- scripts. For example, if G=/parenleftbiggg1 g2/parenrightbigg is a map into E2, where the coordinate functions are g1(x1,x2,t,s) andg2(x1,x2,t,s) , then ifx1=f1(t) andx2=f2(t) , we have ∂G ∂t=/parenleftbigg∂g1 ∂t∂g2 ∂t/parenrightbigg =/parenleftbiggg1,1f/prime 1+g1,2f/prime 2+g1,3 g2,1f/prime 1+g2,2f/prime 2+g1,3/parenrightbigg . Hereg1,1=∂g1/∂x 1, etc. The notation f/prime 1fordf1(t)/dtcould also have been replaced by f1,1—but this is unnecessary here since the fjare functions of one variable. In applications, one commonly meets a problem of the following type. Let u(x,y) be a scalar valued function which satisfies the wave equation uxx−uyy= 0 . IfF:E2→E2 is defined by x=f1(ξ,η) =1 2(ξ,+η) y=f2(ξ,η) =1 2(ξ−η) and ifh=u◦F, that is,h(ξ,η) =u(f1(ξ,η),f2(ξ,η)) , what differential equation does h satisfy? First, we compute hξandhη ∂h ∂ξ=∂u ∂x∂f1 ∂ξ+∂u ∂y∂f2 ∂ξ=ux(1 2) +uy(1 2) =1 2(ux+uy) 9.2. THE DERIVATIVE OF COMPOSITE MAPS (“THE CHAIN RULE”). 379 ∂h ∂η=∂u ∂x∂f1 ∂η+∂u ∂y∂f2 ∂η=ux(1 2) +uy(−1 2) =1 2(ux−uy) In a similar way the second derivatives hξξ,hξηandhηηare found, ∂2h ∂ξ2=∂(hξ) ∂x∂f1 ∂ξ+∂(hξ) ∂y∂f2 ∂ξ =1 2∂ ∂x(ux+uy)1 2+1 2∂ ∂y(ux+uy)·1 2=1 4[uxx+ 2uxy+uyy] ∂2h ∂ξ∂η=∂(hξ) ∂η=∂(hξ) ∂x∂f1 ∂η+∂(hξ) ∂y∂f2 ∂η =1 2∂ ∂x(ux+uy)·1 2+1 2∂ ∂y(ux+uy)·−1 2=1 4[uxx−uyy] ∂2h ∂η2=∂(hη) ∂x∂f1 ∂η+∂(hη) ∂y∂f2 ∂η =1 2∂ ∂x(ux−uy)·1 2+1 2∂ ∂y(ux−uy)(−1 2) =1 4[uxx−2uxy+uyy] Sincehξη=1 4[uxx−uyy] , andusatisfies the wave equation, we see that hsatisfies the equation hξη= 0, so, in fact, the equations for hxiξandhηηare superfluous to obtain the desired result. From this, it is easy to give another procedure for solving the wave equation, indepen- dent of Fourier series. Because hξη= 0 , we know that h(ξ,η) =ϕ(ξ) +ψ(η) , where the functionsϕandψare any twice differentiable functions. However, h(ξ,η) =u(ξ+η 2,ξ−η 2) . Since the equations x=ξ+η 2, y =ξ−η 2may be solved for ξandηin terms of xandy, viz.ξ=x+yandη=x−y, we haveh(x+y,x−y) =u(x,y) . Buth(ξ,η) =ϕ(ξ)+ψ(η) . Consequently u(x,y) =ϕ(x+y) +ψ(x−y). This formula is the general solution of the one space dimensional wave equation. It expresses uin terms of two arbitrary functions ϕandψ. These functions ϕandψcan be chosen so that the function u(x,y) , a solution of the wave equation, has any given initial position u(x,0) =f(x) and initial velocity uy(x,0) =g(x) . Let us do this. From the initial conditions we find f(x) =u(x,0) =ϕ(x) +ψ(x) g(x) =uy(x,0) =ϕ/prime(x)−ψ/prime(x). After differentiating the first expression, one can solve for ϕ/primeandψ/prime, ϕ/prime(x) =f/prime(x) +g(x) 2, ψ/prime(x) =f/prime(x)−g(x) 2. Integrate these: ϕ(x) =ϕ(0) +/integraldisplayx 0f/prime(s) +g(s) 2ds=ϕ(0) +f(x)−f(0) 2+1 2/integraldisplayx 0g(s)ds. 380 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S. ψ(x) =ψ(0) +/integraldisplayx 0f/prime(s) +g(s) 2ds=ψ(0) +f(x)−f(0) 2+1 2/integraldisplayx 0g(s)ds. Thus, u(x,y) =ϕ(x+y) +ψ(x−y) =ϕ(0) +f(x+y)−f(0) 2+1 2/integraldisplayx+y 0g(s)ds+ ψ(0) +f(x−y)−f(0) 2−1 2intx−y 0g(s)ds. Becausef(0) =ϕ(0) +ψ(0) , this simplifies to u(x,y) =f(x+y)−f(x−y) 2s+1 2/integraldisplayx+y x−yg(s)ds, the famous d’Alembert formula for the solution of the one space dimensional wave equa- tion in terms of the initial position f(x) and initial velocity g(x) . Unfortunately, simple formulas like this are exceedingly rare. That is why a different, more generally applicable, procedure was used earlier to solve the wave equation. As was seen in Exercise 6, p. 645, the d’Alembert formula is recoverable from the Fourier series. Exercises (1) For the following function gandf, computed dX(g◦F) and evaluate∂ ∂x1(g◦F) at the pointX0= (2,2) . (a)g(y1,y2) =y1y2−y2e2y1, F:yz= 2x1−x1x2, y 2=x2 1+x2 2 (b)g(y1,y2) = 7 +ey1siny2 F:y1= 2x1x2, y 2=x2 1−x2 2 (c)g(y1,y2,y3) =y2 1−y2 2−3y1y3+y2 F:y1= 2x1−x2, y 2= 2x1+x2, y 3=x2 1 (2) Letϕ(x1,x2,t) :=x2x2−te2x1. IfX=F(t) is defined by x1= 1−t2, x 2= 3t+1 , findd dt(ϕ◦F) att= 1 . (3) Letϕ(x,s,t ) :=xs+xt+st. Ifx=f(t) =t3−7 , compute∂ ∂t(ϕ◦f) att= 3 . Also compute∂2 ∂t2(ϕ◦f) att= 3 . (4) Ifu(x,y) =x2−y2, whileF:= (f1,fx) is given by x=f1(r,θ) =rcosθ, y = f2(r,θ) =rsinθfindhrandhθ, whereh:=u◦F. Also compute, hrr, h rθand hθθ. 9.2. THE DERIVATIVE OF COMPOSITE MAPS (“THE CHAIN RULE”). 381 (5) (a) Let u(x,y) be a scalar valued function and F:E2→E2be defined by the polar coordinate transformation f1(r,0) =rcosθ, f 2(rθ) =rsinθ, Takeh:=u◦F. Findhr,hθ,hrr,hrθ, andhθ,θ. [Answer: h4=uxcosθ+ uysinθ, h rr=−uxxrsinθ+uyy(rcosθ−rsinθ)+uyyrcosθ−uxsinθ+uycosθ] (b) Show that uxx+uyy=hrr+1 r2hθθ+1 rhr. (6) The two space dimensional wave equation is utt=uxx+uyy (a) If the space variables x,y are changed to polar coordinates (ex. 5) while the time variable is not changed, the wave equation reads htt=? whereh(r,θ,t ) =u(rcosθ,rsinθ,t). (b) If a given wave form depends only on the distance rfrom the origin and time t, but not on the angle ∂, how does the wave equation for hsimplify? (c) Consider the equation you found in b. Use the method of separation of variables and seek a solution in the form h(r,t) =R(r)T(t) . What are the resulting ordinary differential equations? Compare the equation for R(r) with Bessel’s differential equation. (7) Ifw=f(x,y,s ) , whilex=ϕ(y,s,t ) andy=ψ(s,t) , find expressions for the partial derivative of the composite function g(ϕ(ψ,s,t ),ψs) with respect to sandt. (8) (a) Let u(x,y) =f(x−y) . Show that usatisfies the partial differential equation ux+uy= 0. (b) Letu(x,y) =f(xy) . Show that usatisfies the equation xux−yuy= 0 . (c) Letu(x,y) =f(x y) . Show that usatisfies the equation xux+yuy= 0. (d) Letu(x,y) =f(x2+y2) , souonly depends on the distance from the origin. Show that usatisfies yux−xuy= 0. (9) Letu(x,y) satisfy the equation xux+yuy= 0 . (a) Change the equation to polar coordinates [Answer: if h(r,θ) :=u(rcosθ,rsinθ) , thenrhr= 0 ]. (b) Solve the equation for h(r,θ) and use it to deduce that u(x,y) =f(x y) for some functionf. (cf. Ex. 8c) 382 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S. (10) Assume u(x,y) satisfies the equation uxx−2uxy−3uyy= 0. (a) Choose the constants α,β,γ , andδso that after the change of variables x= αξ+βη, y =γξ+δη, the equation for h(ξ,η) =u(αξ+βη,γξ +δη) ishξη= 0 . (b) Use the result of part (a) to find the general solution of the equation for u. [Answer:u(x,y) =ϕ(3x−y) +ψ(x+y) ]. (11) Iff(x,y) is a known scalar valued function, find both partial derivatives of the functionf(f(x,y),y) . (12) IfW=G(Y) andY=F(X) are defined by G:/braceleftbiggw1=ey1−y2 w2=ey1+y2, F :/braceleftbiggy1=x2 1−3x2−x3 y2=x1+x2 2+ 3x3, findd dX(G◦F) . (13) Letu(x,y) be a solution of the two dimensional Laplace’s equation uxx+uyy= 0 . (a) Ifudepends only on the distance from the origin u(x,y) =h(r) , wherer= x2+y2, what ordinary differential equation does hsatisfy? Compare your answer with that found in Exercise 5. (b) Solve the resulting equation for hand deduce that all the solutions of the two dimensional Laplace equation which depend only on the distance from the origin are of the form u(x,y) =A+Blog(x2+y2), whereAandBare constants. (c) Now do the same thing all over again for a solution u(x1,x2,...,x n) of then dimensional Laplace equation ux1x1+...+uxnxn= 0 , i.e. find the form of all solutions which only depend on r=/radicalbig x2 1+...+x2n,u(x1,...,x n) =h(r) . [Answer:u(x1,...,x n) =A+B (x2 1+...+x2n)n−2 2=A+B rn−2, n≥3 ]. (14) Iff(t) is a differentiable scalar valued function with the property that f(x+y) = f(x) +f(y) for allx,y∈E1, prove that f(x)≡kxwherek=f(1) . (15) (a) Find the general solution of the partial differential equation ux−2uy= 0 . [Hint: Introduce new variables as in Ex. 10] (b) What is the solution if one requires that u(x,0) =x2? [Answer: u(x,y) = (x+1 2y)2]. Chapter 10 Miscellaneous Supplementary Problems 1. (a)Sn, n= 1,2,..., be a given sequence. Find another sequence ansuch that SN=N/summationdisplay n=1an. In other words, given the partial sums Sn, find a series whose partial sums are Sn. To what extent are the anuniquely determined? (b) Apply part (a) to find an infinite series/summationtextanwhosenth partial sum Snis given by (i)Sn=1 n, (ii)Sn=e−n 2. LetS={x∈R:x∈(−1,1)}. Define addition on Sby the formula x⊕y= x+y 1+xy, x,y∈S, where the operations on the right are the usual ones of arithmetic. Show that the elements of Sform a commutative group with the operation ⊕. 3. (a) Ifan→a, prove thata1+a2+···+an n→aalso. (b) Assume that fis continuous on the interval [0 ,∞] and lim x→∞f(x) =A. Define HN=1 N/integraldisplayN 0f(x)dx. Prove that lim x→∞HNexists and find its value. [Hint: InterpretHNas the average height of the function f]. 4. (a) Suppose that allthe zeroes of a polynomial P(x) are real. Does this imply that all the zeroes of its derivative P/prime(x) are also real? (Proof or counterexample). What can you say about higher derivatives P(k)(x) ? (b) Define the nth Laguerre polynomial by Ln(x) =exdn dxn[xne−x]. Show that Lnis a polynomial of degree n. Prove that the zeroes of Ln(x) are all positive real numbers, and that there are exactly nof them. 383 384 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS 5. Iff(x) has a Taylor series: f(x) =∞/summationdisplay n=0anxn(which converges to ffor|x|< ρ so the remainder does go to zero there) prove that f(cxk) , wherecis a constant and ka positive integer, has the Taylor series f(cxk) =∞/summationdisplay n=0ancnxnk which converges to f(cxk) for |x|<(ρ |c|)1/k. You must show that i) the Taylor coefficients for f(cxk) areancn, that ii) the power series for f(cxk) converges for |x|<(ρ |c|)1/k, and that iii) the remainder tends to zero. Apply the result to obtain the Taylor series for cos(2 x2) from that of cos x. 6. Yet another proof of Taylor’s Theorem. Beginning with equation 9 on p. 98, define the function K(s) by K(s) =f(s)−N/summationdisplay n=0f(n)(x0) n!(s−x0)n−A(s−x0)N+1 (N+ 1)!, whereAis picked so that K(ˆx) = 0 . (a) Verify that K(x0) =K/prime(x0) =...K(N)(x0) = 0 . (b) Use Rolle’s Theorem to prove that if a function K(s) satisfies the properties of a), and ifK(ˆx) = 0 , then there is a ξbetween ˆxandx0such thatK(N+1)(ξ) = 0 . (c) Apply parts a) and b) to prove Taylor’s Theorem. 7. Assume/summationtextanconverges. You are to investigate the convergence of/summationtexta2 nand/summationtext/radicalbig |an|under various hypotheses. (a)anarbitrary complex number (b)an≥0 . (c) lim n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglean+ 1 an/vextendsingle/vextendsingle/vextendsingle/vextendsingle<1 (not= 1). 8. The harmonic series 1+1 2+1 3+···has been said to diverge with “infuriating slowness”. Find a number Nsuch that 1 +1 2+1 3+···+1 Nis at least 100. Compare this with Avogadro’s number ∼6×1023. 9. Consider the series/summationtext∞ n=1an, where the an’s are real. (a) Letb1,b2,b3,...andc1,c2,c3,...denote the positive and negative terms respec- tively from a1,a2,.... If/summationtext∞ n=1anconverges conditionally but not absolutely, prove that both series/summationtext∞ n=1bnand/summationtext∞ n=1cndiverge . (b) Letd1,d2,d3,..., denote the terms a1,a2,a3,... rearranged in any way. Prove Riemann’s theorem, which states that if/summationtext∞ n=1anconverges conditionally but not absolutely, then by picking some suitable rearrangement, the series/summationtext∞ n=1dn can be made to converge to any real number, while using other rearrangements, it can be made to diverge to plus or minus infinity. 385 10. IfAandBare subsets of a linear space V, a) show that span {A∩B} ⊂span{A}∩ span{B}. Give an example showing that span {A∩B}may be smaller than span{A} ∩span{B}. b). Show that if A⊂B⊂span{A}, then span {A} ⊃span{B}. 11. LetA={X1,...,X k}be a set of vectors in a linear space V. Denote by cs A (coset ofA) the set csA={X∈V:X=k/summationdisplay j=1ajXj,wherek/summationdisplay j=1aj= 1}. Prove that cs Ais a coset of V, in fact, the smallest coset of Vwhich contains the vectorsX1,...,X k. 12. (a) Consider the set of real numbers of the form a+b√ 2 , whereaandbare rational numbers. Prove that this set is a vector space over the field of rational numbers. What is the dimension of this vector space? (b) Consider the set of numbers of the form a+bi, whereaandbare real numbers andi=√−1 . Prove that this set is a vector space over the field of realnumbers and find its dimension. 13. IfF1andF2are fields with F1⊂F2, we callF2anextension field ofF1– such asR⊂C. As such, we may think of F2as a vector space over the field F1(see exercise 1l). In other words, take F2as an additive group and take the scalars from F1. If this vector space is finite dimensional, the field F2is called a finite extension ofF1, and the dimension nof this vector space is called the degree of the extension and written n= [F2:F1] . (a) Prove that every element ξ∈F2satisfies an equation anξn+an−1ξn−1+···+a0= 0, where theak∈F1andn= [F2:F1] . [Hint: look at the examples of exercise 1l]. (b) IfF1⊂F2⊂F3are fields with [F2:F1] =n<∞and [F3:F2] =m<∞, prove that [ F3:F1]<∞, in fact, prove [F3:F1] = [F3:F2]]F2:F1] =nm. (c) LetF1be the field of rationals, F2the field whose elements have the form a+b√ 3 , whereaandbare rational, and let F3be the field whose elements have the form c+d√ 5 , wherecanddare inF2. Compute [ F2:F1] and find the polynomial of part a) satisfied by (1 −√ 3)∈F)2. Compute [ F3:F2] and [F3:F1] . Find a basis for F3as a vector space whose scalars are elements of F1. [The ideas in this problem are basic to modern algebra, particularly Galois’ theory of equations.] 386 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS 14. LetPj= (αj,βj), j= 1,...,n,α j/negationslash=αkbe anyndistinct points in the plane R2. One often wants to find a polynomial p(x) =a0+a1x+···+aNxNwhich passes through these npoints,p(αj) =βj, j= 1,...,n . Thus,p(x) is an interpolating polynomial . Given any points P1,...,P n, prove that a unique interpolating polyno- mialp(x) degreen−1(=N) can be found. (More about this is in Exercises 17-18 below). 15. LetL1andL2be linear operators mapping V→V. Then they can be both multiplied and added (or subtracted). The bracket product orcommutator [L1,L2]≡L1L2−L2L1 “measures the non-commutativity”. It is important in mathematics and physics. [In quantum mechanics, the observables - like energy, momentum, and position - are represented by self-adjoint operators. Two observables can be measured at the same time if and only if their associated operators commute]. Prove the identities (a) [L1,L1] = 0,[L1,I] = 0 (b) [L1,L2] =−[L2,L1] (c) [aL1,L2] =a[L1,L2] , a scalar (d) [L1+L2,L3] = [L1,L3] + [L2,L3] (e) [L1,L2,L3] = [L1,L2]L3+L2[L1,L3] (f) [L1,[L2,L3]] + [L2,[L3,L1]] + [L3,[L1,L2]] = 0 (Part f is the Jacobi identity . It has been said that everyone should verify it once in her lifetime.) 16. * Consider the normalized Legendre Polynomials, en(x) =/radicalbigg 2 2n+ 11 2nn!dn dxn(x2−1)n, n = 0,1,2,... which are an orthonormal set of polynomials in L2[−1,1], enbeing of degree n. If f∈C[−1,1] , prove that PNf=N/summationdisplay n=0/angbracketleftf, en/angbracketrighten converges to fin the norm of L2[−1,1] . [Hint: Use the form of the Weierstrass Approximation Theorem (p. 255) and the method of Theorem (p. 241)]. 17. * We again take up the interpolation problem begun in Exercise 13 above. Let Pj= (αj,βj), j= 1,2,...,n benpoints in the plane, αi/negationslash=αj. Although we proved there is a unique polynomial p(x) =a0+a1x+···+an−1xn−1of degreen−1 passing through the npoints, the proof was entirely non-constructive. Here we (or you) explicitly construct the polynomial. (a) Show that the polynomial of degree n−1 ˜pj(x) = Πn k=/lscript k/negationslash=j(x−αk) is zero ifx=αk, k/negationslash=j, but ˜pj(αj)/negationslash= 0 . 387 (b) Construct a polynomial pj(x) with the property pj(αk) =δjk. (c) Show that p(x) =n/summationdisplay j=1βjpj(x) is the desired (unique by Ex. 13) interpolating polynomial. (d) LetP1= (1,1), P2= (2,1), P3= (4,−1), P4= (−1,−2) . Find the interpolating polynomial using the above construction. 18. * Iffis some complicated function, it is often useful to use an interpolating polyno- mial instead of the function. Then the polynomial p(x) will pass through the points Pj= (αm,f(αj)), j = 1,...,n , so by Exercise 16, p(x) =n/summationdisplay j=1f(αj)pj(x). ] How much will pdiffer from fin an interval [ a,b] containing the αj? You must estimate the remainder R=f−p. (a) Assume f∈Cn[a,b] . SinceR(x) =f(x)−p(x) vanishes at x=αj, j= 1,...,n , it is reasonable to write R(x) = (x−α1)···(x−αn)·(?) Fixˆxand define the constant Aby f(ˆx)−p(ˆx) =A(ˆx−α1)···(ˆx−αn). By a trick similar to that used in Taylor’s Theorem (cf. P. 104j Ex. 12), prove thatA=f(n)(ξ)/n! whereξis some point in ( a,b) . Thus, f(ˆx) =p(ˆx) +(ˆx−α1)···(ˆx−αn) n!f(n−1)(ξ), ξ∈(a,b). (b) Letf(x) = 2x, andα1=−1, α2= 0, α3= 1, α4= 2. Find the approximating polynomial and find an upper bound for the error in the interval [ −2,2] . 19. Ifxis irrational and a,b,c, and d are rational (with ad−bc/negationslash= 0) , prove thatax+b cx+d is irrational. 20. Prove by induction that 1 + 3 + 5 + ···+ (2n−1) =n2. 21. (a) If x≥0 , use the mean value theorem to prove ex≥1 +x. 388 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS (b) Ifak≥0 , prove that n/summationdisplay k=1ak≤Πn k=1(1 +ak)≤ePn k=1ak, (where Πn k=1bk=b1b2···bn). (c) Ifak≥0 , prove that the infinite product Π∞ k=1(1 +ak) := lim n→∞Πn k=1(1 +ak) converges if and only if the infinite series/summationtext∞ k=1converges. 22. Letan+1=2 1+an, wherea1>1 . Prove that (a) the sequence a2n+1is monotone decreasing and bounded from below. (b) the sequence a2nis monotone increasing and bounded from above. (c) does lim n→∞anexist? 23. Letak, k= 1,...,n + 1 be arbitrary real numbers which satisfy a1+a2 2+···+ an n+an+1 n+1= 0 . Show that P(x) =a1+a2x+···+anxn−1has at least one zero for x∈(0,1) . 24. Suppose f∈C2in some neighborhood of x0. Prove that lim h→0f(x0+h)−2f(x0) +f(x0−h) h2=f/prime/prime(x0). 25. Lets(x) andc(x) be continuously differentiable functions defined for all x, and having the properties s/prime(x) =c(x), c/prime(x) =s(x) s(0) = 0, c (0) = 1. (a) Prove that c2(x)−s2(x) = 1 . (b) Show that c(s) ands(x) are uniquely determined by these properties. 26. Consider/summationtextanand/summationtextbn. (a) If lim n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglebn an/vextendsingle/vextendsingle/vextendsingle/vextendsingle=K, K /negationslash= 0,∞, then the series both converge or diverge together. (b) If/summationtextanconverges and lim n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglebn an/vextendsingle/vextendsingle/vextendsingle/vextendsingle= 0 , then/summationtextbnconverges. (c) If/summationtextanconverges and lim n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglebn an/vextendsingle/vextendsingle/vextendsingle/vextendsingle=∞, then the series/summationtextbnmay converge or diverge (give examples). (d) Apply these to: (i)∞/summationdisplay n=21 n−√n 389 (ii)∞/summationdisplay n=11 n3−2√n (iii)∞/summationdisplay n=1(−1)nsinπ n. (Hint: as x→0,sinx x→1 ). 27. The following (a weak form of Stirling’s formula ) is an improvement of the result on page 64, Ex. 6. nlogn−(n−1)<logn!<(n+ 1) log(n+ 1)−2 log 2 −(n−1), from which one finds nne−n+1<n!<1 4(n+ 1)(n+1)e−n+1. Prove these. 28. (a) Find the Taylor series expansion for f(x) =e−xaboutx= 0 . (b) Show that the series found in (a) converges to e−xfor allxin the interval [−r,r] , wherer>0 is an arbitrary but fixed real number. 29. Consider the sequence SN=/integraldisplayN 2sinπx xdx. Does lim N→∞SNexist? [Hint: observe that SNcan be written as SN=N−1/summationdisplay 2an, where an=/integraldisplayn+1 nsinπx xdx. Sketch a graph ofsinπx x,x≥2 , to deduce - by inspection - the needed properties of thean’s. Please do not attempt to evaluate the integrals for an]. 30. LetA={p∈P9:p(x) =p(−x)}. (a) Prove that Ais a subspace of P9. (b) Compute the dimension of A. 31. LetXandYbe elements in a real linear space. Prove that /bardblX/bardbl=/bardblY/bardblif and only if (X+Y)⊥(X−Y) . 32. In the space R2, introduce the new scalar product <X,Y > =x1y1+ 4x2y2, whereX= (x1,x2) andY= (y1,y2) . 390 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS (a) Verify that this indeed is a scalar product and define the associated norm /bardblX/bardbl. (b) LetX1= (0,1) andX2= (4,−2) . Using thisnorm and scalar product, find an orthonormal set of vectors e1ande2such thate1is in the subspace spanned byX1. 33. LetHbe a scalar product space with XandYinH. Find a scalar αwhich makes /bardblX−αY/bardbla minimum. For this α, how areX−αYandYrelated? [Hint: Draw a picture in E2]. 34. If∞/summationdisplay n=1anconverges, where an≥0 , does the series∞/summationdisplay 1√a n2also converge? Proof or counterexample. 35. Use the Taylor series about x0= 0 to calculate sin.2 making an error less than .005 . Justify your statements. 36. LetA= span {(1,1,1,1),(1,0,1,0)}be a subspace of E4. Find the orthogonal complement, A⊥, ofAby giving a basis for A⊥. 37. Prove that (a) 1 +1 8<∞/summationdisplay k=11 k3<1 +1 2. (b) 1 +1 2∞/summationdisplay k=11 k2<1 +3 4. 38. Letakbe a sequence of positive numbers decreasing to zero, ak→0 , and letSN= a1+a2+···+aN. (a) Prove that SN≥NaN. (b) Use this to estimate the number, N, of terms needed to make N/summationdisplay k=1k−1/4>1000. 39. Prove or give a counterexample: (a) If∞/summationdisplay n=1bnconverges, then∞/summationdisplay n=1b2nmust converge. (b) If∞/summationdisplay n=1|bn|converges, then∞/summationdisplay n=1|b2n|must converge. 40. LetX1andX2be elements of a scalar product space. (a) IfX1⊥X2, prove that /bardblX1−aX2/bardbl ≤ /bardblX1/bardblfor any real number a. 391 (b) Prove the converse, that is, if /bardblX1−aX2/bardbl ≤ /bardblX1/bardblfor every real number a, thenX1⊥X2. [Hint: After your first approach has failed, try looking at the problem geometrically. How would you pick ato minimize the left side of the inequality?]. 41. LetSn=a1+a2+···+an, wherean→0 asn→ ∞ . Prove that Snconverges if and only if S2n=a1+a2+···+a2n−1+a2nconverges (one could also use S3netc.). 42. Show that the error in approximating the series∞/summationdisplay n=11 nnby the first Nterms is less thanN−N−1. 43. A sample “multiplication” for points X= (x1,x2,x3) andY= (y1,y2,y3) inR3is to define X⊙Y≡(x1y2,x2y2,x3y3). Define a multiplicative identity by yourself. Using these definitions for the multiplica- tive structure and the usual rules for the additive structure, show that the resulting algebraic object is not a field. 44. (a) Assume an≥0 andbn≥0 . Prove that ∠(an+bn) converges if and only if the series ∠anand∠bnbothconverge. (b) What if you allow the bn’s to be negative? 45. (a) Show that the vectors e1= (1√ 2,1√ 2), e2= (1√ 2,−1√ 2) form an orthonormal basis forE2. (b) Write the vector X= (7,−3) in the form X=a1e1+a2e2, using the scalar product to find a1anda2(don’t solve linear equations). 46. Consider the linear space P2as a subspace of L2[0,1] . (a) Ifp(x) = 1−x2, compute /bardblp/bardbl. (b) Find the orthonormal basis for A⊥, whereA= span {2 +x}. (c) Find the polynomial ϕ∈P2such that /angbracketleftp, ϕ/angbracketright=p(1) for all p∈P2, that is, the same ϕshould work for all p’s. 47. Give formal proofs for the following (trivial) properties of a norm on a linear space. Only the axioms may be used. (a)/bardbl −X/bardbl=/bardblX/bardbl (b)/bardblX−Y/bardbl=/bardblY−X/bardbl (c)/bardblX+Y/bardbl ≥ /bardblX/bardbl − /bardblY/bardbl (d)/bardblX1+X2+···+Xn/bardbl ≤ /bardblX1/bardbl+/bardblX2/bardbl+···+/bardblXn/bardbl(I suggest induction here). 392 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS 48. Consider R2with the norms /bardbl /bardbl 1,/bardbl /bardbl 2, and /bardbl /bardbl ∞. (a) Draw a sketch of R2indicating the unit ball for each of these three norms. (The ball may not turn out to be “round”). (b) Which of these three linear spaces have the following property: “given any sub- spaceMand a point X0not inM, then there is a unique point onMwhich is closest to M.” 49. Are the following scalar products the set of functions continuous on [ a,b] ? Proof or counterexample. (a) [f,g] = (/integraldisplayb af(x)dx)(/integraldisplayb ag(x)dx) (b) [f,g] = (/integraldisplayb a|f(x)|dx)(/integraldisplayb a|g(x)|dx) 50. (a) Let dim V=nand{X1,...,X n} ∈V. Prove that {X1,...,X n}are linearly independent if and only if they span V(so in either case, they form a basis for V). (b) Let {e1,...,e n}be an orthonormal set of vectors for an inner product space H. Prove this set of vectors is a complete orthonormal set for Hif and only if n= dimH. (c) Prove that dim V= largest possible number of linearly independent vectors in V. 51. (a) Let XandYbe any two elements in an inner product space. Prove that the parallelogram law holds /bardblX+Y/bardbl2+/bardblX−Y/bardbl2= 2/bardblX/bardbl2+ 2/bardblY/bardbl2 (cf. page 192, Ex. 9). (b) Consider the set of continuous functions on [0 ,1] with the uniform norm, /bardblf/bardbl∞= max 0≤x≤1|f(x)|. Show that this norm cannot arise from an inner product, i.e. there is no inner product such that for all f,/bardblf/bardbl∞=/radicalbig /angbracketleftf, f/angbracketright. [Hint: If there were, the relationship of part a would hold between the norms of various ele- ments. Show that relationship does not, in fact, hold for the function f(x) = 1 andg(x) =x]. 52. (a) Let Hbe a finite dimensional inner product space and /lscript(X) a linear functional defined for all X∈H. Show that there is a fixed vector X0∈Hsuch that /lscript(X) =/angbracketleftX, X 0/angbracketright for allX∈H. This shows that every linear functional can be represented simply as the result of taking the inner product with some vector X0. [Hint: First pick a basis {e1,...,e n}forHand letcj=/lscript(en) . Now use the fact that the ej’s are a basis and that /lscriptis linear]. 393 (b) Consider the linear space P2with theL2[0,1] inner product. This gives an inner product space H. (i) Show that /lscript(p) =p(1 3) is a linear functional. (ii) Find a polynomial p0such that/lscript(p) =/angbracketleftp, p 0/angbracketrightfor allp∈H. 53. Consider the set Sof pairs of real numbers X= (x1,x2) . Define X+Y= (x1+y1,x2+y2), aX = (ax1,x2). IsS, with this definition of vector addition and multiplication by scalars, a vector space? 54. By inspection, place suitable restrictions on the contents a,b,c, ···in order to make the following operator linear: Tu=a[d3u dx3]2+bx2d2u dx2+cudu dx+eu+fsinu+g. 55. Consider the operator D=d dxon the linear space Pnof all polynomials of degree less than or equal to n. Find R(D) and N(D) as well as dim R(D) and dim N(D) . 56. Let A= 1−2 2 0 3 1 , B =/parenleftbigg−1 0 −2 2 1 0/parenrightbigg ,andC= 3 0 2 1 4 −1 0−2 0 . Compute all of the following products which make sense: AB, BA, AC, CA, BC, CB, A2, B2, C2, ABC,CAB. 57. Consider the mapping A:R4→R3which is defined by the matrix A= 1−1 1 1 2 1 1 4 0−3 1 −2  (a) Find bases for N(A) and R(A) . (b) Compute dim N(A) and dim R(A) . 58. LetAbe a square matrix. Consider the system of linear algebraic equations AX=Y0, whereY0is a fixed vector. Assume these equations have two distinct solutionsX1 andX2, AX 1=Y0, AX 2=Y0, X 1/negationslash=X2. (a) Find a third solution X3. 394 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS (b) Does there exist a vector Y1such that the equations AX=Y1 have nosolutions? Why? (c) detA=? 59. LetQbe a parallelepiped in Enwhose vertices Xkare at points with integer coor- dinates, Xk= (a1k,a2k,···ank), a ikintegers. Prove that the volume of Qis an integer. 60. LetAandBbe self-adjoint matrices. Prove that their product ABis self-adjoint if and only if AB=BA. 61. Solve the following initial value problems. (a)u/prime/prime+ 8u/prime+ 16u= 0, u (0) =1 2, u/prime(0) = 0 (b)u/prime/prime+ 10u/prime+ 16u= 0, u (0) = 1, u/prime(0) = 2 (c)u/prime/prime+ 64u= 0, u (0) =1 4, u/prime(0) = 1 (d)u/prime/prime+ 4u/prime+ 5u= 0, u (0) = 2, u/prime(0) = −1 (e) 2u/prime/prime+ 6u/prime+ 5u= 0, u (0) = 0, u/prime(0) = −2 (f) 4u/prime/prime−4u/prime+u= 0, u (1) = −1, u/prime(1) = 0 (g)u/prime/prime+ 8u/prime+ 16u= 2, u (0) =1 2, u/prime(0) = 0 (h)u/prime/prime+ 8u/prime+ 16u=t, u (0) =1 2, u/prime(0) = 0 (i)u/prime/prime+ 8u/prime+ 16u=t−2, u (0) = 0, u/prime(0) = 0 (j)u/prime/prime+ 8u/prime+ 16u=t−2, u (0) =1 2, u/prime(0) = 0 (k)u/prime/prime+ 10u/prime+ 16u=t, u (0) = 1, u/prime(0) = 2 (l)u/prime/prime+ 64u= 64, u (0) =1 4, u/prime(0) = 2 (m)u/prime/prime+ 64u=t−64, u (0) =3 4, u(0) = 0 (n) 2u/prime/prime+ 6u/prime+ 5u=t2, u (0) = 0, u/prime(0) = −2 62. (The complex numbers as matrices). (a) Show that the set of matrices C={/parenleftbigga−b b a/parenrightbigg :aandbare real numbers } is a field. (b) Find a map ϕ:C→complex numbers such that ϕis bijective and such that for allA,B∈C (i)ϕ(A+B) =ϕ(A) +ϕ(B) (ii)ϕ(AB) =ϕ(A)ϕ(B). 395 63. (Quaternions as matrices). A definition: A division ring is an algebraic object which satisfies all of the field axioms except commutativity of multiplication. (a) Show that the set of matrices Q={/parenleftbiggz−¯w w ¯z/parenrightbigg :z,w are complex numbers } form a division ring with the usual definitions of additions and multiplication for matrices. (b) If we write z=x+iy, w =u+ivwherei=√−1 andx,y,u , andvare real numbers, then Qcan be considered as a vector space over the reals with basis 1=/parenleftbigg1 0 0 1/parenrightbigg i=/parenleftbiggi0 0−i/parenrightbigg j=/parenleftbigg0−1 1 0/parenrightbigg k=/parenleftbigg0i i0/parenrightbigg . Compute i2,j2,k2,ij,jk,ki,ji,kj, and ik. (The set Qis called the quater- nions ). 64. Let A= 2−3 1 0 0 2 −3 1 0 0 2 −3 0 0 0 2 . (a) Find det A. (b) FindA−1. (c) SolveAX=Y, whereY= 2 8 8 −16 . (d) LetL:P3→P3be the linear operator defined by Lp=p/prime/prime−3p/prime+ 2p,(u/prime=du dx). Find the matrix eLeforLwith respect to the following basis for P3 e1(x) = 1, e 2(x) =x, e 3(x) =x2 2, e 4(x) =x3 3!. (e) Use the above results to find a solution of Lu= 2 + 8x+ 4x2−8 3x3. [Hint: Express the right side in the basis of part d.]. 65. LetHbe an inner product space, and suppose that Ais a symmetric operator, A∗=A, with the additional property that A2=A. Show that there exist two subspacesV1andV2ofHwith all of the following properties 396 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS (i)V1⊥V2 (ii) IfX∈V1, thenAX=X (iii) IfY∈V2, thenAY= 0 (iv) IfZ∈H, thenZcan be written uniquely as Z=X+YwhereX∈V1and Y∈V2. 66. (a) Find the inverse of the matrix A= 2 1 0 −1 0 1 0−1−1 . (b) Use the result of a) to solve AX=bforXwhereb= (7,−3,2) . 67. LetAandBbe 2×2 positive definite matrices with det A= detB. Prove that det(A−B)<0 . 68. LetL:V1→V2be a linear operator with LX 1=Y1andLX 2=Y2. Give a proof or counterexample to each of the following assertions: (a) IfX1andX2are linearly independent, then Y1andY2must be linearly independent. (b) IfY1andY2are linearly independent, then X1andX2must be linearly independent. 69. Letp0,p1,p2,,... be an orthogonal set of polynomials in [ a,b] wherepnhas degree n. (a) Prove that pnis orthogonal to 1 ,x,x2,...,xn−1. (b) Prove that pnis orthogonal to any polynomial qof degree less than n. (c) Prove that pnhas exactly ndistinct real zeros in ( a,b) . [Hint: Let α1,...,α k be the places in ( a,b) wherepn(x) changes sign, so p(x) =r(x)(x−α1)(x− α2)...(x−αk) wherer(x) is a polynomial of degree n−kwhich does not change sign for xin (a,b) , sayr(x)≥0 . Show that /integraldisplayb ap(x)(x−α1)···(x−αk)dx> 0. Ifk<n , show that this contradicts the result of part b).]. 70. Consider the system of inhomogeneous equations a11x1+···+a1nxn=b, ... ak1x1+···+aknxn=bn. 397 LetA= ((aij)) and letAbdenote the augmented matrix Ab= a11···a1nb1 ... ak1···aknbn  formed by adding the bj’s as an extra column to A. Prove that the given system of equations has a solution if and only if dim R(A) = dim R(Ab) . 71. LetAbe ann×nmatrix. (a) Show that you can not solve the equation A2=−I ifnis odd. (b) Find a 2 ×2 matrixAsuch thatA2=−I. (c) Ifnis even, find an n×nmatrixAsuch thatA2=−I. 72. LetAbe ann×nmatrix such that A2=I. Prove that dim R(A+I)+dim R(A−I) = n. 73. Letf(x,y) = (y−2x2)(y−x2) . Show that the origin is a critical point. Then show that if you approach the origin along a straight line, the origin appears to be a minimum. On the other hand, show that if curved paths are also used, then the origin is a saddle point of f. [The point of this exercise is to illustrate the fact that the nature of a critical point cannot be determined by merely approaching it along straight lines]. 74. (a) Let Abe a diagonal matrix, no two of whose diagonal elements are the same. IfBis another matrix and AB=BA, prove that Bis also diagonal. (b) LetAbe a diagonal matrix, Ba matrix with at least one zero-free column and with the further property that AB=BA. Prove that all of the diagonal elements of Aare equal. 75. (a) If/summationdisplay anconverges, where an≥0 , prove that/summationdisplay√an npconverges if p >1 2. [Hint: Schwarz]. (b) Find an example showing that the series may diverge if p=1 2. 76. Let [X,Y] be an inner product on R3with basis vectors e1,e2,e3, not necessarily orthonormal. Let aij= [ei,ej] . Prove that the quadratic form Q(X) =3/summationdisplay i=13/summationdisplay j=1aijxixj is positive definite. 77. IfAis self-adjoint and AX=λ1X, AY =λ2Ywithλ1/negationslash=λ2, prove that X⊥Y. 398 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS 78. LetSbe a positive definite matrix. Prove that det S > 0 . [Hint: Consider the matrixA(t)≡tS+ (1−t)I, where 0 ≤t≤1 . Show that A(t) is positive definite, so detA(t)/negationslash= 0 . Then use the fact that A(0) =IandA(1) =Sto obtain the conclusion]. 79. Consider the linear space of infinite sequences X= (x1,x2,x3,···) with the usual addition. Define the linear operator S(the right shift operator) by SX= (0,x1,x2,x3,···) (a) DoesShave a left inverse? If so, what is it? (b) DoesShave a right inverse? If so, what is it? 80. Find a right inverse for the matrix A=/parenleftbigg1 0 1 0 1 0/parenrightbigg . CanAhave a left inverse? Why? 81. Which of the following statements are true for allsquare matrices A? Proof or counterexample. (a) IfA2=I, then detA=I. (b) IfA2=A, then detA= 1. (c) IfA2= 0 , then det A= 0 (d) IfA2=I−A, then detA2= 1−detA. 82. LetLbe a linear operator on an inner product space Hwith inner product <,> . Define [X,Y] =/angbracketleftLX, LY /angbracketright. Under what further condition(s) on Lis [X,Y] an inner product too? 83. LetL:H→Hbe an invertible transformation on the inner product space H. If L“preserves orthogonality” in the sense that X⊥YimpliesLX⊥LY, prove that there is a constant αsuch thatR≡αLis an orthogonal transformation. 84. LetHbe an inner product space. If the vectors X1andX2are at opposite ends of a diameter of the sphere of radius rabout the origin, and if Yis any other point on that sphere, prove that Y−X1is perpendicular to Y−X2, proving that an angle inscribed in a hemisphere is a right angle. 85. IfLis skew-adjoint, L∗=−L, prove that /angbracketleftX, LX /angbracketright= 0 for all X. 399 86. LetDnbe an×nmatrix with xon the main diagonal and 1/primeson both the sub- and super-diagonals, so D2=/parenleftbiggx1 1x/parenrightbigg , D 3= x1 0 1x1 0 1x , D 4= x1 0 0 1x1 0 0 1x1 0 0 1x , D 5=···. Ifx= 2 cosθ, prove that det Dn=sin(n+1)θ sinθ. 87. LetAandBbe square matrices of the same size. If I−AB is invertible, prove thatI−BAis also invertible by exhibiting a formula for its inverse. 88. Assume/summationtextanconverges, where an≥0 . Does the series /summationdisplay√anan+1 also converge? Proof or counterexample. 89. LetAbe a square matrix. (a) Prove that AA∗is self-adjoint. (b) IsAA∗always equal to A∗A? Proof or counterexample. 90. Show that C[0,1] is a direct sum of the space V1spanned by e1(x) =xande2(x) = x4, and the subspace V2of all functions ϕ(x) such that 0 =/integraldisplay1 0xϕ(x)dx, 0 =/integraldisplay1 0x4ϕ(x)dx. [Hint: Show that if f∈[0,1] , there are unique constants aandbsuch thatg(x)≡ f(x)−[ax+bx4] belongs to V2]. 91. LetV1be the linear space of all complex-valued analytic functions in the open unit disc, that is, V1consists of all complex-valued functions fof the complex variable zwhich have convergent power series expansions f(z) =∞/summationdisplay 0anzn in the open disc, |z|<1 . LetV2be the linear space of all sequences of complex numbers ( a0,a1,a2,···) with the natural definition of addition and multiplication by constants. DefineL:V1→V2by the rule Lf= (a0,a1,a2,···), where theaj’s are the Taylor series coefficients of f. Answer the following questions with a proof or counterexample. 400 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS (a) IsLinjective? (b) IsLsurjective? (c) Is/lscript2contained in R(L) ? (Note:/lscript2is the subspace of V2such that ∞/summationdisplay k=0|ak|2<∞). 92. Do the following series converge or diverge? (a)∞/summationdisplay n=1/radicalbig 1 + 1/n, (b)∞/summationdisplay n=1(/radicalbig 1 + 1/n2−1). 93. Consider the set of four operators {T1,T2,T3,T4}defined as follows on the set of square invertible matrices. T1A=A, T 2A=A−1 T3A=A∗, T 4A= (A−1)∗. Show that this set of four operators forms a commutative group with the group oper- ation being ordinary operator multiplication. 94. LetSn=a1−a2+a3−a4+a5− ··· . If 0<akand theak’s are increasing, prove that|SN| ≤aN. 95. The Monge-Ampere equation is uxxuyy−u2 xy= 0 . Show that it is satisfied by any u(x,y)∈C2of the form u(x,y) =ϕ(ax+by) , whereaandbare constants. 96. (a) Consider the differential operator Lu=u/prime/prime−4u (i) Find a basis for the nullspace of L. (ii) Find a particular solution of Lu=e2x+1. (iii) Find the general solution of Lu=e2x+1. (b) Consider the differential operator Lu=u/prime/prime+ 4u Repeat part (a), only here use Lu=f, wheref(x) = sec 2x. 97. Find the general solution for each of the following (a) 2u/prime/prime+ 5u/prime−3u= 0 (b)u/prime/prime−6u/prime+ 9u= 0 (c)u/prime/prime−4u/prime+ 5u= 0 401 98. Find the first four non-zero terms in the series solution of 4x2u/prime/prime−4xu/prime+ (3−4x2)u= 0 corresponding to the largest root of the indicial equation. Where does the series converge? 99. Find the complete solution of each of the following equations valid near x= 0 by using power series. (a)x2u/prime/prime+xu/prime−(x2−1 4)u= 0 (b)u/prime/prime+xu/prime−u= 0 (only first five non-zero terms) [Answers: (a)u(x) =Ax−1/2∞/summationdisplay k=0x2k (2k)!+Bx1/2∞/summationdisplay k=0x2k (2k+ 1)!, (b)u(x) =Ax+B(1 +x2 2!−x4 4!+3x6 6!−15x8 8!+···) ]. 100. Consider the matrix A= −1−4−12 0 1 3 6 0 0 0 −1 0 0−4−12 1 . (a) Compute det A. (b) Compute A−1. (c) SolveAX=bwhereb= (1,2,3,−1) . 101. True or false. Justify your response if you believe the statement is false (a counterex- ample is adequate). (a) The set A={X∈R3:x1= 2}is a linear subspace ofR3. (b) The vectors X1= (2,4) andX2= (−2,4)span R2. (c) The vectors X1= (1,2,3), X 2= (−7,3,2), X 3= (2,−1,1) , andX4= (π,e,5) are linearly independent . (d) The set A={u∈C[0,1]:u(x) =a1x+a2ex}is an infinite dimensional subspace of C[0,1] . (e) The functions f1(x) =xandf2(x) =exarelinearly dependent functions in C[0,1] . (f) If {e1,e2,...,e n}are an orthonormal set of vectors in E8, thenn≤7 . (g) The vector Y= (1,2,3) is orthogonal to the subspace of E3spanned by e1= (0,3,−2) ande2= (−1,−1,1) . 402 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS (h) The elements of the set A={u∈C2[0,10]:u/prime/prime+xu/prime−3u= 6x} can be represented as u(x) = ˜u(x) +x3, where ˜u∈S={u∈C2[0,10]:u/prime/prime+xu−3x= 0}. (i) The set of vectors e1= (1 3,0,2 3,−2 3), e 2= (0,0,1√ 2,1√ 2),ande3= (8 9,3 9,−2 9,2 9) constitute a complete orthonormal basis for E4. (j) In the vector space of bounded functions f(x), x∈[0,1] , the functions f1(x) = 1, f 2(x) =/braceleftbigg1,0≤x≤1 2, 0,1 2<x≤1f3(x) =/braceleftbigg0,0≤x≤1 2 1,1 2<x≤1 are linearly independent . (k) The function f(x) =|x|can be represented by a convergent Taylor series about the pointx0= 0 . (l) The function f(x) =x2−x73can be represented by a convergent Taylor series about the point x0=−1 . (m) The function f(x) =|x|can be represented by a convergent Taylor series about the pointx0=−1 . (n) The plane of all points ( x1,x2,x3,x4)∈E4such that 2x1−4x2+ 6x3−5x4= 7 is perpendicular to the vector (2 ,−4,6,−5) . (o) Ife1= (3 5,4 5) ande2= (4 5,−3 5) , thenX= (−1,2) can be written as X= 2e1−e2. (p) The set of all integers (positive, negative, and zero) is a field. (q) Consider the infinite series ∞/summationdisplay k=0ak. If lim k→0|ak|= 0 , then the series must converge . (r) Let {an}be a sequence of rational numbers. If this sequence converges to a, then the limiting value, a, must be a rational number too. (s) The equation x6+ 3 = 0 , where xis an element of an ordered field, has no solutions . (t) It is possible to write√ iin the form a+ib, whereaandbare real numbers. (Herei=√−1 , of course). (u) Letanbe a sequence of complex numbers. If the sequence of absolute values, |an|, converges, then the sequence anmust converge . (v) If ∞/summationdisplay k=0akzk converges at the point z= 3 , then it must converge atz= 1 +i. 403 (w) The linear subspace A={p∈P7:p(x) =a1x+a2x5}is afive dimensional subspace of P7. (x) The linear subspace A={u∈C[−1,1]:u(x) =a1x+a2x5}is an infinite dimensional subspace of C[−1,1] . (y) There is a number αsuch that the vectors X= (1,1,1) andY= (1,α,α2) form a basis forR3. (z) The operator T:C2→C1defined for u∈C2byTu=u/prime−7uis alinear operator. 102. (a) The operator T:C[0,1]→Rdefined for u∈C[0,1] by Tu=/integraldisplay1 0|u(x)|dx is alinear operator. (b) The sequence (1 + i)nconverges to√ 2 . (c) The series ∞/summationdisplay k=1k+ 1 2k+ 1=2 3+3 5+4 7+5 9+··· converges . (d) Iftis real, then/vextendsingle/vextendsingleeit/vextendsingle/vextendsingle= 1 . (e) LetV1andV2be linear spaces and let the operator TmapV1intoV2. If T0 = 0 , then Tis alinear operator. (f) The operator T:C∞[−7,13]→C∞[−7,13] defined by Tu=udu dx islinear . (g) The operator T:C[0,13]→C[0,13] defined by (Tu)(x) =/integraldisplayx 0u(t) sintdt, x ∈[0,13] islinear (h) In the scalar product space L2[0,1] , the functions fandgwhose graphs are a figure goes here areorthogonal . (i) LetLbe a linear operator. IF LX 1=YandLX 2=Y, whereX1/negationslash=X2, then the solution of the homogeneous equation LX= 0 is not unique . (j) LetLbe a linear operator. If X1andX2are solutions of LX= 0 , then 3X1−7X2isalsoa solution of LX= 0 . 404 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS (k) Lete1= (1,1) ande2= (0,1) , and let the linear operator Lwhich maps R2 intoR3satisfy Le1= (1,2,3), Le 2= (1,−2,−1). ThenL(2,3) = (1,1,1). (l) In the space L2[0,1] , iffisorthogonal to the function x2, then either f≡0 orelsefmust be positive somewhere in [0,1] . (m) IfF/prime(X) = (2,3,4) for allX∈E3, thenFis an affine mapping. (n) Iff:E3→E1is such that f: (1,0,0)→1 andf: (0,4,0)→2 , there is a pointZ∈E3such that /bardblf/prime(Z)/bardbl ≥1 5. (o) LetAandBbe square matrices with det A= 7 and det B= 3 . Then detAB= 10. det(A+B) = 10. (p) IfA:R3→R3is given by A=/parenleftbigg2 3 1 1 9 2/parenrightbigg , then dim N(A) = 2 . (q) The function f(x,y,z ) = 9 + 3x+ 4y−7zdoes not take on its maximum value. (r) If the function u(x) has two derivatives in some neighborhood of x= 0 , and satisfies the differential equation 9x2u/prime/prime−28u= 0, thenu(0) = 0 . (s) There are constantsa,bandcsuch that the function u(x) =ex+ 2e2x−e−x is a solution of au/prime/prime+bu/prime+cu= 0. (t) The vector ( xy,x)isthe derivative of some real-valued function f(x,y) . (u) The vector ( y,x) isnotthe derivative of some real-valued function f(x,y) . (v) Given any q×pmatrixA= ((aij(X))) , where X= (x1,···,xp) and where the elements aij(X) are sufficiently differentiable functions, then there is a map F:Rp→Rqsuch thatF/prime(X) =A. (w) IfAis a square matrix and A2=A, thenA=I. (x) IfAis a square matrix and A2= 0 , thenA= 0 . (y) IfAis a square matrix and det A/negationslash= 0 , thenA2=Aif and only if A=I. (z) IfX,Y , andZare three linearly independent vectors, then X+Y, Y +Z, andX+Zare also linearly independent . 103. Define L:P2→P2as follows: if p∈P2 Lp= (x+ 1)dp dx (a) Find the matrix eLerepresenting the operator Lwith respect to the bases e1= 1, e2=x1, e3=x2forP2. 405 (b) IsLan invertible operator? Why? (c) Find dim R(L) and dim N(L) . 104. Let A=/parenleftBigg 1 2−√ 3 2 √ 3 21 2/parenrightBigg , B =/parenleftbigg5√ 3√ 3 3/parenrightbigg . (a) Compute AA∗,ABA∗, and (ABA∗)100. (b) How could you use the result of part (a) to compute B100? 105. Consider the following system of three equations as a linear map L:R2→R3 x1+x2=y1 4x1+x2=y2 x1−2x2=y3 (a) Find a basis for N(L∗) . (b) Use the result of part a) to determine the value(s) of αsuch thatY= (1,2,α) is inR(L) . 106. Find the unique solution to each of the following initial value problems. (a)u/prime/prime+u/prime−2u= 0, u (0) = 3, u/prime(0) = 0 (b)u/prime/prime+ 4u/prime+ 4u= 0, u (0) = 1 u/prime(0) = −1 (c)u/prime/prime−2u/prime+ 5u= 0, u (0) = 2, u/prime(0) = 2 107. Consider the special second order inhomogeneous constant coefficient O.D.E. Lu=f, where Lu≡u/prime/prime−4u, and where fis assumed to be a suitably differentiable function which is periodic with period 2π, f(x+ 2π) =f(x) . (a) Expand fin its Fourier series and seek a candidate, u, for a solution of Lu=f as a Fourier series, showing how the Fourier coefficients of uare determined by the Fourier coefficients of f. (b) Apply the above procedure to the trivial example where f(x) = sin 3x−4 cos 17x+ 3 sin 36x. 108. (a) Find the directional derivative of the function f(x,y) = 2−x+xy at the point (0 ,6) in the direction (3 ,−4) by using the definition of the direc- tional derivative as a limit. Check your answer by using the short method. 406 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS (b) Repeat part (a) for f(x,y) = 1−3y+xy. 109. Find and classify the critical points of the following functions. (a)f(x,y) =x3+y2−3x−2y+ 2 (b)f(x,y) =x2−4x+y2−2y+ 6 (c)f(x,y) = (x2+y2)2−8y2 (d)f(x,y) = (x2−y2)2−8y2 (e)f(x,y) = (x2−y2)2 (f)f(x,y) =x2−2xy+1 3y3−3y 110. Consider the function x3+y2−3x−2y+ 2 . At the point (2 ,1) find the direction in which the directional derivative is greatest. Find the direction where it is least. 111. Letf:E2→Ebe a suitably differentiable function and let X(t) be the equation of a smooth curve CinE2on whichfis identically constant, say, f(X(t))≡4 . Show that on this curve, f/primeis perpendicular to the velocity vector X/prime(t) . [Hint: Do something to ϕ(t) =f(X(t)) . The proof takes but one line.]. 112. Consider the following statements concerning a function f:En→E. (A)fis continuous. (B)fhas a total derivative everywhere. (C)fhas first order partial derivatives everywhere. (D)fhas a total derivative everywhere which is continuous everywhere. (E)fhas first order partial derivatives everywhere and they are continuous functions everywhere. (F)fis an affine function. (G)f/prime≡0 . (a) Which of these statements always imply which others. A sample (possibly incor- rect) answer might look like (A)⇒B,F,··· (B)⇒A,··· (b) Find examples illustrating each case where a given statement does not imply another (the Exercises, pp. 588-95, contain the required examples). 113. Solve the following ordinary differential equations subject to the given auxiliary con- ditions (a)u/prime/prime−u/prime−6u= 0, u (0) = 0, u/prime(0) = 5 (b)xu/prime+u=ex−1, u (1) = 2 (c)u/prime/prime−6u/prime+ 10u= 0 , general solution. 407 114. (a) If u(x,y,t ) =xexy+t2, whilex= 1−t3andy= logt2, then letw(t) = u(x(t), y(t), t) . Finddw dtatt= 1 . (b) IfF:E3→E2andG:E2→E2are defined by F(X) =/parenleftbigg2x1−x2 2+x2x3+ 1 x2 1−x2 3+x2/parenrightbigg , G (Y) =/parenleftbiggy1+y2siny1 −3y1y2+y2 2/parenrightbigg , (i) Why doesn’t F◦Gmake sense? (ii) Compute [ G◦F]/primeat the point X0= (0,1,0) . 115. LetF=E2→E2andG:E3→E2be defined by F(w,z) =/parenleftBigg ew+z2 ez+w2/parenrightBigg G(r,s,t ) =/parenleftbiggr+s2+t3 s+t2+r3/parenrightbigg . (a) FindF/primeandG/prime. (b) Which of F◦GorG◦Fmakes sense? (c) IfG◦Fmakes sense, compute ( G◦F)/primeat (−1,−1) . (d) IfF◦Gmakes sense, compute ( F◦G)/primeat (−1,0,0) . 116. LetF:X→YandG:Y→Zbe defined by F:/braceleftbiggy1=x2−ex1+2x2 y2=x1x2, G :/braceleftbiggw1=y2+y2siny1 w2= (y1+y2)2 (a) Compute F/primeatX0= (−2,1) andG/primeatY0=F(X0) . (b) LetH=G◦F. Compute H/primeatX0= (−2,1) . 117. Consider the map F:E2→E3defined by F:  f1(x,y) =y+ex−y f2(x,y) = sin(x−2y+ 1) f3(x,y) =x−3x2+y2 (a) Find the tangent map at the point X0= (1,1) . (b) Use the result of part (a) to evaluate approximately FatX1= (1.1,.9) . 118. Consider the system of O.D.E.’s u/prime=αu v/prime=αu−βv, whereαandβare constants. If u(0) =Aandv(0) =B, (a) Findu(t) . (b) Findv(t) (remember to consider the case α=βseparately). 408 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS 119. (a) Consider the homogeneous equation u/prime/prime+a(t)u= 0, wherea(t) is continuous and periodic with period P, soa(t+P) =a(t) . (i) Ifa(t)≡1 , show that there is no non-trivial periodic solution by merely solving the equation. (ii) Ifa(t) = cost, show (again by solving the equation) that there is a periodic solutionu(t) with period 2 π. (iii) In general, if u(t) is a solution, not necessarily periodic, show that v(t)≡ u(t+P) is also a solution. (iv) Show that the homogeneous equation has a non-trivial periodic solution of periodPif and only if/integraldisplayP 0a(t)dt= 0 (b) Consider the inhomogeneous equation u+a(t)u=f(t), where both a(t) andf(t) are continuous and periodic with period P. (i) If/integraldisplayP 0a(t)dt=K/negationslash= 0 , show that the inhomogeneous equation has one and only one periodic solution with period P. (ii) If/integraldisplayP 0a(t)dt= 0 , find a necessary condition on fthat the inhomogeneous equation have a periodic solution with period P. 120. Letf:En→Ebe a differentiable function and denote the directional derivative in the direction of the unit vector ebyDef. Prove that D−ef=−Def. 121. Letf:En→Ebe of the form f(a1x1+...+anxn) . Writeα= (a1,···,an) and β= (b1,···,bn) . Ifβis perpendicular to α, prove that β⊥f/prime. 122. LetRdenote the rectangle 0 ≤x1<2π,0≤x2<2π, and define the map f:R→E1by f(x1,x2) = (3 + 2 cos x2) sinx1 Find and classify the critical points of f. (This function is the height function of a torus with major radius 3 and minor radius 2). 123. Consider the constant coefficient differential operator Lu≡au/prime/prime+bu/prime+cu, (a,b,c real, a/negationslash= 0.) Letλ1andλ2denote the roots of the characteristic polynomial p(λ) =aλ2+bλ+c. (a) Ifλ1/negationslash=λ2, find a formula for a particular solution of Lu=f. [Answer:up(x) =1 λ1−λ2/integraldisplayx [eλ1(x−t)−e−λ2(x−t)]f(t)dt. 409 (b) Ifλ1is complex, say, λ1=α+iβ, thenλ2=¯λ1=α−iβ. Show that in this case, the above formula simplifies to up(x) =1 β/integraldisplayx eα(x−t)sinβ(x−t)f(t)dt. (c) Ifλ1=λ2, find a formula for a particular solution of Lu=f. [Answer:up(x) =/integraldisplayx (x−t)eλ1(x−t)f(t)dt]. 124. Consider/integraldisplay/integraldisplay DfdA whereDis the triangle with vertices at ( −1,1),(0,0) , and (3,1) . (a) Set up the iterated integrals in two ways. (b) Evaluate one of the integrals in (a) for the integrand f(x,y) = (x+y)2. 125. When a double integral was set up for the mass Mof a certain plate with density f(x,y) , the following sum of iterated integrals was obtained M=/integraldisplay2 1(/integraldisplayx3 xf(x,y)dy)dx+/integraldisplay8 2(/integraldisplay8 xf(x,y)dy)dx. (a) Sketch the domain of integration and express Mas an iterated integral in which the order of integration is reversed. (b) Evaluate Mif f(x,y) =/radicalbiggx y. 126. Evaluate/integraldisplay1 0/integraldisplay1 0xydxdy . 127. It is difficult to evaluate the integral I=/integraldisplay/integraldisplay DfdA , wheref(x,y) =1 1+x+y2andD is the indicated rectangle. However, you can show that (trivially) 1 3<I <3 2, and, with a bit more effort but the same method, that 1 2<I <3 2. Please do so. 410 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS 128. Consider the integral I=/integraltext/integraltext DfdA , where f(x,h) =3 8 +/radicalbig x4+y4 andDis the domain inside the curve x4+y4= 16 . Show that 2√ 2<I < 6. [Hint: Show that1 4<f <3 8inD. Then approximate the area of Dby an inscribed and circumscribed square. For the record, it turns out that I=3A 4ln(3 2) , whereA is the area =2 πΓ(1 4)2]. 129. (a) Find the derivative matrix for the following mappings Y=F(X) at the given pointX0. (i)F:/braceleftbiggy1=x2 1+ sinx1x2 y2=x2 2+ cosx1x2atX0= (0,0) (ii)F:  y1=x2 1+x3ex2−x3 2 y2=x1−3x2+x1logx3 y3=x2+x3 y4=fx1x2x3atX0= (2,0,1) (b) Find the equation of the tangent plane to the above surfaces at the given point. 130. Consider the following map Ffrom E2→E2, the familiar change of variables from polar to rectangular coordinates. F:/braceleftbiggy1=x1cosx2 y2=x1sinx2 (a) Find the images of (i) the semi-infinite strip 1 ≤x1<∞,0≤x2≤π 2. (ii) the semi-infinite strip 0 ≤x1<∞,0≤x2≤3π 2. (b) Compute F/primeand detF/prime. 131. Given that up(x) =e3x+e−2x−2ex/2is a solution of au/prime/prime+bu/prime+cu=e3x, find the constants a,b, andc. 132. Evaluate the determinants of the following matrices. (a) 1 1 1 1 1−1 1 −1 1−1−1 1 1 1 −1−1 (b) 1 1y z t 2x z t y w2x20 0 0 w3x30 0 0 w4x40 0 0 . 411 133. For what value(s) of xis the following matrix invertible?  1 1 1 1 1 2 2223 1 3 3233 1x x2x3  (Hint: Observe that the determinant is a cubic polynomial all of whose roots are obvious). 134. Letf(x) =n/summationdisplay k=1aksinkx√πandg(x) =n/summationdisplay r=1bsinrx√π. Bydirect integration prove that /integraldisplayπ −πf(x)g(x)dx=n/summationdisplay j=1ajbj. After you are done, compare with Theorem 15, page 206-7 and its proof. 135. Let A= 1 1 0 0 1 0 0 0 0 0 0 −1 0 0 1 0  (a) Find det A. (b) FindA−1. (c) SolveAX=Y, whereY= 2 2 1 3 . (d) LetS={u:u(x) =aex+bex+csinx+dcosx}, wherea,b,c , anddare any real numbers, and define a linear operator L:S→Sby the rule Lu≡u/prime/prime−u/prime+u. Find the matrix eLeforLwith respect to the following basis for S: e1(x) =xex, e2(x) =x, e 3(x) = sinx, e 4(x) = cosx. (e) Use the above results to find a solution of Lu= 2xex+ 2ex+ sinx+ 3 cosx. 136. Letu1andu2be solutions of the homogeneous equation Lu≡a2(x)u/prime/prime+a1(x)u/prime+a0(x)u= 0. 412 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS (a) Show that W(x)≡W(u1,u2)(x) , the Wronskian of u1andu2satisfies the differential equation W/prime=−a1(x) a2(x)W. (b) Find the equation of (a) for the particular operator Lu≡x2u/prime/prime−2xu/prime+ 2u and solve it for Wunder the condition that W(1) = 1 . (c) Given that u1(x) =xis a solution of Lu= 0 for the operator of part (b), use the result of (b) to show that if u2is another solution of Lu= 0 , then u2 satisfies the equation u/prime 2−1 xu2=x, provided that W(x,u 2)(1) = 1 . (d) Solve the equation of part (c) under the assumption that u2(1) = 1 , and thus find a second independent solution of the equation Lu= 0 for the operator of part (b). (e) Generalize the idea of parts (c) - (d) by stating and proving some theorem. 137. Here are some linear transformations defined in terms of matrices. In each case, describe geometrically what the transformation does, by computing the images of the three parallelograms Q1: with vertices at (0 ,0),(2,0),(3,1),(1,1). Q2: with vertices at (1 ,2),(3,2),(4,3),(2,3). Q3: with vertices at (1 ,0),(0,2),(−1,0),(0,−2). (a)Diagonal Maps (Stretchings) L1=/parenleftbigg3 0 0 1/parenrightbigg , L 2=/parenleftbigga0 0 1/parenrightbigg , L 3=/parenleftbigg−4 0 0 6/parenrightbigg , L4=/parenleftbigg1 0 0−1/parenrightbigg , L 5=/parenleftbigg−1 0 0−1/parenrightbigg , L 6=/parenleftbigg−2 0 0 0/parenrightbigg , L7=/parenleftbigga0 0a/parenrightbigg , L 8=/parenleftbigg1 0 0b/parenrightbigg , L 9=/parenleftbigga0 0b/parenrightbigg , (Remember to consider negative values of aandb). (b)Maps with 0 on the diagonal . L1=/parenleftbigg0 1 0 0/parenrightbigg , L 2=/parenleftbigg0a 0 0/parenrightbigg , L 3=/parenleftbigg0 0 −1 0/parenrightbigg , L4=/parenleftbigg0 1 1 0/parenrightbigg , L 5=/parenleftbigg0a 1 0/parenrightbigg , L 6=/parenleftbigg0a b0/parenrightbigg . 413 (c)Upper Triangular Matrices . L1=/parenleftbigg1 1 0 1/parenrightbigg , L 2=/parenleftbigg1−1 0 1/parenrightbigg , L 3=/parenleftbigg1 1 0−1/parenrightbigg , L4=/parenleftbigg1a 0 1/parenrightbigg , L 5=/parenleftbigg−1−1 0 0/parenrightbigg , L 6=/parenleftbigga1 0b/parenrightbigg . (d)Orthogonal Matrices (Rotations and Reflections). L1=/parenleftbigg0−1 1 0/parenrightbigg L3=/parenleftbigg3 54 5 L2=4 5−3 5/parenrightbigg L4=/parenleftBigg −1√ 21√ 21√ 21√ 2/parenrightBigg . 138. Letaandbbe real numbers such that a2+b2= 1 . Let S=/parenleftbigga2−b22ab 2ab b2−a2/parenrightbigg , P =/parenleftbigga2ab ab b2/parenrightbigg , and lete1= (a,b), e 2= (−b,a) , soe1⊥e2. Show that (a)Se1=e2, Pe 1=e1 (b)Se2=−e2, Pe 2= 0 (c)S2=I, P2=P (d) Show that Scan be interpreted as the reflection which leaves the line through e1fixed, and that Pcan be interpreted as the projection onto the line through e1parallel to e2. 139. (a) Consider the following relation defined on the set of allintegers:nRmifnand mare both even integers. Verify that this relation is symmetric and transitive - but not reflexive (since, for example, 1 R/1 ). (b) Let Rbe a symmetric and transitive relation defined on a set A. If, given any elementxinA, there is some element yrelated to it, xRy, prove that the relation Ris also reflexive. (The example in part (a) shows that the assertion will be false if some element is related to no others). 140. Letanbe a decreasing sequence of positive real numbers which satisfy an−1an+1≤ a2 n. If/summationtexta1/n nconverges, prove that/summationtextan an−1converges too. [Hint: Show that (an/an−1)1/n≤an]. 141. (a) Prove that the series/summationtextanznand/summationtexta2 nznhave the same radii of convergence. (b) Prove that the series/summationtextanznand/summationtext(an)kzn, wherek>0 , have the same radii of convergence. 142. LetVbe a linear space and Lan invertible linear map, L:V→V. If{e1,...,e n} is a basis for V, prove that its image {Le1,Le 2,...,Le n}is also a basis for V. 143. LetHbe an inner product space and Ran orthogonal transformation, R:H→ H. If{e1,...,e n}is a complete orthonormal set for H, prove that its image {Re1,...,Re n}is also a complete orthonormal set for H. 414 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS 144. (a) Let Rbe an orthogonal matrix and let ρ1andρ2be any two of its column vectors. Prove that ρ1⊥ρ2. Prove that any two rows of an orthogonal matrix are also orthogonal to each other. (b) Conversely, let Abe a square matrix whose column vectors are orthogonal. Must Abe an orthogonal matrix? Proof or counterexample. 145. LetHbe an inner product space and Athe subspace of Hspanned by the vectors X1,...,X n. The Gram determinant of those vectors is defined as G(X1,...,X n) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/angbracketleftX1, X 1/angbracketright ··· /angbracketleftXn, X 1/angbracketright /angbracketleftX1, X 2/angbracketright · · · · · · · /angbracketleftX1, Xn/angbracketright ··· /angbracketleftXn, Xn/angbracketright/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle (a) Prove that X1,···,Xnare linearly dependent if and only if G(X1,···,Xn) = 0 . [Suggestion: If Z∈A, thenZ=a1X1+···anXn, where the scalars a1,···,an are to be found. This can be done in two ways, by Theorem 31, page 428, or by solving the nequations /angbracketleftZ, X 1/angbracketright=a1/angbracketleftX1, X 1/angbracketright+···+an/angbracketleftXn, X 1/angbracketright · · · · /angbracketleftZ, X n/angbracketright=a1/angbracketleftX1, Xn/angbracketright+···+an/angbracketleftXn, Xn/angbracketright which are obtained from /angbracketleftZ, X j/angbracketright=/angbracketleftaiX1+···+anXn, Xj/angbracketright. Couple both methods to prove the result]. (b) IfX1,···,Xnare an orthogonal set of vectors, compute G(X1,···,Xn) . (c) IfY∈H, prove that the distance of Yfrom the subspace A,/bardblY−PAY/bardbl=δ, is given by the formula δ2=/bardblY−PAY/bardbl2=G(Y,X 1,...,X n) G(X1,...,X n). [Suggestion: Observe that δ2=/bardblY−PAY/bardbl2=/angbracketleftY−PAY, Y/angbracketrightand that /angbracketleftPAY, Y/angbracketright= an/angbracketleftX1, Y/angbracketright+···+an/angbracketleftXn, Y/angbracketright. Now write PAYasZ, use thenequations in a) and the one equation δ2=/angbracketleftY, Y/angbracketright −a1/angbracketleftX1, Y/angbracketright − ··· −an/angbracketleftXn, Y/angbracketrightto solve for δ2by using Cramer’s rule]. (d) Use the fact that G(X1) =/angbracketleftX1, X 1/angbracketrightto prove the Gram determinant of lin- early independent vectors is always positive . In particular, deduce the Cauchy - Schwarz inequality from G(X1,X2)≥0 . (e) InL2[0,1] , letX1= 1+x, andX2=x3. Compute G(X1,X2) . LetY= 2−x4 and compute /bardblY−PAY/bardbl, whereAis the subspace spanned by X1andX2. 415 (f) (Muntz) In L2[0,1] , letAn= span {xj1,xj2,···,xjn}wherej1,···,jnare distinct positive integers. Let Y=xk, wherekis a positive integer by not one of thej’s. Prove that lim n→∞/bardblY−PAnY/bardbl= 0 if and only if/summationdisplay1 jndiverges. 146. (a) Use Theorem 17, page 217 to find linear polynomials PandQsuch that, respectively, (i)/integraldisplay1 −1[x2−P(x)]2dxis minimized, (ii)/integraldisplay1 0[x2−Q(x)]2dxis minimized. (b) WriteP(x) =a+bxand use calculus to again find the values of aandbsuch that/integraldisplay1 −1[x2−P(x)]2dx is minimized. 147. LetZ= (1,1,1,1,1)∈E5and letAbe the subspace of E5spanned by X1= (1,0,1,0,0), X 2= (1,0,0,−1,0) , andX3= (0,1,0,0,1) . Find /bardblZ−PAZ/bardbl. 148. Let Γ 0be a closed planar curve which encloses a convex region, and let Γ rbe the “parallel” curve obtained by moving out a distance of ralong the outer normal. (a) Discover a formula relating the arc length of Γ rto that of Γ 0. [Advise: Examine the special cases of a circle, rectangle, and convex polygon]. (b) Prove the result you conjectured in part a). 149. The hypergeometric function F(a,b;c;x) is defined by the power series F(a,b;c;x) = 1 +a·b 1·cx+a(a+ 1)b(b+ 1) 1·2·c(c+ 1)x2+a(a+ 1)(a+ 2)b(b+ 1)(b+ 2) 1·2·3c(c+ 1)(c+ 2)x2+··· (a) Show that the series converges for all |x|<1 . (b) Show thatd dxF(a,b;c;x) =ab cF(a+ 1,b+ 1;c+ 1;x) . (c) Show that (i) (1 −x)n=F(−n,b;b;x) (ii) (1 +x)n=F(−n,b;b;−x) (iii) log(1 −x) =−xF(1,1; 2,x) (iv) log(1+x 1−x) = 2xF(1 2,1;3 2;x2) (v)ex= lim b→∞F(1,b; 1;x/b) (vi) cosx=F(1 2,−1 2;1 2,sin2x) (vii) sin−1x=xF(1 2,1 2;3 2;x2) (viii) tan−1x=xF(1 2,1;3 2;−x2) 416 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS (d) Show that Fsatisfies the hypergeometric differential equation x(1−x)d2F dx2+ [c−(a+b+ 1)x]dF dx−abF = 0. [This equation is essentially the most general one with three regular singular points - in this case located at 0 ,1 , and ∞]. 150. Let {e1,···,en}be a complete orthonormal set of Enand let {X1,···,Xn}be a set of vectors which are close to the ej’s in the sense that n/summationdisplay j=1/bardblXj−ej/bardbl2<1. Prove that the Xj’s are linearly independent. Give an example in E3of linearly dependent vectors {X1,X2,X3}which satisfy n/summationdisplay j=1/bardblXj−ej/bardbl2= 1. [In fact, one can prove that dimA⊥≤n/summationdisplay j=1/bardblXj−ej/bardbl2, ] whereA= span {X1,···,Xn}]. 151. (a) Show that the function f(z) =ez, z∈C, is never zero. (b) Scrutinize the proof of the Fundamental Theorem of Algebra (pp. 544-548) and find where it breaks down if one attempts to extend it to prove that ezhas at least one zero.