math21-2011 Kazdan Part I
PDF · 425 pages · 1.9 MB
Open PDF file
Lecture notes by J. Kazdan for a second-year Harvard course combining linear algebra with intermediate calculus, with a 1964 preface and a 1966 afterword. Chapters cover review of sets and reals, infinite series and power series, vector spaces, norms and inner products, Fourier series, and linear operators. This is a downloaded copy of someone else's text kept in Phil's math book folder.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
INTERMEDIATE CALCULUS
AND
LINEAR ALGEBRA
Part I
J. KAZDAN
Harvard University
Lecture Notes
ii
Preface
These notes will contain most of the material covered in class, and be distributed before
each lecture (hopefully). Since the course is an experimental one and the notes written
before the lectures are delivered, there will inevitably be some sloppiness, disorganization,
and even egregious blunders—not to mention the issue of clarity in exposition. But we will
try. Part of your task is, in fact, to catch and point out these rough spots. In mathematics,
proofs are not dogma given by authority; rather a proof is a way of convincing one of the
validity of a statement. If, after a reasonable attempt, you are not convinced, complain
loudly.
Our subject matter is intermediate calculus and linear algebra. We shall develop the
material of linear algebra and use it as setting for the relevant material of intermediate
calculus. The first portion of our work—Chapter 1 on infinite series—more properly belongs
in the first year, but is relegated to the second year by circumstance. Presumably this topic
will eventually take its more proper place in the first year.
Our course will have a tendency to swallow whole two other more advanced courses,
and consequently, like the duck in Peter and the Wolf, remain undigested until regurgitated
alive and kicking. To mitigate—if not avoid—this problem, we shall often take pains to
state a theorem clearly and then either prove only some special case, or offer no proof at
all. This will be true especially if the proof involves technical details which do not help
illuminate the landscape. More often than not, when we only prove a special case, the
proof in the general case is essentially identical—the equations only becoming larger.
September 1964
iii
Afterward
I have now taught from these notes for two years. No attempt has been made to revise
them, although a major revision would be needed to bring them even vaguely in line with
what I now believe is the “right” way to do things. And too, the last several chapters
remain unwritten. Because the notes were written as a first draft under panic pressure,
they contain many incompletely thought-out ideas and expose the whimsy of my passing
moods.
It is with this—and the novelty of the material at the sophomore level—in mind, that
the following suggestions and students’ reactions are listed. There are three categories,
A), Material that turned out to be too difficult (they found rigor hard, but not many of
the abstractions), B), changes in the order of covering the stuff, and C), material—mainly
supplementary at this level—which is not too hard, but should be omitted if one ever hopes
to complete the ”standard” topics within the confines of a year course.
(A)It was too hard (unless one took vast chunks of time).
(1) Completeness of reals. Only “monotone sequences converge” is needed for infinite
series.
(2) Term-by-term differentiation and integration of power series. The statement of
the main theorem should be fully intelligible—but the proof is too complicated.
(3) Cosets. This is apparently too abstract. It might be possible to do after finding
general solutions of linear inhomogeneous O.D.E.’s.
(4)L2and uniform convergence of Fourier series. Again, all I ended up doing was to
try to state what the issues were, and not to attempt the proof. The ambitious
student should be warned that my proof of the Weierstrass theorem is opaque
(one should explicitly introduce the idea of an approximate identity).
(5) Fundamental Theorem of Algebra. The students simply don’t believe inequalities
in such profusion.
(6) I you want to see rank confusion, try to teach the class how to compute higher
order partial derivatives using the chain rule. That computation should be one
of the headaches of advanced calculus.
(7) Existence of a determinant function. I don’t know a simple proof except for the
one involving permutations—and I hate that one.
(8) Dual spaces. As lovely as the ideas are, this topic is too abstract, and to my
knowledge, unneeded at this level where almost all of the spaces are either finite
dimensional or Hilbert spaces. One should, however, mention the words “vector”
and “covector” to distinguish column from row vectors. I forgot to do so in these
notes and it did cause some confusion.
(B)Changes in Order and Timing . The structure of the notes is to investigate bare
linear spaces, then linear mappings between them, and finally non-linear mappings
between them. It is with this in mind that linear O.D.E.’s came before nonlinear maps
from Rn→R. The course ended by treating the simplest problem in the calculus
of variations as an example of a nonlinear map from an infinite dimensional space
iv
to the reals. My current feeling is to consider linear andnon-linear maps between
finite dimensional spaces before doing the infinite dimensional example of differential
equations.
The first semester should get up to the generalities on solving LX=Y, p. 319
[incidentally, the material on inverses (p. 355 ff) belongs around p. 319]. Most
students find the material on linear dependence difficult—probably for two reasons:
1) they are not used to formal definitions, and ii) they think they have learned a
technique for doing something, not just a naked definition, and can’t quite figure out
just what they can do with it. In other words, they should feel these definitions about
the anatomy of linear spaces are similar to those describing a football field and of
little value until the game begins—i.e., until the operators between spaces make their
grand entrance.
Because of time shortages, the sections on linear maps from R1→RnandRn→R1,
pp. 320-41 were regrettably omitted both years I taught the course. The notes were
written so that these sections can be skipped.
(C)Supplementary Material . A remarkable number of fascinating and important topics
could have been included—if there were only enough time. For example:
(1) Change of bases for linear transformations (including the spectral theorem).
(2) Elementary differential geometry of curves and surfaces.
(3) Inverse and implicit function theorems. These should be stated as natural gener-
alizations of the problems of a) inverting a linear map, b) finding the null space
of a linear map, and c) generalizing dim D(L) = dimR(L) + dimN(L) all to
local properties of nonlinear maps via the tangent map.
(4) Change of variable in multiple integration. Determinants were deliberately in-
troduced as oriented volume to make the result obvious for linear maps and
plausible for nonlinear maps.
(5) Constrained extrema using Lagrange multipliers.
(6) Line and surface integrals along with the theorems of Gauss, Green, and Stokes.
The formal development of differential forms takes too much time to do here.
Perhaps a satisfactory solution is to restrict oneself to line integrals and these
theorems in the plane, where the topological difficulties are minimal.
(7) Elementary Morse Theory. One can prove the Morse inequalities easily for the
real line, the circle, the plane, and S2merely by gradually flooding these sets
and observing the number of lakes and shore line changes only at the critical
points.
(8) Sturm-Liouville theory. An elegant fusion of the geometry of Hilbert spaces to
differential equations.
(9) Translation-invariant operators with applications to constant coefficient differ-
ence and differential equations. The Laplace and Fourier transforms enter natu-
rally here.
(10) The Calculus of Variations. The formalism of nonlinear functionals on R/multicloseleft, i.e.,
mapsf:R/multicloseleft→R, generalizes immediately to nonlinear functionals defined on
infinite dimensional spaces.
v
(11) The deleted rigor.
(12) Linear operators with finite dimensional (perhaps even compact) range.
One parting warning. When covering intermediate calculus from this viewpoint, it is
all too natural to forget the innocence of the class, to enchant with glitter, and to numb
with purity and formalism. Emphasis should be placed on developing insight and intuition
along with routine computational facility.
My classes found frequent reviews of the mathematical edifice, backward glances at the
previous months’ work, not only helpful but mandatory if they were to have any conception
of the vast canvas which was being etched in their minds over the course of the year. The
question, “What are we doing now and how does it fit into the larger plan?” must constantly
be raised and at least partially resolved.
May, 1966
Contents
0 Remembrance of Things Past. 1
0.1 Sets and Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1
0.2 Relations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
0.3 Mathematical Induction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
0.4 Reals: Algebraic and Order Properties . . . . . . . . . . . . . . . . . . . . . 7
0.5 Reals: Completeness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
0.6 Appendix: Continuous Functions and the Mean Value Theorem . . . . . . . 15
0.7 Complex Numbers: Algebraic Properties . . . . . . . . . . . . . . . . . . . . 22
0.8 Complex numbers: Completeness and Functions . . . . . . . . . . . . . . . 28
1 Infinite Series 33
1.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33
1.2 Tests for Convergence of Positive Series . . . . . . . . . . . . . . . . . . . . 36
1.3 Absolute and Conditional Convergence . . . . . . . . . . . . . . . . . . . . . 41
1.4 Power Series, Infinite Series of Functions . . . . . . . . . . . . . . . . . . . . 43
1.5 Properties of Functions Represented by Power Series . . . . . . . . . . . . . 48
1.6 Complex-Valued Functions, ez,cosz,sinz. . . . . . . . . . . . . . . . . . . 65
1.7 Appendix to Chapter 1, Section 7. . . . . . . . . . . . . . . . . . . . . . . . 70
2 Linear Vector Spaces: Algebraic Structure 75
2.1 Examples and Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 75
a) The Space R2. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 75
b) The Space Rn. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76
c) The Space C[a,b] . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77
d) D. The Space Ck[a,b] . . . . . . . . . . . . . . . . . . . . . . . . . . 77
e) E. The Space l1.. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 78
f) F. The Space L1[a,b] . . . . . . . . . . . . . . . . . . . . . . . . . . . 78
g) G. The Space fn. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79
h) Appendix. Free Vectors . . . . . . . . . . . . . . . . . . . . . . . . . 80
2.2 Subspaces. Cosets. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 84
2.3 Linear Dependence and Independence. Span. . . . . . . . . . . . . . . . . . 88
2.4 Bases and Dimension . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93
3 Linear Spaces: Norms and Inner Products 101
3.1 Metric and Normed Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . 101
3.2 The Scalar Product in E2. . . . . . . . . . . . . . . . . . . . . . . . . . . . 107
3.3 Abstract Scalar Product Spaces . . . . . . . . . . . . . . . . . . . . . . . . . 113
vii
viii CONTENTS
3.4 Fourier Series. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 132
3.5 Appendix. The Weierstrass Approximation Theorem . . . . . . . . . . . . . 140
3.6 The Vector Product in R3. . . . . . . . . . . . . . . . . . . . . . . . . . . . 146
4 Linear Operators: Generalities. V1→Vn,Vn→V1147
4.1 Introduction. Algebra of Operators . . . . . . . . . . . . . . . . . . . . . . . 147
4.2 A Digression to Consider au/prime/prime+bu/prime+cu=f. . . . . . . . . . . . . . . . . . 161
4.3 Generalities on LX=Y. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 170
4.4L:R1→Rn. Parametrized Straight Lines. . . . . . . . . . . . . . . . . . . 177
4.5L:Rn→R1. Hyperplanes. . . . . . . . . . . . . . . . . . . . . . . . . . . . 182
5 Matrix Representation 187
5.1L:Rm→Rn. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 187
5.2 Supplement on Quadratic Forms . . . . . . . . . . . . . . . . . . . . . . . . 210
5.3 Volume, Determinants, and Linear Algebraic Equations. . . . . . . . . . . . 217
a) Application to Linear Equations . . . . . . . . . . . . . . . . . . . . 234
5.4 An Application to Genetics . . . . . . . . . . . . . . . . . . . . . . . . . . . 243
5.5 A pause to find out where we are . . . . . . . . . . . . . . . . . . . . . . . . 246
6 Linear Ordinary Differential Equations 249
6.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 249
6.2 First Order Linear . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 252
6.3 Linear Equations of Second Order . . . . . . . . . . . . . . . . . . . . . . . 258
a) A Review of the Constant Coefficient Case. . . . . . . . . . . . . . . 258
b) Power Series Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . 259
c) General Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 266
6.4 First Order Linear Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . 278
6.5 Translation Invariant Linear Operators . . . . . . . . . . . . . . . . . . . . . 283
6.6 A Linear Triatomic Molecule . . . . . . . . . . . . . . . . . . . . . . . . . . 286
7 Nonlinear Operators: Introduction 293
7.1 Mappings from R1toR1, a Review . . . . . . . . . . . . . . . . . . . . . . 293
7.2 Generalities on Mappings from RntoRm. . . . . . . . . . . . . . . . . . . 295
7.3 Mapping from E1toEn. . . . . . . . . . . . . . . . . . . . . . . . . . . . 300
8 Mappings from EntoE: The Differential Calculus 309
8.1 The Directional and Total Derivatives . . . . . . . . . . . . . . . . . . . . . 309
8.2 The Mean Value Theorem. Local Extrema. . . . . . . . . . . . . . . . . . . 321
8.3 The Vibrating String. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 332
a) The Mathematical Model . . . . . . . . . . . . . . . . . . . . . . . . 333
b) Uniqueness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 334
c) Existence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 336
8.4 Multiple Integrals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 347
9 Differential Calculus of Maps from EntoEm, s. 361
9.1 The Derivative . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 361
9.2 The Derivative of Composite Maps (“The Chain Rule”). . . . . . . . . . . . 373
10 Miscellaneous Supplementary Problems 383
Chapter 0
Remembrance of Things Past.
We shall treat a hodge-podge of topics in a hasty and incomplete fashion. While most
of these topics should have been learned earlier, section 5 on the completeness of the real
numbers has its more rightful place in advanced calculus. Do nottake time to read this
chapter unless the particular topic is needed; then read only the relevant portions. The
chapter is included for reference.
0.1 Sets and Functions
Asetis any collection of objects, called the elements of the set, together with a criterion
for deciding if an object is in the set. For example, I) the set of all girls with blue eyes and
blond hair, and ii) the less picturesque set of all positive even integers. We can also define a
set by bluntly listing all of its elements. Thus, the set of all students in this class is defined
by the list in the roll book.
Sets are often specified by a notation which is best described by examples.
i)S={x:xis an integer }is the set of all integers.
ii)T={(x,y):x2+y2= 1}is the set of all points ( x,y) on the unit circle x2+y2= 1 .
iii)A={1,2,7,−3}is the set of integers 1 ,2,7 and −3 .
Our attitude toward set theory will be extremely casual; we shall mainly use it as a
language and notation. Without further ado, let us introduce some notation.
x∈S, x is an element of the set S, or justxis inS.
x/negationslash∈S, x is not an element of the set S.
Z, the set of all integers, positive, zero, and negative.
Z+, the set of all positive integers, excluding 0.
R the set of all real numbers (to be defined more precisely later).
C, The set of all complex numbers (also to be defined more precisely later).
∅, the set with no elements, the empty ornullset. It is extremely uninteresting.
Definition: Given the two sets SandT, i) the set S∪T, “SunionT”, is the set of
elements which are in eitherSorT, or both.
ii) The set S∩T, “Sintersection T”, is the set of elements in bothSandT.
If we represent Sby one blob and Tby another, S∪Tis the shaded region while
S∩Tis the cross-hatched region. Note that all elements in S∩Tare also in S∪T. Two
sets are disjoint ifS∩T=∅, that is, if their intersection is empty.
1
2 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
Asubset of a set is another way of referring to a portion of a given set. Formally, Ais
the subset of S, writtenA⊂S, if every element in Ais also an element of S. The set
Ais a subset of the set Sif and only if either
A∪S=S,or, equivalently, A∩S=A.
It is possible that A=S, or thatA=∅. If these degenerate cases are excluded, we say
thatAis aproper subset ofS.
Given the two sets SandT, it is natural to form a new set S×T, “ScrossT”,
which consists of all pairs of elements, one from Sand the other from T. For example, if
Sis the set of all men in this class, and Tthe set of all women in this class, then S×T
is the set of all couples, a natural set to contemplate.
Ifx∈Sandy∈T, the standard notation for the induced element in S×Tis (x,y) .
Note that the order in ( x,y) is important. The element on the left is from S, while that on
the right is from T. For this reason the pair of elements ( x,y) is usually called an ordered
pair. The whole set S×Tis called the product ,direct product , orCartesian product ofS
andT, all three names being used interchangeably.
You have met this idea in graphing points in the plane. Since these points, ( x,y) , are
determined by an ordered pair of real numbers, they are just the elements of R×R. From
this example it is clear that even though this set R×Ris the product of a set with itself,
theorder of the pair ( x,y) is still important. For example the point (1 ,2)∈R×Ris
certainly not the same as (2 ,1)∈R×R.
Having defined the direct product of two sets SandTas ordered pairs, it is reasonable
to define the direct product of three sets S, T, andUas the set of ordered triplets ( x,y,z ) ,
wherex∈S, y∈T, andz∈U. The extension to nsets,S1×S2× ··· ×Sn, is done in
the same way.
Let us now recall the ideas behind the notion of a function.
Afunctionffrom the set Xinto the set Bis a rule which assigns to every x∈X
one and only one element y=f(x)∈B. We shall also say that fmapsXintoB, and
write either
f:X→B,orXf→B.
This alternative notation is useful when XandBare more important than the specific
nature off. The setXis the domain off, while the range offis the subset Y⊂Bof
all elements y∈Bwhich are the image of (at least) one point x∈X, soy=f(x) , or in
suggestive notation, Y=f(X).
Automobile license plates supply a nice example, for they assign to every license plate
sold a unique car. The domain is the set of all license plates sold, while the range is not all
cars, but rather the subset of all cars which are driven. Wrecks and museum pieces neither
need nor have license plates since they are not on the roads. Some other examples are i)
the function f(n)≡1
n, n= 1,2,3,... which assigns to every n∈Z+the rational number
1
n, and ii) the function f(n,m) =m
n, n, m = 1,2,2,..., which assigns to every element of
Z+×Z+the rational numberm
n.
Quite often we shall use functions which map part of some set into part of some other
set. In other words the function may be defined on only a subset of a given set and take on
values in a subset of some other set. The function f(n,m)≡m
nof the previous paragraph
is of this nature for we defined it on a subset of Z×Zand takes its values on the positive
subset of the set of all rational numbers.
0.1. SETS AND FUNCTIONS 3
There is some standard nomenclature (or $10 words if you like) associated with map-
ping. Say X⊂Aand the function f:X→B. Note that we know the definition of f
only onX. It may not be defined for the remainder of A.
Definition: i) ifevery element of Bis the image of (at least) one point in X, the map
fis called surjective oronto. In other words f:X→Bis a surjection if the range of f
is all ofB. Thusfis always surjective onto its range.
ii) If the map fhas the property that for every x1,x2∈X, we havef(x1) =f(x2)
when and only when x1=x2, the map is called injective orone to one (1-1). This is the
case if no two different elements in Xare mapped into the same element in B.
iii) If the map fis both surjective and injective, that is, if it is both onto and 1-1, then
fis called bijective .
Examples: For these, we have f:X→BwhereX=B=Z.
(1) The map f(n) = 2nis injective but not surjective since the range does not contain
the odd integers in B.
(2) The map f(n) =/braceleftbiggn
2ifnis even
n+1
2ifnis oddis surjective but not injective since every ele-
ment inBis the image of two distinct elements of X.
(3) The map f(n) =n+ 7 is bijective.
Notational Remark : For functions whose domain is ZorZ+it is customary to indicate
the element of the range by a notation like aninstead off(n) . Thusf(n) =1
n, where
n∈Z+, is written as an=1
n. Such a function is usually called a sequence .
The concepts we have just defined are useful if we try to define what we mean by the
inverse of a function.
Definition: A function f:X→Bisinvertible if to every b∈Bthere is one and only
onex∈Xsuch thatb=f(x) . Thusfis invertible if and only if it is bijective. If fis
invertible, we denote the inverse function by f−1, sox=f−1(b) .
Iff:A→B, andg:B→C, then when composed (put together) these two functions
induce a mapping, g◦f, ofAintoC. Slightly more generally, if B⊂R, andf:A→B
whileg:R→C, th eng◦f:A→C.
You should be able to see why the composed map g◦fis only defined on A, and then
understand that our stipulation that B⊂Ris a convenient requirement.
Ifx∈Aandz∈C, theng◦fmapsxontoz= (g◦f)(x) , or in more familiar
notation,z=g(f(x)) . Now an example. Say the distance syou have walked at time tis
specified by the function s=f(t) , and the amount zof shoe leather worn out by walking
the distance sis given by the function z=g(s) . Then the amount of shoe leather you have
worn out at time tis given by the composed function z=g(f(t)) . Heret∈A, s∈B,
andz∈C. Hopefully you have by now recognized that the “chain rule” for derivatives is
just the procedure for finding the derivative of composed functions from their constituent
parts. In our example the chain rule would be used to finddz
dtfromdg
dsandd f
dt-if these
functions were differentiable.
We conclude this section with more symbols—if you have not yet had enough. These
are borrowed from logic. Although we shall use them only infrequently as a shorthand, they
might have greater use to you in class notes.
∀ “for every”
4 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
∃ “there is”, or “there exists”
/owner “such that”
A⇒B“the truth of statement Aimplies that of statement B”.
A⇔B“statement Ais equivalent to statement B, that is, both A⇒Band
B⇒A.
Exercises
(1) IfR={1,4}, S={1,2,3,4,}, andT={2,3,7}, find the six other sets R∪S, R∩
S, R∪T, R∩T, S∪T,andS∩T. Which of these nine sets are proper subsets of
which other sets?
(2) IfS={x:|x−1| ≤2}andT={x:|x| ≤2}, findS∪TandS∩T. A sketch
is adequate.
(3) IfA, B , andCare any subsets of a set S, prove
(a) (A∪B)∪C=A∪(B∪C) —so that the parenthesis can be omitted without
creating ambiguity.
(b) (A∩B)∩C=A∩(B∩C) —so that again the parentheses are superfluous.
(c) (A∪B)∩C= (A∩C)∪(B∩C).
(d) (A∩B)∪C= (A∪C)∩(B∪C).
Remark: two setsXandYare proved equal by showing that both X⊂Yand
Y⊂X.
(4) If the function fhas domain S, and both A⊂CandB⊂S, prove that
(i)A⊂B⇒f(A)⊂f(B) .
(ii)f(A∩B)⊂f(A)∩f(B) [We cannot hope to prove equality because of coun-
terexamples like: let A={−2,−1,0,1,2,3}andB={−4,−3,−2,−1}.
Then with f(n) =n2, we havef(A) ={0,1,4,9}, f(B) ={1,4,9,16}, and
f(A∪B) ={1,4} /negationslash=f(A)∩f(B) ].
(iii)f(A∪B) =f(A)∪f(B) .
(5) For the following functions f:X→B, classify as to injection, surjection, or bijection,
or none of these.
(i)f(n) =n2withX=Z+andB=Z.
(ii) LetX={all rational numbers },B={all rational numbers }, andf(x) =1
m,
wherex=n
m∈X[Heren
mis assumed to be reduced to lowest terms.]
(iii)f(x) =1
x, wherex∈XandX=B={all positive rational numbers }.
(iv)X={all women born in May }, B={the thirty days in the month of June },
and letfbe the function assigning “her birthday” to each woman born in June.
(v)f(n) =|n|, withX=B=Z.
0.2. RELATIONS 5
0.2 Relations
A relationship often exists between elements of sets. Some common examples are i) a≥b,
ii)a⊥b(perpendicular to), iii) alovesb, and iv)a/negationslash=b. LetSbe a given set, a,
b∈S, and let Rbe a relation defined on S(that is, ∀a, b∈S, eitheraRboraRbwith
no third alternative possible). Most relations have at least one of the following properties.
(i)reflexiveaRa ∀a∈S
(ii)symmetric aRb⇒bRa
(iii) transitive (aRbandbRc)⇒aRc.
Examples:
(1) perpendicular ( ⊥) is only symmetric.
(2) “loves” enjoys none of these (well, maybe it is reflexive).
(3) equality ( = ) has all three properties.
(4) geometric congruence ( ∼=) and geometric similarity ( /similarequal) both have all three.
(5) parallel ( /bardbl) has all three—if we are willing to agree that a line is parallel to itself.
(6) “is less than five miles from” is only reflexive and symmetric.
(7) fora, b∈Z+, the relation “ ais divisible by b” is only reflexive and transitive but
not symmetric.
(8) “less than” ( <) is only transitive.
A relation which is reflexive, symmetric and transitive is called an equivalence relation .
The standard examples are those of algebraic equality and of geometric congruence. An
equivalence relation on a set Spartitions the set into subsets of equivalent elements . Those
terms are illustrated in the following.
Examples:
(1) In the set Sof all triangles, the equivalence relation of geometric congruence parti-
tionsSinto subsets of congruent triangles, any two triangles of Sbeing in the same
subset (or equivalence class as it is called) if and only i f they are congruent.
(2) In the set Pof all people, consider the equivalence relation ”has the same birthday,”
disregarding the year. This relation partitions Pinto 366 equivalence classes. Two
people are in the same equivalence class if their birthdays fall on the same day of the
year.
Notice that any two equivalence classes are either identical or disjoint, that is, they
have either no elements in common or they coincide. This is particularly clear from the
examples with birthdays.
By the fundamental theorem of calculus, we know that the indefinite integral of an
integrable function fcan be represented by any function Fwhose derivative is f. The
6 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
mean value theorem told us that every other indefinite integral of fdiffers from Fby
only a constant. Thus, the indefinite integrals of a given function are an equivalence class
of functions, differing from each other by constants. The equivalence relation is “equal up
to an additive constant”.
Exercises
(1) Ifa, b, c, d ∈Z+, let us define the following equivalence relation between the elements
ofZ+×Z+:
(a,b)R(c,d) if and only if ad=bc.
Verify that Ris an equivalence relation. [In real life, the pair ( a,b) of this example
is written asa
b, so all we have said isa
b=c
dif and only if ad=bc. This equivalence
relation partitions the set of rational numbers into very familiar equivalence classes.
For example the equivalent rational numbers1
2,2
4,3
6,...are in the same equivalence
class, to no one’s surprise].
(2) Explain the fallacy in the following argument by observing that equality “ = ” here is
not the usual algebraic equality , but rather some other equivalence relation.
“Let =/integraltextdx
x. Integration by parts ( p= 1/x,dq=dx), gives
A=x(1
x)−/integraldisplay
x(−1
x2)dx= 1 +A.
Hence 0 = 1 .”
0.3 Mathematical Induction
You are familiar with a variety of proofs, viz. direct proofs and proofs by contradiction.
There is, however, another type of proof which is not encountered very often in elementary
mathematics: proof by induction .
Abstractly, you have a sequence of statements P1, P2, P3,..., and a guess for the nature
of the general statement Pn. A proof by mathematical induction provides a method for
showing the general statement Pnis correct. Here is how it is carried out. First verify
that the statement is true in some special case, say for n= 1 , so you check the validity of
P1.Second you show that ifit is true in some particular case n=k, then it is true for
the next case n=k+ 1 , that is, Pk⇒Pk+1. Now since P1is true, so is P1+1=P2, and
consequently so is P2+1=P3, and so on up. Observe that the procedure does not tell you
how in the world to guess the general statement Pn, but only shows how to verify it.
Let us carry out the procedure for an example. We guess the formula
1 + 2 + ···+n=n(n+ 1)
2(0-1)
step 1. Is the formula true for n= 1 ? Yes, since both sides then equal 1.
step 2. Assuming the formula is true for n=k, we must show this implies the formula is
true forn=k+ 1 .
1 + 2 + ···+k+ (k+ 1) =(k+ 1)(k+ 2)
2.
0.4. REALS: ALGEBRAIC AND ORDER PROPERTIES 7
The formula, assumed to be true, for n=kis
1 + 2 + ···+k=k(k+ 1)
2.
Adding (k+ 1) to both sides we find that
1 + 2 + ···+k+ (k+ 1) =k(k+ 1)
2+ (k+ 1) =(k+ 1)(k+ 2)
2
which is exactly the statement we wanted. This proves that formula (0.3) is true for all
n≥1 .
Exercises
Use mathematical induction to prove the given statements.
(1) 12+ 22+···+n2=n(n+1)(2 n+1)
6
(2)d
dx(xn) =nxn−1(use the formula for the derivative of a product).
(3) LetI(n) =/integraltextπ
2
0sinnxdx
(a) Prove the following formula is correct when nis an odd integer ≥3 ,
I(n) =2·4·6· · · ·(n−1)
1·3·5· · · ·n
(b) Guess and prove the formula when nis an even integer ≥2 .
(4) Let Γ(s) =/integraltext∞
0e−tts−1dt, wheres>0 (this is the famous gamma function ).
(a) Show Γ( s+ 1) =sΓ(s) (Hint: integrate by parts)
(b) Ifn∈Z+, guess and prove the formula for Γ( n+ 1) .
0.4 The Real Numbers: Algebraic and Order Properties.
The set of all real numbers can be characterized by a set of axioms. These properties are of
three different types, i) algebraic properties, ii) order properties, and iii) the completeness
property. Of these, the last is by far the most difficult to grasp. But that is getting ahead
of our story. Let Sbe a set with the following properties.
I. Algebraic Properties
A. Addition. To every pair of elements a,b∈S, is associated another element, denoted
bya+b, with the properties
A - 0. (a+b)∈S
A - 1. Associative: for every a,b,c∈S,a+ (b+c) = (a+b) +c.
A - 2. Commutative: a+b=b+a
A - 3. There is an additive identity , that is, an element ”0” ∈Ssuch that 0 + a=a
for alla∈S.
A - 4. For every a∈S, there is also a b∈Ssuch thata+b= 0 .bis the additive
inverse ofa, usually written −a.
8 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
M. Multiplication. To every pair a,b∈S, there is associated another element, denoted
byab, with the properties
M - 0.ab∈S
M - 1. Associative. For every a,b,c∈S, a(bc) = (ab)c.
M - 2. Commutative. ab=ba.
M - 3. There is a multiplicative identity , that is, an element “ l”∈Ssuch thatla=a
for alla∈S. Moreover 1 /negationslash= 0 .
M - 4. For every a∈S, a/negationslash= 0 , there is also a b∈Ssuch thatab= 1 .bis the
multiplicative inverse ofa, usually written1
aora−1.
D. Connection between Addition and Multiplication .
D - 1. Distributive . For every a,b,c∈S, a(b+c) =ab+ac.
Some sample - and simple—consequences of these nine axioms are i) a+ 0 =a, ii)
a·1 =a, and iii)a+b=a+c⇒b=c.
Any set whose elements satisfy the axioms A-0 to A-4 is called a commutative (or
abelian )group . The group operation here is addition. In this language, we see that the
multiplication axioms just state that the elements of S—with the additive identity 0
excluded—also form a commutative group, with the group operation being multiplication.
These additive and multiplicative structures are connected by the distributive axiom. Most
of high school algebra takes place in this setting; however, the possibility of non-integer
exponents is not yet specifically included; in particular the square root of an element of S
is not necessarily also in S.
Our axioms, or some part of them, are satisfied by sets other than the real numbers. The
set of even integers form a commutative group with the group operation being addition,
while numbers of the form 2n, n∈Z, form a commutative group under multiplication.
The set of rational numbers satisfies all nine axioms. Any such set which satisfies all nine
axioms is called a field. Both the real numbers and the rational numbers (a subset of the
real numbers) are fields. A more thorough investigation of groups and fields is carried out
in courses in modern algebra.
II. Order Axioms
Besides the above algebraic rules, we shall introduce an order relation , intuitively, the
notion of ’greater than”. To do this we need to use an undefined concept of positivity for
elements of Sand use it to state our axioms.
O -1. Ifa∈Sandb∈Sare positive, so are a+bandab.
O -2. The additive identity 0 is not positive.
O - 3. For every a∈S, a/negationslash= 0 , either aor−ais positive, but not both. If −ais
positive, we shall say that ais negative.
Trichotomy Theorem . For any two numbers a,b∈S, exactly one of the following three
statements is true, i) a−bis positive, ii) b−ais positive, or iii) b−ais zero. If the
notationa < b is used to mean “ b−ais positive,” and a > b meansb < a , then this
theorem reads, either a>b, a<b, ora=b. The proof—which you should do—is a simple
consequence of our axioms.
Some other consequences are
a<b andb<c⇒a<c (transitivity of “ <”)
a<b andc>0⇒ac<bc
a/negationslash= 0⇒a2>0 . (Since 1 = 12, this implies 1 >0 ).
The set of rational numbers as well as the set Rof real numbers satisfy all twelve
axioms. Any set which satisfies these twelve axioms is called an ordered field .
0.5. REALS: COMPLETENESS 9
Exercises
(1) LetTbe a set whose elements are of the form a+b√
2 , whereaandbare rational
numbers (and so are elements of a field). Show that Tis also a field.
(2) Consider the set of all integers Zwith the following equivalence relation: m∈Z
andn∈Zare equivalent if they have the same remainder when divided by 2. The
notation for this equivalence is
m≡n(mod2)
This equivalence relation partitions Zinto two equivalence classes which we may
denote respectively by 0 if the number is even, and 1 if the number is odd . Thus
8≡ −22 (mod 2) and 7 ≡13 (mod 2). Prove that the set Zwith ordinary addition
and multiplication but with this equivalence relation forms a field.
(3) Prove the trichotomy theorem.
(4) Prove that if a/negationslash= 0 , thena2>0 . Use it to prove that 1 >0 and then to conclude
that all of the ’positive integers” are, in fact, positive.
0.5 The Real Numbers: Completeness Property.
III. Completeness Axiom.
So far our axioms do not insure that we can take fractional powers like the square root,
of an element of an ordered field Sand still obtain an element of the same field. The issue
here is not merely that of fractional powers or other algebraic operations, but a more serious
one. Imagine the (as yet undefined) real number line. Although the rational numbers are
an infinite number of points on the line, there are many “holes” between the rationals. We
already know of one “hole” at√
2 , there is another at√
3 , atπ, and ate. In fact, in a
sense which can be made precise, almost all of the points on the real number line represent
irrational numbers.
The completeness axiom is designed to eliminate the possibility of ”holes” in the real
number line. It does so by more or less bluntly stating that there are no holes. This is the
“Dedekind cut” form of the completeness axiom. we have chosen it over other equivalent
axioms because it is easy to visualize—even though the “Cauchy sequence” form is perhaps
preferable for more advanced analysis courses. A definition is needed before the axiom can
be stated.
Definition: LetS1andS2be subsets of an ordered field S. Then the set S1precedes
S2if for every a∈S1andb∈S2, we havea≤b.
If you imagine the real number line, “ S1precedesS2” should be thought of as meaning
that all of S1is to the left of all of S2.S1andS2of course might touch, or might just
miss touching.
Completeness Axiom. LetS1andS2be nonempty subsets of an ordered field S. If
S1precedesS2, then there is at least one number c∈Ssuch thatcprecedesS2and is
preceded by S1. In other words, there is (at least) one element of SbetweenS1andS2.
Definition: The set of real numbers ,R, is a set which satisfies the above axioms of algebra,
order, and completeness. Thus, the real numbers is a complete ordered field.
This type of definition of Ramounts to saying “we don’t know or care what the real
numbers are, but in any event they have the required properties.” If we had used the
10 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
Cauchy sequence version of the completeness axiom, we would have begun the rational
numbers—which we do know—and then defined the real numbers as the set of limits of
rational numbers. This would have been somewhat more concrete, but would have involved
the difficult concept of limit before we even get off the ground.
From the picture associated with the completeness axiom, we see that it exactly states
that the real number line has no holes, for - emotionally speaking—if there were a hole, let
S1be the set of real numbers to the left of the hole, and S2the s et to the right of the
hole. Then there would be no real number between S1andS2, since the hole is there,
contradicting the completeness axiom.
Let us use the idea of the last paragraph to show that the rational numbers, an ordered
field, are notcomplete by exhibiting two sets, one preceding the other, which have no
rational number between them. Just let
S1={x:x>0, x2<2}andS2={x:x>0, x2>2}.
The only possible number between S1andS2is√
2 —which is irrational. This construc-
tion is just what we need to prove the following sample.
Theorem 0.1 Every non-negative real number a∈Rhas a unique non-negative square
root.
Proof: Ifa= 0 , then 0 is the square root. If a > 0 , letS1={x:x > 0, x2< a}
andS2={x:x > 0, x2> a}. We first show that neither S1norS2is empty. Since
(1 +a
2)2= 1 +a+a2
4>a, we know that (1 +a
2)∈S2, soS2/negationslash=∅. Also (a
1+a
2)2<a(check
this) so thata
1+a
2∈S1and hence S1/negationslash=∅. BecauseS1precedesS2, by the completeness
axiom there is a c∈RbetweenS1andS2. Notice that c >0 , sincecis preceded by
S1.
It remains to show that c2=a. By the trichotomy theorem, either c2> a, c2< a,
orc2=a. The first two possibilities will be shown to give contradictions. If c2>a, since
a < (c2+a
2c)2< c2, we see thatc2+a
2c∈S2an d precedes c2, contradicting the property
specified in the completeness axiom that c2precedes every element of S2. Similarly the
assumption c2< a, with the inequality c2<(2ac
c2+a)2< a, leads to a contradiction. The
only remaining possibility is c2=a, which shows that cis the desired positive square root
ofa.
Let us now prove that the positive square root cofais unique. Assume that there
are two positive numbers c1andc2such that both c2
1=aandc2
2=a. Then
0 =c2
1−c2
2= (c1−c2)(c1+c2)
Sincec1+c2>0 , we conclude that c1−c2= 0 , soc1=c2, completing the proof of the
theorem.
Definition: The real number Mis an upper bound for the set A⊂Rif for every a∈A,
we havea≤M. The number µ⊂Ris aleast upper bound (l.u.b) forAifµis an upper
bound for Aand no smaller number is also an upper bound for A.Lower bound and
greatest lower bound (g.l.b) are defined similarly. A set A⊂Risbounded if it has both
upper and lower bounds.
Theorem 0.2 Every non-empty bounded set A⊂Rhas both a greatest lower bound and
a least upper bound.
0.5. REALS: COMPLETENESS 11
Proof: Observe first that this theorem utilizes the completeness property in that without
it, there might have been a ”hole” just where the g.l.b. and l.u.b. should be. Since the
proofs for the g.l.b. and l.u.b. are almost identical we only prove there is a g.l.b. Let
S1={x:xprecedesA},andS2=A.
By hypothesis S2/negationslash=∅. SinceAis bounded, it has a lower bound m, m ∈S1soS1/negationslash=∅.
By the completeness axiom, there is a c∈RbetweenS1andS2. It should be obvious
thatcis both greater than or equal to every element of S1, and less than or equal to every
element of S2- so it is the required g.l.b.
Definition: Theclosed interval [a,b] is the set {x∈R:/Gmir≤/archleftdown≤ }.
Theopen interval (a,b) is the set {x∈R:/Gmir</archleftdown<}. All we can do is apologize for
the multiple use of the parentheses in notation. Please note that sets are not like doors.
Some sets, like ( a,b) ={x∈R:/Gmir≤/archleftdown<}are neither open nor closed.
Theorem 0.3 (Nested set property). Let I1,I2,... be a sequence of non-empty closed
bounded intervals, In={x:an≤x≤bn}, which are nested in the sense I1⊃I2⊃I3...,
so each covers all that follow it. Then there is at least one point c∈Rwhich lies in all of
the intervals, that is, cis in their intersection c∈ ∩∞
k=1Ik.
Proof: LetS1={x:xprecedes some In,and so allIk, k≥n}
S2={x:xpreceded by some In,and so allIk, k≥n}.
First, neither S1norS2are empty since a1∈S1andb1∈S2. Thus by the completeness
axiom, there is at least one c∈RbetweenS1andS2. Thiscis the required number
(complete the reasoning).
If the intervals Ikdo not get smaller after, say INbecauseaN=aN+1= . . . and
bN=bN+1= . . . , then the whole interval aN≤x≤bNis caught by the preceding
argument. The more common case is there the ak’s strictly increase and the bk’s strictly
decrease. This is what happens when approximating a real number to successively greater
accuracy by the decimal expansion. In the case of√
2 for example,
I1={x: 1≤x≤2},
I2={x: 1.4≤x≤1.5},
I3={x: 1.41≤x≤1.42},
I4={x: 1.414≤x≤1.415},
and so on, gradually squeezing down on√
2 to any desired accuracy.
Definition: The sequence an∈R,/multicloseleft=/notforces,/notsatisfies,.... of real numbers converges to the real
numbercif, given any /epsilon1>0 , there is an integer Nsuch that |an−c|</epsilon1for alln>N .
We will then write an→c. [In practice no confusion arises for the use of →to denote
both convergence and mappings (cf. 1)].
Again ordinary decimals supply an example, for they allow us to get arbitrarily close
to any real number. We could have defined the real numbers as all decimals; however there
would be a mess avoiding the built-in ambiguity illustrated by 1 .9999....= 2.0000....
Theorem 0.4 Under the hypotheses of the previous theorem, if in addition the length of
Intends to zero, (bn−an)→0, then the number c∈Rfound is unique. Furthermore, if
uk∈Ikfor allk, that is if ak≤uk≤bk, thenuk→ctoo.
Proof: Suppose there were two real numbers cand ˜cin all of the intervals,
ak≤c≤bkandak≤˜c≤bkfor allk.
12 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
Rewriting the second inequality as −bk≤ −˜c≤ −ak, and adding this to the first inequality,
we find that ak−bk≤c−˜c≤bk−ak. Since both sides of this inequality tend to zero, if
c−˜c/negationslash= 0 , we would have a contradiction.
To proveuk→c, repeat the above reasoning with ˜ creplaced by uk. We find that
ak−bk≤c−uk≤bk−ak. Again both sides of this inequality tend to zero. Now let us
fiddle with the /epsilon1,N definition of limit t o complete the proof. Since bn−an→0 , given
any/epsilon1>0 , there is an Nsuch that |an−bn|</epsilon1for alln>N . Thus for any /epsilon1>0 and
the sameN,|un−c|</epsilon1forn>N , which is the definition of un→c.
Theorem 0.5 Bolzano-Weierstrass . Every infinite sequence of real numbers {uk}in
a bounded interval Ihas at least one subsequence which converges to a number c∈R.
Proof: This one is very clever and picturesque. Watch. Bisect Iinto two intervals I1
and ˜I1of equal length. At least one of I1or˜I1must contain an infinite number of the
{uk}’s. Continuing in this way we obtain a set of nested intervals I⊃I1⊃I2⊃. . . each
of which have an infinite number of the {uk}’s, and the length of Intending to zero.
From Theorem 3 we conclude that there must be a c∈Rcommon to all of the intervals.
We must now select the subsequence {ukn}of the {uk}’s which converge to c. Since
eachIncontains an infinite number of points of the sequence, we can certainly pick one,
sayukn∈In. This sequence {ukn}satisfies the hypotheses of Theorem 4. Thus ukn→c.
Remarks: 1. If we also assume Iis closed, then we can further assert that c∈I. IfIis
not closed, cmay be an end point /ownerI.
2. If a sequence ukconverges to a c∈R, then every infinite subsequence uknalso
converges, and to the same number c.
Theorem 0.6 . If the sequence {uk}converges, it is bounded.
Proof: Sayuk→α, and let/epsilon1= 1 in the definition of convergence. Then there is an N
such that |un−α|<1 for alln>N . Thus, when n>N ,
|un|=|un−α+α| ≤ |un−α|+|α|<1 +|a|
Therefore for any kthe number |uk|is bounded by the largest of the N+1 numbers |u1|,
|u2|, . . . , |uN|and (1 + |a|) .
The following theorem shows how to handle algebraic combinations of convergent se-
quences.
Theorem 0.7 Ifan→αandbn→β, then
i)an+bn→α+β
ii)anbn→αβ
iii)an
bn→α
βif bothbn/negationslash= 0, for alln, and ifβ/negationslash= 0.
Proof: Since the proofs are all similar, we only prove ii). Observe that
|anbn−αβ|=|(anbn−αbn) + (αbn−αβ)| ≤ |an−α||bn|+|alpha||bn−β|
By Theorem 6, the |bn|’s are bounded, say by B. Sincean→α, given any ε>0 , there
is anN1such that |an−α|<ε
2Bfor alln>N 1, and since bn→β, for the same εthere
0.5. REALS: COMPLETENESS 13
is anN2such that |bn−β|<ε
2|α|for alln>N 2. Thus, ifnis greater than the larger of
N1andN2, n> max(N1,N2) , we find that
|anbn−αβ|<ε
2+ε
2=ε
which does the job.
Definition: The sequence a1, a2,...of real numbers is said to be monotone increasing if
a1≤a2≤a3≤..., and monotone decreasing a1≥a2≥a3≥.... Both kinds are called
monotone sequences.
Theorem 0.8 Every bounded monotone sequence a1, a2,... of real numbers converges. In
other words, there is an αεRsuch thatan→α.
Proof: We assume the sequence is increasing. The proof for decreasing sequence is iden-
tical. Since the sequence is bounded, by Theorem 2 it has a least upper bound αεR. We
maintainan→α. Given any ε>0 , we know that for all n, a n<a+εbecauseαis an
upper bound. Since α−ε<α , andαis the l.u.b. of the sequence, we can find an Nsuch
thatα−ε<a N. But then, because the sequence is increasing α−ε<a nforalln≥N.
Thus for all n≥N, a−ε<a n<a+ε; that is, |an−α|<εfor alln≥N, proving the
convergence to α.
We shall close this difficult section with a wonderful procedure for computing the square
root of a positive real number. I use it all of the time. It is much easier to understand than
the hair-raising method taught in public school.
Theorem 0.9 For any positive real numbers Aanda0the infinite sequence defined by
an+1=1
2(an+A
an), n= 0,1,2,..., (0-2)
is monotone decreasing and converges to√
A. Moreover, if we let bn=A
an, then thebn’s
are monotone increasing and also converge to√
A:
ba≤b2≤...≤√
A≤...≤a2≤a1
Proof: We first show that a2
k≥Aand thatak+1≤ak,
a2
k−A=1
4(ak−1+A
ak−1)2−A=1
4(ak−1+A
ak−1)2≥0,soa2
k≥A.
From this, it is easy to see that ak+1≤ak, for
ak−ak+1=ak−1
2(ak+A
ak) =a2
k−A
2ak≥0.
Thusa1≥a2≥ ···.≥√
A.
That thea2
kconverge is an immediate consequence of Theorem 8, since the sequence
{a2
k}is a bounded (by A) monotone decreasing sequence. Denoting the limit by α, a2
k→α,
the proof that α=Ais identical to th e reasoning which gave a unique limit in Theorem
4.
Sincebn=A
an, and thean’s decrease and are ≥√
A, then thebn’s increase and are
≤√
A. This also shows that bn≤an. Sincean→√
A, we havebn=A
an→√
Atoo.
14 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
Application: We compute√
8 . Takea0= 3 . Then a1=1
2(3 +8
3) =17
6, and
b1= 8·6
17=48
17. Similarly, a2=577
204, b2=1632
577. This gives1632
577≤√
8≤577
204, or in decimal
form
2.82842<√
8<2.82843,
astounding accuracy after only two steps. I carried the computations one step further and
found
2.828427124....≤√
8≤2.828427124....,
where the dots indicate I gave up on the arithmetic, having obtained the exact value as far
as the approximation went. Digital computers use this method and related ones for similar
computations. It is particularly well adapted to them (and me) since only simple arithmetic
operations are involved.
This Theorem 9 gives another proof that every positive real number has a unique
positive square root. It is valuable to compare this proof with that of Theorem 1. The
main distinction is that the second proof just given is constructive it actually shows a
way to compute successive approximations to the square root of any number. However,
you are justified in asking how we ever found the procedure of equation (0.9) in the first
place. The secret is that this formula is a statement of Newton’s method for finding roots
off(x) = 0 , applied to the particular function f(x) =x2−A. See most calculus books
for more information about this method. Hopefully, we will have time to discuss this topic
later, for it is a constructive way o f proving the existence of a sought after object. The
standard existence theorem for ordinary differential equations is a close relative of Newton’s
method.
Exercises
(1) For the sequences defined below, find which converge, which do not converge but
do have at least one convergent subsequence, and which have neither. In all cases
n∈Z+.
(a)an=1
n+ 1
(b)an=(−1)n
n
(c)bn=en
(d)an=e−2n+1
(e)an= 1 +n
(f)an= 2 + ( −1)n
(g)an=√n+ 1−√n
(h)an=2−3n
5n+1
(i)an=7n
n!(tough, isn’t it?)
(j)sn= 1 +1
2+1
4+1
8+···+1
2n.
(2) Prove that if an→αandbn→β, then (an+bn)→α+β, where all the letters
represent real numbers.
0.6. APPENDIX: CONTINUOUS FUNCTIONS AND THE MEAN VALUE THEOREM 15
(3) a). Prove Bernoulli’s inequality
(1 +h)n>1 +nh, h /negationslash= 0, h>−1, n≥2.
Hereh∈Randn∈Z. I suggest proof by induction.
b). Ifs∈R, use part a) to prove that
an≡sn→/braceleftbigg0 if |s|<1
∞if|s|>1.
[Hint: If |s|<1 , write |s|=1
1+h, h> 0 , while if |s|>1 , write |s|= 1 +h, h> 0] .
0.6 Appendix: Continuous Functions and the Mean Value
Theorem
Definition: : The function f(x) iscontinuous at the point x0if, given any /epsilon1>0 , there
is aδ(/epsilon1)>0 such that
|f(x)−f(x0)|<εwhen 0<|x−x0|<δ(/epsilon1).
Remark: This may be rephrased as
lim
x→x0f(x) =f(x0).
Note that either statement requires
(1)fbe defined at x0.
(2) limf(x) exists.
x→x0
x/negationslash=x0
(3) the limiting value of fatx0is equal to the defined value of fatx0.
If a function is discontinuous at x0, it has at least one of the four troubles
(1) Jump discontinuity
(2) Infinite discontinuity
(3) Infinite oscillations
(4) Removable discontinuity.
Here are examples of each trouble at the point x= 0 .
(1)f(x) =/braceleftbigg1,0≤x
−1, x < 0
(2)f(x) =/braceleftbigg1
xx/negationslash= 0
anything, say 1, x = 0
16 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
(3)f(x) =/braceleftbiggsin1
xx/negationslash= 0
anything, say 0, x = 0
(4)f(x) =/braceleftbiggx, x /negationslash= 0
1, x= 0
Note that a function may oscillate infinitely about a point and still be continuous there.
This is illustrated by the everywhere continuous function
f(x) =/braceleftbiggxsin1
x, x/negationslash= 0
0, x= 0
Theorem 0.10 I. Iff(x)is continuous at x=c, andf(c) =A/negationslash= 0, thenf(x)will
keep the same sign as f(c)in a suitably small neighborhood of x=c.
Proof: : We construct the desired neighborhood. Assume Ais positive. The proof if
A<0 is essentially the same. In the definition of continuity, take /epsilon1=A. Then there is a
δ>0 such that
|f(x)−A|<A when |x−c|<δ,
that is,
0<f(x)<2A,when |x−c|<δ.
In other words, f(x) is positive in the interval |x−c|<δ.
Theorem 0.11 II. Iff(x)is continuous at every point of a closed and bounded inter-
val, then there is a constant Msuch that |f(x)| ≤Mthroughout the interval. Thus a
continuous function in a closed and bounded interval is bounded.
Proof: : By contradiction. If fis not bounded, there is a sequence of points xnsuch
that|f(xn)|>n. From that sequence by Theorem 5 (Bolzano-Weierstrass) we can select
a subsequence xnkwhich converges to some point x0in t he interval, xnk→x0. Thus
|f(xnk)| → ∞.
But we know from the continuity of fthat|f(xnk)| → |f(x0)|. A contradiction.
Moreover, with the same hypotheses , we can conclude more.
Theorem 0.12 III. Iffis continuous at every point of a closed and bounded interval,
then there are points x=αandx=βin the interval where fassumes its greatest and
least values, respectively.
Proof: : We show that fassumes its greatest value. The proof for the least value is
essentially identical. Let Sbe the set of all upper bounds for f. By Theorem II Sis not
empty. Therefore by Theorem 2, Shas a g.l.b., call it M0. SinceM0is the greatest lower
bound of upper bounds for f, there is a sequence xnsuch that lim n→∞f(xn)→M0. Use
Bolzano-Weierstrass to pick a subsequence xnkof thexnsuch that the xnkconverges, say
toc. By continuity of f,limnk→∞f(xnk) =f(x) . Thusf(c) =M0, sofdoes assume its
greatest value at x=c.
Remark: This theorem refers to the absolute maximum andabsolute minimum values.
0.6. APPENDIX: CONTINUOUS FUNCTIONS AND THE MEAN VALUE THEOREM 17
Examples: The following show that the theorem is not necessarily true if any of the
hypotheses are omitted.
(1)f(x) =x,0<x≤1 . No min. (interval not closed).
(2)f(x) =x, x≤0 , andf(x) =1
1+x2,allx, both have no min. (the interval is
unbounded.)
(3)f(x) =/braceleftbiggx, 0≤x<3.No max. (function is discontinuous.)
x−2,3≤x≤4
Theorem 0.13 Iff(x)is continuous at every point of a closed and bounded interval [a,b],
and iff(a)andf(b)have opposite sign, then there is at least one point c∈(a,b)such
thatf(c) = 0 .
Proof: : Sayf(a)<0, f(b)>0 . We find one point c, “the largest xsuch that
f(x) = 0 ”. Let S={x∈[a,b]:f(x)≤0}.
Sincef(a)<0, Sis not empty. It thus has a l.u.b., c. We prove that f(c) = 0 .
Eitherf(c)>0, f(c)<0 , orf(c) = 0 . The first two possibilities cannot happen, since
by Theorem I, if they did, fwould be positive (or negative) in a whole neighborhood of
c-violating the fact that cis the l.u.b. of S.
Corollary 0.14 (intermediate value theorem ). Letf(x)be continuous at every point
of a closed and bounded interval [a,b], withf(a) =A,andf(b) =B. Then ifCis any
number between AandB, there is at least one point c, a≤c<b , such that f(c) =C.
Thus,fassumes every value between AandBat least once.
Proof: : Apply Theorem IV to the function ϕ(x) =C−f(x) .
Remark: The function may assume values other than just those between AandB. An
example is the function f(x) =x2,−1≤x≤3 . The theorem requires that it assume all
values between f(−1) = 1 and f(3) = 9 . Besides those values , this function also happens
to assume all values between 0 and 1.
We can offer another proof of
Corollary 0.15 Every positive number khas a unique positive square root.
Proof: : Consider f(x) =x2−k, which is clearly continuous everywhere. Since f(0)<0 ,
andf(1+k
2) = (1+k
2)2−k= 1+k2
4>0 , Theorem IV shows that fmust vanish somewhere
in the interval 0 < x < 1 +k
2. This is the root. It is the unique positive square root, for
say there were two positive numbers xandysuch thatx2−k= 0 andy2−k= 0 . then
x2−y2= 0 . Thus, 0 = x2−y2= (x−1)(x+y) . Sincex+y>0 , we conclude x−y= 0 ,
orx=y.
Remark : It appears that if a function has the property of Corollary 1, the intermediate
value property, then it must be continuous. This is false . An example is given by the
discontinuous (trouble 3) function
f(x) =/braceleftbiggsin1
x, x/negationslash= 0
0, x= 0
18 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
about the point x= 0 . Ifais any number <0 , andbany number >0 , thenf(x)
assumes every value between f(a) andf(b) , butf(x) is not continuous throughout the
interval since it is not continuous at x= 0 .
Definition: The function f(x) has a relative maximum (minimum) at the point x0, if,
for allxin a sufficiently small interval containing x0as an interior point, we have
f(x)≤f(x0) (f(x)≥f(x0)).
Remark: By convention, we shall agree notto call the possible max (or min) at the end
point of an interval a relative max (or min). This does lead to the possibility of an absolute
max (or min) not being a relative max (or min). However, if the absolute max (or min)
does occur at an interior point of an interval, it is also a relative max (or min).
Definition: The function f(x) isdifferentiable at the point x0if the following limit
lim
x→x0f(x)−f(x0)
x−x0
exists. There are the usual notations: f/prime(x0) ,d f
dx/vextendsingle/vextendsingle/vextendsingle
x=x0,Df(x0) .
Theorem 0.16 Iff(x)is differentiable at x0, then it is continuous there.
Proof: : Now if the limit
lim
x→x0f(x)−f(x0)
x−x0
exists, as we have assumed, then the numerator must approach zero as xtends tox0.
Thusfis continuous at x0.
Theorem 0.17 Iff(x)is differentiable at x0and has a relative maximum or minimum
atx0, thenf/prime(x0) = 0 .
Proof: : Assumefhas a relative min at x0. Then for all xnearx0, f(x)≥f(x0) .
(i) ifx<x 0f(x)−f(x0)
x−x0≤0
so lim x→x0x<x 0f(x)−f(x0)
x−x0≤0
(ii) ifx>x 0f(x)−f(x0)
x−x0≥0
so lim x→x0x>x 0f(x)−f(x0)
x−x0≥0
Because the function is differentiable at x0, the two limiting values are f/prime(x0) . Thus
f/prime(x))≤0 andf/prime(x0)≥0 . Both statements can be true only if f/prime(x0) = 0 . The trick here
was, the slope must be negative to the left, and positive to the right of x0. Since there is
a unique slope (the derivative) at x0, the slope must be zero there. At a relative max., the
same proof holds with obvious modifications.
Examples: 1. Although the function f(x) =|x|has a relative minimum at x= 0 , the
conclusion of the theorem does not hold since fis not differentiable there. Note that both
(i) and (ii) of the proof still do hold.
2. The differentiable function (for all x)
f(x) =/braceleftbiggx4sin1
x, x/negationslash= 0
0, x= 0
has an infinite number of relative max and min in any interval including the origin.
0.6. APPENDIX: CONTINUOUS FUNCTIONS AND THE MEAN VALUE THEOREM 19
Theorem 0.18 (Rolle ). If
(i)f(x)is continuous at every point of the closed and bounded interval [a,b]
(ii)f(x)is differentiable at every point of the open interval (a,b)and
(iii)f(a) =f(b),
then there is at least one point c,a<c<b , wheref/prime(c) = 0 .
Proof: : Iff(x)≡constant throughout [ a,b] , takecto be any point in ( a,b) . Otherwise
f(x) must go either above or below (or both) the value f(a) . Assume it goes above. Then
by Theorem III there is a point x=cwherefhas its absolute maximum. Since we
assumedf(x) goes above f(a) , the point x=cis an interior point. Thus there is
a relative maximum. Since fis differentiable in ( a,b) , we may apply Theorem VI to
conclude that f/prime(c) = 0 . If we had assumed fwent below f(a) , then there would have
been an absolute (and relative) min. etc.
Remarks: 1. From the proof of the theorem, we see that if fhas values both greater and
less thanf(a) , then there would be at least two points in ( a,b) wheref/prime= 0 .
2. You should be able to construct examples showing the theorem is not true if any of
the hypotheses are dropped.
Corollary 0.19 (mean value theorem ) If
(i)f(x)is continuous at every point of the closed and bounded interval [a,b]and
(ii)f(x)is differentiable at every point of the open interval (a,b), then there is at
least one point cin(a,b)where
f/prime(c) =f(b)−f(a)
b−a.
Proof: : “Shift and apply Rolle’s Theorem”. In more detail, consider
F(x) =f(x)−f(a)−x−a
b−a(f(b)−f(a)).
F(x) satisfies all of the assumption of Rolle’s Theorem. Therefore there is a point cwhere
F/prime(c) = 0 . Since
F/prime(x) =f/prime(x)−f(b)−f(a)
b−a,
atx=c, we have
f/prime(c) =f(b)−f(a)
b−a.
Remarks: 1. The function f(x) =|x|in the interval [ a,b], a < 0, b > 0 , shows what
happens if the function fails to be differentiable at even one point of the open interval ( a,b) .
2. An alternative form of the conclusion is: there is a number θ,0<θ< 1 , such that
f(b)−f(a) =f/prime(a+θ(b−a))(b−a).
This is because every point in the interval ( a,b) is of the form a+θ(b−a) , for some
θ,0<θ< 1 .
We shall now give some applications of the Mean Value Theorem. The first one is a
specific example, while the others have great significance in themselves.
20 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
Example: The function f(x) =a1sinx+a2sin 2x+bcosx+b2cos 2xhas at least one
zero in the interval [0 ,2π] , no matter what the coefficients a1,a2,b1andb2are. To show
this, we shall show fis the derivative of a function g(x) which satisfies the hypotheses of
Rolle’s theorem. This function gis just an anti-derivative of f:g/prime(x) =f(x)
g(x) =−acosx−a2
2cos 2x+b1sinx+b2
2sin 2x.
Sincegis clearly continuous and differentiable everywhere, we must only see if g(0) =
g(2π) , which as also easy.
Theorem 0.20 Iff(x)is continuous and differentiable throughout [a,b], and |f/prime|< N
there too, then the δ(/epsilon1)in the definition of continuity can be chosen as δ(c) =/epsilon1
N. Thisδ
works for everyxin[a,b].
Proof: : Use the form of the mean value theorem in Remark 2. Then for any points x, x 0
in (a,b) ,
f(x)−f(x0) =f/prime(˜x)(x−x0),
where ˜xis somewhere between xandx0. Thus
|f(x)−f(x0)| ≤N|x−x0|.
We see now that if δ(/epsilon1) =/epsilon1
N, then for any /epsilon1>0 ,
|f(x)−f(x0)|</epsilon1if|x−x0|<δ.
Theorem 0.21 Iffsatisfies the hypotheses of the mean value theorem and if in addition
f/prime(x)≡0throughout (a,b), thenf(x)≡const.
Proof: : Letx1andx2be any points on ( a,b) . Then by the form of the mean value
theorem in Remark 2
f(x2)−f(x1) = 0·(x2−x1) = 0.
Thusf(x2) =f(x1) for any two points in ( a,b) , that is,fis identically constant.
Corollary 0.22 Iff(x)andg(x)both satisfy the hypotheses of the mean value theorem,
and if in addition f/prime(x)≡g/prime(x)for allxin(a,b), thenf(x) =g(x)+c, wherecis some
constant.
Proof: : consider the function F(x) =f(x)−g(x) . It satisfies the hypothesis of Theorem
VII, soF(x)≡c, cconstant. Thus f(x)−g(x) =c.
Remark: Theorem IX is the converse of the theorem: “the derivative of a constant function
is zero.”
a figure goes here
0.6. APPENDIX: CONTINUOUS FUNCTIONS AND THE MEAN VALUE THEOREM 21
Exercises
(1) Look over all the theorems (and corollaries) here and be sure you can find examples
showing that the theorems are not true if any of the hypotheses are relaxed.
(2) Letf(x) =/braceleftbigg1,ifxis a rational number
0,ifxis an irrational number.
Isfcontinuous anywhere?
(3) Letf(x) be an everywhere differentiable function which is zero at x=aj, j=
1,2,...,n. Find a function which vanishes at least once between each of the zeros of
f.
(4) Use Theorem VIII to find a δ(/epsilon1) for the given functions.
(a)f(x) =x4−7,−2≤x≤3.
(b)f(x) =x2sinx,−4≤x≤3
(c)f(x) =1
1+x2,−2≤x≤1
(d)f(x) =x4
3+ 7,−2≤x≤8
(e)f(x) =x√
x2+ 1,−2≤x≤2
(5) (a) The function f(x) satisfies the following condition
|f(x)−f(x0)| ≤2|x−x0|3
for every pair of points x, x 0in the interval [ a,b] . Provef(x)≡constant in this
interval.
(b) Generalize your proof to the case when fsatisfies
|f(x)−f(x0)| ≤c|x−x0|α,
wherec>0 is some constant and αis any number >1 .
(6) Consider the function f(x) =x2
3, in the interval [ −8,8] .
Sketch a graph. Note that f(−8) =f(8) = 4 but there is no point where f/prime= 0 ;
which hypothesis of Rolle’s theorem is violated?
(7) In a trip, the average speed of a car is 180 miles per hour. Prove that at some time
during the trip, the speedometer must have registered precisely 180 miles per hour.
(8) LetP1:= (x1,y1) andP2:= (x2,y2) be any two points on the parabola y=
ax2+bx+c, and letP3:= (x3,y3) be the point on the arc P1P2where the tangent
is parallel to the chord P1P2. Show that
x3=x1+x2
2.
(9) Prove that every polynomial of odddegree
P(x) =x2n+1+a2nx2n+···+a1x+a0
has at least one real root.
(10) Iffis a nice function and f/prime<0 everywhere, prove that fis strictly decreasing.
22 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
0.7 Complex Numbers: Algebraic Properties
.
In high school, to be able to find the roots of all quadratic equations ax2+2bx+c= 0 ,
we were forced to introduce the symbol i≡√−1 , in other words, introduce a special symbol
for a root of x2+ 1 = 0 . Before going any further, we should prove that no real number c
can satisfy c2+ 1 = 0 . By contradiction, assume that there is such a c. Then necessarily
eitherc>0, c< 0,orc= 0 . Ifc= 0 , we have the immediate contradiction that 1 = 0 . If
c>0 , orc<0,0<c2. Consequently 0 <c2+ 1 too, which again contradicts 0 = c2+ 1 ,
and proves our contention that no real number can satisfy x2+ 1 = 0 .
Observe that our proof also shows that if we introduce a new symbol for a root of
x2+ 1 = 0 , that symbol cannot be an element of an ordered field, for only the ordered field
properties of the real numbers were used in the above proof. we shall see that “ i” is an
element of a field, but not an ordered field.
It is difficult to overestimate the importance of complex numbers for all of mathematics,
both from an esthetic as well as from a practical viewpoint. With them we can prove that
every quadratic polynomial has exactly two roots (which may coincide). What is more
surprising is that every polynomial of order n
anxn+an−1xn−1+···+a1x+a0= 0, an/negationslash= 0,
has exactly ncomplex roots. This result, the fundamental theorem of algebra , was first
proved by Gauss in his doctoral dissertation (1799). It is one of the crown jewels of math-
ematics. The difficult part is proving that every polynomial has at least one complex root,
from which the general result follows using only the “factor theorem” of high school alge-
bra. Later on in the semester we shall discuss this more fully and offer a proof. It is not
simpleminded, for the proof is non-constructive pure existence proof, giving absolutely no
method of finding the roots. Perhaps we shall even prove some more exotic results.
Having gotten carried away, let us retreat and obtain the algebraic rules governing the
setCof complex numbers. In order to reveal the algebraic structure most clearly, we shall
denote a complex number zby an ordered pair of real numbers: z= (x,y), x, y ∈R.
Thus CisR×Rwith the following additional algebraic structure.
Definition: Ifz1= (x1,y1) andz2= (x2,y2) are any two complex numbers, then we
define
Addition: z1+z2= (x1+x2, y1+y2) ,
and
Multiplication: z1·z2= (x1x2−y1y2, x1y2+y1x2) .
Equality:z1=z2if and only if both x1=x2andy1=y2.
Thus, the complex number zero—the additive identity—is (0 ,0) , while the complex
number one—the multiplicative identity—is (1 ,0) . Using the fact that the real numbers
Rform a field, we can now prove the
Theorem 0.23 The complex numbers Cform a field.
Proof: Since the verification of the field axioms are entirely straightforward we give only
a smattering. Note that we shall rely heavily on the field properties of R. Addition is
commutative:
z1+z2= (x1,y1) + (x2,y2) = (x1+x2, y1+y2)
= (x2+x1, y2+y1) = (x2,y2) + (x1,y1) =z2+z1.(0-3)
0.7. COMPLEX NUMBERS: ALGEBRAIC PROPERTIES 23
Additive identity:
0 +z= (0,0) + (x,y) = (0 +x,0 +y) = (x,y) =z.
Multiplicative inverse: For any z∈C, z/negationslash= (0,0) , we must find a ˆ z= (ˆx,ˆy)∈Csuch
thatzˆz= 1 , that is, find real numbers ˆ xand ˆysuch that ( x,y)(ˆx,ˆy) = (1,0) . Using
the definition o f complex multiplication, this means we must solve the two linear algebraic
equations
xˆx−yˆy= 1
yˆx+xˆy= 0/bracerightbigg
x, y∈R,
for ˆxand ˆy∈R. The result is
ˆz= (ˆx,ˆy) = (x
x2+y2,−y
x2+y2).
We will denote this multiplicative inverse, which we have just proved does exist, by1
zor
z−1.
It is interesting to notice that complex numbers of the form ( x,0) have the same
arithmetic definitions as the real numbers, viz.
(x1,0) + (x2,0) = (x1+x2,0)
(x1,0)(x2,0) = (x1x2,0).
We can easily verify that all complex numbers of this form ( x,0) also form a field,
asubfield of the field C. On the basis of these last two equations, we can identify a real
numberxwith the complex number ( x,0) in the sense th at if we perform any computation
with these complex numbers of this form, the result will be the same as if the computation
had been performed with the real numbers alone. Thus, numbers of the form ( x,0)∈Care
algebraically equivalent to the numbers x∈R. The technical term for such an algebraic
equivalence is isomorphic , much as a term for geometric equivalence is congruent. After
identifying the real numbers with complex numbers of the form ( x,0) , we can say that the
field of real numbers Risembedded as a subfield in the field of complex numbers, R⊂C.
After all this chatter, let us at least convince ourselves that every quadratic equation
is solvable if we use complex numbers. First we solve z2+ 1 = 0 , which may be written
as (x,y)(x,y) + (1,0) = (0,0) , or as the two real equations x2−y2=−1,2xy= 0 .
The last equation says that either x= 0 ory= 0 . Now if y= 0 , we are left to solve
x2+ 1 = 0, x∈R, which we know is impossible. Therefore x= 0 and then y2= 1 . Thus
the two complex numbers (0 ,1) and (0,−1) both satisfy z2+ 1 = 0 . The general case,
az2+bz+c= 0 is easily reduced to the special one by completing the square.
One by-product of the above demonstration is that we see it is foolhardy to try to define
an order relation on Cto obtain an ordered field. This is because the equation x2+ 1 = 0
cannot be solved in any ordered field, as was shown earlier, whereas we have just solved it
inC.
Observe that every ( x,y)∈Ccan be written as
(x,y) = (x,0)(1,0) + (y,0)(0,1),
where the complex number (0 ,1) is called the imaginary unit and is denoted by i. If we
now utilize the isomorphism between the real number aand complex numbers ( a,0) , the
24 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
last equation shows that ( x,y) may be thought of as x+iy. Thus, we have obtained the
usual notation for complex numbers. From our development, the algebraic role of ias the
symbol for the imaginary unit (0 ,1) is hopefully clarified. The number xis called the real
part, andytheimaginary part of the complex number z=x+iy. In symbols, x=Re{z}
andy=Im{z}.
Our introduction of complex numbers suggests a geometric interpretation. We have
defined complex numbers Cas ordered pairs of real numbers, elements of R×R, with an
additional algebraic structure. Since the points in the plane are also elements of R×R, it
is clear that there is a one to one correspondence between the complex numbers and the
points in the plane. If we plot the point z= (x,y) , the real number |z|, the “ absolute
value ormodulus ofz” is the distance of the point zfrom the origin. Its value is computed
by the Pythagorean theorem
|z|=/radicalbig
x2+y2.
Here are several formulas which are easily verified:
|z1z2|=|z1|+|z2|
|x| ≤ |z|,|y| ≤ |z|
|z1+z2| ≤ |z1|+|z2|(triangle inequality)
(0-4)
If the line joining the point zto the origin is drawn, the angle θbetween that line
and the positive real (= x) axis is called the argument oramplitude orz. The absolute
valuerand argument θof a complex number determine it uniquely, since we have
z=r(cosθ+isinθ) (0-5)
This is the polar coordinate form of the complex number z. Note that conversely, z
determines its argument only to within an additive multiple of 2 π. This observation will
prove of value to us shortly.
Associated with every complex number, z=x+iythere is another complex number
z=x−iy, the complex conjugate ofz. It is the reflection of zin the real axis. Probably
the main reason for introducing zis that we can solve for xandyin terms of zandz:
x=z+z
2, y =z−z
2i.
Again some simple formulas:
|¯z|=|z|,|z|2=|¯z|2=zz.
(z1+z2) =z1+z2,(z1z2) =z1z2./bracerightbigg
(0-6)
To illustrate the value of this notation, let us leave the main road to prove the interesting
Theorem 0.24 . If the complex number γis a root of the polynomial
P(t) =antn+an−1tn−1+···+a1t+a0,
where the coefficients a0,a1,...,a narerealnumbers, then γis also a root of P(t). In
other words, the roots of realequations occur in conjugate pairs.
0.7. COMPLEX NUMBERS: ALGEBRAIC PROPERTIES 25
Proof: Sinceγis a root, the complex number
P(γ) =anγn+···+a1γ+a0
is zero,P(γ) = 0 . This implies that its conjugate is also 0, P(γ) = 0 . By using equations
(0.7), we have that
P(γ) =anγn+···+a1γ+a0,
since the coefficients ajare real,aj=aj. Thus
0 =P(γ) =anγn+···a1γ+a0=P(γ),
that is, the complex number γis a root of the same polynomial.
Now if the proof looks like it was done with mirrors, go over each step carefully. This
type of reasoning is somewhat typical of modern mathematics in that it yields information
about an object (the roots of a polynomial in this case) without first obtaining an explicit
formula for the object.
After this digression let us return and find a geometric interpretation for the arithmetic
operations on complex numbers. First, addition. The three points z1,z2andz1+z2
together with the origin determine a parallelogram. (check this). Thus addition of complex
numbers is sometimes called the parallelogram rule for additions. Given the points z1and
z2, the point z1+z2can be constructed using compass and straight-edge. Subtraction is
justz1+ (−z2) .
Multiplication is much more difficult to interpret geometrically. We shall use equation
(0.7) and write zj=|zj|(cosθj+isinθj),j= 1,2.Then
z1z2=|z1|(cosθ1+isinθ1)|z2|(cosθ2+isinθ2)
z1z2=|z1z2|[cosθ1+θ2) +isin(θ1+θ2)]. (0-7)
Thus the product of z1andz2has modulus |z1z2|and argument θ1+θ2: multiply the
moduli and add the arguments. This too may be carried out using compass and straight-
edge. Since1
z2=1
|z2|(cosθ2−isinθ2) , division reads
z1
z2=/vextendsingle/vextendsingle/vextendsingle/vextendsinglez1
z2/vextendsingle/vextendsingle/vextendsingle/vextendsingle[cos(θ1−θ2) +isin(θ1−θ2)],
so the moduli are divided while the arguments are subtracted.
We will exploit the multiplication formula (0.7) to find all ncomplex roots of the
specific polynomial
zn=A,
for anyA∈C. This equation is one of the few whose roots can always be found explicitly.
The trick is to write Ain its polar coordinate form
A=|A|[cos(α+ 2kπ) +isin(α+ 2kπ)],
whereαis the argument of Aandkis any integer. Although we get the same Ano matter
whatkis used, as was observed following equation (0.7), we shall retain the arbitrary k
since it is the heart of the process we have in mind. From equation (0.7) we see that
A1
n=|A|1
n[cosα+ 2kπ
n+isinα+ 2kπ
n]
26 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
in the sense that for any value of the integer k,(A1
n)n=A. Askruns through the integers,
we get only ndifferent angles of the formα+2kπ
n, since the other angles differ from these
nangles by multiples of 2 π. For each of these ndifferent angles we obtain a different
complex number A1
n. Thesennumbers for A1
nare the desired nroots ofzn=A. It
is usually convenient to obtain the angles by letting k= 0,1,2,...,n −1 , although any n
integers which do not differ by multiples of nwill do.
An example should help clear the air. We shall find the three cube roots of −2 , that
is, solvez3=−2 . First,
−2 = 2[cos(π+ 2kπ) +isin(π+ 2kπ)],
since the argument of −2 isπwhile its modulus is 2. Thus, the roots are
z= 21
3[cosπ+ 2kπ
3+isinπ+ 2kπ
3], k= 0,t1,t2...
There are only three values of zpossible, no matter what k’s are used. These three
cube roots of −2 are
k= 0,3,6,...z 1= 21
3[cos(π
3) +isin(π
3)] = 21
3(1
2+i√
3
2)k= 1,4,7,...z 2= 21
3[cos(π) +
isin(π)] =−21
3
k= 2,5,8,...z 3= 21
3[cos(5π
3) +isin(5π
3)] = 21
3(1
2+i√
3
2).
It is time-saving to observe that the nroots of unity, that is, of zn= 1 , can be written
down immediately by utilizing the geometric interpretation of multiplication. All of the
roots have modulus 1, and so must lie on the unit circle |z|= 1 . Bisecting the circle into
nequal sectors by the radii, the first beginning on the positive x-axis, we find the roots of
unity,wj, at thensuccessive intersections of these radii with the unit circle. The roots
wj, j= 1,2,3,ofz3= 1 are illustrated in the figure as the intersections of θ= 0, θ=2π
3,
andθ=4π
3with |z|= 1 . Thus w1= cos 0 +isin 0 = 1, w2= cos2π
3+isin2π
3=
−1
2+i√
3
2,w3= cos4π
3+isin4π
3=−1
2−i√
3
2.
Exercises
(1) Express the following complex numbers in the form a+bi.
(a) (1 −i)2
(b) (2 +i)(3−i)
(c)1
i
(d)1+i
2−i
(e)1+i
1+2i
(f)i3+i4+i271
(2) Compute the absolute values of the complex numbers in Ex. 1.
0.7. COMPLEX NUMBERS: ALGEBRAIC PROPERTIES 27
(3) a) Add (1 + i) and (1 + 2 i) using compass and straight-edge.
b) Multiply (1 + i) and (1 + 2 i) using compass and straight-edge.
(4) Express in the form r(cosθ+isinθ),with 0 ≤θ<2π:
(a)i
(b) 2i
(c)−2i
(d) 4
(e)−1
(f)−1 +i
(g) (1 −i)3
(h)1
(1+i)2
(i)1
2(√
3 +i)
(5) Determine the
(a) three cube roots of i,−i, and of 1 + i,
(b) four fourth roots of −1 and +2
(c) six roots of z6= 1 .
(6) LetAbe any complex number, A=|A|[cosα+isinα] , and letw1,...,w nbe the
nroots ofzn= 1 . Prove that the nroots ofzn=Aare
z1=A1
nw1, z2=A1
nw2,...z n=A1
nwn,
where
A1
n=|A|1
n(cosα
n+isinα
n)
is the principal nth root ofA. This shows that the problem of finding the roots of
a complex number is essentially reduced to the simpler problem of finding the roots
of unity.
(7) Draw a sketch of the following sets of points in the complex plane.
(a){z∈C:|z−2| ≤1}
(b){z∈C:|z−1 +i| ≤2}
(c){z∈C:|z−2|>3}
(d){z∈C: 1≤ |z−2| ≤3}
(e){z∈C: 1≤ |z+i|<2}
28 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
0.8 Complex numbers: Completeness Properties, Complex
Functions.
We have just considered the algebraic properties of complex numbers. Now we look at
infinite sequences of complex numbers. To develop the desired properties of C, we shall
utilize those of R.
Definition: The sequence znof complex numbers converges to the complex number zif,
given any/epsilon1>0 , there is an Nsuch that |zn−z|</epsilon1for alln>N . We shall again write
zn→z.
In order to apply the theorem known for real sequences to complex sequences, the
following is vital.
Theorem 0.25 Letzn=xn+iyn, andz=x+iy. Thenznconverges to zif and only
if both the real and imaginary parts converge to their respective limits. In symbols,
zn→z⇐⇒xn→xandyn→y.
Proof: Sincezn→z, given any /epsilon1 >0 , we can find an Netc. for the zn’s. Now by
equation (0.7)
|xn−x| ≤ |zn−z|</epsilon1and|yn−y| ≤ |zn−z|</epsilon1
so bothxn→xandyn→y.
Conversely, given any /epsilon1>0 , we can find an N1for thexn’s and anN2for theyn’s.
LetNbe the larger of N1andN2,N= max (N1,N2) . ThisNworks for both the xn
andyn. But
|zn−z|=|xn+iyn−x−iy| ≤ |xn−x|+|yn−y|<2/epsilon1.
Thereforezn→z, completing the proof.
This theorem states that a definition is equivalent to some other property. We could
thus have used either property as a definition.
Recall that the real numbers were defined so that there would be no“hole” in the real
line. This was the completeness property. It guaranteed that if a sequence of real numbers
an“looked like” they were approaching a limiting value, then indeed th ere is some a∈R
such thatan→a. The issue here was to avoid the problem of a sequence of rational
numbers approaching an irrational number—which is a “hole” if our set just consisted of
the rationals. One consequence of the las t theorem is that the set of complex numbers C
is also complete.
Theorem 0.26 . Every bounded infinite sequence of complex numbers {zk}has at least
one subsequence which converges to a number z∈C. (By bounded, we mean that there is
somer∈Rsuch that |zk|<r for allk).
Proof: Since the {zk}are bounded, we know {xk}and{yk}are also bounded se-
quences of real numbers. The conclusion is now a consequence of the Bolzano-Weierstrass
theorem 5 applied to {xk}and{yk}, and of theorem 12 just proved. There is a fine
point though: how to get a subsequence of the zkwhose real and imaginary parts both
converge. The trick is first to select a subsequence {xkj}={Rez kj}of the {xk}which
converge to some x∈R. Then, from the related subsequence {ykj}={Imz kj}, select
0.8. COMPLEX NUMBERS: COMPLETENESS AND FUNCTIONS 29
a subsequence {ykjn}which converges to some y∈R. Then {xkjn}also converges to
x∈Rsozkjn→z, and we a re done.
With these technical results under our belts, sequences in Cbecome no more difficult
than those in R.
Let us briefly examine the elements of functions of a complex variable. A complex-
valued function f(z) of the complex variable zis a mapping of some subset z U ⊂C
into the complex numbers C, f:U→C. Two examples are f(z) =z2, andf(z) =1
z.
Both the domain and range of f(z) =z2are all of C, while the domain and range of
f(z) =1
zare all of Cwith the exception of 0.
Iffmaps R→R, likef(x) = 1 +xorf(x) =ex, since R⊂C, one asks how
the domain of definition of fcan be extended from RtoC. Of course there are many
possible ways to do this, but most of them are entirely artificial. For f(x) = 1 +x, the
natural extension is f(z) = 1+z, z∈C. Similarly, if P(x) =/summationtextN
k=0akxkis any polynomial
defined for x∈R, the natural extension to z∈CisP(z) =/summationtextN
k=0akzk. We are thus led
to extendf(x) =exforx∈R, toz∈Cby defining f(z) =e2. The only problem is that
we have absolutely no idea what it means to raise a real number, e, to a complex power.
Taylor (power) series are needed to resolve this issue. This will be carried out at the end
of Chapter 1.
Continuity of complex functions is defined in a natural way. Let z0be an interior point
of the setU⊂C(that is,z0is not on the boundary of U).
Definition: The function f:U→Cis continuous at the interior point a0/epsilon1Uif, given
any/epsilon1>0 there is a δ>0 such that |f(z)−f(z0)|</epsilon1for allzin 0<|z−z0|<δ.
Reasonable theorems like, if fandgare continuous at the interior point z0/epsilon1U, so is
the function f+g, are true too - with the same proof as was given for real-valued functions
of a real variable.
Although we could go on and define the derivative and integral for complex-valued
functionsf(z) of a complex variable, the development would take too much work. For
our future purposes, it will be sufficient to define the derivative and integral of a complex-
valued function f(x) of the realvariablex. The first step is to split f(x) into its real
and imaginary parts, that is, find real valued functions u(x) andv(x) such that f(x) =
u(x) +iv(x) . This decomposition c an always be done by taking
u(x) =f(x) +f(x)
2, v(x) =f(x) +f(x)
2i.
Sinceu(x) =u(x) andv(x) =v(x) , bothu(x) andv(x) are real-valued functions. It is
clear thatf(x) =u(x) +iv(x) .
Example: For the functions f(x) = 1 + 2ix, we havef(x) = 1−2ix, so
u(x) =(1 + 2ix) + (1 −2ix)
2= 1, v(x) =(1 + 2ix)−(1−2ix)
2i= 2x.
as expected.
Becausef(x) is a complex number for every xin the domain where fis defined, we
|f(x)|=/radicalbig
u2(x) +v2(x).
With this notion of absolute value, the definitions of continuity and differentiability read
just as iffwere itself real-valued. For example
30 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
Definition: : The complex-valued function f(x) of the real variable xisdifferentiable
at the point x0if
lim
x→x0f(x)−f(x0)
x−x0
exists.
A more convenient way of dealing with the derivative is supplied by the following
Theorem 0.27 . The function f(x) =u(x) +iv(x)is differentiable at a point x0if and
only if both u(x)andv(x)are differentiable there, and
df
dx=du
dx+idv
dx.
Proof: We shall use Theorem 12. Let {xn}be any sequence whose limit is x0. Define
the sequences {an},{αn},and{βn}by
an=f(xn)−f(x0)
xn−x0,
αn=u(xn)−u(x0)
xn−x0,andβn=v(xn)−v(x0)
xn−x0.
We must show that lim n→∞anexists if and only if both limits lim n→∞αnand lim n→∞βn
exist, for the existence of these limits is equivalent to the existence of the respective deriva-
tives. But notice that an=αn+iβn, since
an=f(xn)−f(x0)
xn−x0=u(xn) +iv(xn)−(u(x0) +iv(x0))
xn−x0=αn+iβn.
Thus we can appeal to Theorem 12 to conclude that lim anexists if and only if both lim αn
and limβnexist. The formula f/prime=u/prime+iv/primeis an immediate consequence since
an→f/prime(x0), αn→u/prime(x0),andβn→v/prime(x0)
Examples:
a) Iff(x) = 1 + 2ix,d f
dx=d
dx1 +id
dx2x= 2i
b) Iff(θ) = cos 7θ+isin 7θ+ 2θ−iθ2
d f
dθ=d
dθ[2θ+ cos 7θ] +id
dθ[−θ2+ sin 7θ] = 2−7 sin 7θ+i[−2θ+ 7 cos 7θ]
A related result which is even easier to prove is
Theorem 0.28 . The complex-valued function f(x) =u(x) +iv(x), x/epsilon1Ris continuous at
x0/epsilon1Rif and only if both u(x)andv(x)are continuous at x0.
Proof: An exercise.
Integration is defined more directly.
Definition: Letf(x) =u(x) +iv(x), x/epsilon1R. If the real-valued functions u(x) , andv(x)
are integrable for x/epsilon1[a,b] , we define the definite integral off(x) by
/integraldisplayb
af(x)dx=/integraldisplayb
au(x)dx+i/integraldisplayb
av(x)dx.
0.8. COMPLEX NUMBERS: COMPLETENESS AND FUNCTIONS 31
The standard theorems, like if cis any complex constant, then
/integraldisplayb
acf(x)dx=c/integraldisplayb
af(x)dx,and, ifa≤b,/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayb
af(x)dx/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤/integraldisplayb
a|f(x)|dx
are proved by using the definition above and the corresponding theorems for real functions.
We shall, however, need the more difficult
Theorem 0.29 . If the complex-valued function f(t) =u(t)+iv(t), t/epsilon1R, is continuous for
allt/epsilon1[a,b], then there is a constant Ksuch that |f(t)| ≤Kfor allt/epsilon1[a,b]. Furthermore
ifx, x 0/epsilon1[a,b], the n/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx
x0f(t)dt/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤K|x−x0|. (0-8)
Notice that the left-hand side absolute value is in the sense of complex numbers.
Proof: Sincef(t) is continuous in [ a,b] , by Theorem 15 so are both u(t) andv(t) . But
a real-valued function which is continuous in a closed and bounded interval is bounded.
Thus there are constants K1andK2such that |u(t)| ≤K1,|v(t)| ≤K2for allt/epsilon1[a,b] .
then
|f(t)|=/radicalbig
u2(t) +v2(t)≤/radicalBig
K2
1+K2
2≡K.
To prove the inequality (0.29), we use the inequality mentioned before the theorem to see
that ifx0≤x /vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx
x0f(t)dt/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤/integraldisplayx
x0|f(t)|dt.
Since |f(t)| ≤K, we find that
/integraldisplayx
x0|f(t)|dt≤K|x−x0|.
Combining these last two inequalities, we obtain the desired inequality (0.29) if x0≤x.
The other case, x≤x0, can be reduced to that already proved by observing that
/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx
x0f(t)dt/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle−/integraldisplayx
x0f(t)dt/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx
x0f(t)dt/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤K|x0−x|=K|x−x0|.
Exercises
(1) In the complex sequences below, which ones converge, which do not converge but
have at least one convergent subsequence, and which do neither? In all cases n=
1,2,3,....
(a)zn=i
n+ 3i−4
(b)zn= 2i+ (−1)n
(c)zn=n−i
(d)zn=in
(e)zn= 1 +i√
3−(−1)n
7n
32 CHAPTER 0. REMEMBRANCE OF THINGS PAST.
(f)zn=(4+6 i)n−5
1−2ni.
(2) Write the following complex-valued functions f(x) of the real variable xasf(x) =
u(x) +iv(x) , whereuandvand real-valued.
(a)f(x) =i+ 2(3 −2i)x2,
(b)f(x) = (1 + 2ix)2
(c)f(x) = cos 3x2−(3 +i) sinx
(d)f(x) =1
1+2i−x
(3) (a) Use the definition of the derivative to computed f
dxfor the function in Exercise 2a
above.
(b) Findd f
dxfor all the functions in Exercise 2 above.
(4) Evaluate
(a)/integraltext3
−1(1 + 2ix)dx
(b)/integraltext4
1[x+ (1−i) cos 2x]dx
Chapter 1
Infinite Series
1.1 Introduction
In elementary calculus you have met the notion of the limit of a sequence of numbers (see also
Chapter 0, sections 5 and 7). This concept of limit is just what essentially distinguishes
calculus from algebra. It was crucial in the definition of the derivative as the limit of a
difference quotient and the integral as the limit of a Riemann sum. We now propose to
discuss another limiting process, infinite series, in detail.
An infinite series is a sum of the form
∞/summationdisplay
k=1ak=a1+a2+a3+···, (1-1)
where the ak’s are real or complex numbers. Since there is no added difficulty we shall
suppose the ak’s are complex numbers. One immediate trouble is that it would take us an
infinite amount of time to add an infinite sum. For example, what is
(a)/summationtext∞
k=11 = 1 + 1 + 1 + 1 + ··· = ?
(b)/summationtext∞
k=11 = 1−1 + 1−1 + 1−1··· = ?
(c)/summationtext∞
k=11
2k−1= 1 +1
2+1
4+1
8+1
16+···=?
Thus, we are faced with the realization that be sum (1) is not really well defined, even
in cases where we feel it might make sense.
Our first task is to give a more adequate definition. Let Snbe the sum of the first n
terms:
Sn:=a1+a2+···+an=n/summationdisplay
k=1ak.
Then for each n, we have a complex number Sn, called the nthpartial sum of the series
(1).
Definition: If lim n→∞Sn=S, whereSis a (finite) complex number, we say that the
infinite series converges toS. If the sequence S1,S2,S3,... has no limit, we say that the
infinite series diverges .
For the examples given just above, we have
(a)Sn=/summationtextn
11 =n→ ∞ so the infinite series diverges to ∞.
(b)Sn=/summationtextn
1(−1)n+1=/braceleftbigg1nodd,
0neven,/bracerightbigg
.which does not have a limiting value since
it oscillates between 1 and 0.
33
34 CHAPTER 1. INFINITE SERIES
(c)Sn=/summationtextn
11
2k−1= 2(1 −1
2n)→2 , so the infinite series converges to the number 2
(we found the sum of the series by realizing it is a simple geometric series:
1 +r+r2+···+rN=1−rN+1
1−r) for (r/negationslash= 1).
With an adequate definition of convergence of infinite series, it is clear that we should
develop some tests for determining if a given series converges. That will be done in the next
section. In preparation, let us examine some simple types of series which occur often and
prove a few useful theorems.
There are two types of series whose sums can always be found, and for which the
question of convergence is exceedingly elementary.
Definition: An infinite geometric series is a series of the form
∞/summationdisplay
k=0ark=a+ar+ar2+···.
The partial sums are
Sn=a+ar+···+arn=a1−rn+1
1−rfor (r/negationslash= 1).
Theorem 1.1 The infinite geometric series/summationtext∞
k=0ark, a/negationslash= 0, converges if and only if
|r|<1. Then the sum isa
1−r.
Proof: limn→∞rn+1exists only if |r|<1 . Then the limit is zero so lim n→∞Sn=a
1−r
(the non-convergence when |r|= 1 follow from Theorem 6, p. ?)
Examples:
(a)/summationtext∞
k=0(1 +i)kdiverges since |1 +i|=√
2≥1 .
(b)/summationtext∞
k=0(1+i
2)kconverges since/vextendsingle/vextendsingle1+i
2/vextendsingle/vextendsingle=√
2
2<1.The sum of this series is 1 + i.
(c)/summationtext∞
k=11 diverges since |1|= 1 .
(d)/summationtext∞
k=1(−1)kdiverges since |−1|= 1 .
Definition: An infinite telescopic series is one of the form
∞/summationdisplay
k=1(αk−αk+1) = (α1−α2) + (α2−α3) + (α3−α4) +···.
It is clear that most of the terms cancel each other.
Theorem 1.2 Ifαk→α, then/summationtext∞
k=1(αk−αk+1) =α1−α.
Proof:Sn= (α1−α2) + (α2−α3) +···+ (αn−αn+1) =α1−αn+1→α1−α.
Examples:
(a)1
1·2+1
2·3+1
3·4+···=/summationtext∞
k=11
k(k+1)=/summationtext∞
k=1(1
k−1
k+1) = 1 .
(b)1
4·12−1+1
4·22−1+1
4·32−1+···=1
2/summationtext∞
k=1(1
2k−1−1
2k+1) =1
2
1.1. INTRODUCTION 35
We close this section with some reasonable (and desirable) theorems. The proofs are
immediate consequences of the definition of convergence of infinite series and the related
theorems about limits of sequences of numbers.
Theorem 1.3 . If/summationtext∞
k=1ak→a, andcis any number then/summationtext∞
k=1cak→ca.
Theorem 1.4 If/summationtextn
k=1ak→aand/summationtextn
k=1bk→b, then/summationtextn
k=1(ak+bk)→a+b.
Theorem 1.5 Letak=αk+iβk, whereαkandβkare real. The infinite series/summationtextak
converges if and only if the two real series/summationtextαkand/summationtextβkboth converge. That is, an
infinite complex series converges if and only if both its real and imaginary parts converge.
Proof: We must look at the partial sums. Let σn=/summationtextn
k=1αk, andτn=/summationtextn
k=1βk. Then
Sn=n/summationdisplay
k=1αk=n/summationdisplay
k=1(αk+iβk) =n/summationdisplay
k=1αk+in/summationdisplay
k=1βn=σn+iτn.
But we know from Theorem 12 of Chapter 0 that the complex sequence Snconverges if
and only if both its real part, σn, and imaginary part, τn, both converge—in other words,
if the series/summationtextakand/summationtextβkboth converge.
Two remarks should be made in an attempt to mitigate some confusion. First, the index
kof the series/summationtext∞
k=1akcould have been any other letter. Thus/summationtext∞
k=1ak=/summationtext∞
j=1aj. This
is perhaps indicated most clearly if we left an empty box instead of using any letter at all:
. The connecting line means that the same letter must be used in both boxes. Now you
can fill in any letter that makes you happy. No matter w hat you write, it still means
a1+a2+a3+···. In a similar way, the index need not begin with 1. Thus, for example,/summationtext∞
k=1ak=/summationtext∞
k=17ak−16=a1+a2+···. Although this manipulation looks like unwanted
silliness here, it is sometimes quite useful. Later on this year you will need it. The related
transformation for integrals is illustrated by
/integraldisplay3
21
tdt=/integraldisplay2
11
t+ 1dt.
Exercises
(1) Find a closed form expression for the nthpartial sum of the following infinite series
and determine if they converge.
(a)2
3+2
9+2
27+···+2
3n+···=/summationtext∞
k=12
3k.
(b) 1 +i+i2+i3+···+in+···
(c)1
2!+2
3!+3
4!+···=/summationtext∞
k=2k−1
k!=/summationtext∞
k=2(1
(k−1)!−1
k)
(d) ln1
2+ ln2
3+ ln3
4+···+ ln(n
n+1) +···
(e)/summationtext∞
m=0(3−4i
7)m
(f)/summationtext∞
n=12−3i
n(n+1)
36 CHAPTER 1. INFINITE SERIES
(2) The repeating decimal 1 .565656 ···can be written as
1 +56
102+56
104+56
106+···= 1 + 56∞/summationdisplay
k=1(1
102)k.
Sum the geometric series and find what rational number the repeating decimal repre-
sents. In a similar way, every decimal which begins to repeat eventually is a rational
number. What rational number is represented by 1.4723?
(3) A ball is dropped from a height of 20 feet. Every time it bounces, it rebounds to3
4
of its height on the previous bounce. What is the total distance traveled by the ball?
(4) If/summationtext∞
k=1ak→aand/summationtext∞
k=1bk→b, and ifαandβare any numbers, prove that/summationtext∞
k=1(αak+βbk)→αa+βb.
(5) Ifan>0 and/summationtextanconverges, prove that/summationtext1
andiverges.
(6) Does the convergence of/summationtext∞
n=1animply the convergence of/summationtext∞
n=1(an+an+1) ?
(7) (a) If the partial sums of/summationtextanare bounded, and {bn}is a strictly decreasing
sequence with limit 0, bn/arrowsoutheast0 , prove that/summationtextanbnconverges.
(b) Use (a) to prove that if/summationtext∞
n=1nanconverges then so does the series/summationtext∞
n=1an.
(c) Use (a) to discuss the convergence of/summationtext∞
n=1sinnx
n.
1.2 Tests for Convergence of Positive Series
Tests to determine convergence are of several types, i) those that give sufficient conditions,
ii) those that give necessary conditions, and iii) those that give both necessary and sufficient
conditions. Theorem 1 of the last section governing geometric series was of the last type;
however it is more common to find convergence tests of the first two types since they are
usually easier to come by. You should be careful to observe the nature of a test . A simple
theorem should make the point clear.
Theorem 1.6 . If the series/summationtext∞
k=1ak—whereakmay be complex—converges, then
lim
k→∞|ak|= 0.
Proof: LetSn=a1+a2+···+an. Then |an|=|Sn−Sn−1|. Asn→ ∞ bothSnand
Sn−1tend to the same limit, so |an| →0 .
Returning to the point made before, this theorem states a necessary but not sufficient
(as we shall see) condition for an infinite series to converge. We can apply it to see that/summationtextk
k+1diverges—sincek
k+1→1/negationslash= 0 . Thus this theorem is useful as a quick crude test to
weed out series which diverge badly. But all it tells us about the series/summationtext∞
k=11
k—for which
1
k→0 so the criterion of the theorem is satisfied—is that it might converge. In fact, this
series diverges too, as we shall now prove.
∞/summationdisplay
k=11
k= 1 +1
2+1
3+1
4+1
5+···+1
8+1
9+···+1
16+1
17+···+1
32+···
1.2. TESTS FOR CONVERGENCE OF POSITIVE SERIES 37
1 +1
2+1
4+1
4+1
8+···+1
8+1
16+···+1
16+1
32+···+1
32+···
= 1 +1
2+1
2+1
2+1
2+1
2+···..
ThusS1= 1, S2= 1+1
2, S4>1+1
2+1
2= 1+2 ·1
2, S8>1+3·1
2, S16>1+4·1
2,...,S 2n>
1 +n·1
2.We can easily see that as n→ ∞, S2n→ ∞ , so the series/summationtext1
k, called the
harmonic series , diverges.
For the many series which slip through the test of Theorem 6, more refined criteria are
needed. The criteria we shall present in the remainder of this section are for series with
positive terms,an≥0 . Application of these criteria to series with complex terms will be
made in the next section.
Theorem 1.7 . Ifak≥0for eachk, then the series/summationtext∞
k=1akconverges if and only if
the sequence of partial sums is bounded from above.
Proof: Since all the ak’s are non-negative, Sn+1≥Sn. Thus the Sn’s are a monotone
increasing sequence of real numbers. By Theorems 6 and 8 of Chapter 0, this sequence Sn
converge if and only if it is bounded.
Example: The series/summationtext∞
k=11
k!of positive terms converges, since
1
k!=1
1·2·3···.k≤1
1·2·2·2···.2=1
2k−1
so
Sn=n/summationdisplay
k=11
k!≤n/summationdisplay
k=11
2k−1≤∞/summationdisplay
k=01
2k= 2.
The convergence now follows since Snis bounded from above.
We can extract an exceedingly useful idea from these examples: check the convergence
of a given series by comparing it with another series which we know to converge or diverge.
Theorem 1.8 . (comparison test ) Let/summationtextakand/summationtextbkbe two positive series for which
ak≤bkforn>N . Then
i) if/summationtextbkconverges, so does/summationtextak.
ii) if/summationtextakdiverges, so does/summationtextbk.
Proof: Letsn=/summationtextn
k+N+1akandtn=/summationtextn
k+N+1bk. Thensn≤tnfor alln>N , so i) if
tn→t, thensnis bounded ( sn≤t) , ii) ifsn→ ∞ , thentn→ ∞ too.
Remark: The “n>N ” part of the hypothesis reflects the fact that it is only the infinite
tail of an infinite series that we need to worry about. Any finite number of terms can always
be added later on.
Examples:
(a)/summationtext∞
k=11
2k+1converges since1
2k+1<1
2kand/summationtext1
2kconverges.
(b)/summationtext∞
k=11√
kdiverges since1√
k≥1
k(fork≥1 ) and/summationtext1
kdiverges.
Our next test is based upon comparison with a geometric series/summationtextrn.
38 CHAPTER 1. INFINITE SERIES
Theorem 1.9 . (ratio test ) Let/summationtextanbe a series with positive terms such that the
following limit exists
lim
n→∞an+1
an=L.
Then
i) ifL<1, the series converges
ii) ifL>1, the series diverges
iii) ifL= 1, the test is inconclusive.
Remark: If the assumed limit does not exist, a variant of the theorem is still true but we
shall not discuss it.
Proof: i) IfL < 1 , pick any r, L < r < 1 . Then there is an Nsuch that for all
n≥N,an+1
an< r. Therefore an< ra n−1< r2an−2< ... < rn−NaN, so thatan<
Krn, n≥N, whereK >aN
rN. The series/summationtext∞
n=1an=/summationtextN−1
n=1an+/summationtext∞
n=Nanconsists of a
finite sum plus an infinite tail which is dominated by the geometric series/summationtextKrn. Since
r<1 , the geometric series converges and by the comparison test, so does/summationtextan.
ii) IfL>1 , thenan+1>anfor alln>N ; thus lim n→∞an/negationslash= 0 . By Theorem 6, the
series/summationtextancannot converge.
iii) This is seen from the two examples.
(a)/summationtext1
n,with lim n→∞an+1
an= lim n→∞n
n+1= 1 , which we know diverges.
(b)/summationtext1
n(n+1),with lim n→∞an+1
an= lim n→∞(n+1)(n+2)
n(n+)= 1 , which we know (Theorem
2, Example a) converges.
In both these cases L= 1 . You should notice that the criterion uses the limiting value
ofan+1/an. The divergent harmonic series/summationtext1
n, whose ratio n/n+ 1 is less than one
for finiten, but 1 in the limit shows the mistake you will make if you use the ratio before
passing to the limit.
Examples:
(a)/summationtext1
n!: Since lim n→∞(an+1
an) = lim n→∞(n!
(n+1)!) = lim n→∞1
n+1= 0<1 , the ratio is
less than one so the series converges.
(b)/summationtext10n
n!:Since limn→∞(an+1
an) = lim n→∞(10
n+1) = 0<1 , the series converges.
(c)/summationtextn!
2n:Since limn→∞(an+1
an) = lim n→∞(n+1
2) =∞, the series diverges.
Our last test for series with positive terms is associated with a picture. The crux of the
matter is very simple and clever. We associate an area with the infinite series/summationtext∞
n=1an.
For the term anwe use a rectangle between n≤x≤n+ 1 of height anand base one.
Then the sum of the infinite series is represented by total area under the rectangles. Now
by Theorem 7, if all the an’s are positive we know the series converges if the total area
is finite. Thus, if we can find a function f(x) whose graph lies above the rectangles, and
whose total area is finite, then we know the area contained in the rectangles is finite and so
the series converges.
Theorem 1.10 . (integral test ) Let/summationtext∞
n=ianbe a series of positive decreasing terms:
0< a n+1≤an, andf(x)a continuous decreasing function with f(n) =an. Then the
sequence
SN=N/summationdisplay
n=1anandTN=/integraldisplayN
1f(x)dx
1.2. TESTS FOR CONVERGENCE OF POSITIVE SERIES 39
either both converge or both diverge, in fact, SN−a1≤TN≤SN−1.
Proof: First of all,
/integraldisplayN
1f(x)dx=/integraldisplay2
1+/integraldisplay3
2+···+/integraldisplayN
N−1=N−1/summationdisplay
n=1/integraldisplayn+1
nf(x)dx.
Since in the interval n≤x≤n+ 1 we know that
an=f(n)≥f(x)≥f(n+ 1) =an+1,
we see that
an=/integraldisplayn+1
nf(n)dx≥/integraldisplayn+1
nf(x)dx≥/integraldisplayn+1
nf(n+ 1)dx=an+1.
Adding these up, we find
N−1/summationdisplay
n=1an≥N−1/summationdisplay
n=1/integraldisplayn+1
nf(x)dx≥N−1/summationdisplay
n=1an+1
or
N−1/summationdisplay
n=1an≥/integraldisplayN
1f(x)dx≥N/summationdisplay
n=2an.
Thus
SN−1≥TN≥SN−a1.
From this last inequality, we see that lim n→ ∞TNis finite if and only if lim n→∞SN
is finite. Since the sequences SNandTNare both monotone increasing sequences, by
Theorem 7 the sequences converge or diverge together. And we are done.
Examples:
(a)./summationtext∞
n=11
npconverges if p>1 , diverges if p≤1 . We use the function f(x) =1
xp,
which satisfies the hypothesis of the theorem, and examine the integral
TN=/integraldisplayN
11
xpdx=/braceleftBigg
N1−p−1
1−p, p/negationslash= 1.
lnN , p = 1/bracerightBigg
.
AsN→ ∞,lnN→ ∞ , and so does N1−pifp<1 , whileN1−p→0 ifp>1 . Therefore
TNconverges if and only if p >1 , so by our theorem/summationtext∞
n=11
npconverges if and only if
p > 1 . In the special case p= 1 we have again proven that the harmonic series/summationtext1
n
diverges. Another often seen special case is p= 2,/summationtext1
n2, which converges. Sometime later
we shall prove the amazing/summationtext∞
n=11
n2=π2
6.
(b)/summationtext∞
n=21
nlnndiverges since as N→ ∞,/integraltextN
2dx
xln 2= ln(lnN)−ln(ln 2) → ∞
Exercises
(1) Determine if the following series converge or diverge.
(a)/summationtext∞
n=11
n2+1
40 CHAPTER 1. INFINITE SERIES
(b)/summationtext∞
n=11
2n−1
(c)/summationtext∞
n=11
n(lnn)2
(d)/summationtext∞
n=11
10n2
(e)/summationtext∞
n=1n
n2+1
(f)/summationtext∞
n=11
2n+3
(g)/summationtext∞
n=1cos2n
2n
(h)/summationtext∞
n=1√n
n3+1
(i)/summationtext∞
n=1n2
2n
(j)/summationtext∞
n=1ne−n2
(k)/summationtext∞
n=11√
n(n+1)(n+2)
(l)/summationtext∞
n=1n!
22n
(m)/summationtext∞
n=1|an|
10n,|an|<10
(n)/summationtext∞
n=1npe−n, p∈R
(2) Ifan≥0 andbn≥0 for alln≥1 , and if there is a constant csuch thatan≤cbn,
prove that the convergence of/summationtextbnimplies the convergence of/summationtextan.
(3) Use the geometric idea of the integral test to show lim n→∞[1 +1
2+···+1
n−lnn]
converges to a constant γ, and show that1
2<γ < 1 .γis called Euler’s constant .
(4) If/summationtextanconverges, where an≥0 , prove that/summationtextan
1+analso converges.
(5) (a). If/summationtextanconverges, where an≥0 , andcnhave the property 0 ≤cn≤K, the
sameKfor alln, then prove that/summationtextcnanconverges.
(b). Deduce the result of Exercise 4 from Exercise 5a.
(6) Use the geometric idea behind the integral test to prove that
(a). lnn! = ln 1 + ln 2 + ln 3 + ···+ lnn >/integraltextn
1lnxdx=nlnn−n+ 1 whenn≥2 .
From this deduce that
(b).n!>e(n
e)n, whenn≥2 .
(c). As an application of (b), prove that lim n→∞xn
n!= 0 .
(7) (a). Use the idea in the proof of the divergence of the harmonic series,/summationtext1
n, to prove
the following test for convergence: Let {an}be a positive monotonically decreasing
sequence. Then/summationtextanconverges or diverges respectively if and only if the “condensed”
series/summationtext2na2nconverges or diverges.
(b). Apply the test of part (a) to again prove that/summationtext1
npconverges if p >1 , and
diverges if p≤1 .
(c). Apply the test of part (a) to determine the values of pfor which the series/summationtext∞
n=21
n(lnn)pconverges and diverges.
1.3. ABSOLUTE AND CONDITIONAL CONVERGENCE 41
1.3 Absolute and Conditional Convergence
The tests just given for series with positive terms can be applied to many series with complex
termsanby utilizing the concept of absolute convergence.
Definition: The series/summationtext∞
k=1ak, where the akmay be complex numbers, converges
absolutely if the series of positive numbers/summationtext∞
k=1|ak|converges. It is called conditionally
convergent if/summationtext∞
k=1akconverges but/summationtext∞
k=1|ak|diverges.
Absolute convergence is stronger than ordinary convergence because
Theorem 1.11 . If/summationtext∞
n=1|an|converges, then/summationtext∞
n=1anconverges.
Proof: LetaN=αn+iβn. We shall show that the real series/summationtextαnand/summationtextβnboth
converge. Then by Theorem 5/summationtextanconverges too. To show that/summationtextαnconverges, let
cn=αn+|an|. Since |αn| ≤/radicalbig
(α2n+β2n) =|an|, we know that 0 ≤cn≤2|an|. Thus
the positive series/summationtextcnis bounded,/summationtextcn≤2/summationtext|an|<infty, and so converges by the
comparison test (Theorem 8). Since/summationtextαn=/summationtext(cn− |an|) , and both/summationtextcnand/summationtext|an|
converge, then/summationtextαnalso converges by Theorem 4. Similarly, by taking dn=βn+|an|,
the series/summationtextdnconverges, from which we can conclude that/summationtextβnconverges.
Examples:
(a) The complex series1
12+i
22+i2
32+i3
42+···=/summationtext∞
n=1in
n2converges absolutely since/vextendsingle/vextendsingle/vextendsinglein−1
n2/vextendsingle/vextendsingle/vextendsingle=1
n2and the positive series/summationtext∞
n=11
n2converges.
(b) 1 +1
22−1
23−1
24+1
25+1
26−1
27−1
28+···, which is the geometric series/summationtext1
2nwith
negative signs thrown in, converges absolutely since/summationtext1
2nconverges.
(c)/summationtextrn, rcomplex, converges absolutely if/summationtext|r|nconverges, that is, if |r|<1 .
(d) 1−1
2+1
3−1
4+1
5...=/summationtext∞
n=1(−1)n+1
n, the alternating harmonic series does not converge
absolutely because/summationtext1
ndiverges. It does converge though, as we shall see shortly.
Thus the alternating harmonic series is conditionally convergent.
On the basis of this last theorem, many complex series can be proved to converge by
proving they converge absolutely. Since absolute convergence concerns itself with series
having only positive terms, all the tests for convergence developed in the previous section
may be used. This is the most common way of proving a complex series converges. If it
does not converge absolutely, the proof of convergence will usually be more difficult and use
special ingenuity based on the particular series at hand.
There is one case of conditional convergence which is easy to treat, that of alternating
series.
Definition: A series of real numbers is called alternating if the positive and negative terms
occur alternately. They have the form
∞/summationdisplay
n=1(−1)n−1an=a1−a2+a3−a4+···,
where thean’s are all positive.
42 CHAPTER 1. INFINITE SERIES
Theorem 1.12 . The alternating series/summationtext∞
n=1(−1)n−1an, an>0, converges if i) the an
are monotone decreasing ( an/arrowsoutheast), and ii) limn→∞an= 0. IfSis the sum of the series,
the inequality
0<|S−SN|<aN+1 (1-2)
shows how much the Nthpartial sum differs from the limit S. In words inequality (2)
says that the error which results by using the first Nterms is less than the first neglected
termaN+1.
Proof: The idea is quite simple. Observe that since an/arrowsoutheast, S2n−S2n−2=a2n−1−a2n>0 ,
so theS2n’s increase. Similarly the S2n+1’s decrease. Also both sequences are bounded—
from below by S2and from above by S1(you should check this). Therefore by Theorem
8 Chapter 0, the bounded monotonic sequences S2nandS2n+1converge to real numbers
Sand ˆSrespectively. Let us show that S=ˆS.
ˆS−S= lim
n→∞S2n+1−lim
n→∞S2n= lim
n→∞(S2n+1−S2n) = lim
n→∞a2n+1= 0
Thus the alternating series converges to the unique limit S. All that is left to verify is
inequality (2). Because S2nis increasing and S2n+1is decreasing, we know that
S2n<S andS <S 2n+1
Therefore
0<S−S2n<S 2n+1−S2n=a2n+1 and 0<S 2n−1−S <S 2n−1−S2n=a2n.
These two inequalities are the cases Neven andNodd in (2).
Examples:
(a)/summationtext∞
n=1(−1)n−1
nconverges since it is an alternating sequence and1
ndecreases mono-
tonically to zero. Later we shall show that its sum is ln 2 .
(b)/summationtext∞
n=2(−1)n
lnnconverges since1
lnndecreases monotonically to zero.
(c)/summationtext∞
n=1(−1)n−1n
n+1diverges by Theorem 6 since lim n→∞(−1)n−1n
n+1is not zero.
Exercises
(1) Determine which of the following series converge absolutely, converge conditionally,
or diverge.
(a)/summationtext∞
n=1(−1)n+1
√n
(b)/summationtext∞
n=1(2−3i)n
n!
(c)/summationtext∞
k=2(2k+i)2
ek
(d)/summationtext∞
n=1(−1)n−1lnn
n
(e) 1 −1
2+1
3−1
22+1
5−1
23+1
7−1
24+1
9− ···.
(f)/summationtext∞
n=11
n2+2i
1.4. POWER SERIES, INFINITE SERIES OF FUNCTIONS 43
(g)/summationtext∞
n=11
n+2i
(h)/summationtext∞
n=1(−1)n−1
n+2i
(i)/summationtext∞
n=1(−1)n−1
np, p> 0,
(j)/summationtext∞
n=1(−1)n(1+i)n2
2n2+1
(k)/summationtext∞
n=1cosnθ
n2, θarbitrary.
(2) If/summationtextanand/summationtextbnare absolutely convergent, and αandβare any complex numbers,
prove that/summationtext(αan+βbn) also converges absolutely.
(3) Show that/summationtext∞
n=1nznconverges absolutely if |z|<1 .
(4) Show that for any θ∈R, then/summationtext∞
n=0cosnθdiverges, and that if θ/negationslash= 0,±π,±2π,... ,
then/summationtext∞
n=0sinnθalso diverges.
1.4 Power Series, Infinite Series of Functions
As you will all agree, the simplest functions are polynomials. With infinite series at hand,
it is reasonable to consider an “infinite polynomial”
a0+a1z+a2z2+a3a3+···=∞/summationdisplay
n=0anzn.
Because of the appearance of the powers of z, this is called a power series . The question
of convergence of a power series is trivial at z= 0 , for then we have only the one term a0.
Does this series converge for any other values of z, and if so, for which ones?
The answer depends on the coefficients an, but in any case, the set of complex numbers,
z∈C, for which the series converges is always a disc |z|<ρ—with possibly some additional
points on the boundary |z|=ρ—in the complex pane bC. This number ρis called the
radius of convergence of the power series. We shall first prove that the set z∈Cfor which
a power series converges is always a disc. Then we shall give a way of computing the radius
ρof that disc.
Theorem 1.13 . The setz∈Cfor which the power series/summationtextanznconverges is always
a disc |z|< ρ, inside of which it even converges absolutely. We do not exclude the two
extreme possibilities that the radius of this disc is zero or infinity.
The series might converge at some, none, or all of the points on the boundary of the
disk|z|=ρ.
Proof: We shall show that if the series converges for any ζ∈C, then it converges
absolutely for all complex zwith|z|<|ζ|. Ifζ= 0 , there is nothing to prove, so assume
ζ/negationslash= 0 . Because/summationtextanζnconverges, lim n→∞|anζn| →0 . Thus all the terms are bounded
in absolute value, that is, there is an Msuch that |anζn|<M for alln. Then, since
|anzn|=/vextendsingle/vextendsingle/vextendsingle/vextendsingleanζnzn
ζn/vextendsingle/vextendsingle/vextendsingle/vextendsingle<M/vextendsingle/vextendsingle/vextendsingle/vextendsinglez
ζ/vextendsingle/vextendsingle/vextendsingle/vextendsinglen
for alln,
44 CHAPTER 1. INFINITE SERIES
the series/summationtext|anzn|is dominated by M/summationtext/vextendsingle/vextendsingle/vextendsinglez
ζ/vextendsingle/vextendsingle/vextendsinglen
. But this last series is a geometric series
which does converge since |z|<|ζ|, so/vextendsingle/vextendsingle/vextendsinglez
ζ/vextendsingle/vextendsingle/vextendsingle<1 . Thus by the comparison test/summationtextanzn
converges absolutely for all z∈Cwith|z|<|ζ|.
Therefore, if the power series/summationtextanznconverges for some complex number ζ, then it
converges in the whole disc |z|<|ζ|. The radius of convergence ρis then the radius of
the largest disc |z|<ρfor which the series converges.
See Exercise 3 for examples concerning convergence on the boundary of the disk.
Let us now give a method of computing ρwhich covers most cases arising in practices.
Theorem 1.14 . If limn→∞/vextendsingle/vextendsingle/vextendsinglean+1
an/vextendsingle/vextendsingle/vextendsingle=Lexists, the power series/summationtextanznhas radius of
convergence ρ=1
LifL/negationslash= 0,∞ifL= 0. In other words, if L/negationslash= 0 the series converges
in the disc |z|<1
Land diverges if |z|>1
L. On the circumference |z|= 1/L, anything
may happen (see Exercise 3 at the end of this section). If L= 0, the series converges in
the whole complex plane.
Proof: This is a simple application of the ratio test. The series converges if the limit of
the ratio of successive terms lim n→∞/vextendsingle/vextendsingle/vextendsinglean+1zn+1
anzn/vextendsingle/vextendsingle/vextendsingleis less than one and diverges if it is greater
than one. Thus we have convergence if
lim
n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglean+1z
an/vextendsingle/vextendsingle/vextendsingle/vextendsingle=|z|L<1,i.e. if |z|<1
L,
and divergence if
lim
n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglean+1z
an/vextendsingle/vextendsingle/vextendsingle/vextendsingle=|z|L>1,i.e. if |z|>1
L.
Remark: In the one additional case/vextendsingle/vextendsingle/vextendsinglean+1
an/vextendsingle/vextendsingle/vextendsingle→ ∞ asn→ ∞ , the series diverges for every
|z| /negationslash= 0 , as can easily be seen again by the ratio test.
Examples:
(a)/summationtext∞
n=0znconverges where lim n→∞/vextendsingle/vextendsinglezn+1/zn/vextendsingle/vextendsingle<1 that is; for |z|<1 .
(b)/summationtext∞
n=0nzn
2nconverges where lim n→∞/vextendsingle/vextendsingle/vextendsingle(n+1)zn+1
2n+1/nzn
2n/vextendsingle/vextendsingle/vextendsingle<1 . Since
lim
n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsingle(n+ 1)zn+1
2n+1/nzn
2n/vextendsingle/vextendsingle/vextendsingle/vextendsingle= lim
n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsingle(n+ 1)z
2n/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsinglez
2/vextendsingle/vextendsingle/vextendsingle,
the series converges for all |z|<2 .
(c)/summationtext∞
n=0zn
n!converges where lim n→∞/vextendsingle/vextendsingle/vextendsinglezn+1
(n+1)!/zn
n!/vextendsingle/vextendsingle/vextendsingle<1 . Since
lim
n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglezn+1
(n+ 1)!/zn
n!/vextendsingle/vextendsingle/vextendsingle/vextendsingle= lim
n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglez
n+ 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle= 0,
the series converges for all z∈C, that is, in the whole complex plane.
1.4. POWER SERIES, INFINITE SERIES OF FUNCTIONS 45
(d)/summationtext∞
n=0n!znconverges where lim n→∞/vextendsingle/vextendsingle/vextendsingle(n+1)!zn+1
n!zn/vextendsingle/vextendsingle/vextendsingle<1 . But
lim
n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsingle(n+ 1)!zn+1
n!zn/vextendsingle/vextendsingle/vextendsingle/vextendsingle= lim
n→∞|(n+ 1)z|=∞
unlessz= 0 . Thus the ratio is less than one only at z= 0 , so the series converges
only at the origin.
Only minor modifications are needed for the more general power series
a0+a1(z−z0) +a2(z−z0)2+···=∞/summationdisplay
n=0an(z−z0)n,
wherea0∈C. Again the series converges in a disc in the complex plane, only now the disc
has its center at z0instead of the origin, so if the radius of convergence is ρ, the series
converges for |z−z0|<ρ. An example should make this clear.
Example:/summationtext∞
n=1(z−2i)n
n. By the ratio test, this converges when
lim
n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsingle(z−2i)n+1
n+ 1/(z−2i)n
n/vextendsingle/vextendsingle/vextendsingle/vextendsingle<1,
that is, when |z−2i|<1 . This is a disc with center at 2 iand radius 1.
A few words should be said about real power series/summationtextan(x−x0)nwhere both xand
x0are real (some people only use this phrase if the anare also real). This is a special case
of/summationtextan(z−z0)nwherez0is on the real axis and we only ask for what realzthe series
converges. However we know that/summationtextan(z−z0)nconverges only for those zin the disc
of convergence |z−z0|<ρ—and possibly some boundary points. Thus the realvalues of
zfor which the series/summationtextan(z−z0)nconverges are exactly those points on the real axis
which are also inside the disc of convergence of the complex power series. In particular the
series/summationtextan(x−x0)n, with both xandx0real converges for |x−x0|< ρ, i.e., in the
intervalx0−ρ≤x≤x0+ρ.
Example: For whatx∈Rdoes/summationtext∞
n=01
2n(x−1)nconverge? The related complex series/summationtext∞
n=01
2n(z−1)nconverges in the disc |z−1|<2 . The points on the real axis which are
in this disc are |x−1|<2 , which is −1<x< 3 . A direct check shows the series diverges
at both end points x=−1 andx= 3 . If/summationtextanand/summationtextbnboth converge, can we define
their product in a meaningful way
(∞/summationdisplay
n=0an)(∞/summationdisplay
n=0bn) =∞/summationdisplay
n=0cn?
and if so, does the resulting series converge? The most simple-minded approach is to insert
powers ofz(a bookkeeping device), giving (/summationtextanzn)(/summationtextbnzn) , try long multiplication and
see what happens. A computation shows that
(a0+a1z+axz2+···)(b0+b1z+b2z2+···) =a0b0+ (a0b1+a1b0)z
+(a0b2+a1b1+a2b0)z2+···+ (a0bn+a1bn−1+···+anb0)zn+···.
46 CHAPTER 1. INFINITE SERIES
Motivated by this, we make the following
Definition: The formal product , called the Cauchy product , of the series/summationtextanand/summationtextbn
is defined to be
(∞/summationdisplay
n=0an)(∞/summationdisplay
n=0bn)≡∞/summationdisplay
n=0cn,
where
cn=a0bn+azbn−1+···+anb0=n/summationdisplay
k=0akbn−k.
With this definition we shall answer the question we raised about multiplication of
power series.
Theorem 1.15 . If/summationtext∞
n=0an=Aand/summationtext∞
n=0bn=Bboth converge absolutely, then the
Cauchy product series
(∞/summationdisplay
n=0an)(∞/summationdisplay
n=0bn)≡(∞/summationdisplay
n=0cn),
where
cn=∞/summationdisplay
k=0akbn−k,
also converges absolutely, and to C=AB.
Proof: LetAN=/summationtextN
n=0an, BN=/summationtextN
n=0bn,andCN=/summationtextN
n=0cn. We shall show that by
pickingNlarge enough, |ANBN−CN|can be made arbitrarily small. Since ANBN→
AB, this will complete the proof. Observe that
CN=a0b0+ (a0b1+a1b0) +···+ (a0bN+···+aNb0) =/summationdisplay/summationdisplay
ajbk,
while
ANBN= (a0+···+aN)(b0+···+bN) =N/summationdisplay
j=0N/summationdisplay
k=0ajbk.
Therefore
|ANBN−CN|=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingleN/summationdisplay
j=0N/summationdisplay
k=0ajbk/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤N/summationdisplay
j=0N/summationdisplay
k=0|aj| |bk|.
Sincej+k>N , eitherj >N/ 2 ork>N/ 2 , so
|ANBN−CN| ≤N/summationdisplay
j>N
2N/summationdisplay
k=0|aj| |bk|+N/summationdisplay
j=0N/summationdisplay
k>N
2|aj| |bk|.
Because the original series both converge absolutely, they are bounded,
∞/summationdisplay
j=0|aj|<M and∞/summationdisplay
k=0|bk|<M.
Consequently,
|ANBN−CN| ≤M(∞/summationdisplay
j>N
2|aj|+∞/summationdisplay
k>N
2|bk|).
1.4. POWER SERIES, INFINITE SERIES OF FUNCTIONS 47
Again using the absolute convergence of the original series, we see that for Nlarge,
the right side can be made arbitrarily small.
Since we shall need the ideas later on, let us digress briefly and examine the convergence
of infinite series of functions,/summationtextun(z) . In the special case where un(z) =an(z−z0)n,
this is a power series. Generally, there is little one can say about the convergence of such
series except to apply our general tests and hope for the best. We shall only illustrate the
situation with two
Examples:
(a)/summationtext∞
n=1cosnθ
n2, whereθis any real number. This converges for all θsince it converges
absolutely, that is/summationtext+/vextendsingle/vextendsinglecosnθ
n2/vextendsingle/vextendsingleconverges. We can see this last statement is true
by comparing/summationtext+/vextendsingle/vextendsinglecosnθ
n2/vextendsingle/vextendsinglewith the larger convergent series (since |cosnθ| ≤ 1 )/summationtext∞
n=11
n2.
(b)/summationtext∞
n=1nenx. By the ratio test, converges if lim n→∞/vextendsingle/vextendsingle(n+ 1)e(n+1)x/nenx/vextendsingle/vextendsingle<1 . Since
lim
n→∞/vextendsingle/vextendsingle/vextendsingle(u+ 1)e(n+1)x/nenx/vextendsingle/vextendsingle/vextendsingle= lim
n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglen+ 1
n/vextendsingle/vextendsingle/vextendsingle/vextendsingleex=ex,
the series converges if ex<1 , which happens only when x<0 .
Exercises
(1) Find the disc of convergence of the following power series by finding the center and
radius of the disc.
(a)/summationtext∞
n=0zn
n+1
(b)/summationtext∞
n=0(z−2)n
n
(c)/summationtext∞
n=0in
2n−1zn−1
(d)/summationtext∞
n=0(n+ 1)[z−2 + 3i]n
(e)/summationtext∞
n=0(2z+3)n
n2+2i
(f)/summationtext∞
n=01
lnnzn−2
(g)/summationtext∞
n=0(2n−i)
3nzn
(h)/summationtext∞
n=02nzn
n!(0!≡1)
(i)/summationtext∞
n=0(1
2n+i
3nzn
(j)/summationtext∞
n=0(z+i)n
22n
(k)/summationtext∞
n=0z2n
(2n)!
(l)/summationtext∞
n=0nn(z−1)n
(m)/summationtext∞
n=0z2n
4n
(n)/summationtext∞
n=0(1
n+i
n2+1)(z−√
2i)n
(2) Find the set x∈Rfor which the following series converge.
48 CHAPTER 1. INFINITE SERIES
(a)/summationtext∞
n=0(x−1)n
n2n
(b)/summationtext∞
n=0cosnx
2n
(c)/summationtext∞
n=01
n(x−1
x)n
(d)/summationtext∞
n=0e−n(x+1)
(e)/summationtext∞
n=02n(sinx)n
n
(f)/summationtext∞
n=0(1 +ex)n
(g)/summationtext∞
n=0(1−ex)n
(3) The point of this exercise is to show that a power series might converge at some, none,
or all of the points on the boundary of the disk of convergence.
(a) Show that/summationtext∞
n=0zndiverges at every point on the boundary of its disc of con-
vergence.
(b) Show that/summationtext∞
n=0zn
n+1diverges for z= 1 but converges for z=−1 (in fact, it
converges everywhere on |z|= 1 except at z= 1 ).
(c) Show that/summationtext∞
n=0xn
(n+1)2converges at every point on the boundary of its disc of
convergence.
(4) If/summationtextanzndiverges for z=ζ∈C, prove that it diverges for all z∈Cwith|z|>|ζ|.
(5) For what z∈Cdoes/summationtext∞
n=0z2
(1+z2)nconverge? Find a formula for the nth partial
sumSn(z) . Evaluate lim n→∞Sn(z) . Is the limit function continuous?
(6) Let/summationtext∞
n=0P(n)anznhave radius of convergence rho, and letP(n) be any polyno-
mial. Prove that/summationtext∞
n=0P(n)anznconverges and also has ρas its radius of conver-
gence. (By P(n) w e mean P(n) =Aknk+Ak−1nk−1+···+A1n+A0).
1.5 Properties of Functions Represented by Power Series
Having found that a power series/summationtextan(z−z0)nconverges in some disc, |z−z0|< ρ, it
is interesting to study the function f(z) defined by the power series for zin the disc of
convergence
f(z) =∞/summationdisplay
n=0an(z−z0)n, |z−z0|<ρ.
It turns out that functions f(z) defined by a convergent power series are delightful, as
nicely behaved as functions can be. In particular, they are not only continuous, but also
automatically have an infinite number of continuous derivatives and many other amazing
properties.
This section will be devoted to proving the more elementary properties of functions
represented by power series, while in the next section we will begin with given functions,
like sinx, and see if there is a convergent power series associated with the m, as well as
showing a way of obtaining the coefficients anof that power series. The profound theory
of functions represented by convergent power series is called analytic functions of a complex
variable .
1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 49
Definition: A function f(z) of the complex variable zis said to be analytic in the disc
|z−z0|<ρiff(z) can be represented by a convergent power series in that disc:
f(z) =∞/summationdisplay
n=0an(z−z0)n, |z−z0|<ρ.
Since we have not developed the notion of the derivative,d f
dz, of a complex valued
functionf(z) of the complex variable z, nor have we considered the corresponding theory
of integration,/integraltext
f(z)dz, the scope of our treatment will regrettably have to be narrowed.
However our proofs will have the property that as soon as an adequate theory of differen-
tiation and integration is given, the theorems and proofs remain unchanged.
Instead of considering power series in the complex variable z, we shall restrict our
attention to series in the real variable x
f(x) =∞/summationdisplay
n=0an(x−x0)n,|x−x0|<ρ, (1-3)
still allowing the coefficients anto be complex. Thus, f(x) is a complex-valued function
of the real variable x. The definitions of derivative and integral for such functions were
given in Section 7 of Chapter 0. We shall use that material here . Our aim is the following:
Theorem 1.16 . Suppose that/summationtext∞
n=0anxnhas radius of convergence ρ>0(possibly ∞).
Then
(a)the function f(x)defined by
f(x) =∞/summationdisplay
n=0anxn,|x|<ρ,
has an infinite number of derivatives;
(b)the series/summationtext∞
n=0nanxn−1has the same radius of convergence ρand
f/prime(x) =∞/summationdisplay
n=0nanxn−1,|x|<ρ,
and
(c)the series/summationtext∞
n=0an
n+1xn+1has the same radius of convergence ρ, and
/integraldisplayx
0f(t)dt=∞/summationdisplay
n=0an
n+ 1xn+1,|x|<ρ.
Remark: If we omit f(x) from the picture and write (b) and (c) directly in terms of the
infinite sum, we find
(b)/primed
dx[∞/summationdisplay
n=0anxn] =∞/summationdisplay
n=0nanxn−1
and
(c)/prime/integraldisplayx
0[∞/summationdisplay
n=0antn]dt=∞/summationdisplay
n=0an
n+ 1xn−1.
50 CHAPTER 1. INFINITE SERIES
These two statements are usually abbreviated “a power series may be differentiated
term by term” and “a power series may be integrated term by term” within their domain
of convergence (these statements are notgenerally true for an arbitrary infinite series of
functions/summationtextun(x) , see Exercise 4 below). The generalization to/summationtext∞
n=0an(x−x0)nis
obvious.
Our proof will be given in several parts. We begin with the Lemma 1 . Under the
hypothesis of the theorem, f(x) is continuous for all ˜ xwith|˜x|<ρ.Proof: (This is a
little dull). Given any /epsilon1>0 , we must find a δ>0 such that
|f(x)−f(˜x)|</epsilon1when |x−˜x|<δ.
Let us write fN(x) =/summationtextN
n=0anxnandRN(x) =/summationtext∞
N+1anxn, so thatf(x) =fN(x) +
RN(x) .
Observe that |f(x)−f(˜x)|=|fN(x)−fN(˜x) +RN(x)−RN(˜x)| ≤ |fN(x)−fN(˜x)|+
|RN(x)|+|RN(˜x)|.
We shall show that each of these three terms can be made </epsilon1
3by picking xclose
enough to ˜xandN-which is entirely at our disposal- large enough.
First work with RN(x) andRN(˜x) . Choose rsuch that |˜x|< r < ρ . This is to
insure that we stay away from the boundary |x|=ρwhere the series may diverge. Then/summationtext∞
n=0|anrn|converges absolutely, say to the number S. If we let SN=/summationtextN
0|anrn|,
we know that by picking Nlarge enough,/summationtextN
N+1|anrn|=S−SN</epsilon1
3. But |RN(x)|=/vextendsingle/vextendsingle/summationtext∞
N+1anxn≤/summationtext∞
N+1|anxn|/vextendsingle/vextendsingle, so that if|x|+≤r,by using the same Nfound above, we
have
|RN(x)| ≤∞/summationdisplay
N+1|anrn|=S−SN</epsilon1
3.
Since by the definition of rwe know |˜x| ≤r, this also proves that for this same
N|RN(˜x)|</epsilon1
3. Thus by restricting |x| ≤r, we have seen that both |RN(x)|and|RN(˜x)|
can be made less than/epsilon1
3.
Having fixed N, f N(x) is a polynomial -which we know is continuous. Thus there is a
δ, >0 such that
|fN(x)−fN(˜x)|</epsilon1
3when |x−˜x|<δ 1.
This shows that |f(x)−f(˜x)|</epsilon1ifxis in the intersection of the intervals |x| ≤ |˜x|<
r<ρ ) and |x−˜x|<δ 1. That there is some interval contained in both of these intervals is
easy to see since both contain all points sufficiently close to ˜ x. And the proof is completed.
As you have observed, the proof involves no new ideas but is rather technical.
With this lemma proved, we know that f(x) is continuous -and hence integrable. Thus
we can work with/integraltextx
0f(t)dt. Our next task is to prove a portion of Part (c) of Theorem
16.
Lemma 1.17 If/summationtext∞
n=0anxnhas radius of convergence ρ>0, then
∞/summationdisplay
n=0an
n+ 1xn+1=/integraldisplayx
0f(t)dtfor all |x|<ρ.
Proof: We shall show that
/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx
0f(t)dt−N/summationdisplay
n=0an
n+ 1xn+1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle(1-4)
1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 51
can be made arbitrarily small by choosing Nlarge enough. Write
f(t) =N/summationdisplay
n=0antn+∞/summationdisplay
n=N+1antn.
Then since we can integrate any finite sum term by term, we have
/integraldisplayx
0f(t)dt=N/summationdisplay
n=0an/integraldisplayx
0tndt+/integraldisplayx
0[∞/summationdisplay
n=N+1antn]dt=N/summationdisplay
n=1an
n+ 1xn+1=/integraldisplayx
0[∞/summationdisplay
n=N+1antn]dt,
so that (4) reduces to showing that
/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx
0∞/summationdisplay
n=N+1antndt/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle
can be made small by choosing Nlarge. The idea here is to apply Theorem 16 of Chapter
0. This means we need to estimate the size of the above integrand. By now you should
recognize the method. Because |x|< ρ, we can choose an rsuch that |x|< r < ρ .
Then/summationtextanrnis convergent so its terms are bounded, say M≥ |anrn|for alln, that is,
|an| ≤M
rn. Therefore, since |t|<|x|, we find the inequality
/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle∞/summationdisplay
N+1antn/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤∞/summationdisplay
N+1|an||t|n≤∞/summationdisplay
N+1M
rn|x|n.
But the last series is a geometric series whose sum is/vextendsingle/vextendsinglex
r/vextendsingle/vextendsingleNM|x|
r−x. Thus
/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle∞/summationdisplay
N+1antn/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤/vextendsingle/vextendsingle/vextendsinglex
r/vextendsingle/vextendsingle/vextendsingleNM|x|
r−x.
Applying Theorem 16 of Chapter 0, we find that
/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx
0(∞/summationdisplay
N+1antn)dt/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤/vextendsingle/vextendsingle/vextendsinglex
r/vextendsingle/vextendsingle/vextendsingleNM|x|2
r−x.
that is,/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/integraldisplayx
0f(t)dt−N/summationdisplay
0an
n+ 1xn+1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤/vextendsingle/vextendsingle/vextendsinglex
4/vextendsingle/vextendsingle/vextendsingleNM|x|2
r−x.
Since/vextendsingle/vextendsinglex
r/vextendsingle/vextendsingle<1 , we know that/vextendsingle/vextendsinglex
r/vextendsingle/vextendsingleN→0 asN→ ∞ , which completes the proof of the
lemma.
Incidentally, all we have left to prove of part c of the theorem is that the radius of
convergence of the integrated series is no larger than ρ(since the lemma shows it is at least
ρ). But this will have to wait until after
Lemma 1.18 If/summationtextanxnhas radius of convergence ρ, the series obtained by formally
differentiating term by term,/summationtextnanxn−1, has the same radius of convergence.
52 CHAPTER 1. INFINITE SERIES
Remark: This lemma does notsay that the derived series is equal to the derivative of the
function defined by the original series. It only discusses the radius of convergence, not the
relationship of the functions represented b y the two series.
Proof: Letρ1be the radius of convergence of/summationtextnanxn−1. First we show that ρ1≤ρ. If/summationtextnanxn−1converges for some fixed x, then so does/summationtextnanxn. But the terms of this last
sequence are larger than those of/summationtextanxnsince|nanxn| ≥ |anxn|. Thus by the comparison
test/summationtextanxnalso converges for that x, which shows ρ1≤ρ.
To show that ρ≤ρ1, assume/summationtextanxnconverges for some xand choose rbetween |x|
andρ,|x|<r<ρ . As in the proof of Lemma 2 we find that |an|<Mr−n. Then the terms
in the series/summationtext/vextendsingle/vextendsinglenanxn−1/vextendsingle/vextendsingleare smaller than the corresponding terms in/summationtextnM
r|x|
rn−1. By
the ratio test this last series converges, since |x|<r. Thus the derived series/summationtextnanxn−1
also converges, showing that ρ≤ρ1and completing the proof of the lemma.
Now we can complete the proof of part c of Theorem 16.
Corollary 1.19 If/summationtextanxnhas radius of convergence ρ, then the series obtained by for-
mally integrating term by term,/summationtextan
n+1xn+1also has radius of convergence ρ.
Proof: The series/summationtextanxnis the formal derivative of the series/summationtextan
n+1xn+1, and we have
just seen that these two series have the same radius of convergence.
We shall next prove part (b) of Theorem 16 as
Lemma 1.20 f(x)≡/summationtext∞
n=0anxnhas radius of convergence ρ>0then
df
dx=d
dx[∞/summationdisplay
n=0anxn] =∞/summationdisplay
n=0nanxn−1,
and this series also has radius of convergence ρ.
Proof: In Lemma 3 we proved that the radii of convergence are the same. What we must
prove here is that the derivative of the function is given by the derivative of the series.
This is a more or less immediate consequence of Lemma 2, for let us apply this integration
lemma to the function g(x) defined by
g(x)≡∞/summationdisplay
n=1nanxn−1,|x|<ρ.
Then we find that/integraldisplayx
0g(t)dt=∞/summationdisplay
n=1anxn=f(x)−a0,|x|<ρ.
By the fundamental theorem of calculus, we can take the derivative of the left side, and it
isg(x) . Thus
g(x) =f/prime(x),
that is,
∞/summationdisplay
n=1nanxn−1=d
dxf(x).
This incidentally also proves the otherwise not obvious fact that f(x) , only known to be
continuous so far (Lemma 1) is also differentiable.
1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 53
To complete the proof of Theorem 16, we must prove Lemma 5. If the power series/summationtextanxnconverges for |x|< ρ, then the function f(x) defined by f(x)≡/summationtext∞
n=0anxn
has an infinite number of derivatives. The derivatives are represented by the formal series
obtained by term-by-term differentiation.
Proof: By induction, Lemma 4 shows us that f(x) has one derivative. Assume f(x) has
kderivatives. We shall show that is has k+ 1 . Letf(k)(x) =/summationtextbnxnbe the series for the
kthderivative of f. Applying Lemma 4 to this series we find that f(k)(x) is differentiable.
This proves that fhask+ 1 derivatives and completes the induction proof.
Examples: (a) We know that
1
1 +t=∞/summationdisplay
n=0(−t)n= 1−t+t2−t3+···.
where the geometric series converges for |t|<1 . Applying the theorem, we integrate term
by term to find that
ln(1 +x) =/integraldisplayx
01
1 +tdt=∞/summationdisplay
n=0(−1)n·xn+1
n+ 1,|x|<1,
or
ln(1 +x) =x−x2
2+x3
3−x4
4+x5
5+···.
Thus the function ln(1 + x) is equal to the power series on the right. With a little
more work we can prove that the series, which converges at x= 1 , converges to ln(1 + 1)
and obtain the following interesting formula.
ln 2 = 1 −1
2+1
3−1
4+1
5− ···.
The power series for ln(1 + x) can be used to illustrate the possibilities of computing
with infinite series. If 0 <x< 1 the series for ln(1 + x) is a strictly alternating series to
which we can apply inequality (2) of Theorem 12. For this series it reads
0</vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingleln(1 +x)−k/summationdisplay
n=0(−1)nxn+1
n+ 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle<xk+2
k+ 2, x> 0.
This inequality states that if only the first kterms of the infinite series are used to compute
ln(1 +x) , the error will be less thanxk+2
k+2. Say we want to compute ln(1 +1
4) =
ln5
4to 5 decimal places. Then we want t o choose kso that
1
4k+2
k+ 2<1
1,000,000= 10−6
Cross-multiplying, writing 4 = 22, we wantksuch that 106<(k+ 2)22k+4, sincek+ 2≥
2,22k+5≤(k+ 2)22k+4. Thus, we are done if we can find ksuch that
106≤22k+5.
54 CHAPTER 1. INFINITE SERIES
But since 210= 1024>103, we know 220>106. Thus if 2 k+ 5≥20 , ork= 8 we will
have the desired accuracy. This means that
ln5
4=1
4−1
2(1
4)2+···+1
9(1
4)8+1+ error
where the error is less than 10−6.
From the form of the error estimate, it is clear that the series converges faster if xis
smaller. This power series, valid only if |x|<1 can be used to compute ln(1+ x) if|x|>1
by utilizing the observation illustrated by
ln 6 = 3 ln(3
2) + 2 ln(4
3) = 3 ln(1 +1
2) + 2 ln(1 +1
3),
where both ln(1 +1
2) and ln(1 +1
3) can be computed using the power series. We should
confess that this series converges too slowly to be of much value for that purpose in real
life.
(b) Since1
1+t2is also the sum of a geometric series
1
1 +t2= 1−t2+t4−t6+t8+···=∞/summationdisplay
0(−1)nt2n,|t|<1,
if we integrate term by term, we find
tan−1x=/integraldisplay2
0dt
1 +t2+∞/summationdisplay
n=0(−1)nx2n+1
2n+ 1=x−x3
3+x5
5+···,
which converges if |x|<1 . Further investigation shows that the series also converges
atx= 1 and represents the function at that point. This yields the wonderful formula
(obtained by letting x= 1 )
π
4= 1−1
3+1
5−1
7+···
from which we can compute πto any desired accuracy.
Exercises
(1) Write down an infinite series whose sum is1
1−tand integrate the series term by term
to obtain a power series for ln(1 −x) . For what xdoes the series converge?
(2) Find a power series which converges about x= 0 for the functionx
(1−x)2by recog-
nizing1
(1−x)2as the derivative of a function whose power series in known. For what
xdoes the series converge?
(3) Compute ln9
8to 4 decimal places, proving the error in your approximation is correct.
(4) Show that/summationtext∞
n=1sinn2x
n2converges for all xbut the series obtained by differentiating
term-by-term does not converge, say at x= 0 .
(5) Exercise your ingenuity and apply the theorems of this section to find the function
whose power series is
(a)a+ 2x2+ 4x4+ 6x6+ 8x8+···+ (2n)x2n+···.
(b) 2 + 3 ·2x+ 4·3x2+ 5·4x3+···+ (k+ 2)(k+ 1)xk+···
1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 55
6. Taylor’s Theorem. Representation of a Given Function in a Power Series. The Binomial
Theorem.
In this section we prove Taylor’s Theorem, an important generalization of the mean
value theorem, and use it to investigate the questions i) when does a given function f(x)
have a power series? and ii) if f(x) has a power series about x0, f(x) =/summationtext∞
n=0an(x−x0)n,
how can we find the coefficients an? As a partial answer to i) we know from Theorem 16
of the last section that if f(x) has a power series about x0, it must necessarily have an
infinite number of derivatives at x0. It turns out that this is not enough.
Perhaps it is easiest to begin with question ii).
Assumef(x) has a power series about x0,
f(x) =∞/summationdisplay
n=0an(x−x0)n,
which converges for |x−x0|< ρ. How can we find the coefficients an? By Theorem 16
we know that fhas an infinite number of derivatives at x0. Moreover these derivatives
can be calculated by differentiating the power series term-by-term. F or convenience we let
x0= 0 .
f(x) =a0+a1x+a2x2+a3x3+···+anxn+···,
f/prime(x) =a1+ 2a2x+ 3a3x2+···+nanxn−1+···,
f/prime/primex(x) = 2a2+ 2·3a3x+ 3·4·a4x2+···+n(n−1)anxn−2+···,
f(3)(x) = 2·3a3+ 2·3·4a4x+ 3·4·5a5x2+···+,
f(n)(x) =n!an+ (n+ 1)!an+1x+(n+ 2)
x!an+2x2+···.
By letting x= 0 in each line, we find
a0=f(0), a1=f/prime(0), a2=f/prime/prime(0)
2,...,a n=f(n)(0)
n!.
This proves
Theorem 1.21 Iff(x) =/summationtextan(x−x0)nhas a convergent power series representation
aboutx0, then the coefficients anare equal to f(n)(x0)/n!, so in fact
f(x) =∞/summationdisplay
n=0f(n)(x0)
n!(x−x0)n. (1-5)
This formula (1.21) completely solves the problem of finding the coefficients anof a function
if that function has a power series. A simple consequence is the
Corollary 1.22 A function f(x)has at most one convergent Taylor series about a point
x0.
Proof: By the above theorem, if f(x) =/summationtextan(x−x0)nandf(x) =/summationtextbn(x−x0)n, then
an=f(n)(x0)
n!=bn, so the power series are identical.
Remark: Whenfhas a power series expansion about x0, the series is usually called the
Taylor series offatx0. In the special case x0= 0 , the series is sometimes called the
Maclaurin series forf.
56 CHAPTER 1. INFINITE SERIES
Examples:
(a) Iff(x) =exhas a power series about x= 0 , what is it? Since f(n)(0) =dn
dxnex/vextendsingle/vextendsingle
x=0=
e0= 1 , we know that an−1/n! so that the power series is/summationtext∞
n=01
n!xn. We cannot
yet writeex=/summationtext∞
n=01
n!xnsince we have not proved that exdoes have a power series.
(b) Iff(x) = cosxhas a power series about x= 0 , what is it? f(0) = 1, f/prime(0) =
−sin 0 = 0, f/prime/prime(0) = −cos 0 = −1, f/prime/prime/prime(0) = sin 0 = 0 , f(4)(0) = cos 0 = 1 ,.... All the
odd derivatives at 0 are zero while the even derivatives alternate between +1 and
−1 . Therefore the series is
a−1
2!x2+1
4!x4−1
6!x6+···∞/summationdisplay
n=0(−1)nx2n
(2n)!.
Again we cannot yet claim that this is cos x.
(c) Iff(x) =/braceleftBigg
e−1
x2, x/negationslash= 0
0, x= 0/bracerightBigg
.has a power series about x= 0 what is it?
The computation is somewhat more difficult here. f/prime(x) =2
x3e−1
x2, f/prime/prime(x) = (−6
x4−
4
x6)e−1
x2, and generally f(n)(x) = (α3n
x3n+···+α2n−2
xn+2)e−1
x2where the αkare real
numbers we don’t need to find. If we let x= 0 inf(n)(x) , the resulting expression
has the indeterminate form ∞ ·0 . Thus l’Hˆ ospital’s rule must be invoked. Now
f(n)(x) is the sum of terms of the forme−1/x2
xk, k > 0 . What is lim x→0e−1/x2
xk? Let
t=1
x2, and we must evaluate lim t→0tk/2e−t= lim t→∞tk/2
et. Ifkis an even integer,
applying l’Hˆ ospital’s rule k/2 times leaves a constant in the numerator and etin the
denominator, so the limit is lim t→∞const
et= 0 . Ifkis odd, applying l’Hˆ ospital’s rule
(k+ 1)/2 times leaves a function of the formconst√
tet, which also tends to 0 as t→ ∞ .
What we have just shown is that f(n)(0) = 0 . The power series associated with e−1/x2
is
0 + 0·x+0
2!x2+···0
n!xn+··· ≡ 0.
This function e−1/x2, whose power series about x= 0 is zero, is an example of a function
which is clearly not equal to the power series, 0, associated with it.
To find if a given function has a power series expansion about x0we turn to Taylor’s
Theorem (also known as the extended mean value theorem). Now if a function fdefined
in a neighborhood of x0has a power series expansion there, we know the series is given by
(5). Thus we should investigate
RN(x)≡f(x)−N/summationdisplay
n=0f(n)(x0)
n!(x−x0)n.
To say that fis equal to its series expansion is the same as saying that the remainder,
RN(x) , becomes arbitrarily small as N→ ∞ . We must now seek an estimate of this
remainder RN(x) . Taylor’s theorem is one way of finding an estimate.
1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 57
Theorem 1.23 . (Taylor’s Theorem). Let fbe a real-valued function with N+ 1 con-
tinuous derivatives defined on an interval containing x0andx. There exists a number ζ
betweenx0andxsuch that
f(x) =f(x) +f/prime(x0)(x−x0) +f/prime/prime(x0)
2!(x−x0)2+f/prime/prime/prime(x0)
3!(x−x0)3+···
+f(N)(x0)
N!(x−x0)N+f(N+ 1)(ζ)
(N+ 1)!(x−x0)N+1. (1-6)
In other words,
RN(x) =f(N+1)(ζ)
(N+ 1)!(x−x0)N+1. (1-7)
Remark: 1 The proof will only tell us that such a ζexists but will give us no way
to find it. In practice we often try to find some upper bound Mforf(N+ 1)(ζ) , so/vextendsingle/vextendsinglef(N+ 1)(ζ)/vextendsingle/vextendsingle≤M, for allN, and only use the crude resulting estimate
|RN(x)| ≤M
(N+ 1)!|x−x0|N+1. (1-8)
An example of this is the series for cos x. Assuming the proof of the theorem, we know
that (see Example b above) about x0= 0 ,
cosx= 1−x2
2!+x4
4!+···+(−1)N
(2N)!x2N+RN(x),
where
RN(x) =1
(2N+ 2)![d2N+2
dx2N+2cosx]x=ζx2N+2, ζ/epsilon1(0,x).
Since /vextendsingle/vextendsingle/vextendsingle/vextendsingled2N+2
dx2N+2cosx/vextendsingle/vextendsingle/vextendsingle/vextendsingle
x=ζ≤1,
we find that
|RN(x)| ≤1
(2N+ 2)!|x|2N+2
Because, for fixed x, this remainder tends to 0 as n→ ∞ , we have proved that the power
series for cos xatx0= 0 does converge to cos x, so in the limit
cosx=∞/summationdisplay
n=0(−1)n
(2n)!x2n.
We can apply Theorem 16 and differentiate both sides of this to find the series for sin x.
Remark: 2 Observe that Taylor’s Theorem is only proved for real-valued functionsf. It
is not true if fis complex-valued. However using it we will be able to prove the inequality
(7) for complex-valued f.
Proof: (Taylor’s Theorem). Our proof is short—perhaps a little too slick. The trick is to
appeal to the mean value theorem (really only Rolle’s theorem is used).
58 CHAPTER 1. INFINITE SERIES
Fixxand define the real number Aby
f(x) =N/summationdisplay
n=0f(n)(x0)
n!(x−x0)n+A(x−x0)N+1
(N+ 1)!. (1-9)
Now let
H(t) :=f(x)−/bracketleftBig
f(t) +f/prime(t)(x−t) +f/prime/prime(t)(x−t)2
2!+···+f(N)(t)
N!(x−t)N/bracketrightBig
−A(x−t)N+1
(N+ 1)!.
Thus we are letting x0vary, notx. Observe that H(x) = 0 (obviously) and H(x0) = 0
(by definition of A). Since H(t) satisfies the hypotheses of the mean value theorem, we
conclude that there is some ζbetweenx0andxsuch thatH/prime(ζ) = 0 . But
H/prime(t) =−f/prime(t)−/bracketleftbig
f/prime/prime(t)(x−t)−f/prime(t)/bracketrightbig
− ··· −/bracketleftBigf(N+1)(t)
N!(x−t)N−f(N)(t)
(N−1)!(x−t)N−1/bracketrightBig
−A(x−t)N
N!=(x−t)N
N!/bracketleftbig
A−f(N+1)(t)/bracketrightbig
.
Amazingly, almost all the terms canceled. Since H/prime(ζ) = 0 and ζ/negationslash=x, we now know
thatA=f(N+1)(ζ) . Substitution of this value of Ainto (8) gives us exactly (6), which is
just what we wanted to prove.
As an application let us prove the Binomial Theorem. That is the name given to the
Maclaurin series for (1 + x)α, whereα∈R. The derivatives are easy to compute.
f(x) = (1 +x)α
f/prime(x) =α(1 +x)α−1
f/prime/prime/prime(x) =α(α−1)(1 +x)α−2
...
f(n)=α(α−1)···.(α−n+ 1)(1 +x)α−n.
Thus the power series about 0 associated formally with (1 + x)αis
∞/summationdisplay
n=0α(α−1)···(α−n+ 1)
n!xn.
By the ratio test this series converges for |x|<1 . Does it converge to (1 + x)αwhen
|x|<1 ?
Ifαis a positive integer, α=N, the terms in the power series from n=N+ 1 on all
are zero since they contain the factor ( N−N) . In this case we have only a finite series so
convergence is trivial. The resulting polynomial is the familiar Binomial Theorem of high
school algebra.
Let us therefore assume αis not a positive integer (or 0). Then we have an honest
infinite series. In order to prove that (1 + x)αis equal to the infinite series, we must show
that the remainder
RN(x)≡(a+x)α−N/summationdisplay
n=0α(α−1)···(α−n+ 1)
n!xn
1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 59
tends to zero as N→ ∞ . By Taylor’s Theorem
RN(x) =α(α−1)···(α−N)
(N+ 1)!(1 +ζ)α−N−1xN+1,
whereζis between 0 and x. We shall prove that this tends to 0 as N→ ∞ only when
0≤x<1 . It is also true for −1<x≤0 , but the proof is much longer so we will not give
it [however a different attack yields the proof easily].
Now if 0 ≤x<1 , since 0<ζ <x , then 1<1 +ζ. Therefore for N≥α, we have
(z+ζ)α−N−1<1 . Thus
|RN(x)|</vextendsingle/vextendsingle/vextendsingle/vextendsingleα(α−1)···(α−N)
(N+ 1)!xN+1/vextendsingle/vextendsingle/vextendsingle/vextendsingle
which does tend to zero as N→ ∞ (since it is the N+ 1 st term of the convergent series/summationtext∞α(α−1)···(α−n+1)
n!xn,|x|<1) .
Although we have proved it only if 0 ≤x<1 , we shall state the complete
Theorem 1.24 (Binomial Theorem). The function (1 +x)αis equal to a power series
which converges for |x|<1. It is
(1 +x)α=∞/summationdisplay
n=0α(α−1)···(α−n+ 1)
n!. (1-10)
In practice it is silly to memorize this formula since it is easier to expand (1 + x)α
directly in a Maclaurin series, which we have just shown (partly anyway) is equal to the
function.
We close this section with the generalization of Taylor’s Theorem to complex-valued
functionf(x) .
Theorem 1.25 . Letf(x) =u(x) +iv(x)be a complex-valued function with N+ 1
continuous derivatives defined on an interval containing x0andx. There exists a real
numberMNdepending on Nsuch that
/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglef(x)−N/summationdisplay
n=0f(n)(x0)
n!(x−x0)n/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle≤MN
(N+ 1)!|x−x0|N+1(1-11)
Proof: SincefhasN+ 1 continuous derivatives, so do the real-valued functions u(x)
andv(x) . Applying Taylor’s Theorem to uandv, we find numbers ζ1andζ2, both
betweenx0andx, such that
u(x)−N/summationdisplay
n=0u(n)(x0)
n!(x−x0)n=u(N+1)(ζ1)
(N+ 1)!(x−x0)N+1,
and
v(x)−N/summationdisplay
n=0v(n)(x0)
n!(x−x0)n=v(N+1)(ζ2)
(N+ 1)!(x−x0)N+1.
60 CHAPTER 1. INFINITE SERIES
Thus, by addition, since f(n)=u(n)=iv(n), we find
f(x)−N/summationdisplay
n=0f(n)(x0)
n!(x−x0)n=u(N+1)(ζ1) +iv(N+1)(ζ2)
(N+ 1)!(x−x0)N+1.
However since u(N+1)andv(N+1)are assumed continuous in an interval containing
x0andx, they are bounded there, say by ˆMNand ˜MN. Taking absolute values of the
last equation, we obtain equation (10) where MN=/radicalBig
ˆM2
N+˜M2
N.
Exercises
(1) Find the Taylor series about the specified point x0and determine the interval of
convergence for the following functions. You need not prove that the series do converge
to the functions.
(a) sinx, x 0= 0,
(b) lnx, x 0= 1,
(c)1
x, x0=−1,
(d)√x, x 0= 6,
(e)1
2(ex+e−x), x0= 0
(f)x+i
1+x, x0= 0,
(g) cosx, x 0=π
4,
(h)1
i+x, x0= 0
(i)e−x2, x0= 0,
(j) (1 +x+x2)−1, x0= 0,
(k) cosx+isinx, x 0= 0,
(l)1√1+2x, x0= 0.
(2) Prove that in their interval of convergence about 0 the following power series associated
with the given functions converge to the functions. Do this by proving that the
remainder |RN(x)| →0 asN→ ∞ .
(a) sinx,
(b)1
1+x4,
(c)e−x
(d) coshx[Recall the definition: cosh x=ex+e−x
2].
(3) One often approximates1√
1+x2by 1−x2
2when |x|is small. Give some estimate of
the error if a) |x|+<10−1, b)|x|<10−2, c)|x|<10−4.
(4) Use the Taylor series
e−x2= 1−x2+x4
2!−x6
3!+···+(−1)nx2n
n!+···.
to evaluate/integraltext1
0e−x2dxto three decimal places. I suggest using Theorem 16 and the
error estimate of Theorem 12.
1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 61
(5) Assume the ordinary differential equation y/prime−y= 0 , with y(0) = 1 has a power
series solution y(x) =/summationtext∞
n=0anxnaboutx= 0 . a). Substitute this series directly
into the differential equation and solve for the coefficients an. b). Find when the
series converges; c). justify (a posteriori) the fact that the function defined by the
convergent series does satisfy the differential equation. [We do not yet know that this
is the only solution. All we know is that it is the only solution which has a power
series].
(6) In this exercise you will prove that eis irrational. It all hinges on the series for 3 .
e= 1 + 1 +1
2+1
3!+···+1
n!+···.
(a) Prove that 2 <e< 3 , soeis not an integer (cf. page 58, bottom).
(b) Assume eis rational, e=p
q, wherepandqare integers with no common
factor and q≥2 . Then use the Taylor series with qterms and the remainder
Rqto show that e·q! =N+eζ
q+1, where 0<ζ < 1 , andNis an integer.
(c) From this deduce thateζ
q+1must be an integer, and show that this contradicts
eζ<e/prime<3 , andq+ 1≥3 .
(7) This exercise generalizes the form of the remainder (6’) in Taylor’s Theorem. Fix x
and define the number Bby
f(x) =N/summationdisplay
n=0f(n)(x0)
n!(x−x0)n+B(x−x0)α, α≥1.
Then consider the function H(t) defined by
H(t)≡f(x)−N/summationdisplay
n=0f(n)(t)
n!(x−t)n−B(x−t)α.
Show that there is a ζbetweenx0andxsuch that
B=f(N+1)(ζ)
αN!(x−ζ)N+1−α,
so that
RN=f(N+1)(ζ)
αN!(x−x0)α(x−ζ)N+1−α.
This is Schlomilch’s form of the remainder. In the special case α=N+ 1 , we obtain
Lagrange’s form of the remainder, (6) found previously, while for α= 1 we obtain
Cauchy’s form of the remainder
RN=f(N+1)(ζ)
N!(x−x0)(x−ζ)N.
Here are two applications of Taylor’s Theorem to problems other than infinite series.
The first one deals with max-min. Let f(x) be a sufficiently smooth function (by which
we meanfhas plenty of derivatives—we’ll specify the number later). Now we know that
62 CHAPTER 1. INFINITE SERIES
iffhas a local maximum or minimum at x0, thenf/prime(x0) = 0 , and it is a maximum if
f/prime/prime(x0)<0 , a minimum if f/prime/prime(x0)>0 . But what if f/prime/prime(x0) = 0 ? Consider the examples
f1(x) =x4, f2(x) =−x4, f3(x) =x3, the first of which has a minimum at x= 0 , the
second a maximum at x= 0 , while the third has neither. These three examples suggest
the criterion will depend upon the lowest non-zero derivative being an even or odd derivative,
and on its sign.
a figure goes here
By the definition of local maximum and minimum, the issue is the behavior of f(x) in
a neighborhood of x0, that is, the nature of f(x0+h) for|h|small. We remind you that
fhas a local max at x0iff(x0+h)−f(x0)≤0 for all |h|sufficiently small, and a local
min atx0iff(x0+h)−f(x0)≥0 for all |h|sufficiently small. Since the behavior of f(x)
nearx0is determined by the Taylor polynomial
f(x0+h) =f(x0) +f/prime(x0)h+f/prime/prime(x0)h2
2!+···+f(n)
n!(x0)hn+f(n+1)(ζ)hn+1
(n+ 1)!
whereζis between x0andx0+h, it is natural to look at this polynomial to answer our
question.
Theorem 1.26 Assumefhas (at least) n+ 1 continuous derivatives in some interval
containing x0. Sayf/prime(x0) =f/prime/prime(x0) =...=fn(x0) = 0 butf(n+1)(x0)/negationslash= 0, then
(a) ifnis even, then fhas neither a max nor min at x0.
(b) ifnis odd, then
i)fhas a max at x0iff(n+1)(x0)<0.
ii)fhas a min at x0iff(n+1)(x0)>0.
Proof: We shall use Taylor’s polynomial with n+ 1 terms.
Since the first nderivatives vanish at x0, we havef(x0+h)−f(x0) =f(n+1)(ζ)
(n+1)!hn+1, ζ
betweenx0andx0+h. Becausef(n+1)(x) is assumed continuous at x0, f(n+1)(ζ) must
have the same sign as f(n+1)(x0) in some neighborhood of x0. Restrict your attention to
the neighborhood. Ifnis even ,n+ 1 is odd, so that hn+1is positive if h>0 , negative
ifh<0 . Thusf(x0+h)−f(x0) changes sign in any neighborhood of x0. However ifn
is odd ,hn+1is positive no matter if h>0 orh<0 . Therefore f(x0+h)−f(x0) has the
same sign as f(n+1)(x0) throughout some neighborhood about x0. The precise conditions
are easy to verify now.
Examples:
1.f(x) =x5+ 1 has neither a max nor min at x= 0 , since f/prime(0) =...=f(4)(0) = 0 ,
butf(5)(0) = 5! /negationslash= 0 .
2.f(x) = (x−1)6−7 has a min at x= 1 since f/prime(1) =...=f(5)(1) = 0 , but
f(6)(1) = 6!>0 .
Our second application is a geometrical interpretation of the Taylor polynomial. Given
the function f(x) , consider the polynomial
Pn(x) =f(x0) +f/prime(x0)(x−x0) +f/prime/prime(x0)
2!(x−x0) +···+f(n)(x0)
n!(x−x0)n,
1.5. PROPERTIES OF FUNCTIONS REPRESENTED BY POWER SERIES 63
whose first nderivatives agree with those of fatx=x0.P1(x) =f(x0) +f/prime(x0)(x−x0)
is the equation of the tangent to the curve y=f(x) atx0. It is the straight line which
most closely approximates the curve at x0. Similarly P2(x) is the parabola which most
closely approximates the curve at x0. Generally, Pn(x) is the polynomial of degree n
which most closely approximates the curve y=f(x) at the point x0. Using this Taylor
polynomial, we can define the order of contact of two curves at a point.
Definition: The two curves y=f(x) andy=g(x) have order of contact nat the point
x0if their Taylor polynomials of degree natx0are identical, but their n+ 1 st Taylor
polynomials differ.
An equivalent definition is that f(x0) =g(x0) ,f/prime(x0) =g/prime(x0) , . . . ,f(n)(x0) =
g(n)(x0) , butf(n+1)(x0)/negationslash=g(n+1)(x0) . We have assumed that fandghaven+ 1
continuous derivatives. If fandghave contact natx0, then
f(x0+h)−g(x0+h) =f(n+1)(ζ1)−g(n+1)(ζ2)
(n+ 1)!hn+1.
One interesting consequence of this formula is that if fandghave contact of even
order, then the curves will cross at x0, while if the contact is of odd order, the curves will
notcross in some neighborhood of x0.
We can define the curvature of a curve in the plane by using the concept of contact.
First we define the curvature of a circle (whose curvature had better be constant).
Definition: Thecurvaturekof acircle of radiusRis defined to be1
R, k=1
R.
Thus the smaller the circle, the larger the curvature—a natural outcome. Furthermore,
a straight line—which may be thought of as a circle with infinite radius—has curvature zero.
How can we define the curvature of a given curve? For all non-circles, the curvature will
clearly vary from point to point of the curve. Thus, the concept we want is the curvature
of a given curve y=f(x)at a pointx0. Our definition should appear reasonable.
Definition: Thecurvaturekof a plane curve y=f(x) at the point x0is the curvature
of the circle which has contact of order two at x0.
This circle which has contact of order two is called the osculating circle to the curve
atx0(osculate: Latin, to kiss). Let us convince ourselves that there is only one osculating
circle (for if there were two, the curvature would not b e well defined.) Consider all circles
of contact one to f(x) atx0. These are all circles tangent to f(x) atx0. Their centers
lie on the line lnormal to the curve at x0(“normal” means perpendicular to the tangent
line). It is geometrically clear that of these circles with contact 1, there will be exactly one
with contact 2.
Example: Find the curvature of y=exatx= 0 . The slope of the curve at (0 ,1) is
1. Therefore the equation of the normal is y−1 =−x. Since the center ( x0,y0) of the
osculating circle must lie on this line, and the circle contains the point (0 ,1) , subject to
y0= 1−x0, the value of x0must be determined from the fact that the second derivative
of the circle (0 ,1) must equal the second derivative of y=exatx= 0 , that is, it must
equal 1. But for any circle, ( y−y0)y/prime/prime+y/prime2+ 1 = 0 . In our case y/prime= 1 at (0,1) (recall
the circle is tangent to exat (0,1) ), so that (1 −y0)·1 + 1 + 1 = 0 , or y0= 3 . The
equationy0= 1−x0implies that x0=−2 . Thus the equation of the osculating circle is
(y−3)2+ (x+ 2)2= 8 , and the curvature of y=exatx= 0 isk=1√
8. Later on we will
give another definition of curvature which is applicable not only to plane curves, but also
to curves in space.
64 CHAPTER 1. INFINITE SERIES
Exercises
(1) What is the order of contact of the curves y=e−xandy=1
1+x+1
2sin2xatx= 0 ?
(2) Find the osculating circle and curvature for the curve y=x2atx= 1 .
(3) Show that at x=a, the curve y=f(x) has curvature k=f/prime/prime(a)
[1+f/prime(a)2]3
2and the center
of the osculating circle is at the point ( a−f/prime(a)
f/prime/prime(a)[1 +f/prime(a)2], f(a) +1+f/prime(a)2
f/prime/prime(a)).What
is the messy equation of the osculating circle?
(4) At the given points, the following curves have slope zero. Determine if the curve has
a max, min, or neither there.
(a).y= (x+ 1)4, x=−1,
(b).y=x2sinx, x= 0.
(5) LetP1,P, andP2be three distinct points on the curve y=f(x) , and consider the
circle passing through those three points. Show that in the limit as both P1andP2
approachP, this circle becomes the osculating circle. (Hint: Taylor’s Theorem will
be needed here).
(6) In this problem we outline another derivation of Taylor’s Theorem. Whereas the one
in the notes did not use the fact the f(n+1)was continuous, this proof relies upon
that fact.
(a) Show that
/integraldisplayx
x0(x−t)k−1
(k−1)!f(k)(t)dt=f(k)x0(x−x0)k
k!+/integraldisplayx
x0(x−t)k
k!f(k+1)(t)dt.
(b) Prove by induction that
f(x) =f(x0)+f/prime(x0)(x−x0)+···+f(n)(x0)
n!(x−x0)n+/integraldisplayx
x0(x−t)n
n!f(n+1)(t)dt.
The remainder is expressed as an integral here. It is because f(n+1)is to be
integrated that we require its continuity.
(7) (a) Let g(x) have contact of order nwith the function 0 at the point x=a, and
assume that f(x) has contact of order at least nwith the function 0 at x=a.
Use Taylor’s Theorem to prove that
lim
x→af(x)
g(x)=f(n+1)(a)
g(n+1)(a)
This is l’Hˆ ospital’s Rule .
(b) Apply l’Hˆ ospital’s rule to evaluate
i) lim
x→0x−sinx
x3,ii) lim
θ→π
41−tanθ
θ−π
v
1.6. COMPLEX-VALUED FUNCTIONS, EZ,COSZ,SINZ. 65
(8) Assume fhas two derivatives in the interval [ a,b] , and assume that f/prime/prime≥0 through-
out the interval. Prove that if ζis any point in [ a,b] , then the curve y=f(x) never
falls below its tangent at the point x=ζ, y=f(ζ) . [hint: Use Taylor’s Theorem
with three terms].
(9) Use Cauchy’s form of the remainder (p. 103-4, no. 7) for Taylor’s Theorem to prove
that the binomial series converges to (1 + x)αfor−1< x≤0 . This will complete
the proof of the binomial theorem.
(10) ThenthLegendre polynomial Pn(x) is defined by Pn(x) =1
2nn!dn
dxn[(x2−1)n] . Prove
thatPn(x) is a polynomial of degree nand hasndistinct real zeros in the interval
(−1,1) .
(11) Verify that eaxis a solution of y/prime=ay. Prove that every solution has the form
Aeax, whereAis a constant.
(12) Assume that f(x) has plenty of derivatives in the interval [ a,b] , and that fhas
n+ 1 distinct zeros in the interval. Prove that there is at least one c∈(a,b) such
thatf(n)(c) = 0 .
1.6 Complex-Valued Functions, ez,cosz,sinz.
The task of this section is to answer the following question. Say f(x) is a real or complex
valued function of the realvariablex. How can we define f(z) whereziscomplex ?
For example, if P(x) =a0+a1x+···+anxnis a polynomial, the answer is easily given:
just define P(z) =a0+a1z+···+anzn. Since this function only involves addition and
multiplication of complex numbers, for any complex zthe number P(z) can be computed.
Similarly any rational function,P(x)
Q(x), whereP(x) andQ(x) are both polynomials, can be
defined for complex zasP(z)
Q(z)since both P(z) andQ(z) are defined separately and we
can then take their quotient.
But how do we define ez, or cosz, or (1 +z)α, whereα/epsilon1Ris not a positive integer?
As might have been suspected, the trick is to use infinite series.
Definition: Iff(x), x/epsilon1R, has a convergent Taylor series,
f(x) =∞/summationdisplay
n=0anxn, |x|<ρ,
then we define f(z), z/epsilon1C, by the infinite series
f(z) =∞/summationdisplay
n=0anzn,
and the infinite series converges throughout the disc |z|<ρ.
The assertion that the complex series converges throughout the disc |z|< ρ is an
immediate consequence of Theorem 13 on page ?.
Thus, for example, we define .
E(z) =∞/summationdisplay
n=01
n!zn,
66 CHAPTER 1. INFINITE SERIES
C(z) =∞/summationdisplay
n=0(−1)nz2n
(2n)!,
S(z) =∞/summationdisplay
n=0(−1)nz2n+1
(2n+ 1)!,
and
(1 +z)α=∞/summationdisplay
n=0α(α−1)···(α−n+ 1)
n!zn, α/epsilon1 R
where the first three series converge for all z/epsilon1C, while the last converge for |z|<1 . We
have temporarily used the notation E(z) in place of ez, C(z) for cosz, andS(z) for sinz
so that you do not jump to hasty conclusion s about these functions by merely extrapolating
your knowledge of exetc. For example it is nottrue that |sinz| ≤1 for allz/epsilon1C, even
though |sinx| ≤1 for allx/epsilon1R. All properties of these function s for z/epsilon1Cmust be proved
again beginning with the power series definitions. Known properties of ex, x/epsilon1Rand wishful
thinking don’t prove properties of ez, z/epsilon1C. Let us begin by proving
Theorem 1.27 .
(a)E(iz) =C(z) +iS(z),for allz∈C.
(b)E(−iz) =C(z)−iS(z), for allz∈C.
(c)C(z) =1
2[E(iz) +E(−iz)], for allz∈C.
(d)S(z) =1
2i[E(iz)−E(−iz)], for allz∈C.
Proof: a). b). Just substitute and rearrange the series. For example
C(z) = 1−z2
2!+x4
4!−x6
6!+···.
iS(z) =i[z−z3
3!+z5
5!−z7
7!+···]
so
C(z) +iS(z) = 1 +iz−z2
2!−iz3
3!+z4
4!+iz5
5!− ···,
where the adding of the two series is justified by Theorem 5(page ?). We must compare the
last series with that for E(iz) :
E(iz) = 1 +iz+(iz)2
2!+(iz)3
3!+(iz)4
4!+···= 1 +iz−z2
2!−iz3
3!+z4
4!+···,
which is identical to the series for C(z) +iS(z) .
c)-d). These follow by elementary algebra from a) and b).
The formulas a)-d) of Theorem 21 show there is a close connection between the four
functionsE(iz), E(−iz), C(z),andS(z) . Our next theorem shows that the formula
exey=ex+y, x, y∈R, extends to the function E(z) .
Theorem 1.28 .E(z)E(w) =E(z+w), for allz, w∈C.
1.6. COMPLEX-VALUED FUNCTIONS, EZ,COSZ,SINZ. 67
Proof: We must show that
(∞/summationdisplay
n=0zn
n!)(∞/summationdisplay
n=0wn
n!) =∞/summationdisplay
n=0(z+w)n
n!,
The product of the two series is defined in Theorem 15. Using that definition, we find that
(∞/summationdisplay
n=0zn
n!)(∞/summationdisplay
n=0wn
n!) =∞/summationdisplay
n=0(n/summationdisplay
k=0zk
k!wn−k
(n−k)!).
However, the binomial theorem for positive integer exponents (which only uses the
algebraic rules for complex numbers) states that
(z+w)n=n/summationdisplay
k=0n!
k!(n−k)!zkwn−k.
Upon substituting this into the last equation, we obtain the desired formula.
The formula of this theorem is the key to many results, like the following generalization
of sin2x+ cos2x= 1 .
Corollary 1.29 C(z)2+S(z)2= 1 for allz∈C.
Proof: We use equations a) and b) of Theorem 21 to reduce the question to one of
exponentials.
E(iz)E(−iz) = [C(z) +iS(z)][C(z)−iS(z)] =C2(z) +S2(z).
But by Theorem 22, E(iz)E(−iz) =E(iz−iz) =E(0) . Directly from the power series we
see thatE(0) = 1 . This proves the formula.
Our next corollary states that the addition formulas for sin xand cosxare still valid
forC(z) andS(z) .
Corollary 1.30 C(z+w) =C(z)C(w)−S(z)S(w)andS(z+w) =S(z)C(w)−S(w)C(z)
for allz, w∈C
Proof: A direct algebraic computation does the job.
C(z+w) +iS(z+w) =E(iz+iw) =E(iz)E(iw) = [C(z) +iS(z)][C(w) +iS(w)]
= [C(z)C(w)−S(z)S(w)] +i[S(z)C(w) +S(w)C(z)].
Similarly we find that
C(z+w)−iS(z+w) = [C(z)C(w)−S(z)S(w)]−i[S(z)C(w) +S(w)C(z)].
Addition of these two equations gives the formula for C(z+w) , while subtraction gives the
formula for S(z+w) .
Had we but world enough, and time, we would linger a while. A lovely result we have
not proved is that E(z+ 2πi) =E(z) , the periodicity of E(z) , which is a consequence of
68 CHAPTER 1. INFINITE SERIES
the formulas C(z+ 2π) =C(z) , andS(z+ 2π) =S(z) , the periodicity of C(z) andS(z) ,
by using Theorem 21 (but see pp. ??).
We shall close this chapter by restating the results proved above in the usual language
ofezetc. instead of the temporary notation E(z) etc. we have been using.
eiz= cosz+isinz (1-12)
e−iz= cosz−isinz (1-13)
cosz=1
2(eiz+e−iz) (1-14)
sinz=1
2i(eiz−e−iz) (1-15)
ezew=ez+w(1-16)
sin2z+ cos2z= 1 (1-17)
cos(z+w) = coszcosw−sinzsinw (1-18)
sin(z+w) = sinzcosw+ sinwcosz (1-19)
Generally, all algebraic formulas for sin x,cosx, andexremain valid for sin z,cosz, and
ez. In fact any algebraic relationship between any combination of analytic functions remains
valid as we change the in dependent variable from a real xto the complex z. Inequalities
almost always fall apart in the transition from x∈Rtoz∈C. Exercise 2e below illustrates
this.
One formula which we will use frequently later on is a specialization of (1-12) to the
case when zis real. Then writing the real zasθwe have the famous formula
eiθ= cosθ+isinθ, θ∈R. (1-20)
We cannot resist stating this formula down again for θ=π:
eiπ=−1,
an almost mystical identity connecting the four numbers e, iπ , and −1 . Notice that (1.6)
also implies/vextendsingle/vextendsingleeiθ/vextendsingle/vextendsingle= 1 .
If we write z=x+iy, then using (1.6) we find
ez=ex+iy=exeiy=ex(cosy+isiny). (1-21)
A consequence of this is
|ez|=ex(1-22)
Exercises
(1) Observe that (directly from the power series)
cos(−z) = cosz,and sin( −z) =−sinz.
Use this and the addition formula for cos( z+w) to prove that sin2z+ cos2z= 1 .
(2) If we define sin hx=1
2(ex−e−x) and coshx=1
2(ex+e−x), x∈R, we prove that
1.6. COMPLEX-VALUED FUNCTIONS, EZ,COSZ,SINZ. 69
(a) cosix= coshx,sinix=isinhx
(b) cosz= coshy−isinxsinhy,(z=x+iy)
sinz= sinxcoshy+icosxsinhy
(c)|cosz|2= cos2x+ sinh2y
|cosz|2= cosh2y−sin2x
|sinz|2= sin2x+ sinh2y
|sinz|2= cosh2y−cos2x
(d) Use the identities of part c) to deduce that
|sinhy| ≤ |cosz| ≤coshy
|sinhy| ≤ |sinz| ≤coshy
(e) Prove that there is some z∈Csuch that
|sinz|>1,and|cosz|>1.
(3) Define the derivative of f(z) atz0, wherez, z 0∈C, as
lim
z→z0f(z)−f(z0)
z−z0,
if the limit exists.
(a) By working directly with the power series, show that ezis differentiable for all
z, and that
d
dzeaz=aeaz, a, z∈C,
(b) Apply this to (1-12) and (1-13) to deduce that
d
dzcosz=−sinz,d
dzsinz= cosz
(We cannot appeal to Theorem 16 and differentiate term-by-term since that
theorem assumed the independent variable, x, was real).
(4) Use the results of Exercise 2c to show that the only complex roots z=x+iyof sinz
and coszare at the points on the real axis y= 0 where sin x= 0 and cos x= 0 ,
respectively.
(5) Use the results of this section to prove DeMoirve’s Theorem
(cosθ+isinθ)n= cosnθ+isinnθ, θ ∈R,
wherenis a positive integer.
(6) (a) Show that the sum of the finite geometric series/summationtexteinxis
N/summationdisplay
n=1einx=ei(N+1/2)x−eix/2
eix/2−e−ix/2.
70 CHAPTER 1. INFINITE SERIES
(b) Take the real and imaginary parts of the above formula and prove that for all
x/negationslash= 0, x∈(0,2π) ,
N/summationdisplay
n=1cosx=sin(N+ 1/2)x−sin 1/2x
2 sin1
2x
N/summationdisplay
n=1sinnx=cos1
2x−cos(N+1
2)x
2 sin1
2.
1.7 Appendix to Chapter 1, Section 7.
As a special dessert let us take some time out and prove some interesting results you would
probably never see otherwise. We have in mind to define a specific number α∈Ras the
smallest positive zero of cos x, x∈R—soαhad better turn out as π/2 . Then we prove
that 1) sin( x+ 4a) = sinxetc., 2) the ratio of the circumference to diameter of a circle
is 2αso that 2αdoes equal the πof public school fame. Furthermore, we also present a
way of computing α.
In this section we take sin zand cosz, z∈Cto be defined by their power series,
and use only the properties of these functions which were obtained from the power series
definition.
Lemma 1.31 The setA={x∈R: cosx= 0,0< x < 2}is not empty, that is, the
equation cosx= 0 has at least one real root for x∈(0,2).
Proof: Since cosxis defined by a convergent power series, it is continuous (even infinitely
differentiable); furthermore because x∈Rand the power series has real coefficients, we
know that cos x, x∈Ris real-valued. Observe that cos 0 = 1 >0 , and the following crude
inequality
cos 2 = 1 −22
1·2+∞/summationdisplay
n=2(−1)n22n
(2n)!<−1 +∞/summationdisplay
n=222n
(2n)!
<−1 +24
4!∞/summationdisplay
k=0(2
5)2k=−1 +50
63<0.(1-23)
Thus cos 0 >0 and cos 2 <0 , so there is at least one point in (0 ,2) where the real-valued
continuous function cos xvanishes. This proves the lemma.
Denote the g.l.b of A(which does exist since Ais bounded—say by 0 and 2) by α.
We shall show that α∈A. Sinceαis the g.l.b. of A, there exists a sequence of points
αk∈A(theαkmay just be the same point repeated over and over) such that αk→α
and cosαk= 0 . But since cos xis continuous,
0 = lim
k→∞cosαk= cosα,
so in fact cos α= 0 too ⇒α∈A.
Now cosxmust be positive throughout the interval [0 ,α) , since it is positive at x= 0
andαis the first place it vanishes. Therefore the formulad
dxsinx= cosx—obtained by
differentiating the realpower series for sin xterm by term—shows that sin xis increasing
1.7. APPENDIX TO CHAPTER 1, SECTION 7. 71
forx∈[0,α) . Since sin 0 = 0 , we see that sin x≥0forx∈[0,α) . Thus the formula
d
dxcosx=−sinxtells us that cos xis decreasing in the interval [0,α] . From the formula
1 = sin2α+ cos2α= sin2α,
and the fact that sin α > 0 , we find that sin α= 1 . We can thus conclude from the
addition formulas for sin xand cosxthe:
Theorem 1.32 Letαdenote the smallest zero of cosxforx>0. Then
cosα= 0,cos 2α=−1,cos 3α= 0,cos 4α= 1
sinα= 1,sin 2α= 0,sin 3α=−1,sin 4α= 0,
or more generally
cos(z+α) =−sinz,sin(z+α) = cosz
cos(z+ 4α) = cosz,sin(z+ 4α) = sinz
This proves that the sinzand coszare periodic with period 4α.
As you have guessed, αis another name for π/2 —and serves as our definition of π.
This is based upon power series and is independent of circles or triangles—or even the entire
concept of angle. A simple consequence is the
Corollary 1.33 The function ezis periodic with period 4αi,
ez+4αi=eze4αi=ez.
Proof:ez+4α=eze4iα=ez(cos 4α+isin 4α) +ez(1 +i0) =ez.
Two issues remain to be settled before closing up. We should 1) prove that the ratio
of the circumference Cof a circle to its diameter Disπ, i.e.,C= 2αD, and 2) find
some way of approximating αnumerically (for all we know of alpha so far is that it is
the smallest element in a set and 0 <α< 2 ). The two problems are closely related.
The circle of radius Rhas the equation x2+y2=R2. Consider the portion in the
first quadrant. Then using the familiar formulas for arc length, we find that
C
4=R/integraldisplayR
0dx√
R2−x2=R/integraldisplay1
0dt√
1−t2,
where the change of variable x=Rthas been used to obtain the last integral [this is legal
since the mapping “multiply by R” is a bijection and hence an invertible function]. Thus,
the desired result, C= 2αD= 4αRwill be proved if we can p rove
Theorem 1.34/integraltext1
0dt√
1−t2=α(=π
2)
Corollary 1.35 IfCdenotes the arc length of the circumference of a circle of radius R,
thenC= 4αR.
72 CHAPTER 1. INFINITE SERIES
Proof: of Theorem. We want to make the change of variable t= sinζ, wheret∈[0,1] . In
order to do this we must only check that the function sin ζis differentiable and invertible
function there. We know it is differentiable . Since sin xis continuous and monotone
increasing for x∈[0,α] , and since the end points are mapped into 0 and 1 respectively
( sin 0 = 0,sinα= 1 ), the function f(ζ) = sinζis invertible for x∈[0,α]⇐⇒t∈[0,1] .
The usual formulas are applicable and yield
/integraldisplay1
01√
1−t2dt=/integraldisplayα
0dζ=α
Q.E.D.
To compute π= 2α, it is convenient to introduce tan z= sinz/cosz, for allzwhere
cosz/negationslash= 0 . In particular tan xis defined for all real xin the interval 0 ≤x<α/ 2 . From
the behavior of sin xand cosxin the interval x∈[0,α/2) , it is easy to show that tan x
has infinitely many derivatives and is increasing for x∈[0,α/2) , assuming the values from
0 = tan 0 to 1 = tanα
2. The function tan xis therefore invertible in that interval, so we
can make the natural change of variable t= tanxand obtain
/integraldisplay1
0dt
1 +t2=/integraldisplayα/2
01
1 + tan2x(d
dxtanx)dx=/integraldisplayα/2
0dx=α
2.
But the integral on the left can be approximated readily because of the algebraic identity
1
1 +t2=N/summationdisplay
0(−1)nt2n+(−1)N+1t2N+2
1 +t2,allt/negationslash=i.
Thus
π
4=α
2=/integraldisplay1
0dt
1 +t2=N/summationdisplay
0(−1)n/integraldisplay1
0t2ndt+ (−1)N+1/integraldisplay1
0t2N+2
1 +t2dt,
or
π
4= 1−1
3+1
5−1
7+···+(−1)N
2N+ 1+RN,
where since 2 t≤1 +t2the remainder RNcan be estimated by
|RN|=/integraldisplay1
0t2N+2
1 +t2dt</integraldisplay1
0t2N+2
2tdt=1
4N+ 4
If the first 250 terms in the series are used, N= 250 , we find
π
4= 1−1
3+1
5− ··· +1
251+R250,
where |R250|<1
1004<1
1000, so three decimal accuracy is obtained. This is quite slow—but
it does work. For practical computations, a series which converges much faster is needed.
See exercise 2 below; it is neat.
SinceRN→0 asN→ ∞ , the following formula is a consequence of our effort:
π
4= 1−1
3+1
5−1
7+1
9− ···..
Exercises
1.7. APPENDIX TO CHAPTER 1, SECTION 7. 73
(1) Use the method illustrated here to slow that
ln 2 =/integraldisplay1
01
1 +xdx= 1−1
2+1
3−1
4+1
5− ··· +(−1)N+1
N+RN,
where lim N→∞RN= 0 . Find an Nsuch that |RN|<10−3.
[Hint: Write1
1+x=/summationtextN
0(−1)nxn+(−1)N+1xN+1
1+x, x/negationslash=−1].
(2) To approximateπ
4with fewer terms, the following clever device works. Write
1
1 +t2=N−1/summationdisplay
0(−1)nt2n+(−1)Nt2N
2+ ((−1)Nt2N
2+(−1)N+1t2N+2
1 +t2)
and show that
π
4= 1−1
3+1
5+···+(−1)N−1
2N−1+(−1)N
2(2N−1)+˜RN,
where ˜RN+(−1)N
2/integraltext1
0t2N−t2N+2
1+t2dt.
(a) Prove that/vextendsingle/vextendsingle/vextendsingle˜RN/vextendsingle/vextendsingle/vextendsingle<1
8N2+8N.
(b) What should Nbe to make/vextendsingle/vextendsingle/vextendsingle˜RN/vextendsingle/vextendsingle/vextendsingle<10−3? Amazing saving, isn’t it? The
technique does generalize to other series and can be refined to yield even better
results.
(c) Apply the method given here to problem 1 above to show that ln 2 = 1 −1
2+
1
3+···+(−1)N
N−1+1
2(−1)N+1
N+˜RN, where/vextendsingle/vextendsingle/vextendsingle˜RN/vextendsingle/vextendsingle/vextendsingle<1
(2N+1)(2 N+3). PickNso that/vextendsingle/vextendsingle/vextendsingle˜RN/vextendsingle/vextendsingle/vextendsingle<10−3.
74 CHAPTER 1. INFINITE SERIES
Chapter 2
Linear Vector Spaces: Algebraic
Structure
2.1 Examples and Definition
In order to develop intuition for linear vector spaces, a slew of standard examples are needed.
From them we shall abstract the needed properties which will then be stated as a set of
axioms.
a) The Space R2.
We begin by informally examining a space of two dimensions (whatever that means). It is
constructed by taking the Cartesian Product of Rwith itself. We are thus looking at R×R,
which is denoted by R2. A pointXin this space is an ordered pair, X= (x1,x2) , where
x1∈R, x2∈R.x1andx2are called the coordinates orcomponents of the point x. Let
us propose a reasonable algebraic structure on R×R. IfX= (x1, x2) , andY= (y1, y2)
are any two points, and αis any real number, we define
addition:X+Y= (x1+y1,x2+y2) .
multiplication by scalars: α·X= (αx1,αx 2), α∈R.
equality:X=Y⇐⇒x1=y1, x2=y2
The addition formula states that the parallelogram rule is used to add points, whereas
the second formula states that a point Xis “stretched” by αby stretching each coordinate
byα.
Some immediate consequences of the above definitions are, for all X, Y, Z inR×R,
(1) addition is associative (X+Y) +Z=X+ (Y+Z)
(2) addition is commutative X+Y=Y+X
(3) There is an additive identity , 0=(0,0) with the property that X+ 0 =Xfor anyX.
(4) Every X= (x1,x2)∈R×Rhas an additive inverse (−x1,−x2) , which we denote
by−X. ThusX+ (−X) = 0 . Thus the set of points in R×Rforms an additive
abelian group.
The following additional properties are also obvious, where αandβare arbitrary
real numbers.
75
76 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE
(5)α(βX) = (αβ)X
(6) 1 ·X=X.
and the two distributive laws.
(7) (α+β)X=αX+βX
(8)α(X+Y) =αX+αY.
To insure that you too feel these properties are obvious, let us prove, one, say 7.
(α+β)·X= (α+β)·(x1,x2) = ((α+β)x1,(α+β)x2)
= (αx1+βx1,αx 2+βx2) = (αx1,αx 2) + (βx1,βx 2)
=α·(x1,x2) +β·(x1,x2) =α·X+β·X(2-1)
Example: IfX= (2,1) , then 3X= (6,3) and −2X= (−4,−2) .
Instead of thinking of the elements ( x1,x2) inR2as points, it is sometimes useful to
think of them as directed line segments, from the origin (0,0) directed to the point ( x1,x2) .
The figure at the right illustrates this.
Note that the axes need not be perpendicular to each other in the space R2. They
could just as well veer off at some outrageous angle, as in the diagram. This is because we
have yet to place a metric (distance) structure on R2or introduce any concept of angle
measurement. When we do that, we will have Euclidean 2-space E2. But right now all we
have is R2, which might be thought of as a floppy Euclidean space.
b) The Space Rn
.
This is a simple-minded generalization of R2. A pointXinRn=R×...×Ris an
orderedntuple,X= (x1,x2,...,x n) of real numbers, xk∈R. IfX= (x1,...,x n) and
Y= (y1,...,y n) are any two points in Rn, andαis any real number, we define
addition :λ+Y= (x1+y1,x2+y2,...,x n+yn)
multiplication by scalars :α·X= (αx1,αx 2,...,αx n), α∈R.
equality :X=Y⇐⇒xj=yjfor allj.
Example: The pointX= (1,2,3) , and1
2X= (1
2,1,3
2) inR3are indicated in the figure.
Again the coordinate axes need not be mutually perpendicular.
Properties 1-8 listed earlier remain valid - and with the proofs essentially unchanged
(just add dots inside the parentheses).
Remark: . At this stage, you probably are anxiously waiting for us to define multiplication
inRn, that is, the product of two points in Rn,X·Y=Z∈Rn, possibly using the
multiplication of complex numbers (points in R2) as a guide. Well, we would if we could.
It turns out that it is possible to define such a multiplication only inR1,R2,R4, and in
R8–but in no others . This is a famous theorem. In R2ordinary complex multiplication
does the job. To do it in R4, we have to abandon the commutative law for multiplication.
The result is called quaternions. InR8, the multiplication is neither commutative nor
associative. The result there is the Cayley numbers .
Here we shall not have time to treat this issue. All we shall do (later) is introduce a
“pseudo multiplication” in R3—the so called cross product - obtained from the quaternion
2.1. EXAMPLES AND DEFINITION 77
algebra in R4. The major importance of this pseudo multiplication which holds only in R3
is the fact of life that our world has three space dimensions. This multiplication is extremely
valuable in physics.
c) The Space C[a, b].
Our next example is of an entirely different nature, it is a space of functions, a function
space . The space C[a,b] is the set of all real-valued functions of a real variable xwhich
are continuous for x∈[a,b] . Iffandgare continuous for x∈[a,b] , that is if fand
g∈C[a,b] , and ifαis any real number, we define, in the usual way,
addition: ( f+g)(x) =f(x) +g(x) ,
multiplication by scalars: ( αf)(x) =α[f(x)]. α∈R
equality:f=g⇐⇒f(x) =g(x) for allx∈[a,b].
Notice that the sum of two functions in C[a,b] is again in C[a,b] , and the product of
a continuous function - in C[a,b] —by a constant αis also an element of C[a,b] . We shall
ignore the fact that the product of two continuous functions is also a continuous function.
Properties 1-8 listed earlier are also valid here, that is, if f,g, andhare any elements
inC[a,b] , then
(1)f+ (g+h) = (f+g) +h
(2)f+g=g+f
(3)f+ 0 =f
(4)f+ (−1)f= 0
(5)α(βf) = (αβ)f
(6) (1)f=f 1∈R
(7) (α+β)f=αf+βf
(8)α(f+g) =αf+αg.
Again, 1-4 state that the elements of C[a,b] form an abelian group with the group
operation being addition. When we define the dimension of a vector space, it will turn
out that the space C[a,b] isinfinite dimensional, but don’t let that bother you. This nice
space,C[a,b] , and Rnare the two most useful examples of a vector space.
d) D. The Space Ck[a, b].
The space Ck[a,b] consists of all real-valued functions f(x) which have kcontinuous
derivatives for xin the interval [ a,b]⊂R. Whenk= 0 , this reduces to the space C[a,b] .
Addition and scalar multiplication are defined just as in C[a,b] . The key property is that
the sum of two functions with kcontinuous derivatives of x∈[a,b] is also a function with
kcontinuous derivatives. All of properties 1-8 are valid in Ck[a,b] .
Every function f(x) which has one continuous derivative is necessarily continuous.
This is a basic result from elementary calculus; it may be written as C1[a,b]⊂C[a,b] .
Since the function |x|, x∈[−1,1] is inC[−1,1] but not in C1[−1,1] , we see that C1and
Care not the same, that is C1is a proper subset of C. Similarly, Ck+1[a,b]⊂Ck[a,b]
(see Exercise 7).
78 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE
The space C∞[a,b] consists of all functions with an infinite number of continuous
derivatives for x∈[a,b] . All functions which have a convergent Taylor series for x∈[a,b]
are inC∞[a,b] . In addition, C∞[a,b] contains functions like f(x) =e−1/x2, x/negationslash= 0, f(0) =
0 , which have an infinite number of continuous derivatives (see p. ??) but do not have
convergent Taylor series.
Another example of a function space is the set of analytic functions A(z0,R) , functions
which have a convergent Taylor series in the disc with center at z0∈Cand radius at least
R.
e) E. The Space l1.
The spacel1(tired yet?) consists of all infinite sequences X= (x1,x2,x3,...) which satisfy
the condition∞/summationdisplay
n=1|xn|<∞. Addition and multiplication by scalars are defined in a natural
way. IfXandYare inl1, then
X+Y+ (x1+y1,x2+y2,x3+y3,...)
and, ifαis any complex number
α·X= (αx1,αx2,...).
Equality is defined by
X+Y⇐⇒xj=yjfor allj.
We should show that if XandYare inl1, then so is X+Yandx·X. To prove that
X+Y∈l1, we must show that/summationtext|xn+yn|<∞. But since |xn+yn| ≤ |xn|+|yn|, we
have for any N∈Z+
N/summationdisplay
n=1|xn+yn| ≤N/summationdisplay
n=1|xn|+N/summationdisplay
n=1|yn| ≤∞/summationdisplay
n=1|xn|+∞/summationdisplay
n=1|yn|<∞.
Now letting N→ ∞ on the left, we see that∞/summationdisplay
n=1|xn+yn|<∞. IfX∈l1, it is obvious
thatα·Xis also inl1since
∞/summationdisplay
n=1|αxn|=∞/summationdisplay
n=1|α||xn|=|α|∞/summationdisplay
n=1|xn|<∞.
f) F. The Space L1[a, b].
Yes, the space L1[a,b] does consist of all functions f(x) (possibly complex-valued) with
the property that/integraltextb
a|f(x)|dx<∞. It is the integral analogue of l1. Addition and scalar
multiplication are defined as in C[a,b] , that is, as usual. If fandgare inL1[a,b] , then
so aref+gandαf, whereα∈C, since
/integraldisplayb
a|f(x) +g(x)|dx≤/integraldisplayb
a|f(x)|dx+/integraldisplayb
a|g(x)|dx<∞,
2.1. EXAMPLES AND DEFINITION 79
and
/integraldisplayb
a|αf(x)|dx=|α|/integraldisplayb
a|f(x)|dx<∞.
For example, f(x) =xis inL1[0,1] butf(x) =1
x2isnotinL1[0,1] . It is simple to check
that properties 1-8 are satisfied in L1[a,b] .
g) G. The Space fn.
IfP(x) =a0+a1x+...+anxnis any polynomial of degree nwith real coefficients and
Q(x) =b0+b1x+...+bnxnis another one, then with ordinary addition, multiplication by
real scalars and equality the set fnof all polynomials of degree nsatisfy conditions 1-8.
Since
a0+a1x+...+an−1xn−1=a0+a1x+...+an−1xn−1+ 0xn,
it is clear that fn−1⊂fn.
Enough examples for now. You must have gotten the point. We shall meet more later
on. Let us give the abstract definition of a linear vector space.
Definition: . LetSbe a set with elements X, Y, Z,... andFbe a field with elements
α,β... . The setSis alinear vector space (linear space, vector space )over the field Fif the
following conditions are satisfied.
For any two elements X, Y∈S, there is a unique third element X+Y∈S, such that
(1) (X+Y) +Z=X+ (Y+Z);
(2)X+Y=Y+X;
(3) There exists an element 0 ∈Shaving the property that 0 + X=Xfor allX∈S;
(4) for every X∈S, there is an element −X∈S; such that X+ (−X) = 0 .
Furthermore, if αis any element of the field F, there is a unique element αX∈S
such that, for any α,β∈F,
(5)α(βX) = (αβ)X;
(6) 1 ·X=X.
The additive and field multiplicative structures are related by the following distribu-
tive rules
(7) (α+β)X=αX+βX
(8)α(X+Y) =αX+αY.
Elements of the field Fare called scalars , whereas elements of Sare called vectors .
We shall usually take the real numbers Rfor our field F, although the complex numbers
Cwill be used at times. Exercise 4 shows the need for Axiom 6 (in case you thought it was
superfluous).
All of the examples of this section are linear spaces. For most purposes the simple
example R2will serve you well as a guide to further expectations. The pictures there are
simple. In fact, with a certain degree of cleverness, the “right” proof for R2immediately
generalizes to all other linear spaces - even “infinite dimensional” ones.
80 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE
Since you probably think that everything is a linear space, here is an example to dispel
the delusion. Let Sbe the subset of all functions f(x) inC[0,1] which have the property
f(0) = 1 . Then if fandgare inS, we are immediately stuck since f(0) +g(0) = 2 , so
thatf+gisnotinS. Also, 0 /negationslash∈S.
Both here, and before (p.?) when defining a field, axioms “0” have been used. They all
express roughly the same concept. We have some set Sand an operation * defined on the
set. These axioms all stated that for any x,y∈S, we also have x∗y∈S. In other words,
the setSisclosed under the operation * in the sense that performing that operation does
not take us out of the set. We shall find this concept useful.
h) Appendix. Free Vectors
One more example is needed, an exceedingly important example. There are “physicists’
vectors” or free vectors . I always thought they were easy to define - until today. Twelve
hours and fifty pages later, I begin again on the fifth attempt. The essential idea is easy to
imagine but difficult to convey in a clear and precise exposition.
Say you are given two elements XandYofRn, which we represent by directed line
segments from the origin. Somehow we want to find a directed line segment Vfrom the
tip ofXtothe tip ofY. NowV“looks” like a vector. The problem is that all of the
vectors we have met so far have been directed line segments in Rnbeginning at the origin.
In order to find a way out, it is best to examine the problem for the most simple case
−R1, the ordinary line. Watch closely since we will be so shrewd that all the formalism
will be adequate without change for the general case of Rn.
We are given two points, XandYofR1which we shall represent by directed line
segments from the origin. To make the picture clear, we will draw them slightly above the
line.
a figure goes here
We want a directed line segment Vfrom the tip of Xto the tip of Y. Of course you
recognize this as the problem of solving
X+V=Y
The solution, V=Y−X, is the difference of the two real numbers YandX. But
where should we draw V? If we are stubborn and demand that all real numbers must be
represented by line segments beginning at the origin, we have the picture
a figure goes here
but what we really want to do is place the tail of Vat the tip of Xand add the line
segments. Why not relent and allow ourselves this added flexibility.
a figure goes here
There! Now we have solved our problem. But we have made an important generalization
in doing so. You see, this Vhas been released from its bondage to the origin and is now
free to move along the whole of R.
Although we were led to this Vfrom the pair XandY, the sameVcould have been
generated by a different pair ˜Xand ˜Y, as the diagram below indicates,
2.1. EXAMPLES AND DEFINITION 81
a figure goes here
for we still have ˜X+V=˜Y.
In the first case we might have had X= 2 andY= 3 , so that V= 1 , while in the
second, we might have had X=−4 andY=−3 , and again V= 1 . Even though we
have let this Vgo free, sliding from place to place along R, we still want to say that this
is only one V, and in fact, we want to identify this Vwith theVtied to the origin in
(2). In other words, we would like to say that all three V’s used above are equivalent to
each other.
More formally, the element Visgenerated by an ordered pair, V= [X,Y] , which we
read as the vector fromXtoY, forX, Y∈R. If some ˜Vis generated by another ordered
pair, ˜V= [˜X,˜Y],˜X,˜Y∈R, then we want equality V=˜Vto mean that ˜Y−˜X=Y−X.
Moreover, we want to representV= [X,Y] , the vector from XtoY, by the vector from
the origin 0 to Y−X, V = [0,Y−X] .This representation of Vis unique , since if any
other pair also generates V, V = [˜X,˜Y] , the representative V= [0,˜Y−˜X] = [0,Y−X]
sinceV=Vimplies that ˜Y−˜X=Y−X. Therefore much as each rational number
is an equivalence class, represented by a single rational number - as1
2represents the
equivalence class1
2,2
4,3
6,..., eachVis an equivalence class of ordered pairs V= [X,Y] ,
whereX,Y∈R. It is uniquely represented by an element of R, viz.V=Y−X, the
representation being independent of the particular ordered pair [ X,Y] which generates V.
It is possible to think of Veither as an ordered pair with an equivalence relation, or just
as the representative V= [0,Y−X] of the whole equivalence class, the representation
being written more simply as an element of R:V=Y−X, where here equality is between
elements of R.
The generalization is now easily made
Definition: . (Free vectors). Let XandYbe any elements of Rn. An element V∈Vn,
“physicists’ n-space”, is defined as an equivalence class of ordered pairs of elements in Rn,
V= [X,Y], X, Y ∈Rn,
with the following equivalence relation: If V= [X,Y] and ˜V= [˜X,˜Y] , then
V=˜V⇐⇒ ˜Y−˜X=Y−X,
where the second equality is that of elements in Rn. If we are given XandYinRn, we
speak ofV= [X,Y] as the free vector going fromXtoY.
Previous reasoning also shows that eachV∈Vnis uniquely represented by the ordered
pairV= [0,Y−X] . This representation is independent of the elements [ X,Y] which
generatedV.
We were led to this definition of Vnby examining the situation in the special case of
V1. Since our formal reasoning there was quite algebraic and general, we know that the
definition works algebraically. The geometry works too. An example in V2should make
the general case clear.
LetX= (1,3) andY= (2,1) . These two points in R2generate the ordered pair
V= [(1,3),(2,1)] in V2.Vis the vector going fromX= (1,3)toY= (2,1) . Of all
equivalentV’s, the unique representative which begins at the origin is V= [(0,0),(1,−2)] ,
which we simply write as V= (1,−2) and represent as an ordinary element of R2. On
the same diagram we exhibit the vector from ˜X= (−2,2) to ˜Y= (−1,0) , which is
˜V= [(−2,2),(−1,0)] . The unique representative (of all ˜V’s equivalent of ˜V) which begins
82 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE
from (0,0) is ˜V= [(0,0),(1,−2)] , which we write simply as ˜V= (1,−2) . Comparison of
Vand ˜Vreveals that they are equal, V=˜V. Thus, from the diagram, we see that a
free vector is an equivalence class of directed line segments, with two directed line segments
V,˜Vbeing equivalent as vectors in V2if they are equivalent to the same directed line
segment which begins at the origin. In more geometrical language, V=˜Vif by sliding
them “parallel to themselves”, they can be made to coincide with their representer which
begins at the origin. (We shall not define “parallel” here. It is not needed because we
already have a satisfactory algebraic definition of equivalence.)
Notice that X= (1,3) andY= (2,1) also generates a second ordered pair ˆV=
[(2,1),(1,3)] , the vector fromY= (2,1)toX= (1,3) . Its unique representation which
begins at the origin is ˆV= [(0,0),(−1,2)] , or more simply ˆV= (−1,2) . Comparison with
the previous example shows that ˆV=−V:the vector from YtoXis the negative of the
vector from XtoY. We need the little arrow on our picture of V= [X,Y] to distinguish
it from −V= [Y,X] which is also between the same points but headed in the opposite
direction.
From now on we shall denote a vector V∈VnfromXtoYby its representative
Y−XinRn, soV=Y−X. Hence the vector from (1,3) to (2,1) will be immediately
written as V= (1,−2) . As we have said many times, the representation V=Y−Xas
an element on Rnis independent of which particular pair [ X,Y] happened to generate V.
The following diagram shows a whole bunch of equivalent vectors Vj∈V2,
a figure goes here
Vj=Vk, and their particular representative Vchained to the origin.
In order to justify calling the elements of Vnvectors, we should prove that the elements
ofVndo form a vector space . Addition and scalar multiplication must first be defined, an
easy task. Since every V∈Vnis uniquely represented as an element of Rn, V=Y−X∈
Rn, we use addition and scalar multiplication for elements of Rn—which has already been
defined. Because Rnis known to be a vector space, it is a tedious triviality to prove.
Theorem 2.1 .Vnis a linear vector space.
Proof: . Only a smattering.
(1)Vnis closed under addition. Say V1andV2are in Vn. Then they are represented
as the difference of two elements of Rn, sayV1=Y1−X1andV2=Y2−X2. Thus
V1+V2= (Y1−X1) + (Y2−X2) = (Y1+Y2)−(X1+X2),
so that their sum is generated by [ X1+X2, Y1+Y2] . In other words, there is at
least one pair of elements, [ X3,Y3], X 3=X1+X2andY3=Y1+Y2, inRnwhich
generateV1+V2, so thatV3=V1+V2∈Vn. Of course [0 ,Y3−X3] and many other
pairs also generate V3.
(2)Commutativity.
V1+V2= (Y1−X1) + (Y2−X2) = (Y2−X2) + (Y1−X1) =V2+V1.
(3) (α+β)V1= (α+β)(Y1−X1) =α(Y1−X1) +β(Y1−X1) =αV1+βV1
2.1. EXAMPLES AND DEFINITION 83
Example: IfA= (4,2,−3), B= (0,1,−2), C= (−1,0,1
2) andD= (4,−1
2,1) , find the
vectorV1fromAtoBand the vector from CtoD. Then compute V1+ 2V2and
V1−V2.
solution: V1=B−A= (0,1,−2)−(4,2,−3) = ( −4,−1,1)
V2=D−C= (4,−1
2,1)−(−1,0,1
2) = (5,−1
2,1
2)
V1+ 2V2= (−4,−1,1) + 2(5,−1
2,1
2) = (−4,−1,1) + (10,−1,1) = (6,−2,2)
V1−V2= (−4,−1,1)−(5,−1
2,1
2) = (−4,−1,1) + (−5,1
2,−1
2) = (−9,−1
2,1
2)
Exercises
(1) (a) Find the vector representing the free vectors from the given A∈RntoB∈Rn.
(i)A= (3,1), B= (2,2).
(ii)A= (−3,3), B= (0,4).
(iii)A= (2,2,3), B= (5,2,17)
(iv)A= (0,0,0)B= (9,8,−3)
(v)A= (1,2,3), B= (0,0,−1)
(vi)A= (0,0,−1), B= (1,2,3)
(b) LetV1andV2be the respective vectors of iii) and v) above. Compute V1+
V2, V1−V2, and 2V1−3V2.
(c) Draw a diagram on which you indicate the vector going from A= (3,1) to
B= (2,2) , and indicate the representer of that vector which begins at the
origin. Do the same with the vector from BtoA.
(2) Which of the following subsets of C[−1,1] are linear spaces:
(a) The set of all even functions in C[−1,1] , that is, functions f(x) with the addi-
tional property f(−x) =f(x) , likex2and cosx.
(b) The set of all functions finC[−1,1] with the additional property that |f(x)| ≤
1 .
(c) The set of all functions finC[−1,1] with the property that f(0) = 0 .
(3) In R3, letX= (1,−1,2) andY= (0,4,−3) . FindX+ 2Y, Y−X, and 7X−4Y.
(4) (a) Show that for every X∈R3you can find scalars αj∈Rsuch thatXcan be
written as
X=α1e1+α2e2+α3e3,
wheree1= (1,0,0), e2= (0,1,0), e3= (0,0,1) .
(b) IfX∈R3, can you find scalars αj∈Rsuch that
X=α1θ1+α2θ2+α3θ3,
whereθ1= (1,−1,0), θ2= (−1,1,0), θ3= (0,0,1) , andαj∈R? Proof or
counter-example.
84 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE
(c) Find two polynomials P1(x) andP2(x) in ? 1such that for every polynomial
P(x)∈?1you can find scalars αj∈Rsuch thatPcan be written in the form
P(x) =α1P1(x) +α2P2(x).
(5) LetV=R×Rwith the following definition of addition and scalar multiplication
X+Y= (x1+x2, y1+y2), αX = (αx1,0),
0 = (0,0),−X= (−x1,−x2).
IsVa vector space? Why?
(6) Show that any field can be considered to be a vector space over itself.
(7) Consider the set
S={u∈C2[0,1]:a2u/prime/prime+a1u/prime+a0u= 0},
where theaj(x)∈C[0,1] . IsSa linear space? Note that we do not yet know that
Shas any elements at all. The proof that Sis not empty is the existence theorem
for ordinary differential equations.
(8) By integrating |x|the “right” number of times, find a function which is in Ck[−1,1]
but is not in Ck+1[−1,1] .
2.2 Subspaces. Cosets.
With this section we begin the process of assigning names to the various concepts sur-
rounding the idea of a linear vector space. This name calling will take us the balance of the
chapter. Although the ideas are elementary and theorems simple, do not deceive yourselves
into thinking this must be some grotesque joke that mathematicians have perpetrated. You
see, we are in the process of building a machine. Most of its constituent parts are very
easy to grasp. But when combined, the machine will be equipped successfully to assault a
diversity of problems which appear off hand to be unrelated.
The value of this abstract formalism is that many seemingly distinct complicated specific
problems are just one single problem in a variety of fancy dresses. By ignoring the extraneous
paraphernalia we can concentrate on the essential issues.
a figure goes here
We begin by defining what is meant by a subspace of a vector space W. While reading
the definition, think of a plane through the origin, which is a subspace of ordinary three
dimensional space.
Definition: . A setAis alinear subspace (linear variety, linear manifold) of the linear
spaceWif i)Ais a subset of W, and ii)Ais also a linear space under the operations of
vector addition and multiplication by scalars already defined on V.
2.2. SUBSPACES. COSETS. 85
Examples:
(1) LetA={X∈R3:X= (x1,x2,0)}, that is, the points in R3whose last coordinate
is zero. Since A⊂R3, and a simple check shows that Ais also a linear space,
we see that Ais a linear subspace of R3. Intuitively, this set Acertainly “looks
like” R2. You are right, and recall that the fancy word for this equivalence - of
R2= (x1,x2) and the points in R3of the form ( x1,x2,0) —is isomorphic . Similarly,
the setB={X∈R3:X= (x1,0,x3)}is also a subspace of R3.Bis also
isomorphic to R2.
(2) LetA={X∈Rn:X= (x1,x2,...,x k,0,0,...,0)}, that is, the points in Aare those
points in Rnwhose last n−kcoordinates are zero. It is easy to see that Ais a
linear subspace of Rn, and that Ais isomorphic to Rk.
(3) LetA={f∈C[0,1]:f(0) = 0 }.Ais a subset of the linear space C[0,1] , and is
also a linear space (check this). Thus Ais a linear subspace of C[0,1] .
(4) LetA={f∈C[0,1]:f(0) = 1 }.Ais a subset of C[0,1] , but it is not a linear
subspace since - as we saw in the last section (p. ?)— Ais itself not a linear space.
The following lemma supplies a convenient criterion for checking if a given subset Aof
a linear space Wis a subspace.
Theorem 2.2 . IfAis a non-empty subset of the linear space W, thenAis a linear
subspace of W⇐⇒Ais closed under addition of vectors in Aand multiplication by all
scalars.
Proof: .⇒. SinceAis a subspace, it is itself a linear space. But all linear spaces are,
by definition, closed under addition and multiplication by scalars.
⇐. BecauseAis a subset of W, and properties 1,2,5,6,7, and 8 hold in W, they also
hold for the particular elements in Wwhich happened to be in A. Notice that here we use
the fact that Ais closed under addition. Therefore only the existential axioms 3 and 4 need
be checked. Since Ais not empty, it contains at least one element, say X∈A. Because
Ais closed under multiplication by scalars we see that 0 = 0 ·X∈A. Furthermore, for
everyX∈A, also −X= (−1)·X∈A.
Example: LetA={f∈C1[0,1]:f/prime(0) = 0 }. SinceAis a subset of the linear space
C1[0,1] , all we need show is that Ais closed under addition and multiplication by scalars in
order to prove Aa linear subspace of C1[0,1] . Iff, g∈A, then (f+g)/prime(0) = (f/prime+g/prime)(0) =
f/prime(0) +g/prime(0) = 0 , so f+g∈A. Also, for any α∈R,(αf)/prime(0) =α(f/prime)(0) =α·0 = 0 , so
αf∈A.
Theorem 2.3 . The intersection of two subspaces is also a subspace, but the union of two
subspaces is not necessarily a subspace. More generally, the intersection of any collection
of subspaces is also a subspace.
Proof: . LetA, B be subspaces of W. We show that A∩Bis a subspace. Since
A∩B⊂W, all we need show is the closure properties of A∩B. IfX, Y∈A∩B, then
XandYare both in AandB, soX+Y∈AandX+Y∈B⇒X+Y∈A∩B
too. Similarly for scalar multiples. The proof that A∩B∩C∩...is a subspace is identical
except for a notational mess.
86 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE
For the second part of the theorem we merely exhibit an example of two subspaces
A,B for whichA∪Bis not a subspace. In R2letAbe the linear subspace “horizontal
axis”, that is, A={X∈R2:X= (x1,0)}, whileBis “the vertical axis”, B={X∈
R2:X= (0,x2)}. ThenA∪Bis the “cross” of all points on either the horizontal axis or
the vertical axis. This is not a linear space because points like (1 ,0)∈A,(0,1)∈Bdo
not have their sum (1 ,0) + (0,1) = (1,1) inA∪B. Precisely for this reason R2=R1×R1
was constructed as the Cartesian product of R1with itself; for if it had been constructed
asR1×R1, then only the points situated on the axes themselves would get caught. More
generally - and for the same reason - the Cartesian product is the process always used to
“glue” together a larger space from several linear spaces. Only when A⊂B(orB⊂A)
isA∪Balso a subspace (Exercise 4).
Your image of a linear space should be R3, and a subspace Sis a plane or line in R3.
Note that since every subspace must contain 0, these planes or lines must pass through the
origin .
Example: LetSc={X∈R2:x1+ 2x2=c, creal}. Thus, the set Scis all points
S= (s1,s2)∈R2on the straight line s1+ 2s2=c. For what value(s) of cisSca
subspace? If Scis a subspace, then we must have aS∈Scfor all scalars a, that is
aS= (as1,as 2)∈Sc⇒as1+ 2as2=c. But fora= 0 this states that c= 0 . Therefore
the only possible subspace is S0={X∈R2:x1+ 2x2= 0}. It is easy to check that if S1
andS2are in S0, then so are S1+S2andaS1. Thus S0is a subspace. Similarly, every
straight line through the origin is a subspace.
Our question now is, how can we talk about the other straight lines or planes which do
not happen to pass through the origin? First we answer the question for our example above.
There we have the linear space R2and the subspace S0which will be simply written as
S.Sis a line through the origin. Let X1be any element in R2(think ofX1as a point).
Then the set of all elements of R2which can be written in the form S+X1, whereS∈S,
is the line “parallel” to Swhich passes through X1. This line is written as S+X1. More
explicitly, say X1= (1,3
2) . The set S+X1is the set of all points X= (x1,x2)∈R2of
the form
X=S+X1,whichS∈S,
or
(x1,x2) = (s1,s2) + (1,3
2),wheres1+ 2s2= 0.
Consequently x1=s1+ 1 , andx2=s2+3
2. Using the relation s1+ 2s2= 0 , we find that
x1+2x2= 4 — exactly the equation of the straight line through X1= (1,3
2) and “parallel”
to the subspace S. This subset, S+X1={X∈R2:X=S+X1, whereS∈S}, is called
theX1coset of S. Thus, cosets are the names given to “linear objects” which are not
subspaces. They are subspaces translated to pass through X1. You might prefer to call
them affine subspaces instead of cosets.
Please observe that the cosets S+X1andS+X2, whereX1, X 2∈W, are not
necessarily distinct. In our example, these cosets coincide if and only if X2is on the line
S+X1, that is, if X2∈S+X1. The easiest way to test this is to see if X2−X1∈S.
SayX1= (1,3
2) as before, and that X2= (2,1) . Then the cosets S+X1andS+X2are
the same since the point X2−X1= (1,−1
2) is in S. It should be geometrically clear that
2.2. SUBSPACES. COSETS. 87
the relation of equality among these cosets is an equivalence relation (and so deserving of
the title “equality”). We shall state these ideas formally as we turn from this special - but
characteristic - example to the general situation.
The general problem of describing lines or planes or “higher dimensional linear objects”
which do not pass through the origin - so are not subspaces - is solved similarly.
Definition: . LetWbe a linear space, Sa subspace of V, andX1any element of W.
All elements in Wwhich can be written in the form S+X1, whereS∈S, is called the
X1coset ofS, and written as S+X1.
Our first theorem states that if X2is in theX1coset of S, thenX1is in theX2
coset of S:
Theorem 2.4 .X2∈S+X1⇐⇒X1∈S+X2.
Proof: SinceX2∈S+X1, there is an S∈Ssuch thatX2=S+X1. Therefore
X1= (−S) +X2. Because Sis a linear space, ( −S)∈S. ThusX1has been written as
the sum of X2and an element of S, which means that X1∈S+X2.
By the same argument, one sees that any two cosets S+X1andS+X2are either
identical or are disjoint (have no element in common). Thus the cosets of SpartitionW
in the sense that every element of Wis in exactly one coset, just as for our example, every
point in the plane R2was in exactly one straight line parallel to the subspace determined
byx1+ 2x2= 0 .
Although we were motivated by geometrical considerations, the ideas apply without
alteration to any linear space. This is illustrated by again examining the set
A={f∈C[−1,1]:f(0) = 1 },
which is not a subspace. It is a coset of a subspace SofC[−1,1] which is constructed as
follows. Consider the subspace Swhich is “naturally” associated with A, viz.
S={g∈C[−1,1]:g(0) = 0 }.
ThenAis the coset S+1, A=S+1 . This is true since clearly A⊃S+1 . AlsoA⊂S+1
because for every f∈A,
f(x) = [f(x)−1] + 1 =g(x) + 1,whereg∈S.
ThereforeA=S+1 . Similarly, we could have written AasS+ˆf, where ˆfisanyfunction
inA, for example A=S+ cosx.
Exercises
(1) Find which of the following subsets of Rnare subspaces.
(a){X∈Rn:x1= 0},
(b){X∈Rn:x1≥0},
(c){X∈Rn:x1−x2= 0},
(d){X∈Rn:x1−x2= 1},
(e){X∈Rn:x2
1−x2= 0},
88 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE
(2) In P3, the linear space of all polynomials of degree ≤3 , letA={p(x)∈P3:p(0) =
0}, and letB={p(x)∈P3:p(1) = 0 }.
(a). Show that AandBare subspaces of P3.
(b). FindA∩BandA∪B. Give an example which shows that A∪Bis not a
subspace of P3.
(3) (a) If X1andX2are given fixed vectors in R2then is
A={X∈R2:X=a1X1+a2X2, a1anda2any scalars }
a subspace of R2?
(b) Same as (a) but replace R2by an arbitrary linear space W.
(c) IfX1,X2,...,X k∈W, then is
A={X∈W:X=k/summationdisplay
1ajXj,for any scalars aj},
a subspace of W?
(4) LetAandBbe subspaces of a linear space W. Prove that A∪Bis also a subspace
if and only if either A⊂BorB⊂A, that is, if one of the subspaces contains the
other.
(5) Let SandTbe subspaces of a linear space W, and suppose that Ais a coset
ofSandBis a coset of T. Prove that (a). A⊂B⇒S⊂T, and also (b).
A=B⇒S=T.
(6) (a) Write the plane 2 x1−3x2+x3= 7 as a coset of some suitable subspace S⊂R36.
(b) Write the set A={f∈C[0,4]:f(0) = 1, f(1) = 3 }, as a coset of some suitable
subspace S⊂C[0,4] .
(c) Write the set A={f∈C1[0,4]:f(1) = 1, f/prime(1) = 2 }as a coset of some
suitable subspace S⊂C1[0,4] .
2.3 Linear Dependence and Independence. Span.
IfWis a linear space and X1,X2,...,X k∈W, then we know that, for any scalars aj,
Y=k/summationdisplay
j=1ajXj=a1X1+a2X2+...+akXk
is also in V.Yis alinear combination of theXj’s. Now if 0 can be expressed as a linear
combination of the Xj’s, where at least one of the aj’s is not zero we expect that there is
something degenerate around. In fact, if 0 = a1X1+...+akXkwhere saya1/negationslash= 0 , then we
can solve for X1as a linear combination of X2,X3,...,X k,
X1=−1
a1(a2X2+...+a,Xk).
This leads us to make a definition and state a theorem.
2.3. LINEAR DEPENDENCE AND INDEPENDENCE. SPAN. 89
Definition: . A finite set of elements Xj∈W, j = 1,...,k is called linearly dependent if
there exists a set of scalars aj, j= 1,...,k ,notall zero such that 0 =k/summationdisplay
1ajXj. If theXj
are not linearly dependent, we say they are linearly independent .
Theorem 2.5 . A set of vectors Xj∈W, j = 1,...,k is linearly dependent if and only if
at least one of the Xj’s can be written as a linear combination of the other Xj’s.
To test if a given set of vectors is linearly independent, an equivalent form of Theorem
5 is useful.
Corollary 2.6 A set of vectors Xj∈W, j = 1,...,k is linearly independent if and only if
k/summationdisplay
j=1ajXj= 0 implies that a1=a2=...=ak= 0.
Examples:
(1) The vectors X1= (2,0), X 2= (0,1), X 3= (1,1) in Rare linearly dependent since
0 =X1+ 2X2−2X3. Equivalently, we could have applied the theorem since X3can
be written as a linear combination of X1andX2
X3=1
2X1+X2.
(2) The functions f1(x) =ex, f2(x) =e−x, f3(x) =ex+e−x
2inC[0,1] are linearly depen-
dent since
0 =f1+f2−2f3
(3) The vectors X1= (2,0,1), X 2= (−1,0,0) in R3are linearly independent, since if
for somea1, a2,
0 =a1X1+a+ 2X2= (2a1,0,a1) + (−a2,0,0),
then
0 = (0,0,0) = (2a1−a2,0,a1),
which implies that 2 a1−a2= 0 , anda1= 0 =⇒a1=a2= 0 .
a figure goes here
A simple consequence of these ideas is the following
Theorem 2.7 . IfAandBare any subsets of the linear space Wand ifA⊂B, then
i)Ais linearly dependent ⇒Bis linearly dependent; and the contrapositive: ii) Bis
linearly independent ⇒Ais linearly independent.
We now prove the transitivity of linear dependence.
90 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE
Theorem 2.8 . IfZis linearly dependent on the set {Yj}, j= 1,...,n and eachYj
is linearly dependent on the set {Xl}, l= 1,...,m thenZis linearly dependent on the
{Xl}.
Proof: . This is trivial arithmetic. We know that
Z=a1Y1+...+anYn,
and that
Yj=c1jX1+c2jX2+...+cmjXm
By substitution then
Z=al(cllXl+···+cmlXm) +a2(c12X1+···+cm2Xm)
+···+an(c1mX1+···+cmnXm)
= (a1c11+a2c12+···+ancln)X1+ (a1c21+···+anc2n)X2
+···+ (a1cml+···+cmn)Xm
=γ1X1+···+γmXm,whereγl=n/summationdisplay
j=1ajclj.
More concisely:
Z=n/summationdisplay
j=1ajYj=n/summationdisplay
j=1aj/parenleftBiggm/summationdisplay
l=1cljXl/parenrightBigg
=m/summationdisplay
l=1
n/summationdisplay
j=1ajclj
Xl=m/summationdisplay
l=1γlXl.
LetX1andX2be any elements of a linear space W. Is there a smallest subspace
AofWwhich contains X1andX2? There are two possible ways of answering this,
constructively and non-constructively.
First, constructively. We observe that the desired subspace must contain X1and
X2, and all linear combinations of X1andX2, that is,Amust contain all X∈W
of the form X=a1X1+a2X2for all scalars a1anda2. But observe that the set
B={X∈V:X=a1X1+a2X2}is a linear space, since if XandY∈B, thenaX∈B
for any scalar a, and alsoX+Y∈B. Thus the desired subspace Ais justBitself.
The constructive proof goes as follows: just let Abe the intersection of all subspaces
containing X1andX2. By Theorem 3 the intersection of these subspaces is also a subspace.
It is clearly the smallest one. Do you feel cheated? This type of reasoning is often used in
modern mathematics. Although it reveals little more than the existence of the sought-after
object, it is an extremely valuable procedure when you really don’t want anything more
than to know the object exists. More important, procedures like this are vital when there
is no constructive proof available.
More generally, if S={Xj}, j= 1,...,k , is any finite subset of a linear space W, we
ask for the smallest subspace AofWwhich contains S. There are two proofs - exactly
as in the simple case above (where k= 2 ). From the constructive proof we find that
A={X∈W:X=k/summationdisplay
1ajXj, ajscalars },
soAis the set of all linear combinations of the Xj’s. This set Ais called the span ofS,
and denoted by A= span(S) . We also say that SspansA, or thatAisgenerated byS.
2.3. LINEAR DEPENDENCE AND INDEPENDENCE. SPAN. 91
Examples:
(1) In R3letX1= (1,0,0) andS2= (0,1,0) . Then the span of S={Xj, j= 1,2}is
allX∈R3of the form X=a1X1+a2X2= (a1,a2,0) . If we imagine R3as ordinary
3-space, then the span of X1andX2is the entire x1,x2plane.
(2) In R3, letX1= (1,0,0), X 2= (0,1,0) , andX3= (0,0,1) . Then the span of
T={Xj, j= 1,2,3}is allX∈R3of the form X=a1X1+a2X2+a3X3(a1,a2,a3) .
Since all of R3can be so represented, we have span( T) =R3, that is, the set Bspans
R3. Comparing these two examples, we see that S⊂Tand span(S)⊂span(T) .
(3) In R3, letX1= (1,0) andX2= (0,1) . Then the span of S={X1,X2}is all of
R2, since every X∈R2can be written as X=a1X1+a2X2, wherea1anda2are
scalars. Many other sets also span R2. In fact almost every set of two vectors X1and
X2inR2span R2. This can be seen from the diagram, where we have drawn a net
parallel to X1andX2. ThenX=a1X1+a2X2. Any vectors X1andX2would
do equally well, as long as they do notpoint in the same (or opposite) direction.
We collect some properties of the span
Theorem 2.9 . LetR, S , andTbe subsets of a linear space W. Then
(a)R⊂span(R).
(b)R⊂S=⇒span(R)⊂span(S).
(c)R⊂span(S)andS⊂span(T) =⇒R⊂span(T).
(d)S⊂span(T) =⇒span(S)⊂span(T).
(e) span(span( T)) = span(T).
(f)A vectorXj∈Sis linearly dependent on the other elements of S⇐⇒span(S) =
span(S−{Xj}). (HereS−{Xj}means the set Awith the one vector Xjdeleted).
Proof: These all depend on the representation of span( S) as a linear combination of the
elements of S.
(a) and (b)—Obvious. They really should be if you understand the definitions.
(c). A direct translation of Theorem 7.
(d). This is the special case R= span(S) of part c.
(e). By part (a) span(span( T))⊃span(T) . The opposite inclusion span(span( T))⊂
span(T) is the special case S= span(T) of part (d).
(f).Xjlinearly dependent on S−{Xj}=⇒S⊂span(S−{Xj}) . Thus by part (d),
span(S)⊂span(S−{Xj}) . Inclusion in the opposite direction span( S−{Xj})⊂span(S)
follows from part (b). Therefore span( S) = span(S− {Xj}) means that Xj∈span(S)
can be expressed as a linear combination S− {Xj}, i.e., the other Xk’s.
Now most likely this proof was your first taste of abstract juggling and you find it
difficult. Relax and don’t be impressed with how formidable it appears. Except for parts a
and b, the whole business hinges on the explicit construction of Theorem ?. Since (d) is a
special case of (c), a good exercise is to write out the proof of (d) without relying on (c).
InR2, letX1= (1,0) ,X2= (0,1) , andX3be any vector in R2. Observe that X1
andX2together span R2. ThusX3can be expressed as a linear combination of X1and
X2, so thatX1, X 2, andX3are linearly dependent. The next theorem is a generalization
of this idea.
92 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE
Theorem 2.10 . If a finite set A={Xj, j= 1,...,n }spans a linear space W, then ev-
ery set ˜S={Yj∈V, j= 1,...,m>n }with more than nelements is linearly dependent.
In other words, every linearly independent set has at most nelements.
Proof: Pick anyn+ 1 elements Y1,...Y n+1from ˜Sand throw the rest away. Call the
new setS. We shall show that these n+ 1 elements are linearly dependent. Then, since
S⊂˜S, Theorem ? tells us that ˜Sis also linearly dependent. The only problem is how
to carry out the proof without getting involved in a mess of algebra. By the principle of
conservation of effort, this means that there will be some fancy footwork.
Reasoning by contradiction, assume Sis linearly independent. If we can show that
span(A) = span(S−{Yn+1}) , then span( S)⊂span(S−{Yn+1}) because span( S)⊂V=
span(A) = span(S− {Yn+1})) . Since span( S− {Yn+1})⊂span(S) , we can apply part f
of Theorem ? to conclude that Sis linearly dependent - the desired contradiction.
Thus, assuming S={Y1,...,Y n+1}is linearly independent, we are done if we prove
that span(A) = span(S− {Yn+1}) . Consider the set Bk={Y1,...,Y k, Xk+1,...,X n}.
We know that B0=A, so that span( B0) = span(A) =W. Then by induction we shall
prove that span( Bk) =W=⇒span(Bk+1) =W. Since span( Bk) spansW, thenYk+1
is a linear combination of the elements of Bk. Because the Y’s are assumed linearly
independent, this linear combination must involve at least one of Xk+1,...,X n. Say it
involvesXk+1(if not, relabel the X’s to make it so). Then we can solve for Xk+1as a
linear combination of span( Bk+1) . Therefore W= span(Bk) = span(Bk+1) . Putting this
part together, we find that span( A) =W= span(B0) = span(B1) =...= span(Bn) . But
Bn=S− {Yn+1}. Thus span( A) = span(S− {Yn+1}) , and the proof is completed.
Example . InR2, any three (or more) non-zero vectors are linearly dependent since the two
vectorsX1= (1,0) andX2= (0,1) span R2.
Exercises
(1) (a) In P2p1(x) = 1, p2(x) = 1 +x, p 3(x) =x−x2
(b) In R3, X 1= (0,1,1), X 2= (0,0,−1), X 3= (0,2,3) .
(c) InC[0,π], f(x) = sinx, g(x) = cosx.
(d) In Rn, e1= (1,0,0,..., 0), e2= (0,1,0,0),...,e n= (0,0,..., 0,1) .
(2) Use the result of (d) to show that any set of n+ 1 vectors in Rnmust be linearly
dependent.
(3) (a) Find a set which spans
i)P3, ii)R4
(b) Show that no finite set spans l1.
(4) LetX1,...,X kbe any elements of a linear space V.
(a) Prove that span( {X1,...,X k}) = span( {X1+aXj,X2,...,X k}) , whereais
any scalar and Xjis any of the X2,X3,...,X k,.
(b) Prove that span( {X1,...,X k}) = span( {aX1,X2,...,X k}),a/negationslash= 0 .
2.4. BASES AND DIMENSION 93
(c) In Rn, consider the ordered set of vectors {X1,X2,...,X k}, whereXj=
(x1j,x2j,...,x nj) . They are said to be in echelon form if i) noXjis zero, and ii)
theindex of the first non-zero entry in Xjis less than the index of the first non-
zero entry in Xj+1, for eachj= 1,...,k −1 . ThusX1= (0,1,0), X 2= (0,0,1)
are in echelon form while X1= (0,1,0), X 2= (1,0,1) are not in echelon form.
Prove that any set of vectors in echelon form is always linearly independent. (I
suggest a proof by induction).
(5) For what real value(s) of the scalar αare the vectors ( α,1,0),(1,α,1) and (0,1,α)
inR3linearly dependent?
(6) (a) In R3, letX1= (3,−1,2) . Express ( −6,2,−4) linearly in terms of X1. Show
that (3,4,−7) cannot be expressed linearly in terms of X1. Can (1,2,1) be
expressed linearly in terms of X1?
(b) In R3, letA={X1,X2}, whereX1= (1,3,−2) andX2= (2,1,1) . Express
(3,−1,4) linearly in terms of A. Show that (0 ,0,2) cannot be expressed linearly
in terms of A. Can (0,5,−5) be expressed linearly in terms of A?
(7) (a) In C[0,10] , letf1,...,f 8be defined by
f1(x) =x2−x+ 2, f 5(x) =x3
f2(x) = (x+ 1)2f6(x) = sinx
f3(x) =x+ 3 f7(x) = cosx
f4(x) = 1 f8(x) = sin(x+π/4).
LetA={f1,f2,f3}. Expressf4linearly in terms of A. Show that f5cannot
be expressed linearly in terms of A. Isf6∈span(A) ? Isf8∈span(f6,f7) ? Is
f6∈span(f5,f7,f8) ?
(b) If we let f9(x) = (x−1)3, f10(x) = 2x−1 , determine which of the following sets
are linearly dependent:
(i){f1,f3,f10},
(ii){f1,f5,f9},
(iii){f3,f4,f10},
(iv){f1,f4,f5,f9},
2.4 Bases and Dimension
If the set {X1,...,X m}spans the linear space W, is there any set with less than m
vectors which also spans W? There certainly is if the {X1,...X m}are linearly depen-
dent, for if say Xmdepends linearly upon the {X1,...,X m−1}, then by Theorem ??,
span({X1,...,X m}) = span( {X1,...,X m}) =W, so then {X1,...,X m−1}spanW.
We can continue and eliminate the extra linearly dependent elements until we obtain a set
{X1,...,X n}of linearly independent vectors which still span W.
Definition . A set of vectors Xj∈W,j= 1,...,n which is i) linearly independent, and
ii) spansWis called a basis forW.
Examples .
94 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE
(1) In R2, the vectors X1= (1,0) andX2= (0,1) are linearly independent and span
R2. Therefore X1andX2form a basis for R2. The vectors X3= (3,−1) and
X4= (−2,2) in R2are also linearly independent and span R2. They thus constitute
another basis for R2. Almost any two vectors in R2span R2, as long as they do not
point on the same or opposite direction.
(2) In P2, the polynomials p1(x) = 1 , and p2(x) =x−x2donotform a basis. They
are linearly independent but do not span the space - since for example you can never
obtain the polynomial p(x) =xwhich is in P2. If we add the third polynomial, say
p3(x) =x−2x2, thenp1, p2andp3do form a basis for P2.
Bases have an important property.
Theorem 2.11 . If{X1,...,X n}form a basis for the linear space W, then every X∈
Wcan be expressed uniquely as a linear combination of the Xj’s.
Remark : Every set which spans Whas, by definition, the property that every X∈Wcan
be expressed as a linear combination of the Xj’s. The point here is that for a basis, this
linear combination is uniquely determined.
Proof: Suppose that X=n/summationdisplay
1akXkand alsoX=n/summationdisplay
1bkXk. We must show that ak=bk
for allk. Subtracting the two equations we find that 0 =n/summationdisplay
1ckXk, whereck=ak−bk.
But since the Xk’s are linearly independent, by the Corollary to Theorem 5, the only way
a linear combination can be zero is if ck= 0, k= 1,...,n , that is,ak=bkfor allk.
We have observed that a linear space may have several different bases. Is it possible
that different bases contain a different number of elements? Our next theorem states that
the answer is NO.
Theorem 2.12 . If a linear space Whas one basis with a finite number of elements, say
n, then all other bases are finite and also have exactly nelements.
Proof: We invoke Theorem ?. Let Abe a basis with nelements and Bbe a basis with
melements. Now AspansWand the elements of Bare linearly independent, so the
Theorem ?, m≤n. Reversing the roles of AandBwe find that n≤m. Therefore
n=m.
With this result behind us, we can now define the dimension of a linear space.
Definition . If a linear space Whas a basis with nelements, then we say that the dimension
ofWisn. If a linear space Whas the property that no finite set of elements spans it,
we say it is infinite dimensional .
Remarks . Theorem ? states that the dimension of Wis independent of which basis we
happened to pick. If we want to emphasize the dimension of a finite dimensional space, we
will writeWn.
Announcement . The dimension of Rnisn, for thenelementse1= (1,0,0,..., 0), e2=
(0,1,0,..., 0),..., e n= (0,..., 0,1) are linearly independent and span Rn.
A picture. We have seen that e1= (1,0,0), e2= (0,1,0), e3= (0,0,1) form a basis
inR3. Thus every X∈R3can be expressed uniquely as a linear combination of the ej’s,
X=a1e1+a2e2+a3e3. If we represent e1as a directed line segment from the origin to
(1,0,0) , and similarly for e2ande3, thenXis the geometrical sum of a1e1+a2e2+a3e3,
2.4. BASES AND DIMENSION 95
and is represented as a directed line segment from the origin to ( a1,a2,a3) . In R3, e1is
usually written as i,e2asjande3ask, so that a vector X∈R3is written as
X=a1i+a2j+a3k.
The points in the plane x3= 0 , which is isomorphic to R2, are then represented as
X=a1ˆi+a2ˆj+ 0ˆk=a1ˆi+a2ˆj. We would retain this notation except that one runs out of
letters when considering spaces of higher dimension. For that reason the subscript notation
e1,e2,... is better suited to our purposes.
It behooves us to show that the linear space C[0,1] of functions continuous in the
interval [0,1] is infinite dimensional. This will be done by proving that the functions
f0(x) = 1, f1(x) =ex, f2(x) =e2x,...,f n(x) =enx,... are linearly independent. Assume
that 0 =N/summationdisplay
k=0akekx, whereNis any non-negative integer. We must show that all the ak’s
are zero.
The trick is to use induction. For N= 0 , we know that 0 = a0only ifa0= 0 .
Suppose 1,ex,e2x,...,e(N−1)xare linearly independent. ThenN−1/summationdisplay
k=0akekx= 0 if and only
if all of the ak’s are zero. Let us show that this implies thatN/summationdisplay
k=0akekx= 0 if and only if
all theak’s vanish. Take the derivative. The constant term drops out and we are left with
0 =a1ex+ 2a2e2x+···+NaNeNx.
Factor out ex
0 =ex(a1+ 2a2ex+···+NaNe(N−1)x).
Sinceexis never zero, we know that
0 =a1+ 2a2ex+···+NaNe(N−1)x.
By our induction hypothesis, this linear combination of ! ,ex,...,e(N−1)xcan be zero if
and only if a1=a2=a3=···=aN= 0 . It remains to show that a0= 0 . This is an
immediate consequence ofN/summationdisplay
k=0akekx= 0 and the vanishing of ak, fork≥1 .
Since the functions 1 ,ex,e2x,..., are inCk[a,b] for anykwe have shown that these
spaces are infinite dimensional too. Moreover, the exact same proof also shows that the
set{eα1x,eα2x,...,eαNx}, whereα1,...,a Nare arbitrary distinct complex numbers, is
linearly independent. This fact will be needed later. Perhaps we shall present a different
proof - or several different ones - at that time. All of the other proofs still involve some
calculus - but that should be no surprise since we used calculus to define the exponential
function in the first place.
Not all spaces of functions are infinite dimensional. For example, the function space
A={f∈C[−1,1]:f(x) =a+bex, a, b∈R}has dimension 2. The functions f1(x) = 1
andf2(x) =exconstitute a basis for Abecause every f∈Acan be written in the form
f=a1f1+a2f2, wherea1anda2are real numbers. Another basis for Aisf3(x) = 1+ex
andf4= 2−ex. There are many ways to see this. One is to observe that f3+f4= 3
and 2f3−f4= 3ex. Thus iff(x) =a+bex∈A, thenf=a
3(f3+f4) +b
3(2f3−f4) =
(a
3+2b
3)f3+ (a
3−b
3)f4.
96 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE
The function space B={f∈C[−1,1]:f(x) =asin(x+α), α,a ∈R}, also has
dimension two, since f(x) = (acosα) sinx+ (asinα) cosx=a1sinx+a2cosx. Thus
f1(x) = sinxandf2(x) = cosxform a basis. Actually, we have only shown that f1
andf2spanB, but not that they are linearly independent. You can settle that point
yourselves.
A few more remarks should be added. If Ais a subspace of an ndimensional space
Wn, we would like to enlarge a basis {e1,...,e k}forAto a larger basis {e1,...,e n}for
all ofW. SinceA⊂Wn, it is clear that k= dimA≤n. IfA=Wn, we are done since
{e1,...,e k}already span Wn. Otherwise there is some element ek+1inWnwhich is
not inA. LetA1= span {e1,...,e k+1} ⊂Wn. IfA1=Wn, then {e1,...,e k+1}form
a basis for Wn. Otherwise there is some element ek+2inWnwhich is not in A1. Form
A2=sp{e1,...,e m+2}. Repeat this process until you finally get a basis for all Wn. Only
a finite number of steps are needed since the dimension of Wnis finite. This proves
Theorem 2.13 . IfAis a subspace of (finite dimensional) space W, then any basis for
Acan be extended to a basis that spans all of W.
Consider a subspace Aof a linear space W. Somehow we would like to discuss - and
give a name to - the part A/primeofVwhich is not in A. We would like A/primeto be a subspace
ofVsuch that the only element of VwhichAandA/primeshare is 0, and such that every
element in Vcan be written as the sum of an element in Aand an element in A/prime.
Definition: . LetAbe a subspace of the linear space V. Acomplementary subspace A/prime
ofAis a subset of Vwith the properties
1.A/primeis a subspace of V,
2. IfX∈V, thenX=X1+X2, whereX1∈AandX2∈A/prime.
3.A∩A/prime= 0 . (The zero vector, notthe empty set).
Our first task is to prove
Theorem 2.14 . Every subspace A⊂Vhas at least one complement A/prime.
.
Proof: Let{e1,...,e m}be a basis for A, and {e1,...e m,em+1,...,e n}an extension
to a basis for V. We shall verify that A/prime=sp{em+1,...,e n}satisfies both criteria.
Now ifX∈AandX∈A/prime, then we can write X=a1e1+...+amem∈A, and
X=am+1ee+1+...+anen∈A/prime. Subtracting these equations, we find
0 =a1e1+...+amem−am+1em+1−...−anen.
But since {e1,...,e n}is a basis for V, the elements are linearly independent. Thus
a1=a2=...=am=am+1=...=an= 0 , soX= 0 . Therefore A∩A/prime= 0 .
Furthermore, if X∈Vsince {e1,...,e n}is a basis for V, then
X=n/summationdisplay
j=1cjej=m/summationdisplay
j=1cjej+n/summationdisplay
j=m+1cjej.
Thus we just let X1=c1e1+...+cmem∈AandX2andX2=cm+1em+1+...+cnen.
It is easy to see that the above construction of A/primeisindependent of the basis chosen for
A. This is because the construction of em+1,...,e n(Theorem ??) did not depend on the
particular basis for A. That construction only utilized the fact that we can pick elements
2.4. BASES AND DIMENSION 97
notinA. However, the construction of A/primedoes depend on which elements em+1,...e n
(not inA) we pick. For example, let V=R2, andAbe some one dimensional subspace.
Then we pick e, as any vector in A, ande2as any vector not in A. The resulting
complement A/primeis then the span of e2. But {e1}could have been extended to a basis
forVby choosing another vector ˜ e2/ownerA. This determines a different complement ˜A/primeof
A. A subspace has many possible complements. This ambiguity will not bother us since
we shall only use the properties of a particular complement which do not depend on which
particular complement is chosen. The dimension of the complement is such a property. It
only depends on the dimension of the subspace Aand the larger space V, and has the
reasonable formula dim A/prime= dimV−dimA, which we now prove.
Theorem 2.15 . IfAis a subspace of a linear space Vand ifA/primeis any complement of
A, then
dimA+ dimA/prime= dimV.
Thus, the dimension of A/primeis determined by AandValone.
Proof: The dimAand dimVare given data. We shall compute dim A/prime. Since the
union of a basis for Awith a basis for any A/primespansV(property 2), it is clear that
dimA+ dimA/prime≥dimV. However Aand anyA/primeintersect only at the origin (property
3) and are subspaces of V. Thus the union of their bases can span at most V, that is,
dimA+ dimA/prime≤dimV. These two inequalities prove the theorem.
REMARK. Some people refer to dim A/primeas the codimension ofA(complementary
dimension). In this way they avoid mentioning A/primeat all. The last theorem can be written
as dimA+ codimA= dimV.
A simple result closes the chapter.
Theorem 2.16 . IfAis a subspace of VandA/primeis a complement of A, then forX∈V
the decomposition X=X1+X2, X 1∈A, X 2∈A/primeis unique.
Proof: Assume there are two decompositions, X=X1+X2andX=˜X1+˜X2. Then
˜X1+˜X2=X1+X2or˜X1−X1=X2−˜X2. However the left side of this equation is in
Awhile the right is in A/prime. The only element in both AandA/primeis 0. Thus ˜X1=X1and
˜X2=X2.
a figure goes here
EXERCISES
(1) (a) Let A={X∈R2:x1= 0}. Find a basis for Aand extend it to a basis for all
ofR2. Use this to define a complement A/primeofA. SketchAandA/prime. Extend
the same basis for Ain a different way to a basis for all of R2. Use this to
define another complement ˜A/primeofA. Sketch ˜A/prime.
(b) Find a basis for the subspace A={X∈R2:x1+x2+x3= 0}. Extend
this basis to one for all of R3. Define a complement A/primeofAinduced by this
extension. Write X= (−1,0,7) asX=Y1+Y2whereY1∈AandY2∈A/prime.
(2) (a) Let A={p∈P2:p(0) = 0 }. Find a basis for Aand extend it to a basis for
all of P2. DefineA/primeinduced by this extension. Is the particular polynomial
p(x) = 1 +x2inA? inA/prime? Writep(x) asp(x) =q1(x) +q2(x) where
q1(x)∈A, q 2(x)∈A/prime.
98 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE
(b) LetA={p∈P2:p(1) = 0 }. Find a basis for Aand extend it to a basis for
all of P2.
(3) LetAbe a subspace of a linear space V. Show by an example that a basis for V
need not contain a basis for A.
(4) If dimV=nandV=sp{X1,...,X n}, prove that X1,...,X nare linearly inde-
pendent.
(5) LetV=P4andAthe subspace spanned by 1 ,x2andx4. Find three different
subspaces complementary to A(you may specify a subspace by giving a basis for it).
After all this about bases, it is probably best to notify you that properties of linear
spaces are best defined and proved without introducing a particular basis . As soon as you
define a property of a linear space in terms of a basis, you must then prove that the property
is intrinsic to the space itself and does not depend upon the basis you choose. We met this
problem in defining the dimension in terms of a basis - and were consequently forced to
prove Theorem ? which stated that the property really only depended on the space itself,
not on the basis chosen.
This, in fact, corresponds to one of the major requisites for laws of physics: they should
not depend upon the particular coordinate system you choose (picking a coordinate system
is equivalent to picking a basis). Moreover, the laws should not depend on the units you
choose for each axis of the coordinate system. But these are long, involved questions which
must be investigated deeply to make our remarks precise.
One should, however, distinguish theoretical issues from computational ones. In theo-
retical questions , the rule is never pick a specific basis unless there is no way out . On the
other hand, for computational questions you must always pick a basis . Just as in physics,
on order to perform any measurements, you must pick some specific coordinate system
and specific units. If the theoretical foundations are firm, then you can feel confident that
no matter what choice of basis you make, the essential nature of the results will remain
unchanged.
As an example, let us consider a point Pand two different fixed coordinate systems in
the plane of this paper. You should feel that any motion of the point Pcan be described
adequately in either coordinate system - and that when the observers in the two coordinate
systems get together and discuss the motion of P, they will agree as to what happened. A
common example is the meeting of two people from countries using different units of money.
Exercises
(1) Prove that any n+ 1 elements in a linear space of dimension nmust be linearly
dependent.
(2) Prove that Pnhas dimension n+ 1 .
(3) Since a basis for a linear space of dimension nmust contain exactly nelements,
all one must test is that the nelements which are candidates for a basis are lin-
early independent - or equivalently that they span the space. Show that the vectors
{X1,...,X n}form a basis for Rnif and only if e1,e2,...,e ncan all be expressed
as a linear combination of the {X1,...,X n}.
2.4. BASES AND DIMENSION 99
(4) Use Exercise 3 to determine which of the following sets from bases for R3.
(a)X1= (1,1,0), X 2= (1,0,1), X 3= (0,1,1).
(b)X1= (1,0,1), X 2= (1,1,1).
(c)X1= (1,0,1), X 2= (1,1,0), X 3= (0,−1,1).
(d)X1= (1,1,1), X 2= (1,2,3), X 3= (17,3,9), X 4= (−2,7,−1).
(e)X1= (−1,0,2), X 2= (1,1,1), X 3= (1
2,1
3,−1).
(5) Prove that the subspace of functions in C[0,π] which vanish at x= 0 and at
x=πis infinite dimensional by showing that the functions f1(x) = sinx, f 2(x) =
sin 2x,...,f k(x) = sinkx,... are all linearly independent. [Hint: Assume that 0 =
N/summationdisplay
k=1aksinkx, for arbitrary Nand show that all the ak’s must be zero by multiplying
both sides by sin nxand utilizing the important formula
/integraldisplayπ
0sinnxsinkxdx =/braceleftbigg0, k/negationslash=n
π
2, k=n./bracerightbigg
.]
(6) LetC∗[a,b] denote the set of all complex-valued functions f(x) =u(x) +iv(x)
which are continuous for x∈[a,b] . The complex number field Cis the field of
scalars for C∗. What is the dimension of the subspace A={f∈C∗[−π,π]:f(x) =
aeix+be−ix, a, b∈C}? Show that f1(x) = cosxandf2(x) = sinxconstitute a
basis forA. [Hint: Use (?) on p. ?].
(7) Which of the following sets of vectors form a basis for R4?
(a)X1= (1,0,0,5), X 2= (0,3,2,6), X 3= (0,0,1,2), X 4= (0,0,0,1).
(b)X1= (1,6,7,0), X 2= (−2,2,5,0), X 3= (4,5,6,0), X 4= (7,8,3,0).
(c)X1= (1,2,5,7), X 2= (4,9,11,8), X 3= (6,3,12,2), X 4= (3,−4,7,6),
X5= (0,0,0,1).
(d)X1= (1,2,3,4), X 2= (0,2,3,4), X 3= (0,0,3,4), X 4= (0,0,0,4).
(8) Find a basis for the following subspaces.
(a)A={X∈R2:x1+x2= 0}
(b)B={X∈R3:x1+x2+x3= 0}
(c)C={p∈P3:p(0) = 0 }
(d)D={p∈P3:p(1) = 0 }
(e)E={u∈C1[−1,1]:u/prime−u= 0}
(f)F={u∈C1[−1,1]:u/prime+ 2u= 0}.
100 CHAPTER 2. LINEAR VECTOR SPACES: ALGEBRAIC STRUCTURE
Chapter 3
Linear Spaces: Norms and Inner
Products
3.1 Metric and Normed Spaces
Until now we have been contented with being able to add two elements X1andX2of a
linear space, and to multiply them by scalars, aX. Since only these algebraic operations
have been defined, only algebraic questions could have been raised and answered. Notably
absent were any mention of convergence, because the idea of one element of a linear space
being “close” to another was not defined. In this chapter we shall introduce a distance
ormetric structure into linear spaces. Instead of lingering in the realm of generalities, we
shall define metric and norm in this first section and devote the balance of the chapter
to a particular kind of metric which generalizes the “Pythagorean distance” of ordinary
Euclidean space. Fourier series supply a wonderful and valuable application.
Our first notion of distance, that of a metric , makes sense for elements X, Y, Z of an
arbitrary set S. The idea is to define the distance d(X,Y) between any two elements of
S. This distance is a function which assigns to every pair of points ( X,Y) apositive real
numberd(X,Y) called the “distance between XandY”.
Definition . LetSbe a non-empty set. A metric onSis a real-valued function d:
S×S→R, whereX, Y∈S, which has the three properties:
i)d(X,Y)≥0. d (X,Y) = 0⇐⇒X=Y
ii) (symmetry) d(X,Y) =d(Y,X) ,
iii) (triangle inequality) d(X,Z)≤d(X,Y) +d(Y,Z) .
Well, they certainly are reasonable requirements for any function we intend to think of
as measuring distance.
Examples .
(1) This first example is trivial but acts as an important check on intuition. With it, you
see that every non-empty set can be regarded as a metric space with the following
metric
d(X,Y) =/braceleftbigg0,ifX=Y
1,ifX/negationslash=Y.
A moments reflection will show that this is a metric—but not too useful since it is so
coarse.
101
102 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
(2) For the real line, R, with the usual definition of absolute value we define d(X,Y) =
|X−Y|, which is clearly a metric.
(3) Another less common metric may be given to R. We define d(X,Y) =|X−Y|
1+|X−Y|.
Only the triangle inequality is not evident—and that involves some algebra. This
metric has the property that the distance between any two points is always less than
one,d(X,Y)<1 for allX,Y∈R.
(4)Rncan be endowed with many metrics. Let X= (x1,x2,...,x n) ,Y= (y1,...,y n)
andZ= (z1,...z n) be arbitrary points in Rn. The metric you most expect is the
Euclidean distance
d(X,Y) = [(x1−y1)2+...+ (xn−yn)2]1/2= [n/summationdisplay
k=1(xk−yk)2]1/2
Again, only the triangle inequality is not obvious. It is a consequence of the Cauchy-
Schwarz inequality/parenleftBiggn/summationdisplay
k=1xkyk/parenrightBigg2
≤n/summationdisplay
k=1x2
kn/summationdisplay
k=1y2
k, (3-1)
which in turn is an immediate consequence of the algebraic identity
/parenleftBiggn/summationdisplay
k=1xkyk/parenrightBigg2
=n/summationdisplay
k=1x2
kn/summationdisplay
k=1y2
k−1
2n/summationdisplay
i=1n/summationdisplay
j=1(xiyj−xjyi)2.
And now the triangle inequality. Let ak=xk−yk, andbk=yk−zk. Then
xk−zk=ak+bk. Thus, using Cauchy- Schwarz in the second line below, we find
that
[d(X,Z)]2=n/summationdisplay
k=1(ak+bk)2=n/summationdisplay
k=1a2
k+ 2n/summationdisplay
k=1akbk+n/summationdisplay
k=1b2
k
≤n/summationdisplay
k=1a2
k+ 2/bracketleftBiggn/summationdisplay
k=1a2
kn/summationdisplay
k=1b2
k/bracketrightBigg1/2
+n/summationdisplay
k=1b2
k
=
/parenleftBiggn/summationdisplay
k=1a2
k/parenrightBigg1/2
+/parenleftBiggn/summationdisplay
k=1b2
k/parenrightBigg1/2
2
= [d(X,Y) +d(Y,Z)]2,(3-2)
so
d(X,Z)≤d(X,Y) +d(Y,Z).
Another proof of the Schwarz and triangle inequalities for this metric will be given
later in the chapter.
(5) A second metric for R/multicloseleftis
d(X,Y) =n/summationdisplay
k=1|xk−yk|
The axioms for a metric are easily verified.
3.1. METRIC AND NORMED SPACES 103
(6) A third metric for Rnis
d(X,Y) =/bracketleftBiggn/summationdisplay
k=1|xk−yk|p/bracketrightBigg1/p
, 1≤p<∞.
Example 4 is the special case p= 2 , while example 5 is the special case p= 1 . And
again, all but the triangle inequality are obvious. However the triangle inequality,
called Minkowski’s inequality in this general case, is not simple. We shall not prove it
here. Perhaps it will appear as an exercise later.
(7) The usual metric for C[a,b] is the uniform metric
d(f,g) = max
a≤x≤b|f(x)−g(x)|.
Geometrically, this distance is the largest vertical distance between the graphs of f
andgfor allx∈[a,b] .
(8) The space L1[a,b] of functions whose absolute value is integrable has the “natural”
metric
d(f,g) =/integraldisplayb
a|f(x)−g(x)|dx,
which can be interpreted as the total area between the two curves. Since every function
which is continuous for x∈[a,b] is integrable there, i.e., C[a,b]⊂L1[a,b] , this metric
is another metric for C[a,b] .
(9) For the function space C1[a,b] , the standard metric is
d(f,g) = max
a≤x≤b|f(x)−g(x)|+ max
a≤x≤b/vextendsingle/vextendsinglef/prime(x)−g/prime(x)/vextendsingle/vextendsingle
The metric for Ck[a,b] is defined similarly.
There are many theorems one can prove about metric spaces (a metric space is a set
Son which a metric is defined). Look in any book on general topology (or point set
topology, as it is often called) and you will find more than enough to satisfy you. For
most of our purposes metric spaces are too general. Normed linear spaces will suffice.
The norm /bardblX/bardblof an element Xin a linear space Vis the “distance” of Xfrom
the origin—the 0 element of V.
Definition. LetVbe a linear space over the real or complex field. If to every element
X∈Vthere is associated a real number /bardblX/bardbl, the norm ofX, which has the three
properties
i)/bardblX/bardbl ≥0. /bardblX/bardbl= 0⇐⇒X= 0
ii)/bardblaX/bardbl=|a| /bardblX/bardbl(homogeneity), ais a scalar,
iii)/bardblX+Y/bardbl ≤ /bardblX/bardbl+/bardblY/bardbl, (triangle inequality),
then we say that Vis anormed linear space .
How does a norm differ from a metric?
First of all, a norm is only defined on a linear space (sinceaXandX+Yappear in
the definition) whereas a metric may be defined on any set (cf. example 1 above). But if we
104 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
restrict our attention to linear spaces, how do the concepts of norm and metric differ? Every
normed linear space can be made into a metric space in such a way that /bardblX/bardblis indeed the
distance of Xfrom the origin ,d(X,0) =/bardblX/bardbl. The explicit formula for d(X,Y) should
surprise no one
d(X,Y) =/bardblX−Y/bardbl.
It is easy to check that d(X,Y) is a metric. Thus every normed linear space has a “natural”
metric induced upon it. However, a linear space which has a metric need not be a normed
linear space . For example in R, the linear space of the real numbers, the metric of example
3
d(X,Y) =|X−Y|
1 +|X−Y|
is not associated with a norm because axiom ii) for a norm is not satisfied.
Of the examples considered earlier, all but the first and third metrics arise from norms,
in the sense that
d(X,Y) =d(X−Y,0) =/bardblX−Y/bardbl.
By far the most common norm in Rnis that given by the Pythagorean theorem (ex-
ample 4). Then
/bardblX/bardbl=/radicalBig
x2
1+x2
2+···+k2n=/parenleftBiggn/summationdisplay
k=1x2
x/parenrightBigg1/2
and the induced metric is
d(X,Y) =/bardblX−Y/bardbl=/parenleftBiggn/summationdisplay
k=1(xk−yk)2/parenrightBigg1/2
For obvious historical reasons, we shall refer to R2with this Pythagorean norm as Eu-
clideann-space , and denote it by En. Note that Enis a linear space with a particular way
of measuring length specified. A metric removes the floppiness from Rn, giving the addi-
tional structure needed to investigate those geometrical concepts which utilize the notion
of distance.
Once we have a norm (or metric) it becomes possible to discuss convergence of a se-
quence of elements.
Definition : IfVis a normed linear space, the sequenceXn∈Vconverges toX∈Vif,
given any/epsilon1>0 , there in an Nsuch that
/bardblXn−X/bardbl</epsilon1 for alln>N.
As an example, we shall prove the sample
Theorem 3.1 . A sequence of points Xn= (x(n)
1,x(n)
2,...x(n)
k)inEkconverges to the
pointX= (x1,...,x k)inEkif and only if each component x(n)
jconverges to its respective
limit, lim
n→∞x(n)
j=xj,j= 1,...,k .
Proof: i)Xn→X⇒x(n)
j→xj. This is a consequence of the trivial inequality
/vextendsingle/vextendsingle/vextendsinglex(n)
j−xj/vextendsingle/vextendsingle/vextendsingle≤/radicalBig
(x(n)
1−x1)2+···+ (x(n)
k−xk)2=/bardblXn−X/bardbl;
3.1. METRIC AND NORMED SPACES 105
for if /bardblXn−X/bardbl< /epsilon1forn > N , then/vextendsingle/vextendsingle/vextendsinglex(n)
j−xj/vextendsingle/vextendsingle/vextendsingle< /epsilon1forn > N too. Thus x(n)
j→xj.
If the subscripts are cluttering up the proof, go through it again in a special case, say
x(n)
2→x2.
ii)x(n)
j→xj⇒Xn→X. By hypothesis, given any /epsilon1 > 0 , there are numbers
N1,N2,...N ksuch that/vextendsingle/vextendsingle/vextendsinglex(n)
1−x1/vextendsingle/vextendsingle/vextendsingle< /epsilon1, for alln > N 1,/vextendsingle/vextendsingle/vextendsinglex(n)
2−x2/vextendsingle/vextendsingle/vextendsingle< /epsilon1 for alln >
N2,...,/vextendsingle/vextendsingle/vextendsinglex(n)
k−xk/vextendsingle/vextendsingle/vextendsingle< /epsilon1 for alln > N k. PickN= max(N1,N2,...,N k) . ThisNwill
work for all the x(n)
j’s, that is, for every j,
/vextendsingle/vextendsingle/vextendsinglex(n)
j−xj/vextendsingle/vextendsingle/vextendsingle</epsilon1 for alln>N.
Thus
/bardblXn−X/bardbl=/radicalBig
(x(n)
1−x1)2+...+ (x(n)
k−xk)2
</radicalbig
/epsilon12+...+/epsilon12=/epsilon1√
k,for alln>N.(3-3)
Sincekis a fixed finite number, this shows that /bardblXn−X/bardblmay be made arbitrarily small
by picking nbig enough, so Xndoes converge to X.
Example . In E4, the sequence Xn= (n
n+1,2,−1
n,0) converges to X= (1,2,0,0) since
n
n+1→1,2→2,−1
n→0 , and 0 →0 .
A useful elementary result is
Theorem 3.2 . IfVis a normed linear space, and if Xn→X, Y n→YinV, then for
any scalars aandb, aX n+bYn→aX+bY.
Proof: There are essentially no changes from the case of R1. We must show that /bardblaXn+
bYn−aX−bY/bardblcan be made arbitrarily small by picking nlarge enough. One application
of the triangle inequality
/bardblaXn+bYn−aX−bY/bardbl ≤ /bardblaXn−aX/bardbl+/bardblbYn−bY/bardbl,
and the homogeneity of a norm, yields
≤ |a| /bardblXn−X/bardbl+|b| /bardblYn−Y/bardbl.
BecauseXn→XandYn→Y, ifn>N 1, then /bardblXn−X/bardbl</epsilon1. Also, ifn>N 2, then
/bardblYn−Y/bardbl</epsilon1. PickN= max(N1,N2) . Thus
/bardblaXn+bYn−aX−bY/bardbl<|a|/epsilon1+|b|/epsilon1= (|a|+|b|)/epsilon1, n>N,
and the desired convergence is proved.
For a given linear space V, there may be two (or even more) norms defined, say /bardbl /bardbl
and/bardbl /bardbl 1to distinguish them. Why carry them both around? First of all, a sequence
may converge in one norm and not in the other. Second, even if both norms yield the same
convergent sequences, one norm may be more convenient in some particular computation.
Example . Consider the linear space C[−1,1] of functions f(x) continuous for x∈[−1,1] ,
with the two norms (Examples 7 and 8)
/bardblf/bardbl∞= max
−1≤x≤1|f(x)|;/bardblf/bardbl1=/integraldisplay1
−1|f(x)|dx,
106 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
that is, the uniform norm and the L1norm. We shall exhibit a sequence of functions which
converge in the second norm but not in the first. Let fn(x) be
fn(x) =
0, x ∈[−1,−1
n2]
n3(x+1
n2)x∈[−1
n2,0]
−n3(x−1
n2)x∈[0,1
n2]
0 x∈[1
n2,1]
Then by inspection from the graph ( /bardbl /bardbl 1is the area), we see that /bardblfn/bardbl∞=n, and
/bardblfn/bardbl1=1
n. Asn→ ∞,/bardblfn−0/bardbl1→0 so thatfn→0 in theL1norm. On the other
hand, /bardblfn/bardbl∞→ ∞ so the limit does not exist in the uniform norm. If you look at the
graph,fnis zero except for a spike in the interval [ −1
n2,1
n2] . Asn→ ∞ , the function is
zero in essentially the whole interval, except for the bit around the origin where it blows
up—but it blows up slowly enough that the area under the curve tends to zero.
However, we can prove the
Theorem 3.3 . Letfnandfbe continuous functions, n= 1,2,.... Iffn→fin the
uniform norm, then also fn→fin theL1norm.
Remark . We have just seen that the converse is false.
Proof: An immediate consequence of the
Lemma 3.4 /bardblfn−f/bardbl1≤(b−a)/bardblfn−f/bardbl∞
Proof:
/bardblfn−f/bardbl1=/integraldisplayb
n|fn(x)−f(x)|dx≤/integraldisplayb
a/bardblfn−f/bardbl∞dx
=/bardblfn−f/bardbl/integraldisplayb
adx= (b−a)/bardblfn−f/bardbl∞(3-4)
Exercises
(1) In the set Z, defined(m,n) =|m−n|where |x|is ordinary absolute value. Prove
thatd(m,n) is a metric.
(2) Suppose that d1(X,Y) andd2(X,Y) are both metrics for a set S, whereX,Y∈S.
a). Show that [ d1(X,Y)]2isnot, in general, a metric. b). Prove that d1+d2and/radicalbig
d2
1+d2
2are also metrics for S.
(3) Prove that the function d(X,Y) =|X−Y|
1+|X−Y|, X, Y ∈R, is a metric, but that it is not
a norm on R.
(4) LetX= (x1,...,x k)∈Rk. Define /bardblX/bardbl∞= max
1≤l≤k|xl|.
(a) Prove that /bardblX/bardbl∞is a norm for Rk, and write down the induced metric.
(b) Let /bardblX/bardbl1=k/summationdisplay
l=1|xl|, and /bardblX/bardbl2=/parenleftBiggk/summationdisplay
l=1|xl|2/parenrightBigg1/2
.
Prove
/bardblX/bardbl∞≤ /bardblX/bardbl2≤ /bardblX/bardbl1≤k/bardblX/bardbl∞.
3.2. THE SCALAR PRODUCT IN E2107
(c) Consider the sequence Xn= (1−1
n,−7,1
n2) inR3.
In which of the norms /bardbl /bardbl ∞,/bardbl /bardbl 2,/bardbl /bardbl 1does it converge, and to what?
(5) LetXnbe a sequence of elements in a normed linear space V(not necessarily finite
dimensional). Prove that if Xn→X, then the sequence Xnis bounded in norm (a
sequenceXnin a normed linear space is bounded if there is an M∈Rsuch that
/bardblXn/bardbl ≤Mfor alln). [Hint: Compare with Theorem 6, page ??].
(6) Compute the /bardbl /bardbl 1,/bardbl /bardbl 2,and/bardbl /bardbl ∞(cf. Ex. 4) norms of the following vectors in
R3.
a)X= (1,2,2) , b)Y= (2,−2,1) , c)Z= (0,3,−4) , d)W= (0,−1,0) .
(7) Compute the /bardbl /bardbl 1,/bardbl /bardbl 2,and/bardbl /bardbl ∞norms of the following functions for the interval
[−1,1] .
a)f(x) =−2x+ 3 b) g(x) = sinπx, c) hn(x) =xn
d) Asn→ ∞ , does the sequence hnconverge in any of these three norms?
(8) Which of the following define norms for the given linear spaces?
a) For R3,[X] =x2
1+x2
x+x2
3
b) For P3,[p] = max
0≤x≤1p(x)
c) For P3,[p] = max
0≤x≤1|p(x)|
d) For R3,[X] =|x1|+|x2|
e) For R4,[X] =/radicalbig
1 +x2
1+x2
2+x2
3+x2
4.
(9) Prove that [ X] =/radicalbig
x2
1+x2
2defines a norm for R2(some algebraic fortitude will be
needed to prove the triangle inequality).
3.2 The Scalar Product in E2
In Euclidean space E2—which we remind you is R2with the Euclidean norm /bardblX/bardbl=/radicalbig
x2
1+x2
2—one can introduce many geometric concepts and develop a corresponding geo-
metric theory. Most important of these concepts is that of angle—especially orthogonality
(perpendicularity). It turns out that these ideas generalize almost immediately to all En,
and even to some exceedingly important infinite dimensional spaces. This section is devoted
to the most simple situation: E2. Please look at the pictures.
To begin, we introduce the scalar product (also called the dot product , orinner product )
of two vectors XandY.
Definition: If XandYare two vectors in E2, their scalar product /angbracketleftX, Y/angbracketright(sometimes
writtenX·Y) is defined by
/angbracketleftX, Y/angbracketright=/bardblX/bardbl /bardblY/bardblcosθ,
whereθis the angle between XandY.
Notice that the scalar product of two vectors is a real number, a scalar, notanother
vector. We need not specify the direction in which θis measured, counterclockwise or
clockwise, since cos θ= cos( −θ) . Further, we can use either the acute or obtuse angle
108 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
betweenXandYsince cos(2 π−θ) = cosθ. It is important to observe that the scalar
product of two vectors is defined independent of any coordinate system.
We are immediately led to some simple consequences.
Lemma 3.5 . Two vectors XandYare orthogonal if and only if /angbracketleftX, Y/angbracketright= 0
Proof: IfXandYare orthogonal, the angle θbetween them isπ
2, so/angbracketleftX, Y/angbracketright=
/bardblX/bardbl /bardblY/bardblcosπ
2= 0 . In the other direction, if /angbracketleftX, Y/angbracketright= 0 , then /bardblX/bardbl /bardblY/bardblcosθ= 0 . If
neither /bardblX/bardblnor/bardblY/bardbl= 0 , then cos θ= 0 , that is θ=π
2or3π
2. ThusXis orthogonal
toY. If/bardblX/bardblor/bardblY/bardbl= 0 , then one of them is just the point at the origin, the zero
vector. We agree to say that the zero vector is orthogonal to every other vector. With this
agreement, /angbracketleftX, Y/angbracketright= 0⇒X⊥Y, and the second half of the theorem is proved too.
There is a nice geometric interpretation of the scalar product. A hint of it appeared in
our last lemma. Let ebe a unit vector ,/bardble/bardbl= 1 . Consider /angbracketleftX, e/angbracketright=/bardblX/bardblcosθ(see figure).
This is the length of the projection ofXin the direction of e, or in other words, the length
of the projection of Xinto the subspace spanned by the single vector e. Strictly /angbracketleftX, e/angbracketrightis
not really a “length”, since “length” carries the implication of being positive, whereas the
real number /angbracketleftX, e/angbracketrightwill be negative if the projection “points” in the direction opposite to
e. We shall, however, allow ourselves this abuse of language. The vector U1which is the
projection of Xinto the subspace spanned by eisU1=/angbracketleftX, e/angbracketrighte.
IfYis a (non-zero) vector in E2which is not a unit vector, the above geometric idea
goes through by making the simple observation that given any vector Y/negationslash= 0 , the vector
e=Y//bardblY/bardblis a unit vector in the direction of Y.
Now you are certainly wondering how in the world we compute this scalar product.
You could take out your ruler, protractor and table of cosines—but we will present a more
convenient method. In order to compute this as is always the case, a particular basis must
be chosen. Then the vectors XandYcan be given explicitly in terms of the basis. Since
we want to show that the concepts are independent of any particular basis , you must relax
and be patient. Only after the theory has been exposed will we reveal how to compute in
terms of a given basis.
Theorem 3.6 (a)/angbracketleftX, X/angbracketright=/bardblX/bardbl2
(b)/angbracketleftX, Y/angbracketright=/angbracketleftY, X/angbracketright
(c)/angbracketleftaX, Y /angbracketright=a/angbracketleftX, Y/angbracketrightwherea∈R.
(d)/angbracketleftX, aY /angbracketright=a/angbracketleftX, Y/angbracketright, wherea∈R.
(e)/angbracketleftX+Y, Z/angbracketright=/angbracketleftX, Z/angbracketright+/angbracketleftY, Z/angbracketright
(f)/angbracketleftX, Y +Z/angbracketright=/angbracketleftX, Y/angbracketright+/angbracketleftX, Z/angbracketright
(g)|/angbracketleftX, Y/angbracketright| ≤ /bardblX/bardbl /bardblY/bardbl(Cauchy- Schwarz inequality)
Proof:
(a) Obvious since θ= 0 and cos 0 = 1 .
(b) Obvious since cos( −θ) = cosθ.
3.2. THE SCALAR PRODUCT IN E2109
(c) The vectors XandaXlie along the same line through the origin. There are two
cases,a>0 anda<0 (a= 0 is trivial). Ifa>0 , the angle θbetweenXandY
is identical to that between aXandY. Since /bardblaX/bardbl=a/bardblX/bardblfora>0 , this case is
proved, for
/angbracketleftaX, Y /angbracketright=/bardblaX/bardbl /bardblY/bardblcosθ=a/angbracketleftX, Y/angbracketright.
Ifa<0 , thenaXpoints in the direction opposite to X. Thus the angle θ1between
aXandYequalsπ−θ, whereθis the angle between XandY. The following
computation completes the proof:
/angbracketleftaX, Y /angbracketright=/bardblaX/bardbl /bardblY/bardblcosθ1=|a|/bardblX/bardbl/bardblY/bardblcos(π−θ)
=−|a| /bardblX/bardbl /bardblY/bardblcosθ=a/bardblX/bardbl /bardblY/bardblcosθ=a/angbracketleftX, Y/angbracketright(3-5)
(d) By (b) and (c) and (b) again we are done
/angbracketleftX, aY /angbracketright=/angbracketleftaY, X /angbracketright=a/angbracketleftY, X/angbracketright=a/angbracketleftX, Y/angbracketright.
(e) This is the most subtle part. We shall rely on the interpretation of the scalar product
/angbracketleftU, e/angbracketrightas the length of the projection of Uin the subspace spanned by e. First, let
e=Z//bardblZ/bardblbe the unit vector in the direction of Z. We shall show that /angbracketleftX+Y, e/angbracketright=
/angbracketleftX, e/angbracketright+/angbracketleftY, e/angbracketright. A picture is all that is needed now. Two situations are illustrated,
where both XandYare on the same side of the line perpendicular to eand
a figure goes here
whereXandYare on opposite sides of that line. The vector X+Yis found from
XandYby the parallelogram rule for addition. Interpreting the scalar product of
a vector with eas the length of the projection into the subspace (line) spanned by
e, we see (look) that we must prove
→
OP=→
OQ+→
OM .
But since→
OAand→
BC are on opposite sides of the same parallelogram, know that
→
OM=→
QPboth in magnitude and direction. The natural substitution yields
→
OP=→
OQ+→
QP,
which is indeed all we desired. Thus
/angbracketleftX+Y, e/angbracketright=/angbracketleftX, e/angbracketright+/angbracketleftY, e/angbracketright
To prove the general result for Z=/bardblZ/bardble, multiply the last equation by /bardblZ/bardbl, which
is a scalar. Then by part a we find
/angbracketleftX+Y,/bardblZ/bardble/angbracketright=/angbracketleftX,/bardblZ/bardble/angbracketright+/angbracketleftY,/bardblZ/bardble/angbracketright,
or
/angbracketleftX+Y, Z/angbracketright=/angbracketleftX, Z/angbracketright+/angbracketleftY, Z/angbracketright,
We are done.
110 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
(f) By parts (b), (e) and (b) again we obtain the result.
/angbracketleftX, Y +Z/angbracketright=/angbracketleftY+Z, X/angbracketright=/angbracketleftY, X/angbracketright+/angbracketleftZ, X/angbracketright=/angbracketleftX, Y/angbracketright+/angbracketleftX, Z/angbracketright
(g) Obvious since |cosθ| ≤1 . It is evident that equality occurs when and only when
cosθ±1 , that is, when XandYlie along the same line (possibly pointing in
opposite directions).
Ifeis a unit vector, we know how to find the projection U1of a given vector Xinto the
subspace spanned by e, it isU1=/angbracketleftX, e/angbracketrighte. Similarly, if Yis any vector—not necessarily
of length one, then since Y//bardblY/bardblis a unit vector in the direction of Y, the projection of X
into the subspace spanned by Yis/angbracketleftX, Y/ /bardblY/bardbl/angbracketrightY//bardblY/bardbl=/angbracketleftX, Y/angbracketrightY//bardblY/bardbl2. We can also find
the projection U2ofXinto the subspace orthogonal to the unit vector e. Since the sum
ofU1andU2must add up to X, X =U1+U2, we find that U2=X−U1=X−/angbracketleftX, e/angbracketrighte.
Thus, we have proved
Theorem 3.7 . IfXandYare any two vectors, /bardblY/bardbl /negationslash= 0, thenXcan be decomposed
into two vectors U1andU2, X=U1+U2such thatU1is in the subspace spanned by Y
andU2is in the orthogonal subspace. The decomposition is given by U1=/angbracketleftX, Y/angbracketrightY//bardblY/bardbl2
andU2=X− /angbracketleftX, Y/angbracketrightY//bardblY/bardbl2, so that
X=/angbracketleftX, Y/angbracketrightY
/bardblY/bardbl2+ (X− /angbracketleftX, Y/angbracketrightY
/bardblY/bardbl2).
Without further delay, we shall show how to compute the scalar product of two vectors.
In order to carry this out we must introduce a basis. Let X1andX2be any two vectors in
E2which span E2. Then every vector X∈E2can be written in the form X=a1X1+a2X2,
where the scalars a1anda2are determined uniquely by the vector X. Now it is most
convenient to have a basis whose vectors are i) orthogonal to each other and ii) have unit
length. Such a basis is called an orthonormal basis (orthogonal and normalized to have
unit length). In other words e1ande2are an orthonormal basis for E2if/bardblej/bardbl= 1 and
/angbracketlefte1, e2/angbracketright= 0 . This requirement is most conveniently stated by introducing the Kronecker
symbolδjk
δjk=/braceleftbigg0j/negationslash=k
1j=k.
Then the orthonormality property reads /angbracketleftej, ek/angbracketright=δjk, j,k = 1,2 . The notation is perhaps
excessive for this simple case, but will really be useful in our generalizations.
Therefore, let e1ande2be an orthonormal basis for E2, so that if X∈E2, X=
x1e1+x2e2.Fixthe basis throughout the ensuing discussion. Observe that x1andx2
can be computed in terms of X, and the basis vectors e1ande2, viz/angbracketleftX, e 1/angbracketright=/angbracketleftx1e1+
x2e2, e1/angbracketright=x1/angbracketlefte1, e1/angbracketright+x2/angbracketlefte2, e1/angbracketright=x1, since /angbracketlefte1, e1/angbracketright= 1 and /angbracketlefte1, e2/angbracketright= 0 . Similarly,
/angbracketleftX, e x/angbracketright=x2. Thus we have proved
Theorem 3.8 . If{ej}, j = 1,2, form an orthonormal basis for E2, then every vector
X∈E2can be written as X=2/summationdisplay
j=1xjej, wherexjis the length of the projection of X
into the subspace spanned by ej, x j=/angbracketleftX, e j/angbracketright.
3.2. THE SCALAR PRODUCT IN E2111
IfX=x1e1+x2e2andY=y1e1+y2e2are any two vectors in E2, then
/angbracketleftX, Y/angbracketright=/angbracketleftx1e1+x2e2, y1e1+y2e2/angbracketright
=/angbracketleftx1e1+x2e2, y1e1/angbracketright+/angbracketleftx1e1+x2e2, y2e2/angbracketright
=/angbracketleftx1e1, y1e1/angbracketright+/angbracketleftx2e2, y1e1/angbracketright+/angbracketleftx1e1, y2e2/angbracketright+/angbracketleftx2e2,, y 2/angbracketright
=x1y1/angbracketlefte1, e1/angbracketright+x2yx/angbracketlefte2, e1/angbracketright+x1y2/angbracketlefte1, e2/angbracketright+x2y2/angbracketlefte2, e2/angbracketright
=x1y1+ 0 + 0 +x2y2=x1y1+x2y2.
Now you see how easy it is to compute the scalar product of XandYin terms of the
representation from an orthonormal basis. Let us rewrite our result formally.
Theorem 3.9 . Let {ej}, j= 1,2, form an orthonormal basis for E2. IfX=2/summationdisplay
j=1xjej
andY=2/summationdisplay
j=1vjej, then
/angbracketleftX, Y/angbracketright=2/summationdisplay
j=1xjyj=x1y1+x2y2.
Some numerical examples should reassure you of the basic simplicity of the computation.
As our orthonormal basis in E2, we choose the vectors e1= (1,0) ande2= (0,1) . These
both have unit length, and are perpendicular (one is on the horizontal axis, the other on
the vertical axis). Let X= (−2,3) . ThenX=−2e1+ 3e2. Notice that −2e1and 3e2
are exactly the projections of Xinto the subspaces spanned by e1ande2respectively. If
Y= (1,−2) , then our theorem shows that
/angbracketleftX, Y/angbracketright= (−2)(−1) + (3)( −2) =−2−6 =−8.
From this computation we can reverse the geometric procedure and find the angle θbetween
XandY, for we know the formula
cosθ=/angbracketleftX, Y/angbracketright
/bardblX/bardbl /bardblY/bardbl.
In this example, /angbracketleftX, Y/angbracketright=−8,/bardblX/bardbl=√4 + 9 =√
13 and /bardblY/bardbl=√1 + 4 =√
5 . Thus
θ= cos−1(−8√
65) which can be evaluated by consulting your favorite numerical tables.
It is equally simple to check if two vectors are orthogonal. Let X= (2,−3) and
Y= (6,4) . Then /angbracketleftX, Y/angbracketright= (2)(6) + ( −3)(4) = 0 ; consequently XandYare orthogonal.
Another consequence is the law of cosines. Let X= (x1,x2) andY= (y1,y2) . Then
from the parallelogram construction, the length of the segment joining the tip of Xto the
tip ofYhas length /bardblY−X/bardbl. But
/bardblY−X/bardbl2=/angbracketleftY−X, Y−X/angbracketright
=/bardblX/bardbl2+/bardblY/bardbl2−2/angbracketleftX, Y/angbracketright
=/bardblX/bardbl2+/bardblY/bardbl2−2/bardblX/bardbl /bardblY/bardbl??θ.
One more example. We shall find the distance of the point P= (−3,2) from the coset
A={X= (x1,x2)∈E2:x1−2x2= 2}. Pick some point in X0inA, sayX0= (3,1
2) .
The distance dfromPtoAis then the length of the projection of the segment X0P
onto a line lorthogonal to A. First of all, we can replace the segment X0Pby a vector
from the origin 0 to the point Q=P−X0= (−6,3
2) , for the length of the projection of
¯0Qonto a line lorthogonal to Ais equal to the length of the projection of ¯X0Ponto
112 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
l(see figure). Now we have the vector Q= (−6,3
2) ; all we need to compute the desired
projection is another vector Northogonal to A, for thend=|/angbracketleftY, N//bardblN/bardbl/angbracketright|.
To find a vector Northogonal to A, we realize that Nwill also be orthogonal to
the subspace Sparallel to the coset AsoA=S+X0, where S={X= (x1,x2)∈
E2:x1−2x2= 0}. IfN= (n1,n2) andXis any element of S, sinceN⊥S, we must
have 0 = /angbracketleftX, N/angbracketright=x1n1+x2n2. However X∈Ssox1−2x2= 0 . We want the equation
x1n1+x2n2= 0 to hold for allpoints onx1−2x2= 0 , that is for all X∈S. This is only
possible if n1= 1·candn2=−2·c, wherecis any constant. Thus N=c(1,−2) and
/bardblN/bardbl=|c|√
5 . The distance dbetween the point Pand the coset Ais then
d=|/angbracketleftY, N//bardblN/bardbl/angbracketright|=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/angbracketleft(−6,3
2),c
|c|√
5(1,−2)/angbracketright/vextendsingle/vextendsingle/vextendsingle/vextendsingle
=/vextendsingle/vextendsingle/vextendsingle/vextendsingle(−6)(c
|c|√
5) + (3
2)(−2c
|c|√
5)/vextendsingle/vextendsingle/vextendsingle/vextendsingle=9√
5.(3-6)
This example contained a plethora of ideas. It would be wise to go through it again and list
the constructions and concepts used. The exercises will develop many of them in greater
generality.
Now you should try some problems on your own.
Exercises
(1) IfX= (3,4) andY= (5,−12) are two points in E2, find the angle between→
OX
and→
OY, where 0 is the origin.
(2) IfX= (3,−4) andY= (5,12) are two vectors in E2, find vectors U1∈span(Y)
andU2orthogonal to span( Y) such that X=U1+U2.
(3) Show that the vector N= (a1,a2) is perpendicular to the straight line whose equation
isa1x1+a2x2=c(you will have to supply the natural definition of what it means
for a vector to be perpendicular to a straight line).
(4) (a) Find the distance of the point P= (2,−1) from the coset A={X∈E2:x1+
x2=−2}.
(b) Find the distance between the two “parallel” cosets Adefined above and B=
{X∈E2:x1+x2= 1}. (Hint: Draw a figure and observe that P∈B) .
(5) (a) Prove that the distance dof the point P= (y1,y2) from the coset A={X∈
E2:a1x1+a2x2=c}is given by
d=|a1y1+a2y2−c|/radicalbig
a2
1+a2
2.
(b) Prove that the distance dbetween the two “parallel” cosets A={X∈E2:a1x1+
a2x2=c1}, andB={X∈E2:a1x1+a2x2=c2}is given by
d=|c1−c2|/radicalbig
a2
1+a2
2.
(Hint: If you use part (a) and are cunning, the derivation takes but one line).
3.3. ABSTRACT SCALAR PRODUCT SPACES 113
(6) (a) If it is known that /angbracketleftX, Y 1/angbracketright=/angbracketleftX, Y 2/angbracketright, and that /bardblX/bardbl /negationslash= 0 for a fixedX, can
you “cancel” Xfrom both sides and conclude that Y1=Y2? Reason?
(b) If it is known that /angbracketleftX, Y/angbracketright= 0 for everyX, can you conclude that Y= 0 ?
Reason?
(c) If it is known that /angbracketleftX, Y 1/angbracketright=/angbracketleftX, Y 2/angbracketrightforeveryX, can you conclude that
Y1=Y2? Reason?
(7) (a) Show that the vector
Z=/bardblX/bardblY+/bardblY/bardblX
/bardblX/bardbl+/bardblY/bardbl.
bisects the angle between the vectors XandY.
(b) Show that the vector /bardblX/bardblY+/bardblY/bardblXis perpendicular to the vector /bardblY/bardblX−
/bardblX/bardblY.
(8) Express the angle between an edge and a diagonal of a rectangle in terms of the scalar
product.
(9) Let two of the sides of a parallelogram be given by the vectors XandY. The
parallelogram theorem states that the sum of the squares of the sides is equal to the
sum of the squares of the diagonals, that is,
/bardblX+Y/bardbl2+/bardblX−Y/bardbl2= 2/bardblX/bardbl2+ 2/bardblY/bardbl2.
Prove this in two ways: i) using elementary geometry, and ii) using only the fact that
XandYare elements of a linear space, and the properties of the scalar product
contained in Theorem 4 (using 4a to define /bardbl /bardbl ).
(10) LetXbe any vector in E2, and letebe a unit vector. Define the vector U=ae,
wherea=/angbracketleftX, e/angbracketrightis the length of the projection of Xinto the subspace spanned by
e, andV=αe, whereαis any scalar. Prove that
/bardblX−V/bardbl2≥ /bardblX−U/bardbl2=/bardblX/bardbl2− /bardblU/bardbl2=/bardblX/bardbl2−a2.
This shows that in the subspace spanned by e, the vector closest to Xis the projec-
tionUofXinto that subspace.
(11) IfXis orthogonal to Y, prove the Pythagorean theorem /bardblX+Y/bardbl2=/bardblX/bardbl2+/bardblY/bardbl2
using only /bardblV/bardbl2=/angbracketleftV, V/angbracketrightand the properties of a scalar product in Theorem 4.
(12) LetXandYbe orthogonal elements of E2, with neither /bardblX/bardblnor/bardblY/bardblzero. Prove
thatXandYare linearly independent. Do notintroduce a basis.
3.3 Abstract Scalar Product Spaces
We shall turn the tables around. Whereas in the last section we defined the scalar product
geometrically and deduced its properties, in this section we define a scalar product space as
a linear space upon which a scalar product is defined, and the scalar product is stipulated
to have the properties deduced earlier. After presenting our abstract definition, we shall
give examples—other than E2—of scalar product spaces.
114 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
Definition. A linear space His called a real scalar product space if to every pair of
elementsX,Y∈His associated a real number /angbracketleftX, Y/angbracketright,the scalar product of XandY,
which has the properties
1./angbracketleftX, X/angbracketright ≥0 with equality if and only if X= 0 .
2./angbracketleftX, Y/angbracketright=/angbracketleftY, X/angbracketright
3./angbracketleftaX, Y /angbracketright=a/angbracketleftX, Y/angbracketright, a∈R
4./angbracketleftX+Y, Z/angbracketright=/angbracketleftX, Z/angbracketright+/angbracketleftY, Z/angbracketright
You should observe that the scalar product in E2does have these properties (Theorem
4). Using E2as our model, it is natural to define /bardblX/bardbl=/radicalbig
/angbracketleftX, X/angbracketrightand suspect that /bardbl /bardbl
is indeed a norm on the linear space H. This is true, but proving the triangle inequality
for this norm using only properties 1-4 will take some work. We shall do just that after
presenting
Examples
(1) LetX= (x1,...,x n) andY= (y1,...,y n) be points in the linear space R2. We
define
/angbracketleftX, Y/angbracketright=x1y1+x2y2+...+xnyn.
Only easy algebra is needed to verify that the real number /angbracketleftX, Y/angbracketrightsatisfies all of the
properties of a scalar product. It turns out (after we prove the triangle inequality)
that the natural norm /bardblX/bardbl=/radicalbig
/angbracketleftX, X/angbracketrightis the Euclidean norm, so this is E2.
(2) This example is the first hint that our abstractions are fruitful. Let the functions
f(x) andg(x) be points in the linear function space C[a,b] of real-valued functions
continuous for a≤x≤b. We define
/angbracketleftf, g/angbracketright=/integraldisplayb
af(x)g(x)dx.
You might be surprised; in any event let us verify that the real number /angbracketleftf, g/angbracketrightasso-
ciated with the pair of functions fandgdoes satisfy the four properties of a scalar
product.
(i)/angbracketleftf, f/angbracketright=/integraltextb
af2(x)dx. This is clearly non-negative and f= 0 implies that
/angbracketleftf, f/angbracketright= 0 . All we must show is that if /angbracketleftf, f/angbracketright=/integraltextb
af2(x)dx= 0 , then f= 0 .
By contradiction, assume f(x)/negationslash= 0 . Then there is some point x0∈[a,b] such
thatf(x0) =c/negationslash= 0 . Thus f2(x0) =c2>0 . Sincef—and hence f2—is
continuous, this means that f2is positive in some interval about x0(p. 29b,
Theorem I), so that/integraltextb
af2(x)dx> 0 , the desired contradiction.
(ii)/angbracketleftf, g/angbracketright=/integraltextb
af(x)g(x)dx=/integraltextb
ag(x)f(x)dx=/angbracketleftg, f/angbracketright.
(iii)/angbracketleftαf, g/angbracketright=/integraltextb
aαf(x)g(x)dx=α/integraltextb
af(x)g(x)dx=α/angbracketleftf, g/angbracketright, whereα∈R.
(iv)/angbracketleftf+g, h/angbracketright=/integraltextb
a(f(x) +g(x))h(x)dx
=/integraltextb
af(x)h(x)dx+/integraltextb
ag(x)h(x)dx
=/angbracketleftf, h/angbracketright+/angbracketleftg, h/angbracketright.
There. We did it. After we prove the triangle inequality for an abstract scalar
product space, the natural candidate for a norm /bardblf/bardblis a norm:
/bardblf/bardbl=/radicalBigg/integraldisplayb
af2(x)dx.
3.3. ABSTRACT SCALAR PRODUCT SPACES 115
I like this space very much. You will be meeting it often, becoming much more
intimate with its finer features. We shall—somewhat improperly—refer to this
linear space with the given scalar product as L2[a,b] . The name is improper
sinceL2[a,b] is customarily used for our space but with more general functions
and an extended notion of integration.
(3) Letf(x) andg(x) be inC[0,∞] . This time define
/angbracketleftf, g/angbracketright=/integraldisplay∞
0f(x)g(x)e−xdx.
Sincee−xis continuous and positive for all x, we are assured that /angbracketleftf, f/angbracketright ≥0 , with
equality if and only if f= 0 . The other properties of an inner product follow from
simple manipulations. Do them.
Remark : Complex scalar product spaces are defined similarly. For them, /angbracketleftX, Y/angbracketright
may be a complex number, and complex scalars are admitted. The only change in the
axioms is that property 2 is dropped in favor of
¯2./angbracketleftY, X/angbracketright=¯/angbracketleftX, Y/angbracketright,
where the bar means take the complex conjugate of the complex number /angbracketleftX, Y/angbracketright.
Since we shall not develop the theory far enough, our attention henceforth will be
restricted to real scalar product spaces.
The first order of business is to prove that the natural candidate for a norm /bardblX/bardbl=/radicalbig
/angbracketleftX, X/angbracketrightis in fact a norm for the linear space V. Only properties 1-4 may be used.
(1)/bardblX/bardbl ≥0 , with equality if and only if X= 0 . This follows immediately from the
corresponding property of /angbracketleftX, X/angbracketright.
(2)/bardblaX/bardbl=|a| /bardblX/bardbl. For /bardblaX/bardbl=/radicalbig
/angbracketleftaX, aX /angbracketright=/radicalbig
a2/angbracketleftX, X/angbracketright=|a|/radicalbig
/angbracketleftX, X/angbracketright=
|a| /bardblX/bardbl.
The proof of the triangle inequality
(3)/bardblX+Y/bardbl ≤ /bardblX/bardbl+/bardblY/bardblinvolves more labor. We shall first need to prove the Cauchy-
Schwarz inequality (cf. Theorem 4,g).
Theorem 3.10 (Cauchy-Schwarz inequality).
|/angbracketleftX, Y/angbracketright| ≤ /bardblX/bardbl/bardblY/bardbl.
Proof: If either /bardblX/bardblor/bardblY/bardblis zero, this is immediate. Thus, assume that neither /bardblX/bardbl
nor/bardblY/bardblis zero and define
U=X
/bardblX/bardbl, V =Y
/bardblY/bardbl,
so that both UandVare unit vectors, /bardblU/bardbl=/bardblV/bardbl= 1 . Then
0≤ /bardblU±V/bardbl2=/angbracketleftU±V, U±V/angbracketright
=/angbracketleftU, U/angbracketright ± /angbracketleftU, V/angbracketright ± /angbracketleftV, U/angbracketright+/angbracketleftV, V/angbracketright
=/bardblU/bardbl2±2/angbracketleftU, V/angbracketright+/bardblV/bardbl2,.
Since /bardblU/bardbl= 1 and /bardblV/bardbl= 1 , this shows ±/angbracketleftU, V/angbracketright ≤1 . Substituting for UandV, we
obtain the inequality sought:
|/angbracketleftX, Y/angbracketright| ≤ /bardblX/bardbl/bardblY/bardbl.
116 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
Theorem 3.11 (Triangle inequality) /bardblX+Y/bardbl ≤ /bardblX/bardbl+/bardblY/bardbl.
Proof: This is identical to that given in section 1. /bardblX+Y/bardbl2=/angbracketleftX+Y, X +Y/angbracketright=
/bardblX/bardbl2+ 2/angbracketleftX, Y/angbracketright+/bardblY/bardbl2. By Cauchy-Schwarz, /angbracketleftX, Y/angbracketright ≤ /bardblX/bardbl /bardblY/bardbl, so
/bardblX+Y/bardbl2≤ /bardblX/bardbl2+ 2/bardblX/bardbl /bardblY/bardbl+/bardblY/bardbl2= (/bardblX/bardbl+/bardblY/bardbl)2.
Now take square root of both sides to find
/bardblX+Y/bardbl ≤ /bardblX/bardbl+/bardblY/bardbl.
Nice, eh? See how clean everything is. We have proved
Theorem 3.12 . IfHis a scalar product space and we define /bardblX/bardbl=/radicalbig
/angbracketleftX, X/angbracketrightin terms
of the scalar product, then /bardbl /bardbl is a norm and His a normed linear space with that norm.
This special case where the norm is induced by a scalar product is called a pre-Hilbert space
(an honest Hilbert space has the additional property of being “complete”).
Let us state two easy algebraic consequences of our axioms for a scalar product. The
proofs are identical to those of Theorem 4 in the previous section.
Theorem 3.13
/angbracketleftX, aY /angbracketright=a/angbracketleftX, Y/angbracketright, a∈R (3-7)
/angbracketleftX, Y +Z/angbracketright=/angbracketleftX, Y/angbracketright+/angbracketleftX, Z/angbracketright, (3-8)
Needless to say, we hope you are still thinking in the geometric terms presented earlier.
In particular, the next definition should be reasonable.
Definition Two vectors X,Y are said to be orthogonal if/angbracketleftX, Y/angbracketright= 0 .
The Pythagorean theorem suggests
Theorem 3.14 . IfXandYare orthogonal, then
/bardblX±Y/bardbl2=/bardblX/bardbl2+/bardblY/bardbl2,
and conversely.
Proof: Both parts are an immediate consequence of the identity
/bardblX±Y/bardbl2=/angbracketleftX+Y, X +Y/angbracketright=/bardblX/bardbl2±2/angbracketleftX, Y/angbracketright+/bardblY/bardbl2.
Examples .
(1) LetX= (2,3,−1) andY= (1,−1,−1) be points in E3, where we use the scalar
product of example 1 in this section. Then /angbracketleftX, Y/angbracketright= 2·1 + 3( −1) + (−1)(−1) = 0
soXandYare orthogonal. Similarly X= (2,3,1,−1) andY= (3,−3,3,0) in
E4are orthogonal. A useful example is supplied by the vectors e1= (1,0,0,..., 0) ,
e2= (0,1,0,0,..., 0),...,e n= (0,0,..., 0,1) in En. These are orthonormal since
/angbracketleftek, ek/angbracketright= 1 , but /angbracketleftek, el/angbracketright= 0, k/negationslash=l, that is, /angbracketleftek, el/angbracketright=δkl.
3.3. ABSTRACT SCALAR PRODUCT SPACES 117
(2) Consider the functions Φ k(x) = sinkxinL2[−π,π] , wherek= 1,2,3,.... Then,
since sinθsin Ψ =1
2[cos(θ−Ψ)−cos(θ+ Ψ)] , we find that
/angbracketleftΦk,Φk/angbracketright=/integraldisplayπ
−πsin2kxdx =π
and fork/negationslash=l
/angbracketleftΦk,Φl/angbracketright=/integraldisplayp
i−πsinkxsinlxdx = 0.
as a computation reveals. Thus in L2[−π,π] the function sin kxis orthogonal to the
function sin lxwhenk/negationslash=l. The whole computation may be summarized by
/angbracketleftΦk,Φl/angbracketright=/angbracketleftsinkx,sinlx/angbracketright=πδkl.
It is only the factor πwhich does not allow us to say that the Φ kareorthonormal—
but that is easily patched up. Let ek(x) =sinkx√π. Then
/angbracketleftek, el/angbracketright=/angbracketleftsinkx√π,sinlx√π/angbracketright
=1
π/angbracketleftsinkx,sinlx/angbracketright,
or
/angbracketleftek, el/angbracketright=δkl.
Therefore the functions ek(x) =sinkx√πare orthonormal. Don’t attempt to imagine it.
Just keep on thinking of a big E2and all will be well.
So far we have discussed the notion of two vectors XandYbeing orthogonal. This
can be restated as one vector Xbeing orthogonal to the subspace Aspanned by Y, for
all vectors in Aare of the form aYwhereais a scalar, and /angbracketleftX, aY /angbracketright= 0⇐⇒ /angbracketleftX, Y/angbracketright= 0
since /angbracketleftX, aY /angbracketright=a/angbracketleftX, Y/angbracketright. One can also introduce the concept of a vector Xbeing
orthogonal to an arbitrary subspace A. Think of Aas being a plane (through the origin
of course).
Definition The vector Xisorthogonal to the subspace AifXis orthogonal to every
vector in the subspace A.
In practice, the usual way to check if Xis orthogonal to the subspace Ais as follows.
Pick some basis {Y1,Y2,...}forA. Then every Y∈Ais of the form
Y=/summationdisplay
akYk
(if the basis has an infinite number of elements—that is, if Ais infinite dimensional—
one should worry about convergence; however we shall ignore that issue for now). By the
algebraic rules for the scalar product, we find that
/angbracketleftX, Y/angbracketright=/angbracketleftX,/summationdisplay
akYk/angbracketright=/summationdisplay
ak/angbracketleftX, Y k/angbracketright.
Thus, X is orthogonal to the subspace A if X is orthogonal to every element in some basis
forA/angbracketleftX, Y k/angbracketright= 0 .
For example, if Ais thex1x2plane in E3, andXis the vector (0 ,0,1) , then
we can show that X= (0,0,1) is orthogonal to Aby showing it is orthogonal to both
118 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
the vector e1= (1,0,0) and toe2= (0,1,0) , sincee1ande2form a basis for A. The
computation /angbracketleftX, e 1/angbracketright= 0 and /angbracketleftX, e 2/angbracketright= 0 is immediate. Because Y1= (1,2,0) and
Y2= (1,−1,0) also form a basis for A, we could prove that Xis orthogonal to Aby
showing that /angbracketleftX, Y 1/angbracketright= 0 and /angbracketleftX, Y 2/angbracketright= 0 —which is equally simple.
A less obvious example is supplied by the function Ψ( x) = cosxwhich is orthogonal
to the subspace Aspanned by Φ 1(x) = sinx,Φ2(x) = sin 2x,..., Φn(x) = sinnxin
Lx(−π,π) . The proof is a consequence of the integration formula
/angbracketleftΨ,Φk/angbracketright=/integraldisplayπ
−πcosxsinkxdx = 0 for all k.
Even more general than a vector being orthogonal to a subspace is the idea that two
subspaces A and B are orthogonal , by which we mean that every vector in Ais orthogonal
to every vector in B. IfAis a subspace of a scalar product space H, then it is natural
to define the orthogonal complement A⊥ofAas the set
A⊥={X∈H:/angbracketleftX, Y/angbracketright= 0 for all Y∈A}
of vectorsXorthogonal to A, that is, orthogonal to every vector Y∈A.The setA⊥is
a subspace since it is closed under vector addition and multiplication by scalars (Theorem
2, p. 142).
Without fear of evoking surprise, we define the angle θbetween two vectors Xand
Yby the formula
cosθ=/angbracketleftX, Y/angbracketright
/bardblX/bardbl /bardblY/bardbl.
No matter what XandYare, this defines a real angle since the right side of the equation
is a real number between −1 and +1 (by the Cauchy-Schwarz inequality). To be honest,
there is little use for the concept of angles other than right angles. In E3the formula has
some use, but is totally unused for more general scalar product spaces.
If we are given a set of linearly independent vectors {X1,X2,...}which span a linear
scalar product space H, how can we construct an orthonormal set {e1,e2,...}which also
spans the space? The process is carried out inductively. Let e1=X1
/bardblX1/bardbl. Now we want a
unit vector e2orthogonal to e1. A reasonable candidate is
˜e2=X2− /angbracketleftX2, e1/angbracketrighte1,
which isX2with the projection of X2ontoe1subtracted off (see fig.) This vector ˜ e2is
orthogonal to e1since /angbracketleft˜e2, e1/angbracketright= 0 . We divide by its length to obtain the unit vector e2,
e2=X2− /angbracketleftX2, e1/angbracketrighte1
/bardblX2− /angbracketleftX2, e1/angbracketrighte1/bardbl.
Next we take X2and subtract off both its projection into the subspace spanned by e1and
e2
˜e3=X3−[/angbracketleftX3, e1/angbracketrighte1+/angbracketleftX3, e2/angbracketrighte2].
This vector ˜ e3is orthogonal to both e1ande2. Normalize it to get e3= ˜e3//bardbl˜e3/bardbl.
3.3. ABSTRACT SCALAR PRODUCT SPACES 119
More generally, say we have used the vectors X1, X 2,...,X kto obtain the orthonormal
sete1,e2,...,e k. Thenek+1is given by
ek+1=Xk+1−k/summationdisplay
l=1/angbracketleftXk+1, el/angbracketrightel
/bardblXk+1−k/summationdisplay
l=1/angbracketleftXk+1,el/bardbl
This procedure is called the Gram-Schmidt orthogonalization process . With it we can assert
that if some set of linearly independent vectors spans a linear space A, we might as well
suppose that those vectors constitute an orthonormal set, for if they don’t just use Gram-
Schmidt to construct a set that is orthonormal.
The next result is a useful observation.
Theorem 3.15 . A set {X1,X2,...,X n}of orthogonal vectors, none of which is the zero
vector, is necessarily linearly independent.
Proof: The hypothesis states that /angbracketleftXj, Xk/angbracketright= 0, j/negationslash=kand that /angbracketleftXj, Xj/angbracketright /negationslash= 0 . Assume
there are scalars a1,a2,...a nsuch that
0 =a1X1+a2X2+...+anXn.
We shall show that a1=a2=...=an= 0 . Take the scalar product of both sides with
the vector X1. Then
/angbracketleft0, X 1/angbracketright=a1/angbracketleftX1, X 1/angbracketright+a2/angbracketleftX2, X 1/angbracketright+···+an/angbracketleftXn, X 1/angbracketright.
so that
0 =a1/angbracketleftX1, X 1/angbracketright.
Since /angbracketleftX1, X 1/angbracketright /negationslash= 0 , we conclude that a1= 0 . Similarly, by taking the scalar product with
X2we find that a2= 0 , and so on.
An easy consequence of this theorem is the fact that the functions fn(x) = sinnx, n =
1,2,...,N wherex∈[−π,π] are linearly independent, for they are orthogonal (cf. Exercise
5, p. ???).
Say we are given an orthonormal set of nvectors, {ej}, j = 1,...,n, /angbracketleftej, ek/angbracketright=
δjk, andXan element of the linear space Aspanned by the {ej}. Then
X=n/summationdisplay
j=1xjej,
where thexjare uniquely determined just from the general theory of linear spaces (p. 160,
Theorem 10). In the special case of a scalar product space we can conclude even more.
Theorem 3.16 . Let {ej, j= 1,...,n }be an orthonormal set of vectors which span A.
Then every vector X∈Acan be uniquely written as X=n/summationdisplay
j=1xjej, wherexjis the length
of the projection of Xinto the subspace spanned by ej, that is,xj=/angbracketleftX, e j/angbracketright. Thexjare
the Fourier coefficients of X with respect to the orthonormal basis {ej}.
120 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
Proof: This is identical to Theorem 6 of the last section. Take the inner product of both
sides ofX=n/summationdisplay
n=1xjejwithek. Then
/angbracketleftX, e k/angbracketright=/angbracketleftn/summationdisplay
j=1xjej, ek/angbracketright
=n/summationdisplay
j=1xj/angbracketleftej, ek/angbracketright=n/summationdisplay
j=1xjδjk,(3-9)
so that
/angbracketleftX, e k/angbracketright=xk.
Furthermore,
Theorem 3.17 . Let {ej}, j= 1,...,n be an orthonormal set of vectors which span A.
IfX=n/summationdisplay
j=1xjejandY=n/summationdisplay
j=1yjejare vectors in A, then
/angbracketleftX, Y/angbracketright=n/summationdisplay
j=1xjyj=x1y1+x2y2+···+xnyn.
Proof: Identical to Theorem 7 of the last section.
/angbracketleftX, Y/angbracketright=/angbracketleftn/summationdisplay
j=1xjej,n/summationdisplay
k=1ykek/angbracketright
=n/summationdisplay
j=1xj/angbracketleftej,n/summationdisplay
k=1ykek/angbracketright
=n/summationdisplay
j=1xj(n/summationdisplay
k=1yk/angbracketleftej, ek/angbracketright)
=n/summationdisplay
j=1xj(n/summationdisplay
k=1ykδjk),(3-10)
so that
/angbracketleftX, Y/angbracketright=n/summationdisplay
j=1xjyj.
Remark : We shall see that these two theorems extend to the case n=∞.
Examples .
(1) The vectors e1= (1,0,0), e2= (0,1,0) , ande3(0,0,1) clearly form an orthonormal
basis for E3. LetX= (2,−1,4) . We shall compute the xjin
X=3/summationdisplay
j=1xjej.
3.3. ABSTRACT SCALAR PRODUCT SPACES 121
Sincexj=/angbracketleftX, e j/angbracketright, we findz1=/angbracketleftX, e 1/angbracketright=/angbracketleft(2,−1,4),(1,0,0)/angbracketright= 2·1+(−1)·0+4·0 =
2 , and similarly, x2=−1, x3= 4 as expected. Thus
(2,−1,4) = 2e1−e2+ 4e3.
In the same way, if Y= (7,1,−3) , then
Y= 7e1+e2−3e3.
Also,
/angbracketleftX, Y/angbracketright= (2)(7) + ( −1)(1) + (4)( −3) = 1.
The projection of Xinto the subspace spanned by Yis
/angbracketleftX, Y/ /bardblY/bardbl/angbracketrightY
/bardblY/bardbl=1
59(7,1,−3)
=7
59e1+1
59e2−3
59e3.(3-11)
Another orthonormal basis for E3is ˜e1= (1√
2,1√
2,0),˜e2= (−1√
2,1√
2,0) , and ˜e3=
(0,0,1) , since /angbracketleft˜ej,˜ek/angbracketright=δjk. The expansion for Xin this basis is
X=3/summationdisplay
j=1˜xj˜ej,
where
˜x1=/angbracketleftX,˜e1/angbracketright=/angbracketleft(2,−1,4),(1√
2,1√
2,0)/angbracketright=1√
2, (3-12)
˜x2=−3√
2,and ˜x3= 4. (3-13)
Thus
X=1√
2˜e1−3√
2˜e2+ 4˜e3.
Similarly,
Y=8√
2˜e1−6√
2˜e2−3˜e3.
Therefore
/angbracketleftX, Y/angbracketright= (1√
2)(8√
2) + (−3√
2)(−6√
2) + (4)( −3) = 1.
Notice that the number /angbracketleftX, Y/angbracketrightis the same no matter which basis is used. This is
nota coincidence. Recall that the scalar product /angbracketleftX, Y/angbracketrightwas defined independently
of any basis. Hence its value should not be dependent upon which basis we happen
to choose. If you think of /angbracketleftX, Y/angbracketrightgeometrically in terms of the projection, it should
be clear that the number should not depend upon which particular basis is used to
describe the vectors.
122 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
(2) For our second example, we consider the set of orthonormal functions e1(x) =sinx√π,
e2(x) =sinx√πand letAbe the set in L2(−π,π) which they span. We would like to
expand some function
f(x) +2/summationdisplay
j=1fjej(x).
The only trouble is that Theorems 14 and 15 only allow us to expand functions f
which are in the subspace A, that is, are a linear combination of the basis elements e1
ande2. Since we secretly know that f(x) = sinxcosx(=1
2sin 2x) is such a function,
let us find its expansion. By elementary integration,
f1=/angbracketleftf, e 1/angbracketright=/integraldisplayπ
−π(sinxcosx)sinx√xdx= 0,
and
f2=/angbracketleftf, e 2/angbracketright=/integraldisplayπ
−π(sinxcosx)sin 2x√πdx=√π
2.
Therefore
f= 0·e1+√π
2e2=√π
2e2
or
sinxcosx=√π
2(sin 2x√π) =sin 2x
2,
which we knew was the case from trigonometry.
If the orthonormal set {ej}, j= 1,...,m spans a subspaceAof a linear scalar product
spaceH, and ifX∈H, can any sense be made of the expansion
X?=m/summationdisplay
j=1xjej?
One way to seek an answer is to examine a special case. Again geometry will supply the key.
LetH=E3and letAbe the subspace spanned by the orthonormal vectors e1= (1,0,0)
ande2= (0,1,0) . Then if X∈E3, how can we interpret
X?=2/summationdisplay
j=1xjej=x1e1+x2e2?
Plowing blindly ahead, we take the scalar product of both sides with e1and then with e2.
This gives us xj=/angbracketleftX, e j/angbracketright. Thus the right side, x1e1+x2e2, is the projection ofXinto
the subspace Aspanned by {ej}. It is now clear how our original quandary is resolved.
Definition If the orthonormal set {ej}, j= 1,...,m spans a subspace Aof a linear
scalar product space H, and ifX∈H, then the vectorm/summationdisplay
j=1xjej, wherexj=/angbracketleftX, e j/angbracketright,is
the projection of X into the subspace A.
Remark . It is customary to denote the projection of XintoAbyPAX. Think of PA
as an operator (function) which maps the vector Xinto its projection in A. With this
notation the above definition reads
PAX=m/summationdisplay
j=1xjej,
3.3. ABSTRACT SCALAR PRODUCT SPACES 123
wherexj=/angbracketleftX, e j/angbracketrightand the orthonormal set {ej}spansA.
Since the projection PAXis defined in terms of a particular basis for A, we should
show that this geometrical object is independent of the basis you choose for A. But we
shall not take the time right now. In reality, Theorem 17 below leads us to make a better
definition of projection.
Theorem 3.18 . If the orthonormal set {ej}, j= 1,...,m spans a subspace A⊂H,
and ifXandYare inH, then
a)/angbracketleftPAX, P AY/angbracketright=m/summationdisplay
j=1xjyj,
wherexj=/angbracketleftX, e j/angbracketrightandyj=/angbracketleftY, e j/angbracketright. In particular
b)/bardblPAX/bardbl=/radicaltp/radicalvertex/radicalvertex/radicalbtm/summationdisplay
j=1x2
j.
Furthermore, X−PAX∈A⊥, that is, for every Y∈A
c)/angbracketleftX−PAX, Y/angbracketright= 0
EveryX∈Hcan be written as
d)X=PAX+PA⊥X,wherePA⊥X≡X−PAXis inA⊥.
Proof: Since both vectors PAX=m/summationdisplay
j=1xjejandPAY=m/summationdisplay
j=1yjejare inAitself, a) and
b) are immediate consequences of Theorem 15. Although the equation c) is geometrically
clear, we shall compute it too.
Since theejspanA, this is equivalent to showing it is orthogonal to all the ej. Now
/angbracketleftX−PAX, e j/angbracketright=/angbracketleftX, e j/angbracketright−/angbracketleftPAX, e j/angbracketright=xj−xj= 0 . Since trivially X=PAX+(X−PAX) ,
the only content of part d) is that ( X−PAX)∈A⊥, which is just what part c) proved.
Corollary 3.19
a)/bardblX/bardbl2=/bardblPAX/bardbl2+/bardblX−PAX/bardbl2
/bardblX/bardbl2=/bardblPAX/bardbl2+/bardblPA⊥X/bardbl2/bracerightbigg
(Pythagorean Theorem)
b)/bardblX/bardbl2≥ /bardblPAX/bardbl2=m/summationdisplay
j=1x2
j (Bessel’s Inequality)
Proof: a) is a result of the fact that PAX∈Ais orthogonal to X−PAX∈A⊥and
Theorem 12. The inequality b), Bessel’s inequality, is simply a weaker form of a)—since
/bardblX−PAX/bardbl ≥0 . There is equality if and only if X∈A, for only then does /bardblX−PAX/bardbl= 0 .
Examples :
124 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
(1) LetAbe the subspace of E3spanned by e1= (1,0,0) ande2= (0,1,0) . The
projection of X= (3,−1,7) intoAis represented by
PAX=/angbracketleftX, e 1/angbracketrighte1+/angbracketleftX, e 2/angbracketrighte2= 3e1−e2∈A
Also
PA⊥X=X−PAX= 3e1−e2+ 7e3−(3e1−e2) = 7e3∈A⊥.
Since ˜e1= (1√
2,1√
2,0) and ˜e2= (−1√
2,1√
2,0) also form an orthonormal basis for A,
we can equally well write
PAX=/angbracketleftX,˜e1/angbracketright˜e1+/angbracketleftX,˜e2/angbracketright˜e2=2√
2˜e1−4√
2˜e2.
(2) LetAbe the subspace of L2[−π,π] spanned by the orthonormal functions e1(x) =
sinx√π, e2(x) =sin 2x√π. The projection of the function f(x)≡xintoAis represented
by
PAf=/angbracketleftf, e 1/angbracketrighte1+/angbracketleftf, e 2/angbracketrighte2.
Since an integration by parts shows that
/integraldisplayπ
−πxsinkxdx =−xcoskx
k/vextendsingle/vextendsingleπ
−πcoskxdx
=−xcoskx
k/vextendsingle/vextendsingleπ
−π=−2π
kcoskπ=/braceleftbigg2π
k, kodd
−2π
k, keven/bracerightbigg
= (−1)k+12π
k,
we find
/angbracketleftf, e 1/angbracketright=/angbracketleftx, e 1/angbracketright=/integraldisplayπ
−πxsinx√πdx= 2√π
and
/angbracketleftf, e 2/angbracketright=/angbracketleftx, e 2/angbracketright=/integraldisplayπ
−πxsin 2x√πdx=−√π.
Thus
PAx= 2√πsinx√π−√πsin 2x√π,
or
PAx= 2 sinx= sin 2x.
Also,
/bardblPAX/bardbl2=/angbracketleftf, e 1/angbracketright2+/angbracketleftf, e 2/angbracketright2= 5π.
More generally, we can let ˜Abe the subspace of L2[−π,π] spanned by {ek}, k=
1,2,...,N , whereek(x) =sinkx√π. Then the projection of xonto ˜Ais given by
P˜Ax=N/summationdisplay
k=1/angbracketleftx, e k/angbracketrightek(x).
Since
/angbracketleftx, e k/angbracketright=/integraldisplayπ
−πxsinkx√πdx= (−1)k+12√π
k,
3.3. ABSTRACT SCALAR PRODUCT SPACES 125
we have
P˜Ax=N/summationdisplay
k=1(−1)k+12√π
ksinkx√π
= 2N/summationdisplay
k=1(−1)k+1
ksinkx
= 2(sinx−sin 2x
2+sin 3x
3−...+ (−1)N+1sinNx
N).(3-14)
Furthermore,
/bardblP˜Ax/bardbl2=N/summationdisplay
k=1/angbracketleftx, e k/angbracketright2=N/summationdisplay
k=14π
k2= 4πN/summationdisplay
k=11
k2.
It is from this formula that we eventually intend to obtain the famous formula
∞/summationdisplay
k=11
k2= 1 +1
22+1
32+1
42+···=π2
6.
We will observe that
/bardblf/bardbl2=/bardblX/bardbl2=/integraldisplayπ
−πx·xdx=2π3
3
and prove
lim
N→∞/bardblP˜A⊥x/bardbl= lim
N→∞/bardblx−P˜Ax/bardbl= 0
Then from the Corollary to Theorem 16,
/bardblX/bardbl2= lim
N→∞/bardblP˜Ax/bardbl2,
or
2π3
3= 4π?/summationdisplay
k=11
k2⇒π2
6=∞/summationdisplay
k=11
k2.
Geometry leads us to the next theorem—and the proof too. Let Xbe a given vector
andPAXits projection into the subspace A. Since distance is measured by dropping a
perpendicular, we expect that PAXis the vector in Awhich is closest to X, that is, most
closely approximates X.
Theorem 3.20 . LetXbe a vector in a scalar product space HandAa subspace of
H. Then ifVis any vector in A,
/bardblX−PAX/bardbl ≤ /bardblX−V/bardbl.
Proof: We shall prove the stronger statement (cf. fig. above)
/bardblX−PAX/bardbl2+/bardblV−PAX/bardbl2=/bardblX−V/bardbl2.
Observe that ( V−PAX)∈A, since both terms are in AandAis a subspace. Moreover
X−PAX∈A⊥(Theorem 16c). Therefore X−PAXis orthogonal to PAX−V, so the
identity is a consequence of Theorem 12.
126 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
Remark . With this theorem in mind, we could define the projection PAXinto a subspace
Aas the element in Awhich is closest to X. This definition is independent of any basis,
whereas our original definition was not. One must, however, be somewhat careful when
defining the projection into an infinite dimensional subspace. Although it is clear that the
number /bardblX−V/bardblhas a g.l.b. as Vwanders throughout A, it is notclear that it has an
actual min, that is, there really is a vector U∈Asuch that /bardblX−U/bardbltakes on its g.l.b.
as a min. If there is such a U, we call it PAX. Otherwise there is no projection. When
projecting into a finite dimensional space this difficulty does not arise (but we will stop
without further explanation of this detail).
Some discussion of these results is needed to place the material in its proper perspective.
If you are given an orthonormal set of vectors {ej}which span some subspace Aof a scalar
product space H, then for any XinHyou can find a representation for PAXin terms of
that basis, PAX=/summationtextxjej. If the vector Xhappened to already lie in A, thenPAX=X
soX=/summationtextxjejand/bardblX/bardbl=/radicalBig/summationtextx2
j. This last equation for the length of Xis the
Pythagorean Theorem. If Xdid not lie entirely in A, but “stuck out” of it into the rest
ofH, thenPAX=/summationtextxjejonly represents a piece of X, its projection into A. Since
part ofXhas been omitted, we expect that /bardblX/bardbl>/bardblPAX/bardbl=/radicalBig/summationtextx2
j. This inequality
was the content of the Corollary to Theorem 16. Informally, if no vector X∈Hsticks
out of the linear space spanned by the {ej}, then the set {ej}is said to be complete (do
not confuse this with the complete of Chapter 0; they are entirely different concepts, an
unfortunate coincidence). More precisely,
Definition An orthonormal set is complete for the scalar product space Hif that or-
thonormal set is not properly contained in a larger orthonormal set.
There are many ways to check if a given orthonormal set is complete for H. Geometry
suggests them all.
Theorem 3.21 . Let {ej}be an orthonormal set which spans the subspace Aof the
scalar product space H. The following statements are equivalent
(a)The set {ej}is complete for H.
(b)If/angbracketleftX, e j/angbracketright= 0 for allj, thenX= 0.
(c)A=H.
(d)IfX∈H, thenX=/summationtextxjej, wherexj=/angbracketleftX, e j/angbracketright.
(e)IfXandY∈H, then /angbracketleftX, Y/angbracketright=/summationdisplay
xjyj, wherexj=/angbracketleftX, e j/angbracketrightandyj=/angbracketleftY, e j/angbracketright
(f)IfX∈H, then (Pythagorean Theorem) /bardblX/bardbl2=/summationdisplay
x2
j,wherexj=/angbracketleftX, e j/angbracketright
Proof: We shall use the chain of reasoning a⇒b⇒c...⇒f⇒a.
a⇒b. If/angbracketleftX, e j/angbracketright= 0 butX/negationslash= 0 , thenX//bardblX/bardblis a unit vector orthogonal to all the
ej. This means that {X
/bardblX/bardbl, e1,e2,...}is an orthonormal set which contains {e1,e2,...}
as a proper subset.
b⇒c. If there is an X∈HbutX/ownerA, thenPA⊥X=X−PAX∈A⊥and is not
zero. Since all the ej∈A, we have /angbracketleftPA⊥X, e j/angbracketright= 0 for all jbutPA⊥X/negationslash= 0 , contradicting
b). ThusH⊂A. SinceA⊂Hby hypothesis, this proves that H=A.
3.3. ABSTRACT SCALAR PRODUCT SPACES 127
c⇒d. Since every X∈Ahas the form X=/summationtextxjej(by Theorem 14) and since
H=A, the conclusion is immediate.
d⇒e⇒f. A restatement of Theorem 16 since for every X∈H, we know that
PAX=PHX=X.
f⇒a. If{ej}is not complete, it is contained in a larger orthonormal set. Let e
be a vector in that larger set which is not one of the ej. Then by f), and the fact that
/angbracketlefte, ej/angbracketright= 0 ,
/bardble/bardbl2=/summationdisplay
/angbracketlefte, ej/angbracketright2= 0.
Thereforee= 0 .
Remarks . 1. Because each of the six conditions a-f are equivalent, any one of them could
have been used as the definition of a complete orthonormal set.
2. If the orthonormal set {ej}has a (countably) infinite number of elements, the
theorem is still valid but some convergence questions for d-f arise because of the then
infinite seriesX=∞/summationdisplay
1xjej. The appropriate sense of convergence is that the remainder
afterNterms,∞/summationdisplay
N+1xjej=X−N/summationdisplay
1xjejtends to zero in the norm of the scalar product
space , that is, if
lim
N→∞/bardblX−N/summationdisplay
1xjej/bardbl= 0.
We shall meet this in the next section for the space L2[−π,π] . Condition f) gives us
no convergence problems since the series is an infinite series of positive terms which is
always bounded by /bardblX/bardbl2(Bessel’s Inequality—Corollary b to Theorem 16), and so always
converges. This criterion just asks if the sum of the series actually equals /bardblX/bardbl2(we know
it is no larger).
Examples
(1) The set of orthonormal vectors e1= (1√
2,1√
2,0) ande2= (1√
2,−1√
2,0) are not
complete for E3since any basis for E3must have three elements because its dimension
is 3. This could also be seen geometrically from the fact that, for example X= (1,2,3)
sticks out of the space spanned by e1ande2, or from the fact that e3= (0,0,2)
is a non-zero vector orthogonal to both e1ande2, or in many other ways. The
dimension argument is the easiest to apply if H is finite dimension , for then the
number of elements in a complete orthonormal set {ek}must equal the dimension
of H.
(2) The set {˜en}where ˜en(x) =sinnx√πis an orthonormal set of functions in the scalar
product space L2[−π,π] , but it is nota complete orthonormal set for that space since
the function cos xis a non-zero function in L2[−π,π] which is orthogonal to all the
˜en,
/angbracketleftcosx,˜en/angbracketright=/integraldisplayπ
−πcossinnx√πdx= 0.
Thus, although the set {˜en}has an infinite number of elements, it is still not big
enough to span all of L2[−π,π] . The next section will be devoted to proving that the
128 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
larger orthonormal set, e0,e1,˜e1,e2,˜e2,..., where
e0=1√
2π, en(x) =cosnx√π,˜en(x) =sinnx√π
is a complete orthonormal set for the scalar product space L2[−π,π] . This is a
difficult theorem.
Specific applications of the ideas in this section are contained in the exercises. For
many of them you would be wise if you referred to their corresponding special cases which
appeared in Section 2.
Exercises
(1) LetXandYbe points in Rn. Determine which of the following make Rninto a
scalar product space, and why—or why not.
(a)/angbracketleftX, Y/angbracketright=n/summationdisplay
k=11
kxkyk.
(b)/angbracketleftX, Y/angbracketright=n/summationdisplay
k=1(−1)kxkyk.
(c)/angbracketleftX, Y/angbracketright=/radicaltp/radicalvertex/radicalvertex/radicalbtn/summationdisplay
k=1x2
ky2
k.
(d)/angbracketleftX, Y/angbracketright=n/summationdisplay
k=1akxkyk, whereak>0 for allk.
(2) Letfandgbe continuous real-valued functions in the interval [0 ,1] , sof, g∈
C[0,1] . Determine which of the following make C[0,1] into a scalar product space,
and why—or why not.
(a)/angbracketleftf, g/angbracketright=/integraldisplay1
0f(x)g(x)1
1 +x2dx.
(b)/angbracketleftf, g/angbracketright=/integraldisplay1
0f(x)g(x) sin 2πxdx .
(c)/angbracketleftf, g/angbracketright=/integraldisplay1
0f(x)g2(x)dx
(d)/angbracketleftf, g/angbracketright=/integraldisplay1
0f(x)g(x)ρ(x)dx, whereρ(x) is a fixed continuous function with the
propertyρ(x)>0 .
(e)/angbracketleftf, g/angbracketright=f(0)g(0) .
(3) This is the analogue of L2for sequences. Let l2be the set of all sequences X=
(x1,x2,xe,...) with the property that /bardblX/bardbl=/radicaltp/radicalvertex/radicalvertex/radicalbt∞/summationdisplay
j=1x2
j<∞. Prove that l2is a
normed linear space (cf. the example for l1in Section 1).
3.3. ABSTRACT SCALAR PRODUCT SPACES 129
(4) Use the Cauchy-Schwarz inequality to prove that if∞/summationdisplay
n=1n2a2
n<∞, then∞/summationdisplay
n=1|an|<∞.
(Hint: |an|=1
n|nan|).
(5) Consider the following linearly independent vectors in E3:
X1= (1,0,−1), X 2= (0,3,1), X 3= (2,−1,0).
(a) Use the Gram-Schmidt orthogonalization process to find an orthonormal set of
vectors,e1,e2ande3such thate1is in the subspace spanned by X1.
(b) WriteX= (1,2,3) asX=3/summationdisplay
j=1xjej, where the ejare those of part a). Also,
compute /bardblX/bardbland/bardblPX/bardbl.
(6) Consider the following linearly independent set of functions in L2[−1,1]
f1(x) = 1, f 2(x) =x, f 3(x) =x2.
(a) Use the Gram-Schmidt orthogonalization process to find an orthonormal set of
functionse1(x),e2(x) ande3(x) such that e1is in the subspace spanned by
f1.
(b) Find the projection of the function f(x) = (1+x)3into the subspace of L2[−1,1]
spanned by e1(x),e2(x) , ande3(x) . Also, compute /bardblf/bardbland/bardblPf/bardbl.
(7) LetPn(x) =1
2nn!dn
dxn(1−x2)n,n= 0,1,2,.... These are the Legendre Polynomials .
(a) Prove that /angbracketleftPn, Pm/angbracketright=/integraldisplay1
−1Pn(x)Pm(x)dx= 0, n/negationslash=m, that is, the Pnare
orthogonal in L2[−1,1] by first proving that
/integraldisplay1
−1Pn(x)xmdx= 0, m<n.
(b) Show that /bardblPn/bardbl2=2
2n+1. Thus the functions
en(x) =/radicalbigg
2n+ 1
2Pn(x)
are an orthonormal set of functions for L2[−1,1] . Compute e0(x),e1(x) , and
e2(x) and compare with Exercise 6a.
(8) (a) Show that the vector N= (a1,a2,a3) is orthogonal to the coset (a plane in E3)
A={X∈E3:a1x1+a2x2+a3x3=c}.
(b) Show that the vector N= (a1,...,a n) is orthogonal to the coset (a hyperplane
inEn)A={X∈En:a1x1+...a nxn=c}.
(c) Find the coset A⊂E3which passes through the point X0= (1,−1,2) and is
orthogonal to N= (1,3,2) . In ordinary language, Ais the plane containing
the pointX0which is orthogonal to N.
130 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
(d) Show that the coset A⊂Enwhich passes through the point X0= (˜x1,..., ˜xn)
and is orthogonal to N= (a1,...,a n) is
A={X∈En:/angbracketleftX, N/angbracketright=/angbracketleftX0, N/angbracketright}.
(9) (a) Use Problem 8a to show that the distance dfrom the point P= (y1,y2,y3)∈E3
to the coset a1x1+a2x2+a3x3=cinE3is
d=|a1y1+a2y2+a3y3−c|/radicalbig
a2a+a2
2+a2
3=|/angbracketleftN, P/angbracketright −c|
/bardblN/bardbl
(b) Show that the distance dfrom the point P= (y1,...,y n)∈Ento the coset
a1x1+...+anxn=cinEnis
d=|a1y1+a2y2+···+anyn−c|/radicalbig
a2
1+a2
2+···+a2n=|/angbracketleftN, P/angbracketright −c|
/bardblN/bardbl.
(c) Show that the distance dbetween the “parallel” cosets a1x1+···+anxn=c1
anda1x1+···+anxn=c2inEnis
d=|c1−c2|/radicalbig
a2
1+···+a2n=|c1−c2|
/bardblN/bardbl.
(Hint: Pick a point Pin one of the cosets and apply part b).
(10) Find the angle between the diagonal of a cube and one of its edges.
(11) LetY1andY2be fixed vectors in a scalar product space H.
a). If /angbracketleftX, Y 1/angbracketright= 0 for all X∈H, prove that Y1= 0 .
b). If /angbracketleftX, Y 1/angbracketright=/angbracketleftX, Y 2/angbracketrightfor allX∈H, prove that Y1=Y2.
(12) LetY0be a fixed vector in a scalar product space H. LetA={X∈H:/angbracketleftY, Y 0/angbracketright=
0Rightarrow /angbracketleftX, Y/angbracketright= 0}. Prove that Ais the span of Y0:{X∈ARightarrowX =
cY0}for some scalar c. Make sure to see the geometrical situation for the case
H=E3. [Hint: LetBbe the set of all vectors orthogonal to Y0, soY∈B.
SinceHis composed of two parts, Y0andB, everyX∈Hcan be written as
X=cY0+Z, wherecY0is the projection of Xinto the subspace spanned by Y0
(soc=/angbracketleftX, Y 0/angbracketright//bardblY0/bardbl2) andZ= (X−cY0)∈B. Now show that X∈A⇒Z= 0) .
(13) (a) Let X= (1,3,−1) andY= (2,1,1) . Find a vector Nwhich is orthogonal to
the subspace spanned by XandY.
(b) LetX= (x1,x2,x3) andY= (y1,y2,y3) . Find a vector Nwhich is orthogonal
to the subspace spanned by XandY. [Answer. N=c(x2y3−y2x3,y1x3−
x1y3,x1y2−y1x2) , wherecis any non-zero scalar].
(14) LetAbe the subspace of L2[−π,π] spanned by the orthonormal set {en(x)}, n=
1,2,...,N , whereen(x) =sinnx√π.
(a) Find the projection of f(x) =x2, intoA. (The answer should surprise you).
Compute /bardblf/bardbland/bardblPAf/bardbltoo.
3.3. ABSTRACT SCALAR PRODUCT SPACES 131
(b) Find the projection of f(x) = 1 + sin3xintoA. Compute /bardblf/bardbland/bardblPAf/bardbl.
(c) Iff(x) is an even function, f(x) =f(−x) , show that its projection into Ais
zero. Now look at part (a) again.
(15) (a) If f∈C[a,b] , show that
(/integraldisplayb
af(x)dx)2≤(b−a)/integraldisplayb
af2(x)dx.
[Hint: Write f(x) = 1·f(x) and use the Cauchy-Schwarz inequality for L2[a,b] ].
(b) Iff∈C1[a,b] , prove that
|f(x)−f(a)|2≤(x−a)/integraldisplayb
af/prime(x)2dx, x∈(a,b).
[Hint: Write f(x)−f(a) =/integraltextx
af/prime(t)dtand apply part a)]
(c) Iff∈C1[a,b] andf(a) = 0 , use part b to prove that
/integraldisplayb
af2(x)dx≤(b−a)2
2/integraldisplayb
af/prime(x)2dx.
(16) (a) Let A={h∈C1[a,b]:h(a) =h(b)}and letB={h∈C1[a,b]:/angbracketleft1, h/prime/angbracketright= 0},
whereh/prime=dh
dx. Show that the subspaces AandBare identical h∈A⇐⇒
h∈B.
(b) Letf(x) be any continuous function such that/integraldisplayb
af(x)h/prime(x)dx= 0 for all
h(x)∈C1[a,b] withh(a) =h(b) . Show that f≡constant. [Hint: Use part (a)
and the result of Exercise 12].
(17) Iff(x)∈C[a,b] and satisfies the condition/integraldisplayb
af(x)h(x)dx= 0 for all h(x)∈C[a,b]
which satisfy the conditions
/integraldisplayb
ah(x)dx= 0,/integraldisplayb
axh(x)dx= 0,...,/integraldisplayb
axnh(x)dx= 0,
prove that f∈Pn, that is,fis of the form
f(x) =a0+a1x+···+anxn,
where theajare constants. [Hint: Use Exercise 12].
(18) Determine which of the following orthonormal sets are complete for their respective
spaces.
(a) In E3, e1= (0,1,0), e2= (3
5,0,4
5), e3= (−4
5,0,3
5)
(b) In E4, e1= (1,0,0,0), ex= (0,1,0,0), e3= (0,0,1√
2,1√
2).
(c) In E4, e1,e2,e3as in (b), and e4= (0,0,1√
2,−1√
2).
132 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
(19) Lete1,e2, ande3be an orthonormal basis for E3, and letAbe the subspace
spanned by X1= 3e1−4e3. Find an orthonormal basis for A⊥.
(20) LetAbe a subspace of a scalar product space H. IfX∈H, prove that PA(PAX) =
PAXand interpret this geometrically. This result can be written as P2
A=PA.
(21) LetAbe any operator (not necessarily linear) on a scalar product space. Prove the
polarization identity
2/angbracketleftAX, AY /angbracketright=/bardblAX+AY/bardbl2− /bardblAX/bardbl2− /bardblAY/bardbl2.
3.4 Fourier Series.
Throughout this section we shall only use the scalar product of L2[a,b] ,
/angbracketleftf, g/angbracketright=/integraldisplayb
af(x)g(x)dx.
We begin with the observation that in the interval [ −π,π]
/angbracketleftsinnx,sinmx/angbracketright=/integraldisplayπ
−πsinnxsinmxdx =πδnm, (3-15)
/angbracketleftsinnx,cosmx/angbracketright=/integraldisplayπ
−πsinnxcosmxdx = 0, (3-16)
and
/angbracketleftcosnx,cosmx/angbracketright=/integraldisplayπ
−πcosnxcosmxdx =πδnm, (3-17)
wheren, m = 0,1,2,3,.... Thus the functions
e0(x) =1√
2π, en(x) =cosnx√π,˜cn(x) =sinnx√π
form an orthonormal set:
/angbracketleften,˜em/angbracketright=δnm/angbracketleften,˜em/angbracketright= 0,/angbracketleft˜en,˜em/angbracketright=δnm.
Thus, iff∈Lx[−π,π] , we can find the projection PNfoffinto the subspace spanned
bye0,e1,˜e,...,e N,˜eN.
(PNf) =a0e0+N/summationdisplay
n=1anen+bn˜en, (3-18)
where
ak=/angbracketleftf, ek/angbracketrightandbk=/angbracketleftf,˜ek/angbracketright (3-19)
More explicitly,
(PNf)(x) =a01√
2π+N/summationdisplay
n=1ancosnx√π+bnsinnx√π, (3-20)
3.4. FOURIER SERIES. 133
where
a0=/integraldisplayπ
−πf(x)·1√
2πdx,
and
an=/integraldisplayπ
−πf(x)cosnx√πdx, b n=/integraldisplayπ
−πf(x)sinnx√πdx (3-21)
A natural question arises: as N→ ∞ , does the series converge: PNf→f, in the sense
that/bardblf−PNf/bardbl →0 ? In other words, is the set {ej(x),˜ej(x)}, j= 0,1,2,... acomplete
orthonormal set of functions for L2[−π,π] ? The answer is yes, as we shall prove. Thus for
anyf∈L2[−π,π] ,
f(x) =a01√
2π+∞/summationdisplay
n=1ancosnx√π+bnsinnx√π, (3-22)
where the Fourier coefficients ,an,bnare determined by the formulas (2). The expansion
(3) is called the Fourier series for f.
Historically, Fourier series did not arise from the geometrical considerations we have
developed. Mathematical physics—in particular the vibrations of strings and the flow of
heat in a bar—take the credit for these ideas. Only in recent years has the geometrical
viewpoint been investigated. Later on we shall discuss some of the fascinating problems in
mathematical physics to which Fourier series can be applied.
Beware . The equality which appears in (3) is equality in the L2[−π,π] norm, viz.
/bardblf−PNf/bardbl=/radicalBigg/integraldisplayb
a[f(x)−(PNf)(x)]2dx→0
This is quite different than the convergence of infinite series to which you’re accustomed,
which is the uniform norm
/bardblf−PNf/bardbl∞= max
−π≤x≤π|f(x)−(PNf)(x)|.
In Section 1 (p. 176) you saw one instance of where a sequence of functions converged in
some norm (the L1norm there) but did not converge in the uniform norm. Such is also
the case here. In fact, contrasting the situation in the L2norm, there do exist continuous
functionsfwhose Fourier series (3) does not converge to fin the uniform norm. However
if the function fhas one derivative, then its Fourier series does converge to fin the
uniform norm.
In addition, there are some discontinuous functions whose Fourier series converge. These
ideas will become clearer later on.
You should be warned that our definition (1), (3) of a Fourier series is not the standard
one. Most books do notwork with the orthonormal sete0=1√
2π, en=cosnx√π,˜en=sinnx√π,
but rather use just an orthogonal set which is notnormalized θ0=1
2, θn= cosnx,˜θn=
sinnx. For these people,
f(x) =A0
2+∞/summationdisplay
n=1Ancosnx+Bnsinnx,
where
An=/integraldisplayπ
−πf(x)cosnx
πdx, B n=/integraldisplayπ
−πf(x)sinnx
πdx,
134 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
n= 0,1,2.... As you can see, these differ from our formulas only by factors of√π.
Needless to say, the resulting Fourier series for a given function fdoes not depend which
intermediate formulas you use. We prefer the less standard ones because they are more
intimately tied to geometry (so there is less to remember).
Before discussing the difficult issues of convergence in detail, we will find the Fourier
series associated with some specific functions.
Examples .
(1) Find the Fourier series associated with the functions f(x) =x,−π≤x≤π. We
actually found this in the previous section. A computation (involving integration by
parts) shows that
a0=/angbracketleftf, e 0/angbracketright=/integraldisplayπ
−πx·1√
2πdx= 0
an=/angbracketleftf, en/angbracketright=/integraldisplayπ
−πxcosnx√πdx= 0, n = 1,2,...
bn=/angbracketleftf,˜en/angbracketright=/integraldisplayπ
−πxsinnx√πdx=2(−1)n+1√π
n, n = 1,2,...
Thus, upon substituting into (3) we find that
x=∞/summationdisplay
n=12(−1)n+1
n√πsinnx√π= 2∞/summationdisplay
n=1(−1)n+1
nsinnx
or
x= 2[sinx−sin 2x
2+sin 3x
3−sin 4x
4+···].
Again we remind you that the equality here is in the sense of convergence in L2. For
this particular function, there is also equality in the usual sense of convergence for
infinite series for all x∈(−π,π) . Direct substitution reveals that it does notconverge
in the usual sense at x=±π. These remarks are based upon convergence theorems
we have yet to prove. At x=π
2, this yields
π
4= 1−1
3+1
5−1
7+···
(2) Since the formulas (2)’ make sense even if the function f(x) has a finite number of
discontinuities, we are tempted to find the Fourier series for discontinuous functions
(in contrast, recall that the coefficients of an infinite power series are only defined if
the function had an infinite number of derivatives). We shall find the Fourier series
associated with the discontinuous function
f(x) =/braceleftbigg0,−π≤x≤0
π,0<x<π
3.4. FOURIER SERIES. 135
The computations are particularly simple.
a0=/angbracketleftf, e 0/angbracketright=/integraldisplay0
−π0·1√
2πdx+/integraldisplayπ
0π·1√
2πdx =π2
√
2π,(3-23)
an=/angbracketleftf, en/angbracketright=/integraldisplay0
−π0·cosnx√πdx+/integraldisplayπ
0π·cosnx√πdx= 0, n> 0, (3-24)
bn=/angbracketleftf,˜en/angbracketright=/integraldisplay0
−π0·sinnx√πdx+/integraldisplayπ
0π·sinnx√πdx (3-25)
=√π
n(1−cosnπ) =/braceleftBigg
2√π
n, n odd
0, n even(3-26)
Therefore the Fourier series associated with this function is
f(x) =π2
√
2π·1√
2π+ 2√π/parenleftbiggsinx√π+sin 3x
3√π+sin 5x
5√π+···/parenrightbigg
,
or
f(x) =π
2+ 2(sinx+sin 3x
3+sin 5x
5+···).
As usual, the equality is meant in the sense of convergence in the L2norm. The series
also converges to the function fin the uniform norm in the whole interval except for a
neighborhood of x= 0 . At 0 it hasn’t got a chance because of the discontinuity of f
there. A glance at the series reveals that at x= 0 , the right side is π/2 —the arithmetic
mean between the values of fjust to the left and right of 0. This is the usual case at a
discontinuity: a Fourier series converges to the average of the function values to the right
and left of the point where fis discontinuous . We still offer no proof for these statements.
Observe that the Fourier series (3) for any function f(x) depends only upon the values
ofxin the interval −π≤x≤π. However the series itself is periodic with period 2 π.
If the function f(x) , which we considered only for x∈[−π,π] is defined for all other x
by the formula f(x+ 2π) =f(x) (making fperiodic too), then both sides of the Fourier
series (3) are periodic with period 2 π. Therefore whatever they do in the interval [ −π,π]
is repeated every 2 π.
For example, the function f(x) =x, x∈[−π,π] when continued outside the interval
[−π,π] as a function periodic with period 2 πbecomes
a figure goes here
Since the Fourier series for this particular function converges uniformly for all x∈(−π,π) ,
it also converges uniformly to the periodically continued function for all x∈(kπ,kπ +
2π), k= 1,±1,±2,... . This also makes it clear why the Fourier series for f(x) =x
converges to zero at x=±π, for the series is just converging to the arithmetic mean of its
neighboring values at the discontinuity.
It is pleasant to look at a picture. Let us see how the first four terms of its Fourier
series approximates the function x
x= 2(sinx−sin 2x
2+sin 3x
3−sin 4x
4+···)
a figure goes here
136 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
Notice that as more terms are used, the projection PNx P Nx= 2(sinx−sin 2x
2+···+
(−1)N+1sinNx
N) more and more closely approximate x. This reflects the convergence of the
Fourier series, PNf→f.
One popular interpretation of a Fourier series is as a sum of “waves” which approximate
a given function. Thus the function xis the sum of 2 times the wave sin xplus ( −1) times
the wave sin 2 xand so on. In other words, the Fourier series for the function f(x) =x
represents that function as the superposition of sine waves. The term 2 sin xis spoken of
as the first harmonic , the term – sin 2 xas the second harmonic, the term2
3sin 3xas the
third harmonic, etc.
Although it is difficult to believe, the ear hears by taking the sound wave f(x) which
impinges on the ear drum and splitting it up into its Fourier components (3). It then
analyzes each component anen+bn˜en—only considering the coefficients anandbn. These
Fourier coefficients measure the intensity of the nth harmonic. Particular sounds are then
heard in terms of the intensity of their various harmonics. We recognize familiar sounds by
recognizing that the sound waves have similar Fourier coefficients. Amazing.
It is time to consider the convergence of Fourier series. The question is: does the partial
Fourier series
PNf=a0e0+N/summationdisplay
n=0anen+bn˜en
converge to the function fasN→ ∞ . Since there are several norms, in particular the L2
norm /bardbl /bardbl and the uniform norm /bardbl /bardbl ∞, we must investigate convergence in each norm.
Even though our proofs are reasonably slick, they are neither short nor particularly simple.
A great deal of analytical technique will be needed. The proofs to be presented have been
chosen because each of the devices invoked are important devices in their own right.
We begin with some useful facts which have nothing especially to do with Fourier series.
Theorem 3.22 (Weierstrass Approximation Theorem). If f(x)is continuous in the inter-
val[−π,π]andf(−π) =f(π), then given any /epsilon1>0there is a trigonometric polynomial
TN(x) =α0+N/summationdisplay
n=1αncosnx+βnsinnx
= ˆα0e0+N/summationdisplay
n=1ˆαnen+ˆβn˜en,(3-27)
(whereα0= ˆα0√
2π, α n= ˆαn√π, β n=ˆβn√π), such that
/bardblf−TN/bardbl∞= max
−π≤x≤π|f(x)−TN(x)|</epsilon1.
Note that the numbers αnandβnarenotnecessarily the Fourier coefficients of f. The
proof, which is placed as an appendix at the end of this section, will indicate how they can
be found.
The following theorem states that convergence in the uniform norm implies convergence
in theL2norm.
Theorem 3.23 . Ifθ(x)is any bounded integrable function, then (if b>a )
/bardblθ/bardbl ≤√
b−a/bardblθ/bardbl∞.
3.4. FOURIER SERIES. 137
Proof: Since /bardblθ/bardbl∞= max
x∈[a,b]|θ(x)|we find immediately that
/integraldisplayb
aθ(x)2dx≤/integraldisplayb
a/bardblθ/bardbl2
∞dx=/bardblθ/bardbl2
∞/integraldisplayb
adx= (b−a)/bardblθ/bardbl2
∞
from which the conclusion is obvious. On geometrical grounds the theorem is even easier,
since /bardblθ/bardbl∞is the greatest height of the curve θ(x) .
Although convergence in the L2norm does notimply convergence in the uniform norm
(the example in Section 1 comparing L1convergence and uniform convergence also works
forL2), a useful weaker statement is true.
Theorem 3.24 . (cf. Ex. 15 Section 3). If θ∈C1[a,b]andθ(x0) = 0 , wherex0∈[a,b],
then for every x∈[a,b]
|θ(x)| ≤√
b−a/radicalBigg/integraldisplayb
aθ/primedt=√
b−a/bardblθ/prime/bardbl.
Since the right side is independent of x, this implies that
/bardblθ/bardbl∞=max x∈[a,b]|θ(x)| ≤√
b−a/bardblθ/prime/bardbl=√
b−a/bardblDθ/bardbl
Proof: By the fundamental theorem of calculus,
θ(x) =θ(x)−θ(x0) =/integraldisplayx
x0θ/prime(t)dt.
Thus the Cauchy-Schwarz inequality yields
|θ(x)|2=/parenleftbigg/integraldisplayx
x01·θ/prime(t)dt/parenrightbigg2
≤/integraldisplayx
x012dt/integraldisplayx
x0θ/prime(t)2dt
= (x−x0)/integraldisplayx
x0θ/prime(t)2dt
≤(b−a)/integraldisplayb
aθ/prime(t)2dt.(3-28)
Therefore
|θ(x)|2= (b−a)/bardblθ/prime/bardbl2.
With these preliminaries behind us we turn to the convergence of Fourier series. First
up is convergence in the L2norm.
Theorem 3.25 . Assume fis continuous in the interval [−π,π]andf(−π) =f(π).
Denote the sum of the first Nterms of its Fourier series by PNf. Then
lim
N→∞/bardblf−PNf/bardbl= lim
N→∞/radicalBigg/integraldisplayπ
−π[f(x)−(PNf)(x)]2dx= 0
138 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
Proof: Given any/epsilon1>0 , letTN(x) be the trigonometric polynomial given by Weierstrass
Approximation Theorem. The trick is to apply Theorem 17. Using the NofTN, we know
that
PNf=a0e0+N/summationdisplay
n=1anen+bn˜en
and
TN= ˆa0e0+N/summationdisplay
n=1ˆαnen+ˆβn˜en.
LetAbe the subspace of H=L2[−π,π] spanned by e0,e1,˜e1,...e N,˜eNThen both PNf
andTNare inA. Thus by Theorem 17 of the last section (where slightly different notation
was used),
/bardblf−PNf/bardbl ≤ /bardblf−TN/bardbl,
and by Theorem 20
≤√
b−a/bardblf−TN/bardbl∞<√
b−a/epsilon1.
Thus
lim
N→∞/bardblf−PNf/bardbl= 0,
proving the theorem.
Corollary 3.26 (Parseval’s Theorem). If f(x)is continuous in the interval [−π,π]and
f(−π) =f(π), then
/bardblf/bardbl2= lim
N→∞/bardblPNf/bardbl2,
that is,/integraldisplayπ
−πf2(x)dx=a2
0+∞/summationdisplay
n=1(a2
n+b2
n),
where the Fourier coefficients ajandbjare determined by equations (2) or (2)’.
Proof: The Corollary to Theorem 16 states that
/bardblf/bardbl2=/bardblPNf/bardbl2+/bardblf−PNf/bardbl2.
If we now let N→ ∞ , the second term on the right vanishes by the theorem just proved.
Remark : The theorem and corollary state that the orthonormal set of functions e0=
1√
2π, en(x) =cosnx√π, and ˜en(x) =sinnx√πis a complete orthonormal set for the scalar prod-
uct spaceL2[−π,π] . The formula contained in the corollary is a generalization of the
Pythagorean Theorem to L2[−π,π] .
The proof of convergence in the uniform norm if the function has one continuous deriva-
tive is only slightly more difficult. We shall need a preliminary
Lemma 3.27 . Assumef∈C1[−π,π]. Extend it as a periodic function with period 2π
byf(x+ 2π) =f(x). Let (PNf)be the sum of the first Nterms of its Fourier series.
Then the sum of the first Nterms in the Fourier series for Df=d f
dxisPN(Df), that is
PN(Df) =D(PNf).
This in not necessarily true for other bases in L2[−π,π].
3.4. FOURIER SERIES. 139
Proof: We know that
(PNf)(x) =a01√
2π+N/summationdisplay
n=1ancosnx√π+bnsinnx√π.
Since we can differentiate a finite sum term by term, we find that
D(PNf)(x) =N/summationdisplay
n=1−nansinnx√π+nbncosnx√π,
where theanandbnare found by using formulas (2)’. If
PN(Df) =A01√
2π+N/summationdisplay
n=1Ancosnx√π+Bnsinnx√π,
where theAnandBnare also found by using (2)’, we must show that
A0= 0, A n=nbn,andBn=−nan
But
A0=/integraldisplayπ
−π(Df(x))1√
2πdx=1√
2π[f(π)−f(−π)] = 0 since fis periodic.
Integrating by parts, we further find
An=/integraldisplayπ
−π(Df(x))cosnx√πdx=n/integraldisplayπ
−πf(x)sinnx√πdx=nbn
and
Bn=/integraldisplayπ
−π(Df(x))sinnx√πdx=−n/integraldisplayπ
−πf(x)cosnx√πdx=−nan.
Our result is now only a few steps away.
Theorem 3.28 . Iff∈C1[−π,π]and if both fandf/primeare periodic with period 2π,
then the Fourier series PNfconverges to fin the uniform norm
lim
N→∞/bardblf−PNf/bardbl∞= 0.
Proof: The key observation is that f/primeis a continuous function, so that Theorem 22 can
be applied to its Fourier series. This shows that
lim
N→∞/bardblDf−PN(Df)/bardbl= 0.
By the above lemma,
D(f−Pnf) =Df−D(PNf) =Df−PN(Df).
Thus
lim
N→∞/bardblD(f−PNf)/bardbl= 0. (3-29)
140 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
We would like to apply Theorem 21 to the function θN=f−PNf. In order to do so, we
must only verify that θNvanishes somewhere in [ −π,π] . But the area under θN=f−PNf
is /integraldisplayπ
−πθN(x)dx=√
2π/angbracketleftθN, e0/angbracketright= 0
sincef−PNfis orthogonal to the space spanned by e0,e1,˜e1,...,e N,˜eN(Theorem 16c).
BecauseθN(x) is a continuous function (the difference of the C1) function fand the
infinitely differentiable trigonometric polynomial PNf), the area under it can be zero only
ifθNvanishes somewhere. Thus Theorem 21 is applicable and yields the inequality
/bardblf−PNf/bardbl∞≤√
b−a/bardblD(f−PNf)/bardbl.
We now pass to the limit N→ ∞ and use equation (4) to complete the proof of the
theorem:
lim
N→∞/bardblf−PNf/bardbl∞≤lim
N→∞√
b−a/bardblD(f−PNf)/bardbl= 0.
Remarks . The hypothesis that f∈C1[−π,π] and is periodic with period 2 πhas been
proved a sufficient condition for the Fourier series to converge to the function in the uniform
norm. Much weaker hypotheses also suffice to prove the same result—but mere continuity
is not enough. Convergence of Fourier series or generalizations thereof is a vast and deep
subject, one still the object of intense study.
On the basis of the theorems we have proved, many other problems are reasonably
accessible—like the convergence of the Fourier series for a function which is nice except for
a finite number of jump discontinuities. But there is not time for this pleasant excursion.
a figure goes here
3.5 Appendix. The Weierstrass Approximation Theorem
.
The proof—which is difficult—will be given as a series of lemmas.
Lemma 3.29 . Iff(x)is continuous and periodic with period 2π, then for any a∈R,
the following equality holds
/integraldisplaya+2π
af(x)dx=/integraldisplay2π
0f(x)dx.
Proof: This is clear from a graph of f, since the area under one period of fdoes not
depend upon where you begin measuring. We also offer a computational proof. Write
/integraldisplaya+2π
af(x)dx=/integraldisplay0
af(x)dx+/integraldisplay2π
0f(x)dx+/integraldisplaya+2π
2πf(x)dx.
Letx=t+2πin the last integral and use the fact that f(t+2π) =f(t) . The last integral
is then
−/integraldisplay0
af(t)dt,
which cancels the unwanted term in the last equation and proves the lemma.
3.5. APPENDIX. THE WEIERSTRASS APPROXIMATION THEOREM 141
Lemma 3.30 ./integraldisplayπ/2
0cos2nt dt=1
2cn, wherecn=v1
π2·4·6···(2n)
1·3·5···(2n−1).
Proof: A computation. Integrate by parts to show that
I2n=/integraldisplayπ/2
0cos2ntdt= (2n−1)(I2n−2−I2n).
ThusI2n=2n−1
2nI2n−2. Now induction can be used to do the rest, since by observation
I0=π/2 .
Lemma 3.31 . Assumef(x)is continuous and periodic with period 2π. Let
TN(x) =cN
2/integraldisplayπ
−πf(t) cos2N(t−x
2)dt (3-30)
Then given any /epsilon1>0, there is an Nsuch that
/bardblf−TN/bardbl∞= max
−π≤x≤π|f(x)−TN(x)|</epsilon1.
Proof: How did we guess the formula (4)? We observed that cos2Nxis one atx= 0 ,
and strictly less than one for all other x∈[−π,π] . Thus, for large N,cos2Nxis one at
x= 0 , and decreases sharply thereafter so cos2N(t−x
2) has the same property at x−t= 0 ,
wherex=t. Then essentially the only values of f(t) which will count are those about
t=x, so what comes out will be f(x) . Let us proceed with the details.
Takes=t−x
2. Then
TN(x) =cN/integraldisplayπ/2
−π/2f(x+ 2s) cos2Nsds.
Split the integral into two pieces, from −π
2to 0 and from 0 toπ
2, and then replace sby
−sin the first one. This gives
TN(x) =cN/integraldisplayπ/2
0[f(x) + 2s) +f(x−2s)] cos2Nsds.
From Lemma 2 we know that
f(x) =cN/integraldisplayπ/2
02f(x) cos2Nsds,
sincef(x) is a constant in the integration with respect to s. Therefore
TN(x)−f(x) =cN/integraldisplayπ/2
0[f(x+ 2s)−2f(x) +f(x−2s)] cos2Nsds.
Now given any /epsilon1 >0 , from the continuity of fwe can pick a δ >0 independent of x
such that
|f(x1)−f(x2)|</epsilon1
2when |x1−x2|<δ.
This/epsilon1will be the /epsilon1of our conclusion. Break the integral into two parts, one from 0 to δ
and the other from δtoπ/2 , whereδis theδwe just found. Then in the [0 ,δ] interval,
|f(x+ 2s)−2f(x) +f(x−2s)| ≤ |f(x+ 2s)−f(x)|+|f(x)−f(x−2s)|</epsilon1,
142 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
while in the [ δ,π
2] interval,
|f(x+ 2s)−2f(x) +f(x−2s)| ≤ |f(x+ 2s)|+ 2|f(x)|+|f(x−2s)| ≤4M,
whereM= max
x∈[−π,π]|f(x)|. Hence
|f(x)−TN(x)|<cN[/epsilon1/integraldisplayδ
0cos2Nsds+ 4M/integraldisplayπ/2
0cos2Nsds].
Now we observe that
/integraldisplayδ
0cos2Nsds</integraldisplayπ/2
0cos2Nsds=1
2cN,
and that, since cos sdecreases as sgoes toπ/2 ,
/integraldisplayπ
δcos2Nsds</integraldisplayπ/2
δcos2Nδds<π
2γN,
whereγ= cos2δ<1 . Thus
|f(x)−TN(x)|</epsilon1
2+ 2πMc NγN.
NowπcN=/parenleftBig
2
3·4
5···2N−2
2N−1/parenrightBig
·2N < 2N, so that 2πMc NγN<4MNγN. Becauseγ <1 ,
we know that lim
N→∞NγN= 0 . Thus, pick Nso large that NγN</epsilon1
8M, where this is the
same/epsilon1as before. Consequently, for this N,
|f(x)−TN(x)|</epsilon1.
Since/epsilon1is independent of x,
/bardblf(x)−TN(x)/bardbl= max
x∈[−π,π]|f(x)−TN(x)|</epsilon1
too. A difficult lemma is thereby proved.
The whole proof is completed in the following simple
Lemma 3.32 . The function TN(x)defined by (3)
TN(x) =cN
2/integraldisplayπ
−πf(t) cos2N(t−x
2)dt
is a trigonometric polynomial.
Proof: This can be horribly messy unless one is shrewd. We shall use the formula eiθ=
cosθ+isinθand the binomial theorem (top p. 108). First notice that
cos2Nθ=/parenleftbiggeiθ+e−iθ
2/parenrightbigg2N
=1
22N2N/summationdisplay
k=0(2N)!
(2N−k)k!eikθe−i(2N−k)θ.
3.5. APPENDIX. THE WEIERSTRASS APPROXIMATION THEOREM 143
Letdk= (2N)!/22N(2N−k)!k! Then
cos2Nθ=2N/summationdisplay
k=0dke−i(2N−2k)θ
=2N/summationdisplay
k=0dk[cos(2N−2k)θ−isin(2N−2k)θ].(3-31)
Since cos2Nθis real, the sum of the imaginary terms on the right must be zero. Thus,
replacing 2 θbyt−x, we find that
cos2N(t−x
)2 =2N/summationdisplay
k=0dkcos(N−k)(t−x)
=2N/summationdisplay
k=0dk[cos(N−k)tcos(N−k)x+ sin(N−k)tsin(N−k)x].(3-32)
Split the sum into two parts, one from 0 to N, the other from N+ 1 to 2N, and let
n=N−kin the first, n=k−Nin the second. This gives
cos2N(t−x
2) =N/summationdisplay
n=0dN−n[cosntcosnx+ sinntsinnx]
+N/summationdisplay
n=1dN+n[cosntcosnx+ sinntsinnx],(3-33)
so
cos2N(t−x
2) =dN+N/summationdisplay
n=1(dN+n+dN−n)[cosntcosnx+ sinntsinnx],
which is much more simple than one might have anticipated. Substituting this into (4)
and realizing that the tintegrations just yield constants, we find that TN(x) is indeed a
trigonometric polynomial. Coupled with Lemma 3, the proof of Weierstrass’ Approximation
Theorem is completely proved.
Exercises
(1) Find the Fourier series with period 2 πfor the given functions.
(a)f(x) =/braceleftbigg0,−π≤x≤0
2,0<x<π
(b)f(x) =/braceleftbigg−2,−π≤x<0
2,0≤x<π
(c)f(x) = sin 17x+ cos 2s,−π≤x<π
(d)f(x) = sin2x,−π≤x≤π;
(e)f(x) =x2,−π≤x≤π
144 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
(f)f(x) =/braceleftbiggx+π,−π≤x≤0
−x+π,0≤x≤π(Also, compute /bardblf/bardbl2anda2
0+∞/summationdisplay
n=1(a2
n+b2
n)
for (a)-(f)).
(2) (a) Apply Parseval’s Theorem (Corollary to Theorem 22) to the function f(x) =x
and its Fourier series to deduce that
π2
6= 1 +1
22+1
32+1
42+···
(cf. the example before Theorem 17 of Section 3).
(b) Do the same for the function f(x) =x2(Ex. 1, e above) to evaluate
1 +1
24+1
34+1
44+···=?
(3) A function f(x) iseven iff(−x) =f(x) ,oddiff(−x) =−f(x) . Thus 2 + x2is
an even function, x3−sinxis an odd function, while 1 + xis neither even nor odd.
Letanandbnbe the Fourier coefficients of the piecewise continuous function f(x) .
Prove the following statements.
(a) Iffis an oddfunction,
an= 0, b n= 2/integraldisplayπ
0f(x)sinnx√πdx
(b) Iffis an even function
an= 2/integraldisplayπ
0f(x)cosnx√πdx, b n= 0
(c) A function fdefined in [0 ,π] may be extended to [ −π,π] as either an even or
odd function by the formulas
even extension :f(−x) =f(x), x≥0,
or
odd extension :f(−x) =−f(x), x≥0.
The even extension of f(x) =x, x∈[0,π] isf(x) =|x|, x∈[−π,π] , while its
odd extension is f(x) =x, x∈[−π,π] . The odd extension of f(x) =x2, x∈
[0,π] isf(x) =/braceleftbiggx2, x∈[0,π]
−x2, x∈[−π,0]. Extend the function f(x) = 1, x∈[0,π]
to the interval [ −π,π] as an odd function and sketch its graph. Find its Fourier
series using part (a).
(4) (a) Let f(x) be a given function. Find a solution of the O.D. E. u/prime/prime+λ2u=f, where
λis a real number and u(x) satisfies the boundary condition u(−π) =u(π) = 0 ,
by the following procedure: Expand fin its Fourier series and assume uhas a
Fourier series whose coefficients are to be found. Find a formula for the Fourier
coefficients of uin terms of those for fin the case where λis not an integer.
3.5. APPENDIX. THE WEIERSTRASS APPROXIMATION THEOREM 145
(b) Ifλ=nis an integer, show that there is a solution if and only if 0 = /angbracketleftf,˜en/angbracketright=/integraldisplayπ
−πf(x)sinnx√πdx.
(5) (a) State Parseval’s Theorem for the special cases i) fis a continuous even function
in [−π,π] , and ii)fis a continuous odd function in [ −π,π] .
(b) Iffis a continuous even function in [ −π,π] and
/integraldisplayπ
0f(x) cosnxdx = 0, n = 0,1,2,3,...,
show thatf= 0 in [ −π,π] .
(c) State and prove a theorem similar to (b) in the case of a continuous odd function.
(6) In this exercise you show how a function f∈L2[−A,A] can be expanded in a modified
Fourier series (so far we know only L2[−π,π] ). Lety=πx
A—this maps the interval
[−A,A] onto [ −π,π] —and define g(y) by
f(x) =f(Ay
π) =g(y) =g(πx
A).
Sinceg(y)∈L2[−π,π] , it can be expanded in a Fourier series
g(y) =a01√
2π+∞/summationdisplay
n=1ancosny√π+bnsinny√π,
where theanandbnare given by the usual formulas (2)’.
(a) Prove that f(x)∈L2[−A,A] has the modified Fourier series
f(x) =a01√
2A+∞/summationdisplay
n=1cosnx
Ax+bn√
Asinnπ
Ax,
where
a0=1√
2A/integraldisplayA
−Af(x)dx
an=1√
A/integraldisplayA
−Af(x) cosnπx
Adx, b n=1√
A/integraldisplayA
−Af(x) sinnπx
Adx.
(b) Find the modified Fourier series for f(x) =|x|, in the interval [ −1,1] .
The following exercises all concern the Weierstrass Approximation Theorem.
(7) Prove the following version of the Weierstrass Approximation Theorem. Let f∈
C[a,b] . Then given any /epsilon1>0 , there is a polynomial Q(x) such that
/bardblf−Q/bardbl∞= max
x∈[a,b]|f(x)−Q(x)|</epsilon1.
(Hint: Let y=−π+2(x−a)
b−aπ. This maps [ a,b] into [ −π,π] . Defineg(y), y∈[−π,π]
by
f(x) =f(a+(b−a)
2π(y+π) =g(y) =g(−π+ 2(x−a)
b−aπ).
146 CHAPTER 3. LINEAR SPACES: NORMS AND INNER PRODUCTS
Use the version of the theorem proved to approximate g(y), y∈[−π,π] by a trigono-
metric polynomial TN(y) to within /epsilon1/2 . Then approximate sin nyand cosnyto
withinc/epsilon1(you pickc) by a finite piece of their Taylor series—which are polynomials.
Put both parts together to obtain the complete proof for g(y) . The transition back
tof(x) is trivial.]
(8) (Riemann-Lebesgue Lemma). Let f∈C[a,b] . Prove that
lim
λ→∞/integraldisplayb
af(x) sinλxdx = 0.
[Hint: Integrate by parts to prove it first for all f∈C1[a,b] . For arbitrary f,
approximate fby a polynomial—Ex. 7 above—to within /epsilon1/2 and realize that every
polynomial is in C1[a,b] ].
(9) Iff∈C[0,1] , prove that
lim
n→∞n/integraldisplay1
0f(x)xndx=f(1).
[Hint: Use the hint in Ex. 8].
(10) Iff∈C[a,b] , and if
/integraldisplayb
af(x)xndx= 0, n = 0,1,2,3,...,
show that f= 0 . [ Hint: This implies that/integraldisplayb
af(x)Q(x)dx= 0 , where Qis
anypolynomial. fcan be approximated by some polynomial ˜Q. Now show that/integraldisplayb
af2(x)dx= 0 .]
3.6 The Vector Product in R3
.
As you grasped many years ago, the world we live in has three space dimensions. For
this reason the material in this section is important in many applications. What we intend
to do is define a way to multiply two vectors XandYinR3. Whereas the scalar product
/angbracketleftX, Y/angbracketrightis ascalar , this product X×Y, the vector product , orcross product as it is often
called, is a vector .
For several reasons [i) we shall not cover this in class, and ii) I can probably not do as
good a job as appears in many books] we shall let you read about this topic elsewhere. But
make sure to read about it even though you’ll never be examined on it.
Chapter 4
Linear Operators: Generalities.
V1→Vn, Vn→V1
4.1 Introduction. Algebra of Operators
.
LetVby a linear space. So far we have considered the algebraic structure of such a
space; however most significant reason for studying linear spaces is so that one can study
operators defined on them. Operator is another, more organic, name for function. Thus an
operator
T:A→B
Tmaps elements in its domain Ainto elements of B, whereBcontains the range of T.
IfX∈A, thenT(X) =Y∈B. Think of feeding Xinto the operator T, andYbeing
a figure goes here
whatTsends out in return. It is useful to think of Tas some type of machine or factory,
the input (raw material) is X, and the output is Y. Some examples should illustrate the
situation and its potential power.
Examples:
(1) Let V=R2. IfX= (x1,x2)∈R2, andY= (y1,y2,y3) , we define T(X) =Yby
T(X) =
x1+ 2x2=y1
x1+x2=y2
3x1+x2=y3
,
or
T(X) =T(x1,x2) = (x1+ 2x2, x1+x2,3x1+x2) = (y1,y2,y3) =Y.
This operator Thas the property that to every X∈R2it assigns a Y∈R3. In
other words Tmaps the two dimensional space R2into the three dimensional space
R3
T:R2→R3.
147
148 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
R2is the domain ofT, denoted by D(T) , while the range ofT,R(T) is contained
inR3,
D(T) =R2,R(T)⊂R3.
Sincey1=y2= 0 implies that x1=x2= 0 , which in turn implies that y3= 0 , we
see that the point (0 ,0,1)∈R3isnotin the range of T. Thus,Tisnot surjective
onto R3. It is injective (one-to-one) since every point Y∈R(T) is the image of
exactly one X∈D(T) . This an be seen by observing that y1andy2suffice to
determineX= (x1,x2) uniquely by solving the first two equations
−y1+ 2y2=x1
y1−y2=x2.
Hence ifY=T(X1) and also Y=T(X2) , thenX1=X2. Since the operator
Tis completely determined by the coefficients in the equations, it is reasonable to
represent this Tby the matrix
T=
1 2
1 1
3 1
If you care to think of Xas the input into a paint-making machine, then x1might
represent the quantity of yellow and x2the quantity of blue used. In this case y1,y2
andy3represent the quantities of three different shades of green the machine yields.
For this machine, as soon as you specify the desired quantities of any two of the
greens, say y1andy2, the quantities x1andx2of the input colors are completely
determined, as is the quantity y3of the remaining shade of green.
(2) Let VbeR2again. With X= (x1,x2)∈R2, andY= (y1)∈R1, defineTby
x2
1+x2
2=y1,
or
T(X) =x2
1+x2
2.
This operator Tmaps R2intoR1
T:R2→R1.
It is not surjective onto R1since the negative half of R1is completely omitted from
R(T) . Furthermore, it is not injective either since each point y1∈R(T) other than
zero is the image of infinitely many points—all of those on the circle x2
1+x2
2=y1.
(3) Let VbeC[−1,1] . Iff∈C[−1,1] , we define Tby
T(f) =f(0).
Thus, iff(x) = 2 + cos x, thenTf= 3 . This operator Tis usually denoted by
δand called the Dirac delta functional . It was first used by Dirac in his work on
quantum mechanics and is extremely valuable in modern mathematics and physics.
Tassigns to each continuous function fits value at x= 0 , a real number. Therefore
T:C[−1,1]→R1.
4.1. INTRODUCTION. ALGEBRA OF OPERATORS 149
The operator Tis not injective, since for example the element 2 ∈R1is the image
of bothf(x) = 1 +exandf(x) = 2 . It is surjective since every element a∈R1is
the image of at least one element in C[−1,1] (iff(x)≡a, then clearly T(f) =a).
(4) Let VbeC[−1,1] . Iff∈C1[−1,1] then the differentiation operator Dis defined
by
(Df)(x) =df
dx(x).
It maps each function into its derivative. If f(x) =x2, then (Df)(x) = 2x. Since the
derivative of a continuously differentiable function (a function in C1) is necessarily
continuous, we see that
D:C1[−1,1]→C[−1,1].
Dis not injective since, for example, the function g(x) = 1 is the image of both
f1(x) =xandf2(x) = 2 +x.Dis surjective onto C[−1,1] .
R(D) =C[−1,1],
since ifg(x) is any element of C[−1,1] , thengis the image of the particular function
f∈C1[−1,1] defined by
f(x) =/integraldisplayx
0g(s)ds,
becauseDf=gby the fundamental theorem of calculus.
Throughout this and the next chapter we will study some of the elementary aspects
of linear operators. It is reasonable to denote a linear operator by L.
Definition Let V1andV2both be linear spaces over the same field of scalars. An
operatorLmapping V1intoV2is called a linear operator if for every Xand ˜X
inV1and any scalar a,Lsatisfies the two conditions
1.L(X+˜X) =L(X) +L(˜X)
2.L(aX) =aL(X).
Whenever ambiguity does not arise, we will omit the parentheses and write LX
instead ofL(X) .
An equivalent form of the definition is
Theorem 4.1 .Lis a linear operator ⇐⇒
L(aX+b˜X) =aL(X) +bL(˜X),
whereX,˜X∈V1andaandbare any scalars.
Proof: ⇒
L(aX+b˜X) =L(aX) +L(b˜X) (property 1)
=aLX +bL˜X(property 2) .(4-1)
⇐Property 1 is the special case a=b= 1 . Property 2 is the special case b= 0 .
Remark:
It is useful to observe that always L(0) =L(0·X) = 0L(X) = 0 . This identity is often
the easiest way to test if an operator is notlinear.
150 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
Examples:
(1) The operator Ldefined by example 1 where L:R2→R3is
LX= (x1+ 2x2,x1+x2,3x1+x2),
is linear. Let X= (x1,x2) and ˜X= (˜x1,˜x2) . Then
L(X+˜X) = (x1+ ˜x1+ 2x2+ 2˜x2,x1+ ˜x1+x2+ ˜x2,3x1+ 3˜x1+x2+ ˜x2)
= (x1+ 2x2,x1+x2,3x1+x2) + (˜x1+ 2˜x2,˜x1+ ˜x2,3˜x1+ ˜x2)
=LX+L˜X
and
L(aX) = (ax1+ 2ax2,ax 1+ax2,3ax1+ax2)
=a(x1+ 2x2,x1+x2,3xz+x2)
=aLX.
(4-2)
(2) The operator TX=x2
1+x2
2with domain R2and range R1isnotlinear, since
T(aX) = (ax1)2+ (ax2)2=a2[x2
1+x2
2]/negationslash=aTX
except for the particular scalars a= 0,1 .
(3) The operator Df=d f
dxwith domain C1[−1,1] and range C[−1,1] is linear since if
f1andf2are inC1[−1,1] andaandbare any real numbers, then by elementary
calculus
D(af1+bf2) =d
dx(af1+bf2) =adf1
dx+bdf2
dx
=aDf 1+bDf 2.(4-3)
(4) The operator Ldefined as
Lu=a2(x)u/prime/prime+a1(x)u/prime+a0(x)u,(/prime=d
dx),
whereu(x)∈D(L) =C2, and where a0(x) ,a1(x) , anda2(x) are continuous
functions, is a linear operator,
L:C2→C.
IfAandBare any constants (scalars for C2), then for any u1andu2∈C2,
L(Au1+Bu2) =ax[Au1+Bu2]/prime/prime+a1[Au1+Bu2]/prime+a0[Au1+Bu2]
=a2Au/prime/prime
1+a2B/prime/prime
2+a1Au/prime
1+a1Bu/prime
2+a0Au1+a0Bu2
=A[a2u/prime/prime
1+a1u/prime
1+a0u1] +B[a2u/prime/prime
2+a1u/prime
2+a0u2]
=ALu 1+BLu 2.(4-4)
4.1. INTRODUCTION. ALGEBRA OF OPERATORS 151
(5) The identity operator Iis the operator which leaves everything unchanged. Because
it is so simple, it can be defined on an arbitrary set Sand mapsSintoitselfS→S
in a trivial way. If X∈S, then we define
IX=X.
What could be more simple? If Sis a linear space V(soaXandX1+X2are
defined), then Iis trivially a linear operator, since
I(aX1+bX2) =aX1+bX2=aIX 1+bIX 2
Why are linear operators important? There are several reasons. First, they are much
easier to work with than nonlinear operators. Second, most of the operators which arise
in applications are linear. The feature possessed by linear operators which is central to
applications is that of superposition . IfLu1=fandLu2=g, thenL(u1+u2) =f+g.
In other words, if u1is the response to some external influence fandu2the response to
g, then the response to f+gis found by adding the separate responses.
The special case of a linear operator whose range is the real number line R1arises often
enough to receive a name of its own.
Definition: A linear operator whose range is R1is called a linear functional ,/lscriptV→R1.
The Dirac delta functional is such an operator. So is the operator
l(f) =/integraldisplay1
0f(x)dx,
which assigns to every continuous function f∈C[0,1] the real number equal to the area
between the graph of fand thex-axis. Check that /lscriptis linear.
If the linear operator L:V1→V2the range of L—a subset of the linear space V2—has
a particularly nice structure. In fact, R(L) is not just any clump of points in V2but
Theorem 4.2 . The range of a linear operator L:V1→V2is a linear subspace of V2.
Remark: Even more is true. We shall prove (p. 312-3) that dim R(L)≤dimD(L) so that
no matter how large V2is, the range has at most the same dimension as the domain.
Proof: The range of Lconsists of all elements Y∈V2of the form Y=LX where
X∈V1. We know that R(L) is a subset of the linear space V2. The only task is to prove
that it is actually a subspace. Since V2is a linear space, it is sufficient to show that the
setR(L) is closed under multiplication by scalars, and under addition of vectors. i) R(L)
is closed under multiplication by scalars. If Y∈R(L) , there is an X∈V1=D(L) such
thatY+LX. We must find some ˜XinV1such thataY=L˜X, whereais any scalar.
SinceaY=aLX =L(aX) , we take ˜X=aX.
ii)R(L) is closed under addition of vectors. If Y1andY2are in R(L) , there are elements
X1andX2inV1=D(L) such that Y1=LX 1andY2=LX 2. We must show that
Y1+Y2∈D(L) , that is, find some ˜X∈V1such thatY1+Y2=L˜X. ButY1+Y2=
LX 1+LX 2=L(X1+X2) . Thus we can take ˜X=X1+X2.
Before moving further on into the realm of special linear operators, we shall take this
opportunity to define algebraic operations (addition and multiplication) for linear operators.
But first we define equality, L1=L2, in a straightforward way.
152 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
Definition: (equality ) IfL1andL2both mapV1intoV2, whereV1andV2are linear
spaces, and if L1X=L2Xfor allXinV1, thenL1equalsL2. Thus, two operators are
equal if they have the same effect on any vector.
Addition is equally simple.
Definition: (addition ). IfL1:V1→V2andL2:V1→V2then their sum, L1+L2, is
defined by the rule
(L1+L2)X=L1X+L2X, X ∈V1
Examples:
(1) LetL1:R2→R3be defined by
L1(X) = (x1+x2,x1+ 2x2,−x2), X = (x1,x2)∈R2
andL2:R2→R3be defined by
L2X= (−3x1+x2,x1−x2,x1), X = (x1,x2)∈R2.
ThenL1+L2is defined, and is
(L1+L2)X+L1X+L2X= (x1+x2,x1+ 2x2,−x2) + (−3x1+x2,x1−x2,x1)
= (−2x1+ 2x2,2x1+x2,x1−x2)
(4-5)
(2) LetD:C1→Cbe defined by
Du=du
dxu∈C1,
andL:C1→Cbe defined by
Lu=/integraldisplay1
0ex−tu(t)dt u ∈C1
=ex/integraldisplay1
0e−tu(t)dt.(4-6)
(In reality, Lmay be defined on a much larger class of functions— u∈Cis plenty, while
its image is the smaller space, constant ex⊂C. We have decided on the smaller domain
and larger image space so that the sum D+Lis defined). Then for any u∈C1.
(D+L)u=Du+Lu=du
dx+/integraldisplay1
0ex−tu(t)dt.
The following theorem is a statement of some simple facts about the sum of two linear
operators.
Theorem 4.3 . LetL1,L2,L3,... be any linear operators which map V1→V2, so that
their sums are defined. Then
0.L=L1+L2is a linear operator
(1)L1+ (L2+L3) + (L1+L2) +L3,
4.1. INTRODUCTION. ALGEBRA OF OPERATORS 153
(2)L1+L2=L2+L1
(3)Let 0 be the operator which maps every element of V1into 0∈V2, so 0X= 0.
Then
L1+ 0 =L1.
(4)L1+ (−L1) = 0 . Here −L1is the operator which maps every element X∈V1
into−(L1X).
Proof: These are just computations. Let X1,X2∈V1.
0.
L(aX1+bX2) = (L1+L2)(aX1+bX2)
=L1(aX1+bX2) +L2(aX1+bX2)
=aL1X1+bL1X2+aL2X1+bL2X2
=a(L1X1+L2X1) +b(L1X2+L2X2)
=a(L1+L2)X1+b(L1+L2)X2
=aLX 1+bLX 2.(4-7)
(1) (L1+(L2+L3))X=L1X+(L2+L3)X=L1X+L2X+L3X= (L1+L2)X+L3X=
((L1+L2) +L3)X.
(2) (L1+L2)X=L1X+L2X=L2X+L1X= (L2+L1)X. The step L1X+L2X=
L2X+L1Xis justified on the grounds that the vectors Y1:=L1XandY2:=L2X
are elements of V2—which is a linear space—so that Y1+Y2=Y2+Y1.
(3) (L1+ 0)X=L1X+ 0X=L1X+ 0 =L1X
Note that the 0 in 0 Xis an operator, while the 0 in the next step is an element of
V2. This ambiguity causes no trouble once you understand it.
(4) (L1+ (−L1))X=L1X+ (−L1)X=L1X−L1X= 0
The crucial step ( −L1)X=−L1Xis the definition of the operator ( −L1) .
Remark: This theorem states that the set of all linear operators mapping one linear space
V1into another V2form an abelian group under addition.
Multiplication of operators is not much more difficult. If L1andL2are linear opera-
tors, then their product L2L1in that order is defined by the rule L2L1X=L2(L1X) . In
other words, firstoperate on XwithL1giving a vector Y=L1X.Then operate on this
new vector YwithL2, givingL2Y=L2(L1X) . It is clear that in order for this to make
sense, for every X∈D(L1) , the new vector Y=L1Xmust be in the domain of L2. Thus
to form the product L2L1, we require that R(L1)⊂D(L2) .
Look at our machine again.
a figure goes here
154 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
The multiplication L2L1means sending the output from L1as input into L2. In order
to join the machines in this way, surely one necessary requirement is that L2is equipped
to act on the output from R(L1) , that is, R(L1)⊂D(L2) . Of course the L2machine
might be able to digest input other than what L1sends out. But all we care is that L2
can digest at least whatL1sends it.
Definition: (multiplication ). LetL1:V1→V2andL2:V3→V4. If the range of L1
is contained in the domain of L2,R(L1)⊂D(L2) , then the productL2L1is definable by
the composition rule
L2L1X=L2(L1X),whereX∈V1=D(L1).
The product L2L1maps the input V1forL1into the output V4forL2, L2L1:V1→
V3→V4.
We exhibit a little diagram (cf. p. ???).
a figure goes here
The way to get from V1toV4usingL2L1is to first use L1to reachV2. Then use L2
to get toV4.
Remarks: IfL2L1is defined, it is notnecessarily true that L1L2is defined (Example
1 below). Furthermore, even if L1L2is also defined, it is only a rare coincidence that
multiplication is commutative. Usually L2L1/negationslash=L1L2when both products are defined.
Thus the orderL2L1isimportant .
Examples:
(1) LetL1:R2→R3be defined as
L1X= (x1−x2,x2,−x1−2x2),whereX= (x1,x2)∈R2,
and letL2:R3→R1be defined as
L2Y= (y1+ 2y2−y3),whereY= (y1,y2,y3)∈R3.
Then R(L1)⊂R3=D(L2) so that the product L2L1is definable and L2L1:R2L1→
R3L2→R1. Consider what L2L1does to the particular vector X0= (−1,2)∈R2.
L2L1X0=L2(L1X0) =L2(−3,2,−3) = ( −3 + 4 + 3 = 4)
ThusL2L1maps ( −1,2)∈R2into 4 ∈R1. More generally, if Xis any vector in
R2,
L2L1X=L2(L1X) =L2(x1−x2,x2,−x1−2x2)
= (x1−x2+ 2x2+x1+ 2x2) = 2x1+ 3x2∈R1.(4-8)
ThusL2L1maps (x1,x2)∈R2into 2x1+ 3x2∈R1.
Since R(L2) =R1andD(L1) =R2,R(L2) not ⊂D(L1) so that the product L1L2
isnotdefined. You might be thinking that R1is part of R2. What you mean is that
R2has one dimensional subspaces. It certainly does—an infinite number of them, all
of the straight lines through the origin. Because there are so many subspaces of R2
4.1. INTRODUCTION. ALGEBRA OF OPERATORS 155
which are one dimensional, there is no natural way of regarding R1as being contained
inR2. [On the other hand, there is a natural way in which C1can be regarded as
contained in C. We used this above in our second example for addition of linear
operators].
(2) Define L1:R2→R2by the rule L1X= (2x1−3x2,−x1+x2) andL2:R2→R2by
the ruleL2X= (2x2,x1+x2) Then R(L1) =R2=D(L2) so thatL2L1is defined.
It is given by
L2L1X=L2(2x1−3x2,−x1+x2) = (−2x1+ 2x2,x1−2x2)
In particular, L2L1mapsX0= (1,2) into (2,−3) . Now R(L2) =R2=D(L1) , so
thatL1L2is also definable. It is given by
L1L2X=L1(2x2,x1+x2)
= (2·2x2−3·(x1+x2),−2x2+ (x1+x2))
= (−3x1+x2,x1−x2).(4-9)
In particular, L1L2mapsX0= (1,2) into ( −1,−1) . SinceL1L2andL2L1map
the pointX0= (1,2) into two different points, it is clear that L1L2/negationslash=L2L1, the
operators do not commute.
(3) LetAbe the subspace of R2spanned by some unit vector e1andBbe the subspace
spanned by another unit vector e2. Consider the projection operators PAandPB.
They are linear since, for example,
PA(aX1+bX2) =/angbracketleftaX1+bX2, e1/angbracketrighte1
=a/angbracketleftX1, e1/angbracketrighte1+b/angbracketleftX2, e1/angbracketrighte1
=aPAX1+bPAX2.(4-10)
BecausePA:R2→R2andPB:R2→R2, both products PAPBandPBPAare
defined. We have
PAPBX=PA(PBX) =PA(/angbracketleftX, e 2/angbracketrighte2)
=/angbracketleftX, e 2/angbracketrightPAe2=/angbracketleftX, e 2/angbracketright/angbracketlefte2, e1/angbracketrighte1.(4-11)
Also,
PBPAX=PB(PAX) =PB(/angbracketleftX, e 1/angbracketrighte1)
=/angbracketleftX, e 1/angbracketrightPBe1=/angbracketleftX, e 1/angbracketright/angbracketlefte1, e2/angbracketrighte2.(4-12)
SincePAPBX∈A⊂R2, whilePBPAX∈B⊂R2, it is clear that usually PAPB/negationslash=
PBPA. They will happen to be equal if A=B, or ifA⊥B(for thenPAPB=
PBPA= 0 ). See the figure at the beginning of this example—and draw some more
special cases for yourself.
(4) LetL:C∞→C∞(C∞is the space of infinitely differentiable functions) be defined
by
(Lu)(x) =xu(x), u∈C∞,
156 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
andD:C∞→C∞be defined by
(Du)(x) =du
dx(x), u∈C∞.
Then R(L) =D(D) so that the product DLis definable by
DLu =D(Lu) =D(xu) =d
dx(xu(x)) =xu/prime+u.
Also, R(D) =D(L) soLDis definable by
LDu =L(Du) =L(u/prime) =xu/prime.
Notice that LD/negationslash=DLunlessu= 0 .
We collect some properties of multiplication.
Theorem 4.4 . IfL1:V1→V2, L2:V3→V4, andL3:V5→V6, whereV1⊂V3and
V4⊂V5, then
0. The operator L=L2L1is a linear operator.
1.L3(L2L1) = (L3L2)L1—Associative law.
Proof: 0.
L(aX1+bX2) =L2/parenleftbig
L1(aX1+bX2)/parenrightbig
=L2/parenleftbig
aL1X1+bL1X2/parenrightbig
=L2(aL1X1) +L2(bL1X2)
=aL2L1X1+bL2L1X2
=aLX 1+bLX 2.(4-13)
(1) By definition of the product,
[L3(L2L1)]X=L3[(L2L1)X] =L3[L2(L1X)]
and
[(L3L2)L1]X= (L3L2)(L1X) =L3[L2(L1X)].
Now match the ends.
Notice that the commonly occurring special case V1=V2=V3=V4=V5=V6is
included in this theorem. In this special case, even more can be proved. For then the
identity operator I, defined by IX=Xfor allX∈Vcan be used to multiply any other
operator. Moreover, addition, L1+L2also makes sense.
Theorem 4.5 . If the linear operators L1,L2,L3all mapVintoV, then representing
any one of these by L,
(1)LI=IL=L.
(2)For any positive integer n, we define Lninductively by the rule Ln+1=LLn,
andL0=I. Then for any non-negative integers mandn,
Lm+1=LmLn.
4.1. INTRODUCTION. ALGEBRA OF OPERATORS 157
(3) (L1+L2)L3=L1L3+L2L3.
(4)L3(L1+L2) =L3L1+L3L2(This is needed in addition to 3 because of the non-
commutativity).
Proof:
(1) IfX∈V,
(LI)X=L(IX) =LX
(IL)X=I(LX) =LX
(2) We shall prove Lm+n=LmLnby induction on m. The statement is true, by
definition, for m= 1 . Assume it is true for m=k, soLk+n=LkLn. Our job is to
prove the statement for m=k+ 1 . By the definition and the induction hypothesis,
we have
Lk+n+1=LLk+n=L(LkLn).
Since multiplication is associative, we find that
L(LkLn) = (LLk)Ln.
But, by definition,
LLk=Lk+1.
Thus,
Lk+n+1=Lk+1Ln.
This completes the induction proof.
(3) IfX∈V,
[(L1+L2)L3]X= (L1+L2)(L3X)
LetL3X=Y∈V. Then (L1+L2)Y=L1Y+L2Y.Thus
[(L1+L2)L3]X=L1(L3X) +L2(L3X) = (L1,L3)X+ (L2L3)X.
(4) Same proof as 3.
Remark: IfV1andV2are two linear spaces, the set of all linear operators which
mapV1intoV2is usually denoted by Hom( V1,V2) —Hom rhymes with Mom and
Tom. In this notation, the last theorem concerned Hom( V,V) . The abbreviation
Hom is for the impressive word “homomorphism”. Tell your friends.
Examples: ConsiderD:C∞→C∞defined by ( Du)(x) =du
dx(x) . ThenDn=
dn
dxn.
Exercises
(1) Determine which of the following are linear operators.
158 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
(a)T:R2→R2
TX= (x1+x2,x1−x2),
whereX= (x1,x2)∈R2.
(b)T:R2→R2,
T(X) = (x1+x2+ 1,x1−x2)
(c)T:R3→R2
T(X) = (x1+x1x2,x2)
(d)T:R3→R1
T(X) = (x1+x2−x3)
(e)T:R3→R1
T(X) + (x1+x2−x3+ 2)
(f)D:P2→P1. IfP(x) =a2x2+a1x+a0∈P2then
D(P) = 2a2+a1∈P1.
(g)T:C1[−1,1]→R1. Ifu(x)∈C1[−1,1] , then
T(u) =u(0) +u/prime(0).
(h)T:C[2,3]→C[2,3] . Ifu∈C[2,3] ,
(Tu)(x) =/integraldisplay3
2ex−tu(t)dt
(i)T:C[2,3]→C[2,3] ,
(Tu)(x) = 1 +/integraldisplay3
2ex−tu(t)dt
(j)T:C[2,3]→C[2,3],
(Tu)(x) =/integraldisplay3
2ex−tu2(t)dt
(k)S1:C[0,∞]→C[0,∞]
(S1u)(x) =u(x+ 1)−u(x)
(l)L:A→C[0,∞] ,
whereA={u∈C[0,∞]:/integraltext∞
0|u(t)|dt<∞},
(Lu)(x) =/integraldisplay∞
0e−xtu(t)dt,
[Our restriction on Ais just to insure that the integral exists. Luis usually
called the Laplace transform ofu].
4.1. INTRODUCTION. ALGEBRA OF OPERATORS 159
(m)T:C[0,∞]→C[0,∞]
(Tu)(x) =a2u(x2) +a1u(x+ 1) +a0u(x),
where theak(x) are continuous functions.
(n)T:C[0,1]→C[0,1] .
(Tu)(x) = 2xu(x).
(o)T:R2→R1
TX=|x1+x2|,whereX= (x1,x2)∈R2.
[Answers : a,d,f,g,h,k,l,m,n are linear].
(2) (a) If l(x) is a linear functional mapping R1→R1, prove that l(x) =αx, where
α=l(1) .
(b) Ifl(X) is a linear functional mapping Rn→R1, prove that l(X) =n/summationdisplay
k=1αkxk,
whereX= (x1,...,x n) .
(3) LetL1:R1→R2be defined by
L1X= (x1,3x1),whereX= (x1)∈R1,
andL2:R2→R2be defined by
L2Y= (y1+y2,y1+ 2y2),whereY= (y1,y2)∈R2.
ComputeL2L1X0, whereX0= 2∈R1. IsL1L2defined?
(4) LetA:R2→R2be defined by
AX= (x1+ 3x2,−x1−x2), X = (x1,x2)∈R2
andB:R2→R2by
BX= (−x1+x2,2x1+x2).
a). Compute ABX,BAX,B2X,A2BX, and (A+B)X.
b). Find an operator Csuch thatCA=I. [hint: LetCX= (c11x1+c12x2,c21x1+
c22x2) and solve for c11,c12, etc.]
(5) Consider the operators D:C∞→C∞,(Du) =u/primeandL:C∞→C∞,(Lu)(x) =/integraltextx
0u(t)dt.
(a) Show that DL=I, LD =I−δ, whereδis the delta functional.
(L2u)(x) =/integraldisplayx
0/parenleftbigg/integraldisplay2
0u(t)dt/parenrightbigg
ds.
Integrate by parts to conclude that
(L2u)(x) +/integraldisplayx
0(x−t)u(t)dt.
160 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
(b) Observe that D2L2=D(DL)L=DIL =DL=I. Use this observation to find
a solution of the differential equation D2u=fforu, wheref∈C∞. Solve
the particular equation ( D2u)(x) =1
1+x2
(6) LetA:R2→R2be defined by
AX= (a11x1+a12x2,a21x1+a22x2),
andB:R2→R2be defined by
BX= (b11x1+b12x2,b21x1+b22x2).
(a) Compute AB.
(b) Find a matrix Bsuch thatAB=I, that is, determine b11,b12,... in terms of
a11,a12,...such thatAB=I. [In the course of your computation, I suggest in-
troducing a symbol, say ∆ , for a11a12−a12a21when that algebraic combination
crops up.]
(7) In the plane E2, consider the operator Rwhich rotates a vector by 90oand the
operatorPprojecting onto the subspace spanned by e(see fig). (a) Prove that R
is linear. (b). Let X= (x1,x2) be any point on E2. Compute PRX andRPX .
Draw a sketch for the special case X= (1,1) .
(8) In R3, letAdenote the operator of rotation through 90oabout the x1-axis (so
A: (0,1,0)→(0,0,1) ),Bthe operator of rotation through 90oabout thex2-axis
andCthe operator of rotation through 90oabout thex3-axis (see fig.) Prove these
operators are linear (just do it for A). Show that A4=B4=C4=I, AB /negationslash=BA,
and thatA2B2=B2A2. Is it true that ABAB =A2B2?
(9) Let Pdenote the linear space of all polynomials in x. Forp∈P, consider the
operatorsDp=dp
dxandLp=xp. Show that DL−LD=I.
(10) (a) If L1L2=L2L1, prove that
(L1+L2)2=L2
1+ 2L1L2+L2
2.
(b) IfL1L2/negationslash=L2L1, then (L1+L2)2=?
(11) IfL1andL2are operators such that L1L2−L2L1=I, prove the formula Ln
1L2−
L2Ln
1=nLn−1
1, wheren= 1,2,3,....
(12) IfL1is a linear operator, L1:V1→V2[orL1∈Hom(V1,V2) ], andais any scalar,
define the operator L=aL1by the rule LX= (aL1)X=a(L1X) , whereX∈V1.
Prove
(0).L=aL1is a linear operator, L:V1→V1.
(5).a(bL1) = (ab)L1, wherea,bare any scalars.
(6). 1 ·L1=L1.
(7). (a+b)L1=aL1+bL1.
(8).a(L1+L2) =aL1+aL2, whereL2∈Hom(V1,V2) .
4.2. A DIGRESSION TO CONSIDER AU/prime/prime+BU/prime+CU=F. 161
Coupled with Theorem 3, this exercise proves that the set of all linear operators map-
ping one linear space in to another linear is itself a linear space , that is, Hom ( V1,V2)
is a linear space .
(13) (a). In E2, letLdenote the operator which rotates a vector by 90o. ThenL:E2→
E2. IfX= (x1,x2) =x1e1+x2e2, wheree1= (1,0) ande2= (0,1) , writeLas
LX= (a11x1+a12x2,a21x1+a22x2),
That is, find the coefficients a11,a12,.... This gives two ways to represent L, as
a rotation (geometrically), and by linear equations in terms of a particular basis
(algebraically).
(b). In E2, consider the operator Lof rotation through an angle α. Show that
Le1= (cosα,sinα), Le 2= (−sinα,cosα),
and then deduce that if X= (x1,x2) =x1e1+x2e2,
LX= (x1cosα−x2sinα, x 1sinα+x2cosα).
(14) Consider the space Pnof all polynomial of degree n. DefineL:Pn→Pnas the
translation operator ( Lp)(x) =p(x+ 1) , and D:Pn→Pnas the differentiation
operator, ( Dp)(x) =dp
dx(x) . Show that
L=I+D+D2
2!+···+Dn−1
(n−1)!+Dn
n!
(15) Consider the linear operators L1=a1D2+b1D+c1I, andL2=a2D2+b2D+e2I.
BothL1andL2map the linear space of infinitely differentiable function into itself,
Lj:C∞→C∞. If the coefficients a1,a2,b1,... areconstants , prove that L1L2=
L2L1.
4.2 A Digression to Consider au/prime/prime+bu/prime+cu=f.
Essentially the only linear equation youcan solve explicitly are linear algebraic equations,
like two equations in two unknowns. Since our theory applies to much more general situa-
tions, we shall develop a different example for you to keep in the back of your minds along
with that of linear algebraic equations. The example we have chosen has the additional
virtue that it contains most of the solvable differential equations which arise anywhere.
Watch closely because we shall be brief and with a high density of valuable ideas.
Problems concerning vibration or oscillatory phenomena are among the most important
and significant ones which arise in applications. The simplest case is that of a simple
harmonic oscillator . We have
a figure goes here
162 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
a massmattached to a spring. Pull the mass back a little and watch it move back and
forth, back and forth. These are oscillations. To make the situation simple, we assume
that the spring has no mass and that the surface upon which the mass rests is frictionless.
Letu(t) denote the displacement of the center of gravity of the mass from the equilibrium
position. Two experimental results are needed from physics.
1.Newton’s Second Law :m...u=/summationtextF, where/summationtextFmeans the resultant of all the
forces on the center of gravity of the mass (we assume all forces are acting horizontally).
2.Hooke’s Law : If a spring is not stretched too far, then the force it exerts is propor-
tional to the displacement,
F=−ku, k> 0.
We chose the minus sign since if a spring is displaced, the force it exerts is in the direction
opposite to the displacement. [Under larger displacements, actually
F(u) =a1u+a2u2+a3u3+...
-wherea0=F(0) = 0 . If the displacement uis small, the lowest term in the Taylor series
forF(u) gives an adequate approximation. This is a more precise statement of Hooke’s
Law].
Putting these two results together, we find that
m¨u=−ku+F1, (notation : ¨ u=d2u
dt2)
whereF1represents all of the remaining forces on the mass. One possible force (so far
incorporated into F1) is a so-called viscous damping force . It is of the form Fv=−µ˙uwhere
µ>0 ; at low velocities, this force is experimentally found to account for air resistance. It
is directed opposite to the velocity, and increases as the speed does (speed = /bardblvelocity /bardbl).
[Again,Fv=b1¨u+b2˙u2+..., that isFv( ˙u) is given by a Taylor series with Fv(0) = 0 . At
low speeds, the higher order terms can be neglected to yield a reasonable approximation.]
Thus, to our approximation,
m¨u=−ku−µ˙u+F2,
whereF2represents the forces yet unaccounted for. Let us assume that these remaining
forces do not depend on the motion and are applied by the outside world. Then the force
F2depends only on time, F2=f(t) . It is called the applied orexternal force . Newton’s
law gives
m¨u=−ku−µ˙u+f(t),
or
Lu: =a¨u=b˙u+cu=f(t),
wherea=m,b=µ, andc=k. For the purposes of our discussion, we shall assume that
kandµdo not depend on time. Then a,bandcare non negative constants.
In order to determine the motion of the mass, we must solve the ordinary differential
equationLu=fforu. Have we given enough information to determine the solution? In
other words, is the solution unique? For any physically reasonable problem, we expect the
mathematical model has a unique solution since (neglecting quantum mechanical effects)
once we let the mass go, it will certainly move in one particular way, the same way every
time we perform the same experiment. It is clear that the motion will depend on the initial
4.2. A DIGRESSION TO CONSIDER AU/prime/prime+BU/prime+CU=F. 163
positionu(t0) . But if two masses have the same initial position, the resulting motion will
still be different if their initial velocities ˙ u(t0) are different. Thus we must also specify the
initial velocity ˙ u(t0) as well as the initial position u(t0) . Are these sufficient to determine
the motion? Yes, however that requires proof. What must be proved is that if we have two
solutionsu1(t) andu2(t) of the same ordinary differential equation (1) and if their initial
positions and velocities coincide, then the solutions coincide, u1=u2for all later time,
t≥t0.
Theorem 4.6 (Uniqueness). Let u1(t)andu2(t)be two solutions of the ordinary differ-
ential equation
Lu: =a¨u+b˙u+cu=f(t),
wherea, b, andcare constants, a >0, b≥0, c≥0. Ifu1(t0) =u2(t0), and ˙u1(t0) =
˙u2(t0), thenu1(t) =u2(t)for allt≥0, in other words, the solution is uniquely determined
by the initial position and velocity.
Remark: The theorem is true under much more general conditions - as we shall prove in
Chapter 6.
Proof: Letw(t) =u2(t)−u1(t) . We shall show that w(t)≡0 for allt≥t0. Now
Lw=L(u2−u1) =Lu2−Lu1=f−f= 0,
that is,
a¨w+b˙w+cw= 0 (4-14)
Furthermore
w(t0) = 0 and ˙w(t0) = 0, (4-15)
sincew(t0) =u2(t0)−u1(t0) = 0 , and ˙ w(t0) = ˙u2(t0)−˙u1(t0) = 0 . This reduces the
question to showing that if Lw= 0 , and if whas zero initial position and velocity, then
in factw≡0 .
The trick is to introduce a new function, E(t) , associated with (2) (which happens to
be the total energy of the system)
E(t) =1
2a˙w2+1
2cw2.
How does this function change with time? We compute its derivative.
˙E(t) =a˙w¨w+cw˙w= ˙w(a¨w+cw).
Using (2) we know that a¨w+cw=−b˙w. Therefore
˙E(t) =−b˙w2≤0 (since b≥0)
[Thus energy is dissipated ( b>0 ) - or conserved ˙E= 0 in the special case of no damping
(b= 0 ).] Consequently
E(t)≤E(t0) for allt≥t0 (4-16)
Now observe that for the mechanical system associated with w, we haveE(t0) =a
w˙w2(t0)+
c
2w2(t0) = 0 . Furthermore, it is obvious from the definition of E(t) (sinceaandcare
positive) that 0 ≤E(t) . Substitution of this information into (4) reveals
0≤E(t)≤0 for all t≥t0.
164 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
This proves E(t)≡0 for allt≥t0, which in turn implies w(t)≡0 —again from the
definition of E(t) . Our proof is completed. We have taken some care since all of our
uniqueness proofs will use essentially no additional ideas. A more general case ( a, bandc
still constants but not necessarily positive) will be treated in Exercise 9.
Having proved that there is at most one solution of the initial value problem
Lu: =a¨u+b˙u+cu=f(t) (differential equation)
u(t0) =αand ˙u(t0) =β (initial conditions)
we must now prove there is at least one solution. This is the question of existence. For the
special equation (5), the solution is shown to exist by explicitly exhibiting it. In the case of
more complicated equations we are not as fortunate and must content ourselves with just
showing that a unique solution exists but cannot exhibit it in closed form.
It is easiest to begin with the homogeneous equation Lu= 0 , that is, find a solution of
a¨u+b˙u+cu= 0 withu(t0) =α,and ˙u(t0) =β.
Without motivation, let us see what the substitution u(t) =eλtyields. Here λis a
constant. We must compute Leλt.
Leλt= (aλ2+bλ+c)eλt.
Canλbe chosen so that eλtis a solution of Lu= 0 ? Since eλt/negationslash= 0 for any t, this means,
is it possible to pick λso thataλ2+bλ+c= 0 ? Yes. In fact that “quadratic equation
formula” yields two roots
λ1=−b+√
b2−4ac
2a, λ 2=−b−√
b2−4ac
2a
of the characteristic polynomial p(λ) =aλ2+bλ+c. Notice that we have assumed a/negationslash= 0 .
Thus, two solutions of the homogeneous equation are
u1(t) =eλ1tandu2(t) =eλ2t.
Since the operator Lis linear, every linear combination of solutions is also a solution,
L(Au1+Bu2) =ALu 1+BLu 2= 0 . Therefore u(t) =Au1(t) +Bu2(t) is a solution of the
homogeneous equation Lu= 0 for any choice of the scalars AandB.
What about the initial conditions u(t0) =α,˙u(t0) =β; can they be satisfied by picking
the constants AandBsuitably? Let us try. We want to pick AandBso that
Aeλ1t0+Beλ2t0=α (u(t0) =α)
Aλ1eλ1t0+Bλ2eλ2t0=β( ˙u(t0) =β).
These equations can be solved as long as
0/negationslash=λ2e(λ1+λ2)t0−λ1e(λ2+λ1)t0= (λ2−λ1)e(λ1+λ2)t0,
which means λ1/negationslash=λ2orb2−4ac/negationslash= 0 . [The linear equations Ar1+Bs1=α, Ar 2+Bs2=β
can be solved for AandBifr1s2−r2s1/negationslash= 0 ]. Before dealing with the degenerate case
b2−4ac= 0 , let us consider an
4.2. A DIGRESSION TO CONSIDER AU/prime/prime+BU/prime+CU=F. 165
Example: Solve ¨u+ 3 ˙u+ 2u= 0 with the initial conditions u(0) = 1 and ˙ u(0) = 0 . If
we seek a solution of the form u(t) =eλt, the characteristic polynomial is λ2+ 3λ+ 2 = 0 ,
and has roots λ1=−1, λ2=−2 . Therefore u(t) =Ae−1+Be−2tis a solution. Since
λ1/negationslash=λ2, we can solve for AandBby using the initial conditions. We find
A+B= 1 (u(0) = 1),
−A−2B= 0 ( ˙u(0) = 0).
These two equations yield A= 1, B=−1 . Thus
u(t) = 2e−t−e−2t
is the unique solution of our initial value problem.
The degenerate case b2−4ac= 0 must be discussed separately. In this case λ1=λ2=
−b
2a, so the two solutions eλ1tandeλ2tare really the same solution. Without motivation
(but see Exercise 12) we claim that teλ1tis also a solution. This is easy to verify by a
calculation.
L(teλ1t) =a(tλ2
1eλ1t+ 2λ1eλ1t) +b(eλ1t+λ1teλ1t) +cteλt
= (aλ2
1+bλ1+c)teλt+ (2aλ1+b)eλ1t(4-17)
Since (aλ2
1+bλ1+c) = 0 by definition of λ1, andλ1=−b
2ain our special case, both terms
on the right vanish. Hence both u1(t) =eλ1tandu2(t) =teλ1tare solutions of Lu= 0
(ifb2−4ac= 0) , sou(t) =Aeλ1t+Bteλ1tis a solution for any choice of AandB. It is
possible to pick AandBto satisfy arbitrary initial conditions u(t0) =α,˙u(t0) =β.
Aeλ1t0+Bt0eλ1t0=α, (u(t0) =α)
Aλ1eλ1t0+B(1 +λ1t0)eλ1t0=β ( ˙u(t0) =β).
These can be solved for AandBsince
0/negationslash= (1 +λ1t0)e2λ1t0−λ1t0e2λ1t0=e2λ1t0.
Example: Solve ¨u+6 ˙u+9u= 0 with the initial conditions u(1) = 2 , ˙u(1) = −1 . Seeking
a solution in the form eλt, we are led to the characteristic equation λ2+ 6λ+ 9 = 0 , which
hasλ1=−3, λ2=−3 , as roots. Therefore u1(t) =e−3tis a solution of Lu= 0 . Since
λ1=λ2, another solution is u2(t) =te−3t. Thusu(t) =Ae−35+Bte−3tis a solution for
anyAandB. To solve for AandBin terms of the initial conditions, we must solve the
algebraic equations
Ae−3+B·1·e−3= 2, (u(1) = 2),
−3Ae−3+B(1−3)e−3=−1, ( ˙u(1) = −1).
We find that A=−3e3andB= 5e3. Thus
u(t) =−3e3e−3t+ 5e3te−3t,
or, equivalently,
u(t) =−3e−3(t−1)+ 5te−3(t−1).
Our results will now be collected as
166 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
Theorem 4.7 . The initial value problem
a¨u+b˙u+cu= 0, a/negationslash= 0,withu(t0) =α,˙u(t0) =β,
wherea,b, andcare constants, has a unique solution.
i) Ifb2−4ac/negationslash= 0, it is of the form
u(t) =Aeλ1t+Beλ2t.
ii) Ifb2−4ac= 0, soλ1=λ2, it is of the form
u(t) =Aeλ1t+Bteλ1t.
Hereλ1andλ2are the roots of the characteristic equation aλ2+bλ+c= 0 , and the
constantsAandBare determined from the initial conditions.
Remark: We have omitted the condition a>0 ,b≥0 ,c≥0 from our theorem since the
construction presented to find a solution did not depend on this. Uniqueness for that case
is treated as exercise 9, as we mentioned earlier.
Another
Example: Solve ¨u−2 ˙u+ 2u= 0 , with the initial conditions u(0) = 1,˙u(0) = 1 . The
characteristic polynomial is λ2−2λ+2 = 0 . Its roots are λ1= 1+i, andλ2= 1−i. Since
λ1/negationslash=λ2, the solution is of the form u(t) =Ae(1i)t+Be(1−i)t. From the initial conditions,
we find that
A+B= 1, (u(0) = 1),
(1 +i)A+ (1−i)B= 1, ( ˙u(0) = 1).
ThusA=1
2, B=1
2, so
u(t) =1
2e(1+i)t+1
2e(1−i)t.
Recalling that ex+iy=ex(cost+isiny) , this solution may be written in a more familiar
form:
u(t) =1
2et(cost+isint) +1
2et(cost−isint),
that is,
u(t) =etcost.
What has been done can be summarized elegantly in the language of linear spaces. We
have sought a solution of a second order linear O.D.E., which we write as Lu= 0 . It was
found that every solution of this equation could be expressed as a linear combination of
two specific solutions u1andu2, u(t) =Au1(t) +Bu2(t) , where the constants AandB
are uniquely determined from u(t0) and ˙u(t0) . Thus, the set of functions u which satisfy
Lu= 0 form a two dimensional subspace of D(L) =C2. The functions u1andu2span
that subspace. If we call the set of all solutions of Lu= 0 the nullspace of L,N(L) , then
our result simply reads “dim N(L) = 2 ”. A particular solution of Lu= 0 is found by
specifyingu(t0) and ˙u(t0) .
The inhomogeneous equation Lu=fis treated by finding a coset of the nullspace
ofL. For ifu0is a particular solution of the inhomogeneous equation Lu0=f, then
4.2. A DIGRESSION TO CONSIDER AU/prime/prime+BU/prime+CU=F. 167
u= ˜u+u0; where ˜u∈N(L) , is also a solution since Lu=L(˜u+u0) =L˜u+Lu0= 0+f=f.
Therefore, if one solution u0of the inhomogeneous equation Lu=fis found, the general
solution is u= ˜u+u0where ˜u∈N(L) . In particular, the solution ˜ u∈N(L) can be
chosen so that arbitrary initial conditions for u,u(t0) =α,˙u(t0) =β, can be met. We
shall defer (until our systematic treatment of linear O.D.E.’s) presenting a general method
for finding a solution u0of the inhomogeneous equation. In our example, the particular
solution will be found by guessing.
Example: SolveLu: = ¨u−u= 2t, with the initial conditions u(0) = −1,˙u(0) = 3 . The
homogeneous equation Lu= 0 has the general solution ˜ u(t) =Aet+Be−t. We observe
that the function u0(t) =−2tis a particular solution of the inhomogeneous equation,
Lu= 2t. Thusu(t) =Aet+Be−t. The initial conditions lead us to solve the following
equations for AandB,
A+B=−1 (4-18)
A−B−2 = 3. (4-19)
A computation gives A= 2 ,B=−3 . Thus the solution of our problem is
u(t) = 2et−3e−t−2t
It is routine to verify that this function u(t) does satisfy the O.D.E. and initial conditions
(you should verify the solutions to check for algebraic mistakes).
Exercises
(1) Solve the following homogeneous initial value problems,
(a). ¨u−u= 0, u (0) = 0,˙u(0) = 1 .
(b). ¨u+u= 0, u (0) = 1,˙u(0) = 0.
(c). ¨u−4 ˙u+ 5u= 0, u(0) = −1,˙u(0) = 2.
(d). ¨u+ 2 ˙u−8u= 0, u(2) = 3,˙u(2) = 0.
(e). ¨u= 0, u (0) = 7,˙u(0) = 3.
(2) Solve the following inhomogeneous initial value problems by guessing a particular
solution of the inhomogeneous equation. Check your answers.
(a) ¨u−u=t2, u (0) = 0,˙u(0) = 0
hint: Tryu0(t) =a1t2+a2t+a3and solve for a1,a2,a3.]
(b) ¨u−4 ˙u+5u= sint, u (0) = 1,˙u(0) = 0 [ hint: Tryu0(t) =a1sint+a2cost.]
(3) Consider an undamped harmonic oscillator with a sinusoidal forcing term, ¨ u+n2u=
sinγt. Find the general solution if γ2/negationslash=n2[tryu0(t) =a1sinγt+a2cosγtfor a
particular solution]. What happens if γ→+−n? This is called resonance .
(4) You shall discuss damping in this problem. Consider the equation ¨ u+ 2µ˙u+ku= 0 ,
whereµ>0 , andk>0 . We shall let γ=/radicalbig
|µ2−k|.
168 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
(a)Light damping (µ2<k) . Show that the solution is
u(t) =e−µt(Acosγt+Bsinγt),
and sketch a rough graph for the case A= 1, B= 0 . This is the kind of
oscillation you want for a pendulum clock, with µsmall.
(b)Heavy damping (µ2>k) . Show that the solution is
u(t) =e−µt(Aeγt+Be−γt).
Show that u(t) vanishes at most once. Sketch a graph for the two cases A=
B= 1 andA=−1,B= 3 . The first describes the oscillation of an ideal screen
door, while the second describes the ideal oscillation of a slammed car door.
(5) It is often useful to study the oscillations described by ¨ u+ 2µ˙u+ku= 0 by sketching
the solution in the u,˙uplane - or phase space as it is called. Investigate the curves
for heavily and lightly damped oscillators. Show that the curve for a heavily damped
oscillator will be a straight line through the origin for special initial conditions. What
does the phase space curve look like for an undamped oscillator ( µ= 0, k> 0) ?
(6) Consider the linear operator Lu=a¨u+b˙u+cu, wherea,b,c are constants. We have
seen thatLert=p(r)ertwherep(r) is the characteristic polynomial.
(a) Ifrisnotone of the roots of the characteristic polynomial, observe that you
can find a particular solution of Lu=ert. What is it?
(b) If neither r1norr2is a root of the characteristic polynomial, find a particular
solution of Lu=a1er1t+a2er2t, wherea1anda2are specified constants.
(c) Use this procedure to find a particular solution of
i)¨u−4u= cosht, ii )¨u+ 4u= sint
(7) (a) Imitate our procedure and develop a theory for the first order homogeneous
O.D.E.Lu: = ˙u+bu= 0 , where bis a constant. In particular, you should prove
that there exists aunique solution satisfying the initial condition u(t0) =α, and
give a recipe for finding it. Use your recipe to solve ˙ u+ 2u= 0, u(0) = 3 .
(b) And now you will show us how to find a particular solution of the inhomogeneous
equationLu=f, wheref(t) is some given continuous function and Lu: =
˙u+bu. [hint: Try to find a function µ(t) such that µ( ˙u+bu) =d
dt(µu) . Then
integrated
dt(µu) =µf, and solve for u]. Use your method to find a particular
solution for ˙ u+2u=x, and then a solution of the same equation which satisfies
the initial condition u(0) = 1 .
(8) Find a solution of u/prime/prime/prime−2u/prime/prime−u/prime+ 2u= 0 which satisfies the initial conditions
u(0) =u/prime(0) = 0, u/prime/prime(0) = 1 . [ hint: The cubic equation γ3−2γ2−γ+ 2 has roots
+1,−1 and 2].
(9) You will prove the uniqueness theorem for the equation ¨ u+b˙u+cu= 0 , where band
care any constants (we have let a= 1 , because if it is not 1, just divide the whole
equation by a). The trick is to reduce this to the special case b≥0, c≥0 , already
done.
4.2. A DIGRESSION TO CONSIDER AU/prime/prime+BU/prime+CU=F. 169
(a) Show that in order to prove the solution of
¨u+b˙u+cu=f,whereu(t0) =α,˙u(t0) =β
is unique, it is sufficient to prove that the only solution of
¨w+b˙w+cw= 0, w(t0) = 0,˙w(t0) = 0
isw(t)≡0 .
(b) Define ϕ(t) byw(t) =eγtϕ(t) . Observe: to prove w= 0 , it is sufficient to
proveϕ≡0 ( hereγis any constant). Use the differential equation and initial
conditions for wto find the differential equation and initial conditions for ϕ.
Show that γcan be picked so that the D.E. for ϕis
¨ϕ+˜b˙ϕ+˜0ϕ= 0,
where ˜band ˜care positive. Deduce that ϕ≡0 , and from that, that w≡0 ,
completing the proof.
(10) A boundary value problem for the equation
u/prime/prime+bu/prime+cu= 0
is to find a solution of the equation with given boundary values, say u(0) =αand
u(1) =β. Assumebandcare real numbers.
(a) Show that a solution of the boundary value problem always exists if b2−4c≥0
(the caseb2−4c= 0 will have to be done separately).
(b) Prove that if b2−4c≥0 , the solution is unique too. [I suggest letting u(t) =
eγtv(t) , and then choosing γso that the equation satisfied by vis of the form
v/prime/prime+ ˜cv= 0 , where ˜ c≤0 . The case ˜ c= 0 is trivial. If˜(c)<0 , can the solution
have a positive maximum or negative minimum?]
(11) If a spring is hung vertically and a mass mplaced at its end, an external force of
magnitude mgdue to gravity is placed on the system. Assume there are no dissipative
forces of any kind.
(a) Set up the differential equation of motion. Remember that you must specify
which is the positive direction.
(b) If the tip of the spring is displaced a distance dby placing the mass on it (no
motion yet), so the equilibrium position isdbelow the unstretched end of the
spring, show that the spring constant kis given by k=mg/d .
(c) Let the body weigh 32 pounds, and dbe 2 feet. Find the subsequent motion if
the body is initially displaced from rest one foot below its equilibrium position.
[Take |g|= 32 ft/sec2].
(12) * Consider au/prime/prime+bu/prime+cu= 0 . Ifγ1/negationslash=γ2, are the roots of the characteristic equation,
observe that the function
˜u(t) =eγ1t−eγ2t
γ1−γ2
170 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
is also a solution (it is a linear combination of eγ1tandeγ2t). Now pass to the
limitγ2→γ1(leaveγ1fixed and let γ2move) by using the Taylor series for eγt.
The function you get is then a “guess” for a second solution in the degenerate case
γ1=γ2. This supplies some motivation for the guess made earlier.
(13) * Consider Lu: =u/prime/prime+ 2u=f, wherefis given. You know how to solve
Lu=Asinnx(Exercise 6). Find a particular solution to the general inhomoge-
neous equation in the interval [ −π,π] by expanding fin a Fourier series and then
use superposition. Apply this to solve u/prime/prime+ 2u=x.
(14) Consider an undamped harmonic oscillator, whose motion is specified by u(t) , where
mu/prime/prime+ku= 0,k> 0 . Show that the solution u(t) =A1cos/radicalbigg
k
mt+B1sin/radicalbigg
k
mtmay
be written in the form
u(t) =Asin(wt+θ),
whereAis the amplitude of the oscillation, w= 2πv, v is the frequency , andθis
thephase . Show that u(t) is periodic, u(t+T) =u(t) , where the periodT= 1/v.
Interpret the amplitude and phase and determine A,w, andθin terms of A1, B1, k
andm. [I suggest looking at a specific example and its graph first].
4.3 Generalities on LX =Y.
Undoubtedly the fundamental problem in the theory of linear (and nonlinear) operators is
to determine the nature of the range of an operator L. One particular aspect of this is the
vast problem of solving the equation
LX=Y
forXwhenYis given to you. The question here is, “is a given Yin the range of L?”,
or “can we find some Xsuch thatLX=Y?” If one can solve the problem uniquely for
anyY, then the solution is written as
X=L−1(Y),
whereL−1is the operator inverse to L, in the sense that L−1L=I(so to solve LX=Y,
applyL−1, X=L−1LX=L−1Y).
Let us give some examples, familiar and unfamiliar, of problems of the form LX=Y,
whereYis given.
1.LX= (2x1+ 3x2,x1+ 2x2), X ∈R2,
L:R2→R2.
The problem of solving LX=YwhereY= (−1,2)∈R2is that of solving the two
equations
2x1+ 3x2=−1
x1+ 2x2= 2
for two unknowns ( x1,x2) =X.
4.3. GENERALITIES ON LX=Y. 171
2.Lu=u/prime/prime+ 2u/prime+ 3u,whereu∈C2,
L:C2→C.
The problem of solving L(u) =xis that of solving the inhomogeneous ordinary differential
equation
Lu: =u/prime/prime+ 2u/prime+ 3u=x
foru(x) .
3.Lu=/integraldisplayπ
0cos(x−t)u(t)dt, u ∈C[0,π].
You should check that Lis a linear operator. The problem of solving L(u) = sinxis that
of solving the integral equation
Lu: =/integraldisplayπ
0cos(x−t)u(t)dt= sinx
for the function u. In this example, it is instructive to examine the range more closely.
Since cos(x−t) = cosxcost+ sinxsintand since functions of xare constant with respect
totintegration, we see that Lumay be written as
Lu: = cosx/integraldisplayπ
0costu(t)dt+ sinx/integraldisplayπ
0sintu(t)dt,
or
Lu: =α1cosx+α2sinx,
where the numbersα1andα2are
α1=/integraldisplayπ
0(cost)u(t)dt;α2=/integraldisplayπ
0(sint)u(t)dt.
Thus, the range of Lis the linear space spanned by cos xand sinx, which has dimension
two. This linear operator Ltherefore maps the infinite dimensional space C[0,π] into a
finite (two) dimensional space. In order to even have a chance of solving Lu=ffor this
operatorL, we first check to see if feven lies in this two dimensional subspace (for if it
doesn’t, it is futile to go further). The particular function sin xdoes, so it is reasonable
to look for a solution - which we shall not do right now (however there are infinitely many
solutions, among them u(x) =2
πsinx).
One particularly important equation which arises frequently is the homogeneous equa-
tion
LX= 0,
which is the special case Y= 0 of the inhomogeneous equation ,
LX=Y.
SinceLis a linear operator, there is no problem of our finding onesolution of LX= 0
forX= 0 is a solution, the so-called trivial solution of the homogeneous equation. The
problem is to find a non-trivial solution, or better yet, all solutions. In the previous section,
this question was answered fully for the particular operator Lu=au/prime/prime+bu/prime+cu, where
a, b, andcare constants. Many of our results there generalize immediately, as we shall
see now.
172 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
Definition: The set of all solutions of the homogeneous equation LX= 0 where Lis a
linear operator is called the nullspace ofL. This nullspace of L,N(L) , consists of all X
in the domain of Lwhich are mapped into zero by L,
L:N(L)→0.N(L)⊂D(L).
We have called N(L) the null space ofL, not the null setbecause of
Theorem 4.8 . The nullspace of a linear operator L:V1→V2is a linear space, a subspace
of the domain of L.
Proof: Since the domain of L,D(L) =V1, is a linear space and N(L)⊂D(L) , by
Theorem 2, p.142 all we need show is that the set N(L) is closed under multiplication by
scalars and under addition of vectors. Say X1andX2∈N(L) . ThenLX 1= 0 and
LX 2= 0 . We must show that L(aX1) = 0 for any scalar a, and that L(X1+X2) = 0 .
ButL(aX1) =aL(X1) =a·0 = 0 , and L(X1+X2) =LX 1+LX 2= 0 + 0 = 0 . Thus
N(L) is a subspace of D(L) =V1.
One important reason for examining the null space of a linear operator is because
ifN(L) is known, and if any onesolution of the inhomogeneous equation is known, say
LX 1=Y(whereYwas given and X1is the solution we know), then every solution of the
inhomogeneous equation is of the form ˜X+X1, where ˜X∈N(L) . In other words every
solution of LX=Yis inN(L) +X1, theX1coset of the subspace N(L) .
Theorem 4.9 . LetL:V1→V2be a linear operator. If X1andX2are any two solutions
of the inhomogeneous equation LX=Y, whereYis given, then X2−X1∈N(L), or
X2=˜X+X1where ˜X∈N(L).
Proof: Let ˜X=X2−X1. We shall show that ˜X∈N(L) .
L˜X=L(X2−X1) =LX 2−LX 1=Y−Y= 0.
By using this theorem, we see that if allsolutions of the homogeneous equation LX=
0 are known - the nullspace of L—and if onesolution of the inhomogeneous equation
LX 1=Yis known, then allof the solutions of the inhomogeneous equation are known.
This solution set of the inhomogeneous equation is the X1coset of N(L) .
Example: 1 LetL:R2→R2be defined by
LX= (x1+x2,x1−x2)∈R2
Then N(L) is the set of all points in R2such thatLX= 0 , that is, which satisfy the
equations
x1+x2= 0
x1−x2= 0
Thus the nullspace of Lconsists of the intersection of the two lines x1+x2= 0, x1−x2= 0 .
The only point on both lines is 0. Thus N(L) is just the point 0. To solve the inhomogeneous
equationLX=Y, whereY= (1,1) .
x1+x2= 1, x1−x2= 1,
4.3. GENERALITIES ON LX=Y. 173
we find one solution of it, X1= (1,0) . Then every solution of the inhomogeneous equation
is of the form X=˜X+X1, where ˜X∈N(L) . But since ˜X+0 is the only point in N(L) ,
every solution is of the form X= 0 +X1=X1. Thus every solution is exactly X1, which
is the unique solution of LX=Y. This situation is a general one. Again, we also saw this
forLu=au/prime/prime+bu/prime+cu.
Theorem 4.10 . If the nullspace N(L)of the linear operator consists only of 0, then the
solution of the inhomogeneous equation LX=Y(if a solution exists) is unique. (Thus, if
the nullspace contains only 0, then Lis injective).
Proof: Say there were two solution X1andX2. ThenLX 1=YandLX 2=Y, which
impliesL(X2−X1) =LX 2−LX 1=Y−Y= 0 . Therefore ( X2−X1)∈N(L) . Since the
only element of N(L) is 0, X 2−X1= 0 , or,X1=X2. In other words, the two solutions
are the same.
Example: 2 LetL:C2→Cbe defined on functions u∈C2by
Lu: =a(x)u/prime/prime+b(x)u/prime+c(x)u.
The nullspace of Lconsists of all solutions of the homogeneous equation Lu= 0 . It
turns out (see chapter 6) - as in the constant coefficient case - that every solution of this
homogeneous O.D.E. has the form u=Au1+Bu2, whereu1andu2are any two linearly
independent solutions of the equation, and where AandBare constants. Thus N(L)
is a two dimensional space spanned by u1andu2. Ifu1is a particular solution of the
inhomogeneous equation Lu1=f, then allthe solutions of Lu=fare just the elements
of theu1coset of N(L) , that is, functions of the form u= ˜u+u1, where ˜u∈N(L) .
With every linear operator L:V1→V2, V1=D(L) , we have associated two other
linear spaces, the nullspace N(L)⊂D(L) =V1and range R(L)⊂V2. There is a valuable
and elegant way to connect D(L),N(L) and R(L) . The result we are aiming at is certainly
the most important theorem of this section.
We know that R(L)⊂V2. The space V2may be of arbitrarily high dimension.
However, since R(L) is the image of D(L) , we suspect that R(L) can take up “no more
room” then D(L) . To be more precise,
dimR(L)≤dimD(L).
Thus, for example, if L:R2→R17, we expect that the range of Lis a subspace of
dimension no more than two in R17. Not only is this a justifiable expectation, but even
more is true.
If Dim R(L) = Dim D(L) , essentially all of D(L) is carried over under the mapping.
But if Dim R(L)<DimD(L) , what has happened to the remainder of D(L) ? Let us
look at N(L)⊂D(L) . The elements of N(L) are all squashed into the zero element of
V2. In other words, a set of dim N(L) inV1=D(L) is mapped into a set of dimension
zero inV2. DoesLdecompose D(L) +V1into two parts, N(L) and a complement N(L)/prime
such thatLmaps N(L) into zero and the dimension of the remainder, N(L)/prime, is preserved
underL(so dim N(L)/prime= dim R(L)) . Of course,
a figure goes here
174 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
Theorem 4.11 . Let the linear operator LmapV1=D(L)intoV2. IfD(L)has finite
dimension, then
dimD(L) = dim R(L) + dim N(L).
Proof: LetN(L)/primebe a complement of N(L) (cf. pp. 163a-d). Since dim N(L) +
dimN(L)/prime= dim D(L) , it is sufficient to prove that dim N(L)/prime= dim R(L) .
ForX∈V1, we can write X=X1+X2, whereX1∈N(L) andX2∈N(L)/prime. Now
LX=LX 1+LX 2, so the image of D(L) is the same as the image of N(L)/prime. In addition, if
X2∈N(L)/prime, thenLX 2= 0 if and only if X2= 0 , merely because N(L)/primeis a complement
of the nullspace. Let {θ1,...,θ k}be a basis for N(L/prime) . IfX2∈N(L)/prime, we can write X2=
k/summationdisplay
j=1ajθj, andLX 2=k/summationdisplay
j=1ajLθj. LetLθ1=Y1, Lθ 2=Y2,...,Lθ k=Yk. Since the image
ofN(L)/primeisR(L) , the vectors Y1,...,Y kspanR(L) . Thus, dim R(L)≤k= dim N(L)/prime.
To show that there is equality, dim R(L) = dim N(L)/prime, we prove that Y1,...,Y kare
linearly independent. If c1Y1+···+ckYk= 0 , then 0 = c1Lθ1+···+ckLθk=L(c1θ1+
···+ckθk) =L˜Xwhere ˜x=c1θ1+···+ckθk∈N(L)/prime. However for any ˜X∈N(L)/prime,
we knowL˜x= 0 implies that ˜X= 0 . The linear independence of θ1,...,θ kfurther
shows that c1=c2=···=ck= 0 . The hypothesis c1Y1+···+ckYk= 0 has led us to
conclude that the cj’s are all zero, that is, the Yj’s are linearly independent. Therefore
dimR(L) = dim N(L)/prime. Coupled with our first relationship, this proves the result.
Corollary 4.12 :dimR(L)≤dimD(L).
Proof: dimN(L)≥0 .
Two examples, one an illustration, the other an application.
Example: 1 Consider a projection operator, PA, mapping vectors from Eninto a subspace
AofEn, where the dim A=m < n . Let us first show that PAis a linear operator. If
e1,...e mis an orthonormal basis for A, then for any XandYinEn,
PA(X+Y) =m/summationdisplay
k=1/angbracketleftX+Y, e k/angbracketrightek=m/summationdisplay
k=1(/angbracketleftX, e k/angbracketright+/angbracketleftY, e k/angbracketright)ek
=m/summationdisplay
k=1/angbracketleftX, e k/angbracketrightek+m/summationdisplay
k=1/angbracketleftY, e k/angbracketrightek=PAX+PAY.
Similarly,PA(aX) =aPAXfor every scalar a. Thus the projection operator is a linear
operator. Since R(PA) =Aand dimA=m, while dim En=n, we conclude that
dimN(PA) =n−m. This could have been arrived at immediately since PAwill certainly
map everything perpendicular to A, that isA⊥, into 0 (see fig. illustrating the case
E2→A, whereAis a line). Thus N(PA) +A⊥, so dim N(PA) = dimA⊥=n−m.
Example: 2 DefineL:Rn→Rkby,
LX= (a11x1+a12x2+···+a1nxn, a21x1+···+a2nxn,···,aklx1+ak2x2+···+aknxn)
whereX= (x1,x2,...x n)∈Rn. If we let Y= (y1,...,y k)∈Rk, then writing Y=LX,
the linear operator Lmay be defined by the kequations (for y1,...,y k) inn“unknowns”
4.3. GENERALITIES ON LX=Y. 175
(x1,...,x n) ,
a11x1+a12x2+···+a1nxn=y1
a21x1+a22x2+···+a2nxn=y2
...
ak1x1+ak2x2+···+aknxn=yk.
The problem of solving LX=Y, whereYis given, is that of solving kequations with n
“unknowns”.
Consider the special case k <n , when there are less equations than unknowns. Since
the range of Lis contained in Rk,R(L)⊂Rk, then dim R(L)≤dimRk=k. Because
D(L) =Rn, we also know that dim D(L) = dim Rn=n. Thus
dimN(L) = dim D(L)−dimR(L)≥n−k>0.
However ifdimN(L)>0 , then N(L) must contain something other than zero. Thus there
is at least one non-trivial solution ˜Xof the homogeneous equation ,L˜X= 0 . Since a˜Xis
also a solution, where ais any scalar, there are, in fact an infinite number of solutions .
Notice that the above was a non-constructive existence theorem. We proved that a
solution does exist but never gave a recipe to obtain it. One consequence of this result is
that, if dim N(L)>0 , and if a solution of the inhomogeneous equation LX=Yexists, it
is not unique ; for ifLX 1=Y, then also L(X1+˜X) =Y, where ˜Xis any solution of the
homogeneous equation.
In the special case n=k, and dimN(L) = 0 a fascinating (and non-constructive)
theorem falls out of Theorem 11: the inhomogeneous equation LX=Yalways has a
solution and the solution is unique . Put in more conventional terms, if there are the same
number of equations as unknowns, and if the only solution of the homogeneous equation
is zero, then the inhomogeneous equation always has a unique solution. Thus, if n=k,
uniqueness implies existence .
Since dim N(L) = 0 , then dim R(L) = dim D(L) =n. However L:Rn→Rnin this
case (n=k) . Since R(L)⊂Rnand dim R(L) =n, we see that R(L) must be all of Rn,
that is, every Y∈Rnis in the range of L, which means that the inhomogeneous equation
LX=Yis solvable for every Y∈Rn. Theorem 10 gives the uniqueness. We shall obtain
a better theorem later.
Remark: Some people refer to dim R(L) as the rank of the linear operator L. We shall,
however, refer to it as the dimension of the range of L.
IfL1:V1→V2andL2:V2→V3, it is easy to make a few statements about
dimR(L2L1) .
Theorem 4.13 . IfL1:V1→V2andL2:V3→V4, whenV2⊂V3, (soL2L1) is
defined), then
dimR(L2L1)≤min(dim R(L1),dimR(L2)).
Proof: The last corollary states that an operator is like a funnel with respect to dimension:
the dimension can only get smaller or remain the same. After passing through two funnels,
we obtain no more than the smallest allowed through. One might think that there should be
equality in the formula. That this is not the case can be seen from the possibility illustrated
in the figure. Only the shaded stuff gets through.
176 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
Exercises
(1) LetL:Rn→Rnbe defined by
LX= (x1,x2,...,x k,0,..., 0),
whereX= (x1,x2,...x n)∈Rn. Describe R(L) and N(L) . Compute dim R(L)
and dim N(L) .
(2) (a) Describe the range and nullspace of the linear operator L:R3→R3defined by
LX= (x1+x2−x3,2x1−x2+x3, x2−x3), X= (x1,x2,x3)∈R3.
(b) Compute dim R(L) and dim N(L) .
(c) Is (1,2,0)∈R(L) ? Is (1,2,1)∈N(L) ?
Is (1,2,2)∈N(L) ? Is (0,−1,−1)∈N(L) ?
(3) LetA={u∈C2[0,2]:u(0) =u(1) = 0 }, and define L:A→C[0,1] byLu=
u/prime/prime+b(x)u/prime−u, whereb(x) is some continuous function. Prove N(L) = 0 . [ hint: If
u∈N(L) , canuhave a positive maximum or negative minimum?]
(4) Consider the linear operator L:C[0,1]→C[0,1] defined by
(Lu)(x) =u(x) + 2/integraldisplay1
0ex−tu(t)dt
(a) Find the nullspace of L.
(b) SolveLu= 3ex. Is the solution unique?
(c) Show that the unique solution of Lu=f, wheref∈C[0,1] is
u(x) =f(x)−cex,wherec=2
3/integraldisplay1
0e−tf(t)dt.
(5) LetL:V→V(soLkis defined for k= 0,1,2...). Prove that
(a)R(L)⊂N(L) if and only if L2= 0 .
(b)N(L)⊂N(L2)⊂N(L3)⊂...
(c)N(L)/prime⊃N(L2)/prime⊃N(L3)/prime⊃....
(6) IfL1:V1→V2andL2:V3→V4whereV2⊂V3, Theorem 12 gives an upper bound
for dim R(L2L1) .
(a) Prove the corresponding lower bound
dimR(L2L1)≥dimR(L1) + dim R(L2)−dimV3.
[hint: Prove the equivalent inequality dim R(L1)≤dimR(L2L1) + dim N(L2)
by letting ˜V=R(L1) and applying Theorem 11 to L2defined on ˜V].
(b) Prove: if dim N(L2) = 0 , then
dimR(L2L1) = dim R(L1).
4.4.L:R1→RN. PARAMETRIZED STRAIGHT LINES. 177
(c) If dim N(L1) = 0 , is it then true that dim R(L2L1) = dim R(L2) ? Proof or
counterexample.
(d) If dimV1= dimV2= dimV3and dim N(L1) = 0 , is it true that dim R(L2L1) =
dimR(L2) ? Proof or counterexample.
(7) IfL1andL2both mapV1→V2, prove
|dimR(L1)−dimR(L2)| ≤dimR(L1+L2).
(8) Consider the operator L:C2[0,∞)→C[0,∞) defined by
Lu: =u/prime/prime+ 3u/prime+ 2u.
(a) Describe N(L) . What is dim N(L) ? Isf(x) = sinx∈R(L) ?
(b) Consider the same operator Lbut mapping AintoC[0,∞] , whereA={u∈
C2[0,∞):u(0) +u/prime(0) = 0 }. Answer the same questions as part a).
(c) Same as bbutA={u∈C2[0,∞):u(1) +u/prime(1) = 0 }this time.
4.4 L:R1→Rn. Parametrized Straight Lines.
Our study of particular linear operators begins with the most simple case: a linear operator
which maps a one- dimensional space R1into anndimensional space Rn. Since the
dimension of the range of Lis no greater than that of the domain R1and dim R1= 1 ,
then
dimR(L)≤1.
This proves
Remark: IfL:R1→Rn, then the dimension of the range of Lis either one or zero.
The case dim R(L) = 0 is trivial, for then Lmust map all of R1into a single point, and
that single point must be the origin since the range of Lis a subspace. Thus, ifdimR(L) =
0, thenLmaps every point into 0 . Without change, the same holds if L:V1→V2(where
V1andV2are any linear spaces) and dim R(L) = 0 . Not very profound.
If dim R(L) = 1 , then the subspace R(L) inRnis a one dimensional subspace in the
ndimensional space Rn, this is, R(L) is a “straight line” through the origin of Rn. This
straight line is determined if any one point P/negationslash= 0 on it is known. Then there is a point
X1∈R1such thatLX 1=P. Since R1is one dimensional it is spanned by any element
other than zero, so every X∈R1can be written as X=sX1. Therefore, if Xis any
element of R1,
LX=L(sX1) =sLX 1=tP.
In other words, this last equation states that the range of Lis a multiple of a particular
vectorP, that is, a straight line through the origin.
Example:
IfL:R1→R2such that the point X1= 2∈R1is mapped into the point P=
(1,−2)∈R2, then
L:X=s2→(s,−2s),
In particular, the point X= 3(s=3
2) is mapped into the point (3
2,−3) .
178 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
a figure goes here
In applications, the domain R1usually represents time, while the range represents the
position of a particle. Then L:R1→Rnis an operator which specifies the position of a
particle at a given time. Since Lis linear and L0 = 0 , the path of the particle must be
a straight line which passes through the origin at t= 0 . Later on in this section we shall
show how to treat the situation of a straight line not through the origin, while in Chapter
7 we shall examine curved paths (non-linear operators).
Example: This is the same example as before. L:R1→R2is such that at time t= 2∈R1
a particle is at the point (1 ,−2) (while at t= 0 it is at the origin). At any time t=s2 ,
the particle is at ( s,−2s) . In particular, at t= 3(s=3
2) , the particle is at (3
2,−3) . It is
also convenient to rewrite the position ( s,−2s) directly in terms of the time. Since t= 2s,
the position at time tis (t
2,−t) . Thus we can write
L:t→(t
2,−t),
which clearly indicates the position at a given time. If a point in the space R2is specified
byY= (y1,y2)∈R2, then the operator can be written as
y1=1
2t
y2=−t.
All of these are useful ways to write the operator L. In some situations, one might be more
useful than another.
This brings us to an issue which perhaps seems a bit pedantic but can serve you well
in times of need. How can we represent the operator in a picture? There are three distinct
ways. Some clarity can be gained by distinguishing them carefully. The same ideas carry
over immediately to nonlinear operators.
Our first picture has two parts. If L:R1→Rn, then the first part is a diagram of R1,
the second part a diagram of Rn, and between them are arrows to indicate the image of
each point in R1. The picture below the first example was of this type. All of the arrows
get in the way, so a more convenient picture is needed. That comes next.
The second picture is the graph of an operator L. The graphL:R1→Rnis the set
of points (X,LX ) in the Cartesian product space R1×Rn. Thus, ifV1is time, and Rn
space with Lassigning a position to every time, then the points on the graph are points
in time - space ( X,LX ) . For the previous example, these are the points ( t,(t
2,−t)) in
R1×R2, a straight line in time-space ( or space-time if you prefer). To each time, there
is a unique point in space. In a sense, this second picture, the graph, associated with an
operator results from gluing together the two pieces of the first picture. By using the graph
of an operator, we avoid the arrow mess of the first picture.
The third picture just indicates the range of an operator (when thinking of pictures,
the range is often referred to as the path of the operator). In terms of the time- position
example, this picture only shows the path of a particle in space and ignores when a particle
had a given position. Thus, this picture is the second half of the first picture. From our
physical interpretation, it is clear that two different operators might have the same path (for
two particles could travel the same path without having the same position at every time).
Thus, this picture is an incomplete representation of an operator.
4.4.L:R1→RN. PARAMETRIZED STRAIGHT LINES. 179
Example: If˜L:R1→R2such that the point X1= 1∈R1is mapped into the point
P= (1,−2)∈R2, then
˜L:X=s·1→(s,−2s).
In particular, the point X= 3(s= 3) is mapped into the point (3 ,−6) . The graph of˜L
is the set of points ( s,s,−2s) , which is a straight line in R1×R2. Compare this with the
operatorLconsidered previously (we remind you that L:X= 2s→(s,−2s)) . The graph
ofLwas the set of point (2 s,s,−2s) . These two sets of points the graphs of ˜LandL,
respectively, do not coincide since the operators are the same. On the other hand, the path
of˜Lis the set of points ( s,−2s) , which is exactly the same set of points as the path of
L. Shortly, we shall ask the question, how can we describe a straight line in Rn? One way
is to find an operator whose path is that straight line. Since many operators have the same
path, there will be many possible ways to describe the straight line. All we need do is pick
one, any one will do.
LetL:R1→RnandY0be some fixed point in Rn. Consider the operator MX :=
LX+Y0. SinceM0 =L0 +Y0=Y0/negationslash= 0 , we see that Mis not a linear operator; it is
called an affine operator oraffine mapping . The range of Mis the subspace translated by
the vector Y0, a straight line which does not necessarily pass through the origin ( it will if
and only if Y0∈R(L) ). In other words, R(M) is theY0coset of the subspace R(L) .
Example: TakeLto be the same as before, so L:X= 2s→(s,−2s) orL(2s) = (s,−2s) .
LetY0= (−3,2) . ThenMX :=LX+Y0= (s,−2s) + (−3,2) = (s−3,−2s+ 2) , where
X= 2s. In particular, Mmaps the point X= 3∈V1(s=e
2) into ( −3
2,−1)∈R2. The
figure shows the path of LandM. SinceX= 2s, we can eliminate sfrom the above
formula and write
MX = (1
2X−3,−X+ 2), X ∈R1.
If we denote by Y= (y1,y2) a general point in R2, thenMmay be written in the standard
form
y−1 =1
2X−3
y−2 =−X+ 2.
Of course, one could eliminate Xfrom these too and be left with 2 y1+y2=−4 , which is
the equation of the path and could come from any mapping with the same path.
It is instructive to investigate the reverse question, given two points PandQinRn,
find an equation for the straight line passing through them. Any mapping whose path is
the desired line will do. We have learned that MX =LX+Y0is the general equation
of a straight line through Y0. There is complete freedom in specifying which points are
mapped into PandQ, so we would be foolish not to pick the most simple case. Let
M: 0→PandM: 1→Q. ThenP=M(0) =L(0) +Y0=Y0, soY0=P, and
Q=M(1) =L(1) +Y0=L(1) +P, soL: 1→P−Q. This completely determines M
(sinceLis determined once the image of one point is known, L: 1→P−Q, and the
vectorY0is also determined, Y0=P).
Example: Find an equation for the straight line passing through the two points P=
(1,2,−3,−4) ,Q= (−1,3,2,−2) in R4. SayPis the image of 0 and Qis the image of
1, soM: 0→PandM: 1→Q. Then since MX =LX+Y0⇒P=M(0) =Y0so
Y0= (1,2,−3,−4) . AlsoQ=L(1)+Y0⇒L(1) =Q−Y0=Q−P= (−2,1,5,2) . Because
180 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
everyX∈R1can be written as X=s·1⇒LX=L(s·1) =sL(1) =s(−2,1,5,2) , or
LX= (−2s,s,5s,2s) , whereX=s·1∈R1. ThusMX =LX+Y0= (−2s,s,5s,2s) +
(1,2,−3,−4) , or
MX = (−2s+ 1,s+ 2,5s−3,2s−4),whereX=s· ∈R1.
If we useY= (y1,y2,y3,y4) to indicate a general point in R4, thenM:R1→R4can be
written as four equations.
y1=−2s+ 1
y2=s+ 2
y3= 5s−3
y4= 2s−4
whereX=s·1∈R1. For example, the image of X= 2(s= 2) in R1is the point
Y= (−3,4,7,4)∈R4.
The discussion before the example contained most of the proof of Theorem 13 . LetP
andQbe two points in Rn. Then the affine mapping
MX =P+s(Q−P),
has as its path the straight line passing through PandQ.
Remark: 1 The affine mapping ˜MX =P+ks(Q−P) , wherek/negationslash= 0 is some constant,
has the same path too. The only change is that while M: 0→PandM: 1→Qthis
mapping ˜M: 0→Pand ˜M:ks→Q. In other words for ˜Mwe have chosen to take ks
(nots) as the pre-image of Q. This pre-image of Qwas entirely arbitrary anyway.
Remark: 2 The equation MX =P+s(Q−P) ofM:R1→Rn, whereX=s·1∈R1
is called a parametric equation of the straight line which passes through PandQin
Rn, andsis called the parameter . Other parametric equations of the same line arise if
X=ks·1∈R1(cf. Remark 1), where kis some non-zero constant.
In order to introduce the slope of a straight line, let us paraphrase the last few para-
graphs in terms of particle motion. If PandQare two points in Rn, thenMt=
P+t(Q−P)M:R1→Rn, wheret∈R1describes the position of the particle at time
t. Att= 0 the particle is at P, while att= 1 the particle is at Q. Another particle
movingktimes as fast has the position ˜Mt=P+kt(Q−P) . This other particle is also
atPwhent= 0 , but takes time t=1
kto reach the point Q. It still has the same path
as the first particle. If we denote by Y= (y1,y2,...,y n) an arbitrary point in Rn, then
the position Yat timetis
y1=p1+kt(q1−p1)
y2=p2+kt(q2−p2)...
yn=pn+kt(qn−pn).
Now consider the mapping Mt=P+kt(Q−P) . The derivative att=t1is
dM
dt/vextendsingle/vextendsingle
t=t1= lim
t2→t1M(t2)−M(t1)
t2−t1
4.4.L:R1→RN. PARAMETRIZED STRAIGHT LINES. 181
It represents the velocity at t=t1. To have this make sense, we must introduce a norm
inRnso that the limit can be defined. Use the Euclidean norm (although any other one
could be used, for it turns out that there is no need for a limit in the case of a straight line).
SinceM(t2)−M(t1) =P+kt2(Q−P)−[P+kt1(Q−P)] =k(t2−t1)(Q−P) , we have
M(t2)−M(t1)
t2−t1=k(Q−P),
so
dM
dt(t) =k(Q−P).
Because this is independent of t, it is the derivative at anytimet. Thus, the derivative is
a vector,k(Q−P) . The derivative represents the velocity of a particle moving on the line.
The speed is the length of the velocity vector, speed = /bardblk(Q−P)/bardbl. What is the slope of
the line? Since the line is the path of a mapping, it should not depend on which mapping
is used. In terms of mechanics, the slope should not depend on the speed of the particle
moving along the line, but only that it moved along the straight line, that is its velocity
vector was along the line. Thus we define the slope as a unit vector in the direction of the
velocity. In our case, slope = Q−P//bardblQ−P/bardbl. This is a unit vector from PtoQand
only depends upon the mapping to specify a positive direction (orientation) for the line.
Example: A particle moves on a straight line from P= (1,−2,1) att= 0 toQ=
(3,1,−5) att= 2 . Find the position of the particle as a function of time, the velocity and
speed of the particle, and slope of the path.
The equation of the path is Mt=P+kt(Q−P) , wherekis determined from
Q=M(2) =P+ 2k(Q−P) , sok=1
2. ThusM(t) = (1,−2,1) +1
2t(2,3,−6) =
(1 +t,2 +3
2t,1−3t) . Velocity =1
2(Q−P) = (1,3
2,−3) . Speed = /bardblvelocity /bardbl=7
2. Slope
= velocity /speed = (2
7,3
7,−6
7) .
A glance at the formulas which precede the example reveals that the position of a
particle which moves along a straight line through Pcan be written in any of the forms
1. M (t) =P+kt(Q−P).
whereQis another point on the path and the particle is at Qwhent=1
k,
2. M (t) =P+dM
dtt
or
3. M (t) =P+Vt,
whereVis the velocity. See Exercise 5 too.
Exercises
(1) (a) If L:R1→R2such that the point X1= 3∈R1is mapped into P= (1,0) ,
which of the following points are in R(L) i) (2,0) , ii) (1,2) , iii) ( −1,0) ?
(b) Sketch two pictures, one of the graph of L, the other of the path of L.
(c) Find another operator ˜L:R1→R2whose path is the same as that for L.
182 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
(2) Find a mapping whose path is the straight line passing through the points (2 ,−1,3)
and (1,−3,−5) . Find its slope too.
(3) If a point is at (1 ,−1,0) att= 0 and at (2 ,3,8) att= 3 , find the position as a
function of time if the particle moves along a straight line. What is the velocity and
speed of the particle?
(4) If a particle is initially at (0 ,1,0,1) and has constant velocity (1 ,−2,3,−1) , find its
position as a function of time. Where is it at t= 3 ?
(5) A particle moves along a straight line in such a way that at t=t0it is at ˜P, while
att=t1it is at ˜Q.
(a) Show that its position M(t) as a function of time is
M(t) =˜P+ (t−t0)˜Q−˜P
t1−t0
(b) What is the velocity?
(c) Show that
M(t) =M(t0) +dM
dt(t−t0).
(6) Two straight lines are parallel if they have the same slope. If M(t) =P+t(Q−P) is
a parametric equation of one line, find an equation for the parallel line which passes
through the point ˜P.
4.5 L:Rn→R1. Hyperplanes.
Whereas in the previous section we examined linear mappings from a one-dimensional linear
space into an ndimensional space, now we shall look at the opposite extreme, linear
mappings from an ndimensional space into a one- dimensional space.
LetL:Rn→R1. We would like to find a representation theorem for this linear
operator. The most natural way to do this is to work with a basis {e1,...,e n}forRn.
Then every X∈Rncan be written as X=n/summationdisplay
1xkek. Consequently,
LX=L/parenleftBiggn/summationdisplay
1xkek/parenrightBigg
=n/summationdisplay
1L(xkek) =n/summationdisplay
1xkL(ek).
It is clear that LXis determined once we know all the numbers Lek. In other words, the
linear mapping Lis determined by the effect of the mapping on a basis for the domain of
the operator. This proves
Theorem 4.14 . LetL:Rn→R1linearly. If {ek}is a basis for the domain of L,Rn,
then
LX=a1x1+a2x2+...+anxn=n/summationdisplay
k=aakxk,
4.5.L:RN→R1. HYPERPLANES. 183
whereX=n/summationdisplay
1xkekandak=Lek. Notice that the akare scalars since they are in the
range ofL—and the range of LisR1by hypothesis.
Examples:
(1) Consider the linear operator L:R3→R1, which maps L:e1= (1,0,0)→1, L:e2=
(0,1,0)→0 , andL:e3= (0,0,1)→0 . Since the ekconstitute a basis for R3, the
mappingLis completely determined by using Theorem 14. If X= (x1,x2,x3)∈R3,
thenX=x1e1+x2e2+x3e3. Thus
LX=x1Le1+x2Le2+x3Le3=x1−x2
or
LX=x1.
For example, L: (2,1,7)→2 . The nullspace of L—those points X∈R3such that
LX= 0 —are the points X= (x1,x2,x3)∈R3such thatx1= 0 which is the x2x3
plane.
(2) LetL:R4→R1such thatLe1= 1 ,Le2=−2 ,Le3= 5 ,Le4=−3 , where
e1= (1,0,0,0) ,e2= etc. Then if X= (x1,x2,x3,x4)∈R4, we have
LX=x1−2x2+ 5x3−3x4.
The nullspace of Lis again a hyperplane, the hyperplane x1−2x2+ 5x3−3x4= 0
inR4.
So far we have not given any attention to the range of L, all of our pictures being in the
domain of L. Since the range is R1, its picture is a simple straight line which is not very
interesting. However the graph of Lis interesting. Let L:Rn→R1andY= (y)∈R1.
Then
y=a1x1+...+anxn.
The graph of Lis the set of points ( X,LX ) [or (X,Y) whereY=LX] inRn×R1∼=
Rn+1. A point (X,Y) = (x1,...,x n,y)∈Rn×R1is on the graph if the coordinates satisfy
the equation y=a1x1+...+anxn. This equation can be written as 0 = a1x1+...+
anxn+ (−1)ywhich is a hyperplane in Rn+1.
Thus we have found two ways to associate a hyperplane with L:Rn→R1,
i) AllXsuch thatLX= 0 , which is the nullspace of L, a linear space of dimension
n−1 (since dim N(L) = dim D(L)−dimR(L) =n−1) .
ii) The graph of L, that is, all points of the form ( X,LX ) , is a linear space of dimension
n+ 1 .
Although this is confusing, both ways are used in practice, whichever is most convenient
for the problem at hand. For the remainder of this section, we shall confine our attention
to hyperplanes defined in the first way.
Since linear mappings L:Rn→R1all have the form LX=a1x1+...+anxn, and
since it is natural to think of the sum as the scalar product of the vectors N= (a1,...,a n)
andX= (x1,...,x n) . Theorem 14 may be rephrased as
Theorem 4.15 . IfL:Rn→R1, thenLX=/angbracketleftN, X/angbracketright, whereNis the vector N=
(Le1,...,Le n)and{ek}form a basis for Rn.
184 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
Remark: The vector Nis an element of the so-called dual space ofRn. From the above,
it is clear that the dual space of Rnalso has dimension n.
Theorem 14’ is a “representation theorem”. It states that every linear mapping L:Rn→
R1may be represented in the form LX:=/angbracketleftN, X/angbracketrightfor some vector Nwhich depends on
L. You may wish to think of Nas a vector perpendicular to the hyperplane LX= 0 (cf.
Ex. 8, p. 225).
Example: Consider the operator Lof Example 2 in this section. For it, LX=/angbracketleftN, X/angbracketright
whereNis the particular vector N= (1,−2,5,3) .
Recall that a linear functional is a linear operator lwhose range is R1. Since the
operatorsL:Rn→R1we are considering have range R1, they are all linear functionals.
We may again rephrase Theorem 14 in this language. It states that every linear functional
defined on Rnmay be represented in the form l(X) =/angbracketleftN, X/angbracketright, whereNdepends on the
functional lat hand. This is just a restatement of Theorem 14 with the realization that
ourL’s are linear functionals. Don’t let the excess language bewilder you.
So far in this section, we have concentrated our attention on the algebraic representation
of a linear operator (functional) L:Rn→R1. Let us turn to geometry for a bit. In passing
we observed that the nullspace of the operator was a hyperplane in the domain of L(a
hyperplane in a linear space Vis a “flat” subset of Vwhose dimension is one less than V,
that is, of codimension one). These hyperplanes, {X∈Rn:LX= 0}, all passed through
the origin of Rn. A plane parallel to this one which passes through the particular point
X0∈Rnhas the form
L(X−X0) = 0.
It is clear that the point X=X0does satisfy the equation. From the representation
theorem,
L(X−X0) =a1(x1−x0
1) +a2(x2−x0
2) +...+an(xn−x0
n) = 0,
is the equation of this hyperplane, where X= (x1,x2,...,x n) andX0= (x0
1,x0
2,...,x0
n) .
If we again write N= (a1,a2,...,a n) , then the equation of the hyperplane is
/angbracketleftN, X−X0/angbracketright= 0,
all vectors Xsuch thatX−X0is perpendicular to N.
Examples:
(1) Find the equation of a plane which passes through the point X0= (1,2,−5) and is
parallel to the plane −2x1+ 7x2+ 4x3= 0 .
Solution : HereN= (−2,7,4) ,X= (x1,x2,x3) , so the plane has the equation
0 =/angbracketleftN, X−X0/angbracketright=−2(x1−1) + 7(x2−2) + 4(x3+ 5),
which may be written as
−2x1+ 7x2+ 4x3=−8.
The equation has been cooked up so that X0= (1,2,−5) does satisfy it.
(2) Find the equation of a plane which passes through the point X0= (1,2,−5) and is
parallel to the plane −2x1+ 7x2+ 4x3= 37 .
Solution : Since this plane is also parallel to the plane −2x1+ 7x2+ 4x3= 0 , the
solution is that of Example 1.
4.5.L:RN→R1. HYPERPLANES. 185
(3) Find the equation of the plane in R4which is perpendicular to the vector N=
(1,−2,3,1) and passes through the point X0= (1,0,1,−1) . Easy. The plane is all
pointsXsuch that
/angbracketleftN, X−X0/angbracketright= 0,
that is
(x1−1)−2(x2−0) + 3(x3−1) + (x4+ 1) = 0,
or
x1−2x2+ 3x3+x4= 3.
(4) Find the equation of the plane in R3which passes through the three points
X1= (7,0,0), X2= (1,0,−2), X3= (0,5,1).
We shall find this by using the general equation of a plane,
a1(x1−x0
1) +a2(x2−x0
2) +a3(x3−x0
3) = 0.
HereX0= (x0
1,x0
2,x0
3) is a particular point on the plane. We may use any of X1,X2,
orX3for it. Since X1is simplest, we take X0= (7,0,0) . All that remains is to
find the coefficients a1,a2, anda3in
a1(x1−7) +a2x2+a+ 3x3= 0.
SinceX2andX3are in the plane (and so must satisfy its equation), the substitution
X=X2andX=X3yields two equations for the coefficients,
a1(1−7) +a20 +a3(−2) = 0
a1(0−7) +a2(5) +a3(1) = 0.
These two equations in three unknowns may be solved for any two in terms of the
third. We find a3=−3a1anda2= 2a1, so the equation is
a1(x1−7) + 2a1x2−3a1x3= 0.
Factoring out the coefficient a1, we obtain the desired equation
x1−7 + 2x2−3x3= 0.
(It is clear from the general equation of a plane that the coefficients are determined
only to within a constant multiple).
Exercises
(1) LetL:R2→R1map
L: (1,0)→3, L : (0,1)→ −2.
WriteLXin the form Lx=a1x1+a2x2.L: (7,3)→?
186 CHAPTER 4. LINEAR OPERATORS: GENERALITIES. V1→VN,VN→V1
(2) LetL:R2→R1map
L: (2,1)→1, L : (0,3)→ −2.
WriteLXin the form LX=a1x1+a2x2.L: (7,3)→?
(3) Find the equation of a plane in R3which passes through the point (3 ,−1,2) and is
parallel to the plane x1−x2−2x3= 7 .
(4) Find the equation of a plane in R5which is perpendicular to the vector N=
(6,2,−3,1,−1) and contains the point (1 ,1,1,1,4) .
(5) Find the equation of a plane in R4which contains the four points X1= (2,0,0,0) ,
X2= (1,0,2,0) ,X3= (0,−1,0,−1) ,X4= (3,0,1,1) .
(6) In this problem, you will have to use the norm induced by the scalar product.
a). Show that the distance between the point Y∈Rnand the plane A={X∈
Rn:/angbracketleftN, X−X0/angbracketright= 0}is
d(Y,A) =/vextendsingle/vextendsingle/angbracketleftN, Y−X0/angbracketright/vextendsingle/vextendsingle
/bardblN/bardbl.
b). Prove that the distance between the parallel planes A={X∈Rn:/angbracketleftN, X−X1/angbracketright=
0}andB={X∈Rn:/angbracketleftN, X−X2/angbracketright= 0}is
d(A,B) =/vextendsingle/vextendsingle/angbracketleftN, X2−X1/angbracketright/vextendsingle/vextendsingle
/bardblN/bardbl.
Chapter 5
Matrices and the Matrix
Representation of a Linear
Operator
5.1 L:Rm→Rn.
The simplest example of a linear operator Lwhich maps RmintoRnis supplied by n
linear algebraic equations with mvariables. Let X= (x1,...,x m)∈Rm. Then we define
LX=
a11x1+a12x2+···+a1mxm
a21x1+a22x2+···+a2mxm
............................
an1x1+······ +anmxm
(5-1)
Notice the right side of this equation is a (column) vector with ncomponents. If we let
Y= (y1,...,y n) , then the equation LX=Yor
m/summationdisplay
j=1aijxj=yi, i = 1,2,...,n,
determines a vector YinRnfor everyXinRm. Since the operator Lis essentially
specified by the coefficients a11,a12,...,a nm, it is convenient to represent it by the notation
L=
a11a12···alm
a21a22···a2m
...................
anlan2···anm
,
and use the notation
LX=
a11a12···alm
a21a22···a2m
...................
anl........ a nm
x1
x2
.
xm
(5-2)
The ordered array of m×ncoefficients is called a matrix associated with L, and the
numbersaijare called the elements of the matrix. The first index irefers to the rowwhile
187
188 CHAPTER 5. MATRIX REPRESENTATION
the second index jrefers to the column . We may also write L= ((aij)) as a shorthand
to refer to the whole matrix. Since we shall only use linear operators in this chapter it
is convenient to drop the letter Lfor the operator and use A= ((aij)) instead. This
will facilitate the notation when referring to other matrices B= ((bin) , etc. since there
will be enough subscripts without adding to the confusion by using L1,L2, etc. for linear
operators.
In this section we shall work out the meaning of operator algebra applied to the special
case of operators L:Rm→Rnwhich are represented by matrices. It turns out that every
operatorL:Rm→Rncan be represented by a matrix (proved later in this very section).
Let us first i) define equality, ii) exhibit the matrices for the zero operator O(X) = 0
(additive identity). If A= ((aij)) andB= ((bij)) both map Rm→Rn, then by definition,
A=Bif and only if AX=BX for everyX∈Rm, that is, for all X= (x1,x2,...,x m) ,
ai1x1+ai2x2+···+aimxm=bilxl+···+bimxm, i= 1,2,...,n
orm/summationdisplay
j=1aijxj=m/summationdisplay
j=1bijxj, i = 1,2,...,n.
Subtracting, we find that
m/summationdisplay
j=1(aij−bij)xj= 0, i = 1,2,...,n
must hold for any choice of X= (x1,x2,...,x m) . From the particular choice X=
(1,0,0,...0) , we see that
ai1−bi1= 0, i = 1,2,...,n,
that is,
a11=b11,a21=b21,...,a nl=bnl.
Similarly, by using other vectors X, we conclude
Theorem 5.1 1 (equality ). IfA= ((aij))andB= ((bij))both map Rm→Rn, then
A=Bif and only if the corresponding elements of their matrices are equal,
aij=bij, i = 1,2,...,n, j = 1,2,...,m.
It is clear that the n×mmatrix all of whose elements are zero
0 =
0 0 ··· 0
0 0 ··· 0
............
0 0 ··· 0
has the property that it maps every X∈Rminto zero, and thus satisfies the conditions for
the zero matrix. That this is the only such matrix follows from Theorem 1, since any other
matrix which acts the same way on every vector X∈Rmmust have the same elements -
all zeroes.
5.1.L:RM→RN. 189
Theorem 5.2 2. The zero matrix 0:Rm→Rnis uniquely represented by a matrix with
nrows andmcolumns, all of whose elements are zero.
How is the identity matrix Idefined? Since I:Rn→Rnmaps every vector into
itself,IX=X, the linear equations (1) must have the property that given any vector
X= (x1,x2,...,x n)∈Rn, then
n/summationdisplay
j=1δijxj=xi, i = 1,2,...,n
Ifaij=δij(the Kronecker delta), so a11=a22=···=ann= 1 while aij= 0, i/negationslash=j,
then indeed
n/summationdisplay
j=1δijxj=xi, i = 1,2,...n
is satisfied. Thus, the coefficients of the identity matrix are I= ((δij)) . This is a square
(nxm) matrix,
I=
1 0 ··· 0
0 1 ··· 0
............
0 0 ··· 1
with ones along the main diagonal and zeroes elsewhere.
Theorem 5.3 3. The identity matrix I:Rn→Rnis uniquely represented by a square
(n×n)matrix whose elements are I= ((δij)).
We turn to addition. Let A= ((aij)) andB= ((bij)) be twon×mmatrices, so they
both represent operators mapping RmintoRn. Their sum C=A+Bis defined as the
operator which acts upon Xaccording to the rule (p. 268)
CX=AX+BX, X ∈Rm.
The elements cijof the matrix Cconsequently satisfy
m/summationdisplay
j=1cijxj=m/summationdisplay
j=1aijxj+m/summationdisplay
j=mbijxj, j = 1,2,...,n.
or
=m/summationdisplay
j=1(aij+bij)xj, j = 1,2,...,n
for allX= (x1,x2,...,x m) . Thus, the cijare in fact aij+bij(by Theorem 1)
Theorem 5.4 4. IfA= ((aij))andB= ((bij))both map RmintoRn, then their sum
C=A+Bhas elements
cij=aij+bij.
190 CHAPTER 5. MATRIX REPRESENTATION
Remark: From this it follows that the zero matrix is actually the additive identity, for if
A= ((aij)) , thenC=A+ 0 has elements cij=aij+ 0 =aij, that is,A+ 0 =A.
Example: 1 LetAandBwhich map R3→R4be represented by the matrices
A=
−3 0 1
7 2 −1
5 4 −3
0 1 1
;B=
2 2 2
−3 0 0
−4−2 2
0−1−1
.
Then
A+B=
−1 2 3
4 2 −1
1 2 −1
0 0 0
.
Example: 2. LetAandBbe the operators on p. 268 (called L1andL2there) which
mapR2→R3. Then
A=
1 1
1 2
0−1
, B =
−3 1
1−1
1 0
,
so
A+B=
−2 2
2 1
1−1
which agrees with the sum obtained there.
IfA= ((aij)) , is there a matrix ˜Asuch thatA+˜A= 0 ? Clearly the matrix ˜A
defined by ˜A= ((−aij)) does the job since
A+˜A= ((aij)) + (( −aij)) = ((0))
by definition of addition. We shall denote the matrix with elements (( −aij)) by “ −A”
sinceA+ (−A) = 0 . This matrix “ −A” is the additive inverse to A.
Example: If
A=
1−1
−π 2
0−1
then −A=
−1 1
π−2
0 1
.
Since a linear operator which is represented by a matrix is still a linear operator,
Theorem 3 (p. 269) certainly holds for matrix addition. We shall rewrite it.
Theorem 5.5 LetA,B,C,... be matrices which map RmintoRn(so they are n×
mmatrices). The set of all such matrices forms an abelian (commutative) group under
addition, that is,
1.A+ (B+C) = (A+B) +C
2.A+B=B+A
3.A+ 0 =A
4. For every A, there is a matrix ( −A) such that
A+ (−A) = 0.
5.1.L:RM→RN. 191
Proof: No need to do this again since it was carried out in even greater generality on
p. 270. For practice, you might want to write out the proof in the special case of 2 ×3
matrices and see how much more awkward the formulas become when you use the specific
elements instead of proceeding more abstractly as we did in the proof on p. 270.
Ifαis a scalar and A= ((aij)) is ann×mmatrix which represents a linear operator
mapping Rm→Rn, the operator αAisdefined by the rule
(αA)X=A(αX)
whereXis any vector in Rm. In terms of the elements (( aij)) , this means that the
elements ((˜ aij)) ofαAare given by
m/summationdisplay
j=1˜aijxj=m/summationdisplay
j=1aij(αxj), i = 1,2,...,n
=m/summationdisplay
j=1(αaij)xj, i = 1,2,...,n
so ˜aij=αaij. Thus, the matrix αAis found by multiplying each of the elements of Aby
α,
α
a11a12···a1m
a21........ a 2m
...................
anl... ... a nm
=
αa11αa12···αa1m
αa21αa22···αa2m
.......................
αanl...···αanm
Example:
−1
7 1 3
−2−1 4
9 6 5
−3 1 −1
=
−14−2−6
4 2 −8
−18−12−10
6−2 2
.
The following theorem concerns multiplication of matrices by scalars. It is proved either
by direct computation - or more simply by realizing that it is a special case of Exercise 12,
p. 284.
Theorem 5.6 . IfAandBare matrices which map Rm→Rn, and ifα,β are any
scalars, then
1.α(βA) = (αβ)A
2.1·A=A
3.(α+β)A=αA+βA
4.α(A+B) =αA+αB.
Remark: Theorems 5 and 6 together state that the set of all matrices which map Rminto
Rnforms a linear space . It is easy to show that the dimension of this space is m·n(by
exhibitingm·nlinearly independent matrices which span the whole space).
Now we get more algebraic structure and see how to multiply. Let AmapR?intoRn
andB= ((bij)) map RrintoRs. By definition of operator multiplication (p. 271-2), the
productABis defined on an element X∈Rr=D(B) by the rule
ABX =A(BX).
192 CHAPTER 5. MATRIX REPRESENTATION
Since the vector BX∈Rsmust be fed into A, we find that BX∈Rmtoo. Thus, in
order for the product AB of ann×mmatrixAwith as×rmatrixBto make sense,
we must have ? =m, that is, the range of Bmust be contained in the domain of A,
a figure goes here
IfC= ((cij)) =AB, then for every X∈Rr
CX=A(BX)
or
r/summationdisplay
k=1cijxk=s/summationdisplay
j=1aij/parenleftBiggr/summationdisplay
k=1bjkxk/parenrightBigg
, i= 1,2,...,n
so
=r/summationdisplay
k=1
s/summationdisplay
j=1aijbjk
xk, i= 1,2,...,n.
Therefore, the elements cikof the product ABare given by the formula
cik=s/summationdisplay
j=1aijbjki= 1,2,...,n
k= 1,2,...,r.
Since the summation signs have probably overwhelmed you, we repeat it in a special
case. LetBbe determined by the linear equations
b11x1+b12x2+b13x3=y1
b21x1+b22x2+b23x3=y2.
ThenB:R3→R2. Also letA:R2→R2be determined by
a11y1+a12y2=z1
a21y1+a22y2=z2.
The product AB maps a vector X∈R3first intoY=BX∈R2and then into Z=
ABX ∈R2.
a figure goes here
Ordinary substitution yields Z=ABX as a function of X:
a11(b11x1+b12x2+b13x3) +a12(b21x1+b22x2+b21x3) =z1
a21(b11x1+b12x2+b13x3) +a22(b21x1+b22x2+b23x3) =z2,
or
(a11b11+a12b21)x1+ (a11b12+a12b22)x2+ (a11b13+a12b23)x3=z1
(a21b11+a22b21)x1+ (a21b12+a22b22)x2+ (a21b13+a22b23)x3=z2.
5.1.L:RM→RN. 193
If we write this in the matrix form
/parenleftbiggc11c12c13
c21c22c23/parenrightbigg
x1
x2
x3
=/parenleftbiggz1
z2/parenrightbigg
,
we find
c11=a11b11+a12b21, c 12=a11b12+a12b22
etc., just as was dictated by the general formula for the multiplication of matrices.
Theorem 5.7 . IfA= ((aij))andB= ((bij))are matrices with B:Rr→Rsand
A:Rs→Rn, then the product C=AB is defined and the elements of the product C=
((cij))are given by the formula
cik=s/summationdisplay
j=aaijbjk, i= 1,2,...,n ;k= 1,2,...,r.
Remark: Since this formula for matrix multiplication is impossible to remember as it
stands, it is fortunate that there is an easy way to remember it. We shall work with the
example of matrices A:R2→R2andB:R3→R2discussed earlier. Then
AB=/parenleftbigga11a12
a21a22/parenrightbigg /parenleftbiggb11b12b13
b21b22b23/parenrightbigg
=/parenleftbiggc11c12c13
c21c22c23/parenrightbigg
.
To compute the element cik, we merely observe that
cik=2/summationdisplay
j=1aijbjk=ai1b1k+a12b2k
cikis the scalar product of theith row inAwith thekth column in B(see fig.). Thus,
the element c21inC=ABis the scalar product of the 2nd row of Awith the 1st column
ofB. Do not be embarrassed to use two hands to multiply matrices. Everybody does.
Examples:
(1) (cf. p. 274 where this was done without matrices). If
A=/parenleftbigg2−3
−1 1/parenrightbigg
, B =/parenleftbigg0 2
1 1/parenrightbigg
,
then
AB=/parenleftbigg2−3
−1 1/parenrightbigg/parenleftbigg0 2
1 1/parenrightbigg
=/parenleftbigg−3 1
1−1/parenrightbigg
and
BA=/parenleftbigg0 2
1 1/parenrightbigg/parenleftbigg2−3
−1 1/parenrightbigg
=/parenleftbigg−2 2
1−2/parenrightbigg
.
Notice that even though AB andBA are both defined, we have AB/negationslash=BA—the
expected noncommutativity in operator multiplication.
194 CHAPTER 5. MATRIX REPRESENTATION
(2) (cf. p. 272 bottom where this was done without matrices). If
A=
1−1
0 1
−1−2
, B = (1,2,−1),
then
BA= (1,2,−1)
1−1
0 1
−1−2
= (2,3).
However the product ABdoes not make sense.
From the general theory of linear operators (Theorem 4, p. 276) we can conclude
Theorem 5.8 . Matrix multiplication is associative, that is, if
RkA→RlB→RmC→Rn,
so the products C(BA)and (CB)Aare defined, then
C(BA) = (CB)A.
Thus the parenthesis can be omitted without risking chaos.
Remark: Returning to linear algebraic equations, you will observe that the matrix
notationAX there [eq(2)] can now be viewed as matrix multiplication of the n×m
matrixA= ((aij)) with the m×1 matrix (column vector) X.
In developing the algebra of matrices - and operators in general - we have been ne-
glecting one important issue, that of an inverse operator. If L:V1→V2, can we find
an operator ˜L:V2→V1which reverses the effect of L, that is, if LX=Y, where
X∈V1andY∈V2, is there an operator ˜Lsuch that ˜LY=X? If so, then
˜LLX =˜LY=X,
and we write
˜LL=I.
This operator ˜Lis the left(multiplicative) inverse of L. Similarly, an operator ˆL
such thatLˆL=Iis the right (multiplicative) inverse ofL. We shall shortly prove
that ifan operator Lhas both a left inverse Land a right inverse ˆL, then they are
equal, ˆL=˜L, so without ambiguity one can write L−1fortheinverse.
a figure goes here
5.1.L:RM→RN. 195
To begin, we compute the inverse of the matrix
A=/parenleftbigg5−2
3−1/parenrightbigg
associated with the system of linear equations
5x1−2x2=y1
3x1−x2=y2.
These equations specify a mapping from R2intoR2. They map a point XintoY.
Finding the inverse of Ais equivalent to answering the question, if we are given a point
Y, can we find the Xwhence it came?
AX=Y, X =A−1Y.
Finding the Xin terms of Ymeans solving these two equations, a routine task. The
answer is
x1=−y1+ 2y2
x2=−3y1+ 5y2.so/parenleftbiggx1
x2/parenrightbigg
=/parenleftbigg−1 2
−3 5/parenrightbigg/parenleftbiggy1
y2/parenrightbigg
Thus,
X=A−1Y,
where
A−1=/parenleftbigg−1 2
−3 5/parenrightbigg
.
The matrix A−1is the matrix inverse to A. It is easy to check that
AA−1=/parenleftbigg5−2
3−1/parenrightbigg/parenleftbigg−1 2
−3 5/parenrightbigg
=/parenleftbigg1 0
0 1/parenrightbigg
=I
and
A−1A=/parenleftbigg−1 2
−3 5/parenrightbigg/parenleftbigg5−2
3−1/parenrightbigg
=/parenleftbigg1 0
0 1/parenrightbigg
=I.
Thus, this matrix A−1is both the right and left inverse of A.
Our second example is of a more geometric nature. We shall consider a matrix Rwhich
represents rotation of a vector in E2through an angle α.
a figure goes here
Ris represented by the matrix (cf. Ex. 13b p. 285)
R=/parenleftbiggcosα−sinα
sinα cosα/parenrightbigg
.
It is geometrically clear that in inverse of this operator Ris an operator which rotates
through an angle −α, unwinding the effect of R. Thus, immediately from the formula for
R, we find
R−1=/parenleftbiggcos(−α)−sin(−α)
sin(−α) cos( −α)/parenrightbigg
=/parenleftbiggcosαsinα
−sinαcosα/parenrightbigg
.
196 CHAPTER 5. MATRIX REPRESENTATION
To check that geometry has not deceived us, we should multiply out RR−1andR−1R.
Do it. You will find RR−1=R−1R=I. One could also have found R−1by solving linear
algebraic equations as was done in the first example.
The problem of finding the matrix inverse to any square matrix,
A=
a11a12···a1n
a21a22···a2n
......
an1an2···ann
is equivalent to the dull problem of solving nlinear algebraic equations in nunknowns
a11x1+··· +a1nxn=y1
a21x1+··· +a2nxn=y2
......
an1x1+··· +annxn=yn
forXin terms of Y, X =A−1Y. Forn= 2 the computation is not too grotesque, and
yields the formulas
x1=a22
∆y1−a12
∆y2
x2=−a21
∆y1+a11
∆y1
where ∆ = a11a22−a12a21(= determinant of A, for those who have seen this before).
From this formula we read off that the inverse of the 2 ×2 matrix
A=/parenleftbigga11a12
a21a22/parenrightbigg
isA−1=1
∆/parenleftbigga22−a12
−a21a11/parenrightbigg
.
As a check, one computes that
AA−1=A−1A=I.
Thus the 2×2matrixAhas an inverse if and only if ∆ :=a11a22−a12a21/negationslash= 0 .
a figure goes here
Fortunately, one rarely needs the explicit formula for the inverse of a square n×n
matrix other than the reasonable cases n= 2 andn= 3 . The inverse of a matrix has
greater conceptual use as the inverse of an operator.
Having relegated the computation of the inverse of a matrix to the future, let us see
what can be said about the inverse without computation. This will necessarily be a bit
more abstract. Since the issues involve solving systems of linear algebraic equations, we
shall invoke the theory concerning that which was developed in Chapter 4 Section 3. For
this discussion, it is convenient to use the following definition (cf. p. 6).
Definition: An operator A:V1→V2isinvertible if it has the two properties
i) IfX1/negationslash=X2thenAX 1/negationslash=AX 2(injective, 1-1)
ii) To every Y∈V2, there is at least one X∈V1such thatAX=Y(surjective,
onto).
5.1.L:RM→RN. 197
Thus, an operator is invertible if and only if it is bijective. An invertible matrix is usually
called non-singular , while a matrix which is not invertible is called singular .
To show that this definition is identical with the previous one, we must show that every
invertible linear operator Ahas a right and left inverse. A more pressing matter though,
is
Theorem 5.9 . If the linear operator A:V1→V2whereV1andV2are finite dimen-
sional, is invertible, then dimV1= dimV2, so a matrix must necessarily be square for an
inverse to exist (but being square is not sufficient, as was seen in the 2×2case where the
additional condition a11a22−a21a12/negationslash= 0 we needed). In other words, you haven’t got a
chance to invert a matrix unless it is square, but being square is not enough.
Proof: Condition i) states that N(A) = 0 , for if X1/negationslash= 0 , thenAX 1/negationslash= 0 . Therefore
dimR(A) = dim D(A)−dimN(A) = dimV1−0 = dimV1.
On the other hand, condition ii) states that V2⊂R(A) . SinceA:V1→V2, we know that
R(A)⊂V2. Therefore R(A) =V2. Coupled with the first part, we have
dimV1= dim R(A) = dimV2.
Theorem 5.10 . Given an operator Awhich is invertible, there is a linear operator A−1
such thatAA−1=A−1A=I.
Proof: If˜Y∈V2, there is an ˜X∈V1such thatA˜X=˜Y(by property ii), and that ˜X
is unique (property i). Therefore without ambiguity we can define A−1˜Y=˜X. A similar
process defines the operator A−1for everyY∈V2. From our construction, it is clear (or
should be) that
AA−1=A−1A=I.
All that remains is to show A−1is linear. If A˜X=˜YandAˆX=˜Y, then since Ais linear,
A(a˜X+bˆX) =aA˜X+bAˆX=a˜Y+bˆY. ThusA−1(a˜Y+bˆY) =a˜X+bˆX=aA−1˜Y+bA−1ˆY.
Remark: Glancing over this proof, it should be observed that finite dimensionality (or
even the concept of dimension) never entered - so the result is true for infinite dimensional
spaces. Furthermore, linearity was only used to show that A−1was linear. Thus the
theorem (except for the claim that A−1is linear) is true for nonlinear operators as well.
Needless to say, this construction of A−1one point at a time is useless as a method for
findingA−1(since even in the simplest case A:R1→R1it involves an infinite number of
points).
This theorem shows that if an operator Ais invertible, then there are right and left
inverses which are equal AA−1=A−1A=I. We can reverse the theorem and prove
Theorem 5.11 . Given the linear operator A:V1→V2, if there are linear operators ˆA
(right inverse) and ˜A(left inverse) such that
AˆA=˜AA=I,
thenAis invertible and A−1=ˆA=˜A.
198 CHAPTER 5. MATRIX REPRESENTATION
Proof: Verify condition i: If AX 1=AX 2, then ˜AAX 1=˜AAX 2. Since ˜AA=I, this
impliesX1=X2.
Verify condition ii. If Yis any element in V2, letX=ˆAY. ThenAX=AˆAY=Y,
so thatYis the image of Xunder the mapping.
The proof that A−1=˜A=ˆAis delightfully easy. Only the associative property of
multiplication is used:
ˆA= (A−1A)ˆA=A−1(AˆA) =A−1= (˜AA)A−1=˜A(AA−1) =˜A.
Examples:
(1) The identity operator Ion every linear space is invertible, for it trivially satisfies
both criteria. Not only that, but it is its own inverse for II=I.
(2) The zero operator is never invertible, for even though X1/negationslash=X2, we always have
0(X1) = 0 = 0(X2) .
(3) The 2 ×2 matrix
A=/parenleftbigg1 3
−2−6/parenrightbigg
is not invertible since, from the formula
AX=/parenleftbigg1 3
−2−6/parenrightbigg/parenleftbiggx1
x2/parenrightbigg
=/parenleftbiggx1+ 3x2
−2x1−6x2/parenrightbigg
,
we see that the vector ( −3,1)/negationslash= 0 is mapped into zero by A(whereas criterion i).
states that only 0 can be mapped into 0 by an invertible linear operator). Another
way to see that Ais not invertible is to observe that ∆ = a11a22−a12a21= 0 . thus
violating the explicit condition for 2 ×2 matrices found earlier.
In this last example, we observed that if a linear operator Ais invertible, then by
property i) the equation AX= 0 has exactly one solution X= 0 . IfA:V1→V2on a
finite dimensional space, and dim V1= dimV2the converse is true also.
Theorem 5.12 If the linear operator Amaps the linear space V1intoV2anddimV1=
dimV2<∞, then
Ais invertible ⇐⇒AX= 0impliesX= 0.
Proof: ⇒A restatement of condition i) in the definition.
⇐A restatement of lines 7-10 on page 316.
Corollary 5.13 . A square matrix A= ((aij))is invertible if and only if its columns
A1=
a11
a21
·
·
·
an1
,A2=
a12
a22
·
·
·
an2
,...,An=
a1n
a2n
·
·
·
ann
are linearly independent vectors.
5.1.L:RM→RN. 199
Proof: To test for linear independence, we examine
xaA1+x2A2+···+xnAn= 0,
and try to prove that x1=x2=···=xn= 0 . But writing the equation in full, it reads
a11x1+a12x2+··· +a1nxn= 0
a21x1+a22x2+··· +a2nxn= 0
· · ·
· · ·
· · ·
an1x1+an2x2+··· +annxn= 0,
or
AX= 0.
By the theorem, Ais invertible if and only if the equation AX= 0 has only the solution
X= 0 . Thus Ais invertible if and only if the only solution of
x1A1+x2A2+···+xnAn= 0
isx1=x2=···=xn= 0 .
We close our discussion of invertible operators with
Theorem 5.14 . The set of all invertible linear operators which map a space into itself
constitutes a (non- commutative) group under multiplication; that is, if L1,L2,... are
invertible operators which map Vinto itself then they satisfy
0. Closed under multiplication ( L1L2is an invertible linear operator which maps V
into itself).
(1)L1(L2L3) = (L1L2)L3- Associative
(2)There is an identity Isuch that
IL=LI=L.
(3)For every operator Lin the set, there is another operator L−1for which
LL−1=L−1L=I.
Proof: 0)L1L2is a linear operator which maps Vinto itself by part 0. of Theorem 4
(p. 276). It is invertible since its inverse can be written in the explicit form (an important
formula)
(L1L2)−1=L−1
2LL−1
1,
as we will verify:
(L1L2)(L−1
2L−1
1) =L1(L2L−1
2)L−1
1=L1IL−1
1=L1L−1
1=I
connect these?? ( L−1
2L−1
1)(L1L2) =L−1
2(L−1
1L1)L2=L−1
2IL2=L−1
2IL2=L−1
2L2=I.
(1) Part 1 of Theorem 4 (p. 276).
200 CHAPTER 5. MATRIX REPRESENTATION
(2) Part 1 of Theorem 5 (p. 277)
(3) A direct restatement of the fact that our set consists only of invertible operators.
Closely associated with a matrix A:Rn→Rm
A=
a11a12···a1n
a21··· ··· ·
· ·
· ·
· ·
am1am2···amn
.
is another matrix A∗, the transpose oradjoint ofA, which is obtained by interchanging
the rows and columns of A, viz.
A∗=
a11a21···am1
a12a22··· ·
· ·
· ·
· ·
a1n··· ···amn
.
For example,
ifA=
1 2
4−2
5−2
,thenA∗=/parenleftbigg1 4 5
2−2−1/parenrightbigg
.
IfA= ((aij)) , thenA∗= ((aij)) . The adjoint of an m×nmatrix is an n×mmatrix.
Thus, ifA:Rn→RmthenA∗:Rm→Rn, and for any Z∈Rm, we have
A∗Z=
a11a21···am1
a12a22··· ·
· ·
· ·
· ·
a1n··· ···amn
z1
·
·
·
zm
=
a11z1+a21z2+··· +am1zm
a12z1+··· ··· +am2zm
· ·
· ·
· ·
a1nz1+··· ··· +amnzm
,
so thejth component ( A∗Z)jof the vector A∗Z∈Rnis
(A∗Z)j=m/summationdisplay
i=1aijzi=a1jz1+a2jz2+···+amjzm.
Beware : The classical literature on matrices uses the term “adjoint of a matrix” for an
entirely different object. Our nomenclature is now standard in the theory of linear operators.
A real square matrix Ais called symmetric orself-adjoint ifA=A∗. For example,
A=
7 2 −3
2−1 5
−3 5 4
=A∗.
For a symmetric matrix A, we haveaij=aji.
5.1.L:RM→RN. 201
The significance of the adjoint of a matrix (as well as its relation to the more general
conception of the adjoint of an arbitrary operator) arises in the following way. If A:En→
Em, then for any XinEnthe vectorY=AX is a vector in Em. We can form the scalar
product of this vector Y=AX with any other vector ZinEm(becauseYandZare
both in Em
/angbracketleftZ, Y/angbracketright=/angbracketleftZ, AX /angbracketright.
SinceA∗:Em→En, andZ∈Em, thenA∗Zmakes sense, and is a vector in En, so
/angbracketleftA∗Z, X/angbracketrightis a real number for any X∈En.Claim :
/angbracketleftZ, AX /angbracketright=/angbracketleftA∗Z, X/angbracketright.
This is easy to verify. Let A= ((aij)) . Then
(AX)i=n/summationdisplay
j=1aijxjand (A∗Z)j=m/summationdisplay
i=1aijzi,
so that
/angbracketleftZ, AX /angbracketright=m/summationdisplay
i=1zi(AX)i=m/summationdisplay
i=1zi(n/summationdisplay
j=1aijxj)
=m/summationdisplay
i=1n/summationdisplay
j=1ziaijxj.
In the same way,
/angbracketleftA∗Z, X/angbracketright=n/summationdisplay
j=1(A∗Z)jxj=n/summationdisplay
j=1(m/summationdisplay
i=1aijzi)xj
=m/summationdisplay
i=1n/summationdisplay
j=1ziaijxj.
Comparison reveals we have proved
Theorem 5.15 . IfA:En→Em, then for any X∈Enand anyZ∈Em,
/angbracketleftZ, AX /angbracketright=/angbracketleftA∗Z, X/angbracketright,
whereA∗is the adjoint of A.
Remark: From a more abstract point of view, the operator A∗is usually defined as the
operator which has the above property. If this definition is adopted, one must use it to prove
the adjoint A∗of a matrix Ais found by merely interchanging the rows and columns (try
to do it!).
It is remarkably easy to obtain some properties of the adjoint by using Theorem 14.
Our attention will be restricted to square matrices (although the results are still true with
but minor modifications for a rectangular matrix).
202 CHAPTER 5. MATRIX REPRESENTATION
Theorem 5.16 . LetAandBben×nmatrices (so the products AB,BA,B∗A∗,
A+Betc. are all defined). Then
0.I∗=I(becauseIis symmetric)
1.(A∗)∗=A
2.(AB)∗=B∗A∗
3.(A+B)∗=A∗+B∗.
4.(cA)∗=cA∗, cis a real scalar.
5.Ais invertible if and only if A∗is invertible, and
(A∗)−1= (A−1)∗.
6.Ais invertible if and only if the rows of Aare linearly independent.
Proof: We could use subscripts and the aijstuff - but it is clearer to use the result of
Theorem 14. In order to do so, an important preliminary result is needed.
Theorem 5.17 . IfC:En→Em, then the equation
/angbracketleftCx, Y /angbracketright= 0 for allXinEnandYinEm
⇐⇒Cis the zero operator, C= 0. Thus ifC1andC2mapEninto itself, the equation
/angbracketleftC1X, Y/angbracketright=/angbracketleftC2X, Y/angbracketrightfor allX,Y∈En⇐⇒C1=C2.
Proof: ⇒By contradiction, if C/negationslash= 0 there is some X0such that 0 /negationslash=CX 0∈En. Now
just pickY0=CX 0. Then
0 =/angbracketleftCX 0, Y0/angbracketright=/angbracketleftCX 0, CX 0/angbracketright=/bardblCX 0/bardbl2>0
because by assumption CX 0/negationslash= 0 . A glance at this line reveals the desired contradiction.
⇐Obvious.
The last assertion of the theorem follows by subtraction,
0 =/angbracketleftC1X, Y/angbracketright − /angbracketleftC2X, Y/angbracketright=/angbracketleftC1X−C2X, Y/angbracketright=/angbracketleft(C1−C2)X, Y/angbracketright
and letting C=C1−C2.
Now we return to the
Proof of Theorem 15 : The vectors X,Z will be in En.
(0) Particularly clear because Iis symmetric. You should try constructing another
proof patterned on those below.
(1) Two successive interchanges of the rows and columns of a matrix leave it unchanged.
Again, try to construct another proof patterned on those below.
(2)/angbracketleft(AB)∗Z, X/angbracketright=/angbracketleftZ, ABX /angbracketright=/angbracketleftZ, A(BX)/angbracketright=/angbracketleftA∗Z, BX /angbracketright
=/angbracketleftB∗(A∗Z), X/angbracketright=/angbracketleft(B∗A∗)Z, X/angbracketright
for allX,Z inEn. Application of Theorem 16 yields the result.
5.1.L:RM→RN. 203
(3)/angbracketleft(A+B)∗Z, X/angbracketright=/angbracketleftZ,(A+B)X/angbracketright=/angbracketleftZ, AX +BX/angbracketright
=/angbracketleftZ, AX /angbracketright+/angbracketleftZ,BX, =/angbracketright/angbracketleftA∗Z, X/angbracketright+/angbracketleftB∗Z, X/angbracketright
=/angbracketleftA∗Z+B∗Z, X/angbracketright=/angbracketleft(A∗+B∗)Z, X/angbracketright.
And apply Theorem 16.
(4)/angbracketleft(cA)∗Z, X/angbracketright=/angbracketleftZ, cAX /angbracketright=c/angbracketleftZ, AX /angbracketright
=c/angbracketleftA∗Z, X/angbracketright=/angbracketleft(cA∗)Z, X/angbracketright.
Apply Theorem 16.
(5) IfAis invertible, then AA−1=A−1A=I. An application of parts 0 and 2 shows
(A−1)∗A∗= (AA−1)∗=I∗=I.
Similarly,A∗(A−1)∗=I. ThusA∗has a left and right inverse, so it is invertible by
Theorem 11. The above formulas reveal ( A∗)−1= (A−1)∗.
In the other direction, assume A∗is invertible. Since A∗∗=A(part 1) the matrix
Ais the adjoint of A∗. But we just saw that if a matrix is invertible then its adjoint
is too. Thus the invertibility of A∗implies that of A.
(6) By the Corollary to Theorem 12, A∗is invertible if and only if its columns are linearly
independent. Since the columns of A∗are the rows of A, we find that A∗is invertible
if and only if the rows of Aare linearly independent. Coupled with Part 5, the proof
is completed.
In our later work we shall need an inequality. Why not insert it here for future reference.
Theorem 5.18 . IfA= ((aij))is anm×nmatrix, so A:En→Em, then for any X
inEnandYinEm
/bardblAX/bardbl ≤k/bardblX/bardbl
and
|/angbracketleftY, AX /angbracketright| ≤k/bardblX/bardbl /bardblY/bardbl,
where
k2=m/summationdisplay
i=1n/summationdisplay
j=1a2
ij.
Proof: By definition
/bardblAX/bardbl2=m/summationdisplay
i=1(AX)2
i=m/summationdisplay
i=1(n/summationdisplay
j=1aijxj)2,
where (AX)iis theith component of the vector AX. The Schwarz inequality shows
(n/summationdisplay
j=1aijxj)2≤n/summationdisplay
j=1a2
ijn/summationdisplay
j=1x2
j=/bardblX/bardbl2n/summationdisplay
j=1a2
ij.
204 CHAPTER 5. MATRIX REPRESENTATION
Thus,
/bardblAX/bardbl2≤ /bardblX/bardbl2m/summationdisplay
i=1(n/summationdisplay
j=1a2
ij) =k2/bardblX/bardbl2,
which proves the first part. The second part follows from this and one more application of
Schwarz:
|/angbracketleftY, AX /angbracketright| ≤ /bardblY/bardbl/bardblAX/bardbl ≤k/bardblX/bardbl/bardblY/bardbl.
After all of this detailed discussion of matrices as an example of a linear operator L
mapping one finite dimensional space into another, our next theorem will show why matrices
are so ubiquitous. You see, we shall prove that every such linear operator L:V1→V2can
be represented as a matrix after bases forV1andV2have been selected.
Theorem 5.19 . (Representation Theorem) Let Lbe a linear operator which maps one
finite dimensional space into another
L:V1→V2.
Let{e1,e2,...,e n}be a basis for V1, and {θ1,θ2,...θ m}be a basis for V2. Then in
terms of these bases Lmay be represented by the matrix θLewhosejth column is the
vector (Lej)θ, that is, the vector Lej(which is a vector in V2) written in terms of the θ
basis forV2. Pictorially we have
θLe= ((Le1)θ···(Len)θ).
Proof: Finding the representation of Lin terms of given bases for V1andV2means:
given a vector XinV1which is represented in the ebasis forV1(write it as Xe) to find
a matrix θLesuch that the image vector θLeXeis the image ( LX)θofXwritten in the
θbasis forV2. We have used the cumbersome notation θLeto make explicit the fact that
it maps vectors written in the ebasis forV1into vectors written in the θbasisV2.
To avoid even further notation, we shall carry out the details only for the particular case
where the domain V1is two dimensional with basis {e1,e2}andV2is three dimensional
with basis {θ1,θ2,θ3}. The general case is proved in the same way.
Since the vectors Le1andLe2are inV2, they can be written in the θbasis, say,
Le1=a1θ1+b1θ2+c1θ3, Le 2=a2θ1+b2θ2+c2θ3,
so
(Le1)θ=
a1
b1
c1
and (Le2)θ=
a2
b2
c2
.
GivenXinV1, it can be written in the ebasis forV1,
X=x1e1+x2e2,soXe=/parenleftbiggx1
x2/parenrightbigg
.
Then
LX=L(x1e1+x2e2) =x1Le1+x2Le2
=x1(a1θ1+b1θ2+c1θ3) +x2(a2θ1+b2θ2+c2θ3)
= (a1x1+a2x2)θ1+ (b1x1+b2x2)θ2+ (c1x1+c2x2)θ3.
5.1.L:RM→RN. 205
If we write LXas a column vector in the θbasis it is
(LX)θ=
a1x1+a2x2
b1x1+b2x2
c1x2+c2x3
which is recognized as a product
θLe=
a1a2
b1b2
c1c2
/parenleftbiggx1
x2/parenrightbigg
=
a1a2
b1b2
c1c2
Xe
therefore, the matrix we want is
θLe=
a1a2
b1b2
c1c2
((Le1)θ(Le2)θ),
a matrix whose jth column is the vector Lejwritten in the θbasis forV2.
Example: Consider the integral operator L: =/integraltextx
0as a map of the two dimensional space
P1into the three dimensional space P2. Any bases for P1andP2will do, however we
must simply fix our attention to specific bases. Say
basis for P1:={e1(x) = 1, e 2(x) =x}
basis for P2:={θ1(x) =1+x
2, θ 2(x) =1−x
2, θ 3(x) =x2}.
Then
Le1=/integraldisplayx
01dt=x=θ1−θ2
and
Le2=/integraldisplayx
0tdt=x2
2=1
2θ3.
Therefore
(Le1)θ=
1
−1
0
,and (Le2)θ=
0
0
1
2
,
so
θLe= ((Le1)θ(Le2)θ) =
1 0
−1 0
01
2
is the matrix representing Lin terms of the given ebasis for P1andθbasis for P2. To
make you believe this, let us evaluate
LP=/integraldisplayx
0P
for some polynomial p∈P1by using the matrix. For example, p(x) = 3−x= 3e1−e2,
so in theebasis for P1, P 3=/parenleftbigg3
−1/parenrightbigg
. Its image under Lin terms of the θbasis for
P2is then
(Lp)θ=θL(pe)
e=
1 0
−1 0
01
2
/parenleftbigg3
−1/parenrightbigg
=
3
−3
−1
2
;
206 CHAPTER 5. MATRIX REPRESENTATION
that is,
Lp= 3θ1−3θ2−1
2θ3= 3(1 +x
2)−3(1−x
2)−1
2(x2) = 3x−1
2x2
which, of course, agrees with
/integraldisplayx
0p(t)dt=/integraldisplayx
0(3−t)dt= 3x−1
2x2.
WARNING : If we had used a different basis for either P1orP2, the resulting matrix
representing Lwould be different. For example, if the same basis were used for P1but a
different basis for P2,
˜θbasis for P2:={˜θ1(x) = 1,˜θ2(x) =x,˜θ3(x) =x2},
then
Le1=x=˜θ2 andLe2=x2
2=1
2˜θ3,
so
(Le1)˜θ=
0
1
0
, (Le2)˜θ=
0
0
1
2
.
Therefore the matrix ˜θLewhich represents Lin terms of the ebasis for P1and the ˜θ
basis for P2is
˜θLe=
0 0
1 0
01
2
.
Again, ifp(x) = 3−x= 3e1−e2, then in the ˜θbasis
(Lp)˜θ=˜θLePe=
0 0
1 0
01
2
/parenleftbigg3
−1/parenrightbigg
=
0
3
−1
2
;
that is,
Lp= 0˜θ1+ 3˜θ2−1
2˜θ3= 3x−1
2x2,
to no one’s surprise.
Observe that the matrices θLeand ˜θL3both represent L—but with respect to dif-
ferent basis. The second matrix ˜θLeis somewhat simpler that the first since it has more
zeroes. It is often useful to pick bases in order that the representing matrix be as simple as
possible. We shall not discuss that issue right now.
There is a simple class of operators (transformations) which are not linear, but enjoy
most of the properties which linear ones do. They are affine operators, or affine transfor-
mations. To define them, it is best to first define the translation operator.
Definition: IfVis any linear space and Y0a particular element of V, then the operator
T:V→Vdefined by
TY=Y+Y0, Y∈V,
is the translation operator . It translates a vector Yinto the vector Y+Y0.
5.1.L:RM→RN. 207
Definition: Anaffine transformation Ais a linear transformation Lfollowed by a trans-
lation. ifL:V1→V2andY0∈V2, it has the form
AX:=LX+Y0. X ∈V1, Y 0∈V2.
Affine transformations can be added and multiplied by the same definition which gov-
erned linear transformations. Thus, if AandBare affine transformations mapping V1
intoV2,
(A+B)X:=AX+BX.
In particular, if AX=L1X+Y0andBX=L2X+Z0, whereY0andZ0are inV2, then
(A+B)X=AX+BX=L1X+Y0+L2X+Z0
= (L1+L2)X+ (Y0+Z0).
Similarly, if A:V1→V2andB:V3→V4, whereV2⊂V3, then
(BA)X:=B(AX) =B(L1X+Y0) =L2(L1X+Y0) +Z0
=L2L1X+L2Y0+Z0,
whereY0∈V2andZ0∈V4.
You will carry out the (straightforward) proofs of the algebraic properties for affine
transformations in Exercise 23.
The curtain on this longest of sections will be brought down with a brief discussion
of the operators which characterize rigid body motions, or Euclidean motions, as they are
often called.
Definition: The transformation R:En→Enis an isometric transformation , (orEuclidean
transformation or rigid body transformation) if the distance between two points is preserved
(invariant) under the transformation. Thus, Ris an isometry if
/bardblRX−RY/bardbl=/bardblX−Y/bardbl
for allXandYinEn.
It is interesting to think for a moment how all these names originated. The phrase
rigid body transformation arises from the idea that any motion of a rigid body (such as a
translation or rotation) does not alter the distance between any two points in the body. In
the framework of Euclidean geometry the whole notion of congruence is defined to be just
those properties of a figure which are invariant under isometries. By allowing deformations
other than isometries, one obtains geometries, so affine geometry is the study of properties
invariant under all affine motions.
The study of isometric transformations is mainly contained in that of a special case,
orthogonal transformations . These are isometries which leave the origin fixed, R0 = 0 . It
should be clear from our next theorem (part 3) that the idea of an orthogonal transformation
generalizes the idea of a rotation to higher dimensional space. Reflections (mirror images)
are also orthogonal transformations. Theorem 20 states that every isometric transformation
is the result of an orthogonal transformation followed by a translation.
Example: The matrix R=/parenleftbigg1 0
0−1/parenrightbigg
defines an orthogonal transformation since if
X=/parenleftbiggx1
x2/parenrightbigg
,thenRX=/parenleftbigg1 0
0−1/parenrightbigg/parenleftbiggx1
x2/parenrightbigg
=/parenleftbiggx1
−x2/parenrightbigg
,
208 CHAPTER 5. MATRIX REPRESENTATION
and if
Y=/parenleftbiggy1
y2/parenrightbigg
,thenRY=/parenleftbiggy1
−y2/parenrightbigg
.
Consequently /bardblRX−RY/bardbl=/bardblX−Y/bardbl=/radicalbig
(x1−y1)2+ (x2−y2)2, soR, being isometric
and linear is an orthogonal transformation. It represents a reflection across the x1axis.
Our definition of an orthogonal transformation does not presume its linearity. This is
because the linearity is a consequence of the given properties. A proof is outlined in Ex.
16, p. 390. For convenience, the linearity will be assumed in the following theorem where
we collect the standard properties of orthogonal transformations.
Theorem 5.20 . LetR:En→Enbe a linear transformation. The following properties of
Rare equivalent.
(1)Ris an orthogonal transformation, that is
/bardblRX−RY/bardbl=/bardblX−Y/bardblandR0 = 0.
(2)/bardblRX/bardbl=/bardblX/bardbl
(3)/angbracketleftRX, RY /angbracketright=/angbracketleftX, Y/angbracketright(so angles are preserved)
(4)R∗R=I
(5)Ris invertible and R−1=R∗. (Only in this part do we use the finite dimension-
ality of En).
Proof: We shall prove the following chain of implications: 1 = ⇒2 =⇒3 =⇒4 =⇒5 =⇒
4 =⇒1
1 =⇒2 . Trivial, for
/bardblRX/bardbl=/bardblRX−R0/bardbl=/bardblX−0/bardbl=/bardblX/bardbl.
2 =⇒3 . By linearity and part 2) applied to the vector X+Y, we have
/bardblRX+RY/bardbl=/bardblR(X+Y)/bardbl=/bardblX+Y/bardbl.
Now square both sides and express the norm as a scalar product:
/angbracketleftRX+RY, RX +RY/angbracketright=/angbracketleftX+Y, X +Y/angbracketright.
Upon expanding both sides, we find that
/bardblRX/bardbl2+ 2/angbracketleftRX, RY /angbracketright+/bardblRY/bardbl2=/bardblX/bardbl2+ 2/angbracketleftX, Y/angbracketright+/bardblY/bardbl2.
Since by part 2) /bardblRX/bardbl=/bardblX/bardbland/bardblRY/bardbl=/bardblY/bardbl, we are done.
3 =⇒4 . By part 3) and Theorem 14 (p. 369),
/angbracketleftR∗RX, Y /angbracketright=/angbracketleftRX, RY /angbracketright=/angbracketleftX, Y/angbracketright.
Thus, an application of the second part of Theorem 16 (p. 371) gives us R∗R=I.
4 =⇒5 . SinceX=R∗RX, we see that RX= 0 implies X= 0 , consequently, Ris
invertible (Theorem 12, p. 364). Moreover R∗R=IsoR∗=R−1.
5 =⇒4 . Clear, since R∗=R−1.
5.1.L:RM→RN. 209
5 =⇒1 . Because Ris linear,R0 = 0 . It remains to show that /bardblRX−RY/bardbl=/bardblX−Y/bardbl,
an easy computation. /bardblRX−RY/bardbl2=/bardblR(X−Y)/bardbl2=/angbracketleftR(X−Y), R(X−Y)/angbracketright, so using 4)
=/angbracketleftR∗R(X−Y), X−Y/angbracketright=/angbracketleft(X−Y),(X−Y)/angbracketright=/bardblX−Y/bardbl2.
Done.
Earlier in this section (p. 357-8) we considered a matrix Rwhich represented the
operator which rotates a vector in E2through an angle α. This matrix is the simplest
(non-trivial) example of a rigid body transformation which leaves the origin fixed, that is,
an orthogonal transformation.
R=/parenleftbiggcosα−sinα
sinα cosα/parenrightbigg
.
To prove that Ris an orthogonal matrix, by Theorem 19 part 3, it is sufficient to verify
/angbracketleftRX, RY /angbracketright=/angbracketleftX, Y/angbracketrightfor allXandYinE2. A calculation is in order here.
RX=/parenleftbiggcosα−sinα
sinα cosα/parenrightbigg/parenleftbiggx1
x2/parenrightbigg
=/parenleftbiggx1cosα−x2sinα
x1sinα+x2cosα/parenrightbigg
.
Similarly for RY, just replace x1andx2byy1andy2respectively. Then
/angbracketleftRX, RY /angbracketright= (x1cosα−x2sinα)(y1cosα−y2sinα) + missing?
(x1sinα+x2cosα)(y1sinα+y2cosα)
=x1y1cos2α−(x1y2+x2y1) sinαcosα+x2y2sin2α
+x1y1sin2α+ (x1y2+x2y1) sinαcosα+x2y2cos2α
=x1y1+x2y2=/angbracketleftX, Y/angbracketright.Done.
We previously found an expression for R−1(p. 358) by geometric reasoning. It is
reassuring to notice R−1=R∗, just as part 5 of our theorem states.
The most general rotation in E3may be decomposed into a product of these simple
two dimensional rotations. For a brief discussion - complete with pictures - open Goldstein,
Classical Mechanics to pp. 107-9.
Now to the last theorem of this section.
Theorem 5.21 . IfR:En→Enis a rigid body transformation, then for every X∈En
RX=R0X+X0,
whereR0is an orthogonal transformation (rotation) and X0is a fixed vector in En. Thus,
every rigid body motion is composed of a rotation (by R0and a translation (through X0).
Proof: LetR0X=RX−R0 . Since
R00 =R0−R0 = 0,
the operator R0has the property R00 = 0 . Furthermore, for any XandYinEn,
/bardblR0X−R0Y/bardbl=/bardblRX−R0−RY+R0/bardbl
=/bardblRX−RY/bardbl=/bardblX−Y/bardbl.
ThereforeR0satisfies the definition of an orthogonal transformation. The proof is com-
pleted by defining X0to be the image of the origin under R, X 0=R0 . Then
R0X=RX−X0,
or
RX=R0X+X0.
210 CHAPTER 5. MATRIX REPRESENTATION
5.2 Supplement on Quadratic Forms
Quadratic polynomials of the form
Q(X) =αx2
1+βx1x2+γx2
2, X = (x1,x2)
and the generalization to nvariablesX= (x1,x2,...,x n)
Q(X) =n/summationdisplay
in/summationdisplay
j=1αijxixj
often arise in mathematics. They are called quadratic forms and can always be represented
in the form /angbracketleftX, SX /angbracketrightwhereSis a self adjoint matrix. For example, the first quadratic
form can be written as
Q(X) = (x1,x2)/parenleftbiggαβ
2β
2γ/parenrightbigg/parenleftbiggx1
x2/parenrightbigg
=/angbracketleftX, SX /angbracketright,
whereSis the matrix indicated.
The procedure for finding the elements (( aij)) of the matrix Sis simple. First take
care of the diagonal terms by letting aiibe the coefficient of x2
1inQ(X) . Realizing that
xixj=xixj, collect the terms αijxixjandαjixjxiinQ(X) , getting ( αij+αji)xixj.
Then let
aij=aji=1
2(αij+αji)i/negationslash=j.
Example: Q(X) =x2
1−2x1x3−x2
2+ 6x1x2+ 4x3x1. Rewrite this as Q(X) =x2
1−x2
2+
6x1x2+ 2x1x3. Then
S=
1 3 1
3−1 0
1 0 0
and
Q(X) =/angbracketleftX, SX /angbracketright.
as you can easily verify.
Definition: A quadratic form Q(X) ispositive semi definite ifQ(X)≥0 for allXand
positive definite ifQ(X)>0, x/negationslash= 0. Q(X) isnegative semi definite ornegative definite
if, respectively, Q(X)≤0 , orQ(X)<0, X/negationslash= 0 . IfSis the self adjoint matrix associated
with the quadratic form Q(X) , thenSis positive semi definite, positive definite, etc., if
Q(X) has the respective property.
We may think of Q(X) as representing a quadratic surface. Thus, if Sis diagonal,
for example
S=
2 0 0
0 1 0
0 0 3
,
with positive diagonal elements, then the equation Q(X) = 1 , where Q(X) =/angbracketleftX, SX /angbracketright=
2x2
1+x2
2+3x2
3, represents an ellipsoid. This matrix Sis positive definite since by inspection
Q(X)>0, X /negationslash= 0 .
It is easy to see if a diagonal matrix Sis positive semi definite, negative semi definite,
positive definite, or negative definite.
5.2. SUPPLEMENT ON QUADRATIC FORMS 211
Example: The diagonal matrix
S=
γ1...0
...
0...γ n
is
(a) positive semi definite if and only if γ1,...,γ nare all non-negative,
(b) positive definite if and only if γ1,...,γ nare all positive (not zero), and the obvious
statements for negative semi definite and negative definite.
The problem of determining if a non diagonal symmetric matrix is positive etc. is more
subtle. We shall find necessary and sufficient conditions for the two variable case, but only
necessary conditions for the general case.
Consider the 2 ×2 self-adjoint matrix
S=/parenleftbigga b
b c/parenrightbigg
and the associated quadratic form
Q(X) =ax2+ 2bxy+cy2.
There are several cases.
(i)Ifa= 0 , then
Q(X) = (2bx+cy)y.
Ifb/negationslash= 0 , by choosing xandyappropriately, we can make Q(X) assume bothpositive
and negative values. Thus, for a= 0, b/negationslash= 0, Q can be neither a positive nor a negative
semi-definite form. On the other hand, if a= 0 , andb= 0 , thenQis positive (negative)
semi definite if and only if c≥0 (c≤0) . Ifa= 0, Q can never be positive definite or
negative definite since if X= (x,0) wherex/negationslash= 0 , thenQ(X) = 0 butX/negationslash= 0 .
(ii)Ifa/negationslash= 0 , thenQcan be written as
Q(X) =1
a[(ax+by)2= (ac−b2)y2].
We can immediately read off the conditions from this. Qis positive semi definite (definite)
if and only if a>0 andac−b2≥0 (ac−b2>0) , and negative semi definite (definite) if
and only if a<0 andac−b2≥0 (ac−b2>0) .
In summary, we have proved
Theorem 5.22 A. LetQ(X) =ax2+ 2bxy+cy2, andSbe the associated symmetric
matrix. Then
(a)Qis positive semi definite if and only if a≥0andac−b2≥0(this implies c≥0
too).
(b)Qis positive definite if and only if a >0andac−b2>0(this implies c >0
too).
212 CHAPTER 5. MATRIX REPRESENTATION
The general case of a quadratic form in nvariables is much more difficult to treat.
There are known necessary and sufficient conditions, but they are not too useful in practice,
especially for a large number of variables. We shall only prove one necessary condition for a
quadratic form to be positive semi-definite (or positive definite), a condition which is both
transparent to verify in practice and even easier to prove.
THEOREM B. If the self adjoint matrix S= ((aij)) is positive definite, then the
diagonal elements must all be positive, a11,a22,...,a nn>0 . Similarly, if Sis negative
definite then the diagonal elements must all be negative.
Proof:Q(X) =/angbracketleftX, SX /angbracketright=n/summationdisplay
i,j=1aijxixj. SinceQis positive definite, Q(X)>0 for all
X/negationslash= 0 . In particular, Q(ek)>0, k= 1,...,n , whereekis thekth coordinate vector
ek= (0,0,..., 0,1,0,..., 0) . ButQ(ek) =akk. Thusakk>0, k= 1,...,n , just what we
wanted to prove.
Examples: 1. The quadratic form Q(X) = 3x2+743xy−y2+4z2+xzis positive definite
or semi definite since the coefficient of y2is negative. It is not negative definite or semi
definite since the coefficient of x2is positive.
2. The quadratic form Q(X) =x2−5xy+y2+ 2z2satisfies the necessary conditions
of Theorem B, but the conditions of Theorem D were not sufficient conditions for positive
definiteness. Thus, we cannot conclude this Q(X) is positive definite. In fact, this Q(X) is
notpositive definite or semi definite since, for example, if X= (1,1,1) , thenQ(X) =−1 .
It is clearly not negative definite or semi definite.
Exercises
(1) Find the self-adjoint matrix Sassociated with the following quadratic forms:
(a)Q(X) =x2
1−2x1x2+ 4x2
2.
(b)Q(X) =−x2
1+x1x2−x1x3+x2
2−3x2x1−2x3x2+ 3x2
3
(c)Q(X) = 2x1x2−3x3x2+ 4x2x4+x3x4+ 7x2
2
[Answers: (a)/parenleftbigg1−1
−1 4/parenrightbigg
, (b)
−1−1−1
2
−1 1 −1
−1
2−1 3
, (c)
0 1 0 0
1 7 −3
22
0−3
201
2
0 21
20
(2) Use Theorem AorBto determine which of the following quadratic forms in two
variables are positive or negative definite, or semi definite, or none of these.
(a)Q(X) =x2
1−2x1x2+ 4x2
2
(b)Q(X) =−x2
1+x1x2−4x2
2
(c)Q(X) =x2
1−6x1x2−4x2
2
(d)Q(X) =x2
1−6x1x2+ 4x2
2
(e)Q(X) =x2
1−6x1x2+ 4x2x3−x2
2+ 4x2
3
(3) If the self-adjoint matrix Sis positive definite, prove it is invertible. Give an example
of an invertible self-adjoint matrix which is neither positive nor negative definite.
5.2. SUPPLEMENT ON QUADRATIC FORMS 213
(4) Find all real values for λfor which the quadratic form
Q(X) = 2x2+y2+ 3z2+ 2λxy+ 2xz
is positive definite. [Hint: Q(X) = (5
3−λ2)x2+ (λx+y)2+ (√
3z+1√
3x)2]
(5) Let the integer nbe≥3 . If the quadratic form
Q(X) =n/summationdisplay
i,j=1aijxixj, a ij=aji
is the product of two linear forms
Q(X) = (n/summationdisplay
i=1λixi)(n/summationdisplay
j=1µjxj),
show that det A= det((aij)) = 0 .
(6) If the self-adjoint matrix Sis positive definite or semi-definite, prove the generalized
Schwarz inequality :
|/angbracketleftY, SX /angbracketright|2≤ /angbracketleftY, SY /angbracketright/angbracketleftX, SX /angbracketright
for allXandY. [Hint: Observe [ X,Y] :=/angbracketleftY, SX /angbracketrightsatisfies all the axioms for a
scalar product].
(7) If the self-adjoint matrix Sis positive definite (so S−1exists by Exercise 3), prove
thatS−1is also positive definite. [Hint: Use the generalized Schwarz inequality,
Exercise 6, with Y=S−1Xand the inequality /angbracketleftX, SX /angbracketright ≤k2/bardblX/bardbl2of Theorem 17,
p. 373].
(8) Proof or counterexample:
(a) If a matrix A= ((aij)) is positive definite, then all of its elements are positive,
aij>0 for alli,j.
(b) If a matrix Ais such that all of its elements are positive, aij>0 , then the
matrix is positive definite.
Exercises
(1) Write out the matrices associated with the operators AandBin Exercise 4a, p.
281, and carry out the computation there using matrices.
(2) Write out the matrices RA,RB, andRCfor the rotation operators A,B , andC
in Exercise 8 p. 281 and complete that problem using matrices. [Ans. RA=
1 0 0
0 0 −1
0 1 0
in terms of the basis e1= (1,0,0), e 2= (0,1,0), e 3= (0,0,1) ].
(3) Prove Exercise 2b (p. 281) as a corollary of Theorem 18.
214 CHAPTER 5. MATRIX REPRESENTATION
(4) If
A=/parenleftbigg0 0
0 1/parenrightbigg
, B =/parenleftbigg0 1
0 0/parenrightbigg
computeAB, BA , andB2.
(5) Compute A−1if
(a).A=/parenleftbigg1 2
3 4/parenrightbigg
[ans.A−1=/parenleftbigg−2 1
3
2−1
2/parenrightbigg
]
(b).A=
4 0 5
0 1 −6
3 0 4
[ans.A−1=
4 0 −5
−18 1 24
−3 0 4
]
(c)A=
1 1 0 0
0 1 1 0
0 0 1 1
0 0 0 1
[ans.A−1=
1−1 1 −1
0 1 −1 1
0 0 1 −1
0 0 0 1
]
(6) IfAis the matrix of 5a) above, from the definition compute directly ,
(a)−6A−1+1
2A∗[ans./parenleftbigg25
2−9
2
−8 5/parenrightbigg
] .
(b) (A∗)−1and (A−1)∗. Compare them. State and prove a general theorem.
(c)AA∗andA∗A.
(7) If
A=/parenleftbigg1 2
3 4/parenrightbigg
, B =/parenleftbigg1 1
2−1/parenrightbigg
,
compute (AB)∗, A∗B∗, andB∗A∗. Compare ( AB)∗andB∗A∗and explain the
outcome.
(8) Prove that I∗=IandA∗∗=Ausing only Theorems 14 and 16 (cf. Parts 2-4 of
Theorem 15).
(9) IfA:Rn→RmandB:Rm→Rnwheren>m , prove that BA(ann×nmatrix)
is singular. Is ABnecessarily singular? (Proof or counterexample).
(10) Given two square matrices AandBsuch thatAB= 0 , which of the following
statements are always true. Proofs or counterexamples are called for. [I suggest you
confine your search for counterexamples to the case of 2 ×2 matrices.]
(a).A= 0 .
(b).B= 0.
(c).Aand/orBare (is) singular (not invertible).
(d).Ais singular.
(e).B−1exists.
(f). IfA−1exists, then B= 0 .
(g). IfBis nonsingular, then A=C.
5.2. SUPPLEMENT ON QUADRATIC FORMS 215
(h).BA= 0 .
(i). IfA/negationslash= 0 andB/negationslash= 0 , then neither AnorBare invertible.
(11) (a). If Ais a square matrix which satisfies
A2−2A−I= 0,
findA−1in terms of A. [Hint: Find a matrix Bsuch thatAB=BA=I.]
(b). IfAis a square matrix which satisfies
An+an−1An−1+an−2An−2+...+a1A+a0I= 0, a 0/negationslash= 0,
wherea0,a1,...,a n−1are scalars, prove that Ais invertible and find A−1in terms
ofA.
(12) (a). If L:En→Em, prove N(L∗) =R(L)⊥
[Hint: Show (in two lines) that X∈N(L∗)⇐⇒ /angbracketleftX, LZ /angbracketright= 0 for all Z∈En—from
which the result is immediate.]
(b). Use part (a) to show that dim R(L) = dim R(L∗) .
(c). Do exercise 19, page 441.
(13) (a). If T:En→Enis a translation, TX=X+X0, proveTis invertible by explicitly
findingT−1(which is a trivial task). [Answer: T−1X=X−X0.]
(b). IfR:En→Enis a rigid body transformation, show that Ris always invertible
by exhibiting R−1. [Answer: If RX =R0X+X0, thenRcan be written as
Rx= (TR 0)X. R−1=R∗
0T−1.]
(14) IfAis anyn×nmatrix, find matrices A1andA2such thatAis decomposed into
the two parts
A=A1+A2
whereA1is symmetric and A2isanti-symmetric , i.e.,A∗
2=−A2. [Hint: Assume
there is such a decomposition and use it to find A1andA2in terms of AandA∗.
Then verify that these work.]
(15) Consider the operator D=d
dxonP5. Prove that Dis not invertible (return to the
definition p. 360) but exhibit an operator Lwhich is a right inverse, DL=I.
(16) This problem proves that Ris orthogonal if and only if Ris linear and isometric.
(a) Prove that if Ris linear and isometric, then it is orthogonal. (Trivial!).
(b) IfRis orthogonal, prove that
i)/bardblRX/bardbl=/bardblX/bardbl
ii)/angbracketleftRX, RY /angbracketright=/angbracketleftX, Y/angbracketright(Hint: Use /bardblRX−RY/bardbl2=/bardblX−Y/bardbl2)
iii)R(aX) =aRX (Hint: Prove /bardblR(aX)−aRX/bardbl2= 0 )
iv)R(X+Y) =RX+RY(Hint: Prove /bardbl“something” /bardbl2= 0 )
v)Ris linear and isometric
[Warning: If you assume linearity in b), you’ll vitiate the whole problem].
216 CHAPTER 5. MATRIX REPRESENTATION
(17) (a). Let Abe a square matrix such that A5= 0 . Verify that ( I+A)−1=I−A+
A2−A3+A4.
(b). IfA7= 0 , then ( I−A)−1= ?
(18) Consider the matrices
(a)./parenleftbiggα1
2
−1
2δ/parenrightbigg
, (b)./parenleftBigg1√
2β
γ1√
2/parenrightBigg
,
(c)./parenleftbigg0β
γ0/parenrightbigg
, (d)./parenleftbigg1β
0 2/parenrightbigg
.
For what value(s) of α, β, γ andδdo these matrices represent orthogonal transfor-
mations?
(19) IfA= ((aij)) is a square ( n×n) matrix, the trace ofAis defined as the sum of the
elements on the main diagonal, tr A: =a11+a22+...+ann. Prove
(a). tr(αA) =αA, whereαis a scalar.
(b). tr(A+B) = trA+ trB, whereBis also ann×nmatrix.
(c). tr(AB) = tr(BA) .
(d). tr(I) =?
(20) Assume that A:En→Enis anti-symmetric, A∗=−A.
(a). Prove A−Iis invertible. [By Theorem 12, it is sufficient to show ( A−I)X=
0⇒X= 0 . Use the property of Ato prove it AX =X, then /angbracketleftX, AX /angbracketright=
/bardblX/bardbl2,/angbracketleftA∗X, X/angbracketright=−/bardblX/bardbl2,and/angbracketleftX, AX /angbracketright=/angbracketleftA∗X, X/angbracketright.]
(b). IfU= (A+I)(A−I)−1, thenUis an orthogonal transformation.
(21) LetAnbe the orthogonal matrix which rotates vectors in E2through an angle of
2π/n.
(a). Find a matrix representing An(use the standard basis for E2).
(b). LetBdenote the orthogonal matrix of reflection across the x1axis (p. 382).
Show that BAb=A−1
nB. [The group of matrices generated by AnandBand all
possible products is the dihedral group of order n].
(22) Prove that the set of all orthogonal transformations of EnintoEnforms a (non-
commutative) group under multiplication.
(23) An affine transformation AX =LX+X0of a linear space into itself is called
non-singular if the linear transformation Lis non-singular. Prove that the set of
all such non-singular affine transformations form a (non-commutative) group under
multiplication.
(24) LetA= ((aij)) be a square matrix. Find all such matrices with the property that
tr(AA∗) = 0 (see Ex. 19 for the definition of the trace).
5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 217
(25) Consider the linear space
S={f(x):f(x) =a+bcosx+csinx}.
with the scalar product
/angbracketleftf, g/angbracketright=a˜a+1
2(b˜b+c˜c),
whereg(x) = ˜a+˜bcosx+ ˜csinx. Define the linear transformation R:S→Sby the
rule
(Rf)(x) =f(x+α), α real.
(a) Show that Ris an orthogonal transformation by proving that /angbracketleftRf, Rg /angbracketright=/angbracketleftf, g/angbracketright
for allf,ginS.
(b) Choose a basis for Sand exhibit a matrix eRewhich represents Rwith respect
to that basis for both the domain and target.
(26) LetA:En→Em. Prove:Ais surjective (= onto) if and only if A∗is injective (=
one to one).
(27) Define A:P3→R3by
A[p(x)] = (p(0), p(1), p(−1)) where p∈P3.
Find the matrix for this transformation with respect to the basis e1= 1e2=
(x+ 1)2, e3= (x−1)2, e4=x3forP3; and the standard basis for R3.
(b). Find the matrix representing Ausing the same basis for R3but using the basis
ˆe1= 1,ˆe2=x,ˆe3=x2and ˆe4=x3forP3.
(28) IfAandBboth map the linear space Vinto itself, and if Bis the only right
inverse ofA, AB =I, proveAis invertible. [Hint: Consider BA+B+I].
(29) LetA:En→Embe represented by the matrix (( aij)) , andB:Em→Enby ((bij)) .
If
/angbracketleftY, AX /angbracketright=/angbracketleftBY, X /angbracketright
for allX∈Enand allY∈Em, proveB=A∗. This proves the statement made in
the remark following Theorem 14.
(30) LetL:R4→R4be defined by LX= (x1,0,x3,0) , whereX= (x1,x2,x3,x4) . Find
a matrix representing Lin terms of some basis. You may use the same basis for both
the domain and the target.
5.3 Volume, Determinants, and Linear Algebraic Equations.
Often we have stated that thus and so is true if and only if a certain set of vectors are
linearly independent. But we still have no adequate criteria for determining if a set of
vectors is linearly independent. What would be an ideal criterion? One superb criteria
would be as follows. Find a function which assigns to a set of nvectorsX1,X2,...,X nin
Rna real number, with the property that this number is zero if and only if the vectors are
linearly dependent.
218 CHAPTER 5. MATRIX REPRESENTATION
There is a geometric way of solving this problem. For clarity we shall work in two
dimensions, E2. IfX1andX2are any two vectors in E2, then intuition tells us X1and
X2are linearly dependent if and only if the area of the parallelogram (see fig.) is zero.
Thus, once we define the analogue of volume for ndimensional parallelepipeds in Rn, the
appropriate criterion appears to be that a set of nvectorsX1,...,X ninEmis linearly
dependent if and only if the volume of the parallelepiped they span is zero.
The major hurdle is constructing a volume function which behaves in the manner dic-
tated by two and three dimensional intuition. Our program is to state a few (four to be
exact) desirable properties of a volume function Vfor parallelepipeds, then construct a
simpler related function - the determinant D, and observe that V=|D|(absolute value
ofD) is a volume function. This determinant function will prove useful in the theory of
linear algebraic equations.
LetX1andX2be any two vectors in R2. We define the parallelogram spanned by
X1andX2to be the set of points XinR2which have the form
X=t1X1+t2X2, 0≤t1≤1,0≤t2≤1
You can check that these points are precisely those in the parallelogram drawn above. The
volume function (really area in this case) V(X1,X2) which assigns to each parallelogram
its volume should have the properties
1.V(X1,X2)≥0.
2.V(λX1,X2) =|λ|V(X1,X2), λ scalar.
3.V(X1+X2,X2) =V(X1,X2) =V(X1,X1+X2) .
4.V(e1,e2) = 1. e 1= (1,0), e2= (0,1).
The second property states that if one side is multiplied by λ1then the volume is
multiplied by |λ|(see fig.).
The third property is more subtle. It states that the volume of the parallelogram
spanned by X1andX2is the same as the parallelogram spanned by X1andX1+X2.
This is clear from the figure since both parallelograms have the same base and height.
The last property merely normalizes the volume. It states that the unit square has
volume 1.
Our first task is to define a parallelepiped in En.
Definition: Thendimensional parallelepiped inEnspanned by a linearly independent
set of vectors X1,X2,...,X nis the set of all points XinRnof the form
X=t1X1+t2X2+···+tnXn, 0≤tj≤1.
It is a straightforward matter to write the axioms for the volume V(X1,X2,...,X n)
for thendimensional parallelepiped in En.
V-1.V(X1,X2,...,X n)≥0.
V-2.V(X1,X2,...,X n) is multiplied by |λ|if someXjis replaced by λXjwhereλ
is real.
V-3.V(X1,X2,...,X n) does not change if some Xjis replaced by Xj+Xk, where
j/negationslash=k.
V-4.V(e1,e2,...,e n) = 1 , where e1= (1,0,0,..., 0) , etc.
These axioms are amazingly simple. It is surprising that the volume function Vin
uniquely determined by them; that is, there is only one function which satisfies these axioms.
You might wonder why we did not add the reasonable stipulation that volume remains
5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 219
unchanged if the parallelepiped is subjected to a rigid body transformation. The reason
is that this axiom would be redundant, for this invariance of volume under rigid body
transformation will be one of our theorems.
The most simple way to obtain the volume function is to first obtain the determinant
functionD(X1,X2,...,X n) . We define thedeterminant functionD(X1,X2,...,X n) ofn
vectorsX1,X2,...,X ninRnby the following axioms (selected from those for V).
D-1.D(X1,X2,...,X n) is a real number.
D-2.D(X1,X2,...,X n) is multiplied by λif someXjis replaced by λXjwhereλ
is real.
D-3.D(X1,X2,...,X n) does not change if some Xjis replaced by Xj+Xk, where
j/negationslash=k.
D-4.D(e1,e2,...,e n) = 1 , where e1= (1,0,0,..., 0) etc.
Remarks :
(1) IfA= ((aij)) is a (square) n×nmatrix,
A=
a11a12···a1n
a21a22··· ·
· ·
· ·
· ·
an1an2···ann
we can consider it as being composed of ncolumn vectors A1,A2,...,An, and define
the determinant of the square matrix Ain terms of the determinant of these vectors
detA=D(A1,A2,...,An) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11a12···a1n
a21 a2n
· ·
· ·
· ·
an1··· ···ann/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle.
(2) Although we have written a set of axioms for D, it is not at all obvious that such a
function exists. Rest assured that we will prove the existence of such a function.
(3) Observe: if we define
V(X1,X2,...,X n) :=|D(X1,X2,...,X n)|, X j∈En,
thenVdoes satisfy the axioms for volume.
Granting existence of D, we derive some algebraic consequences of the axioms.
Theorem 5.23 . LetDbe a function which satisfies axiom D-1 to D-3 (not necessarily
D-4).
(1)IfXjis replaced by Xj=/summationdisplay
k/negationslash=jλkXkthenDdoes not change.
220 CHAPTER 5. MATRIX REPRESENTATION
(2)If one of the vectors Xjis zero, then D= 0.
(3)If the vectors X1,X2,...,X nare linearly dependent then D= 0. In particular
D= 0 if two vectors are equal.
(4)Dis a linear function of each of its variables, that is
D(...,λY +µZ,... ) =λD(...,Y,... ) +µD(...,Z,... )
(soDis a multilinear function).
(5)If any two vectors XiandXjare interchanged, then Dis multiplied by −1.
D(...,X i,...,X j,...) =−D(...,X j,...,X i,...)
Proof: These proofs, like the statements above, are conceptually simple but notationally
awkward. Notice that only Axioms 1-3 but not Axiom 4 will be used. We shall need this
fact shortly.
(1) We prove this only if Xjis replaced by Xj+λXk, j/negationslash=kandλ/negationslash= 0 . The general
case is a simple repetition of this until the other Xk’s are used up. It is simplest to
work backward. By Axiom 2,
D(...,X j+λXk,...,X k...) =1
λD(...,X j+λXk,...,λX k,...)
so by axiom 3 (since λXkis now a vector in D)
=1
λD(...,X j,...,λX k,...)
and axiom 2 again
=D(...,X j,...,X k,...).
(2) Write the vector Xj= 0 as 0Xjwhere 0 is now a scalar. This scalar may be brought
outsideDby axiom 2. Since Dis a real number, 0 ·D= 0 .
(3) LetXj=/summationdisplay
k/negationslash=jakXk. By part 1, Ddoes not change if Xjis replaced by Xj+/summationdisplay
k/negationslash=jλkXk.
Chooseλk=−ak. This gives a Dwith one vector zero, Xj−/summationdisplay
k/negationslash=jakXk= 0 . Thus
Dis zero by part 2.
(4) The trickiest part. Axiom 2 immediately reduced this to the special case λ=µ= 1 .
For notational convenience, let Y+Zby in the last slot. We have to prove
D(X1,X2,...,Y +Z) =D(X1,X2,...,Y ) +D(X1,X2,...,Z ).
IfX1,X2,...,X n−1(which appear in all three terms above) are linearly dependent,
we are done by part 3. Thus assume they are linearly independent. Since our linear
space Rnhas dimension n, thesen−1 vectors can be extended to a basis for Rn
5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 221
by adding one more, ˜Xn. Now we can write YandZas a linear combination of
these basis vectors
Y=a1X1+···+an−1Xn−1+an˜Xn, Z =b1X1+···+bn−1Xn−1+bn˜Xn.
Substituting this into Dwe obtain
D(X1,...,Y +Z) =D(X1,...,...,n−1/summationdisplay
1(aj+bj)Xj+ (an+bn)˜Xn).
But by part 1,
=D(X1,..., (an+bn)˜Xn)
and axiom 1 results in
= (an+bn)D(X1,..., ˜Xn).
However, again by part 1,
D(X1,...,Y ) =D(X1,...,n−1/summationdisplay
1ajXj+an˜Xn)
=D(X1,...,...,a n˜Xn) =anD(X1,..., ˜Xn).
Similarly
D(X1,...,Z ) =bnD(X1,..., ˜Xn).
Adding these two expressions and comparing them with the above, we obtain the
result.
(5) To avoid a mess, indicate only the ith andjth vectors. Our task is to prove
D(...,X i,...,X j,...) =−D(...,X j,...,X i,...).
This is clever. Watch: By the multilinearity (part 4)
D(...,X i+Xj,...,X i+Xj,...)
=D(...,X i,...,X i,..., ) +···+D(...,X i,...,X j,...)
+D(...,X j,...,X i,...) +···+D(...,X j,...,X j,...).
However part 2 states that the left side as well as the first and last terms on the right
are zero. Thus
0 =D(...,X i,...,X j,...) +D(...,X j,...,X i,...).
Transposition of one of the terms to the other side of the equality sign completes the
proof. You should also be able to fashion an easy proof of this part which uses only
the axioms directly (and uses none of the other parts of this theorem).
Instead of moving on immediately, it is instructive to compute D[X1,X2] whereX1
andX2are vectors in R2, X 1= (a,b), X 2= (c,d) . Then we are computing
D/bracketleftbigg/parenleftbigga
b/parenrightbigg
,/parenleftbiggc
d/parenrightbigg/bracketrightbigg
,
222 CHAPTER 5. MATRIX REPRESENTATION
which is, equivalently, the determinant of the matrix/parenleftbigga c
b d/parenrightbigg
.
D/bracketleftbigg/parenleftbigga
b/parenrightbigg
,/parenleftbiggc
d/parenrightbigg/bracketrightbigg
=aD/bracketleftbigg/parenleftbigg1
b
a/parenrightbigg
,/parenleftbiggc
d/parenrightbigg/bracketrightbigg
(axiom 2)
=aD/bracketleftbigg/parenleftbigg1
b
a/parenrightbigg
,/parenleftbiggc
d/parenrightbigg
−c/parenleftbigg1
b
a/parenrightbigg/bracketrightbigg
(Theorem 21 part 1)
=aD/bracketleftbigg/parenleftbigg1
b
a/parenrightbigg
,/parenleftbigg0
ad−cb
a/parenrightbigg/bracketrightbigg
(algebra)
= (ad−bc)D/bracketleftbigg/parenleftbigg1
b
a/parenrightbigg
,/parenleftbigg0
1/parenrightbigg/bracketrightbigg
(axiom 2)
= (ad−bc)D/bracketleftbigg/parenleftbigg1
b
a/parenrightbigg
−b
a/parenleftbigg0
1/parenrightbigg
,/parenleftbigg0
1/parenrightbigg/bracketrightbigg
(Theorem 21 part 1)
= (ad−bc)D/bracketleftbigg/parenleftbigg1
0/parenrightbigg
,/parenleftbigg0
1/parenrightbigg/bracketrightbigg
(algebra)
= (ad−bc)D[e1,e2] =ad−bc (axiom 4).
Thus |Area|=|(a+c)(b+d)−2bc−cd−ab|=|ad−bc|
You can indulge in a bit of analytic geometry (or look at my figure) to show that the
area of a parallelogram spanned by X1andX2is|ad−bc|. From our explicit calculation,
the existence and uniqueness of the determinant of two vectors in R2has been proved.
There are several ways to prove the general existence and uniqueness of a determi-
nant function. Our procedure is to first prove there is at most one determinant function
(uniqueness). Then we shall define a function inductively, and verify it satisfies the axioms.
By uniqueness, it must be the only function. Two interesting and important preliminary
propositions are needed.
The following lemma shows how to evaluate the determinant if all of the elements above
the principal diagonal are zeroes (that is, the determinant of a lower triangular matrix).
LEMMA : LetX1,···,Xnbe the columns of a lower triangular matrix
a11 0 0... 0
a21a22 0 ·
· · 0
· · ·
· · ·
an1an2··· ···ann
Then
D(X1,···,Xn) =a11a22···annD(e1,···,en) =a11a22···ann,
that is, the determinant of a triangular matrix is the product of the diagonal elements.
Proof: If any one of the principal diagonal elements are zero, then the determinant is
zero. For example, if ajj= 0 , then the n−j+ 1 vectors Xj,···,Xnall have their
firstjcomponents zero, and hence can span at most an n−jdimensional space. Since
n−j+ 1> n−j, these vectors must be linearly dependent. Therefore, by Theorem 21,
part 3, the determinant is zero, as the theorem asserts. [If you didn’t follow this, look at a
3×3 or 4 ×4 lower triangular matrix and think for a moment].
5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 223
If none of the diagonal elements are zero, we can carry out the following simple recipe.
The recipe gives a procedure for reducing the problem to evaluating a matrix which is zero
everywhere except along the diagonal.
First, we get all zeros to the left of a22in the second row by multiplying the second
column,X2, by−a21/a22and adding the resulting vector to X1. This gives a new
first column with i= 2, j = 1 element zero. Moreover, the new matrix has the same
determinant as the old one (Theorem 21, part 1). It looks like
a11 0 0 0
0a22 0·
˜a31a32a33·
· · · 0
· · · ·
· · · ·
˜an1an2···ann
.
Only the first column has changed. Repeat the same process to get all zeros to the left
ofa33. Thus, multiply the third column by −˜a31/a33and−a31/a33and add the result
to the first and second columns respectively. This gives a new matrix, again with equal
determinant, but which looks like
a11 0 0 0 ··· 0
0a22 0 0 0
0 0a33 0 ·
ˆa41ˆa42a43a44 ·
· · 0
· · ·
· · ·
ˆan1ˆan2an3· ·ann
.
Moving on, we gradually eliminate all of the terms to the left of the diagonal but keep
thesame diagonal ones. The final result is
a11 0···0
0a22 ·
· · 0
· · ·
· · ·
0 0 0 ann
.
It has the same determinant as the original matrix, so
D(X1,...,X n) =D(a11e1,...,a nnen)
=a11···annD(e1,...,e n),
where Axiom 2 has been used to pull out the constants. Now Axiom 4, D(e1,...,e n) = 1 ,
can be used to complete the proof. Observe that Axiom 4 is not used until the very last
step. Thus, the formula D= (something) D(e1,...,e n) depends only on Axioms 1-3. We
shall need this soon.
The above theorem shows how easy it is to evaluate the determinant of a lower trian-
gular matrix. It becomes particularly valuable when coupled with the next theorem which
224 CHAPTER 5. MATRIX REPRESENTATION
shows how the determinant of an arbitrary matrix can be reduced to that of a lower tri-
angular matrix. The reduction procedure given here is the best practical way of evaluating
a determinant . There is a peculiar criss-cross method for evaluating 3 ×3 determinants
which is taught in many high schools. Forget it. The method is not very practical and does
not generalize to 4 ×4 or larger determinants.
Theorem 5.24 . The evaluation of the determinant D(X1,...,X n)can be reduced to the
evaluation of a lower triangular matrix - and hence has the form
D= (something )D(e1,...,e n).
The proof gives a way of computing “something” in terms of the original matrix.
Remark : In the above formula, we did not utilize the fact that
D(e1,···,en) = 1
since this one step in the proof is the only place where Axiom 4 would be used, so we can
(and shall) use the fact that this result holds for any function which only satisfies Axioms
1-3.
Proof: This is just a recipe for carrying out the reduction. It essentially is a repetition
of the last part of the preceding lemma. Instead of waving our hands at the procedure, we
shall work out a representative
Example: Evaluate
D=D(X1,X2,X3,X4) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 2 −1 0
−1−2 3 1
0−1 4 −3
2 5 0 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle
by reducing it to a lower triangular determinant.
First we get all zeros to the right of the diagonal in the first row, that is, except in the
a11slot, by multiplying X1by the constants −2,1 and 0 and adding the resulting vectors
toX2,X3, andX4, respectively. We obtain
D=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 2 −1 0
−1−2 3 1
0−1 4 −3
2 5 0 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 −1 0
−1 0 3 1
0−1 4 −3
2 1 0 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 0 0
−1 0 2 1
0−1 4 −3
2 1 2 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle
Now we get all zeros to the right of the diagonal in the second row. Since the new a22ele-
ment above is zero, interchange the second and third columns (one could have interchanged
the second and fourth). This introduces a factor of −1 (by Theorem 21, part 5). Then
multiply the new second column by the constants 0 and −1
2, respectively, and add to the
last two columns, respectively. This gives
D=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 0 0
−1 0 2 1
0−1 4 −3
2 1 2 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 0 0
−1 2 0 1
0 4 −1−3
2 2 1 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 0 0
−1 2 0 0
0 4 −1−5
2 2 1 0/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle
5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 225
And on the third row, where we again want all zeros to the right of the diagonal, so multiply
the new third column by −5 and add it to the fourth column:
D=−/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 0 0
−1 2 0 0
0 4 −1−5
2 2 1 0/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=−/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 0 0
−1 2 0 0
0 4 −1 0
2 2 1 −5/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=−(1)(2)( −1)(−5) =−10,
where we have used the lemma about determinants of lower triangular matrices to evaluate
the last determinant.
Uniqueness is now elementary.
Theorem 5.25 . There is at most one function
D(X1,···,Xn), X k∈Rn,
which satisfies the 4 axioms for a determinant function.
Proof: Assume there are two such functions,
D(X1,···,Xn)
and
˜D(X1,···,Xn).
Let
∆(X1,···,Xn) =D(X1,···,Xn)−˜D(X1,···,Xn).
We shall show ∆( X1,···,Xn) = 0 for any choice of X1,···,Xn. Since both Dand ˜D
satisfy Axioms 1-4, we have
1). ∆ =D−˜Dis real valued.
2). ∆(...,λX j,...) =D(...,λX j,...)−˜D(...,λX j,...)
=λD(...,X j,...)−λ˜D(...,X j,...)
=λ∆(...,X j,...).
3). ∆(...,X j+Xk,...) =D(...,X j+Xk,...)−˜D(...,X j+Xk,...)
=D(...,X j,...−˜D(...,X j,...)
= ∆(...,X j,...), j/negationslash=k.
4). ∆(e1,...,e n) =D(e1,...,e n)−˜D(e1,...,e n) = 1−1 = 0 . Thus, ∆ satisfies the
same first three axioms but ∆( e1,...,e n) = 0 in place of Axiom 4. Because the proof of
Theorem 22 and its predecessors never used Axiom 4, we know that
∆(X1,...,X n) = (something) ∆( e1,...,e n) = 0.
Thus ∆(X1,···,Xn) = 0 for any vectors Xj.
If it exists, the determinant function is known to be unique. We intend to define the
determinant of order n, that is, of nvectors in Rn, in terms of determinants of order
n−1 . The key to such an approach is a relationship between a determinant of order nand
226 CHAPTER 5. MATRIX REPRESENTATION
determinants of order n−1 . To motivate our definition, we first examine the case n= 3
and utilize the intimate relation between determinant and volume.
LetX1,X2andX3be three vectors in R3. To find the determinant D(X1,X2,X3) ,
we can resolve one of the vectors, say X1, into its components X1=a11e1+a21e2+a31e3.
Since the determinant function is linear (Theorem 21, part 4),
D=D(X1,X2,X3) =a11D(e1,X2,X3) +a21D(e2,X2,X3) +a31D(e3,X2,X3).
How can we interpret D(a11e1,X2,X3) ,
D(a11e1,X2,X3) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11a12a13
0a22a23
0a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle?
By subtracting suitable multiples of the first column from the other two, we have
D(a11e1,X2,X3) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11 0 0
0a22a23
0a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle.
Consider the related volume function. The vectors in the last matrix span a parallelepiped
whose base is the parallelogram spanned by (0 ,a22,a32) and (0,a23,a33) , while the height
isa11. Thus, we expect the volume to be a11times the area of the base. Since the area of
the base is/vextendsingle/vextendsingle/vextendsingle/vextendsingledet/vextendsingle/vextendsingle/vextendsingle/vextendsinglea22a23
a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle, we hope
/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11 0 0
0a22a23
0a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=a11/vextendsingle/vextendsingle/vextendsingle/vextendsinglea22a23
a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle.
except possibly for a factor of ±1 . This last formula is the connection between determinants
of order three and those of order two.
Notice that the determinant on the right in the last equation is obtained from that of
D=D(X1,X2,X3) by deleting both the first row and first column. It is called the 1,1
minor ofD, and written D11. More generally, the i,jminorDijofDis the determinant
obtained by deleting the ith row and jth column of D. IfDis of order n, then each
Dijis of order n−1 .
In this notation, we expect from the expansion of D(X1,X2,X3) that
D(X1,X2,X3) =±?a11D11±?a21D21±?31D31,
or/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11a21a13
a21a22a23
a31a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=±?a11/vextendsingle/vextendsingle/vextendsingle/vextendsinglea22a23
a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle±?a21/vextendsingle/vextendsingle/vextendsingle/vextendsinglea12a13
a32a33/vextendsingle/vextendsingle/vextendsingle/vextendsingle±?a31/vextendsingle/vextendsingle/vextendsingle/vextendsinglea12a13
a22a23/vextendsingle/vextendsingle/vextendsingle/vextendsingle.
where ? indicates our doubt as to the signs. Explicit evaluation of both sides (using
Theorem 22) reveals that the correct sign pattern is + ,−,+ .
Having examined this special case (and the 4 ×4 case too), we are tentatively led to
Suspicion (Expansion by Minors). If D(X1,...,X n) is a determinant function, that is, if
it satisfies the axioms, then
D(X1,X2,...,X n) =n/summationdisplay
i=1(−1)i+jaijDij, (5-3)
5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 227
whereXj= (a1j,a2j,...,a nj) .
For the case n= 3,j= 1 this is the formula we found above. To verify that the
formula is correct, we must verify that the function satisfies our axioms for a determinant.
The reasoning goes as follows: we know exactly what determinants of order two are by a
previous computation, so the formula gives a candidate for the determinant of order three,
which in turn gives a candidate for a determinant function of order four, and so on. Thus,
by induction, let us assume that determinants of order k−1 are known. We must prove
Theorem 5.26 . The previous function D(X1,...,X k)defined by the above formula is a
determinant function, that is, it satisfies the axioms.
Proof: 1).D(X1,...,X k) is real valued since, by our induction hypothesis, each of the
Dij, determinants of order k−1 , is real valued.
2).D(...,λX l,...) =λD(...,X l,...) . There are two cases. Ifl=j, thenλXj
means that a1j,a2j,..., is multiplied by λ. Thus
D(...,λX j,...) =n/summationdisplay
i=1(−1)i+jλaijDij=λD(...,X j,...),
so the axiom is satisfied. Ifl/negationslash=j, then some vector Xother than Xjis multiplied by λ,
so
D(...,λX l,...,X j,...) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle···λa1l···a1j···
···λa2la2j···
· ·
· ·
· ·
···λaklakj···/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle
SinceDijis formed by deleting the ith row and jth column of D, andl/negationslash=j, one
column in minor Dij will have the factor λappearing in it. By the induction hypothesis,
the factor can be pulled out of each one, and hence from any linear combination of them.
Because the expansion formula for Dis a linear combination of the minors, the axiom is
verified in this case too.
3). Omitted. This one is just plain messy. If you don’t care to try the general case for
yourself, at least try the case n= 3 and verify it there.
4). To prove D(e1,...,e n) = 1 . Of the coefficients a1j,a2j,...,a nj, onlyajj/negationslash= 0 ,
andajj= 1 . Thus D(e1,...,e n) = (−1)j+jajjDjj=Djj. But by the induction hypoth-
esis,Djj= 1 since it has only ones on its main diagonal and zero elsewhere. Therefore
D(e1,...,e n) = 1 , as desired.
This theorem completes (except for one segment) the proof that a unique determinant
function exists. The uniqueness was proved directly, while the existence was obtained from
the known existence of 2 ×2 determinant functions (the simpler case of 1 ×1 determi-
nants could also have been used) and proving inductively that a candidate for the n×n
determinant function does satisfy the axioms.
Emerging from the jungle of the existence proof, we are fully equipped with the powerful
determinant function and the associated volume function. It will be relatively simple to
prove the remaining theorems involving determinants. The trick in most of them is to make
clever use of the fact that the determinant function is unique. We shall expose this trick in
its bare form.
228 CHAPTER 5. MATRIX REPRESENTATION
Theorem 5.27 . Let ∆(X1,...,X n)be a function of nvectors in Rnwhich satisfies
axioms 1-3 for the determinant. Then for every set of vectors X1,...,X n
∆(X1,...,X n) = ∆(e1,...,e n)D(X1,...,X n).
Thus, the function ∆differs from Donly by a constant multiplicative factor, which is the
number ∆assigns to the unit matrix (geometrically, the unit cube) in Rn.
Proof: If ∆(e1,...,e n) = 1 , then ∆ satisfies Axiom 4 also, so by the uniqueness theorem,
it must be Ditself. If ∆( e1,...,e n)/negationslash= 1 , consider
˜D(X1,...,X n) :=D(X1,...,X n) = ∆(X1,...,X n)
1−∆(e1,...,e n).
Note that the denominator is a fixed scalar which does not depend on X1,...,X n. It is a
mental calculation to verify that ˜Dsatisfies all of Axioms 1-4. Therefore ˜D(X1,...,X n) :=
D(X1,...,X n) by uniqueness. Solving the last equation for ∆( X1,...,X n) yields the
formula.
ConsiderD(X1,...,X n) . IfB= ((bij)) is a square n×nmatrix representing a linear
transformation from RntoRn, how are D(X1,...,X n) andD(BX 1,BX 2,...,BX n)
related? The answer to this question is vital if we are to find how volume varies under a
linear transformation B. IfA= ((aij)) is the matrix whose columns are X1,...,X n, and
C= ((cij)) is the matrix whose columns are BX 1,BX 2,...,BX n, thenC=BA[since,
for example, c11—the first element in the vector BX 1—is
c11=b11a11+b12a21+b13a31+···+b1nan1.]
BecauseD(X1,...,X n) = detAandD(BX 1,...,BX n) = detC, our question becomes
one of relating det C= det(BA) to detA. The result is as simple as one could possibly
expect.
Theorem 5.28 . IfAandBare twon×nmatrices, then
det(BA) = (detB)(detA) = (detA)(detB) = det(AB)
or, ifX1,...,X nare the column vectors of A, then this is equivalent to
D(BX 1,...,BX n) =D(Be1,Be 2,...,Be n)D(X1,...,X n)
(since the matrix whose columns are Be1,...,Be nis justB).
Proof: Let ∆(X1,...,X n) :=D(BX 1,...,BX n) . This function clearly satisfies Axiom
1. We shall verify Axioms 2 and 3 at the same time.
∆(...,λX j+µXk,...) =D(...,B (λXn+µXk),...)
BecauseBis a linear transformation, we have
=D(...,λBX j+µBX k,...).
By the linearity of D(Theorem 21, part 4)
=λD(...,BX j,...) +µD(...,BX k,...).
5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 229
Ifj/negationslash=k, then the vector BXkin the second term on the right also appears as another
column in the same determinant. Hence the second term vanishes. Thus if j/negationslash=k,
∆(...,λX j+µXk,...) =λD(...,BX j,...).
The special case µ= 0 shows Axiom 2 holds for ∆ , while the case λ=µ= 1 verifies
Axiom 3. Therefore ∆ satisfies Axioms 1-3. Applying the preceding Theorem (25), we
have
∆(X1,...,X n) = ∆(e1,...,e n)D(X1,...,X n).
By definition, ∆( e1,...,e n) :=D(Be1,...,Be n) . Substitution verifies our formula. The
commutativity
(detB)(detA) = (detA)(detB)
follows from the fact that det Aand detBare real numbers - which do commute under
multiplications.
Corollary 5.29 . IfAis an invertible matrix, then
det(A−1) =1
detA.
Proof: SinceAA−1=I, and detI= 1 , we find
(detA)(detA−1) = det(AA−1) = detI= 1.
Ordinary division completes the proof.
Our next theorem is also a corollary, but because of its importance, we call it
Theorem 5.30 . The vectors X1,...,X ninRnare linearly independent if and only if
D(X1,...,X n)/negationslash= 0.
Proof: ⇐IfD(X1,...,X n)/negationslash= 0 , then the vectors X1,...,X nare linearly independent,
since if they were dependent, then D= 0 by part 3 of Theorem 21.
⇒. IfX1,...,X nare linearly independent vectors in Rn, then the Corollary to
Theorem 12 (p. 364) shows that the matrix Awhose columns are the Xjis invertible.
LetA−1be its inverse. From the computation in the corollary preceding this theorem,
(detA)(detA−1) = 1.
Thus the real number det Acannot be zero. The equivalent form of our theorem is also a
consequence of the Corollary to Theorem 12.
Example: (cf. p. 157, Ex. 1b). Are the vectors
X1= (0,1,1), X 2= (0,0,−1), X 3= (0,2,3)
linearly dependent? We compute the determinant
D(X1,X2,X3) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle0 0 0
1 0 2
1−1 3/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle.
230 CHAPTER 5. MATRIX REPRESENTATION
If we knew that “the determinant of a matrix was equal to the determinant of its adjoint”
(a true theorem to be proved below), then taking the adjoint we get a matrix with one
column zero 0 which gives D= 0 . Since the quoted theorem is not yet proved, we proceed
differently and reduce our 3 ×3 determinant to 2 ×2 determinants expanding by minors
(p. 411). The simplest column to use is the second.
/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle0 0 0
1 0 2
1−1 3/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle= (−1)1+2 0/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 2
1 3/vextendsingle/vextendsingle/vextendsingle/vextendsingle+ (−1)2+2 0/vextendsingle/vextendsingle/vextendsingle/vextendsingle0 0
1 3/vextendsingle/vextendsingle/vextendsingle/vextendsingle+ (−1)3+2(−1)/vextendsingle/vextendsingle/vextendsingle/vextendsingle0 0
1 2/vextendsingle/vextendsingle/vextendsingle/vextendsingle
=/vextendsingle/vextendsingle/vextendsingle/vextendsingle0 0
1 2/vextendsingle/vextendsingle/vextendsingle/vextendsingle= 0·2−1·0 = 0
by the explicit formula for evaluating 2 ×2 determinants. Thus D= 0 so the vectors
X1,X2,X3are linearly dependent.
That nice theorem we could have used in the above example is our next target.
Theorem 5.31 . IfAis ann×nmatrix, then
detA∗= detA.
Proof: LetA1,...,Anbe the columns of AandB,...,Bnits rows,
A=
a11a12···a1n
a21··· ···a2n
· ·
· ·
· ·
an1··· ···ann
B1
}B2
·
·
·
}Bn.
Consider the function
D(B1,...,Bn) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11···an1
a12 ·
· ·
· ·
a1nann/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle}A1
·
·
·
}An= detA∗.
since the rows of Aare the columns of A∗. Let us define a new function
ˆD(A1,...,An) :=D(B1,...,Bn).
Our task is to verify that ˆD(A1,...,An) satisfies all of Axioms 1-4. Then by uniqueness
detA∗:=ˆD(A1,...,An) =D(A1,...,An) = detA.
(1) ˆD(A1,...,An) is a real number since det A∗, the determinant of the matrix A∗is
a real number.
5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 231
(2) We must show ˆD(...,λAj,...) =λˆD(...,Aj,...) , that is,
/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11a2j···an1
·
·
λa1jλa2j... λa nj
·
·
·
a1n... ... a nn/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=λ/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11···anl
·
·
a1j···anj
a1n···ann/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle
(a fact we only know so far if a column is multiplied by a scalar). Trick: observe that
jth row...
1 0 0 0
0 1 0
0···λ1 0
0 0...0 1
a11···anl
· ·
· ·
a1j···anj
· ·
· ·
a1n···ann
=
a11···anl
· ·
· ·
a1j···anj
· ·
· ·
a1n···ann
The matrix on the left is the identity matrix Iexcept for a λin itsjth row and
jth column. Its determinant is λ(since you can factor λfrom thejth column and
are left with the identity matrix). By Theorem 26, the determinant of the product
on the left is λˆD(A1,...,An) while the right is ˆD(A1,...,An) , proving ˆDsatisfies
Axiom 2.
(3) The proof of Axiom 3 involves a similar trick. We have to show ˆD(...,Aj+Ak,...) =
ˆD(...,Aj,...) wherej/negationslash=k, that is, to show
/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11 ···an1
· ·
· ·
a1j+a1k···anj+ank
· ·
· ·
· ·
a1n ···ann/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11···an1
·
·
a1j···anj
· ·
· ·
· ·
a1n···ann/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle,j/negationslash=k.
232 CHAPTER 5. MATRIX REPRESENTATION
Observe that
1 0 ··· 0
0 1 ···
··· ··· ···
0 0 ··· 0
0 0 1
a11···an1
·
·
a1j···anj
·
·
·
a1n···ann
=
a11 ···an1
· ·
· ·
a1j+a1k···anj+ank
·
·
·
a1n ···ann
,
where the matrix on the left is the identity matrix with an extra 1 in the jth row,kth
column. Since the determinant of this matrix is one (check by a mental computation),
the rule for the determinant of a product of matrices shows that Axiom 3 is satisfied.
(4) Easy, for
ˆD(e1,...,e n) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0 ··· ··· 0
0 1 ··· ··· 0
··· ··· 1 0
0··· ··· 0 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle=D(e1,...,e n) = 1.
This verification of the four Axioms coupled with the remarks at the beginning of the
proof completes the proof.
Corollary 5.32 . The column operations of Theorem 21 are also valid as row operations.
Proof: Every row operation on a matrix A(like adding two rows) can be split up to : i)
takeA∗so the rows become columns, ii) carry out the operation on the column of A∗and
iii) take the adjoint again. Since the determinant does not change under these operations,
we are done.
Corollary 5.33 . IfRis an orthogonal matrix then
detR=±1.
Proof: IfRis orthogonal, then R∗R=Iby Theorem 19 (p. 383). Thus,
a= detI= det(R∗R) = (detR∗)(detR) = (detR)2,
where Theorems 25 and 27 were invoked once each. Now take the square root of both sides.
The orthogonal matrices
R1=/parenleftbigg1 0
0 1/parenrightbigg
andR2=/parenleftbigg0 1
1 0/parenrightbigg
,
for which det R1= 1 and det R2=−1 show that both signs are possible. IfdetR=−1 ,
then the orthogonal transformation has not only been a rotation but also a reflection . The
transformation given by R2is
a figure goes here
5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 233
which can be thought of as the composition (product) of a rotation by +900followed
by a reflection (mirror image). In fact, R2may be factored into ˆR˜R=r2, where
/parenleftbigg1 0
0−1/parenrightbigg/parenleftbigg0 1
−1 0/parenrightbigg
=/parenleftbigg0 1
1 0/parenrightbigg
=R2.
Pictorially
a figure goes here
Our theorems about determinants also imply the following valuable result about volume.
Theorem 5.34 . LetX1,...,X nspan a parallelepiped QinEnand the matrix Amap
EnintoEn. Then the volume is magnified by |detA|, that is,
V[AX 1,...,AX n] =|detA|V[X1,...,X n].
If we denote the image of QbyA(Q), then this theorem reads
Vol[A(Q)] =|detA|Vol[Q].
Proof: [We should first prove that there is at most one volume function Vsatisfying its
four axioms. Since V:=|D|is a volume function, assume there is another volume function
V∗and define ˜D(X1,...,X n) by
˜D(X1,...,X n) :=/braceleftBigg
V∗(X1,...,X n)D(X1,...,X n)
|D(X1,...,X n)|ifD/negationslash= 0
0 if D= 0.
It is simple to check that ˜Dsatisfies the axioms for a determinant. By uniqueness, ˜D=D.
Solving the last equation, we find V∗(X1,...,X n) =|D(X,...,X n)| ≡V(X1,...,X n) , so
the volume function is also unique.]
The theorem is easily proved. Since V=|D|, an application of Theorem 26 tells us
that
V[AX 1,...,AX n] =|D(AX 1,...,AX n)|
=|D(Ae1,...,Ae n)| |D(X1,...,X n)|
=|detA|V[X1,...,X n].
Done.
Corollary 5.35 . Volume is invariant under an orthogonal transformation.
V(RQ) =V(Q)
Proof: IfRis an orthogonal transformation, |detR|= 1 .
Remark 1 . Since we eventually want to define the volume of suitable sets by approximating
the sets by parallelepipeds, this theorem will allow us to conclude the same results about how
the volume of some set changes under a linear transformation in general and an orthogonal
transformation in particular.
234 CHAPTER 5. MATRIX REPRESENTATION
Remark: 2 We define the determinant of a linear transformation Lwhich maps Rninto
Rnas the determinant of a matrix which represents L. This definition makes it mandatory
to prove: “the determinant of two different matrices which represent L(different because
of a different choice of bases) are equal.” However the theorem is an immediate consequence
of the following fact we never proved: “if AandBare matrices which represent the same
linear transformation Lwith respect to different bases then there is a nonsingular matrix
Csuch thatB=CAC−1.” The matrix Cis the matrix expressing one set of bases vectors
in terms of the other bases. Using this theorem, we find
detB= det(CAC−1) = (detC)(detA)(detC−1) = detA.
How does volume change under a translation T, TX =X+X0? A little thought is
needed. Imagine a parallelepiped Qspanned by X1,...,X n. The crux of the matter is to
realize that the parallelepiped has the origin as one of its vertices and X1,...,X nat the
others. Under the translation T, not only do the Xj’s get translated through X0, but so
does the origin , 0→X0, X 1→X1+X0, X 2→X2+X0, etc.
a figure goes here
In terms of free vectors, the edge from 0 to Xjbecomes the edge from X0toXj+X0
(see figure). Thus the free vector representing this edge is ( Xj+X0)−X0, that is, it is
stillXj! This motivates the
Definition: The volume of a parallelepiped is defined to be the volume of the parallelepiped
after translating one vertex to the origin.
Theorem 5.36 . The change in volume of a parallelepiped Qunder an affine transforma-
tionAX=LX+X0, Llinear, is given by:
Vol[A(Q)] =|detL|Vol[Q].
In particular, volume is invariant under a rigid body transformation (for then Lis an
orthogonal transformation).
Proof: The affine transformation may be factored into A=TL, a linear transformation
followed by a translation (p. 380). Since Lchanges volume by |detL|while translation
preserves the volume, the net result is a change by |detL|as claimed.
a) Application to Linear Equations
What have our geometrically motivated determinants in common with the determinants
of high school fame - where they were used to solve systems of linear algebraic equations?
Everything, for they are the same. Since determinants are defined only for square matrices,
they are applicable to linear algebraic equations only when there are the same number of
equations as unknowns. At the end of this section, we shall make some remarks about the
case when the number of equations and unknowns are not equal.
Consider the system of equations
a11x1+···+a1nxn=y1
a21x1+···+a2nxn=y2
......
an1x1+···+annxn=yn,
5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 235
which we can write as
x1A1+···+xnAn=Y,
where Ajis thejth column of the matrix A= ((aij)) andYis the obvious column
vector. The problem is to find numbers x1,...,x nsuch thatx1A1+···+xnAn=Y,
whereYis given.
Theorem 5.37 . LetA= ((aij))be a square n×nmatrix and Ya given vector. The
system of linear algebraic equations AX=Ycan always be solved for Xif and only if
detA/negationslash= 0. This can be rephrased as, Ais invertible if and only if detA/negationslash= 0.
Proof: LetAjbe thejth vector of A. Each Ajis a vector in Rn. If detA/negationslash= 0 ,
then the An’s are linearly independent by Theorem 27, p. 417. But since they are linearly
independent and there are nof them, A1,···,An, they must span Rn. Thus, any Y∈Rn
can be written as a linear combination of the Aj’s. The numbers x1,···,xnare just the
coefficients in this linear combination.
Conversely, if the equations AX=Ycan be solved for anyY∈Rn, then the vec-
torsA1,···,Anspan Rn. But ifnvectors span Rn, these vectors must be linearly
independent, so det A/negationslash= 0 , again by Theorem 27, page 417.
Theorem 5.38 . LetAbe a square matrix. The system of homogeneous equations AX=
0has a non-trivial solution if and only if detA= 0.
Proof: By Theorem 27, Page 417, det A= 0 if and only if the column vectors A1,...,An
are linearly dependent. Now if the column vectors A1,...,Anare linearly dependent,
then there are numbers x1,...,x n, not all zero, such that x1A1+...+xnAn= 0 . The
vectorX= (x1,...,x n) is then a non-trivial solution of AX= 0 . Conversely, if there is
a non-trivial solution of AX= 0 , then x1A1+···+xnAn= 0 , so the Aj’s are linearly
dependent. Hence det A= 0 .
In contrast to the above theorems which give no hint of a procedure for finding the
desired vector X, the next theorem gives an explicit formula for the solution of AX=Y.
Theorem 5.39 (Cramer’s Rule). Let A= ((aij))be a square n×nmatrix with columns
A1,...,An. Assume detA/negationslash= 0. Then for any vector Y, the solution of AX=Yis
x1=D(Y,A2,...,An)
D(A1,...,An), x 2=D(A1,Y,A3,...,An)
D(A1,...,An)
...
xn=D(A1,...,An−1,Y)
D(A1,...,An).
For example, in detail, the formula for x2is
x2=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11y1a13···a1n
............
an1ynan3···ann/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle
/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11a12a13···a1n
...............
an1an2an3···ann/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle.
236 CHAPTER 5. MATRIX REPRESENTATION
Proof: A snap. Since det A/negationslash= 0 , by Theorem 31 we know a solution X= (x1,...,x n)
exists. Thus x1A1+···+xnAn=Y. Let us obtain the formula for x2as a representative
case. Observe that
D(A1,Y,A3,···,An) =D(A1,x1A1+···+xnAn,A3,···,An).
SinceDis multilinear, we can expand the above to
=x1D(A1,A1,A3,···,An) +xnD(A1,A2,A3,···,An) +···+xnD(A1,An,A3···An).
Now all of these determinants, except the second one, vanishes since each has two identical
columns (part 5 of Theorem 21, page 400). Thus
D(A1,Y,A3,·,An) =x2D(A1,A2,···,An).
Because det A=D(A1,···,An)/negationslash= 0 , we can divide to find the desired formula for x2.
Done.
Remark: This elegant formula is mainly of theoretical use. It is not the most efficient
procedure for solving such equations. That honor belongs to the method of reducing to
triangular form which was outlined in the proof of Theorem 22. To be more vivid, if
Cramer’s rule were used to solve a system of 26 equations, approximately (23 + 1)! ≈1028
multiplications would be required. Reduction to triangular form, on the other hand, would
only require about (1 /3)(23)3≈6000 multiplications. Think about that.
For non-square matrices, determinants are not applicable. Given a vector Y, one would
still like a criterion to determine if one can solve AX=Y, that is, one would like a criterion
to see ifY∈R(A) .
Theorem 5.40 . LetL:V1→V2be a linear operator. Then
R(L)⊥=N(L∗);
or equivalently (for finite dimensional spaces)
R(L) =N(L∗)⊥.
Proof: IfX∈V, andY∈R(L)⊥, then for all X
0 =/angbracketleftY, LX /angbracketright=/angbracketleftL∗Y, X/angbracketright.
This means L∗Yis orthogonal to all X, consequently, L∗Y= 0 , soY∈N(L∗) . The
converse is proved by observing that our steps are reversible.
Application . For what vectors Y= (y1,y2,y3) can you solve the equations
2x1,+3x2=y1
x1−x2=y2
x1+ 2x2=y3 ?
5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 237
If the equations are written as AX=Y, then by the above theorem Y∈R(A) if and
only ifY⊥N(A∗) . Let us find a basis for N(A∗) . This means solving the homogeneous
equationsA∗Z= 0 ,
2z1+z2+z3= 0
3z1−z2+ 2z3= 0.
If we letz1=α, and solve the resulting equations for z2andz3, we find that z3=
−5α/3 andz2=−11α/3 . Consequently, all vectors Z∈(A∗) have the form Z=
(3α,−11α,−5α) . A basis for N(A∗) ise= (3,−11,−5) . Therefore, Y⊥N(A∗) if and only
if 3y1−11y2−5y3= 0 . By the above reasoning, the equation AX=Ycan be solved for
only these Y’s.
Remark: The use of Theorem 34 as a criterion for finding if Y∈R(L) is much more
valuable in infinite dimensional spaces, for it quite often turns out that N(L∗) is still finite
dimensional while R(L) is infinite dimensional. For more on these ideas, see page 389,
Exercise 12 and page 501 Exercises 27- 29.
Exercises
(1) Evaluate the following determinants as you see fit:
a)./vextendsingle/vextendsingle/vextendsingle/vextendsingle7 3
2−1/vextendsingle/vextendsingle/vextendsingle/vextendsingle, b)./vextendsingle/vextendsingle/vextendsingle/vextendsingle1
25
−3 4/vextendsingle/vextendsingle/vextendsingle/vextendsingle.
c)./vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle−10−2 3
−3 2 1
5 0 −1/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle, d)./vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle53 17 29
36 12 39
69 23 75/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle, e)./vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 2 0 1
1 3 4 0
0 1 −5 6
1 2 3 4/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle
f)./vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle2 1 1 1
1 2 1 1
1 1 2 1
1 1 1 2/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle, g)./vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea1 0 0 0
b1 0 0 0
c0 0 1 −b
c0 0 1 −a
d e 1f g/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle
[Answers: a) −13 , b) 17, c) −14 , d) 6, e) 5, g) −(b−a)2].
(2) IfAandBare the matrices whose respective determinants appear in #1 a) and
b), compute det( AB) by first finding AB. Compare with (det A)(detB) .
(3) a). Use Cramer’s rule (Theorem 33) to solve the equation AX=Y, where A is given
below. Then observe you have computed A−1, so exhibit it.
A=
1 1 1
2−3−1
4 9 1
. [A−1=1
30
6 8 2
−6−3 3
30−5−5
].
b). Use the formula for A−1to solve the equations
AX=YwhereY= (1,2,0).
238 CHAPTER 5. MATRIX REPRESENTATION
(4) a). Find the volume of the parallelepiped QinE3which is spanned by the vectors
X1= (1,1,1), X 2= (2,−1,−3) andX3= (4,1,9) . [Answer: Volume = 30].
b). The matrix A,
A=
−10−2 3
−3 2 1
5 0 −1
−(cf. #1,c)
maps E3into itself. Find the volume of the image of Q, that is, the volume of A(Q) .
[Answer: 420].
(5) LetB=A−λIwhereAis a square matrix. The values λfor whichBis singular
are called the eigenvalues ofA. Find the eigenvalues for
a).A=/parenleftbigg3 2
2−1/parenrightbigg
, b). A=/parenleftbigg3 2
1−1/parenrightbigg
.
c).A=/parenleftbigga b
c d/parenrightbigg
.
[Hint: IfBis singular, then 0 = det B= det(A−λI) . Now observe that det( A−λI)
is a polynomial in λ. The answer to c) is λ=1
2(a+d±/radicalbig
(a+d)2−4(ad−bc))].
(6) For what value(s) of αare the vectors
X1= (1,2,3), X 2= (2,0,1), X 3= (0,α,−1)
linearly dependent?
(7) IfX1,X2,X3andY1,Y2,Y3are vectors in R3, prove that
D[X1,X2,X3]−D[Y1,Y2,Y3]
=D[X1−Y1,X2,X3] +D[X1,X2−Y2,X3] +D[X1,X2,X3−Y3].
[Hint: First work out the corresponding formula for the 2 ×2 case.]
(8) Here you shall compute the derivative of a determinant if the coefficients of A= ((aij))
depend ont,aij(t) . LetX1(t),...,X n(t) be the vectors which constitute the columns
ofA. The problem is to compute
dD(t)
dt=d
dtD[X1,...,X n](t) =d
dt/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea11(t),···, a 1n(t)
·
·
·
an1(t),···, a nn(t)/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle
a). Use Exercise 7 (generalized to n×nmatrices) to show
D(t+ ∆t)−D(t)≡D[X1(t+ ∆t), X 2(t+ ∆t),...]−D[X1(t), X 2(t),...]
=n/summationdisplay
j=1D[X1(t),...,X j−1(t), Xj(t+ ∆t)−Xj(t),Xj+1(t+ ∆b),...]
5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 239
[Hint: Do the cases n= 2 andn= 3 first].
b). Use part a to show that
dD
dt= lim
∆t→0/bracketleftbiggD(t+ ∆t)−D(t)
∆t/bracketrightbigg
=n/summationdisplay
j=1D[X1,...,X j−1,dXj
dt,Xj+1,...,X n],
so the derivative of a determinant is found by taking the derivative one column at a
time and adding the result.
(9) Letu1(t) , andu3(t) be solutions of the differential equation
u/prime/prime+a1(t)u/prime+a0(t)u= 0.
Consider the Wronski determinant
W(u1,u2)(t) :=/vextendsingle/vextendsingle/vextendsingle/vextendsingleu1(t)u2(t)
u/prime
1(t)u/prime
2(t)/vextendsingle/vextendsingle/vextendsingle/vextendsingle
(a) Use Exercise 8 to prove
dW
dt=−a1(t)W.
(b) Consequently, show
W(t) =W(t0) exp/braceleftbigg
−/integraldisplayt
t0a1(s)ds/bracerightbigg
.
(c) Apply this to show that if the vectors ( u1(t),u/prime
1(t)) and (u2(t),u/prime
2(t)) are linearly
independent at t=t0, then they are always linearly independent.
(d) Letu1(t)...,u n(t) be solutions of the differential equation
u(n)+an−1(t)u(n−1)+···+a1(t)u/prime+a0(t)u= 0.
Consider the Wronski determinant of u1,...,u n
W(u1,...,u n) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingleu1u2···un
u/prime
1u/prime
2···u/prime
n
·
·
·
u(n−1)
1a(n−)
2 ···u(n−1)
n/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle
ProvedW
dt=−an−1(t)W,
so again
W(t) =W(t0) exp/braceleftbigg
−/integraldisplayt
t0an−1(s)ds/bracerightbigg
.
240 CHAPTER 5. MATRIX REPRESENTATION
(e) Use part d) to conclude that the nvectors
(u1,u/prime
1,...,u(n−1)
1),(u2,u/prime
2,...,u(n−1)
2,),···(un,u/prime
n,...,u(n−1)
n)
(where the ujare solutions of the O.D.E.) are linearly independent for all tif
and only if they are so at t=t0.
(10) A matrix Aisupper (lower) triangular if all the elements below (above) the main
diagonal are zero,
A=
a11a12···an
0a22··· ·
0 0 ·
0 0 ···ann
.
IfAis upper (or lower) triangular, prove again that
detA=a11a22...a nn.
by expanding by minors. What is the relation of this result to the exercise ( #4 , p.
157) on echelon form?
(11) LetX1,...,X nbe vectors in Rnand let ˆD(X1,...,X n) be a real valued function
which has properties 1 and 4 of Theorem 21. Thus ˆDis skew-symmetric, and is
linear in each of its columns. Prove ˆDnecessarily satisfies Axioms 2 and 3 for the
determinant, and conclude that
ˆD(X1,...,X n) =kD(X1,...,X n),
where the constant k=D(e1,...,e n) .
(12) Letu1(t),...,u n(t) be sufficiently differentiable functions ( Cn−1is enough). Define
the Wronskian as in Exercise 9 part d. Prove that if the functions u1,...,u nare
linearly dependent, then W(t)≡0 . Thus, if W(t0)/negationslash= 0 , the functions are linearly
independent in any interval containing t0. [Do nottry to apply the result of Exercise
9 for it is not applicable].
(13) (a) If Iis then×nidentity matrix, evaluate det( λI) whereλis a constant.
(b) IfAis ann×nmatrix, prove
det(λA) =λndetA.
(c) IfAorBaren×nmatrices, is
det(A+B)?= detA+ detB?
Proof or counterexample.
(14) For what value of αdoes the system of equations
x+ 2y+z= 0
−2x+αy+ 2z= 0
x+ 2y+ 3z= 0
have more than one solution?
5.3. VOLUME, DETERMINANTS, AND LINEAR ALGEBRAIC EQUATIONS. 241
(15) A matrix is nilpotent if some power of it is zero, that is, AN= 0 for some positive
integerN. Prove that if Ais nilpotent, then det A= 0 .
(16) (a) Solve the systems of equations
i)x+y= 1 ,x−.9y=−1
and
ii)x+y= 1 ,x−1.1y=−1 ,
and compare your solutions, which should be almost the same.
(b) Solve the systems of equations
i)x+y= 1 ,x+.9y=−1,
and
x+y= 1 ,x+ 1.1y=−1.
and again compare your solutions. Explain the result in terms of the theory in
this section.
(c) Consider the solution of the systems of equations
x+y= 1
x+αy=−1
as the point where the lines x+y= 1 andx+αy=−1 intersect. Sketch
the graph of these lines for αnear−1 and then for αnear +1 . Use these
observations to again explain the phenomena in parts a) and b).
(17) Let ∆ nbe then×ndeterminant of a matrix with a’s along the main diagonal and
b’s on the two “off diagonals” directly above and below the main diagonal. Thus
∆5=/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglea b 0 0 0
b a b 0 0
0b a b 0
0 0b a b
0 0 0b a/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle.
(a) Prove ∆ n=a∆n−1−b2∆n−2.
(b) Compute ∆ 1and ∆ 2by hand. Then use the formula to compute ∆ 3and ∆ 4.
(c) Ifa2/negationslash= 4b2, can you show
∆n=1√
a2−4b2
/parenleftBigg
a+√
a2−4b2
2/parenrightBiggn+1
−/parenleftBigg
a−√
a2−4b2
2/parenrightBiggn+1
?
Later, we shall give a method for obtaining this directly from the equation of part a).
[p. 522-523].
(18) Prove Part 5 of Theorem 21 using only the axioms and no other part of Theorem 21.
242 CHAPTER 5. MATRIX REPRESENTATION
(19) Apply the result of Exercise 12 on page 389. Try to prove the following. Ais asquare
matrix.
a). dim N(A) = dim N(A∗).
Thus, the homogeneous equation AX= 0 has the same number of linearly indepen-
dent solutions as does the equation A∗Z= 0 .
b). LetZ1,...,Z kspanN(A∗) . Then the inhomogeneous equation
AX=Y
has a solution, that is, Y∈R(A) , if and only if
/angbracketleftZj, Y/angbracketright= 0, j = 1,2,...,k.
In other words, the equation AX=Yhas a solution if and only if Yis orthogonal
to the solutions of the homogeneous adjoint equation.
c). Consider the system of linear equations
2x−3y+z= 1
−3x+ 2y−4z=α
x−4y−2z=β.
LetAbe the coefficient matrix. Find a basis for N(A∗) . [Answer: dim N(A∗) = 1
andZ1= (2,1,−1) is a basis]. For what value(s) of the constants α,β can you solve
the given system of equations? [Answer: There is a solution if and only if β−α= 2 .]
Find a solution if α= 1 andβ= 3 .
d). Repeat part c) for the system of equations
x−y= 1
x−2y=−1
x+ 3y=α.
[Answer: dim N(A) = 1 and Z1= (−5,4,1) is a basis. There is a solution if and
only ifα=−1 ].
(20) Use the result of Exercise 12 to prove that each of the following sets of functions are
linearly independent everywhere.
a)u1(x) = sinx, u 2(x) = cosx
b)u1(x) = sinnx, u 2(x) = cosmx, wheren/negationslash= 0.
c)u1(x) =ex, u2(x) =e2x, u3(x) =e3x.
d)u1(x) =eax,u2(x) =ebx, u3(x) =ecx, wherea,b, andcare distinct numbers.
e)u1(x) = 1, u2(x) =x, u 3(x) =x2, u4(x) =x3
f)u1(x) =ex, u2(x) =e−x, u3(x) =xex, u4(x) =xe−x.
5.4. AN APPLICATION TO GENETICS 243
5.4 An Application to Genetics
A mathematical model is developed and solved. Although this particular model will be
motivated by genetics, the resulting mathematical problem also arises in sociometrics and
statistical mechanics as well as many other places. In the literature you will find these
mathematical ideas listed under the title Markov chains .
Part of the value you should glean from our discourse is insight into the process of
going from vague qualitative phenomena to setting up a quantitative model. One part of
this scientific process we shall not have time to investigate in detail is the very important step
of comparing the quantitative results with experimental data. Furthermore, we shall never
delve into the fertile realm of generalizing our accumulated knowledge to more complicated
- as well as more interesting and realistic - situations.
In bisexual mating, the genes of the resulting offspring occur in pairs, one gene in
each pair being contributed by each parent. Consider the simplest case of a trait which is
determined by a single pair of genes, each of which is one of two types gandG. Thus, the
father contributes Gorgto the pair, and the mother does likewise. Since experimental
results show that the pair Ggis identical to the pair gG, the offspring has one of the three
pairs
GG Gg gg.
The geneGdominatesgif the resulting offspring with genetic types GGandGg“appear”
identical but both are different from gg. In this case, an individual with genetic type GG
is called dominant , while the types ggandGgare called recessive andhybrid , respectively.
An offspring can have the pair GG(resp.gg) if and only if bothparents contributed
a gene of type G(resp. g) while the combination Ggoccurs if either parent contributed
Gand the other g. A fundamental assumption we shall make is that a parent with genetic
typeabcan only contribute a gene of type aor of typeb. This assumption ignores such
things as radioactivity as a genetic force. Thus, a dominant parent, GGcanonlycontribute
a dominant gene, G, a recessive parent, gg, can only contribute g, and a hybrid parent
Ggcan contribute eitherGorg(with equal probability). Consequently, if two hybrids
are mated, the offspring has probability1
2of getting Gorgfrom each parent, so the
probability of his having genetic type GGofggis1
4each, while the probability of having
genetic type Ggis1
2.
We introduce a probability vector V= (v1,v2,v3) , withv1representing the probability
of being genetic type GG,v2of being type Gg, andv3of being type gg. Thus for an
offspring of two hybrid parents ,V= (1
4,1
2,1
4) . Observe that, by definition of probability,
0≤vj≤1, j= 1,2,3 , andv1+v2+v3= 1 (since with probability one - certainty - the
offspring is either GG,Gg, orgg).
Consider the issue of mating an individual whose genetic type is unknown with an
individual of known genetic type (dominant, hybrid or recessive). To be specific, assume
the known person is of dominant type. Then the following matrix of transition probabilities
D=
1 1/2 0
0 1/2 1
0 0 0
describes the probability of the offspring’s genetic type in the following sense: if the unknown
parent had genetic type V0(soV0= (1,0,0) if unknown was dominant, V0= (0,1,0) if
244 CHAPTER 5. MATRIX REPRESENTATION
hybrid, and V0= (0,0,1) if recessive), then
V1=DV 0,
is the probability vector of the offspring. For example, if the unknown parent was hy-
brid, then V1=DV 0= (1
2,1
2,0) . Thus the offspring can, with equal likelihood, be either
dominant or hybrid, but cannot be recessive.
Notice that the matrix Dembodies the fact that one of the parents is dominant.
If the individual of unknown genetic type were crossed with an individual of hybrid
type, then the corresponding matrix His
H=
1
21
40
1
21
21
2
01
41
2
,
while if the person of unknown type were crossed with the individual of recessive type, then
R=
0 0 0
11
20
01
21
.
It is of interest to investigate the question of genetic stability under various circum-
stances. Say we begin with an individual of unknown genetic type and cross it with a
dominant individual, then cross that offspring with another dominant individual, and so
on, always mating the resulting offspring with a dominant individual. Let Vnrepresent the
genetic probability vector for the offspring in the nth generation. Then
Vn=DVn−1=D2Vn−2=···=DnV0,
whereV0is the unknown vector for the initial parent (of unknown genetic type). Without
knowingV0, can we predict the eventual ( n→ ∞ ) genetic types of the offspring? Intu-
itively, we expect that no matter what the type of the initial parent, the repeated mating
with a dominant individual will produce a dominant strain. The question we are asking is,
does lim
n→∞Vnexist, and if so, what is it?
Assume for the moment that the limit does exist and denote it by V. ThenV=DV
since
V= lim
n→∞Vn= lim
n→∞Vn+1= lim
n→∞DVn=D( lim
n→∞Vn) =DV
Armed with the equation DV =V, we can solve linear equations for the vector V=
(v1,v2,v3)
v1+1
2v2+ 0 =v1
0 +1
2v2+v3=v2
0 + 0 + 0 = v3.
Clearlyv1=v2=v3= 0 is a trivial solution. A non-trivial one can be found by transposing
thevj’s to the left side and solving. We find v1= 1,v2= 0,v3= 0(v1= 1 since
v1+v2+v3= 1 ). Thus, ifthe limitVnexists, the limit must be V= (1,0,0) . In
genetic terms, this sustains our feeling that the offspring will eventually become genetically
dominant.
5.4. AN APPLICATION TO GENETICS 245
But does the limit exist? To prove it does, we must show for any probability vector
V0= (v1,v2,v3) , wherev1+v2+v3= 1 , that the limit
lim
n→∞Vn= lim
n→∞DnV0,
exists and equals V= (1,0,0) . By evaluating D,D2,andD3explicitly, we are led to
guess
Dn=
1 1−1
2n1−1
2n−1
01
2n1
2n−1
0 0 0
,
which is then easily verified using mathematical induction. Thus
Vn=DnV0=
v1+ (1 −1
2n) + (1 −1
2n−1)v3
0 +1
2nv2 +1
2n−1v3
0 + 0 + 0
=
v1+v2+v3−1
2n(v2+ 2v3)
1
2nv2+1
2n−1v3
0
Sincev1+v2+v3= 1 , we find
Vn=DnV0=
1
0
0
+1
2n(v2+ 2v3)
−1
1
0
.
It is now clear that the limit as n→ ∞ does exist, and is V= (1,0,0) . Consequently, if we
begin with a random individual (you) and mate that individual and the successive offspring
with a dominant gene bearer, then the resulting generations will tend to all dominant
individuals. Moreover, the process proceeds exponentially because the “damping factor” is
essentially1
2for each generation (see above formula).
Were there enough time, you would see a second application of matrices to the special
theory of relativity. Given your knowledge of linear spaces, it is possible to present an
elegant exposition of the theory. The Lorentz transformation would appear as an orthogonal
transformation - a rotation - in world space orMinkowski’s space as it is often called. This
is a four dimensional space three of whose dimension are those of ordinary space, while
the fourth dimension is an imaginary (i=√−1) time dimension. Goldstein’s Classical
Mechanics contains the topic. Regrettably, he does not begin with the Michelson - Morley
experiment but rather plunges immediately into mathematical technicalities.
Exercises
1. If you begin with an individual of unknown genetic type and cross it with a hybrid
individual and then cross the successive offspring with hybrids, does the resulting strain
approach equilibrium? If so, what is it?
2. Same as 1 but you mate an individual of unknown type with a recessive individual.
3. Beginning with an individual of unknown genetic type, you mate it with a dominant
individual, mate the offspring with a hybrid , mate that offspring with a dominant, and
continue mating alternate generations with dominants and hybrids respectively. Does the
246 CHAPTER 5. MATRIX REPRESENTATION
resulting strain approach equilibrium? If so, what is it? (You will need to define equilibrium
to cope with this problem. There are several reasonable definitions.)
4. a). The city Xhas found that each year 5% of the city dwellers move to the suburbs,
while only 1% of the suburbanites move to the city. Assuming the total population of the
city plus suburb does not change, show that the matrix of transition probabilities is
P=/parenleftbigg.95.01
.05.99/parenrightbigg
,
where a vector V= (v1,v2) = (proportion of people in city, proportion of people in suburb).
b). Given any initial population distribution V, does the population approach an
equilibrium distribution? If so, find it.
5. A long queue in front of a Moscow market in the Stalin era sees the butcher whisper to
the first in line. He tells her “Yes, there is steak today.” She tells the one behind her and so
on down the line. However, Moscow housewives are not reliable transmitters. If one is told
“yes”, there is only an 80% chance she’ll report “yes” to the person behind her. On the
other hand, being optimistic, if one hears “no”, she will report “yes” 40% of the time. If
the queue is very long, what fraction of them will hear “there is no steak”? [This problem
can be solved without finding a formula for Pn, although you might find it a challenge to
find the formula].
5.5 A pause to find out where we are
.
We all know the homily about the forest and the trees. The next few pages are about
the forest.
In the beginning we introduced dead linear spaces with their algebraic structure (Chap-
ter II). Then we investigated the geometry induced by defining an inner product on a linear
space and saw how easily many of the results in Euclidean geometry generalize (Chapter
III).
Our next step was to consider mappings, linear mappings, between linear spaces (Chap-
ter IV). Not much could be said in general, so we began investigating a particular case, linear
maps between finite dimensional spaces. Two important special cases of this
L:R1→Rn,
and
L:Rn→R1,
were treated before the general case,
L:Rn→Rm.
A key theorem which facilitates the theory of linear mappings between finite dimensional
spaces is the representation theorem (page 374): every such map can be represented as a
matrix.
What next? There are two equally reasonable alternatives:
5.5. A PAUSE TO FIND OUT WHERE WE ARE 247
(A) We can continue with linear maps ,
L:V1→V2,
and consider the case where V1orV2, or both are infinite dimensional. The general
theory here is in its youth and still undeveloped. Only one of the sources of difficulty
is that a generalization of the representation theorem (page 374) remains unknown
- except for some special cases. Thus, many special types of mappings have to be
investigated individually. We shall consider only one type of linear mapping between
infinite dimensional spaces, those defined by linear differential operators (Chapter VI
and Chapter VII, Section 3).
(B) The second alternative is to continue our study of mappings between finite dimensional
spaces, only now switch to non linear mappings. This theory should parallel the
transition in elementary calculus from the analytic geometry of straight lines,
f(x) =a+bx,
that is, affine mappings, to genuine non linear mappings, as
f(x) =x2−7√x
or
f(x) =x3−esinx.
You recall, one important idea was to approximate the graph of a function y=f(x)
at a pointx0by its tangent line at x0, since forxnearx0, the curve and the tangent
line there approximately agree. For example, one easily proves that at a maximum or
minimum, the tangent line must be horizontal, f/prime= 0 .
In generalizing this to functions of several variables,
Y=F(X) =F(x1,···,xn),
the role of the derivative at X0is assumed by the affine map,
A(X) =Y0+LX,
which is tangent to FatX0. Thus, linear algebra appears as the natural extension
of analytic geometry to higher dimensional spaces. See Chapters VII - IX for this.
248 CHAPTER 5. MATRIX REPRESENTATION
Chapter 6
Linear Ordinary Differential
Equations
6.1 Introduction
.
Adifferential equation is an equation relating the values of a function u(t) with the
values of its derivatives at a point,
F(t,u(t),du
dt,...,dnu
dtn) = 0 (6-1)
The order of the equation is the order, n, of the highest derivative which appears. For
example, the equations/parenleftbiggd2u
dt2/parenrightbigg3
−7du
dt+t2u2−sint= 0
du
dt−tsinu2= 0
are of order two and one respectively. A function u(t) is a solution of the differential
equation if it has at least as many derivatives as the order of the equation, and if substitution
of it into the equation yields an identity. Thus, the equation
/parenleftbiggdu
dt/parenrightbigg2
+u2= 1
has the function u(t) = sintas a solution, since for all t
/parenleftbiggd
dtsint/parenrightbigg2
+ (sint)2= 1.
A differential equation (1) for the unknown function u(t) islinear if it has the form
Lu:=an(t)dnu
dtn+an−1(t)dn−1
dtn−1+···+a0(t)u= 0 (6-2)
You should verify that this coincides with the notion of a linear operator used earlier. Equa-
tion (2) is sometimes called linear homogeneous to distinguish it from the inhomogeneous
equation
Lu=f(t), (6-3)
249
250 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
that is
an(t)dnu
dtn+···+a0(t)u=f(t). (6-4)
The subject of this chapter is linear ordinary differential equations with variable coeffi-
cients (to distinguish them from the special case where the aj’s are constants). This opera-
torLdefined by (2) has as its domain the set of all sufficiently differentiable functions— n
derivatives is enough. These functions constitute an infinite dimensional linear space. Thus,
the differential operator Lacts on an infinite dimensional space, as opposed to a matrix
which acts on a finite dimensional space.
Differential equations abound throughout applications of mathematics. This is because
most phenomena are described by laws which relate the rate of change of a function - the
derivative - at a given time (or point) to the values of the function at that same time.
For example, we have seen that at any time the acceleration of a harmonic oscillator is
determined by its position and velocity at the same time,
¨u=−µ˙u−ku.
When confronted by a differential equation, your first reaction should be to attempt to
find the solution explicitly. We were able to do this for linear constant coefficient equations
(Chapter 4, Section 2). One of the main goals of this chapter is to show you how to solve
as many linear ordinary differential equations as possible. However, it is naive to expect
to solve an arbitrary equation which crops up in terms of the few functions we know:
xα,ex,logx,sinx, and cosx. In fact, to even solve the elementary equation
du
dx=1
x,
appearing in elementary calculus, we were forced to define a new function as the solution
of this equation
u(x) = logx+c
and obtain the properties of this function and its inverse exdirectly from the differential
equation. Many many functions arise which cannot be expressed in terms of the few elemen-
tary functions we know and love. Most of these functions - like Bessel’s functions, elliptic
functions, and hypergeometric functions, arise directly because they are the solutions of
differential equations nature has forced us to consider.
How do we know these strange sounding functions are solutions of the differential equa-
tions? Well, we somehow prove a solution exists and then simply give a name to the solution
- much as babies are given names at birth. Furthermore, as is the case with babies, their
actual “names” are the least important aspect.
To summarize briefly, we shall solve as many equations as we can. For the remaining
ones (which include most equations), we shall attempt to describe a few of the main prop-
erties so that if one arises in your work, you will have a place to begin the attack. Later
on, we shall again return to the more complicated situation of nonlinear equations. Much
less can be said there. Only very few general results are known.
Lest you get the wrong idea, we shall cover but a fraction of the known theory for just
linear ordinary differential equations. In the next chapter, we shall only look at one partial
differential equation (the wave equation for a vibrating violin string). The general theory
there is too complicated to allow discussion for more than one particular equation.
6.1. INTRODUCTION 251
Exercises
1. Assume there exists aunique functionE(x) which satisfies the following differential
equation for all xand satisfies the initial condition
du
dx=u, u (0) = 1.
(a) Use the “chain rule” and uniqueness to prove for any a∈R
E(x+a) =E(a)E(x)
[Hint: Prove ˜E(x) :=E(x+a) is also a solution of the equation. Then apply the
uniqueness to the function ˜E(x)/E(a) ].
(b) Prove
E(−x) =1
E(x).
(c) Prove for any x
E(nx) = [E(x)]n, n∈Z.
In particular, show
E(n) = [E(1)]n, n∈Z
and
E(1
m) = [E(1)]1/m, m ∈Z+
(d) Prove
E(n
m) = [E(1)]n/m, n∈Z, m ∈Z+
[Thus, the function E(x) is defined for all rational x=n
mas the number E(1) to
the powern/m . SinceE(x) is continuous (even differentiable by definition, we can
extend the last formula to irrational xby continuity: if rjis a sequence of rational
numbers converging to the real number x(which may or may not be rational) then
by continuity
E(x) = lim
j→∞E(rj) = lim
j→∞[E(1)]rj=E(1)x.
Consequently, E(x) is the familiar exponential function ex].
2. Find the general solutions of the following equations by any method you can.
(a)du
dx−2u= 0
(b)du
dx=x2+ sinx
(c)/parenleftbigdu
dx/parenrightbig2+ 4u2= 1
(d)du
dx=x
u+1
(e)du
dx=x2eu
(f)d2u
dx2+ 3du
dx−4u= 4
252 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
6.2 First Order Linear
.
Except for those differential equations which can be solved by inspection, the next most
simple equation is one which is linear and first order, the homogeneous equation
du
dx+a(x)u= 0, (6-5)
and the inhomogeneous equation
du
dx+a(x)u=f(x). (6-6)
The homogeneous equation can be solved by first writing it in the form
1
udu
dx=−a(x)
and then integrating both sides
logu(x) =−/integraldisplayx
a(s)ds+C1.
Thus
u(x) =Ce−/integraldisplayx
a(s)ds (6-7)
is the solution of equation (4) for any constant C. In the very special case a(s)≡constant,
the solution does have the form found earlier (Chapter 4, Section 2) for a linear equation
with constant coefficients.
How can we integrate the inhomogeneous equation (5)? A useful device is needed.
Multiply both sides of this equation by an unknown function q(x)
q(x)du
dx+q(x)a(x)u=q(x)f(x),
Ifwe can find q(x) so that the left side is a derivative,
q(x)du
dx+q(x)a(x)u=d
dx(q(x)u), (6-8)
then the equation reads
d
dx(q(x)u) =q(x)f(x),
which can be integrated immediately,
q(x)u(x) =/integraldisplayx
q(s)f(s)ds+c, (6-9)
and then solved for u(x) by dividing by q(x) .
Thus, the problem is reduced to finding a q(x) which satisfies (7). Evaluating the right
side of (7), we find
qdu
dx+qau=udq
dx+qdu
dx,
6.2. FIRST ORDER LINEAR 253
soq(x) must satisfy
dq
dx=q(x)a(x).
It is easy to find a function q(x) which satisfies this - for it is a homogeneous equation of
the form (4). Therefore
q(x) =eRxa(t)dt,
the reciprocal of the solution (6) to the homogeneous equation, does satisfy (7). Notice
we have ignored the arbitrary constant factor in the solution since all we want is any one
functionq(x) for (7).
Now we can substitute into (8) to find the solution of the inhomogeneous equation
u(x) =1
q(x)/integraldisplayx
q(s)f(s)ds+c
q(x), (6-10)
whereq(x) is given by the formula at the top of the page. If it makes you happier,
substitute the expression for q(x) into (9) to obtain the messy formula. We have left some
room.
a figure goes here
Examples: 1.du
dx+2
xu= (1 +x3)17, x/negationslash= 0.
First,
q(x) = exp(/integraldisplayx2
sds) = exp(2 ln x) = exp(lnx2) =x2.
Thusd
dx(x2u) =x2(1 +x3)17.
Integrating both sides we find
x2u(x) =(1 +x3)18
54+C.
Therefore
u(x) =1
54(1 +x3)18
x2+C
x2, x/negationslash= 0.
2.du
dx+ 2xu=x
First,
q(x) = exp(/integraldisplayx
2sds) = expx2.
Thus,
d
dx(ex2u) =ex2x.
Integrating both sides, we find
ex2u(x) =1
2ex2+C,
so
u(x) =1
2+Ce−x2.
This formula could have been guessed much earlier since we know the general solution of
the inhomogeneous equation can be expressed as the sum of a particular solution to that
254 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
equation plus the general solution of the homogeneous equation. The particular solution
u0(x) =1
2can be obtained by inspection of the D.E.
Let us summarize our results.
Theorem 6.1 . Consider the first order linear inhomogeneous equation
Lu:=du
dx+a(x)u=f(x).
Ifa(x)andf(x)are continuous functions, the equation has the solutions
u(x) = ˜u(x)/integraldisplayxf(s)
˜u(s)ds+C˜u(x) (6-11)
where
˜u(x) = exp( −/integraldisplayx
a(s)ds)
is a non-trivial solution of the homogeneous equation. Moreover, if we specify the initial
conditionu(x0) =α, then the solution which satisfies this initial condition is unique.
Proof: Theexistence follows from the explicit formula (9) or (10) and from the fact that
a continuous function is always integrable.
Uniqueness . This will be quite similar to the proof carried out in Chapter 4. If u1(x)
andu2(x) are two solutions of the inhomogeneous equation Lu=f, with the same initial
conditions, then the function
w(x) :=u1(x)−u2(x)
satisfies the homogeneous equation
Lw:=w/prime+a(x)w= 0,
and is zero at x0,
w(x0) =u1(x0)−u2(x0) = 0.
Our task is to prove w(x)≡0 . Multiply the equation (20) by w(x) . Then
ww/prime=−a(x)w2,
or1
2d
dxw2=−a(x)w2.
Sincea(x) is continuous, for any closed and bounded interval [ A,B] there is a constant k
(depending on the interval) such that −a(x)≤kfor allx∈[A,B] . Consequently,
1
2d
dxw2≤kw2,
ord
dxw2−2kw2≤0.
Now we need an important identity which can be verified by direct computation: for any
smooth function g, and any constant α,g/prime+αg=e−αx(eαxg)/prime. We apply this to the
above inequality with g=w2andα=−2kto conclude that
e2kxd
dx[e−2kxw2]≤0.
6.2. FIRST ORDER LINEAR 255
Becausee2kxis always positive, by the mean value theorem this inequality states that
e−2kxw2is a decreasing function of x. Thus
e−2kxw2(x)≤e−2kx0w2(x0), x≥x0,
or
w2(x)≤e2k(x−x0)w2(x0), x≥x0.
But sincew(x0) = 0 andw2(x)≥0 this means that
0≤w2(x)≤0.
Thereforew(x)≡0x≥x0.
To provew(x)≡0 forx≤x0, merely observe that the equation (11) has the same
form ifxis replaced by −x. Thus the above proof applies and shows w(x)≡0 forx≤x0
too.
Remark: Although a formula has been exhibited for the solution, this does not mean that
the integrals which occur can be evaluated in terms of elementary functions. These integrals
however can be at least evaluated approximately using a computer if a numerical result is
needed.
Exercises
(1) . Find the solution of the following equations with given initial values
(a)u/prime+ 7u= 3, u(1) = 2
(b) 5u/prime−2u=e3x, u(0) = 1.
(c) 3u/prime+u=x−2x2, u(−1) = 0.
(d)xu/prime+u= 4x3+ 2, u(1) = −1.
(e)u/prime+ (cotx)u=ecosx+ 1, u(π
2) = 0.[/integraltext
cotxdx= ln(sinx)] .
(2) . The differential equation
Ldu
dt+Ru=Esinωt, L,R,E constants
arises in circuit theory. Find the solution satisfying u(0) = 0 and show that it can
be written in the form
u(t) =ωEL
R2+ω2L2e−Rt/L+E√
R2+ω2L2sin(ωt−α)
where
tanα=ωL
R.
(3) Bernoulli’s equation is
u/prime+a(x)u=b(x)uk, k a constant.
256 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
(a) Use the substitution v(x) =u(x)1−kto transform this nonlinear equation to the
linear equation
v/prime+ (1−k)a(x)v= (1−k)b(x).
(b) Apply the above procedure to find the general solution of
u/prime−2exu=exu3/2.
(4) . Consider the equation
u/prime+au=f(x),
whereais a constant, fis continuous in the interval [0 ,∞] , and |f(x)|< M for
allx.
(a) Show that the solution of this equation is
u(x) =e−axu(0) +e−ax/integraldisplayx
0eatf(t)dt
(b) Prove (if a/negationslash= 0 )
/vextendsingle/vextendsingleu(x)−e−axu(0)/vextendsingle/vextendsingle≤M
a[1−e−ax].
(5) (a) Show the uniqueness proof yields the following stronger fact. If u1(x) andu2(x)
are both solutions of the same equation
u/prime+a(x)u=f(x)
but satisfy different initial conditions
u1(x0) =α, u 2(x0) =β,
then
|u1(x)−u2(x)| ≤ek(x−x0)|α−β|, x≥x0
for allx∈[A,B] , where −a(x)< k in the interval. Thus, if the initial values
are close, then the solutions cannot get too far apart.
(b) Show that if a(x)≤A < 0 , whereAis a constant, then as x→0 any two
solutions of the same equation - but with possibly different initial values - tend
to the same function.
(6) . Show that the differential equation
y/prime=a(x)F(y) +b(x)G(y)
can be reduced to a linear equation by the substitution
u=F(y)/G(y) oru=G(y)/F(y)
if (FG/prime−GF/prime)/Gor (FG/prime−GF/prime)/F, respectively, is a constant. Use this substitution
to again solve Bernoulli’s equation.
6.2. FIRST ORDER LINEAR 257
(7) . LetS={u∈C/prime:u(0) = 0 }, and define the operator LfromStoCby
Lu=u/prime+u.
ProveLis injective and R(L) =C.
(8) . Set up the differential equation and solve. The rate of growth of a bacteria culture
at any time tis proportional to the amount of material present at that time. If there
was one ounce of culture in 1940 and 3 ounces in 1950, find the amount present in the
year 2000. The doubling time is the interval it takes for a given amount to double.
Find the doubling time for this example.
(9) . Find the general solution of x2u/prime+ 3xu= sinx.
(10) . Assume that a body decreases its temperature u(t) at a rate proportional to the
difference between the temperature of the body and the temperature Tof the sur-
rounding air. A body originally at a temperature of 1000is placed in air which is
kept at a temperature of 500. If at the end of one hour the temperature of the body
has fallen 200, how long will it take for the body to reach 600?
(11) . Here is one simple mathematical model governing economic behavior. Think of
yourself as a widget manufacturer for now. Let
i)S(t) be the supply of widgets available at time t. This is the only function you
can control directly.
ii)P(t) be the market price of a widget at time t.
iii)D(t) is the demand for widgets at time t—the number of widgets people want
to buy at time t. You cannot control this given function.
It has been found that the market price P(t) changes at a rate proportional to the
difference between demand and supply,
dP
dt=k(D(t)−S(t)),
wherek>0 is a fixed constant.
You decide to vary the supply so that it is a fixed constant S0plus an amount
proportional to the market price,
S(t) =S0+αP(t), α> 0.
(a) Set up the differential equation for S(t) in terms of the given function D(t) and
solve it.
(b) Analyze the solution and give an argument making it plausible that the market
for widgets behaves roughly in this way. What criticisms can you make of the
model?
(c) How does the market behave if the demand increases for a long time and then
levels off at some constant value, D(t) =D(t1) fort≥t1? A qualitative
description of S(t) andP(t) is called for here. In particular, say whether price
increases without bound (bringing the evils of inflation) or whether it, too, levels
off.
258 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
(12) It is found that a juicy rumor spreads at a rate proportional to the number of people
who “know”. If one person knows initially, t= 0 , and tells one other person by
the next day, t= 1 , approximately how long does it take before 4000 people know?
Analyze the mathematical model as t→ ∞ and state why it is, in fact, the wrong
model. (The question to ask yourself is, “how long will it take before everyone even
remotely concerned knows?”). The same mathematical model applies to the spreading
of contagious diseases - and many other similar phenomena.
6.3 Linear Equations of Second Order
In this section we will consider a portion of the general theory of second order linear
O.D.E.’s, with variable coefficients,
Lu:=a2(x)d2u
dx2+a1(x)du
dx+a0(x)u=f(x).
Although allof the results obtained generalize immediately to linear equation of order n,
only the special case n= 2 will be treated. This special case has the advantage of clearly
illustrating the general situation and supplying proofs which generalize immediately - while
avoiding the inevitable computation complexities inherent in the general case.
There are three parts:
A). a review of the constant coefficient case,
B). power series solutions, and
C). the general theory.
Whereas the first two parts are concerned with obtaining explicit formulas for the solutions,
the last resigns itself to some statements which can be made without finding the solution
explicitly.
a) A Review of the Constant Coefficient Case.
Here we have the operator
Lu:=a2u/prime/prime+a1u/prime+a0u, (6-12)
wherea0,a1, anda2are constants. In order to solve the homogeneous equation
Lu= 0,
the function eλxis tried. Substitution yields
L(eλx) = (a2λ2+a1λ+a0)eλx=p(λ)eλx. (6-13)
labeleq:13 The polynomial p(λ) is called the characteristic polynomial forL. Ifλ1is a
root of this polynomial, p(λ1) = 0 , then u1(x) =eλ1xis a solution of the homogeneous
equationLu= 0 . Ifλ2is another root of this polynomial λ1/negationslash=λ2, u 2(x) =eλ2xis
another solution. Then every function of the form
u(x) =Au1(x) +Bu2(x) =Aeλ1x+Beλ2x, (6-14)
whereAandBare constants, is a solution of the homogeneous equation. The uniqueness
theorem showed that every solution of Lu= 0 is of the form (14).
6.3. LINEAR EQUATIONS OF SECOND ORDER 259
Ifthe two roots ofp(λ)coincide , then a second solution is u2(x) =xeλ1x, and every
function of the form
u(x) =Au1(x) +Bu2(x) =Aeλ1x+Bxeλ1x(6-15)
whereAandBare constants, is a solution of the homogeneous equation. Again the
uniqueness theorem showed that every solution of Lu= 0 is of the form (15).
In both (14) and (15), the constants AandBcan be chosen to find a unique function
u(x) which satisfies the homogeneous equation
Lu= 0
as well as the initial conditions
u(x0) =α, u/prime(x0) =β,
whereαandβare specified constants.
It turns out that the inhomogeneous O.D.E.
Lu=f,
wherefis a given continuous function, can always be solved once two linearly independent
solutionsu1andu2of the homogeneous equation Luj= 0 are known. Since the procedure
for solving the inhomogeneous equation also works if the coefficients in the differential
operatorLare not constant, it is described later in this section in the more general situation
(p. 487-8, Theorem 8). Somewhat simpler techniques can be used for the constant coefficient
equation if the function fis a linear combination of functions of the form xkerx, wherek
is a nonnegative integer and ris some real or complex constant (cf. Exercise 6, p. 300).
Because both sin nxand cosnxare of this form, Fourier series can be used to supply a
solution for any function fwhich has a convergent Fourier series (cf. Exercise 13, p. 303).
Section 5 of this chapter contains an interesting generalization of the theory for constant
coefficient ordinary differential operators to operators which are “translation invariant”.
b) Power Series Solutions
.
Many ordinary differential equations (linear and nonlinear) can be solved by merely
assuming the solution can be expanded in a power series u(x) =/summationtextcnxn, and plugging into
the differential equation to find the coefficients cn. A simple example illustrates this.
Example: Solveu/prime/prime−2xu/prime= 0 with the initial conditions u(0) = 1,u/prime(0) = 0 .
Solution: We try
u(x) =c0+c1x+c2x2+···+cnxn+···.
Then
u/prime(x) =c1+ 2c2x+ 3c3x2+···+ncnxn−1+···
so
2xu/prime(x) = 2c1x+ 4c2x2+···+ 2ncnxn+···
260 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
Also
u/prime/prime(x) = 2c2+ 2·3c3x+ 3·4c4x2+·+ (n−1)ncnxn+···
Addingu/prime/prime−2xu/prime−uand collecting like powers of xwe find that 0 = u/prime/prime−2xu/prime−u=
[2c2−c0] + [2·3c3−2c1−c1]x+ [3·4c4−4c2−c2]x2
+···+ [(k+ 1)(k+ 2)ck+2−2kck−ck]xk+···
If the right side, a Taylor series, is to be zero (= the left side), then the coefficient of each
power ofxmust vanish because the only convergent Taylor series for zero is zero itself.
The coefficient of
x0is 2 c2−c0
x1is 6 c3−3c1
x2is 12 c4−5c2
xkis ( k+ 1)(k+ 2)ck+2−(2k+ 1)ck
Equating these to zero we find that
c2=c0
2, c 3=c1
2, c 4=5c2
12=5
24c0
and, more generally,
ck+2=2k+ 1
(k+ 2)(k+ 1)ck. (6-16)
Thus, for this example eeven is some multiple of c0whilecoddis some multiple of c1.
Sinceu(0) =c0andu/prime(0) =c1, the constants c0andc1are determined by the initial
conditions.
c0= 1, c 1= 0.
Consequently, all of the odd coefficients c3,c5,... vanish, while
c2=1
2, c 4=5
24, c 6=3
10c4=1
16, c 8=...,
so the first few terms in the series for u(x) are
u(x) = 1 +1
2x2+5
24x4+1
16x6+··· (6-17)
We should investigate if this formal power series expansion converges. Using (16), the ratio
of successive terms in the series for u(x) is
/vextendsingle/vextendsingle/vextendsingle/vextendsingleck+2xk+2
ckxk/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle(2k+ 1)
(k+ 2)(k+ 1)x2/vextendsingle/vextendsingle/vextendsingle/vextendsingle
Therefore the ratio test shows the formal power series actually converges for all x. By
Theorem 16, p. 82, the series can be differentiated term by term and does satisfy the
equation.
Although the computation is lengthy, the series (17) is a solution. Since there is no
way of finding the solution in terms of elementary functions, we must be contented with the
power series solution. You have seen (Chapter 1, Section 7) how properties of a function
can be extracted from a power series definition.
This example is typical.
6.3. LINEAR EQUATIONS OF SECOND ORDER 261
Theorem 6.2 . If the differential equation
a2(x)u/prime/prime+a1(x)u/prime+a0(x)u= 0
has analytic coefficients about x= 0, that is, if the coefficients all have convergent Taylor
series expansions about x= 0, and ifa2(0)/negationslash= 0, then given any initial values
u(0) =α, u/prime(0) =β,
there is a unique solution u(x)which satisfies the equation and initial conditions. Moreover,
the solution is analytic about x= 0 and converges in the largest interval [−r,r]in which
the series for a1/a2anda0/a2both converge.
Outline of Proof . There are two parts: i) find a formal power series u(x) =/summationtextcnxn, and ii)
prove the formal power series converges. Since explicit formulas can be found for the cn’s
(cf. Exercise 30a) the first part is true. Proof of the second part is sketched in the exercises
too (Exercise 30b).
From the explicit formulas mentioned above for the cn’s, it is clear there is at most
oneanalytic solution. But because the general uniqueness proof (p. 510, Theorem 9) states
there is at most one solution which is twice differentiable and since u(x) is certainly such
a function - the uniqueness of u(x) among all twice differentiable functions follows as soon
as Theorem 9 is proved.
The restriction a2(0)/negationslash= 0 which was made in Theorem 3 is very important. If a2(0) = 0
then the differential equation
a2(x)u/prime/prime+a1(x)u/prime+a0(x)u= 0
is degenerate at x= 0 because the coefficient of the highest order derivative vanishes there.
Then the point x= 0 is called a singularity of the differential equation. A simple example
illustrates the situation. The function u(x) =x5/2satisfies the differential equation
4x2u/prime/prime−15u= 0
and the initial conditions u(0) = 0,u/prime(0) = 0 . However u(x)≡0 is also a solution. Thus
it will be impossible to prove any uniqueness theorem at x= 0 for this equation. Perhaps
the singular nature of this equation at x= 0 is more vivid if the equation is written as
u/prime/prime−15
4x2u= 0.
Although the possibility of a uniqueness result is ruled out for equations with singular-
ities, it is important to be able to find the non-zero solutions of these equations, important
because many of the equations which arise in practice do happen to have singularities
(Bessel’s equation, Legendre’s equation, the hypergeometric equation, ...). In all of the
commonly occurring cases, the coefficients a0(x), a1(x) , anda2(x) ,
a2u/prime/prime+a1u/prime+a0u= 0,
are analytic functions. Thus the only obstacle to applying Theorem 3 is the condition
a2/negationslash= 0 . We persist, however, in the belief that a power series, or some modification of it,
262 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
should work. The modification must allow for such solutions as u(x) =x3/2which do not
have Taylor expansions about x= 0 . Undoubtedly the most naive candidate for a solution
is to try
u(x) =xρ∞/summationdisplay
n=0cnxn, (6-18)
whereρmay be any real number. The particular choice ρ= 3/2, c0= 1,c1=c2=c3=
...= 0 does yield the function u(x) =x3/2. It turns out that (18) is usually the correct
guess.
Again, we turn to an example. Bessel’s equation of ordern,
x2u/prime/prime+xu/prime+ (x2−n2)u= 0,
which arises in the study of waves in a two dimensional circular domain, like those on
tympani, in a tea cup, or on your ear drum. Let us find a solution to Bessel’s equation of
order one,
x2u/prime/prime+xu/prime+ (x2−1)u= 0 (6-19)
This equation does have a singularity at the origin, x= 0 . Ifuhas the form (18), then
u(x) =∞/summationdisplay
n=0cnxn+ρ,
u/prime(x) =∞/summationdisplay
n=0(n+ρ)cnxn+ρ−1,
and
u/prime/prime(x) =∞/summationdisplay
n=0(n+ρ)(n+ρ−1)cnxn+ρ−2.
Substituting this into the differential equation (19), we find
∞/summationdisplay
n=0(n+ρ)(n+ρ−1)cnxn+p+∞/summationdisplay
n=0(n+ρ)cnxn+ρ
+∞/summationdisplay
n=0cnxn+ρ+2−∞/summationdisplay
n=0cnxn+ρ= 0. (6-20)
We must equate the coefficients of successive powers of xto zero. The lowest power of x
which appears is xρ, the nextxρ+1, and so on.
xρ:ρ(ρ−1)c0+ρc0−c0= 0
xρ+1: (ρ+ 1)ρc1+ (ρ+ 1)c1−c1= 0
xρ+2: (ρ+ 2)(ρ+ 1)c2+ (ρ+ 2)c2+c0−c2= 0
xρ+3: (ρ+ 3)(ρ+ 2)c3+ (ρ+ 3)c3+c1−c3= 0
·
·
·
·
xρ+n: (ρ+n)(ρ+n−1)cn+ (ρ+n)cn+cn−2−cn= 0.
6.3. LINEAR EQUATIONS OF SECOND ORDER 263
From the equation for the power xρ, we find
(ρ2−1)c0= 0
The polynomial q(ρ) =ρ2−1 which appears in the coefficient of the lowest power of x
in (20) is called the indicial polynomial since it will be used to determine the indexρ. If
c0/negationslash= 0 , the equation ( ρ2−1)c0= 0 can be satisfied only if ρis a root of the indicial
polynomial. Thus ρ1= 1, ρ2=−1 .
Consider the largest root ρ1= 1 . Then the equation for the coefficients of xρ+1in
(20) is
xρ+1=x2: 3c1= 0⇒c1= 0,
while the equation for the coefficient of xρ+nin (20) is
xρ+n=x1+n: (n+ 1)ncn+ (n+ 1)cn+cn−2−cn= 0,
or
cn=−cn−2
n(n+ 2), n = 2,3,...
Sincec1= 0 , this equation implies codd= 0 and determines the cevenin terms of c0,
c2=−c0
2·4, c 4=−c2
4·6=c0
2·42·6, c 6=−c4
6·8=c0
2·42·62·8
c2k=(−1)kc0
2·42·62·82···(2k)2(2k+ 2)=(−1)kc0
22kk!(k+ 1)!.
Thus, the formal series we find for the solution, J1(x) , of the Bessel equation of first order
corresponding to the largest indicial root, ρ1= 1 is
J1(x) =1
2x1(1−x2
2·4+x4
2·42·6− ··· )
or
J1(x) =1
2x∞/summationdisplay
k=0(−1)kx2k
22kk!(k+ 1)!(6-21)
since it is customary to choose the constant c0forJ1(x) asc0=1
2(and the constant c0
forJn(x) as 1/2nn! whennis a positive integer).
The other (smaller) root, ρ2=−1 , is much more difficult to treat. If the above steps
are imitated (which you should try), division by zero needed to solve for c2fromc0. It
turns out that the solution corresponding to the smaller root ρ2=−1 is not of the form
(18). We shall not enter into this matter further except to note that the difficulty occurs
because the two rootsρ1andρ2differ by an integer . If the two roots ρ1andρ2donot
differ by an integer, the above method yields two different solutions of the form (18) for the
equation. In any case, this method always gives a solution of the form (18) for the largest
root of the indicial equation.
It is easy to check that the power series (21) does converge for all xand is therefore
a solution to Bessel’s equation of the first order. From the power series, with considerable
effort one can obtain a series of identities for Bessel functions which exactly parallels those
for the trigonometric functions. The functions Jn(x) behaving in many ways similar to
sinnxor cosnx. Here is a graph of J1(x) :
264 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
a figure goes here
Forxvery large, J1(x) is asymptotically
J1(x)∼/radicalbigg
2
πcos(x−3π/4)√x,
which is a cosine curve whose amplitude decreases like 1 /√x. For good reason this curve
resembles the height of surface waves on a lake after a pebble has been dropped into the
water, or those on the surface of a cup of tea.
Having worked out this example in detail, we shall state a definition in preparation for
our theorem.
Definition: The differential equation
a2(x)u/prime/prime+a1(x)u/prime+a0(x)u= 0,
where theaj(x) are analytic about x= 0 , it has a regular singularity atx= 0 if it can
be written in the form
x2u/prime/prime+A(x)xu/prime+B(x)u= 0,
where the functions A(x) andB(x) are analytic about x= 0 . Otherwise the singularity
isirregular .
Examples:
(1) .x2(1 +x)u/prime/prime+ 2(sinx)u/prime−exu=−has a regular singularity at x= 0 since the
equation may be written as
x2u/prime/prime+2(sinx)
1 +xu/prime−−ex
1 +xu= 0,
where the coefficients 2 sin x/(1 +x)xandex/1 +xdo have convergent Taylor series
aboutx= 0 . (Here we observed thatsinx
x= 1−x2
3!+···).
(2) .xu/prime/prime−7u/prime+3
cosxu= 0 has a regular singularity at x= 0 since it can be written in
the form
x2u/prime/prime−7xu/prime+3x
cosxu= 0,
where the coefficients −7 and 3x/cosxare analytic about x= 0 .
(3) .x2u/prime/prime−2u/prime+xu= 0 has an irregular singularity at x= 0 since it cannot be written
in the desired form.
(4) .x3u/prime/prime−2xu/prime+u= 0 has an irregular singularity at x= 0 .
Theorem 6.3 . (Frobenius) Consider the equation with a regular singularity at x= 0
a2(x)u/prime/prime+a1(x)u/prime+a0(x)u= 0,
so it can be written in the form
x2u/prime/prime+A(x)xu/prime+B(x)u= 0,
6.3. LINEAR EQUATIONS OF SECOND ORDER 265
where the analytic function A(x)andB(x)have convergent power series for |x|<r. Let
ρ1andρ2be the roots of the indicial polynomial
q(ρ) =ρ(ρ−1) +A(0)ρ+B(0),
whereρ1≥ρ2(orReρ 1≥Reρ 2if roots are complex). Then the differential equation has
one solution u1(x)of the form
u1(x) =xρ1∞/summationdisplay
n=0cnxn(c0/negationslash= 0),
the series converging for all |x|<r. Moreover, if ρ1−ρ2is not an integer (or zero), there
is a second solution u2(x)of the form
u2(x) =xρ2∞/summationdisplay
n=0˜cnxn(˜c0/negationslash= 0),
where this series also converges in the interval |x|< r. In the special case ρ1−ρ2=
integer, there may not be a solution of the form (18) - see Exercise 19c. Notice: although
the power series do converge at x= 0, the functions u1(x)andu2(x)may not be solutions
at that point because the functions xρmay not be twice differentiable (for example, if ρ=1
2
then√xhas no derivatives at x= 0).
Outline of Proof . Like Theorem 2, this proof also has two parts; i) finding the coefficients
cnfor the formal power series, and ii) proving the formal power series converges. As in
Theorem 3, part i) is proved by exhibiting formulas for the cn’s, while part ii) is proved by
comparing the series/summationtextcnxnwith another convergent series/summationtextCnxnwhose coefficients
are larger, |cn| ≤Cn.
To illustrate the procedure of part i), we will obtain the stated formula for the indicial
polynomial q(ρ) . LetA(x) =∞/summationdisplay
n=0αnxnandB(x) =∞/summationdisplay
n=0βnxnbe the power series expan-
sions ofA(x) andB(x) . Then assuming u(x) has a solution in the form (18), we find by
substituting these formulas into the differential equation that
∞/summationdisplay
n=0(ρ+n)(ρ+n−1)cnxρ+n+ (∞/summationdisplay
n=0αnxn)(∞/summationdisplay
n=0(ρ+n)cnxρ+n)
+(∞/summationdisplay
n=0βnxn)(∞/summationdisplay
n=0cnxρ+n) = 0.
The lowest power of xappearing is xρ, then comes xρ+1,....
xρ:ρ(ρ−1)c0+α0ρc0+β0c0= 0
xρ+1: (ρ+ 1)ρc1+ [α1ρc0+α0(ρ+ 1)c1] + [β1c0+β0c1] = 0
·
·
·
xρ+n: (ρ+n)(ρ+n−1)cn+n/summationdisplay
k=0αn−k[(ρ+k)ck] +n/summationdisplay
k=0βn−kck= 0,
266 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
the last formula arising from the formula for the coefficients in the product of two power
series (p. 76). If c0/negationslash= 0 , the first equation states
q(ρ) :=ρ(ρ−1) +α0ρ+β0= 0,
whereq(ρ) is the indicial polynomial. Since α0=A(0) andβ0=B(0) , this is precisely
the formula given in the theorem.
c) General Theory
We begin immediately by stating
Theorem 6.4 (Existence and Uniqueness). Consider the second order linear O.D.E.
Lu:=a2(x)u/prime/prime+a1(x)u/prime+a0(x)u=f(x),
where the coefficients a0,a1, anda2as well asfare continuous functions, and a2(x)/negationslash= 0.
There exists a unique twice differentiable function u(x)which satisfies the equation and the
initial conditions
u(x0) =α, u/prime(x0) =β,
whereαandβare arbitrary constants.
If time permits, the existence proof will be carried out in the last chapter as a special
case of a more general result. The uniqueness will be proved later too, as a special case of
Theorem 9, page 510 - in the next section. We will not be guilty of circular reasoning.
Now what? Although this theorem appears to make further study unnecessary, there
are several general statements which can be made because the equation is linear . Two
other theorems are particularly nice; the first is dim N(L) = 2 , while the second gives a
procedure for solving the inhomogeneous equation once two linearly independent solutions
of the homogeneous equation are known.
A preliminary result on linear dependence and independence of functions is needed. If
the differentiable functions u1(x) andu2(x) are linearly dependent, there are constants c1
andc2not both zero such that
c1u1(x) +c2u2(x)≡0.
Differentiating this equation, we find
c1u/prime
1(x) +c2u/prime
2(x)≡0.
Since the two homogeneous algebraic equations for c1andc2have a non-trivial solution,
by Theorem 32 (page 428), the determinant
W(x) :=W(u1,u2)(x) :=/vextendsingle/vextendsingle/vextendsingle/vextendsingleu1(x)u2(x)
u/prime
1(x)u/prime
2(x)/vextendsingle/vextendsingle/vextendsingle/vextendsingle= 0
must vanish. This determinant is called the Wronskian ofu1andu2. We have proved
6.3. LINEAR EQUATIONS OF SECOND ORDER 267
Theorem 6.5 . If the differentiable functions u1(x),u2(x)are linearly dependent in the
interval [α,β], then necessarily W(x)≡0throughout [α,β]. Thus, if W/negationslash= 0, theuj’s
are independent.
Remark: The condition W= 0 is necessary for linear dependence but not sufficient in
general, as can be seen from the example
u1(x) =/braceleftbiggx2, x≥0,
0, x< 0u2(x) =/braceleftbigg0, x≥0
x2, x< 0,
for whichW(u1,u2)≡0 for allxbutu1andu2are linearly independent. However it is
sufficient if u1andu2are solutions of a second order linear O.D.E., Luj= 0 . An even
stronger statement is true in this case. All we need require is that Wvanish at one point
x0.
Theorem 6.6 . Letu1andu2both be solutions of
Lu:=a2u/prime/prime+a1u/prime+a0u= 0,
wherea2/negationslash= 0. IfW(x0) = 0 at some point x0, thenu1andu2are linearly dependent -
which implies by Theorem 6 that W(x)≡0for allx. In other words, if W(x0)/negationslash= 0, then
u1andu2are linearly independent.
Proof: SinceW(x0) = 0 , the homogeneous algebraic equations
c1u1(x0) +c2u2(x0) = 0
c1u/prime
1(x0) +c2u/prime
2(x0) = 0
have a non-trivial solution c1,c2. Let
v(x) =c1u1(x) +c2u2(x).
We went to prove v(x)≡0 . Observe Lv= 0 . Moreover v(x0) = 0 andv/prime(x0) = 0 . Thus
by uniqueness, v(x)≡0 , establishing the linear dependence of u1andu2.
The same type of reasoning proves
Theorem 6.7 . LetLu:=a2u/prime/prime+a1u/prime+a0u, wherea2(x)/negationslash= 0. Then
dimN(L) = 2.
Proof: We exhibit two special solutions φ1andφ2ofLu= 0 and prove they constitute
a basis for N(L) . Let
φ1(x) satisfy Lφ1= 0 with φ1(x0) = 1, φ/prime
1(x0) = 0
φ2(x) satisfy Lφ2= 0 with φ2(x0) = 0, φ/prime
2(x0) = 1.
There are such functions by the existence theorem.
268 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
i) They are linearly independent.
W(x0) =W(φ1,φ2)(x0) =/vextendsingle/vextendsingle/vextendsingle/vextendsingleφ1(x0)φ2(x0)
φ/prime
1(x0)φ/prime
2(x0)/vextendsingle/vextendsingle/vextendsingle/vextendsingle=/vextendsingle/vextendsingle/vextendsingle/vextendsingle1 0
0 1/vextendsingle/vextendsingle/vextendsingle/vextendsingle= 1/negationslash= 0.
Thus by Theorem 7, φ1andφ2are linearly independent.
ii) They span N(L) . Letu(x) be any element in N(L) and consider the function
v(x) =u(x)−[u(x0)φ1(x) +u/prime(x0)φ2(x)].
ThenLv= 0 andv(x0) = 0,v/prime(x0) = 0 . By uniqueness, v(x)≡0 . Thus every u∈N(L)
can be written as
u(x) =Aφ1(x) +Bφ2(x),
where the constants AandBareA=u(x0),B=u/prime(x0) .
All of our attention has been on the homogeneous equation Lu= 0 . Let us solve the
inhomogeneous equation. This is particularly simple for a linear differential equation once
we have a basis for N(L) .
Theorem 6.8 (Lagrange). Let u1(x)andu2(x)be a basis for N(L), whereLu:=
a2(x)u/prime/prime+a1(x)u/prime+a0(x)u, witha2/negationslash= 0. Then the inhomogeneous equation Lu=f
has the particular solution
up(x) =u1(x)/integraldisplayxW1(s)
W(s)f(s)ds+u2(x)/integraldisplayxW2(s)
W(s)f(s)ds,
whereW(s) :=W(u1,u2)(s)andWj(s)is obtained from W(s)by replacing the jth
column (uj,u/prime
j)ofWby the vector (0,1/a2).
Remark: If we let
G(x;s) =u1(x)W1(s) +u2(x)W2(s)
W(s)
then the above formula assumes the elegant form
up(x) =/integraldisplayx
G(x;s)f(s)ds.
Proof: A device (due to Lagrange) called variation of parameters is needed. We already
used a form of this device to solve the inhomogeneous first order linear equation (5, p. 457).
The trick is to let
up(x) =v1(x)u1(x) +v2(x)u2(x)
where the functions v1(x) andv2(x) are to be found. This attempt to find upis reminiscent
of writing the general solution of the homogeneous equation as Au1+Bu2. Differentiate:
u/prime
p(x) =v1u/prime
1+v2u/prime
2+ [v/prime
1u1+v/prime
2u2].
The functions v1andv2will be chosen to make
v/prime
1u1+v/prime
2u2= 0.
6.3. LINEAR EQUATIONS OF SECOND ORDER 269
Using this, we differentiate again
u/prime/prime
p(x) =v1u/prime/prime
1=v2u/prime/prime
2+ [v/prime
1u/prime
1+v/prime
2u/prime
2]
Now multiply u/prime/prime
pbya2, u/prime
pbya1, upbya0, and add to find
Lup=v1Lu1+v2Lu2+a2[v/prime
1u/prime
1+v/prime
2u/prime
2]
=a2[v/prime
1u/prime
1+v/prime
2u/prime
2].
If we can choose v1andv2so thata2[ ] =f, then indeed Lup=f, sou0=v1u1+v2u2
is a particular solution. It remains to see if v1andv2can be found which satisfy the two
needed conditions
v/prime
1u1+v/prime
2u2= 0
v/prime
1u/prime
1+v/prime
2u/prime
2=f
a2.
These two linear equations for v/prime
1andv/prime
2may be solved by Cramer’s rule (Theorem 33,
page 429),
v/prime
1=/vextendsingle/vextendsingle/vextendsingle/vextendsingle0u2
f/a 2u/prime
2/vextendsingle/vextendsingle/vextendsingle/vextendsingle
W=f/vextendsingle/vextendsingle/vextendsingle/vextendsingle0u2
1/a2u/prime
2/vextendsingle/vextendsingle/vextendsingle/vextendsingle
W=W1
Wf
v/prime
2=/vextendsingle/vextendsingle/vextendsingle/vextendsingleu1 0
u/prime
1f/a 2/vextendsingle/vextendsingle/vextendsingle/vextendsingle
W=f/vextendsingle/vextendsingle/vextendsingle/vextendsingleu10
u/prime
11/a2/vextendsingle/vextendsingle/vextendsingle/vextendsingle
W=W2
Wf
Integration of these equations yields v1andv2, which, when substituted into up=u1v1+
u2v2, do give the stated result
With this theorem, knowing the general solution of the homogeneous equation L˜u= 0
allows us to find a particular solution of the homogeneous equation Lup=f. The general
solutionuof the inhomogeneous equation Lu=fis then the upcoset of N(L) , that is,
all functions of the form
u=up+ ˜u.
This puts the burden on finding the general solution of the homogeneous equation.
Examples:
(1) . The homogeneous equation x2u/prime/prime−3xu/prime+ 3u= 0, x/negationslash= 0 , has the two linearly
independent solutions u1(x) =x, u 2(x) =x3—which might have been found by the
power series method. Therefore a particular solution of the inhomogeneous equation
x2u/prime/prime−3xu/prime+ 3u= 2x4
can be found by the variation of parameters. We try
up=v1x3+v2x
and are led to the equations
v/prime
1=−2x4
x2x3
2x3, v/prime
2=2x4
x2x
2x3
270 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
or
v/prime
1=−x2, v/prime
2= 1.
Thus
v1(x) =−x3
3, v 2(x) =x.
Therefore
up(x) =x(−x3
3) +x3(x) =2
3x4
The general solution to the inhomogeneous equation is found by adding the general
solution of the homogeneous equation to this particular solution,
u(x) =Ax+Bx3+2
3x4.
(2) The homogeneous equation u/prime/prime+u= 0 has the linearly independent solutions u1(x) =
cosx, u 2(x) = sinx. Let us solve
u/prime/prime+u=f(x),
wherefis an arbitrary continuous function. Trying
up(x) =v1cosx+v2sinx,
we are led to
v/prime
1=−fsinx
1, v/prime
2=fcosx
1.
Thus
v1(x) =−/integraldisplayx
f(s) sinsds, v 2(x) =/integraldisplayx
f(s) cossds.
Therefore
up(x) =−cosx/integraldisplayx
f(s) sinsds+ sinx/integraldisplayx
f(s) cossds
=/integraldisplayx
f(s)[−sinscosx+ cosssinx]ds
=/integraldisplayx
f(s) sin(x−s)ds.
Consequently, the handsome formula
u(x) =Asinx+Bcosx+/integraldisplayx
f(s) sin(x−s)ds
is the general solution of the inhomogeneous equation u/prime/prime+u=f.
Exercises
(1) Solve the following initial value problems any way you can. Check your answers by
substituting back into the differential equation.
6.3. LINEAR EQUATIONS OF SECOND ORDER 271
(a)u/prime+ 2u= 0, u(1) = 2
(b)u/prime/prime+ 3u/prime+ 2u= 7, u(0) = 0, u/prime(0) = 0
(c)u/prime/prime+ 3u/prime+ 2u= 2ex, u(0) = 0, u/prime(0) = 1
(d)u/prime/prime+ 3u/prime+ 2u=e−2x, u(0) = 1, u/prime(0) = 0
(e) (tanx)du
dx+u−sin2x= 0, u(π
6) = 1
(f)u/prime/prime+u= tanx, u (0) =u/prime(0) = 1, x/epsilon1(−π
2,π
2).
(g)u/prime/prime/prime−8u= 0, u(0) = 1, u/prime(0) = 2, u/prime/prime(0) = 3
(h)u/prime/prime/prime/prime−k4u= 0.General solution.
(i)u/prime/prime−6u/prime+ 10u=x2+ sinx, u (0) =u/prime(0) = 0.
(j)u/prime/prime/prime/prime−7u/prime/prime/prime−8u/prime/prime= 0, u(0) = 3, u/prime(0) = 8, u/prime/prime(0) = 65, u/prime/prime/prime(0) = 511.
(k)xu/prime+u=x3, u(1) = 1.
(l)u/prime/prime+ 4u= 4x2+ cos 2x, u (0) = 0, u/prime(0) = 1
(m)u/prime/prime/prime−u/prime=ex. General solution.
(n)u/prime/prime/prime= 3u/prime/prime+ 3u/prime−u= 0, u(0) = 1, u/prime(0) = 2, u/prime/prime(0) = 3
(o)u(5)−u(4)+ 3u(3)−3u(2)−4u(1)+ 4u= 0 . General solution.
[Hint:λ5−λ4+ 3λ3−3λ2−4λ+ 4 = (λ2−1)(λ2+ 4)(λ−1) ].
(2) Find the first four non-zero terms (if there are that many) in the power series solutions
aboutx= 0 for the following equations.
(a)u/prime/prime−xu/prime−u= 0, u(0) =u/prime(0) = 1
(b)u/prime/prime−2xu/prime+ 2u= 0, u(0) = 0,u/prime(0) = 1.
(c)u/prime/prime−2xu/prime−2u= 0, u(0) = 1, u/prime(0) = 0.
(d)u/prime/prime+xu= 0, u(0) = 1, u/prime(0) = −1.
(e)u/prime/prime/prime−xu= 0, u(0) = 1,u/prime(0) =u/prime/prime(0) = 0.
(f)u/prime/prime−x2u=1
1−x2, u(0) = 0, u/prime(0) = 0.[Hint:1
1−x2= 1 +x2+x4+···]
(g)u/prime/prime−1
1−xu= 0, u(0) = 0, u/prime(0) = 1.[Hint:1
1−x=?]
(3) a) - e) Find where the power series in Ex. 2 a-e converge.
(4) Find the first four non-zero terms (if there are that many) in the power series solutions
corresponding to the larger root of the indicial polynomial.
(a) 2x2u/prime/prime−3xu/prime+ 2u= 0
(b)xu/prime/prime+ 2u/prime−xu= 0.[Answer:u(x) =c0∞/summationdisplay
0x2n
(2n+ 1)!] .
(c) 4xu/prime/prime+ 2u/prime+u= 0.
(d)xu/prime/prime+ (sinx)u/prime+x2u= 0, u(0) = 0, u/prime(0) = 1.
(e)xu/prime/prime+u/prime=x2.
(5) (a-e). Investigate the convergence of the series solutions found in Exercise 4 above.
272 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
(6) Find the power series solution about x= 0 for the nth order Bessel equation corre-
sponding to the highest root of the indicial polynomial. The answer is:
Jn(x) = (x
2)n∞/summationdisplay
k=0(−1)k
k!(k+n)!(x
2)2k,
where we have chosen c0= 1/2nn! .
(7) Find two linearly independent power series solutions of
u/prime/prime+xu/prime+u= 0
and prove they are linearly independent. Find all solutions.
(8) The Hermite equation is
u/prime/prime−2xu/prime+ 2αu= 0.
For which value(s) of the constant αare the solutions polynomials - that is, a solution
with a finite Taylor series. These are the Hermite polynomials .
(9) Find the first three non-zero terms in the power series about x= 0 for two linearly
independent solutions of
2x2u/prime/prime+xu/prime+ (x−1)u= 0.
(10) The homogeneous equation Lu:= 2x2u/prime/prime−3xu/prime−2u= 0 has the two linearly inde-
pendent solutions u1(x) =x2, u2(x) =√x(see Ex. 20c below). Find the general
solution of the inhomogeneous equation Lu= log(x3) .
(11) LetLu= (1−x2)u/prime/prime−2xu/prime+n(n+ 1)uwherenis an integer. Show that Lu= 0
has a polynomial solution - the Legendre polynomial. Compute this for n= 3 . (cf.
page 104l Ex. 10).
(12) LetJ0(x) be a solution of the zerothorder Bessel equation. ProvedJ0
dxis a solution
of the first order Bessel equation. [Hint: Work directly with the equation itself, not
with power series].
(13) Consider the equation
a2(x)u/prime/prime+a1(x)u/prime+a0(x)u= 0.
(a) Letu(x) :=u1(x)v(x) . Show that the result arranged as an equation for v(x)
is
a2u1v/prime/prime+ (2a2u/prime
1+a1u1)v/prime+ (a2u/prime/prime
1+a1u/prime
1+a0u1)v= 0
(b) Ifu1is known to be one solution of the equation, show that the second solution
isu2(x)
u2(x) =u1(x)/integraldisplay
w(x)dx
wherew(x) is a solution of the firstorder equation
a2u1w/prime+ (2a2u/prime
1+a1u1)w= 0.
6.3. LINEAR EQUATIONS OF SECOND ORDER 273
Thus, if one solution of a second order linear O.D.E. is known, the problem of finding
a second solution is reduced to the problem of solving a first order linear O.D.E. -
which can always be solved by separation of variables.
(14) Apply Exercise 13 to the following:
(a) One solution of 2 x2u/prime/prime−3xu/prime+ 2u= 0 isu1(x) =x2. Find another.
(b) One solution of x2u/prime/prime−xu/prime+u= 0 isu1(x) =x. Find another.
(c) One solution of (1 + x)xu/prime/prime−xu/prime+u= 0 isu1(x) =x. Find another, and then
write down the general solution.
(d) One solution of the equation x2u/prime/prime+2xu/prime= 0 is clearly u1(x) = 1 . Find another.
Prove the solutions are linearly independent for x>0 . Find the general solution
ofx2u/prime/prime+ 2xu/prime= 1 .
(15) Consider the O.D.E. u/prime/prime+a(x)u/prime+b(x)u= 0 , where aandbare continuous about
x0. If the graphs of two solutions are tangent at x=x0, are these two solutions
linearly dependent? Explain: Can you make an even stronger deduction?
(16) (a) Let Lbe a constant coefficient differential operator with characteristic polyno-
mialp(λ) . Ifp(λ) =p(−λ) , prove
L(sinkx) =p(ik) sinkx
(b) Apply this to find a particular solution of u/prime/prime/prime/prime−u= sin 2x
(17) Find a particular solution of the equation
u/prime/prime−n2u=f, n /negationslash= 0.
[You will need: sin h(α−β) = sinhαcoshβ−sinhβcoshα].
[Answer:u(x) =1
n/integraldisplayx
0f(s) sinhn(x−s)ds.]
(18) Use the method of variation of parameters to find a particular solution to u/prime/prime=f.
Compare with Exercise 5, p. 282.
(19) Consider the differential operator
Lu:=x2u/prime/prime+axu/prime+bu,
whereaandbare constants. This is called Euler’s equation . It is the simplest
equation with a regular singularity at x= 0 .
(a) Show that Lxρ=q(ρ)xρ, whereq(ρ) is the indicial polynomial for L.
(b) If the roots of q(ρ) = 0 are distinct, find two solutions of Lu= 0, x > 0 , and
prove the solutions are linearly independent for x>0 .
(c) If the roots ρ1andρ2ofq(ρ) = 0 coincide, take the derivative with respect to ρ
of the equation in a) - holding xfixed - to obtain the candidate u2(x) =xρ1lnx
for a second solution. Verify by substitution that u2is a solution in this case
and prove the two solutions
u1(x) =xρ1, u2(x) =xρ1lnx, x> 0
are linearly independent for x/negationslash= 0 .
274 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
(20) Apply the method of Exercise 19 to find two linearly independent solutions for each
of the following Euler equations
a).x2u/prime/prime+xu/prime= 0.
b). 2x2u/prime/prime−3xu/prime+ 2u= 0.
c). 2x2u/prime/prime−3xu/prime−2u= 0.
d).x2u/prime/prime−xu/prime+u= 0.
(21) (a) Use the result of Ex. 19 a) to find a particular solution of the equation Lu=xα,
where
Lu:=x2u/prime/prime+axu/prime+bu,
withaandbconstant, and where αisnota root of the indicial polynomial
q(ρ) (cf. Ex. 6, p. 300).
(b) If neither αnotβare roots of q(ρ) , find a particular solution to the inhomo-
geneous equation
Lu=Axα+Bxβ.
(c) Apply this procedure to find the general solution of
2x2u/prime/prime−3xu/prime−2u= 3x−4x1/3.
(d) How can you solve Lu=xαifαis a root of the indicial polynomial?
(22) (a) If uhasnderivatives and λis a constant, prove
Dn[eλxu] =eλx(D+λI)nu.
Thus (D+λI)nu=e−λxDn[eλxu] .
(b) LetL= (D−a)nbe a constant coefficient differential operator with charac-
teristic polynomial p(λ) = (λ−a)n. Showu(x) is a solution of the equation
Lu= 0 if and only if u(x) has the form
u(x) =eaxQ(x),
whereQ(x) is a polynomial of degree ≤n−1 .
(23) Consider the O.D.E. Lu=f, whereLis a second order constant coefficient operator,
and letλ1andλ2be the characteristic roots of L1. Assume i) Reλ 1<0 and
Reλ 2<0 , and ii)there is some constant Msuch that |f(x)| ≤Mfor allx∈[0,∞] .
(a) Prove every solution of Lu=fis bounded for x∈[0,∞] .
(b) If lim
x→∞f(x) = 0 , prove that as x→ ∞ , every solution of Lu=ftends to zero.
(24) Consider the operator Lu:=a2(x)u/prime/prime+a1(x)u/prime+a0(x)u, where the aj’s are continuous
forx∈[α,β] . Letu1,u2andφ1,φ2both be bases for N(L) . Prove there is a
constantk/negationslash= 0 such that
W(u1,u2)(x) =kW(φ1,φ2)(x) for all x∈[α,β].
6.3. LINEAR EQUATIONS OF SECOND ORDER 275
(25) (a) Generalize the procedure of Ex. 21b and show how the inhomogeneous Euler
equationLu=fcan be solved if fhas a power series expansion. You will have
to assume that no root of the indicial polynomial is a positive integer.
(b) Apply a) to find a particular solution (as a power series) of
2x2u/prime/prime+ 3xu/prime−u=1
1−x.
(26) Given the equation Lu:=u/prime/prime+a(x)u/prime+b(x)u= 0 has solutions u1(x) = sinx,
u2(x) = tanx, find the general solution of the inhomogeneous equation
Lu=cosx
1 + sin2x.
(27) (a) If Lu:=a2u/prime/prime+a1u/prime+a0uandL∗v:= (a2v)/prime/prime−(a1v)/prime+a0v, prove the Lagrange
identity
vLu−uL∗v=d
dx[a2(u/primev−v/primeu) + (a1−a/prime
2)uv],
where the functions ajare assumed to be sufficiently differentiable. The operator
L∗is the adjoint ofL.
(b) Show that Lisself-adjoint ,L=L∗, if and only if a/prime
2=a1. Write the Lagrange
identity in this case.
(c) Ifc1u1(x) +c2u2(x) is the general solution of the equation Lu= 0 find the
general solution of the adjoint equation L∗v= 0 . [Answer: v=c3u1+c4u2
u1u/prime
2−u/prime
1u2] .
(d) Letube a twice differentiable function which vanishes at αandβ. Show the
adjoint operator L∗has the property that for all such functions uandv,
/angbracketleftv, Lu/angbracketright=/angbracketleftL∗v, u/angbracketright
where
/angbracketleftf, g/angbracketright:=/integraldisplayβ
αf(x)g(x)dx.
(28) (a) Let Lbe a self-adjoint operator,L=L∗. IfLX 1=λ1X1andLX 2=λ2X2,
whereλ1andλ2are real number, λ1/negationslash=λ2, proveX1andX2are orthogonal
/angbracketleftX1, X 2/angbracketright= 0.
[Hint: Compare /angbracketleftX2, LX 1/angbracketright=λ1/angbracketleftX2, X 1/angbracketrightwith/angbracketleftLX 2, X 1/angbracketright=λ2/angbracketleftX2, X 1/angbracketright].
(b) LetL=d2
dx2. For what values of λcan you find a non-zero solution uof the
equationLu=λuwhereusatisfies the boundary conditions u(0) =u(π) = 0 ?
(c) Apply parts a) and b) as well a Ex. 27d to prove
/angbracketleftsinnx,sinmx/angbracketright=/integraldisplayπ
0sinnxsinmxdx = 0,
wherenandmare unequal integers.
276 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
(29) . Consider the boundary value problem
Lu:=u/prime/prime+u=f, u (0) = 0, u(π) = 0,
wherefis continuous in [0 ,π] .
a). Show that if a solution exists, it is not unique.
b). Show a solution exists if and only if
/integraldisplayπ
0f(x) sinxdx= 0.
[Hint: First find the general solution of the homogeneous equation].
Remark: In the notation of Ex. 27, we have L=L∗. Moreover, N(L∗) =
span{sinx}. The conclusions of b) states that R(L) =N(L∗)⊥, and illustrates
how Theorem 34, p. 431, is used in infinite dimensional spaces.
(30) . A proof of Theorem 3. Since a2(x)/negationslash= 0 , the equation can be written as
u/prime/prime+a(x)u/prime+b(x)u= 0.
If
a(x) =∞/summationdisplay
n=0αnx2, b (x) =∞/summationdisplay
n=0βnxn,
let
u(x) =∞/summationdisplay
n=0cnxn,whereu(0) =c0, u/prime(0) =c1,
(a) Imitate the example to prove the remaining cn’s must satisfy
cn+2=−n/summationdisplay
k=0[αn−k(k+ 1)ck+1+βn−kck]
(n+ 2)(n+ 1).
Show that if c0andc1are known, then the remaining cn’s are determined
inductively by the above formula.
(b) Because the series for a(x) andb(x) converge for |x|<r, ifRis any number
less thanr, there is a constant Msuch that for all n,|αn| ≤M
Rnand|βn| ≤M
Rn
(cf. p. 72, line 2). Define constants Cnas
C0=|c0|, C1=|c1|,
and forn≥0
Cn+2=M
Rnn/summationdisplay
k=0[(k+ 1)Ck+1+Ck]Rk+MC n+1R
(n+ 2)(n+ 1).
(i) Prove |cn| ≤Cn, n = 0,1,2,3,...
6.3. LINEAR EQUATIONS OF SECOND ORDER 277
(ii) Prove/vextendsingle/vextendsingle/vextendsingle/vextendsingleCn+1xn+1
Cnxn/vextendsingle/vextendsingle/vextendsingle/vextendsingle=n(n−1) +MnR +MR2
R(n+ 1)n|x|.
(iii) Prove∞/summationdisplay
n=0Cnxnconverges for |x|< R, whereRis any number less than
r.
(iv) Prove∞/summationdisplay
n=0cnxnconverges for |x|< R, whereRis any number less than
r.
(31) .
(a) Letu(x) andv(x) be solutions of the equations L1u:=u/prime/prime+a(x)u= 0 ,
andL2v:=v/prime/prime+b(x)v= 0 respectively, in some interval, where aandbare
continuous. If b(x)≥a(x) throughout the interval, prove there must be a zero
ofvbetween any two zeroes of u. This is the Sturm oscillation theorem . [Hint:
Supposeαandβare consecutive zeroes of uandu>0 in (α,β) . Prove
0 =/integraldisplayβ
α(vL1u−uL2v)dx=vu/prime/vextendsingle/vextendsingleβ
α−/integraldisplayβ
α(b−a)uvdx,
and show, because u/prime(α)>0, u/prime(β)<0 , there is a contradiction if vdoes not
vanish somewhere in ( α,β) .]
(b) Letu1(x) andu2(x) be two linearly independent solutions of u/prime/prime+a(x)u= 0 .
Prove between any two zeroes of u1, there is a zero of u2and vice verse. Thus,
the zeroes interlace.
(c) Apply b) to the solutions sin γxand cosγxof the equation u/prime/prime+γ2u= 0 to
conclude a well-known fact.
(d) Ifb(x)≥δ >0 , whereδis a constant, prove every solution of v/prime/prime+b(x)v=
0 must have an infinite number of zeros by comparing vwith a solution of
u/prime/prime+γ2u= 0 , where γis an appropriate constant.
(e) Apply d) to prove every solution of
v/prime/prime+ (1−3
4x2)v= 0,
has an infinite number of zeroes for x≥1 .
(f) Letu1(x) be a solution of the first order Bessel equation. Take v(x) =u1(x)√x
and show that vsatisfies the equation in e). Deduce that J1(x) has infinitely
many zeroes.
(32) LetL1andL2be linear constant coefficient differential operators with characteristic
polynomials p1(λ) andp2(λ) respectively.
(a) If there is a function u(x), u(x)/negationslash≡0 , which satisfies both L1u= 0 andL2u= 0 ,
prove the polynomials p1andp2have a common root.
(b) Ifp1andp2have no common roots, prove the solution of L1L2u= 0 are exactly
all functions of the form c1u1+c2u2whereu1is a solution of L1u1= 0 , andu2
ofL2u2= 0 . Thus N(L1L2) may be decomposed into the two complementary
subspaces N(L1) and N(L2),N(L1L2) =N(L1)⊕N(L2) .
278 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
(33) Imitate Exercise 30 and prove Theorem 3. Make sure to observe the trouble in trying
to find the solution corresponding to the lower root of the indicial polynomial if the
roots differ by an integer.
(34) The purpose of this exercise is to show that an equation with an irregular singular
point may have a formal power series at that point which does not converge to the
solution.
Try to find a solution of the form (18) for the following equation which has an irregular
singularity at x= 0 ,
x6u/prime/prime+ 3x5u/prime−4u= 0.
What happened? Two linearly independent solutions for x/negationslash= 0 are
u1(x) =e−1/x2andu2(x) =e1/x2.
How does this explain the situation (cf. p. 95-6)?
(35) Consider the equation 2 x2u/prime/prime+ 3xu/prime+u=/radicalbig
(x) . Two linearly independent solutions
of the homogeneous equation are x−1/2andx−1. Find the general solution of the
homogeneous equation.
(36) Consider the equation u/prime/prime+b(x)u/prime+c(x)u= 0 , where bandcare continuous
functions and c(x)<0 . Prove that a solution cannot have a positive maximum or
negative minimum.
6.4 First Order Linear Systems
Quite often in applications you must consider systems of differential equations. We shall
consider a linear system of the form
du1
dx+a11(x)u1+a12(x)u2+···+a1n(x)un=f1(x) (6-22)
du2
dx+a21(x)u1+a22(x)u2+···+a2n(x)un=f2(x) (6-23)
...... (6-24)
dun
dx+an1(x)u1+an2(x)u2+···+ann(x)un=fn(x), (6-25)
where the functions aij(x) andfj(x) are continuous. If we anticipate the next chapter
and write the derivative of a vector U= (u1,...,u n) as the derivative of its components,
d
dxU(x) =/parenleftbiggdu1
dx,du2
dx,···,dun
dx/parenrightbigg
,
then the above system can be written in the clean form
dU
dx+A(x)U=F(x), (6-26)
where,
A(x) = ((aij)), F = (f1,f2,...,f n)
6.4. FIRST ORDER LINEAR SYSTEMS 279
and
U(x) = (u1,u2,...,u n).
The initial value problem for the system of differential equations (22) is to find a vector
U(x) which satisfies the equation as well as the initial condition
U(x0) =U0, (6-27)
whereU0is a vector of constants.
It is useful to observe that the initial value problem for a single linear equation of order
n
u(n)+an−1(x)u(n−1)+···+a0(x)u=f(x)
u(x0) =α1, u/prime(x0) =α2,...,u(n−1)(x0) =αn,
can be transformed to the conceptually simpler problem (22)-(23). Let u1(x) :=u(x) ,
u2(x) :=u/prime(x),..., andun(x) =u(n−1)(x) . Then the components of the vector U(x) =
(u1,u2,...,u n) must obviously satisfy the relations
du1
dx=u2
du2
dx=u3
·
·
·
dun−1
dx=un
dun
dx=−a0u1−a1u2− ··· −an−1un+f(x),
which may be written as
U/prime=MU+F,
where
M(x) =
0 1 0 ··· 0
0 0 1 ··· 0
......
0 0 0 ··· 1
−a0−a1−a2··· −an−1
,
and
F= (0,0,..., 0,f).
The initial conditions read
U(x0) = (α1,α2,...,α n).
Conversely, if Uis any solution of this system of equations with the proper initial conditions,
then the first component u1(x) is a solution of the single nth order equation. Thus, the
general theory of a single nth order linear O.D.E. is completely subsumed as a portion
of the theory of a system of first order linear O.D.E.’s. You should be warned that this
generalization is mainly of theoretical value and is of little use if you are seeking an explicit
solution.
Both the existence and uniqueness theorems are true for systems, and supply an example
where the theoretical advantages of systems become clear. To illustrate this, we shall prove
the uniqueness theorem. Our proof is patterned directly after the uniqueness proof for a
single equation (Theorem 1).
280 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
Theorem 6.9 (Uniqueness). Let A(x)be a matrix whose coefficients aij(x)are bounded
|aij(x)| ≤Mforxin some interval, and let F(x)be a continuous function. Then there
is at most one solution U(x)of the initial value problem
U/prime+AU=F, U (x0) =U0.
Remark: The existence theorem states, if Ais nonsingular and each element is integrable
there is at least one solution. Thus, there is then exactly one solution.
Proof: AssumeU1andU2are both solutions. Let
W=U1−U2.
ThenWsatisfies the homogeneous equation and is zero at x0,
W/prime+AW= 0, W (x0) = 0.
Take the scalar product of this with W,
/angbracketleftW, W/prime/angbracketright+/angbracketleftW, AW /angbracketright= 0.
But
/angbracketleftW, W/prime/angbracketright=w1w/prime
1+ww/prime
2+···+wnw/prime
n
=1
2d
dx(w2
1+w2
2+···+w2
n)
=1
2d
dx/bardblW/bardbl2.
Thus,
1
2d
dx/bardblW/bardbl2=−/angbracketleftW, AW /angbracketright.
By Theorem 17, p. 173 and the hypothesis |aij(x)| ≤M, we know
/vextendsingle/vextendsingle/angbracketleftW, AW /angbracketright/vextendsingle/vextendsingle≤/bracketleftBign/summationdisplay
i,j=1|aij|2/bracketrightBig1/2
/bardblW/bardbl2≤nM/bardblW/bardbl2.
so that
1
2d
dx/bardblW/bardbl2≤nM/bardblW/bardbl2.
Therefore, as on p. 462-3
d
dx(/bardblW/bardbl2)−2nM/bardblW/bardbl2≤0,
or
e2nMxd
dx[e−2nMx/bardblW/bardbl2]≤0.
Becausee2nMxis always positive, by the mean value theorem the quantity [ ] is a
decreasing function. Its value for x>x 0is then less than at x0,
e−2nMx/bardblW(x)/bardbl2≤e−2nMx 0/bardblW(x0)/bardbl2, x ≥x0
6.4. FIRST ORDER LINEAR SYSTEMS 281
Consequently
/bardblW(x)/bardbl ≤enM(x−x0)/bardblW(x0)/bardbl, x ≥x0.
SinceW(x0) = 0 and the norm is non negative, we have
0≤ /bardblW(x)/bardbl ≤0, x ≥x0,
which implies
/bardblW(x)/bardbl= 0, x ≥x0.
Therefore,
W(x)≡0x≥x0.
By replacing xwith−xin the original equation, the same statement is true for x≤x0.
Thus, throughout the interval where |aij(x)| ≤M, we have proved W(x)≡0 , that is,
U1(x)≡U2(x) , so the solution is indeed unique.
Because a single linear nth order O.D.E. can be replaced by an equivalent system of
equations, this theorem implies the uniqueness theorem for a single O.D.E. of order nif the
coefficients aj(x) are bounded in some interval - which is certainly true in every interval if
theaj’s are continuous.
With this theorem, a short section closes. Further developments in the theory of systems
of linear O.D.E.’s make elegant use of linear operators in general and matrices in particular.
As you might well accept, the exercises contain a few of the more accessible results.
Exercises
(1) . Find functions u1(x),u2(x) which satisfy
u/prime
1=u1
u/prime
2=u1−u2,
with the initial conditions U(0) := (u1(0),u2(0) = (1,0) . Find the general solution
too. [Hint: Solve the equation u/prime
1=u1first, then substitute. Answer: General
solution is U(x) = (γ1ex,γ1
2ex+γ2e−x) ].
(2) Consider the system
u/prime
1= 2u1−u2
u/prime
2= 3u1−2u2,
that is,
U/prime=AU, whereA=/parenleftbigg2−1
3−2/parenrightbigg
.
Letφ1(x) =au1+bu2, φ2(x) =cu1+du2, wherea,b,c anddare constants. Thus,
Φ =SU,
where
S=/parenleftbigga b
c c/parenrightbigg
, Φ = (φ1,φ2).
282 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
(a) By direct substitution, find the differential equations satisfied by the φj’s and
show they can be written as
Φ/prime=SAS−1Φ.
(b) Pick the coefficients of Sso the matrix SAS−1is a diagonal matrix,
SAS−1=/parenleftbiggλ10
0λ2/parenrightbigg
≡Λ
(c) Solve the resulting equation Φ/prime= ΛΦ . [Solution: φ1=αex, φ 2=βe−x—you
might have φ1andφ2interchanged].
(d) Use this solution to solve the original equations for U. [hint: RecallU=
S−1Φ ].
(3) By only a slight modification of Exercise 2, solve
v/prime/prime
1= 2v1−v2
v/prime/prime
2= 3v1−2v2.
[Hint: Everything, even the algebra, is identical. The only difference is in part c) you
have to solve Φ/prime/prime= ΛΦ . Then V=S−1Φ as before].
(4) A bathtub initially contains Q1gallons of gin and Q2gallons of vermouth, where
Q1+Q2=Q, Q being the capacity of the tub. Pure gin enters from one faucet at
a constant rate of R1gallons per minute, while pure vermouth enters from another
faucet at a constant rate R2gallons per minute. The well stirred mixture of martinis
leaves the drain at a rate R1+R2gallons per minute (so the total amount of fluid in
the tub remains constant at Qgallons). Let G(t) be the quantity of gin in the tub
at timetandV(t) be the quantity of vermouth.
(a) Show
dG
dt=R1−G
Q(R1+R2)
dV
dt=R2−V
Q(R1+R2).
(b) Integrate this simple system of equations to find G(t) andV(t) . Also find their
ratioP(t) :=G(t)/V(t) which is the strength of the martinis at time t.
(c) Prove
lim
t→∞P(t) =R1
R2.
Compare this with your intuitive expectations.
(d) IfQ1= 20,Q2= 0,R1=R2= 1 gal/min, how long must I wait to get a perfect
martini (for me, perfect is 5 parts gin to 1 part vermouth). [Needless to say, the
mathematical model is applicable to many problems in the mixing of chemicals
which do not react with each other. If the chemicals do interact, the model must
be changed to account for the interaction].
6.5. TRANSLATION INVARIANT LINEAR OPERATORS 283
(5) Consider the homogeneous equation U/prime=A(x)U, whereAis non-singular (so
detA/negationslash= 0 ). Assuming the validity of the existence theorem, prove there exists
nlinearly independent vectors U1(x),U2(x),...,U n(x) which are solutions, U/prime
k=
AUk, k= 1,...,n . [Hint: Construct nsolutions which are linearly independent at
x=x0, and then prove a set of nsolutions are linearly independent in an interval
if and only if they are linearly independent at x=x0, wherex0is a point in the
interval].
(6) LetLU:=U/prime−A(x)Uas in Exercise 5. Prove dim N(L) =n.
(7) LetLU:=U/prime−A(x)U. If a basis U1,...,U n, forN(L) is known, prove the inho-
mogeneous equation LU=Fcan be solved by variation of parameters. That is, seek
a particular solution UpofLU=Fin the form
Up=n/summationdisplay
i=1Uivi
where thevi(x) are scalar-valued functions ( notvectors).
(a) Compute U/prime
pand substitute into the O.D.E. to conclude Upis a particular
solution ifn/summationdisplay
i=1Uiv/prime
i=F.
(b) LetUbe then×nmatrix whose columns are U1,U2,...,U n. ProveUis
invertible and show
v/prime
i(x) = (U−1F)ith component.
(c) Show
Up(x) =n/summationdisplay
i=1Ui(x)/integraldisplayx
[U−1(s)F(s)]ids.
This may also be written in the form
Up(x) =U(x)/integraldisplayx
U−1(s)F(s)ds
(d) Apply this procedure to find the general solution of
u/prime
q=u1+e2xcf. Ex 1
u/prime
2=u1−u2+ 1.
6.5 Translation Invariant Linear Operators
This section develops various extensions and applications of the procedure used to solve
linear ordinary differential equations with constant coefficients. The results will be proved
as a series of exercises interspersed by various remarks.
Definition: Thetranslation operator Ttacting on functions u(x) is defined by the prop-
erty
(Ttu)(x) =u(x−t). x,t ∈R.
284 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
A linear operator Listranslation invariant if
LTt=TtL
for everyt, that is, if
L(Ttu) =Tt(Lu)
for everytand for every function ufor which the operators are defined.
Example: 1 Let (Lu)(x) := 3u(x)−2u(x−1) . Then
[Tt(Lu)](x) = 3u(x−t)−2u(x−t−1),
and
[L(Ttu)](x) =Lu(x−t) = 3u(x−t)−2u(x−t−1).
Thus,
LTt=TtL,
so the operator Lis translation invariant.
2. Let (Lu)(x) := 3xu(x).Then
[Tt(Lu)](x) = 3(x−t)u(x−t),
and
[L(Ttu)](x) =Lu(x−t) = 3xu(x−t).
Thus
LTt/negationslash=TtL,
so this operator is nottranslation invariant.
Exercises
(1) Which of the following linear operators (verify!) are also translation invariant?
(a) (Lu)(x) :=cu(x), c ≡constant
(b) (Lu)(x) :=u(x+h)−u(x)
h, h≡constant /negationslash= 0 .
(c) (Lu)(x) :=/integraldisplayx
−∞k(x−s)u(s)ds
(d) (Lu)(x) := (x−1)u(x)
(e) (Lu)(x) =du
dx(x).
(f) Any linear ordinary differential operator with constant coefficients,
Lu:=anu(n)+an−1u(n−1)+···+a0u, a kconstants.
(g) Any linear ordinary differential operator with variable coefficients.
(h) (Lu)(x) =n/summationdisplay
k=1aku(x−γk), a kandγkconstants.
[Answers: All but d) and g) are translation invariant].
6.5. TRANSLATION INVARIANT LINEAR OPERATORS 285
(2) IfL1andL2are translation invariant operators which map some linear space into
itself, then so are
a).AL1+BL 2, A,B constants
b).L1L2andL2L1
c). If in addition Lis invertible, then L−1is also translation invariant.
Theorem 6.10 . IfLis a translation invariant linear operator, then
L(eλa) =φ(λ)eλx.
Proof: We know so little about Lthat all we can hope to do is compute TtL(eλx)
andLTt(eλx) and see what happens. Let Leλx=ψ(λ;x) , whereψis some unknown
function whose value depends on both λandx. Then
TtL(eλx) =ψ(λ;x−t),
while
LTteλx−Leλ(x−t)=L(e−λteλx)
=e−λtLeλx=e−λtψ(λ;x).
SinceTtL=LTt, we find
e−λtψ(λ;x) =ψ(λ;x−t),
or
ψ(λ;x) =ψ(λ;x−t)eλt.
Because the left side does not contain t, the right side must not depend on which
value oftis chosen. Using this freedom, we let t=xand conclude
ψ(λ;x) =ψ(λ; 0)eλx.
By setting φ(λ) =ψ(λ,0) , we find
Leλx=ψ(λ;x) =φ(λ)eλx
as desired.
Exercises
(3) By direct substitution, find φ(λ) for those operators in Exercise 1 which are trans-
lation invariant. [Answers: a) φ(λ) =c, b)φ(λ) = (e−ah−1)/hc)φ(λ) =/integraldisplay0
−∞k(−s)eλsds, d)φ(λ) =cλ, f)φ(λ) =n/summationdisplay
k=0akλk(the characteristic polynomial),
h)φ(λ) =n/summationdisplay
k=1ake−λγk].
286 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
(4) With the same assumptions and notation as in the theorem, if φ(λ) = 0 is a poly-
nomial equation with Ndistinct roots λ1,λ2,...,λ N, soφ(λj) = 0, j= 1,...,N ,
prove any linear combination of the function eλjxis inN(L) , that is,
Lu= 0 where u(x) =N/summationdisplay
1cjeλjx.
(5) Apply the theorem to find the solution of Exercise 4 for the equation Lu= 0 , where
(a)Lu:=u/prime/prime−u/prime−u.
(b) (Lu)(x) =u(x+ 2)−u(x+ 1)−u(x) .
(c) Find a special solution of b) which satisfies the “initial conditions” u(0) =u(1) =
1 . Compute u(2),u(3) andu(4) directly from b). The integers u(n), n∈Z+
are called the Fibonacci sequence . [Answer: u(2) = 2,u(3) = 3,u(4) = 5 , and
surprisingly ,
u(n) =1√
5
/parenleftBigg
1 +√
5
2/parenrightBiggn+1
−/parenleftBigg
1−√
5
2/parenrightBiggn+1
].
(6) Solveu(x)−au(x−1) +b2u(x−2) = 0 with the initial conditions u(1) =a,u(2) =
a2−b2. Compare with Exercise 17, p. 440.
(7) Extend Exercises 5(b - c) and 6 to develop a theory of second order difference equations
with constant coefficients . Thus
Lu:=a2u(x+ 2) +a1u(x+ 1) +a0u(x), a 2/negationslash= 0, x ∈Z.
In particular, you should,
(a) Find two linearly independent solutions of Lu= 0 . Remember the degenerate
casea2
1−4a0a2= 0 .
(b) Prove there is at most one solution of the initial value problem Lu=f,u(0) =
α0,u(1) =α1.
(c) Prove dim N(L) = 2 .
Remarks: The ideas presented above generalize immediately to the case where X∈Rn
instead of just R1, as well as to the case where the u’s are vectors and not scalars. These few
concepts lie at the heart of any treatment of many linear operators with constant coefficients,
especially ordinary and partial differential operators. This mildly abstract formulation
manages to penetrate through the obscuring details of particular cases to observe a rather
simple structure unifying many seemingly different problems.
6.6 A Linear Triatomic Molecule
A molecule composed of three atoms is called a triatomic . Consider a triatomic molecule
whose equilibrium configuration is a straight line with two atoms of equal mass msituated
on either side of a central atom of mass M.
6.6. A LINEAR TRIATOMIC MOLECULE 287
a figure goes here
To simplify the situation further, we shall only consider the motion along the straight
line (axis) of these atoms, and shall assume the inter-atomic forces can be approximated by
springs with equal spring constants k.u1(t),u2(t) andu3(t) will denote the displacements
of the atoms (see fig.) from their equilibrium position.
Newton’s second law, m¨u=/summationtextF, will give the equations of motion. The atom on
the left only “feels” the force due to the spring attached to it, the force being equal to the
spring constant ktimes the amount that spring is stretched, u2−u1. Thus,
m¨u1=k(u2−u1).
The central atom “feels” two forces, one from each side, with the resulting equation of
motion
M¨u2=−k(u2−u1) +k(u3−u2).
In the same way, the equation of motion for the remaining atom is
m¨u3=−k(u3−u2).
Collecting our equations, we have
¨u1=−k
mu1+k
mu2
¨u2=k
Mu1−2k
Mu2+k
Mu3
¨u3=k
mu2−k
mu3.
These are a system of three linear ordinary differential equations with constant coefficients.
They cannot be integrated as they stand since each equation involves functions from the
other equations, that is, the equations are copied (not surprising since we are considering
coupled oscillators . Now we can integrate such a system immediately if they are in the
simple form
¨φ1=λ1φ1
¨φ2=λ2φ2
¨φ3=λ3φ3
by integrating each equation separately. By using an important method, we will be able to
place our system in this special form.
Before doing so, it is suggestive to rewrite the system in matrix form
¨u1
¨u2
¨u3
=
−k
mk
m0
k
M−2k
Mk
M
0k
m−k
m
u1
u2
u3
.
LettingAdenote the 3 ×3 matrix, our hope is to somehow change Ainto a diagonal matrix
(one with zeroes everywhere except along the principal diagonal), for then the differential
equations will be in a form mentioned above which can be immediately integrated.
288 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
The trick is to replace the basis u1,u2,u3by some other basis in which the matrix
assumes a diagonal form. The differential equation can be written in the form
¨U=AU,
whereU= (u1,u2,u3) , and the derivative of a vector being defined as the derivative of
each of its components. Let φ1(t),φ2(t) , andφ3(t) be three other functions - which we
plan to use as a new basis. Then the φj’s can be written as a linear combination of the
uj’s,
φ=s11u1+s12u2+s13u3
φ2=s21u1+s22u2+s23u3
φ3=s31u1+s32u2+s33u3,
wheresijareconstants . WritingS= ((sij)) and Φ = ( φ1,φ2,φ3) , this last equation reads
Φ =SU.
Taking the derivative of both sides (or going back to the equations defining φjin terms of
theuk’s), we find
¨Φ =S¨U.
Because both u1,u2andu3as well asφ1,φ2, andφ3are bases for the solution, the matrix
Smust be non-singular (its inverse expresses the φ/prime
jsin terms of the uj’s). Thus
¨Φ =SAS−1Φ.
The problem has been reduced to finding a matrix Ssuch that the matrix SAS−1is a
diagonal matrix ,
SAS−1=
λ10 0
0λ20
0 0λ3
≡Λ.
Multiply by S−1on the left:
AS−1=S−1Λ.
Since this equation is equally between matrices, their corresponding columns must be equal.
Thus, if we denote by ˆSi, theith column of S−1, the above equation then reads
AˆSi=λiˆSi,
or
(A−λiI)ˆSi= 0.
For eachithis is a system of three linear algebraic equations for the three components of
ˆSi. If there is to be a non-trivial solution, we know
det(A−λiI) = 0.
Since
det(A−λiI) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle−k
m−λik
m0
k
M−2k
M−λik
M
0k
m−k
m−λi/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle,
6.6. A LINEAR TRIATOMIC MOLECULE 289
(algebra later)
=−λi(k
m+λi)[λi+ (2
M+1
m)k]
We see the three possible values of λfor det(A−λiI) = 0 are
λ1= 0,λ2=−k
m, λ 3=−k(2
M+1
m).
These numbers λiare the eigenvalues of A. The non-trivial solution ˆSiof the homo-
geneous equations ( A−λiI)ˆSi= 0 corresponding to the ith eigenvalue is called the
eigenvalue ofAcorresponding to the eigenvalue λi. For example, ˆS2is the solution of
(A−λ2I)S2= 0 corresponding to λ2=−k/m,
0ˆs12+k
mˆs22+ 0ˆs32= 0
k
Mˆs12−(2k
M−k
m)ˆs22+k
Mˆs32= 0
0ˆs12+k
mˆs22+ 0ˆs32= 0.
We see ˆs22= 0 while ˆs12=−ˆs32. Thus, one solution is
ˆS2= (1,0,−1)
Similarly we find one solution for ˆS1is
ˆS1= (1,1,1),
while one solution for ˆS3is
ˆS3= (1,−2m
M,1).
The computation is over. All that remains is to put the parts together and interpret
the solution. If you got lost, presumably this recapitulation will help. We have found a
transformation Sto new coordinates ( φ1,φ2,φ3) such that the differential equations for
theφj’s are in diagonal form, ¨φm=λjφj,
¨φ1= 0
¨φ2=−k
mφ2
¨φ3=−k(2
M+1
m)φ3.
The solutions are
φ1(t) =A1+B1t,
φ2(t) =A2cos/radicalbigg
k
mt+B2sin/radicalbigg
k
mt
φ3=A3cos/radicalbigg
k(2
M+1
m)t+B3sin/radicalbigg
k(2
M+1
m)t.
290 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
Since Φ =SU, and the ˆSjare the columns of S−1,
S−1=
1 1 1
1 0 −2m
M
1−1 1
,
we haveU=S−1Φ ,
u1(t) =φ1(t) +φ2(t) +φ3(t)
u2(t) =φ1(t) −2m
Mφ3(t)
u3(t) =φ1(t)−φ2(t) +φ3(t)
Although the solutions φ1(t),φ2(t) , andφ3(t) can now be substituted into the first
set of equations for the uj’s, it is more instructive to leave that step to your imagination
and analyze the nature of the solution.
(1) Ifφ1(t)/negationslash= 0 but φ2(t) =φ3(t) = 0,then
u1(t) =u2(t) =u3(t) =A1+B1t.
Thus all three atoms - the whole molecule - moves with a constant velocity B1.
This is the trivial translation motion of the molecule, simply moving without internal
oscillations at all.
(2) Ifφ2(t)/negationslash= 0 butφ1(t) =φ3(t) = 0 , then
u1(t) =φ2(t) =−u3(t),andu2(t) = 0.
Thus, the two outside atoms vibrate in opposite directions with frequency/radicalbig
k/m
while the center atom remains still:
a figure goes here
(3) Ifφ3(t)/negationslash= 0 butφ1(t) =φ2(t) = 0
u1(t) =u3(t) =φ3(t), u 2(t) =−2m
Mφ3(t).
A bit more complicated. The two outside atoms move in the same direction with same
frequency/radicalBig
k(2
M+1
m) , while the center atom moves in a direction opposite to them and
with the same frequency but a different amplitude (to conserve linear momentum m˙u1+
M˙u2+m˙u3= 0 ). In the figure we take m=M.
a figure goes here
These three simple motions are called the normal modes of oscillation of the molecule.
They are the oscillations determined by the φ1,φ2, andφ3. Every motion of the system is
a linear combination of the normal modes of oscillation, the particular oscillation depending
on what initial conditions are given. By an appropriate choice of the initial conditions, one
or another of the normal modes will result. Otherwise some less recognizable motion will
result.
Exercises
Consider the simpler model of a diatomic molecule
6.6. A LINEAR TRIATOMIC MOLECULE 291
a figure goes here
which we will represent as two masses joined by a spring with spring constant k.
(a) Show the equations of motion are
m¨u1=k(u2−u1)
M¨u2=−k(u2−u1)
(b) Introduce new variables, Φ = SU,
φ1=s11+s12u2
φ2=s21u1+s22u2,
and findSso that the equation
¨Φ =SAS−1Φ
is in diagonal form.
(c) Solve the resulting equation and find the normal modes of oscillation. Interpret your
results with a diagram.
292 CHAPTER 6. LINEAR ORDINARY DIFFERENTIAL EQUATIONS
Chapter 7
Nonlinear Operators: Introduction
7.1 Mappings from R1toR1, a Review
.
The subject of this section is one you presumably know well. Our intention is to briefly
review the more important results, stating them in a form which suggests the generalizations
we intend to develop.
Consider a function y=f(x), x∈R. This function assigns to each number xanother
real number y. Thus we may write
f:R→R.
fis a scalar-valued function of a scalar. What are the simplest such functions? Linear
ones of course,
f(x) =ax+b.
In keeping with our more sophisticated terminology, this should be called an “affine” func-
tion (mapping, operator, ...) since it is linear only if b= 0 . We shall, however, be abusive
and refer to such functions as linear mappings. The study of linear functions in one variable,
x, is carried out in elementary analytic geometry.
At an early age we enlarged our vocabulary of functions from linear ones to a more
general class which includes, for example,
f1(x) =ax2+bx+c, f 2(x) = sinx, f 3(x) =√x.
These functions are all examples of nonlinear functions. They map the reals (only the
positive reals in the case of f3) into the reals. The portion of the reals for which they are
defined is called their domain of definition ,D(f) . Thus
D(f1) =R1,D(f2) =R1,D(f3) ={x∈R1:x>0}.
The class of all real valued functions of a real variable is too large to consider. For
most purposes it is sufficient to restrict oneself to the class of continuous or sufficiently
differentiable functions.
Here is an outline of the basic definitions and theorems from elementary calculus. In
our prospective generalization from the simplest case of a function (operator) fwhich
maps numbers to numbers, f:R1→R1, to the case of a function from vectors to vectors
f:Rn→Rm, all of these concepts and results will need to be extended.
293
294 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION
Definition: akconverges to a, a kanda∈R1.
Definition: Continuity.
Theorem 7.1 The set of continuous functions forms a linear space.
Definition: The derivative: limit of difference quotient.
Theorem 7.2 1.d
dx(af+bg) =ad f
dxf+bdg
dx(linearity)
2.d
dx(fg) =fdg
dx+ (d f
dx)g(Product rule)
3.d
dx(f◦g) =d f
dgdg
dx(Chain rule)
Theorem 7.3 The Mean Value Theorem.
Definition: The integral.
Theorem 7.4 1./integraldisplayb
af(x)dx=−/integraldisplaya
bf(x)dx
2./integraldisplayb
1f(x)dx+/integraldisplayc
bf(x)dx=/integraldisplayc
af(x)dx
3./integraldisplayb
a[αf(x) +βg(x)]dx=α/integraldisplayb
af(x)dx+β/integraldisplayb
ag(x)dx(linearity)
4./integraldisplayb
a(f◦φ)(x)dφ
dxdx=/integraldisplayφ(b)
φ(a)f(x)dx(Change of variable in an integral)
Theorem 7.5 1./integraldisplayb
adf
dx(x)dx=f(b)−f(a)
2.d
dx/integraldisplayx
af(t)dt=f(x)
3./integraldisplayb
af(x)dg
dxdx=fg/vextendsingle/vextendsingleb
a−/integraldisplayb
adf
dxg(x)dx(Integration by parts).
Remark: These theorems contain essentially all of elementary calculus. What are missing
are specific formulas for the derivatives and integrals of the basic functions as well as the
application of these theorems to compute maxima, area, etc.
Exercises
(1) Use the definition of the derivative (as the limit of a difference quotient) to compute
the derivatives of the following functions at the given point.
a). 3x2−x+ 1, x0= 2
b).1
x+1, x 0= 2
c).x
1+x, x 0= 2
d).x
1−x, x =x0/negationslash= 1.
7.2. GENERALITIES ON MAPPINGS FROM RNTORM. 295
(2) Use the definition of the integral to evaluate
/integraldisplay2
0x2dx.
You should approximate the area by rectangular strips and evaluate the limit as the
width of the thickest strip tends to zero. [Hint: 12+22+32+···+n2=n(n+1)(2 n+1)
6].
(3) Prove that
.6<log 2<.8 (log 2 = 0 .693)
by using the definition of the integral to find upper and lower bounds for
log 2 =/integraldisplay2
11
xdx.
(4) Find the equation of the straight line which is tangent to the curve f(x) =x7/3+ 1
atx= 1 . Draw a sketch indicating both the curve and tangent line. Use the tangent
line to approximately evaluate (1 .01)7/3. Find some estimate for the error in your
approximation.
7.2 Generalities on Mappings from RntoRm.
A function, or operator, Fwhich maps RntoRm, is a rule which assigns to each vector
XinRnanother vector Y=F(X) inRm. It is a function from vectors to vectors, a
vector-valued function of a vector. We have already discussed the case when Fis an affine
operator,
Y=F(X) =b+LX
or in coordinates,
y1=b1+a11x2+···+a1nx1n
y2=b2+a21x2+···+a2nxn
·
·
·
ym=bm+am1x2+···+amnxn
Linear algebra can be thought of as the study of higher dimensional analytic geometry, the
affine transformations taking the role of the straight line y=b+cx.
But now it is time to consider more complicated mappings from RntoRm. Here is an
Example:/braceleftbiggy1=x1+x2sinπx3
y2=e1−x1−√x2.
This transformation maps vectors X= (x1,x2,x3)∈R3to vectorsY= (y1,y2)∈R2. Note
the second function is only defined for x2≥0 . Thus the domain of the transformation F
is
D(F) =/braceleftbig
X∈R3:x2≥0/bracerightbig
.
For example, Fmaps the point (1 ,4,1
6) into the point (3 ,−1) .
296 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION
It is usual to write a transformation Fwhich maps a set A⊂Rnto a setB⊂Rmin
terms of its components ,
y1=f1(x1,...,x n) =f1(X)
y2=f2(x1,...,x n) =f2(X)
·
·
·
ym=fm(x1,...,x n) =fm(X),
or more concisely as
Y=F(X).
To discuss continuity etc. for nonlinear mappings from RntoRm, it is necessary that
the distance between points be defined. We shall use the Euclidean norm - although any
other norm could also be used. If X= (x1,...,x k) is a point (or vector, if you like) in Rk,
then/bardblX/bardbl=/radicalBig
x2
1+···+x2
k. To review briefly, a sequence of pointsXjinRkconverges
to a pointXinRkif, given any /epsilon1>0 , there is an integer Nsuch that
/bardblXj−X/bardbl</epsilon1 textforall j ≥N.
Anopen ball inRkof radiusrabout the point X0is the setB(X0;r) ={X∈Rk:/bardblX−
X0/bardbl<r}.
Aclosed ball inRkis
¯B(X0;r) ={X∈Rk:/bardblX−X0/bardbl ≤r}.
The only difference is the open ball does not contain the boundary of the ball. In two
dimensions, R2, the names open and closed discare often used.
A setD⊂Rkisopen if each point X∈Dis the center of some ball contained entirely
withinD. The radius may be very tiny. Every open ball is open, as can be seen in the
figure. A closed ball is not open since there is no way of placing a small ball about a point
on the boundary in such a way that the small ball is inside the larger one. A set Ais
closed if it contains allof its limit points , that is, if the points Xj∈Aconverge to a point
X, X j→X, thenXis also inA. An open ball is not closed, for a sequence of points
in the ball may converge to a point on the boundary, and the boundary points are not in
the ball. For the special case of R1, these notions coincide with those of open and closed
intervals. Again, sets - like doors - may be neither open nor closed.
A point set Disbounded if it is contained in some ball (of possibly large radius). The
pointXisexterior toDifXdoes not belong to Dand if there is some ball about X
none of whose point are in D.Xisinterior toDifXbelongs to Dand there is some
ball about Xall of whose points are in D.Xis aboundary point ofDif it is neither
interior nor exterior to D. Note that a boundary point of Dmay or may not belong to
D. For example, the boundaries of the open and closed balls B(0;r),¯B(0;r) are the same.
The boundary of a set Dis denoted by ∂D. It is evident that a set is open if and only
if every point is an interior point, and a set is closed if and only if it contains all of its
boundary points.
Definition: LetAbe a set in RnandCa set in Rm. The function F:A→Cis
continuous at the interior point X0∈Aif, given any radius /epsilon1>0 , there is a radius δ>0
7.2. GENERALITIES ON MAPPINGS FROM RNTORM. 297
such that
/bardblF(X)−F(X0)/bardbl</epsilon1 textforall /bardblX−X0/bardbl<δ.
[Observe the norm on the left is in Rmwhile that on the right is in Rn].
It is easy to prove
Theorem 7.6 . An affine mapping F(X) =b+LX from RntoRmis continuous at
every point X0∈Rm.
Proof: First,F(X)−F(X0) =b+LX−b−LX 0=L(X−X0) . Thus,
/bardblF(X)−F(X0)/bardbl=/bardblL(X−X0)/bardbl.
Let ((aij)) be a matrix representing Lwith respect to some bases for RnandRm. Then
by Theorem 17, p. 373
/bardblL(X−X0)/bardbl2=/angbracketleftL(X−X0), L(X−X0)/angbracketright ≤k/bardblL(X−X0)/bardbl/bardblX−X0/bardbl,
where
k2=m/summationdisplay
i=1n/summationdisplay
j=1a2
ij.
Therefore
/bardblL(X−X0)/bardbl ≤k/bardblX−X0/bardbl.
It is now clear that if X→X0, thenL(X−X0)→0 . More formally, given any /epsilon1>0 , if
δ=/epsilon1
k+1, we have
/bardblF(X)−F(X0)/bardbl</epsilon1 textforall /bardblX−X0/bardbl<δ.
The following theorems have the same proofs as were given earlier for special cases. (See a
first year calculus book and our Chapter 0).
Theorem 7.7 . LetF1andF2mapA⊂RnintoC⊂Rm. IfF1andF2are continuous
at the interior point X0∈A, then
1.aF1+bF2is continuous at X0.
2./angbracketleftF1, F2/angbracketrightis continuous at X0.
Theorem 7.8 . LetF= (f1,...,f m)mapA⊂RnintoC⊂Rm. ThenFis continuous
at the interior point X0∈Aif and only if each of the fj, j= 1,...,m , is continuous at
X0.
Theorem 7.9 . LetF:A→C, whereAis a closed and bounded (= compact) set. If F
is continuous at every point of A, then it is bounded; that is, there is a constant Msuch
that/bardblF(x)/bardbl ≤Mfor allX∈A. Moreover, if M0is the least upper bound, then there is
a pointX0∈Asuch that /bardblF(X0)/bardbl=M0. Similarly, if m0is the greatest lower bound
for/bardblF/bardbl, then there is a point X1∈Asuch that /bardblF(X1)/bardbl=m0.
There is nothing better than to close this otherwise unauspicious section with one of
the crown jewels of mathematics - the Fundamental Theorem of Algebra, all of whose proofs
require the non-algebraic notion of continuity. Let
p(z) =a0+a1z+···+anzn, (n≥1),
where theaj’s are complex numbers and anis not zero.
For every complex number z, the value of the function p(z) is a complex number.
Thusp:C→C. We want to prove there is at least one z0∈Csuch thatp(z0) = 0 .
298 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION
Lemma 7.10 .p(z)is a continuous function for every z∈C.
Proof: Identical to the proof that a real polynomial is continuous everywhere.
Lemma 7.11 LetDbe a set in the complex plane in which p(z)/negationslash= 0. The minimum
modulus of p(z), that is, the minimum value of |p(z)|, cannot occur at an interior point
ofD. It must occur on the boundary ∂DofD.
Proof: Letz0be any interior point of D. Rewritep(z) in the form
p(z) =b0+b1(z−z0) +···+bn(z−z0)n.
Sincep(z0)/negationslash= 0 , we know b0/negationslash= 0 . Also, because pis not identically constant, at least one
coefficient following b0is not zero. Take bkto be the first such coefficient. We must write
b0, bkandz−z0in polar form,
b0=ρ0euαbk=ρ1eiβz−z0=ρeiθ,
whereρ0=|p(z0)|,ρ1andρare positive real numbers. Here we are restricting zto a
point on a circle of radius ρaboutz0, after taking ρsmall enough to insure this circle is
interior to D. Then
p(z) =ρ0eiα+ρ1eiβρkeikθ+bk+1(z−z0)k+1+···+bn(z−z0)n
=ρ0eiα+ρ1ρkei(β+kθ)+ (z−z0)k+1[bk+1+···+bn(z−z0)n−k−1].
Pick the particular point ˆ zon the circle whose argument θis given by β+kθ=α+π.
Thenei(β+kθ)=ei(α+π)=−eiα, so
p(ˆz) = (ρ0−ρ1ρk)eiα+ (ˆz−z0)k+1[bk+1+···+bn(ˆz−z0)n−k−1].
By the triangle inequality we find
|p(ˆz)| ≤/vextendsingle/vextendsingle/vextendsingleρ0−ρ1ρk/vextendsingle/vextendsingle/vextendsingle+ρk+1[|bk+1|+···+|bn|ρn−k−1].
Choose the radius ρso small that ρ0−ρ1ρk≥0 . Then
|p(ˆz)| ≤ρ0−ρ1ρk+ρk+1[|bk+1|+···+|bn|ρn−k−1].
By choosing ρsmaller yet, if necessary, we can make the term ρ[|bk+1|+···+|bn|ρn−k−1]<
1
2ρ1. Consequently,
|p(ˆz)| ≤ρ0−ρ1ρk+1
2ρ1ρk=ρ0−1
2ρ1ρk
<ρ 0=|p(z0)|.
Thus, ifz0is any interior point of a domain Din whichpdoes not vanish, then there is
a point ˆzalso interior to Dsuch that |p(z)|<|p(z0)|. Therefore, the minimum of |p(z)|
must occur on the boundary of any set in which pdoes not vanish.
Lemma 7.12 . Given any real number M, there is a circle |z|=Ron which |p(z)|>M
for allz,|z|=R.
7.2. GENERALITIES ON MAPPINGS FROM RNTORM. 299
Proof: Forz/negationslash= 0 , we can write the polynomial p(z) as
p(z)
zn=an+an−1
z+···+a0
zn.
From the triangle inequality written in the form |f1+f2| ≥ |f1| − |f2|, we find
/vextendsingle/vextendsingle/vextendsingle/vextendsinglep(z)
zn/vextendsingle/vextendsingle/vextendsingle/vextendsingle≥ |an| −/vextendsingle/vextendsingle/vextendsinglean−1
z+···+a0
zn/vextendsingle/vextendsingle/vextendsingle.
If|z|is taken large enough, |z| ≥R0, it is possible to make the second term on the right
less than |an|/2 ,
/vextendsingle/vextendsingle/vextendsinglean−1
z+···+a0
zn/vextendsingle/vextendsingle/vextendsingle</vextendsingle/vextendsingle/vextendsinglean
2/vextendsingle/vextendsingle/vextendsingle, texton |z|=R>R 0
Therefore, for |z|=R≥R0
/vextendsingle/vextendsingle/vextendsingle/vextendsinglep(z)
zn/vextendsingle/vextendsingle/vextendsingle/vextendsingle≥ |an| −/vextendsingle/vextendsingle/vextendsinglean
2/vextendsingle/vextendsingle/vextendsingle=1
2|an|,
so
|p(z)| ≥1
2|an|Rn, texton |z|=R.
It is now clear that by choosing Rsufficiently large, |p(z)|can be made to exceed any
constantMon the circle |z|=R.
Theorem 7.13 (Fundamental Theorem of Algebra). Let
p(z) =a0+a1z+···+anzn, a n/negationslash= 0,n≥1,
be any polynomial with possibly complex coefficients, a0,a1,...,a n. Then there is at least
one number z0∈Csuch thatp(z0) = 0 . In other words, every polynomial has at least one
complex root.
Proof: By Lemma 3, we can find a large circle |z|=R, on which |p(z)|>2|a0|for all
|z|=R. Sincep(z) is a continuous function, by Theorem 4 there is a point z0in the closed
and bounded disc |z| ≤Rfor which |p|attains its minimum value m0,|p(z0)|=m0. If
p(z0) = 0 , we are done. However if pdoes not vanish inside the closed disc, by the
important Lemma 2 its minimum value is attained only on the boundary, so z0is on the
circle |z0|=R. But on the circle we know |p(z0)|>2|a0|= 2|p(0)|, so the minimum is
not atz0after all. The assumption that pdoes not vanish in the disc |z| ≤Rhad led us
to a contradiction. Notice the proof does not give a procedure for finding the root whose
existence has been proved.
Exercises
1. Prove Theorem 2, part 1.
2. Use the Fundamental Theorem of Algebra along with the “factor theorem” of high
school algebra to prove that a polynomial of degree nhas exactly nroots (some of which
may be repeated roots).
300 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION
7.3 Mapping from E1toEn
.
As a particle moves along a curve γinEnits position F(t) at timetcan be specified
by a vector
X=F(t) = (f1(t),f2(t),...,f n(t)),
wherexj=fj(t) is thejthcoordinate of the position at time t. Thus, the curve is
specified by F(t) , a mapping from numbers to vectors, F:A⊂E1→En, whereAis the
domain of definition of F.
For example, the mapping
F:t→(cosπt,sinπt,t), t∈(−∞,∞)
which may also be written as
F(t) = (cosπt,sinπt,t)
can be thought of as describing the motion of a particle along a helix.
It is natural to ask about the velocity, which means derivative must be defined.
Definition: LetF(t) define a curve γfortin the interval A= [a,b] . Consider the
difference quotient
F(t+h)−F(t)
h, t textand t +hinA,
wheretis fixed. If this vector has a limit as htends to zero, then Fis said to have a
derivativeF/prime(t) att,
F/prime(t) = lim
h→0F(t+h)−F(t)
h,
while the curve has slopeF/prime(t) att. Some other common notations are
˙F(t),dF
dt, D tF.
The curve γis called smooth if i) the derivative F/prime(t) exists and is continuous for each t
in [a,b] , and if ii) /bardblF/prime(t)/bardbl /negationslash= 0 for any point tin [a,b] .
Iftrepresents time, then F/prime(t) is the velocity of the particle at time twhile /bardblF/prime(t)/bardbl
is the speed .
IfF(t) is given in terms of coordinate functions, F:t→(f1(t),...,f n(t)) , how can
the derivative of Fbe computed?
Theorem 7.14 . IfF(t) = (f1(t),...,f n(t))is a differentiable mapping of A⊂E1into
En, then the coordinate functions are differentiable and
dF
dt=/parenleftbiggdf1
dt,df2
dt,···,dfn
dt/parenrightbigg
.
Conversely, if the coordinate functions are differentiable, then so is F(t)and the derivative
is given by the above formula.
7.3. MAPPING FROM E1TOEN301
Proof: Iftandt+hare both in A, then
F(t+h)−F(t)
h=1
h[(f1(t+h),...,f n(t+h))−(f1(t),...,f n(t))]
=/parenleftbiggf1(t+h)−f1(t)
h,···,fn(t+h)−fn(t)
h/parenrightbigg
Since the limit as h→0 of the expression on the left exists if and only if all of the limits
lim
h→0fj(t+h)−fj(t)
h, j = 1,...,n
exist, the theorem is proved.
Examples:
(1) IfF:t→(cosπt,sinπt,t), t∈(−∞,∞), F is differentiable for all tsince each of
the coordinate functions are differentiable. Also,
F/prime(t) = (−πsinπt,π cosπt,1).
In addition, the curve - a helix - which Fdefines is smooth since F/primeis continuous
and
/bardblF/prime(t)/bardbl −/radicalbig
π2sin2πt+π2cos2πt+ 1 =/radicalbig
π2+ 1/negationslash= 0,
(2) LetF:t→(a1+b1t,a2+b2t,a3+b3t) =P+QtwhereP= (a1,a2,a3) and
Q= (b1,b2,b3) are constant vectors. Then the curve Fdefines is a straight line
which passes through the point P= (a1,a2,a3) att= 0 .Fis differentiable for all
t, since each of the coordinate functions are differentiable. Furthermore,
F/prime(t) =Q= (b1,b2,b3),
aconstant vector pointing in the direction Q= (b1,b2,b3) , as is anticipated for a
straight line. Because
/bardblF/prime(t)/bardbl=/bardblQ/bardbl=/radicalBig
b2
1+b2
2+b2
3,
this curve is smooth except in the degenerate case b1=b2=b3= 0 , that is, Q= 0 ,
when the curve degenerates to a single point, F(t) = (a1,a2,a3) =P.
(3) The curve defined by the mapping F:t→(t,|t|) is differentiable everywhere and
/bardblF/prime(t)/bardbl /negationslash= 0 except at t= 0 . It is not differentiable there since the second coordinate
function,f2(t) =|t|is not differentiable at t= 0 . Thus, the curve is smooth except
att= 0 .
(4) The curve defined by the mapping F:t→(t3,t2) is differentiable everywhere, and
F/prime(t) = (3t2,2t).
However, /bardblF/prime(t)/bardbl=√
9t4+ 4t2, so the curve is smooth everywhere except at t= 0 ,
which corresponds to a cusp at the origin in the x1,x2plane.
302 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION
It is elementary to compute the derivative of the sum of two vectors. The derivative
of a product can be defined for the inner product, and for the product with scalar-valued
function.
Theorem 7.15 . IfF(t)andG(t)both map an interval A⊂E1intoEn, and are both
differentiable there, then for all t∈A,
1.d
dt[aF+bG] =adF
dt+bdG
dt(linearity of the derivative).
2.d
dt/angbracketleftF, G/angbracketright=/angbracketleftF/prime, G/angbracketright+/angbracketleftF, G/prime/angbracketright, (in “dot product” notation:d
dt(F·G) =F/prime·G+F·G/prime).
Proof: Since these are identical to the proofs of the corresponding statements for scalar-
valued functions, we prove only the second statement.
d
dt/angbracketleftF(t), G(t)/angbracketright= lim
h→01
h[/angbracketleftF(t+h), G(t+h)/angbracketright − /angbracketleftF(t), G(t)/angbracketright]
= lim
h→01
h[/angbracketleftF(t+h)−F(t), G(t+h)/angbracketright+/angbracketleftF(t), G(t+h)−G(t)/angbracketright]
= lim
h→0/bracketleftbigg/angbracketleftF(t+h)−F(t), h/angbracketright
G(t+h)+/angbracketleftF(t),G(t+h)−G(t)
h/angbracketright/bracketrightbigg
=/angbracketleftF/prime(t), G(t)/angbracketright+/angbracketleftF(t), G/prime(t)/angbracketright.
An interesting and simple consequence is the fact that if a particle moves on a curve
F(t) which remains a fixed distance from the origin, /bardblF(t)/bardbl ≡ constant = c, then the
velocity vector F/primeis always orthogonal to the position vector F. This follows from
c2=/bardblF(t)/bardbl2=/angbracketleftF(t), F(t)/angbracketright,
so taking the derivative of both sides we find
0 =/angbracketleftF/prime, F/angbracketright+/angbracketleftF, F/prime/angbracketright= 2/angbracketleftF, F/prime/angbracketright.
Thus /angbracketleftF, F/prime/angbracketright= 0 for all t, an algebraic statement of the orthogonality. As a particular
example, the mapping
F(t) = (cosπ
1 +t2,sinπ
1 +t2)
has the property /bardblF(t)/bardbl= 1 for all t. You can see the path of the particle in the figure.
Att= 0 the particle is at ( −1,0) . As time increases, the particle moves along an arc of
the unit circle toward (1 ,0) , reaching (0 ,1) att= 1 . The velocity at time tis
F/prime(t) =2πt
(1 +t2)2(sinπ
1 +t2,−cos−π
1 +t2).
From this expression, it is evident the particle slows down as it approaches (1 ,0) . In fact,
the particle never does manage to reach (1 ,0) .
We would like to define the notion of a straight line which is tangent to a smooth
curve at a given point. There is one touchy issue. You see, the curve may intersect itself,
thus having two or more tangents at the same point. Once acknowledged, the difficulty is
resolved by realizing that for each value of t, there is a unique point F(t) on the curve.
X0is a double point if F(t1) =F(t2) =X0.
7.3. MAPPING FROM E1TOEN303
By picking one value of t, there will be a unique tangent line to the curve for this value
oft. Thus, we define the tangent line for t=t1to the curve defined by a differentiable
functionF(t) as the straight line whose equation is
A(t) =F(t1) +F/prime(t1)(t−t1).
Att=t1, the curves defined by F(t) andA(t) have the same value F(t1) =X0and the
same derivative (slope), F/prime(t) .
Example: Consider the curve defined by the mapping F:t→(3 +t3−t,t2−t), t∈
(−∞,∞) . The point (3 ,0) is a double point since F: 0→(3,0) andF: 1→(3,0) .
Thus, the line tangent to the point (3 ,0) whent= 1 is defined by
A(t) = (3,0) + (2,1)(t−1) = (3,0) + (2(t−1),(t−1))
or
A(t) = (1,−1) + (2t,t).
Since we are still working with functions F(t) of one real variable t, the mean value
theorem and chain rule follow immediately by applying the corresponding theorems for
scalar valued functions to each of the components f1(t),...,f n(t) ofF(t) .
Theorem 7.16 (Approximation Theorem and Mean Value Theorem). If the vector valued
functionF(t)is continuous for t∈[a,b]and differentiable for t∈(a,b)then fort0∈
(a,b),
1.F(t) =F(t0) +dF
dt/vextendsingle/vextendsingle
t0(t−t0) +R(t,t0)|t−t0|where
lim
t→t0/bardblR(t,t0)/bardbl= 0.
2. There is a point τbetweentandt0such that
/bardblF(t)−F(t0)/bardbl ≤ /bardblF/prime(τ)/bardbl|t−t0|.
3. IfF= (f1,...,f n), there are points τ1,...,τ nbetweentandt0such that
F(t) =F(t0) +L(t−t0),
whereLis the linear transformation
L= (f/prime
1(τ1),f/prime
2(τ2),...,f/prime
n(τn))
Remark: Although 1 and 3 follow from the one variable case f(t) —and will be proved
again in greater generality later on - the proof of 2 is difficult under our weak hypothesis. If
the stronger assumption, Fis continuously differentiable, is made, then 2 becomes easy, and
the factor /bardblF/prime(τ)/bardblcan be replaced by a constant M= max
τ∈[a,b]/bardblF/prime(τ)/bardbl, since a continuous
function /bardblF/prime(τ)/bardbldoes assume its maximum if τis in a closed and bounded set, τ∈[a,b] .
Corollary 7.17 . IfFsatisfies the hypotheses of Theorem 8 and if F/prime(t)≡0for all
t∈[a,b], thenFis a constant vector.
Proof: Just look at 2 or 3 above to see that for any points t,t0in [a,b] , we have
F(t) =F(t0) .
304 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION
Theorem 7.18 (Chain Rule). Consider the vector-valued function F(t)which is differ-
entiable for t∈(a,b), and the scalar valued function φ(s)which is differentiable for
s∈(α,β). If the range of φis contained in (a,b),R(φ)⊂(a,b), then the composed func-
tionG(s) = (F◦φ)(s) =F(φ(s))is differentiable as a function of sfor allsin(α,β)
and
G/prime(s) =F/prime(φ(s))φ/prime(s),
that is,
dG
ds(s) =dF
dφ(φ)dφ
ds(s) =dF
dt(t)/vextendsingle/vextendsingle/vextendsingle/vextendsingle
t=φ(s)dφ
ds(s).
IfF(t) = (f1(t),...,f n(t)), then
G(s) =F((s)) = (f1(φ(s)),...,f n(φ(s)),and
G/prime(s) =)f/prime
1(φ)φ/prime(s),...,f/prime
n(φ)φ/prime(s))
= (f/prime
1(φ),...,f/prime
n(φ))φ/prime(s).
Proof not given here . It is the same as that given in elementary calculus for n= 1 . A more
general theorem containing this one is proved later (p. 701).
Examples:
1. IfF(t) = (1 −t2,t3−sinπt) andφ(s) =e−s, thenG(s) = (F◦φ)(s) = (1 −
e−2s,e−3w−sinπe−2) . We compute G/prime(s) in two distinct ways, using the chain rule, and
directly from the formula for G(s) . By the chain rule:
G/prime(s) =F/prime(t)/vextendsingle/vextendsingle
t=φ(s)φ/prime(s)
= (−2t,3t2−πcosπt)/vextendsingle/vextendsingle
t=e−s(−)e−s
=−(−2e−s,3e−2s−πcosπe−s)e−s,
In particular, at s= 0 , since t= 1 whens= 0 , we find
G/prime(0) = −(−2,3 +π) = (2,−3,−π)
Directly from the formula for G(s) = (1 −e−2s,−3e−3s−sinπe−s),we find
G/prime(s) = (2e−2s,−3e−3s+πe−scosπe−2),
which agrees with the chain rule computation.
Since the derivative F/prime(t) of a function F(t) from numbers to vectors, F:E1→En,
is also a function of the same type, the second and higher order derivatives can be defined
inductively;
d2
dt2F(t) :=d
dtF/prime(t),dk+1
dtk+1F(t) :=d
dtF(k)(t).
7.3. MAPPING FROM E1TOEN305
Example: IfF:t→(cosπt,sinπt,t) , then
F/prime/prime(t) =d
dt(−πsinπt,π cosπt,1)
= (−π2cosπt,−π2sinπt,0).
IfF(t) represents the position of a particle at time t, thenF/prime/prime(t) is the acceleration
of the particle at time t. All of these ideas were used in the last two sections in Chapter 6
where linear systems of ordinary differential equations were encountered. Time permitting,
a second application to a non-linear system of O.D.E.’s will be treated in Section of Chapter.
There another of the crown jewels in the intellectual history of mankind will be discussed:
Newton’s incredible solution of “the two body problem”, that is, to determine the motion
of the heavenly bodies.
Recall that the length of a curve is defined to be the limit of the lengths of inscribed
polygons which approximate the curve as the length of the longest subinterval tends to
zero - if the limit does exist. Let the curve γ, which we assume is smooth, be determined
by the function F(t),t∈[a,b] . Then the length of the straight line joining F(tj) to
F(tj+ ∆tj),tj+1=tj+ ∆tj, is
/bardblF(tj+ ∆tj)−F(tj)/bardbl=/bardblF(tj+ ∆tj)−F(tj)
∆tj/bardbl∆tj
Adding up the lengths of these segments and letting the largest ∆ tjtend to zero, we find
the length of γis given by
L(γ) =/integraldisplayb
1/bardblF/prime(t)/bardbldt.
If the function Fis defined through coordinates, F(t) = (f1(t),...,f n(t)) , this formula
reads
L(γ) =/integraldisplayb
a/radicalBig
f/prime2
1+f/prime2
2+···+f/prime2ndt.
You will recognize the special case where F(t) = (x(t),y(t))
L(γ) =/integraldisplayb
a/radicalbig
x2+ ˙y2dt.
Example: Find the length of the portion of the helix γdefined by F(t) = (cost,sint,t) ,
fort∈[0,2π] . This is one “hoop” of the helix. Since F/prime(t) = (−sint,cost,1) , we have
/bardblF/prime(t)/bardbl=/radicalbig
sin2t+ cos2t+ 1 =√
2 , so the length is
L(γ) =/integraldisplay2π
0√
2dt= 2π√
2.
For eacht∈[a,b] , we can define an arc length function s(t) , the arc length from ato
t, by
s(t) =/integraldisplayt
a/bardblF/prime(τ)/bardbldτ.
Note we are using a dummy variable of integration τ. By the fundamental theorem of
calculus, we have
ds
dt=/bardblF/prime(t)/bardbl
306 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION
Sinceds/dt can be thought of as the rate of change of arc length with respect to time, it
is the speed of a particle moving along the curve , the tangential speed .
The integral used in arc length is the integral of a scalar- valued function /bardblF/prime(t)/bardbl.
How can we define the integral of a vector-valued function F(t) = (f1(t),...,f n(t)) ? Just
integrate each component, assuming they are all integrable of course,
/integraldisplayb
aF(t)dt:= (/integraldisplayb
af1(t)dt,...,/integraldisplayb
afn(t)dt).
For example, if F(t) = (t−3t2,1−√
2t,e3t) , then
/integraldisplay2
0F(t)dt= (/integraldisplay2
0(t−3t2)dt,/integraldisplay2
0(1−√
2t)dt,/integraldisplay2
0e3tdt)
= (−4,−2
3,e6−1).
We give no physical interpretation of the integral (as an area or the like) except in the case
whereF(t) represents the velocity of a particle. Then/integraldisplayb
aF(t)dtis the vector pointing
from the position at t=ato the position at t=b.
Exercises
(1) (a) Describe and sketch the images of the curves F:E1→E2defined by
(i)F(t) = (2t,3−t)
(ii)F(t) = (2t,|3−t|)
(iii)F(t) = (t2,1 +t2)
(iv)F(t) = (2t,sint)
(v)F(t) = (t2,1 +t4)
(b) Which of the above mappings are differentiable and for what value(s) of t? Find
the derivatives if the functions are differentiable. Which of the curves defined by
these mappings are smooth, and where are they not smooth?
(2) Use the definition of the derivative to find F/prime(t) att= 2πfor the functions
a).F(t) = (2t,3−t)t∈(−∞,∞).
b).F(t) = (1 +t2,sin 2t). t∈(−∞,∞).
(3) Find the lengths of the curves γdefined by the mappings
a).F(t) = (a1+b1t,a2+b2t,...,a n+bnt),=P+Qt, t ∈[0,1].
b).F(t) = (sin 2t,1−3t,cos 2t,2t3/2),t∈[−π,2π]
(4) Consider the curve defined by the equation
F(t) = (t−t2,t4−t2+ 1), t∈(−∞,∞)
a). Sketch the curve.
b). Where does the curve intersect itself?
c). Find the line tangent to the curve at the image of t= 1 .
7.3. MAPPING FROM E1TOEN307
(5) IfF:A⊂E1→Enis twice continuously differentiable and F/prime/prime(t)≡0 for allt∈A,
what can you conclude? Please prove your assertion. [Hint: First consider the special
case where F:E1→E1].
(6) LetF(t) be a twice differentiable function which maps a set in E1intoEnand
satisfies the ordinary differential equation F/prime/prime+µF/prime+kF= 0 , where kandµare
positive constants. Define the energy as
E(t) =1
2/bardblF/prime/bardbl2+1
2k/bardblF/bardbl2
(a) Prove E(t) is a non-increasing function of t(energy is dissipated). [Hint:
dE/dt =? ].
(b) IfF(0) = 0 and F/prime(0) = 0 , prove E(t)≡0 .
(c) Prove there is at most one function which satisfies the given differential equation
as well as the initial conditions F(0) =A, F/prime(0) =B, whereAandBare
given vectors.
(7) IfF(t) = (1 −e2t,t3,1
1+t2) , andφ(x) =1
1+x, x > −1 , computed
dx(F◦φ)(x) by
using the chain rule.
(8) Compute d2F/dt2for the function F(t) in Exercise 7.
(9) (a) Show that the equation of a straight line which passes through the point P1at
t= 0 andP2att= 1 is
F(t) =P1+ (P2−P1)t.
(b) Find the equation of a straight line which passes through the point P1= (1,2,3)
att= 0 andP2= (1,−5,0) att= 1 .
(c) Find the equation of a straight line which passes through the point P1att=t1
andP2att=t2.
(d) Apply this to find the equation of a straight line which passes through P1=
(−3,1,−2) att=−1 andP2= (0,2,1) att= 2 . What is the slope of this
line?
(10) Given a smooth curve all of whose tangent lines pass through a given point, prove
that the curve is a straight line.
(11) LetF:E1→Endefine a smooth curve which does not pass through the origin. Show
that the position vector F(t) is orthogonal to the velocity vector at the point of the
curve which is closest to the origin. Apply this to prove anew the well known fact
that the radius vector to any point on a circle is perpendicular to the tangent vector
at that point. [Hint: Why is it sufficient to minimize ϕ(t) =/angbracketleftF(t), F(t)/angbracketright?]
308 CHAPTER 7. NONLINEAR OPERATORS: INTRODUCTION
Chapter 8
Mappings from Ento E: The
Differential Calculus
8.1 The Directional and Total Derivatives
.
Throughout this and the next chapter we shall consider functions which map Enor a
portion of it A, into E. By the statement
f:A→E, A ⊂En
we mean that to every vector XinA, the function (operator, map, transformation) assigns
a unique real number w. Thusw=f(X) in this case is a map from vectors to numbers.
Two particular examples prove helpful in thinking conceptually about mappings of this
type.
(1)The temperature function .f:A→E, where the set A⊂E3is the room in which
you are sitting. To every point Xin the room, A, this function fassigns a number
- the temperature f(X) atX, w =f(X) .
(2)The height function .f:A→E, where the set Ais some set in the plane E2. To every
pointXinA, this function fassigns a number - the height f(X) of a surface (or
manifold)Mabove that point. Thus, the set of all pairs ( X,f(x)), X∈A, defines
a portion of a surface, a surface in E2×E∼=E3.
From the second example, it is clear that every function f:A⊂En→Emay be
regarded as the graph of a surface in En×E∼=En+1, the surface being regarded as all
points in En+1of the form ( X,f(X)) , where X∈A. For example, the temperature
function can be thought of as the graph of a surface in E4, the height of the surface
w=f(X) aboveXbeing the temperature at X. (Compare with the discussion from p.
322 bottom, to p. 324).
In concrete situations, the point X∈Enis specified by giving its coordinates with
respect to some fixed bases for EnandE. The particular coordinate system used depends
on the geometry of the problem at hand. Rectangular symmetry calls for the standard
309
310CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
rectangular coordinates, while polar coordinates are well suited to problems with circular
symmetry. We shall meet these issues head-on a bit later.
IfX= (x1,...,x n) with respect to some coordinates for En, then we write w=
f(X) =f(x1,...,x n) . The points ( X,f(X)) on the graph are (x1,...,x n, f(x1,...,x n)) ,
which we may also write as ( x1,...,x n,f) or else as ( x1,...,x n,w) . For low dimensional
spaces, E2orE3, it is convenient to avoid subscripts. In these situations we shall write
w=f(x,y) andw=f(x,y,z ) for mappings with domains in E2andE3, respectively.
We now examine some more specific examples.
Examples:
(1)w=−1
2x+y−1 . This function assigns to every point X= (x,y) inE2a numberw
inE. We can represent the function, an affine mapping from E2→E, as the graph
of a plane in E3. The linear nature of the plane reflects the fact that the mapping
is an affine mapping - a linear mapping except for a translation of the origin. More
generally, the function w=α+a1x1+a2x2+···+anxn, an affine mapping from
En→E, represents a plane in En+1. In fact, this can be taken as the algebraic
definition of a plane in En+1. These affine functions are the simplest functions which
mapEnintoE. Although we shall not, it is customary to abuse the nomenclature
and refer to affine mappings as being linear. This is because they share most of
the algebraic and geometric properties of proper linear mappings, as opposed to the
honestly nonlinear mappings we will be treating as in the next examples.
(2)w=x2+y2. This function assigns to every point X= (x,y) inE2a real number
w∈E. We can represent the function as the graph of a paraboloid of revolution,
obtained by rotating the parabola w=x2about the waxis. If this paraboloid is
cut by a plane parallel to the x,yplane, say w= 2 , the intersection of these two
surfaces is the circle x2+y2= 2 .
(3)w=−x2+y2. This function can be represented as the graph of a very fancy surface
- ahyperbolic paraboloid . If this surface is cut by a plane parallel to the x,yplane,
w=c, the intersection is the curve c=−x2+y2. Forc>0 , this curve is a hyperbola
which opens about the yaxis, while if c<0 , the curve is a hyperbola which opens
about the xaxis. Forc= 0 we obtain two straight lines, x=+−y(see fig). The
intersection of the surface with the plane x=cis a parabola which opens upward
in they,w plane. Similarly, the intersection of the surface with the plane y=cis
a parabola which opens downward in the xwplane. This curve is rightly called a
saddle , and the origin (0 ,0,0) a saddle point (or mountain pass) since a particle can
remain at rest at that point, or ii) move on the surface in one direction and go up, or
iii) move on the surface in another direction and go down.
Letf(X) be a function from vectors to numbers,
f:A⊂En→E.
How can we define the notion of derivative for such functions? The derivative should
measure the rate of change of f(X) asXmoves about. But if you think of f(X)
as the temperature function, it is clear that the temperature will change at different
rates depending which direction you move. Thus, if you move across the room in
8.1. THE DIRECTIONAL AND TOTAL DERIVATIVES 311
the direction of the door, the temperature may decrease, while if you move up to the
ceiling, the temperature will likely increase. Thus, the natural notion of a derivative
is the rate of change in a particular direction - a directional derivative .
LetX0denote your position and f(X0) the temperature there. Take ηto be a free
vector, which we shall think of as pointing from X0toX0+η. We want to define the rate
at which the temperature changes as you move from X0in the direction ηtowardX0+η.
Since all points on the line joining X0toX0+ηare of the form X0+λη, whereλis a
real number, the difference f(X0+λη)−f(X0) is the difference between the temperatures
atX0+ληand atX0.
Definition: Letf:A⊂En→E. The derivative offat the interior point X0∈Awith
respect to the vector ηis
f/prime(X0;η) = lim
λ→0f(X0+λη)−f(X0)
λ,
if the limit exists.
In the special case when η=eis aunit vector ,/bardble/bardbl= 1 , we see that λ=/bardblλe/bardbl. Then
Def(X0) :=f/prime(X0;e) is the instantaneous rate of change of fper unit length asXmoves
fromX0towardX0+ 3 . This normalization to using only unit vectors is necessary to
have a meaningful definition of a directional derivative. Thus, the directional derivative of
fatX0∈Ain the direction of the unit vector eis the derivative with respect to the unit
vectore. It measures how fchanges as you move from X0to a point on the unit sphere
aboutX0. For theoretical purposes, the derivative of fwith respect to any vector ηis
useful, while for practical purposes, the more restrictive notion of the directional derivative
is needed.
Example: 1 Find the directional derivative of f(X) =x2
1−2x1x2+ 3x2atX0= (1,0) in
the direction η= (−1,1) . Note that ηis not a unit vector. The unit vector is e=η
/bardblη/bardbl=
(−1√
2,1√
2) . Then
X0+λe= (1,0) +λ(−1√
2,1√
2) = (1 −λ√
2,λ√
2),
so
f(X0+λe) = (1 −λ√
2)2−2(1−λ√
2)(λ√
2) + 3(λ√
2)
= 1−λ√
2+3
2λ2.
Thus,
f(X0+λe)−f(X0)
λ=1−λ√
2+3
2λ2−1
λ=−1√
2+3
2λ.
Therefore, the directional derivative Defis
Def(X0) = lim
λ→0f(X0+λe)−f(X0)
λ=−1√
2.
In words, the rate of change of fatX0in the direction of the unit vector eis−1√
2. One
qualitative conclusion we arrive at is that f(X) decreases as Xmoves from X0in the
directione.
312CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
2. Compute f/prime(X;η) iff(X) =/angbracketleftX, AX /angbracketright, whereAis a self-adjoint transformation.
f(X+λη) =/angbracketleftX+λη, A (X+λη)/angbracketright
=/angbracketleftX, AX /angbracketright+λ/angbracketleftη, AX /angbracketright+λ/angbracketleftX, Aη /angbracketright+λ2/angbracketleftη, Aη/angbracketright
and sinceAis self-adjoint,
=/angbracketleftX, AX /angbracketright+ 2λ/angbracketleftAX, η /angbracketright+λ2/angbracketleftη, Aη/angbracketright.
Thus,
f/prime(X,η) = lim
λ→0f(X+λη)−f(X)
λ= 2/angbracketleftAX, η /angbracketright.
In particular when A=Iis the identity operator, f(X) =/bardblX/bardbl2, we findf/prime(X;η) =
2/angbracketleftX, η/angbracketright
The directional derivatives of fin the particular direction of the coordinate axes e1=
(1,0,...), e2= (0,1,0,...) have special names. They are called the partial derivatives of
f. For example, the partial derivative of f(X) =f(x1,x2,...,x n) atX0with respect to
x2is
∂f
∂x2(X0) :=f/prime(X0;e2) = lim
λ→0f(X0+λe2)−f(X0)
λ
There are many other competing notations, all of them being used. We shall list them
shortly, after observing there is a simple way to compute these partial derivatives. Consider
f(X) =f(x−1,x2,x3) . Then
∂f
∂x2(X) = lim
λ→0f(X+λe1)−f(X)
λ
SinceX+λe1= (x1,x2,x3) +λ(1,0,0) = (x1+λ,x 2,x3) we have
∂f
∂x2(X) = lim
λ→0f(x1+λ,x 2,x3)−f(x1,x2,x3)
λ.
But this is the ordinary derivative of fwith respect to the single variable x1, while holding
the other variables x2andx3fixed. Thus, ∂f/∂x 1can be computed by merely taking the
ordinary one variable derivative of fwith respect to x1, pretending the other variables
are constants.
Example: Iff(X) =x2
1+x1ex1x2, find the rate of change of fat the point X0in the
directionse1= (1,0) ande2= (0,1) . Thus, we want to compute∂f
∂x1(X0) and∂f
∂x2(X0).
∂f
∂x1= 2x1+ex1x2+x1x2ex1x2
∂f
∂x2=x2
1ex1x2
At the pint X0= (2,−1) , we have
∂f
∂x1/vextendsingle/vextendsingle/vextendsingle/vextendsingle
2,−1= 4−e−2,∂f
∂x2/vextendsingle/vextendsingle/vextendsingle/vextendsingle
(2,−1)= 4e−2.
8.1. THE DIRECTIONAL AND TOTAL DERIVATIVES 313
Some common notation. If w=f(x1,x2) , then
∂w
∂x1=∂f
∂x1=D1f=f1=fx1=wx1
∂w
∂x2=∂f
∂x2=D2f=f2=fx2=wx2
Iff:A⊂En→E, then∂f/∂x jis another function of X= (x1,x2,...,x n) . It is
then possible to take further partial derivatives.
Example: Letw=f(X) =x2
1+x1ex1x2as in the previous example. Then
w11=f11=fx1x1=∂2f
∂x2
1=∂
∂x1(∂f
∂x1) = 2 +x2ex1x2=x2ex1x2+x1x2
2ex1x2
w12=f12=fx1x2=∂2
∂x1∂x2=∂
∂x2(∂f
∂x1) =x1ex1x2=x1ex1x2+x2
1x2ex1x2
w21=f21=fx2x1=∂2f
∂x2∂x1=∂
∂x1(∂f
∂x2) = 2x1ex1x2=x2
1x2ex1x2=f12
w22=f22=fx2x2=∂2f
∂x2
2=∂
∂x2(∂f
∂x2) =x3
1ex1x2.
And even higher derivatives can be computed too, like
f221=fx2x2x1=∂3f
∂x2
2∂x1=∂
∂x1(∂2f
∂x2
2) = 3x2
1ex1x2+x3
1x2ex1x2.
Remark: From this one example, it appears possible that we always have f12=f21,
that is∂2f
∂x1∂x2=∂2f
∂x2∂x1. This is indeed the case ifthe second partial derivatives of fare
continuous, but for lack of time we shall not prove it (see Exercise 6).
So far we have defined the directional derivative of a function f:En→Eand called
particular attention to those in the direction of the coordinate axes - the partial derivatives
off. Although the actual computation of the partial derivatives has been reduced to
the formal procedure of computing ordinary derivatives, the computation of the directional
derivative in an arbitrary direction must still be done by using the definition: the limit of
a difference quotient. We shall now reduce the computation of all directional derivatives
to a simple formal procedure. In order to do so, we shall introduce the concept of the
total derivative for functions f:A⊂En→E1. This derivative will not be a directional
derivative, but rather a more general object.
The motivating idea here is the important one of approximating a non-linear function
fat a pointX0by a linear function. If we think of the function f(X) as defining a surface
MinEn+1with point ( X,f(X)) , then the picture is that of approximating the surface
MnearX0by a plane (or hyperplane) tangent to the surface at X0. We want to write
f(X)∼f(X0) +L(X−X0),
whereLis a linear operator, L:En→E, which may depend on the “base point” X0.
Of course, as X→X0we want the accuracy to improve in the sense that the tangent
plane should be a better approximation the closer Xis toX0. AtX=X0, the tangent
314CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
planef(X0) +L(X−X0) and surface Mtouch since they both pass through the point
(X0,f(X0)) . Notice that the function f(X)) +L(X−X0) is affine, so it does represent a
plane surface.
Motivated by the above considerations, we can now make a reasonable
Definition: Letf:A⊂En→EandX0be an interior point of A.fisdifferentiable
atX0if there exists a linear transformation L:En→Esuch that
lim
/bardblh/bardbl→0/bardblf(X0+h)−f(X0)−Lh/bardbl
/bardblh/bardbl= 0,
for any vector hin some small ball about X0(sof(X0+h) is defined). The operator L
will usually depend on the base point X0. Iffis differentiable at X0, we shall use the
notation
df
dX(X0) =f/prime(X0) =L(X0)=L,
and refer to f/prime(X0) as the total derivative offatX0. [The notation ∇f(X0) and grad
f(X0) , for gradient , are also used]. If L=f/prime(X0) , a linear operator from EntoE, exists
and depends continuously on the base point X0for allX0∈A, thenfis said to be
continuously differentiable in A, writtenf∈C1(A) .
Remark: The condition that fbe differentiable at X0can also be written in the following
useful form:
f(X0+h) =f(X0) +Lh+R(X0,h)/bardblh/bardbl, (8-1)
where the remainder R(X0,h) has the property
lim
/bardblh/bardbl→0R(X0,h) = 0.
This abstract operator Lhas the delightful property that it can be computed easily.
But before telling you how, we should first prove for a given fthere can be at most one
linear operator Lwhich is the total derivative.
Theorem 8.1 . (Uniqueness of the total derivative). Let f:A→Ebe differentiable at
the interior point X0∈A. IfL1andL2are linear operators both of which satisfy the
conditions for the total derivative of fatX0, thenL1=L2.
Proof: LetL=L1−L2. We shall show Lis the zero operator. Since
Lh=L1h−L2h= [f(X0+h)−f(X0)−L2h]−[f(X0+h)−f(X0−L1h],
by the triangle inequality we have
/bardblLh/bardbl ≤ /bardblf(X0+h)−f(X0)−L2h/bardbl+/bardblf(X0+h)−f(X0)−L1h/bardbl.
Consequently,
lim
/bardblh/bardbl→0/bardblLh/bardbl
/bardblh/bardbl= 0.
8.1. THE DIRECTIONAL AND TOTAL DERIVATIVES 315
To complete the proof, a trick is needed. Fix η/negationslash= 0 . Ifλis a constant, λ→0 , then
/bardblλη/bardbl →0 so
lim
/bardblλ/bardbl→0/bardblL(λη)/bardbl
/bardblλη/bardbl= 0.
But sinceLis linear, /bardblLλη/bardbl=/bardblλLη/bardbl=|λ| /bardblLη/bardbl, so the factor λcan be canceled in
numerator and denominator. Thus the last equation is independent of λ, so/bardblLη/bardbl//bardblη/bardbl= 0 .
Becauseη/negationslash= 0 , this implies /bardblLη/bardbl= 0 . Therefore Lmust be the zero operator.
Next, we give a method for computing L. Not only that, but we also find an easy way
to compute the directional derivatives.
Theorem 8.2 . Letf:A→Ebe differentiable at the interior point X0∈A. Then a) the
directional derivative of fatX0exists for every direction eand is given by the formula
Def(X0) =Le.
b) Moreover, if fis given in terms of coordinates, f(X) =f(x1,...,x n), thenLis
represented by the 1×nmatrix
f/prime(X0) =L= (fx1(X0),...,f xn(X0)).
c) Consequently, the directional derivative is simply the product of this matrix Lwith the
unit vector e, which can also be thought of as the scalar product of the 1×nmatrix, a
vector, and the vector e,
Def(X0) =/angbracketleftf/prime(X0), e/angbracketright.
Proof: This falls out of the definitions. First
Def(X0) = lim
λ→0f(X0+λe)−f(X0)
λ.
= lim
λ→0f(X0+λe)−f(X0)−L(λe) +L(λe)
λ.
SinceL(λe) =λLe and/bardblλe/bardbl=λ
= lim
/bardblλe/bardbl→0f(X0+λe)−f(X0)−L(λe)
/bardblλe/bardbl+Le.
Becausefis differentiable at X0, the first term tends to zero. Thus proving the first part.
To prove the last part, it is sufficient to observe that if e=ejis one of the coordinate
vectors, then by definition Dejf(X0) :=fxj. Thus, ifhis any vector h= (h1,...,h n) =
h1e1+···+hnen, by the linearity of Lwe have
Lh=L(h1e1+···+hnen) =h1Le1+···+hnLen
=h1fx1(X0) +···+hnfxn(X0)
= (fx1(X0),...,f xn(X0))
h1
·
·
·
hn
.
316CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
Sincehis any vector, we have shown Lis represented by the given matrix.
Remark: The theorem states that iffis differentiable, then all the partial derivatives
exist andf/prime(X0) :=Lis represented by the above matrix. It does notstate that if the
partial derivatives exist, then fis differentiable. This is false (see Exercise 16). However,
if the partial derivatives of fexist and are continuous, then fis differentiable. The last
statement will be proved as Theorem 3.
Example: The same one worked before (p. 573). Find the directional derivative of f(X) =
x2
1−2x1x2+ 3x2atX0= (1,0) in the direction η= (−1,1) .
Sinceηis not a unit vector, we let e=η
/bardblη/bardbl= (−1√
2,1√
2) . Now at a point X,
L=f/prime(X) = (fx1,fx2) = (2x1−2x2,−2x1+ 3).
In particular, at X=X0= (1,0) ,
L= (2,−2 + 3) = (2,1).
Therefore
Def(X0) =Le= (2,1)/parenleftBigg
−1√
21√
2/parenrightBigg
=−1√
2,
which checks with the answer found previously.
Consider the mapping w=f(X), X∈A⊂En, w∈Eas defining a surface M⊂En+1.
It is now evident how to define the tangent plane to Mat the point ( X0,f(X0)) , where
X0∈A.
Definition: LetF:A⊂En→Ebe a differentiable mapping, thus defining a surface M
with points ( X,f(X)), X∈A. The tangent plane toMat the point ( X0,f(X0)) , where
X0∈A, is the surface defined by the affine mapping
Φ(X) =f(X0) +f/prime(X0)(X−X0).
or
Φ(X) =f(X0) +L(X−X0),whereL=f/prime(X0),
Thus, the tangent plane to the surface defined by fis merely the “affine part” of fat
X0.
Example: Consider the function w=f(X) = 3 −x2
1−x2
2. This function defines a
paraboloid (see fig.). Let us find the tangent plane to this surface at ( X0,f(X0)) , where
X0= (1,−1) , sof(X0) = 3−12−(−1)2= 1 . Also
fx1(X) =−2x1,fx2(X) =−2x2.
Thus
f/prime(X0) = (fx1(X0),fx2(X0)) = (−2,2).
SinceX−X0= (x1,x2)−(1,−1) = (x1−1,x2+ 1) we find the equation of the tangent
plane is
Φ(X) = 1 + ( −2,2)/parenleftbiggx1−1
x2+ 1/parenrightbigg
= 1−2(x1−1) + 2(x2+ 1),
8.1. THE DIRECTIONAL AND TOTAL DERIVATIVES 317
or
Φ(X) = 5−2x1+ 2x2.
This tangent plane is the unique plane with the property
Φ(X0) =f(X0),and Φ/prime(X0) =f/prime(X0).
Although we have given necessary conditions that a function be differentiable (all directional
derivatives exist, in particular, all partial derivatives exist), we have not given sufficient
conditions. The next theorem gives sufficient conditions for a function to be continuously
differentiable.
Theorem 8.3 . Letf:A⊂En→E, whereAis an open set. Then fis continuously
differentiable throughout Aif and only if all the partial derivatives of fexist and are
continuous.
Proof: ⇒Iffis continuously differentiable, then the partial derivatives exist by Theorem
2. Furthermore, for any XandYinA,
fxi(X)−fxi(Y) =/angbracketleftf/prime(X), ei/angbracketright − /angbracketleftf/prime(Y), ei/angbracketright=/angbracketleftf/prime(X)−f/prime(Y), ei/angbracketright.
Thus, applying the Schwartz inequality we find
|fxi(X)−fxi(Y)| ≤ /bardblf/prime(X)−f/prime(Y)/bardbl.
The statement fis continuously differentiable means the vector f/prime(X) is a continuous
function of X. Therefore, given any /epsilon1>0 is aδ >0 such that /bardblf/prime(X)−f/prime(Y)/bardbl</epsilon1for
all/bardblX−Y/bardbl<δ. For any/epsilon1>0 , the inequality above shows |fxi(X)−fxi(Y)|is also less
than/epsilon1for the same δ. Consequently, fxiis continuous.
⇐. A little more difficult. The idea is to use the mean value theorem for functions of
one variable. Let XandYbe points on A. To prove continuity at X, it is sufficient
to restrictYto being in some ball about Xwhich is entirely in A(some ball does exist
sinceAis open). For notational convenience, we take n= 2 . Then
f(Y)−f(X) =f(Y)−f(Z) +f(Z)−f(X),
whereZis a point in Awhose coordinates, except the first, are the same as Xand
whose coordinates, except the second, are the same as Y. By the one variable mean value
theorem, there is a point ˜XbetweenXandZand point ˆXbetweenYandZsuch
that
f(Z)−f(X) =∂f
∂x1(˜X)(y1−x1), f(Y)−f(Z) =∂f
∂x2(ˆX)(y2−x2).
Therefore
f(Y)−f(X) =∂f
∂x1(˜X)(y1−x1) +∂f
∂x2(ˆX)(y2−x2),
so
f(Y)−f(X)−[fxi(X)(y1−x1) +fx2(X)(y2−x2)]
= [fxi(˜X)−fxi(X)](y1−x1) + [fx2(ˆX)−fx2(X)](y2−x2).
Therefore
/bardblf(Y)−f(X)−L(Y−X)/bardbl ≤/vextendsingle/vextendsingle/vextendsinglefxi(˜X)−fxi(X)/vextendsingle/vextendsingle/vextendsingle|y1−x1|/vextendsingle/vextendsingle/vextendsinglefx2(ˆX)−fx2(X)/vextendsingle/vextendsingle/vextendsingle|y2−x2|
318CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
where we have written L= (fx1(X),fx2(X)) . Since |yj−xj| ≤ /bardblY−X/bardbl, we see that
/bardblf(X)−f(Y)−L(Y−X)/bardbl
/bardblY−X/bardbl≤/vextendsingle/vextendsingle/vextendsinglefx1(˜X)−fx1(X)/vextendsingle/vextendsingle/vextendsingle+/vextendsingle/vextendsingle/vextendsinglefx2(ˆX)−fx2(X)/vextendsingle/vextendsingle/vextendsingle.
Becausefx1andfx2are continuous and /bardbl˜X−X/bardbl</bardblY−X/bardbl,/bardblˆX−X/bardbl</bardblY−X/bardbl,
by making /bardblY−X/bardblsufficiently small the right side of the above inequality can be made
arbitrarily small. This proves the limit as /bardblY−X/bardbl →0 of the expression on the left - exists
and is zero. Since Lis linear, the proof that fis differentiable is complete. The continuous
differentiability is an immediate consequence of the linearity of Land the continuity of its
components - the partial derivatives fxi.
Exercises
(1) i) Use the definition of the directional derivative to compute the given directional
derivatives, ii) Check your answer by computing the directional derivative using the
procedure of the Corollary to Theorem I.
(a)f(x1,x2) = 1−2x1+ 3x2, at (2,−1) in the direction (3 ,4) . [Answer: +6
4].
(b)f(x,y) =ex+2y, at (3,−2) in the direction (1 ,1) . [Answer: 3 e−1/√
2 ].
(c)f(u,v,w ) = 3uv+uw−v2, at (1,1,1) in the direction (1 ,−2,2) .
(d)f(x,y) = 1−3y+xyat (0,6) in the direction (3
5,−4
5) .
(2) i) Compute allof the first and second partial derivatives for the following functions.
(a)f(x1,x2) =x1+x1sin 2x1
(b)f(x1,x2,x3) =x2
1x2+ 2x1√x3−x3
(c)f(x,y) =xy
(d)f(x1,x2,...,x n) =a+a1x1+a2x2+...+anxn.
(e)f(x1,x2,...,x n) =n/summationdisplay
i,j=1aijxixj=/angbracketleftX, AX /angbracketright, whereaij=aji(first try the cases
n= 2 andn= 3 to see what is happening).
ii) Find the 1 ×nmatrixf/prime(x) .
(3) For the surfaces defined by the functions f(X) listed below, find the equation of the
tangent plane to the surface at the point ( X0,f(X0)) . Draw a sketch showing the
surface and its tangent plane.
(a)f(X) =x2
1+ 3x2
2+ 1, X 0= (0,0).
(b)f(X) =ex1x2, X 0= (0,1)
(c)f(X) =x2
1sinπx2, X 0= (−1,1
2)
(d)f(X) =−1
2x1+x2+ 1, X 0= (2,1)
(e)f(X) =x2
1+ 2x2
2−x1x3+x1, X 0= (1,−2,−1).
Why can’t you sketch the surface defined by this function?
8.1. THE DIRECTIONAL AND TOTAL DERIVATIVES 319
(4) Letf(X) andg(X) both map A⊂En→E1. Iffandgare differentiable for all
X∈A, prove
(a)d
dX[af(X) +bg(X)] =ad f
dX(X) +bdg
dX(X) (Linearity), where aandbare con-
stants.
(b)d
dX[f(X)g(X)] =f(X)dg
dX(X) +g(X)d f
dX(X)
(c)d
dX/bracketleftBig
f(X)
g(X)/bracketrightBig
=g(X)f/prime(X)−f(X)g/prime(X)
g2(X), ifg(X)/negationslash= 0 .
(5) Use the rules (a-c) of Exercise 4 to computed
dX[2f−3g],d
dX[f·g] , andd
dX[f
g] , where
f(X) =f(x1,x2) = 1−x1+x1x2, andg(X) =g(x1,x2) =ex1−x2.
(6) Letf(X) =f(x,y) =/braceleftBigg
xy(x2−y2)
x2+y2, X = (x,y)/negationslash= 0
0 X= 0
Prove
(a)f,fx,fyare continuous for all X∈E2. [Hint: Prove and use 2 xy≤x2+y2].
(b)fxyandfyxexist for all X∈E2, and are continuous except at the origin.
(c)fxy(0) = 1,fyx(0) = −1 , sofxy(0)/negationslash=fyx(0) (cf. Remark p. 577).
(7) Letf:A⊂En→Ebe a differentiable map. Prove it is necessarily continuous. [Hint:
This is a simple consequence of the definition in the form (1)].
(8) Letf:A⊂En→Ebe a continuous map. We say fhas a local maximum at the
pointX0interior to Aiff(X0)≥f(X) for allXin some sufficiently small ball
aboutX0. If we assume fis continuously differentiable, more can be said.
(a) Iffas above has a local maximum at the point X0, prove /angbracketleftf/prime(X0, X−X0)/angbracketright+
R(X0,X)/bardblX−X0/bardbl ≤0 for allXis some small ball about X0.
(b) Use the property of R(X0,X) to conclude the stronger statement
/angbracketleftf/prime(X0),(X−X0)/angbracketright ≤0.
for allXin some small ball about X0.
(c) Observe the statement must also hold for the vector X0−X, which points in
the direction opposite to X−X0, to conclude
/angbracketleftf/prime(X0),(X−X0)/angbracketright ≥0,
and hence that in fact
/angbracketleftf/prime(X0), Z/angbracketright= 0,
for all vectors Z=X−X0.
(d) Finally, show that at a maximum,
f/prime(X0) = 0.
(9) (a) Find the equation of the plane which is tangent at the point X0= (2,6,3) to
the surface consisting of the points ( X,f(X)) , where
f(X) =f(x,y,z ) = (x2+y2+z2)1/2.
320CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
(b) Use the tangent plane found above to find the approximate value of
((2.01)2+ (5.98)2+ (2.99)2)1/2.
(10) Assume the continuously differentiable function f(X) has a zero derivative, f/prime(X)≡
0 , forXin some ball in En. Prove that f(X)≡constant throughout the ball.
(11) (a) Show the following functions satisfy the two dimensional Laplace equation
∂2u
∂x2+∂2u
∂y2= 0
i)u(x,y) =x2−y2−3xy+ 5y−6
ii)u(x,y) = log(x2+y2) , except at the origin, ( x,y) = 0.
iii)u(x,y) =exsiny
(b) Show the following functions satisfy the one (space) dimensional wave equation
utt=c2uxx, c≡constant
[Heretis time and xis space;cis the velocity of light, sound, etc.]
i)u(x,y) =ex−ct−2ex+ct
ii)u(x,y) = 2(x+ct)2+ sin 2(x−ct).
(12) Letf:A⊂En→Ebe continuously differentiable throughout A. IfX0∈Ais
not a critical point of f, sof/prime(X0)/negationslash= 0 , prove the directional derivative at X0is
greatest in the direction emax:=f/prime(X0)//bardblf/prime(X0)/bardbl, and least in the opposite direction,
emin:=−emax. [Hint: Use the Schwarz inequality.]
(13) Consider the function f(X) =f(x,y) =/braceleftbiggxy
x2+y2, X = (x,y)/negationslash= 0
0, X = 0
SinceFis the quotient of two continuous functions, it is continuous except possibly
at the origin, where the denominator vanishes. Show that f(X) isnotcontinuous at
the origin by finding lim f(X) asX→0 along paths 1 and 2, and showing that
lim
X→0
path1f(X)/negationslash= lim
X→0
path2f(X).
(14) LetLbe the partial differential operator defined by
Lu=∂2u
∂x2−5∂2u
∂x∂y+ 6∂2u
∂y2.
Show that
L[eαx+βy] =p(α,β)eαx+βy,
wherep(α,β) is a polynomial in αandβ. Find a solution of the linear homogeneous
partial differential equation Lu= 0 . Find an infinite number of solutions of Lu= 0 ,
one for each value of α, by choosing αto depend on βin a particular way. [Answer:
e2βx+βyande3βx+βyare solutions for any β].
8.2. THE MEAN VALUE THEOREM. LOCAL EXTREMA. 321
(15) The two equations
x=eucosv
y=eusinv
defineu=f(x,y) andv=g(x,y) . Find the functions fandgforx>0 . Compute
f/prime(X) andg/prime(X) and show f/prime(X)⊥g/prime(X) .
(16) This exercise gives an example in which the first partial derivatives of a function exist
but the function is not continuous, let alone differentiable. Let
f(X) =f(x,y) =/braceleftBigg
xy2
x2+y4, X = (x,y)/negationslash= 0
0, X = 0.
(a) If cosα/negationslash= 0 , prove the directional derivative at the origin in the direction e=
(cosα,sinα) exists and is
Def(0) =2 sin2α
cosα, cosα/negationslash= 0
while if cos α= 0 ,
Def(0) = 0, cosα= 0.
(b) Provefis discontinuous at the origin by showing lim
X→0f(X) has two different
values along the two paths in the figure. Then appeal to exercise 7 to conclude
fis not differentiable.
(17) (a) Let P(X),X∈En, be a polynomial of degree N, that is,
P(X) =/summationdisplay
k1+k2+···+kn≤Nak1,...,k nxk1
1xk2
2···xknn,
wherek1,k2,...,k nare all non-negative integers. Prove P(α) is continuously
differentiable. [Hint: How do you prove a polynomial in one variable is continu-
ously differentiable.]
(b) LetR(X),X∈En, be a rational function - that is, the quotient of two polyno-
mials. Prove R(X) is continuously differentiable whenever the denominator is
not zero.
(18) Iff:E1→E1, show that the definition of differentiability on page 578 coincides with
the usual one.
8.2 The Mean Value Theorem. Local Extrema.
Although the full “chain rule” will not be proved until Chapter 10, we shall need a very
special and elementary case to develop the main features of the theory of mappings from En
toE. Letf:A⊂En→Ebe a continuously differentiable function at all interior points
ofA. TakeXandZto be fixed interior points of A. Letφ(t) =f(X+tZ) . We want
to compute
d
dtφ(t) =d
dtf(X+tZ)
that is, the rate of change of f(X) at the point X+tZasXvaries along the line joining
XtoZ.
322CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
Theorem 8.4 . Letf:A→Ebe a differentiable function throughout A. IfXandZ
are two interior points of A, and if the line segment joining them is in A, then
d
dtf(X+tZ) =f/prime(X+tZ)Z, t ∈(0,1).
By the product f/prime(Y)Zwe mean matrix multiplication.
Proof: For fixedXandZ, the function φ(t) :=f(X+tZ) an ordinary scalar valued
function of the one variable t. Thus
d
dtφ(t) = lim
λ→0φ(tλ)−φ(t)
λ
= lim
λ→0f(X+tZ+λZ)−f(X+tZ)
λ
= lim
λ→0f(X+tZ+λZ)−f(X+tZ)−f/prime(X+tZ)(λZ) +f/prime(X+tZ)(λZ)
λ
Sincefis differentiable at X+tZ, then asλ→0 the first three terms tend to zero. The
factorλin the last term cancels. Therefore
d
dtf(X+tZ) = lim
λ→0f/prime(X+tZ)Z=f/prime(X+tZ)Z,
as claimed.
An easy consequence is
Theorem 8.5 (The Mean Value Theorem). Let f:A→E, whereAis an open convex
set in En, that is, if XandYare any points in Z, then the straight line segment joining
XandYis inAtoo. Iffis differentiable in A, there is a point Zon the segment
joiningXandYsuch that
f(Y)−f(X) =f/prime(Z)(Y−X).
If, moreover, f/primeis bounded by some constant C,/bardblf/prime(X)/bardbl ≤Cfor allX∈A, then
|f(Y)−f(X)| ≤C/bardblY−X/bardbl
a figure goes here
Proof: Every point on the segment joining XandYis of the form X+t(Y−X) , where
t∈[0,1] . Consider the function φ(t) of one variable,
φ(t) =f(X+t(Y−X)).
Theorem 4 states φis differentiable. Therefore, by the one variable mean value theorem,
there is a number t0in the interval (0 ,1) such that φ(1)−φ(0) =φ/prime(t0) . Butφ(1) =
f(Y), φ(0) =f(X) and, by Theorem 4, φ/prime(t0) =f/prime(X+t0(Y−X))(Y−X) . Letting
Z=X+t0(Y−X) , a point on the segment joining XtoY, we conclude
f(Y)−f(X) =f/prime(Z)(Y−X).
8.2. THE MEAN VALUE THEOREM. LOCAL EXTREMA. 323
The second part of the theorem follows by applying the Schwarz inequality to the function
f/prime(Z)(Y−X) which can be written as /angbracketleftf/prime(Z),(Y−X)/angbracketright. Then
/angbracketleftf/prime(Z), Y−X/angbracketright ≤ /bardblf/prime(Z)/bardbl/bardblY−X/bardbl.
Therefore if /bardblf/prime(Z)/bardbl ≤Cfor allZ∈A, we find
|f(Y)−f(X)| ≤C/bardblY−X/bardbl.
Corollary 8.6 Letf:A→Ebe a differentiable map and Aan open connected set in En
(by a connected open set we mean it is possible to join any two points in Aby a polygonal
curve contained in A). If f/prime(X)≡0for everyX∈A, that is, if fx1(X) =...=fxn(X) =
0, thenf(X)≡c,ca constant.
Proof: IfAis convex, say a ball, this is an immediate consequence of the second part of
the mean value theorem, for /bardblf/prime(X)/bardbl= 0 so |f(Y)−f(X)|= 0 . Thus f(Y) =f(X) =
constant for any two points XandY. The requirement that Ais connected is to exclude
the possibility that Aconsists of two (or more) disjoint sets, in which case, all we can
conclude is that fis constant on each connected part, but not necessarily the same constant.
However, if Ais connected, then any two points in Acan be joined by a polygonal curve
which is contained in A. Consider some straight line segment in this curve. By the mean
value theorem, fmust be constant on it. In particular, it has the same value at both
end points. Checking the beginning and end of the whole polygonal curve, we find that
f(X) =f(Y) . Because XandYwere any points, we are done.
It is not at all difficult to generalize the mean value theorem to Taylor’s theorem and
then to power series for functions of several variables. The only problem is one of notation,
and that is a problem. As a compromise, we will prove the Taylor theorem - but only the
first two terms for functions of three variables f(x,y,z ) .
Just as in the mean value theorem, the idea is to reduce the problem to a function
φ(t) of one real variable, because we do know the result for these functions. Let fbe
differentiable in some open set A⊂E3andX0a point inA. IfX0+his also inA, we
would like to express f(X0+h) in terms of fand its derivatives at X0. FixX0andh
and consider the real valued function φ(t) of one variable defined by
φ(t) =f(X0+th), t ∈[0,1].
Then by Theorem 4,
φ/prime(t) =f/prime(X0+th)h=fx(X0+th)h1+fy(X0+th)h2+fz(X0+th)h3,
whereh= (h1,h2,h3) . Since each of the partial derivatives are maps from AtoE, they
can be differentiated in the same way fwas. So can a sum of such functions. Thus
φ/prime/prime(t) =d
dt[fx(X0+th)h1+···+fz(X0+th)hn]
=fxx(X0+th)h1h1+fxy(X0+th)h1h2+fxz(X0+th)h1h3
+fyx(X0+th)h2h1+fyy(X0+th)h2h2+fyz(X0+th)h2h3
+fzx(X0+th)h3h1+fzy(X0+th)h3h2+fzz(X0+th)h3h2.
324CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
If we introduce a matrix H(X) , the Hessian matrix , whose elements are∂2f(X)
∂xi∂xj, φ/prime/prime(t)
can be written as
φ/prime/prime(t) =/angbracketlefth, H(X0+th)h/angbracketright.
We remark that if fis sufficiently differentiable (two continuous derivatives is enough),
then the Hessian matrix is self-adjoint since fxixj=fxjxi, as we mentioned - but did not
prove - earlier. If φ(t) is twice differentiable, by Taylor’s theorem for functions of one
variable, we know that
φ(1) =φ(0) +φ/prime(0) +1
2!φ/prime/prime(τ), τ∈(0,1).
Substituting into this formula, we find
f(X0+h) =f(X0) +f/prime(X0)h+1
2!/angbracketlefth, H(X0+τh)h/angbracketright.
Let us summarize. We have proved
Theorem 8.7 ( Taylor’s Theorem with two terms) . Letf:A→E, whereAis
an open connected set in En. Assumefhas two continuous derivatives - that is, all the
second partial derivatives of fexist and are continuous. If X0is inAandX0+his in
a ball about X0inA, then
f(X0+h) =f(X0) +f/prime(X0)h+1
2!/angbracketlefth, H(X0+τh)h/angbracketright,
whereH(X) = ((∂2f
∂x1∂xj))is then×nHessian matrix and τ∈(0,1).
LettingX=X0+handZ=X0+τh, Z being a point on the line segment joining
X0toX, this reads
f(X) =f(X0) +f/prime(X0)(X−X0) +1
2!/angbracketleftX−X0, H(Z)(X−X0)/angbracketright,
or, in more detail,
f(X) =f(X0) +n/summationdisplay
i=1∂f(X0)
∂xi(xi−x0
i) +1
2n/summationdisplay
i=jn/summationdisplay
j=1∂2f(Z)
∂xi∂xj(xi−x0
i)(xj−x0
j).
Example: Find the first two terms in the Taylor expansion for the function f(X) =
f(x,y) = 5 + (2x−y)3about the point X0= (1,3) .
We compute
fx(X) = 6(2x−y)2, f y(X) =−3(2x−y)2
fxx(X) = 24(2x−y),fxy(X) =fyx(X) =−12(2x−y),fyy(X) = 6(2x−y).
Thereforef(X0) = 4,fx(X0) = 6,fy(X0) =−3 , so
f(X) = 4 + (6,−3)/parenleftbiggx−1
y−3/parenrightbigg
+1
2(x−1,y−3)/parenleftbigg2ξ−η −12(2ξ−η)
−12(2ξ−η) 6(2ξ−η)/parenrightbigg/parenleftbiggx−1
y−3/parenrightbigg
8.2. THE MEAN VALUE THEOREM. LOCAL EXTREMA. 325
whereZ= (ξ,η) is a point on the segment between X0= (1,3) andX= (x,y) . Written
out, the above equation reads,
f(x,y) = 4 + 6(x−1)−3(y−3) +1
2[fxx(x−1)2+ 2fxy(x−1)(y−3) +fyy(y−x)2],
where the second derivatives are evaluated at Z= (ξ,η) .
We are now in a position to examine the extrema of functions of several variables.
Finding the maxima and minima of functions is important for several reasons. First of all,
there is the vague emotional feeling that all patterns of action should maximize or minimize
something. Second, we can investigate a complicated geometrical object by the relatively
easy procedure of finding the local maxima and minima. Without further mention, for the
balance of this section f(X) will be a twice continuously differentiable function which maps
the open set A⊂EnintoE.
Definition: A function f:A→Ehas a local maximum at the interior point X0∈Aif,
for allXin some open ball about X0
f(X)≤f(X0).
fhas a local minimum atX0if for allXin some open ball about X0
f(X)≥f(X0).
Iffhas a local maximum or minimum at X0, isf/prime(X0) = 0 ? Certainly.
Theorem 8.8 . Iffhas a local maximum or minimum at X0, thenf/prime(X0) = 0 . In
coordinates, this means all the partial derivatives vanish at X0,
∂f
∂x11(X0) =∂f
∂x2(X0) =···=∂f
∂xn(X0) = 0.
Proof: Letηbe any fixed vector. Then the function φ(t) of one variable
φ(t) =f(X0+tη)
has a local maximum or minimum at t= 0 . Consequently φ/prime(0) = 0 . But by Theorem 4,
φ/prime(0) =f/prime(X0)ηwhich we may write as /angbracketleftf/prime(X0), η/angbracketright. Thus /angbracketleftf/prime(X0), η/angbracketright= 0 , so the vector
f/prime(X0) is orthogonal to η. Sinceηwas any vector, we conclude that f/prime(X0) = 0 .
The derivative f/prime(X0) may vanish at points other than maxima or minima. An example
is the “saddle point” of the hyperbolic paraboloid at the beginning of Section 1. All points
wheref/primevanishes are called critical points orstationary points off. Let us give a precise
definition of a saddle point. fhas a saddle point atX0ifX0is a critical point of f
and if every ball about X0contains points X1andX2such thatf(X1)< f(X0) and
f(X2)>f(X0) . Thus, every critical point is either a local maximum, minimum, or saddle
point.
There is a more intuitive way to prove Theorem 7. If eis a unit vector, then by
Theorem 2, the directional derivative at Xin the direction eisDef(X) =/angbracketleftf/prime(X), e/angbracketright. In
what way should you move so fincreases fastest? By the Schwartz inequality, we find
|Def(X)| ≤ /bardblf/prime(X)/bardbl /bardble/bardbl=/bardblf/prime(X)/bardbl,
326CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
with equality if and only if the vectors eandf/prime(X) are parallel. Thus, the directional
derivative is largest when e has the same direction as f/prime(X), and smallest when e has the
opposite direction ,emax=f/prime(X)//bardblf/prime(X)/bardbl, emin=−emax,
Demaxf(X) =/bardblf/prime(X)/bardbl, Deminf(X) =−/bardblf/prime(X)/bardbl.
IfX0is a local maximum of f, thenf/prime(X0) must be zero, for otherwise you could move in
the direction of f/prime(X0) and increase the value of f. Similarly, if X0is a local minimum,
f/prime(X0) must be zero.
Once we know X0is a critical point of f, f/prime(X0) = 0 , an effective criterion is needed
to determine if X0is a local maximum, minimum, or saddle point for f. In elementary
calculus, the sign of the second derivative was used. Our next theorem generalizes this test.
The idea is essentially the same as in the one variable case (p. 104a-c). If fhas a
local maxima or minima, the tangent plane to the surface whose points are ( X,f(X)) is
horizontal, that is, f/prime(X0) = 0 . Thus, near X0the quadratic terms - the next lowest power
in the Taylor expansion of faboutX0—will determine the behavior of fnearX0. Let
X0be the origin and take f(X) =f(x,y) to be a function of two variables with f(0) = 0 .
Then near X0= 0 , by Taylor’s theorem, we have
f(x,y)∼1
2[ax2+ 2bxy+cy2],
wherea=fxx(0), b=fxy(0) , andc=fyy(0) . The nature of the quadratic form
Q(X) =ax2+ 2bxy+cy2
has already been determined. If Q(X) is positive definite, then Q(X)>0 forX/negationslash= 0 .
Sincef(x,y)∼Q(X) , this means f(x,y) is positive near the origin. Because f(0,0) = 0 ,
this implies the origin is a minimum for f.
Instead of completing and rigorously justifying this special case, we shall immediately
treat the general situation.
Theorem 8.9 . Assume the twice continuously differentiable function f:A→Ehas a
critical point at an interior point X0ofA⊂En, f/prime(X0) = 0 . LetH(X0)be the Hessian
matrix/parenleftbigg/parenleftbigg∂2f
∂xi∂xj(X0)/parenrightbigg/parenrightbigg
evaluated at X0.
(a)IfH(X0)is positive definite, then fhas a local minimum at X0.
(b)IfH(X0)is negative definite, then fhas a local maximum at X0.
(c)If at least two of the diagonal elements of H(X0), fx1x1(X0),...,f xnxn(X0)have
different signs, then X0is a saddle point.
(d)Otherwise the test fails.
Proof: IfX0is a critical point for f, then Taylor’s theorem (Theorem 6) states
f(X0+η) =f(X0) +1
2/angbracketleftη, H(Z)η/angbracketright
whereZis between X0andX0+η. The linear term has been dropped since f/prime(X0) = 0 .
8.2. THE MEAN VALUE THEOREM. LOCAL EXTREMA. 327
As in the proof of Taylor’s theorem, let
φ(t) =f(X0+tη).
Then
φ/prime/prime(t) =/angbracketleftη, H(X0+tη)η/angbracketright.
Since the second derivatives of fare assumed to be continuous, the function φ/prime/prime(t) is a
continuous function of t. Consequently, if φ/prime/prime(0) is positive then φ/prime/prime(t) is also positive
for alltsufficiently close to zero (Theorem I p. 29b). Because φ/prime/prime(0) = /angbracketleftη, H(X0)η/angbracketrightand
φ/prime/prime(τ) =/angbracketleftη, H(Z)η/angbracketright, whereZ=X0+τη, this implies if His positive definite at X0, it
is also positive definite at ZwhenZis close toX0.
AssumingH(X0) is positive definite, we see that for all ηsufficiently small, H(Z) is
positive definite. Therefore,
f(X0+η)−f(X0) =1
2/angbracketleftη, H(Z)η/angbracketright>0, η/negationslash= 0,
that is
f(X0+η)−f(X0)>0
for allηis some small ball about X0. Thusfhas a local minimum at X0.
IfH(X0) is negative definite, the same proof with trivial modifications works. Another
way to complete the proof is to apply part a) to the function g(X) :=−f(X) . The Hessian
forgatX0will be −H(X0) which is positive definite (since H(X0) was negative definite).
Thusghas a local minimum at X0sof: =−ghas a local maximum at X0.
If any two of the diagonal elements of H(X0) have opposite sign, say fx1x1(X0)>0
andfx2x2(X0)<0 , then for η=λe1= (λ,0,0,..., 0) ,λany real number, we find
/angbracketleftη, H(X0)η/angbracketright=λ2fx1x1(X0)>0 , while for η=λe2= (0,λ,0,..., 0)/angbracketleftη, H(X0)η/angbracketright=
λ2fx2x2(X0)<0 . Therefore the quadratic form /angbracketleftη, H(X0)η/angbracketrightassumes positive and negative
values in any ball about X0, provingX0is a saddle point.
Since this theorem reduces the investigation of the nature of a critical point to testing if
a matrix is positive or negative definite, it would do well in this context to repeat Theorem
A (p. 386d) which tells us when a 2 ×2 matrix is positive definite.
Corollary 8.10 . LetX0be a critical point for the function of two variables f(x,y)with
Hessian matrix
H(X0) =/parenleftbiggfxx(X0)fxy(X0)
fxy(X0)fyy(X0)/parenrightbigg
.
(a)IfdetH(X0)>0andfxx(X0)>0, thenfhas a local minimum at X0.
(b)IfdetH(X0)>0andfxx(X0)<0, thenfhas a local maximum at X0.
(c)IfdetH(X0)<0, thenfhas a saddle point at X0(this is a stronger statement
than part c of Theorem 8).
Proof: Since these merely join Theorem A (p. 386d) with Theorem 8, the proof is done.
Examples:
328CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
(1) Find and classify the critical points of the function w=f(x,y) := 3 −x2−4y2+ 2x.
A sketch of the surface with points ( x,y,f (x,y)) , a paraboloid, is at the right. At a
critical point f/prime(X) = 0 , that is, fx= 0, fy= 0 . Since
fx=−2x+ 2,fy=−8y,
at a critical point
−2x+ 2 = 0,−8y= 0.
There is therefore only one critical point, X0= (1,0) . We look at the Hessian to
determine the nature of the critical point. Because fxx=−2, fxy=fyx= 0, fyy=
−8 ,
H(X0) =/parenleftbigg−2 0
0−8/parenrightbigg
.
Since detH(X0) = 16>0 andfxx(X0) =−2<0,H(X0) is negative definite so
X0= (1,0) is a local maximum for the function, and at that point f(X0) = 4 .
(2) Find and classify the critical points of w=f(x,y) =−x2+y2.
The surface ( x,y,f (x,y)) is a hyperbolic paraboloid. We expect a saddle point at
the origin. At a critical point
fx=−2x= 0, f y= 2y= 0.
Thus the origin (0 ,0) is the only critical point. Since
H(x,y) =/parenleftbigg−2 0
0 2/parenrightbigg
,
and detH(0,0) =−4<0 , the origin is a saddle point. This also follows from the
observation that the diagonal elements have different signs.
(3) Find and classify the critical points of
w=f(x,y) = [x2+ (y+ 1)2][x2+ (y−1)2].
At a critical point,
fx= 2x[x2+ (y−1)2] + 2x[x2+ (y+ 1)2] = 0
and
fy= 2(y+ 1)[x2+ (y−1)2] + 2(y−1)[x2+ (y+ 1)2] = 0.
The first equation implies x= 0 . Substituting this into the second we find y= 0,y=
1,y=−1 . Thus there are three critical points
X1= (0,0), X 2= (0,1), X 3= (0,−1).
We must evaluate the Hessian matrix at these points. Since
fxx= 12x2+ 4y2+ 4, f xy= 9xy, f yy= 4x2+ 12y2=−4,
H(X1) =/parenleftbigg4 0
0−4/parenrightbigg
, H (X2) =/parenleftbigg8 0
0 8/parenrightbigg
=H(X3).
Because det H(X1) =−16<0, X 1= (0,0) is a saddle point. Because det H(X2)>
detH(X3) = 64>0 andfxx(X2) =fxx(X3) = 8>0 , bothX2= (0,1) and
X3= (0,−1) are local minima. To complete the computation, we find f(X1) =
1,f(X2) = 0,f(X3) = 0 . A sketch of the surface is at the right.
8.2. THE MEAN VALUE THEOREM. LOCAL EXTREMA. 329
(4) Find and classify the critical points of
w=f(x,y,z ) = 1−2x+ 3x2−xy+xz−z2+ 4z+y2+ 2yz.
At a critical point,
fx=−2 + 6x−y+z,fy=−x+ 2y+ 2z,fz=x−2z+ 4 + 2y.
Solving these equations, we find only one critical point, X0= (0,−1,1),where
f(X0) = 3 . Since
fxx= 6, fxy=−1, fxz= 1fyy= 2,fyx= 2,fzz=−2,
then
H(X) =
6−1 1
−1 2 2
1 2 −2
.
Because the diagonal elements 6 ,2,−2 are not all of the same sign, by part c of the
theorem, the critical point X0= (0,−1,1) is a saddle point.
(5) Find and classify the critical points of w=f(x,y) :=x2y2. At a critical point,
fx= 2xy2= 0, f y= 2x2y= 0.
Thus the points where either x= 0 ory= 0 are all critical points. Since
fxx= 2y2, fxy= 4xy, f yy= 2x2,
we find
H(X) =/parenleftbigg2y24xy
4xy2x2/parenrightbigg
If eitherx= 0 ory= 0 , then det H= 0 so none of our tests apply to determine the
nature of the critical point. However, a glance at the function f(x,y) =x2y2reveals that
all of the points where either x= 0 ory= 0 are clearly local minima, since f= 0 there,
whilef >0 elsewhere.
Exercises
(1) Find and classify the critical points of the following functions.
(a)f(x,y) =x2−3x+ 2y2+ 10
(b)f(x,y) = 3−2x+ 2y+x2y2
(c)f(x,y) = [x2+ (y+ 1)2][4−x2−(y−1)2]
(d)f(x,y) =x3−3xy2(figure on next page)
(e)f(x,y) =xy−x+y+ 2
(f)f(x,y) =xcosy
(g)f(x,y,z ) = 2x2+ 3xz+ 5z2+ 4y−y2+ 7
330CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
(h)f(x,y,z ) = 5x2+ 4xy+ 2y2+z2−4z+ 31
(2) LetX1,...,X NbeNdistinct points in En. Find a point X∈Ensuch that the
function
f(X) =/bardblX−X1/bardbl2+···+/bardblX−XN/bardbl2
is a minimum. [Answer: X=1
N/summationtextN
j=1Xj, the center of gravity.]
(3) (a) Find the minimum distance from the origin in E3to the plane 2 x+y−z= 5.
(b) Find the minimum distance from the origin in Ento the hyperplane a1x1+
a2x2+···+anxn=c.
(c) Find the minimum distance between the fixed point X0= (˜x1,..., ˜xn) and the
hyperplane a1x1+...+anxn=c.
(d) Find the minimum distance between the two parallel planes a1x1+···+anxn=c1
anda1x1+···+anxn=c2,
(4) Iff(x,y) has two continuous derivatives, use Taylor’s Theorem (Theorem 6) to prove
f(x+h1,y+h2) =f(x,y) +fx(x,y)h1+fy(x,y)h2
+1
2[fxx(x,y)h2
1+ 2fxy(x,y)h1h2+fyy(x,y)h2
2] + (h2
1+h2
2)R,
whereRdepends on x,y,h 1andh2, and lim
h1→0
h2→0R= 0 .
(5) (a) If u(x,y) has two continuous derivatives, use the result of Exercise 4 to prove
uxx(x,y) =u(x+h1,y)−2u(x,y) +u(x−h1,y)
h2
1+h1˜R
and
uyy(x,y) =u(x,y+h2)−2u(x,y) +u(x,y−h2)
h2
2+h2ˆR,
where lim
h1→0˜R= 0 and lim
h2→0ˆR= 0 .
(b) Use part a) to deduce that if h1=h2=hthen
uxx(x,y) +uyy(x,y)
=4
h2[u(x,y)−u(x+h,y) +u(x−h,y) +u(x,y+h) +u(x,y−h)
4] +h2R
(8-2)
where lim
h→0R= 0 .
(c) Use part b) to deduce that if his small, the solution of the partial differential
equationuxx+uyy= 0 , Laplace’s equation, approximately satisfies the difference
equation
u(x,y) =u(x+h,y) +u(x−h,y) +u(x,y+h) +u(x,y−h)
4
This difference equation states that the value of uat the center of a cross equals
the arithmetic mean (“average”) of its values at the four ends of the cross. One
could use the difference equation to solve Laplace’s equation numerically.
8.2. THE MEAN VALUE THEOREM. LOCAL EXTREMA. 331
(d) Prove that any function which satisfies the above difference equation in some set
cannot have a maxima or minima inside that set. [Do not differentiate! Reason
directly from the difference equation. No computation is necessary.]
(6) If all the second partial derivatives of a function f(X) vanish identically in some open
connected set, prove that fis an affine function.
(7) (The Method of Least Squares). Let Z1,...,Z NbeNdistinct points in En, and
w1,...,w Na set ofNnumbers. We imagine the points ( Zj,wj)∈En+1to be points
on a surface MinEn+1. Find a hyperplane
w=φ(X) =c+ξ1x1+···,+ξnxn≡c+/angbracketleftξ, X/angbracketright
which most closely approximates the surface Min the sense that the error E(ξ)
E(ξ) :=N/summationdisplay
j=1|φ(Zn)−wj|2=
is minimized. Note that you are to find the coefficients ξ1,...,ξ nin the equation of
the hyperplane.
(8) (a) Let u(x,y) be a twice continuously differentiable function which satisfies the
partial differential equation
Lu:=uxx+uyy+aux+buy−cu= 0
in some open set D, where the coefficients a(x,y),b(x,y) , andc(x,y) are con-
tinuous functions. If c >0 throughout D, prove that u(x,y) cannot have a
positive maximum or negative minimum anywhere in D.
(b) Extend the result of part a) to functions u(x1,...,x n) which satisfy
Lu:=n/summationdisplay
i=1∂2u
∂x2
1+n/summationdisplay
j=1aj∂u
∂xj−cu= 0,
in some open set D, wherec>0 throughout D.
(c) Ifu(x,y) satisfies the equation of part a) and uvanishes on the boundary of
D, u ≡0 on∂D, prove that u(x,y)≡0 throughout D.
(d) Assume u(x,y) andv(x,y) both satisfy the same equation Lu= 0, Lv = 0 ,
whereLis the operator of part a). If u(x,y)≡v(x,y) on the whole boundary
ofD, prove that u(x,y)≡v(x,y) throughout the interior of D.
(9) Letfbe a twice continuously differentiable function throughout the open set A.
Prove that
(a) iffhas a local minimum at X0∈A, then its Hessian H(X0) is positive definite
or semi-definite there.
(b) iffhas a local maximum at X0∈A, then its Hessian H(X0) is negative
definite or semi-definite there.
332CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
(10) LetAbe a square n×nself-adjoint matrix and Ya fixed vector in En, and let
f(X) =/angbracketleftX, AX /angbracketright −2/angbracketleftX, Y/angbracketright.
(a) Iff(X) has a critical point at X0, proveX0satisfies the equation
AX 0=Y.
(b) IfAis positive definite and X0satisfies the equation AX 0=Y, provef(X)
defined above has a minimum at X0. [The results of this problem remain valid
ifAis any positive definite linear operator - possibly a differential operator.
The nonlinear function f(X) defines a variational problem associated with the
equationAX=Y.]
(11) Iff:A→Ehas three continuous derivatives in the open set A⊂E2containing the
origin, state precisely and prove Taylor’s Theorem with three terms about the origin.
The resulting expression will be
f(X) :=f(x,y) =f(0) +fx(0)x+fy(0)y+1
2![fxx(0)x2+ 2fxy(0)xy+fyy(0)y2]
+1
3![fxxx(Z)x3+ 3fxxy(Z)x2y+ 3fxyy(Z)xy2+fyyy(Z)y3]
whereZis on the line segment between 0 and X= (x,y) .
(12) Ifu(x,y) has the property uxy(x,y) = 0 for (x,y) in some open set, prove u(x,y) =
φ(x) +ψ(y) , whereφandψare functions of one variable.
(13) Compute the direction(s) at X0in which the following functions f
i) increase most rapidly,
ii) decrease most rapidly,
iii) remain constant.
(a)f(x1,x2) = 3−2x1+ 5x2 atX0= (2,1)
(b)f(x,y) =e2x+yatX0= (1,−2)
(c)f(x,y,z ) = 2x2+ 3xy+ 5z2+ 4y−y2+ 7 at X0(1,0,−1)
(d)f(u,v) =uv−u+v+ 2 at ( −1,1) .
8.3 The Vibrating String.
Waves. You have been hearing about them your whole life. Waves are the term used to
describe the oscillatory behavior of continuous media; water waves and sound waves being
the most familiar. We shall give a mathematical description of a very simple type of wave
- those in an oscillating violin string. The resulting mathematical model will be a second
order linear partial differential equation - the wave equation - with both initial and boundary
conditions.
8.3. THE VIBRATING STRING. 333
a) The Mathematical Model
Consider a string of length /lscriptstretched along the xaxis. Imagine the string vibrating
in the plane of the paper and let u(x,t) denote the vertical displacement of the point x
at timet. In order to end up with a tractable mathematical model several reasonable
simplifying assumptions will be made. We assume the tension τand density ρof the
string are constant throughout the motion, while the string is taken to be perfectly flexible
so the tension force in the string acts along the tangential direction. Dissipative effects (air
resistance, heating, etc.) are entirely neglected. One more assumption will be made when
needed. It essentially states that the oscillations are small in some sense.
Newton’s second law, ma=/summationtextF, is where we begin. Draw your attention to a small
segment of the string whose length, at rest, is ∆ x=x2−x1. The mass of the segment is
ρ∆x. By Newton’s second law the segment moves in such a way that the product of its
center of gravity equals the resultant of the forces acting on it. For the vertical component,
this means
ρ∆x∂2u
∂t2(˜x,t) =Fv,
where ˜x∈(x,x+ ∆x) is the horizontal coordinate of the center of gravity of the segment,
andFvmeans the vertical component of the resultant force.
There are two types of forces. One is the tension acting at both ends of the segment.
The other is gravity acting down with a force equal to the weight of the segment, ρg∆x. To
evaluate the tension forces, let θ1andθ2be the angles the string makes with the horizontal
at either end of the segment (see figure above). Then the vertical component of the tension
force is
τsinθ2−τsinθ1.
The signs indicate one force is up while the other is down. Adding the tension force to the
gravitational force and substituting into Newton’s second law, we find
ρ∆x∂2u
∂t2(˜x,t) =τ(sinθ2−sinθ1)−ρg∆x.
The dependence of θ1andθ2on the displacement can be brought out by using the
relation
sinθ=ux/radicalbig
1 +u2x,
which follows from the relation ux= tanθfor the slope of the string. Using this, we obtain
the equation
ρ∆x∂2u
∂t2(˜x,t) =τ/bracketleftBigg
ux/radicalbig
1 +u2x/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsinglex=x2−ux/radicalbig
1 +u2x/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle
x=x1/bracketrightBigg
−ρg∆x.
A simplifying assumption is badly needed. If the function ux//radicalbig
1 +u2xis expanded in
a Taylor series,
ux/radicalbig
1 +u2x=ux−1
2u3
x+···,
334CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
we see that if the slope uxis small, essentially only the linear term in this series counts.
Therefore, we do assume the slope uxis small (this is the same assumption made in treating
the simple pendulum). With this simplification, the equation of motion is
ρ∆x∂2u
∂t2(˜x,t) =τ[ux(x2,t)−ux(x1,t)]−ρg∆x.
Divide both sides of this equation by ∆ x=x2−x1and let the length of the interval shrink
to zero. Since
lim
(x2−x1)→0ux(x2,t)−ux(x1,t)
x2−x1=∂
∂xux(x,t) =∂2u
∂x2(x,t),
wherexis the limiting value of x1andx2, we find
ρ∂2u
∂t2(x,t) =τ∂2u
∂x2(x,t)−ρg
Because the length of the interval has been shrunk to one point x, the center of gravity is
now atxtoo.
It is customary to let τ/ρ=c2. The constant chas units of velocity, and, in fact, is
just the speed with which waves travel along the string. Thus
Lu:=utt−c2uxx=−g.
This is the wave equation , a second order linear inhomogeneous partial differential equation.
As was the case with linear ordinary differential equations, it is easier to attempt first to
solve the homogeneous equation
Lu:=utt−c2uxx= 0.
On physical grounds, we expect the motion u(x,t) of the string will be determined if
the initial position u(x,0) and initial velocity ut(x,0) are known, along with the motion of
both end points u(0,t) andu(/lscript,t) . However the mathematical model must be examined
to see if these four facts do determine the subsequent motion (which it should if the model
is to be of any use). Thus we must prove that given the
initial position u(x,0) =f(x), x ∈[0,/lscript]
initial velocity ut(x,0) =g(x), x ∈[0,/lscript]
motion of left end u(0,t) =φ(t)t≥0
motion of right end u(/lscript,t) =ψ(t), t≥0,
then a solution u(x,t) of the wave equation
utt−c2uxx= 0
does exist which has these properties, and there is only one such solution. Existence and
uniqueness theorems must therefore be proved.
b) Uniqueness
.
This is almost identical to all uniqueness theorems encountered earlier, especially that
for the simple harmonic oscillator in Chapter 4, Section 2.
8.3. THE VIBRATING STRING. 335
Theorem 8.11 (Uniqueness). There exists at most one twice continuously differentiable
functionu(x,t)which satisfies the inhomogeneous wave equation
Lu:=utt−c2uxx=F(x,t)
and the subsidiary
initial conditions: u(x,0) =f(x),ut(x,0) =g(x), x∈[0,/lscript]
boundary conditions: u(0,t) =φ(t),u(/lscript,t) =ψ(t), t≥0,
whereF,f,g,φ , andψare given functions.
Proof: Assumeu(x,t) andv(x,t) both satisfy the same equation and the same subsidiary
conditions. Let w(x,t) =u(x,t)−v(x,t) . ThenLw=Lu−Lv=F−F= 0 , sowsatisfies
the homogeneous equation
Lw:=wtt−c2wxx= 0
and has zero subsidiary data
initial conditions: w(x,0)≡0, wt(x,0)≡0, x∈[0,/lscript]
boundary conditions: w(0,t)≡0, w(/lscript,t)≡0, t≥0
We want to prove w(x,t)≡0 . Notice that wsatisfies the equation for a vibrating string
which is initially at rest on the xaxis, and whose ends never move. Therefore our desire
to prove the string never moves, w(x,t)≡0 , is certain physically reasonable.
For this function w,define the new function E(t)
E(t) =1
2/integraldisplay/lscript
0[w2
t+c2w2
x]dx.
We have named the function E(t) since it actually happens to be the energy in the string
associated with the motion w(x,t) at timet, except for a factor of ρ. Assume it is “legal”
to differentiate under the integral sign (it is). Upon doing so, we get
dE
dt=/integraldisplay/lscript
0[wtwtt+c2wxwxt]dx.
But an integration by parts reveals that
/integraldisplay/lscript
0wxwxtdx=wxwt/vextendsingle/vextendsingle/lscript
0−/integraldisplay/lscript
0wtwxxdx.
Because the end points are held fixed, w(0,t) = 0 and w(/lscript,t) = 0 , the velocity at those
points is zero too, wt(0,t) = 0 andwt(/lscript,t) = 0 . This drops out the boundary terms in the
integration by parts. Substituting the last expression into that for dE/dt , we find that
dE
dt=/integraldisplay/lscript
0wt[wtt−c2wxx]dx.
Butwsatisfies the homogeneous wave equation wtt−c2wxx= 0 . Therefore dE/dt ≡0 ,
so
E(t)≡constant = E(0),
that is, energy is conserved . Now
E(0) =1
2/integraldisplay/lscript
0[w2
t(x,0) +c2w2
x(x,0)]dx.
336CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
Since the initial position is zero, w(x,0) = 0 , its slope is also zero, wx(x,0) = 0 . The
initial velocity wt(x,0) is also zero, wt(x,0) = 0 . Thus
E(t)≡E(0)≡0,
that is,
0 =E(t) =1
2/integraldisplay/lscript
0[w2
t(x,t) +c2w2
x(x,t)]dx.
Because the integrand is positive, we conclude wt(x,t)≡0 andwx(x,t)≡0 . Consequently
w(x,t)≡constant. Since w(0,t) = 0 , that constant is the zero constant,
w(x,t)≡0.
Therefore
u(x,t)−v(x,t)≡w(x,t)≡0,
sou(x,t)≡v(x,t) : the solution is unique.
c) Existence
For the simple one (space) dimension wave equation, there are many ways to prove a solution
exists. The one to be given here is not the simplest (see Exercise 6 for the result of that
method), but it does generalize immediately to many other problems. It makes no difference
how we find a solution, for once found, by the uniqueness theorem it is the only possible
solution. To avoid complications, we shall consider only the homogeneous equation and
assume the end points are tied down. Thus, we want to solve
Wave equations: utt−c2uxx= 0.
Initial conditions: u(x,0) =f(x), ut(x,0) =g(x).
Boundary conditions: u(0,t) = 0, u(/lscript,t) = 0.
The idea is first to find special solutions u1(x,t), u2(x,t),..., which satisfy the bound-
ary conditions but do not necessarily satisfy the initial conditions. Then, as was done for
linear O.D.E.’s, we build the solution which does satisfy the given initial conditions as a
linear combination of these special solutions,
u(x,t) =/summationdisplay
Ajuj(x,t),
that is, by superposition.
Let us seek special solutions in the form of a standing wave ,
u(x,t) =X(x)T(t).
HereX(x) andT(t) are functions of one variable. Our procedure is reasonably called
separation of variables . Substitution of this into the wave equation gives
¨T(t)X(x)−c2X/prime/prime(x)T(t) = 0,
or
X/prime/prime(x)
X(x)=1
c2¨T(t)
T(t).
8.3. THE VIBRATING STRING. 337
Since the left side depends only on x, while the right depends only on t, both sides must
be constant (a somewhat tricky remark; think it over). Let that constant be −γ(using
−γinstead ofγis the result of hindsight, as you shall see).
X/prime/prime
X=1
c2¨T
T=−γ.
This leads us to the two ordinary differential equations
X/prime/prime(x) +γX(x) = 0, ¨T(t) +γc2T(t) = 0.
Sinceu(0,t) = 0 and u(/lscript,t) = 0 and u(x,t) =X(x)T(t) , the function X(x) must also
satisfy the boundary conditions
X(0) = 0, X (/lscript) = 0.
There are several ways to show γmust be positive. Perhaps the simplest is to observe
that ifγ <0 orγ= 0 , the only function X(t) which satisfies the differential equation
X/prime/prime+γX= 0 and boundary conditions X(0) =X(/lscript) = 0 is the zero function X(x)≡0 .
Since for this function u(x,t) =X(x)T(t)≡0 , it is devoid of further interest.
Another way to show γis positive is to multiply the ordinary differential equation
X/prime/prime+γX= 0 byX(x) and integrate over the length of the string,
/integraldisplay/lscript
0[X(x)X/prime/prime(x) +γX2(x)]dx= 0.
Upon integrating by parts, we find that
/integraldisplay/lscript
0?(x)X/prime(x)dx=XX/prime/vextendsingle/vextendsingle/vextendsingle/lscript
0−/integraldisplay/lscript
0X/prime2(x)dx.
SinceX(0) =X(/lscript) = 0 , the boundary terms drop out. Substituting this into the above
equation, we find that/integraldisplay/lscript
0X/prime2(x)dx=γ/integraldisplay/lscript
0X2(x)dx.
IfX(x) is not identically zero, this can be solved for γ
γ=/integraldisplay/lscript
0X/prime2(x)dx
/integraldisplay/lscript
0X2(x)dx,
and clearly shows γ >0 .
Enough for that. The solution of X/prime/prime+γX= 0, γ > 0 , is
X(x) =Acos√γx+Bsin√γx.
The boundary condition X(0) = 0 implies A= 0 , while the boundary condition at the
other end point X(/lscript) = 0 , implies
0 =Bsin√γ/lscript.
338CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
IfB= 0 too, then X(x)≡0 , sou(x,t)≡0 . This is of no use to us. The only alternative
is to restrict γso that sin√γl= 0 . This means√γ/lscriptis a multiple of π,√γ/lscript=nπ, n =
1,2,...,
√γ=nπ
/lscript, n = 1,2,....
There is then one possible solution X(x) for each integer n,
Xn(x) =Bnsinnπ
/lscriptx,
where the constants Bnare arbitrary.
Remark: There is a similarity of deep significance for mathematics and physics between
the work in these last few paragraphs and that done for the coupled oscillators in Chapter
6. There (p. 528-9), we had an operator Aand wanted to find nonzero vectors Snand
numbersλsuch that
ASn=λnSn.
The numbers found λnwere called the eigenvalues of A, andSnthe corresponding eigen-
vectors.
Here, we were given the operator A=−d2
dx2and wanted to find nonzero functions
Xn(t)∈ {X∈C2[0,/lscript]:X(0) =X(/lscript) = 0}which satisfy the equation
AXn=γnXn
The numbers found, γn=n2π2//lscript2, are also called the eigenvalues ofA, and the function
Xn(t) = sinnπ
/lscriptx, the eigenfunction ofAcorresponding to the eigenvalue γn.
Associated with each possible eigenvalue γn, there is a solution of the time equation,
¨T+γc2T= 0 ,
Tn(t) =Cncosncπ
/lscriptt+Dnsinncπ
/lscriptt.
We therefore have found one special solution, un(x,t)−Xn(t)Tn(t) , for each value of
the indexn,
un(x,t) = sinnπx
/lscript(αncosncπt
/lscript+βnsinncπt
/lscript).
The arbitrary constants have been lumped in this equation. These special solutions are the
“natural” vibrations of the string, or normal modes of vibration . A snapshot at t=t0of
the string moving in the nth normal mode would reveal the sine curve
un(x,t0) =Csinnπx
/lscript,
the constant Caccounting for the remaining terms, which are constant for tfixed. In
music, the integer nrefers to the octave. The fundamental tone is the case n= 1 , while
the tone for n= 2 , the second harmonic orfirst overtone , is one octave higher.
a figure goes here
8.3. THE VIBRATING STRING. 339
The time frequency Vnof thenth normal mode is Vn=ncπ
/lscript, this is the number of
oscillations in 2 πunits of time. It is the time frequency which we usually associate with
musical pitch. The (time) periodτnof thenth normal mode is 2 π/Vn, that isτn= 2l/nc.
Another name you will want to know is the wave length λnof thenth normal mode,
λn= 2/lscript/n(see figures above). Notice that Vnλn=c, an important relationship.
Having found the special normal mode solutions, un(x,t) , we hope that arbitrary
constantsαnandβncan be chosen so a linear combination
u(x,t) =∞/summationdisplay
n=1un(x,t) =∞/summationdisplay
n=1(αncosncπt
/lscript+βnsinncπt
/lscript) sinnπx
/lscript
will satisfy the given initial conditions. Every function u(x,t) of this form automatically
satisfies the boundary conditions u(0,t) = 0, u(/lscript,t) = 0 since each of the un’s satisfy
them.
Ifu(x,0) =f(x) andut(x,0) =g(x) , then from the above equation, we must have
f(x) =∞/summationdisplay
n=1un(x,0) =∞/summationdisplay
n=1αnsinnπx
/lscript
and
g(x) =∞/summationdisplay
n=1∂un
∂t(x,0) =∞/summationdisplay
n=1nπc
/lscriptβnsinnπx
/lscript.
Thus, the coefficients αnare the coefficients in the Fourier sine series for f, while the βn
are essentially the coefficients in the Fourier sine series for g. In fact, this is how Fourier
was led to the series bearing his name. These formulas for u(x,y), f(x) , andg(x) become
easier on the eye if the length of the string is π,/lscript=π. Then
u(x,y) =∞/summationdisplay
n=1(αncosnct+βnsinnct) sinnx,
while
f(x) =∞/summationdisplay
n=1un(x,0) =∞/summationdisplay
n=1αnsinnx, (8-3)
and
g(x) =∞/summationdisplay
n=1∂un
∂t(x,0) =∞/summationdisplay
n=1ncβnsinnx.
Finding the coefficients αnandβnis particularly simple if fandgcan be represented
by finite series.
Examples: Find the solution u(x,t) of the wave equation for a string of length π, l=π,
which is pinned down at its end points, u(0,t) =u(π,t) = 0 , and satisfies the given initial
conditions.
(1)u(x,0) =f(x) = 2 sin 3x, u t(x,0) =g(x) =1
2sin 4x. We have to find αnandβnfor
the two series
2 sin 3x=∞/summationdisplay
n=1αnsinnx
340CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
1
2sin 4x=∞/summationdisplay
n=1ncβnsinnx.
For these simple functions, just match coefficients, giving
α3= 2, αn= 0, n/negationslash= 3,andβ4=1
8c, βn= 0, n/negationslash= 4.
Therefore, the sum of the two waves
u(x,t) = 2 cos 3ctsin 3x+1
8csin 4ctsin 4x
is the (unique!) solution of this example.
(2)u(x,0) =f(x) =1
2sin 3x−sin 17xand
ut(x,0) =g(x) =−9 sinx+ 13 sin 973 x.
We have to find αnandβnfor the two series
1
2sin 3x−sin 17x=∞/summationdisplay
n=1αnsinnx
and
−9 sinx+ 13 sin 973 x=∞/summationdisplay
n=1ncβnsinnx.
By matching again, we find α3=1
2, α17=−1 , andαn= 0 forn/negationslash= 3 or 17 . Also,
β1=9
c,β973=13
973c, andβn= 0 forn/negationslash= 1 or 973 . The (unique) solution is then a
sum of four waves
u(x,t) =−9
3sinctsinx+1
2cos 3ctsin 3x
−cos 17ctsin 17x+13
973csin 973ctsin 973x.
Sincefandgare not usually given in the simple form of these examples, the full
Fourier series is needed. Recall that the string is pinned down at both ends. Therefore
both the initial position function f(x) and velocity function g(x) have the property
f(0) =f(π) = 0 , and g(0) =g(π) = 0 , where we have taken the length of the string
to beπ. It is now possible to extend both fandg, assumed continuous in [0 ,π] ,
to the whole interval [ −π,π] as continuous odd functions,
a figure goes here
that is, ifx∈[0,π] , we can define
f(−x) =−f(x) andg(−x) =−g(x),
since the right sides, −f(x) and −g(x) , are known functions for x∈[0,π] .
As odd functions now on the whole interval [ −π,π] , the functions fandghave
Fourier sine series (cf. p. 252, Exercise 3a).
8.3. THE VIBRATING STRING. 341
f(x) =∞/summationdisplay
n=1bnsinnx√π
g(x) =∞/summationdisplay
n=1˜bnsinnx√π
where
bn= 2/integraldisplayπ
0f(x)sinnx√πdx, ˜bn= 2/integraldisplayπ
0g(x)sinnx√πdx (8-4)
Comparing with the previous formulas (3) for fandg, we find
αn=bn/√π,andβn=˜bn/nc√π
Consequently
u(x,t) =∞/summationdisplay
n=1(bncosnct√π+˜bn
ncsinnct√π) sinnx (8-5)
the coefficients bnand˜bnbeing determined from the initial conditions by equation
(4).
Thus, we have almost proved
Theorem 8.12 . Iff(x)is twice continuously differentiable and g(x)once continuously
differentiable for x∈[0,π]and both functions vanish at x= 0 andx=π, then the
functionu(x,t)defined by equation (5) is a solution of the homogeneous wave equation
utt−c2uxx= 0
and satisfies the
initial conditions: u(x,0) =f(x), ut(x,0) =g(x), x∈[0,π],
as well as the
boundary conditions: u(0,t) = 0, u(π,t) = 0, t≥0,
wherebnand ˜bnare determined from fandgthrough equations (4). Moreover, this
solution is unique (by Theorem 9).
Outline of Proof . If it is possible to differentiate the infinite series (5) term by term u(x,t)
would satisfy the wave equation since each special solution un(x,t) does. In any case, the
initial condition u(x,0) =f(x) is clearly satisfied. However, checking the other initial
conditionut(x,0) =g(x) also involves differentiating the infinite series term by term.
Thus, we must only justify the term by term differentiation of an infinite Fourier series.
For power series, we found (p. 82-3, Theorem 16) we can always differentiate term by
term within its disc of convergence. Such is not the case with Fourier series. For example,
the Fourier series∞/summationdisplay
n=1sinn2x
n2converges for all x, but the series obtain by differentiating
formally,∞/summationdisplay
n=1cosn2xdiverges at x= 0 . However, if a function is sufficiently smooth, its
Fourier series can be differentiated term by term and does converge to the derivative of the
342CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
function. Since the details of a complete proof are but a rehash of the proof carried out for
power series (p. 82ff), we omit it.
Example: Find the displacement u(x,t) of a violin string of length πwith fixed end
points which is plucked at its midpoint to height h. The initial position is then
f(x) =/braceleftbiggxh, x ∈[0,π/2]
(π−x)h, x∈[π/2,π],
and the initial velocity, g(x) , is zero.
We must find the coefficients bnand˜bnin the series (5). After mentally continuing f
andgto the interval [ −π,π] as odd functions, the formulas (4) give us bnand˜bn,
bn= 2/integraldisplayπ
0f(x)sinnx√πdx=2h√π/braceleftBigg/integraldisplayπ/2
0xsinnxdx +/integraldisplayπ
π/2(π−x) sinnxdx/bracerightBigg
.
Integrating and simplifying, we find that
bn=4h√πn2sinnπ
2=
0, neven
1, n= 1,5,9,13,...
−1, n= 3,7,11,15.
Fromg(x)≡0 , it is immediate that βn= 0 for all n. Thus,
u(x,t) =4h
π∞/summationdisplay
n=11
n2sinnπ
2cosnctsinnx
=4h
π[cos 3ctsinx
1−cosctsin 3x
32+cos 5ctsin 5x
52+···]
is the desired solution.
Exercises
(1) (a) Find a solution u(x,t) of the homogeneous wave equation for a string of length
πwhose end points are held fixed if the initial position function is
u(x,0) =1
2sin 4x−sin 7x,
while the initial velocity is
ut(x,0) = sin 3x+ sin 73x.
(b) Same problem as a), but
u(x,0) = sin 5x+ 12 sin 6x−7 sin 9x
ut(x,0) =−sinx+ 91 sin 273 x.
8.3. THE VIBRATING STRING. 343
(2) Find a solution u(x,t) of the homogeneous wave equation for a string of length π
whose end points are held fixed if the string is initially plucked at the point x=π/4
to the height h.
(3) Consider a vibrating string of length /lscriptwhose end points are on rings which can slide
freely on poles at 0 and /lscript. Then the boundary conditions at the end points are
ux(0,t) = 0, ux(/lscript,t) = 0
that is, zero slope.
(a) Use the method of separation of variables to find the form of special standing
wave solutions. [Answer: un(x,t) = cosnπx
/lscript(αncosncπt
/lscript+βnsinncπt
/lscript) ].
(b) Use these to find a solution with the initial conditions
u(x,0) = cosx−6 cos 3x (let/lscript=π)
ut(x,0) =1
2cos 2x.
(4) Letu(x,t) satisfy the homogeneous wave equation. Instead of keeping the end points
fixed, we either put them on rings (cf. Exercise 3) or attach them by elastic bands,
in which case the boundary conditions become
ux(0,t)−c1u(0,t) = 0, ux(π,t) +c2u(π,t) = 0, c1,c2≥0.
(a) Define the energy as before, and prove that energy is dissipated with these bound-
ary conditions, unless c1andc2vanish.
(b) Prove there is at most one function u(x,t) which satisfies the inhomogeneous
wave equation utt−c2uxx=F(x,t) with initial conditions as before, but with
elastic boundary conditions
ux(0,t)−c1u(0,t) =φ(t), ux(π,t) +c2u(π,t) =ψ(t),
wherec1andc2are non-negative constants.
(5) To account for the effect of air resistance on a vibrating string, one common assump-
tion is that the resistance on a segment of length ∆ xis proportional to the velocity
of its center of gravity,
Fres=−k∆xut(˜x,t), k> 0,
wherekis a numerical constant. This is analogous to the standard viscous resistance
force on a harmonic oscillator.
(a) Find the equation of motion ignoring gravity. [Answer:1
c2utt+kut=uxx]
(b) Find the form of the special standing wave solutions, assuming, the end points
are held fixed.
(c) Write a formula giving the probable form for the general solution u(x,t) .
(d) If the end points are pinned down, what do you expect the behavior of the string
will be as t→ ∞ ? Does the formula found in part c) verify your belief (it
should).
344CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
(e) Define the energy E(t) as before and show that energy is dissipated if the ends
are held fixed.
(f) Use the result of e) to prove ˙E(t) + 2kE(t)≥0 , and conclude that E(t)≥
E(0)e−2ktfort≥0 . This shows that the energy is not dissipated too rapidly.
(6) It is possible to write the solution of the homogeneous wave equation for a string of
lengthπwith fixed end points in a simple closed form by using the trigonometric
identities
2 sinnxcosnct= sinn(x−ct) + sinn(x+ct).
2 sinnxsinnct= sinn(x−ct)−cosn(x+ct).
(a) Do this and obtain d’Alembert’s formula
u(x,t) =f(x−ct) +f(x+ct)
2+1
2c/integraldisplayx+ct
x−ctg(ξ)dξ.
(b) Solve the example of a plucked string (p. 641) again using this formula. Draw
two sketches, one indicating the position of the string at time t=π
2cand another
att=π
c.
(7) (a) Prove the wave operator L:=∂2
∂t2−c2∂2
∂x2,ca constant, is translation invariant,
that is, ifT:u(x,t)→u(x+x0, t+t0),prove (LT)u= (TL)ufor all values of
x0andt0, and for all functions ufor which the operators make sense.
(b) Find the function φ(a,b) in the formula
Leax+bt=φ(a,b)eax+bt.
(c) Use part b) to show that if ais any constant, the four functions
ea(x+ct),e−a(x+ct),ea(x−ct),e−a(x−ct)
are solutions of the homogeneous wave equation Lu= 0 .
(d) Use the fact that each of the above functions satisfies the ordinary differential
equationv/prime/prime(x) =a2v(x) to conclude that if linear combinations of these func-
tions are to satisfy the boundary conditions v(0) =v(/lscript) = 0 , then necessarily
a2<0 , so the constant ais pure imaginary and we can write a=iγ, whereγ
is real.
(e) Letu(x,t) be a linear combination of the four functions part c) with a=iγ.
Show that u(x,t) may be written in the form
u(x,t) = sinγx [Acosγct+Bsinγct].
(f) Ifu(0,t) =u(/lscript,t) = 0 , show that γn=nπ
/lscript. Find an infinite set of special solu-
tionsun(x,t) which satisfy the homogeneous wave equation with zero boundary
values [From here on, one proceeds as before to find the general solution. This
problem has shown how the idea of translation invariance can also be used to
lead one to the special solutions un].
8.3. THE VIBRATING STRING. 345
(8) (a) By inspection , find a particular solution for the solution of the inhomogeneous
wave equations
Lu:=utt−c2uxx=g, g ≡constant.
(b) How can this particular solution be used to find the solution of the equation
Lu=gwhich has given initial conditions and zero boundary conditions?
(9) Flow of heat in a thin insulated rod on the xaxis is governed by the heat equation
ut(x,t) =k2uxx(x,t),
whereu(x,t) represents the temperature at the point xat timet, andk2, the
diffusivity , is a constant depending on the material. The “energy” in a rod of length
/lscript,0≤x≤/lscript, is defined as
E(t) =1
2/integraldisplayl
0u2(x,t)dx.
(a) If the ends of the rod have zero temperature, u(0,t) =u(/lscript,t) = 0 , prove “energy”
is dissipated, ˙E(t)≤0 , by showing
dE(t)
dt=−k2/integraldisplay/lscript
0u2
x(x,t)dx.
(b) Given a rod whose ends have zero temperature and whose initial temperature is
zero,u(x,0) = 0 , prove that the temperature remains zero, u(x,t)≡0 .
(c) Prove the temperature of a rod is uniquely determined if the following three data
are known:
initial temperature: u(x,0) =f(x), x∈[0,/lscript].
boundary conditions: u(0,t) =φ(t), u(/lscript,t) =ψ(t), t≥0.
(d) Use the method of separation of variables to find an infinite number of special
solutions of the heat equation for a thin rod whose end points have zero temper-
ature for all t≥0 . [Answer: un(x,t) =cne−n2k2π2
/lscript2tsinnπ
/lscriptx, n = 1,2,...]
(e) If the ends of a rod have zero temperature for all t≥0 , what do you intuitively
expect the temperature u(x,t) will be as t→ ∞ ? Is this borne out by the
formulas for the special solutions?
(f) Find the temperature distribution in a rod of length πif the ends have zero
temperature and if the initial temperature distribution in the rod is
u(x,0) = sinx−4 sin 7x,
(10) If the temperature at the ends of the bar of length /lscriptis constant but not necessarily
zero, say
u(0,t) =θ1, u (/lscript,t) =θ2,
the temperature distribution can be found be splitting the solution into two parts,
u(x,t) = ˜u(x,t) +up(x,t) , whereup(x,t) is a particular solution having the correct
temperature at the ends of the bar and u(x,t) is a general solution which has zero
temperature at the ends.
346CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
(a) Find a particular solution of the homogeneous heat equation ut=k2uxxwhich
satisfiesu(0,t) = 200, u(/lscript,t) = 500, but does not necessarily satisfy any
prescribed initial condition. [Answer: Many possible solutions - for example
up(x,t) = 20 + 30x
/lscript, orup(x,t) = 20 + 30 sinπx
2/lscript] .
(b) Find the temperature distribution in a rod of length πif the initial temperature
isu(x,0) = 2 sinx−sin 4x, while the boundary conditions are as in part a).
(11) If the ends of a bar of length /lscriptare insulated instead of being kept at zero, the
boundary conditions are
ux(0,t) =ux(/lscript,t) = 0.
(a) Use the method of separation of variables to find an infinite number of spe-
cial solutions for the homogeneous heat equation with insulated ends. [Answer:
un(x,t) =cne−n2k2π2
/lscript2tcosnπx
/lscript, n= 0,1,2,...].
(b) What is the temperature distribution in a rod whose ends are insulated if the
initial temperature distribution is
u(x,t) = 3 cos2πx
/lscript−1
5cos5πx
/lscript.
(12) In this exercise you will find a quantitative estimate for the rate of decrease of energy
for the heat in a rod of length /lscriptwith zero temperature at the ends.
(a) Use the result of Exercise 9a to prove the differential inequality
dE
dt≤ −cE(t),
wherecis a positive constant. [Hint: Look at p. 227 Exercise 15c].
(b) Conclude that
E(t)≤E(0)e−ct, t≥0.
This is the desired estimate for the decrease of energy in the rod.
(13) The linear partial differential equation
uxx−u=ut
governs the temperature distribution in a rod of length /lscriptmade up of a material which
uses up heat to carry out a chemical process. Define the energy E(t) in the rod as
in Exercise 9.
(a) Prove that if the ends of the rod have zero temperature, then the energy is
dissipated, ˙E(t)≤0 .
(b) Given a rod whose ends have zero temperature and whose initial temperature
u(x,0) is zero, use a) to prove that the temperature remains zero, u(x,t)≡
0, t≥0 .
(c) Use part b) to prove that the temperature of the rod described above is uniquely
determined if the following three data are known
u(x,0) forx∈[0,/lscript], u(0,t) andu(/lscript,t) fort≥0.
8.4. MULTIPLE INTEGRALS 347
(14) In setting up the mathematical model for the vibrating string, we never examined the
horizontal components of the forces.
(a) Show that the net horizontal force is
Fh=τcosθ2−τcosθ1
(b) Under our assumption uxis small, show that the net horizontal force is zero -
so there is no horizontal motion of the string. This justifies the statement that
the motion of the string is entirely vertical.
(15) Use the formula Vn=nπc//lscript (page 635) for the frequency and the relationship be-
tweenc,T andρ(page 624) to derive a formula for Vnin terms of the physical
constants/lscript,T, andρfor a vibrating string. Interpret the effect on the frequency,
Vn, if the physical constants are changed. Does this agree with your experience in
tuning stringed instruments?
8.4 Multiple Integrals
How can we extend the notion of integration from functions of one variable to functions of
several variables? That is the problem we shall face in this section.
Letw=f(X) =f(x1,...,x n) be a scalar-valued function defined in C⊂En. For
the purposes of this section it will be convenient to think of fas either the height function
for a surface MinEn+1overD, or as the mass density of D. In the first case./integraldisplay/integraldisplay
Df
should be the volume of the solid contained between MandD(see fig.), whereas in the
second case,/integraldisplay/integraldisplay
Dfshould be total mass of the set D.
Two problems have to be solved. First, define the integral in En. Second, give a
reasonable procedure for explicitly evaluating the integral in sufficiently simple situations.
More so than for the single integral, the problem of defining the multiple integral bristles
with technical difficulties. However, after this is done the evaluation of integrals in Encan
be reduced to the evaluation of repeated integrals, that is, a sequence of nintegrals in E1,
which is in turn effected not by using the definition of the integral, but rather by recourse
to the fundamental theorem of calculus.
Before starting the formalities, it is well advised to see where some difficulties lie.
Suppose we are given a density function fdefined on some domain Dand want to find
the total mass of D. To make things even simpler, assume for the moment that the density
is constant and equal to 1, for all X∈D⊂En. Then the mass coincides with the volume
of the domain. For the special case of functions of one variable D⊂E1is an interval so
the “volume” of D(really the length of D) is trivial
a figure goes here
to compute, Vol ( D) =b−a. However if Dhas two or more dimensions, even finding the
volume ofD(area ifD⊂E2) is itself difficult.
The problem is that a connected set DinE1can only be a line segment, whereas a
connected open set in En, n≥Zcan be much more complicated topologically. In E1, the
348CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
closed “cube” and closed “ball” are both intervals [ a,b] , and every other connected set is
also an interval. In E2, not only do the cube and ball become distinct, but also a slew of
other possibilities arise. Dmay be riddled with holes and its
a figure goes here
boundary wild (contrasted to the boundary of a connected set in E1which is always just
two points, the end points of the interval). It should be clear that the notion of volume of
a setDmay only be definable if the boundary of Dis sufficiently smooth.
As you should be anticipating, the volume of a set Dwill be defined by filling it up
with little cubes of volume ∆ x1∆x2...∆xn= ∆V, and then proving that as the size of
the cubes becomes small, the sum of volumes of the cubes approaches a limit (here is where
the smoothness of θDenters). In two dimensions, D⊂E2, this roughly reads
Area (D) = lim
∆x→0
∆y→0/summationdisplay/summationdisplay
∆x∆y=/integraldisplay
Ddxdy.
Only after the volume of a domain is defined can the more general notion of mass
of a setDfor a density function fbe defined. The procedure here is straightforward,
however it is important that the density fbe “essentially” continuous. Using the same
approximating cubes, we assign to each little cube its approximate density, say by using the
value of the density fat the center of the little cube. Adding up the masses of these little
cubes and passing to the limit again, we find the total mass of the solid Dwith density f.
Again, in two dimensions this roughly reads
Mass (D) = lim
∆x→0
∆y→0/summationdisplay/summationdisplay
f(xi,yj)∆x∆y=/integraldisplay/integraldisplay
Df(x,y)dxdy.
Because of the technical complications, we shall only state a series of propositions which
give the existence of the integral. The proofs of several crucial - but believable - results will
not be carried out, but can be found in many advanced calculus books. For convenience,
the geometric language of the plane, E2, will be used. The ideas extend immediately to
higher dimensions. Now some terminology.
Definition: Ashaved rectangle is a rectangle with its bottom and left sides omitted, that
is, a set of the form
Q={X= (x1,x2):aj<xj≤bj, j= 1,2}.
Arectangular complex is a finite union of shaved rectangles, which can always be assumed
disjoint, that is, non-overlapping. This should more accurately be called a shaved rectan-
gular complex, but is not for the sake of euphony.
IfDis a set, the characteristic function of D,XDis defined by
XD(X) =/braceleftbigg1, X ∈D
0, X /ownerD.
Astep function s(X) is a finite linear combination of characteristic functions of shaved
rectangles. The graph of this function looks like its name implies.
8.4. MULTIPLE INTEGRALS 349
a figure goes here
A function fhascompact support if it is identically zero outside some sufficiently large
rectangle. The support of a particular function f, written supp f, is the smallest closed
set outside of which fis zero. Thus, it is the set of all points Xwheref(X)/negationslash= 0 and the
limit points of those points.
We take the area of a shaved rectangle Qas a known quantity - the height times base,
anddefine the integral as
I(XQ) =/integraldisplay/integraldisplay
E2XQdA=/integraldisplay
DdA≡Area (Q),
where the Area ( Q) is defined in the natural way as length ×width. You may wish to think
ofdAas representing an “infinitesimal element of area”. We however assign no meaning
to the symbol and use it only as a reminder. Some prefer to do without it altogether and
write /integraldisplay/integraldisplay
E2XQ= Area (Q).
Our task is to define
I(f)≡/integraldisplay/integraldisplay
E2fdA
for density functions other than XQ’s. For example, if Dis some set, for the function XD
we want to define
Area (D) =/integraldisplay/integraldisplay
E2XDdA=/integraldisplay/integraldisplay
DdA
But this will not make sense unless it is shown that the set Ddoes have a number associated
with it which has the properties of area. It is easy to define the integral of a step function
S. Let
S(X) =n/summationdisplay
j=1ajXQj(X),
where the Qj’s are disjoint. Then/integraldisplay/integraldisplay
SdA should represent the total mass of a plate
composed of rectangles Q1,...,Q nwith respective densities a1,...,a n. Thus, we define
I(S) =/integraldisplay/integraldisplay
SdA≡a1Area (Q1) +...+anArea/prime,(Qn) =n/summationdisplay
j=1aj/integraldisplay
XQjdA.
The integrals of step functions clearly satisfy the following
Lemma 8.13 . IfS1(X)andS2(X)are step functions, then
a).I(aS1+bS2) =aI(S1) +bI(S2).
b).S1(X)≤S2(X)impliesI(S1)≤I(S2).
c). IfS(X)is bounded by M, S (X)≤M, then
I(S)≤cM,
wherecis the area of the support of S.
350CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
The integral of any other more complicated function is defined by using step functions.
Definition: A function f:E2→EisRiemann integrable if given any /epsilon1>0 , there are step
functionssandSwiths(X)≤f(X)≤S(X) for allX∈E2such thatI(S)−I(s)</epsilon1,
that is /integraldisplay/integraldisplay
E2SdA−/integraldisplay/integraldisplay
E2sdA</epsilon1.
Intuitively, a function is Riemann integrable if it can be trapped between two step
functionsSandsin such a way that the integrals of Sandsdiffer by an arbitrarily
small amount.
Definition: Iffis Riemann integrable, let Snandsnbe a trapping sequence , forf,
that is,sn(X)≤f(X)≤Sn(X) andI(Sn)−I(sn)<1
n. Then the Riemann integral of
f, I(f) is defined as (cf. page 21, for the definition of l.u.b. = least upper bound, and of
g.l.b.).
I(f)≡l.u.b. n→∞I(sn)
We could have equivalently defined I(f) asI(f) = g.l.b.n→∞I(Sn) . Since both limits
are the same, it is irrelevant. However, it is important to show that I(f) has the same
value if any other trapping sequence ˆSn(X),ˆsn(X) is used. This is the content of
Lemma 8.14 . Iffis Riemann integrable, then I(f)does not depend on which trapping
sequences are used. Proof not given.
Now we exhibit a class of functions which are Riemann integrable. The issue boils down
to finding functions which can be approximated well by step functions.
Lemma 8.15 . Iffis a continuous function and Dis a closed and bounded set, then
fcan be approximated arbitrarily closely from above and below by step functions Sands
throughout D. Thus, given any /epsilon1>0, there are step functions Sandssuch that
0≤S(X)−f(x)</epsilon1, and 0≤f(X)−s(X)</epsilon1 for allX∈D.
Proof not given.
Theorem 8.16 . Iffis a continuous function with compact support, then it is Riemann
integrable.
Proof: LetS(X) ands(X) be as in the lemma where Dis the support of f. Then
s(X)≤f(X)≤S(X)
and
S(X)−s(X) = [S(X)−f(X)] + [f(X)−s(X)]<2/epsilon1.
Thus by Lemma 1,
I(S)−I(s) =I(S−s)<2c/epsilon1,
wherecis the area of the set (supp S)∪(supps).
Becausefhas compact support, the constant cis bounded. Therefore the factor
2c/epsilon1can be made arbitrarily small by choosing /epsilon1small. This verifies all the conditions for
integrability.
8.4. MULTIPLE INTEGRALS 351
We have disposed of the problem of integrating continuous functions with compact
support. Notice that the above procedure is identical to that used for functions of one
variable (see figure.)
We still do not know how to find the area of a domain D. Although we anticipate that
Area (D) =I(XD) , this does not yet make sense (except for rectangular complexes) since
thediscontinuous functionXDis not covered by Theorem 1). Let us remedy this now.
The problem is to show the boundary ∂Ddoes not have any area.
Definition: A set in E2hascontent zero if it can be enclosed in a rectangular complex
whose total area is arbitrarily small. Thus, if a set has content zero, given any /epsilon1>0 , there
is a rectangular complex Rcontaining ∂Dsuch that
Area (R) =I(XR)</epsilon1.
It should be clear that any set with a finite number of points has content zero (since
each point can be enclosed on a square of side /epsilon1, so the total area of Nsuch squares is
N/epsilon12, which can be made arbitrarily small.) One would also expect that curves will have
zero content. This is not necessarily true unless the curve is not too badly behaved.
Lemma 8.17 . If a curve is composed of a finite number of smooth curves, then it has
zero content. In particular, if the boundary ∂Dof a bounded domain Dis such a curve,
it has zero content. Proof not given.
Theorem 8.18 . If the boundary ∂Dof a domain D⊂E2has content zero, then the
functionXDis Riemann integrable. Consequently, the area of Dis definable and given by
Area (D) =/integraldisplay/integraldisplay
E2XDdA=/integraldisplay/integraldisplay
DdA.
Proof: Almost identical to that for Theorem 11. Let /epsilon1 > 0 be given and let Rbe
the rectangular complex which encloses the boundary ∂D, whereRhas area less than
/epsilon1, I(XR)< /epsilon1. Then the part of Dwhich is enclosed by R, D −=D−R∩D, is a
rectangular complex as is D+=R∪D−andD+−D−=R. SinceD+⊃D⊃D−, we
have
XD−(X)≤XD(X)≤XD+(X) for all X.
Also,
I(XD+)−I(XD−) =I(XR)</epsilon1.
ThusXDis trapped by the step functions S=XD+ands=SD−andI(S)−I(s)</epsilon1,
proving the theorem.
It is now possible to define /integraldisplay/integraldisplay
DfdA
for continuous functions fwhereDis not necessarily the support of f.
Theorem 8.19 . Iffis continuous in a closed and bounded set Dwhose boundary ∂D
has content zero, then the function fXDis Riemann integrable and
/integraldisplay/integraldisplay
DfdA≡I(fXD).
352CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
Proof: LetRbe the rectangular complex which encloses ∂Dand has area less than
/epsilon1, I(XR)</epsilon1. TakeD−=D−R∩DandD+=R∪D−as in Theorem 12. Further let
S1ands1be step functions which trap fwithin/epsilon1for allX∈D−(this is possible by
Lemma 3)
0≤S1(X)−f(X)</epsilon1,0≤f(X)−s1(X)</epsilon1 for allX/epsilon1D −,
so
0≤S1(X)−s1(X)<2/epsilon1for allX∈D.
LetMbe an upper bound for |f|onD,|f(X)| ≤Mfor allX∈D−. Then define
S=S1+MX Rands=s1−MX R.
These functions Sandstrapfon all ofD,
s(X)≤f(X)≤S(X) for all X∈D,
that is,
s≤fXD≤Sfor allX.
Furthermore
I(S−s) =I(S1−s1) + 2MI(XR)
<2c/epsilon1+ 2M/epsilon1= (2c+ 2M)/epsilon1,
wherecis the area of D−. SinceSandsare step functions which trap f, and since
I(S−s) can be made arbitrarily small, the proof that fXDis Riemann integrable is
completed. We follow custom and write
I(fXD)≡/integraldisplay/integraldisplay
DfdA.
Except for the three unproved lemmas, this completes the proof of the existence of the
integral. The next theorem summarizes some important properties of the integral.
Theorem 8.20 . Iffandgare Riemann integrable, then
a).I(af+bg) =aI(f) +bI(g), a,b constants
b).f≤gimpliesI(f)≤I(g).
c).|I(f)| ≤I(|f|)
Proof:
a) and b) are immediate consequences of the corresponding statements for step functions
(Lemma 1) and the definition of the Riemann integral as the limit of step functions. To
prove c), we first observe that if fis integrable, so is |f|. Since −|f| ≤f≤ |f|, by parts
a and b
−I(|f|)≤I(f)≤I(|f|),
which is equivalent to the stated property.
Although the approximate value of the integral/integraldisplay/integraldisplay
DfdA can be evaluated by using the
procedures of the above theorems, we have as yet no routine way of evaluating the integral
8.4. MULTIPLE INTEGRALS 353
iffandDare simple. Some notation will suggest the method. Write dA=dxdy and
think ofdxdy as the area of an “infinitesimal” rectangle. Then
/integraldisplay/integraldisplay
DfdA =/integraldisplay/integraldisplay
Df(x,y)dxdy.
IfDis the domain in the figure, it is reasonable to evaluate the double integral, which we
shall think of as the mass of Dwith density f, by first finding the mass of a horizontal
strip
g(y) =/integraldisplayγ2
γ1f(x,y)dx,
and then adding up the horizontal strips to find the total mass
/integraldisplay/integraldisplay
Df(x,y)dxdy =/integraldisplayγ4
γ3g(y)dy=/integraldisplayγ4
γ3/parenleftbigg/integraldisplayγ2
γ1f(x,y)dx/parenrightbigg
dy.
The integral on the right is called an iterated orrepeated integral. In a similar way, one
could begin with mass of vertical strips
h(x) =/integraldisplayγ4
γ3f(x,y)dy
and add these up
/integraldisplay/integraldisplay
Df(x,y)dxdy =/integraldisplayγ2
γ1h(x)dy=/integraldisplayγ2
γ1/parenleftbigg/integraldisplayγ4
γ3f(x,y)dy/parenrightbigg
dx.
For most purposes, it is sufficient to consider domains which are of the two types
pictured
a figure goes here
that is,Dis bounded on two sides by straight line segments. More complicated domains
can be treated by decomposing them into domains of these two types, where one or both
of the straight line segments might degenerate to a point.
Theorem 8.21 . Iffis continuous on a domain D1(respectively D2) as above, then
the iterated integral
/integraldisplayb
a/parenleftBigg/integraldisplayφ2(x)
φ1(x)f(x,y)dy/parenrightBigg
dx [resp./integraldisplayβ
α/parenleftBigg/integraldisplayφ2(y)
φ1(y)f(x,y)dx/parenrightBigg
dy]
exists and equals /integraldisplay/integraldisplay
DfdA.
Proof not given. It is rather technical.
Remark: If a domain Dhappens to be of both types (as, for example, rectangles and
triangles are ) then either iterated integral can be used and yield the same result - since
they are both equal/integraldisplay/integraldisplay
DfdA . See Examples 1 and 3 below (Example 2 could also have
been done both ways).
Examples:
354CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
(1) Evaluate/integraldisplay/integraldisplay
DfdA wheref(x,y) =x2yandDis the rectangle in the figure. We
shall integrate with respect to xfirst.
/integraldisplay/integraldisplay
DfdA =/integraldisplay2
1/parenleftbigg/integraldisplay3
1(x2+xy)dx/parenrightbigg
dy.
The inner integral is the mass of a strip. Think of yas being the fixed height of the
strip. Then
/integraldisplay3
1(x2+xy)dx=x3
3+x2y
2/vextendsingle/vextendsingle/vextendsinglex=3
x=1= 9 +9y
2−1
3−y
2=26
3+ 4y
Therefore, adding up all the strips we find
/integraldisplay/integraldisplay
DfdA =/integraldisplay2
1(26
3+ 4y)dy= (26
3y+ 2y2)/vextendsingle/vextendsingle/vextendsingley=2
y=1=26
3+ 6 =44
3
Let us evaluate this again, now integrating first with respect to y.
/integraldisplay/integraldisplay
DfdA =/integraldisplay3
1/parenleftbigg/integraldisplay2
1(x2+xy)dy/parenrightbigg
dx.
First/integraldisplay2
1(x2+xy)dy= (x2y+xy2
2)/vextendsingle/vextendsingle/vextendsingley=2
y=1=x2+3
2x
so/integraldisplay/integraldisplay
DfdA =/integraldisplay3
1(x2+3
2x)dx= (x2
3+3
4x2)/vextendsingle/vextendsingle/vextendsinglex=3
x=1=44
3,
which agrees with the previous computation. Instead of imagining fas the density
ofD, one can also take fto be the height function of a surface above D. Then the
integral/integraldisplay/integraldisplay
DfdA is the volume of the solid whose base is Dand whose “top” is the
surfaceMwith points ( x,y,f (x,y)) . In this case, the volume is 44/3.
(2) Evaluate/integraldisplay/integraldisplay
DfdA wheref(x,y) =x2+xy+ 2 andDis the domain bounded by
the curves φ1(x) = 2x2, φ2(x) = 4 +x2, andx= 0 .
Integrate first with respect to y. Thenyvaries between 2 x2and 4 +x2, whilex
varies between the two straight lines x= 0 andx= 2 .
/integraldisplay/integraldisplay
DfdA =/integraldisplay2
0/parenleftBigg/integraldisplay4+x2
2x2(x2+xy+ 2)dy/parenrightBigg
dx
=/integraldisplay2
0(x2y+xy2
2+ 2y)/vextendsingle/vextendsingle/vextendsingley=4+x2
y=2x2dy
/integraldisplay2
0(8 + 8x+ 2x2+ 4x3−x4−3
2x5)dy=464
15
8.4. MULTIPLE INTEGRALS 355
(3) Evaluate/integraldisplay/integraldisplay
DfdA wheref(x,y) = (x−2y)2andDis the triangle bounded by
x= 1,y=−2 , andy+ 2x= 6.
We shall integrate first with respect to x. Thenxvaries between x= 1 and
x=−1
2y+ 2 , while yvaries between the lines y=−2 andy= 2 .
/integraldisplay/integraldisplay
DfdA =/integraldisplay2
−2/integraldisplay1
2y+2
1(x−2y)2dxdy
Since
/integraldisplay−1
2y+2
1(x−2y)2(x−2y)2dx=1
3(x−2y)3/vextendsingle/vextendsingle/vextendsinglex=−1
2y+2
x=1=1
3(2−5
2y)3−1
3(1−2y)3,
we find /integraldisplay/integraldisplay
DfdA =1
3/integraldisplay2
−2[(2−5
2y)3−(1−2y)3]dy=164
3.
One can also integrate first with respect to y. Thenyvaries between y=−2 and
y=−2x+ 6 , while xvaries between the lines x= 1 andx= 3 .
/integraldisplay/integraldisplay
DfdA =/integraldisplay3
1/parenleftbigg/integraldisplay−2x+4
−2(x−2y)2dy/parenrightbigg
dx.
Since
/integraldisplay−2x+4
−2(x−2y)2dy=−1
6(x−2y)3/vextendsingle/vextendsingle/vextendsingley=−2x+4
y=−2=−1
6[(5x−8)3−(x+ 4)3]
we again find
/integraldisplay/integraldisplay
DfdA =−1
6/integraldisplay3
1[(5x−8)3−(x+ 4)3]dx=164
3.
(4) Find the volume of the pyramid Pbounded by the four planes x= 0,y= 0,z= 0,
andx+y+z= 1 . The easiest way to do this is to let z=f(x,y) = 1−x−ybe
the height function of the tilted plane which we shall take as the top of the pyramid
which lies above the triangle D(in thexyplane) which is bounded by the three
linesx= 0,y= 0 , andx+y= 1 . Then
Volume (P) =/integraldisplay/integraldisplay
Df(x,y)dxdy
One can integrate with respect to either xoryfirst. We shall do the xintegration
first.
/integraldisplay/integraldisplay
DfdA =/integraldisplay1
0/parenleftbigg/integraldisplay1−y
0(1−x−y)dx/parenrightbigg
dy.
Since /integraldisplay1−y
0(1−x−y)dx=−1
2(1−x−y)2/vextendsingle/vextendsingle/vextendsinglex=1−y
x=0=1
2(1−y)2
356CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
we find
Volume (P) =/integraldisplay/integraldisplay
DfdA =1
2/integraldisplay1
0(1−y)2dy=−1
6(1−y)3/vextendsingle/vextendsingle/vextendsingle1
0=1
6.
This agrees with the usual formula for the volume of a pyramid
Vol =1
3altitude ×area of base .
The identical methods work for triple integrals. All of the theorems and proofs remain
unchanged. Again the integral /integraldisplay/integraldisplay/integraldisplay
DfdV
can either be interpreted as the mass of a solid Dwith density f, or as the “vol-
ume” of a four dimensional solid whose base is Dand top in the surface with points
(x,y,z,f (x,y,z )) . Because of conceptual difficulties, one usually thinks of fas a density.
Calculation of triple integrals is done by evaluating three integrals, as
/integraldisplay/integraldisplay/integraldisplay
DfdV =/integraldisplay/parenleftbigg/integraldisplay/parenleftbigg/integraldisplay
f(x,y,z )dz/parenrightbigg
dy/parenrightbigg
dx,
where the limits in the iterated integral on the right are determined from the domain D.
An example should illustrate the idea adequately,
Example: Evaluate/integraldisplay/integraldisplay/integraldisplay
DfdV wheref(x,y,z )≡candDis the solid bounded by the
two planes z≡0, y≡2 , and the surface z≡ −x2+y2. We have to evaluate/integraldisplay/integraldisplay/integraldisplay
DcdV
which is the mass of the solid Dwith constant density c, that isctimes the volume of
D. It is convenient to carry out the zintegration first, then the xintegration
/integraldisplay/integraldisplay/integraldisplay
DcdV =c/integraldisplay2
0/parenleftBigg/integraldisplayy
−y/parenleftBigg/integraldisplay−x2+y2
0dz/parenrightBigg
dx/parenrightBigg
dy.
Thexlimits of integration have been found by looking at the region of integration in the
xyplane beneath the surface z=−x2+y2. This region, found by setting z= 0 , consists
of the points between the straight lines 0 = −x2+y2, that is between the lines x=yand
x=−y. Then/integraldisplay/integraldisplay/integraldisplay
DfdV =c/integraldisplay2
0/parenleftbigg/integraldisplayy
−y(−x2+y2)dx/parenrightbigg
dy
=c/integraldisplay2
0(−x3
3+xy2)/vextendsingle/vextendsingle/vextendsinglex=y
x=−ydy=c/integraldisplay2
04
3dy=16
3c.
By letting c= 1 , the volume of the solid is seen to be 16 /3 .
Exercises
(1) Evaluate/integraldisplay/integraldisplay
Dxydxdy for the following domains Din two ways:/integraldisplay
(/integraldisplay
xydx )dy
and/integraldisplay
(/integraldisplay
xydy )dx.
8.4. MULTIPLE INTEGRALS 357
(a)Dis the rectangle with vertices at (1 ,1),(1,5),(3,1) and (3,5) .
(b)Dis the triangle with vertices at (1 ,1),(3,1) and (3,5) .
(c)Dis the region enclosed by the lines x= 1, y= 2 , and the curve y=x3(a
curvilinear triangle).
(d)Dis the region enclosed by the curves y=x2andy=√x.
(2) Evaluate/integraldisplay/integraldisplay
Dsinπ(2x+y)dxdy,
whereDis the triangle bounded by the lines x= 1, y= 2 andx−y= 5 .
(3) Evaluate/integraldisplay/integraldisplay
D(xy−y3)dxdy,
whereDis the region enclosed by the lines x=−1, x= 1, y=−2 and the curve
y= 2−x2.
(4) Evaluate/integraldisplay
D(xy+z)dxdydz,
whereDis the rectangular parallelepiped bounded by the six planes x=−2, y=
1, z= 0, x= 1, y= 2, z= 3 .
(5) Evaluate/integraldisplay/integraldisplay/integraldisplay
Dxyzdxdydz,
whereDis the solid enclosed by the paraboloid z=x2+y2and the plane z= 4 .
(6) Find the volume of an octant of the ball x2+y2+z2≤a2in two ways;
(a) by evaluating/integraldisplay/integraldisplay
Df(x,y)dxdy
wherefis a suitable function and Da suitable domain
(b) by evaluating/integraldisplay/integraldisplay/integraldisplay
Ddxdydz,
whereDis the ball.
(7) Iff(x,y)>0 is the density function of a plate D, thexandycoordinates of the
center of mass (¯x,¯y) are defined by
x=/integraltext/integraltext
Dxf(x,y)dxdy/integraltext/integraltext
Df(x,y)dxdy,y=/integraltext/integraltext
Dyf(x,y)dxdy/integraltext/integraltext
Df(x,y)dxdy.
Find the center of mass of a triangle whose vertices are at the points (0 ,0),(0,4) ,
and (2,0) , and whose density is f(x,y) =xy+ 1 .
358CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
(8) The moment of inertia with respect to a point p= (ξ,η) of a plate Dwith density
f(x,y) is defined by
Jp(D) =/integraldisplay/integraldisplay
D[(x−ξ)2+ (y−η)2]f(x,y)dxdy.
(a) Find the moment of inertia of the plate in Exercise 7, with respect to the point
p= (1,0) .
(b) IfDis any plate (with sufficiently smooth boundary), prove that the moment
of inertia is smallest if the point f= (ξ,η) is taken to be the center of mass of
D. [Hint: Consider Jas a function of the two variables ξandηand showJ
has a minimum at (¯ x,¯y) .]
(9) (a) Show that
/integraldisplay/integraldisplay
Dfxy(x,y)dxdy =f(p1)−f(p2) +f(p3)−f(p4),
whereDis a rectangle with vertices at p1,p2,p3,p4(see fig.).
(b) Use the result of part (a) to again evaluate the integral in Ex. 1a.
(c) IfU(x,y) satisfies the partial differential equation Uxy= 0 for 0<y <x and
U(x,x) = 0 while U(x,0) =xsinx, findU(x,y) for all points ( x,y) in the
wedge 0<y<x . [Answer: U(x,y) =xsinx−ysinyfor 0<y<x ].
(10) Letf(x,y) be a bounded function which is continuous except as a set of points
of content zero, and suppose fhas compact support. Prove that fis Riemann
integrable. This again proves Theorem 13.
(11) LetD1andD2be domains whose boundaries have zero content and whose intersec-
tionD1∩D2has zero content.
(a) Iffis continuous on D1∪D2, prove that the integral/integraldisplay/integraldisplay
D1∪D2fdA exists and
that /integraldisplay/integraldisplay
D1∪D2fdA =/integraldisplay/integraldisplay
D1fdA +/integraldisplay/integraldisplay
D2fdA.
(b) Give an example showing the above equality does not hold if D1∩D2has non-
zero content.
(12) (a) By an explicit construction, show that the region D={(x,y)/epsilon1E2:|x|+|y| ≤1}
has boundary with zero content.
(b) By an explicit construction, show that the circle ? = {(x,y)/epsilon1E2:x2+y2= 1}
has zero content.
(13) (a) By interchanging the order of integration, show that
/integraldisplayx
0(/integraldisplays
0f(t)dt)ds=/integraldisplayx
0(x−t)f(t)dt.
(b)/integraldisplayx
0(/integraldisplay2
0(/integraldisplayr
0f(t)dt)dr)ds=?
8.4. MULTIPLE INTEGRALS 359
(14) LetDbe a plate in the x,yplane with density fand total mass M. Ifp= (ξ,η)
is an arbitrary point in the plane and ¯ p= (¯x,¯y) is the center of mass of D, prove
Jp(D) =J¯p(D) +M/bardblp−p0/bardbl2,
where the notation of Exercise 8 has been used. This is the parallel axis theorem . It
again proves the result of Exercise 8b.
360CHAPTER 8. MAPPINGS FROM ENTOE: THE DIFFERENTIAL CALCULUS
Chapter 9
Differential Calculus of Maps from
Ento Em, s.
9.1 The Derivative
.
Now we generalize the ideas of Chapters 7 and 8 and consider nonlinear mappings from
a setDinEntoEm, F:D⊂En→Em, orY=F(X) , whereX∈DandY∈Em. In
coordinates, these functions look like
y1=f1(x1,...,x n)
·
·
·
ym=fm(x1,...,x n)
where the functions fjare scalar-valued. The special case n= 1, marbitrary, was treated
in Chapter 7, section 3, while the special case m= 1, narbitrary, was treated in Chapter
8.
One interpretation of maps F:D⊂En→Emis as a geometric transformation from
some subset DofEninto all or part of Em.
EXAMPLES.
(1) The affine map Y=F(X) defined by
y1= 2 +x1−2x2
y2= 1 +x1+x2
maps E2intoE2. Under this map, the origin goes into (2 ,1) , thex1axis (i.e. the
linex2= 0 ) goes into the line y1−y2= 1 ,
a figure goes here
361
362 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S.
while thex2axis goes into the line y1+ 2y2= 4 . The shaded region indicates the
image of the indicated square.
(2) The map Y=F(X) defined by
y1=x1−x2
y2=x2
1+x2
2
maps all of E2onto the upper half y1y2plane (since y2≥0 ). Let us see what
happens to a rectangle under this mapping. Consider the rectangle Rin the figure.
Thex1axis,x2= 0 , goes into the parabola y2=y2
1, and the line x2= 1 into
y2= 1 + (y1+ 1)2.
a figure goes here
Similarly, the line x1= 1 is mapped into y2= 1+(y1−1)2, whilex1= 2 is mapped
intoy2= 4 + (y1−2)2. By following the images of the boundary ∂R, we now see
that the interior of Ris mapped into the shaded curvilinear “parallelogram”. This
mapping, though injective when restricted to our rectangle, is not injective for all
(x1,x2)∈E2, since, for example, the points X1= (1,2) andX2= (−2,−1) are
both mapped into the same point ( −1,5) .
(3) The function w=x2
1+x2
2whose graph is a paraboloid, is a map from E2into
E1. It can also be regarded as a map from E2intoE3by a useful artifice. Let
y1=x1, y2=x2, andy3=w=x2
1+x2
2. Then
y1=x1
y2=x2
y3=x2
1+x2
2
is a mapFfrom E2intoE3. The image of the unit square (see figure) is then the
shaded region in the figure above the image ( y1,y2) of the square R
a figure goes here
(4) The map F:E2→E3defined by (cf. example 2)
y1=x1−x2
y2=x2
1+x2
2
y3=x1+x2
also represents a surface M. In fact, since y2
1+y2
3= 2y2, this surface is a paraboloid
opening out on the y2axis. Again, we investigate where the rectangle Rof example 2
is mapped. Since the y1andy2components of the mapping are the same as before, the
image ofRwill lie on the surface Mabove the image ( y1,y2) of (x1,x2) . Thus the image
of the rectangle Ris a patch of the surface M.
9.1. THE DERIVATIVE 363
From these examples, we see it is natural to regard any map F:D⊂E2→Em
as an ordinary surface, or two dimensional manifold, embedded in Em, much as a map
F:D⊂E1→Emwas regarded as an ordinary curve. In the case m= 1 , the surface
F:D⊂E2→E1was representable as the graph of the function F. Form= 2 and
higher, this surface is seen as the range of the map. In the same way, an ndimensional
surface, or manifold, embedded in IEmis a mapF:D⊂En→Em. You might want to
think ofnas being the number of “degrees of freedom” on the manifold. In a strict sense,
the mapF:D⊂En→Emis not annmanifold embedded inEmunless Emis big enough
to hold an mmanifold, i.e. m≥n. However by either using the graph of F, a subset of
Em+n, or by using the trick of example 3 we can always think of the map F:En→Emas
anndimensional surface. For m≥n, this surface can be embedded as a subset of En.
There are several valuable physical interpretations of these vector valued functions of a
vector,Y=F(X) . Consider a fluid flowing through a domain DinE3. The fluid could
be air and Das the outside of an airplane, or the fluid could be an organic fluid, and D
as some portion of the body.
The velocity Vof a particle of fluid is a three vector which depends upon the space co-
ordinate (x1,x2,x3) as well as the time coordinate tof the particle, V=F(x1,x2,x3,t) =
F(X,t) . This velocity vector V(X,t) atXpoints in the direction the fluid is moving.
Thus, the velocity function is an example of a mapping from space-time E3×E1∼=E4into
vectors in E3. In this case, we think of the velocity vector V=F(X,t) as having its foot
at the point X∈Dand imagine the mapping as the domain Dalong with a vector V
attached to each point of D(see fig. above). One calls this a vector field defined on the
domainD, since it assigns a vector to each point of D.
A very common vector field is a field of forces. By this we mean that to every point
Xof a domain D, we associate a vector F(X) equal to the force an object at X“feels”.
If the forces are time dependent, then the force field is written F(X,t), X∈D. You are
most familiar with the force field due to gravity. If e3is the direction toward the center
of the earth, and say e1points east and e2north along the surface of the earth (other
coordinates must be chosen for the north and south poles), then the gravitational force
is usually written as F= (0,0,g) , a constant vector pointing down to the center of the
earth. For more precise purposes, one must take into account the fact that gdoes vary
from place to place of the earth’s surface. Then F(x) = (0,0,g(X)) . In even more accurate
experiments - or in outer space - must further account for the effect of the other heavenly
bodies. This brings in the other components of force as well as a time dependence due to the
motion of the earth, F(X,t) = (f1(X,t), f2(X,t), f3(X,t)) . The force field is imagined as
a vector attached to each point Xin space, the vector having the magnitude and direction
of the net force Fthere.
An entirely different example of a mapping Ffrom EntoEmis a factory - or an even
larger economic system. The vector X= (x1,x2,...,x n) might represent the quantities
x1,x2,... of different raw materials needed. Y=F(X) could then represent the output
from the factory, the number yjbeing the quantity of the jth product produced from the
inputX.
Turning to the quantitative mathematical aspect of the mappings F:En→Em, we
define the derivative. The definition will be formal, patterned directly on the definition of
the total derivative given previously (p. 578-9).
Definition: LetF:D⊂En→EmandX0be an interior point of D.Fisdifferentiable
364 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S.
atX0, if there exists a linear transformation L(X0):En→Em, depending on the base
pointX0, such that
lim
/bardblh/bardbl→0/bardblF(X0+h)−F(X0)−L(X0)h/bardbl
/bardblh/bardbl= 0
for any vector hin some sufficiently small ball about X0. IfFis differentiable at X0,
we shall use the notationsdF
dX(X0) =F/prime(X0) =L(X0)
and refer to them as the derivative of FatX0. IfF/prime(X0) depends continuously on the
base point X0for allX0inD, thenFis said to be continuously differentiable inD,
writtenF∈C1(D) .
Many of the results from Chapter 8 Sections 1 and 2 generalize immediately to the
present situation.
Proposition 9.1 . The function F:D⊂En→Emis differentiable at the interior point
X0∈Dif and only if there is a linear operator L(X0):En→Emand a function R(X0,h)
such that
F(X0+h) =F(X0) +L(X0)h+R(X0,h)/bardblh/bardbl,
where the remainder R(X0,h)has the property
lim
/bardblh/bardbl→0/bardblR(X0,h)/bardbl= 0.
Proof: ⇐IfFis differentiable at X0, letL(X0)be the derivative and take R(X0,h) =
[F(X0+h)−F(X0)−L(X0)h]//bardblh/bardbl. Then this L(X0)andR(X0,h) do satisfy the above
conditions.
⇒IfL(X0)andR(X0,h) are as above, then
lim
/bardblh/bardbl→0/bardblF(X0+h)−F(X0)−L(X0)h/bardbl
/bardblh/bardbl= lim
/bardblh/bardbl→0/bardblR(X0,h)/bardbl= 0.
SinceL(X0)is linear, this proves Fis differentiable at X0.
There is at most one derivative operator L(X0), that is
Proposition 9.2 . (Uniqueness of the derivative). Let F:D⊂En→Embe differentiable
at the interior point X0∈D. IfˆL(X0)and ˜L(X0)are linear operators both of which satisfy
the conditions for the derivative of FandX0, then ˆL(X0)=˜L(X0).
Proof: Word for word the same as the proof of Theorem 1, page 579-80.
If the map F=F(X) is given in terms of coordinates,
y1=f1(x1,...,x n)
y2=f2(x1,...,x n)
· ·
· ·
· ·
ym=fm(x1,...,x n),
how is the derivative computed, and what is its relationship to the derivative of the indi-
vidual coordinate functions fj? The answer is contained in
9.1. THE DERIVATIVE 365
Theorem 9.3 . LetFmapD⊂EnintoEmbe given in terms of the coordinate functions
fj(X), j = 1,...,m
y1=f1(X)f1(x1,...,x n)
·
·
·
ym=fm(X) =fm(x1,...,x m).
(a) ThenFis differentiable or continuously differentiable at the interior point X0∈D
if and only if all of the fj’s are respectively differentiable or continuously differentiable.
(b) Moreover, if Fis differentiable at X0, then the derivative in these coordinates is
given by the m×nmatrix of partial derivatives
L(X0):=F/prime(X0) =
f/prime
1(X0)
·
·
·
f/prime
m(X0)
=
∂f1
∂x1(X0),...,∂f1
∂xn(X0)
·
·
·
∂fm
∂x1(X0),...,∂fm
∂xn(X0)
.
The matrix is sometimes called the Jacobian matrix.
Proof: (a) Observe that the limit
lim
/bardblh/bardbl→0/bardblF(X0+h)−F(X0)−L(X0)h/bardbl
/bardblh/bardbl= 0
exists if and only if each of its components tend to zero,
lim
/bardblh/bardbl→0/bardblfjX0+h)−fj(X0)−Lj(X0)h/bardbl
/bardblh/bardbl= 0, j = 1,2,...,m.
Thus, ifFis differentiable at X0, each of the coordinate functions fjare differentiable
and have total derivative Lj(X0). Conversely, if each of the coordinate functions are dif-
ferentiable at X0, all of the above limits exist so the vector valued function Fis also
differentiable.
(b) Since the differentiability of Fimplies that of the coordinate vectors, we have
F/prime(X0) =
f/prime
1(X0)
·
·
·
f/prime
m(X0)
.
The result now follows by writing out each of the derivatives
f/prime
1(X0) = (∂f1(X0)
∂x1,...,∂f1(X0)
∂xn)
f/prime
2(X0) =...etc. and then inserting these in the expression for F/prime(X0) .
Corollary 9.4 . A function F:D⊂En→Emis continuously differentiable in Dif and
only if all the partial derivatives of its components ∂fi/∂x jexist and are continuous.
366 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S.
Proof: This follows from this theorem and Theorem 3, p. 585.
EXAMPLES.
1. LetFbe an affine map from EntoEm
F(X) =Y0+BX,
whereBis a linear operator from EntoEm(which you may choose to think of as an
m×nmatrix with respect to some coordinate system) and Y0=F(0) is a fixed vector in
Em. ThenFis differentiable at every point of Enand it given by the eminently reasonable
formula
F/prime(X0) =B,
where the operator Bdoes not depend on X0. For proof, we observe that
F(X0+h)−F(X0) =Y0+B(X0+h)−[Y0+BX 0] =Bh.
Thus
lim
/bardblh/bardbl→0/bardblF(X0+h)−F(X0)−Bh/bardbl
/bardblh/bardbl= lim
/bardblh/bardbl→00
/bardblh/bardbl= 0.
SinceBis linear, this shows the derivatives exists and is B. Let us do this again in
coordinates. If B= ((bij)) the function Fis
f1(X) =y01+b11x1+b12x2+...+b1nxn
f2(X) =y02+b21x1+... +b2nxn
·
·
·
fm(X) =y0m+bm1x1+... +bmnxn.
Therefore each of the functions fjis clearly differentiable and
f/prime
1= (∂f1
∂x1,...,∂f1
∂xn) = (b11,...,b 1n)
· · ·
· · ·
· · ·
f/prime
m= (∂fm
∂x1,...,∂fm
∂xn) = (bm1,...,b mn).
Consequently,
F/prime(X0) =
f/prime
1(X0)
·
·
·
f/prime
m(X0)
=
b11, ..., b 1m
·
·
·
bm1, ..., b mn
=B,
which agrees with the result obtained without coordinates.
2. LetF:E2→E3be defined by
f1(x1,x2) = 2−x1+x2
2
9.1. THE DERIVATIVE 367
f2(x1,x2) =x1x2−x3
2
f3(x1,x2) =x2
1−3x1x2.
Since each of the coordinate functions fjare continuously differentiable, so is F. Because
f/prime
1(X) = (−1,2x2), f/prime
2(X)−(x2,x1−3x2
2), f/prime
3(X) = (2x1−3x2,−3x1),
we find that at X0= (3,1)
F/prime(X0) =
f/prime
1(X0)
f/prime
2(X0)
f/prime
3(X0)
=
−1 2
1 0
3−9
.
IfXis nearX0, then by Proposition 1 with h=X−X0
F(X) =F(X0) +f/prime(X0)(X−X0) + remainder
=
0
2
3
+
−1 2
1 0
3−9
/parenleftbiggx1−3
x2−1/parenrightbigg
+ remainder ,
where the remainder term becomes less significant the closer Xis toX0.
Motivated by our previous work, it is natural to formally define the tangent map as
follows.
Definition: LetF:D⊂En→Embe differentiable at the interior point X0∈D. The
tangent map atF(X0) to the (hyper) surface defined by Fis defined to be the affine
mapping
Φ(X) =F(X0) +f/prime(X0)(X−X0).
Examples:
(1) LetFbe the function of Example 2 above. Then the tangent map at X0= (3,1) is
Φ(X) =
0
2
3
+
−1 2
1 0
3−9
/parenleftbiggx1−3
x2−1/parenrightbigg
.
(2) LetFbe the function of Example 4 (page 679). Then
F/prime(X) =
1−1
2x12x2
1 1
.
Thus the tangent map at (2 ,1) is
Φ(X) =
1
5
3
+
1−1
4 2
1 1
/parenleftbiggx1−2
x2−1/parenrightbigg
If we letY= Φ(X) , then the target plane in the tangent space is found from
y1= 1 + (x1−2)−(x2−1)
y2= 5 + 4(x1−2) + 2(x2−1)
y3= 3 + (x1−2) + (x2−1)
By eliminating x1andx2from these equations, we find y2=−5 +y1+ 3y3. A graph of
the surface Mand the tangent plane can now be drawn.
368 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S.
a figure goes here
The next result is the generalization of the mean value theorem.
Theorem 9.5 . (Mean Value Theorem). Let F:D⊂En→Embe differentiable at every
point ofD, whereDis an open convex set in En. IfF/prime(X)is bounded in D, that is, if
there is a constant γ <∞such that/vextendsingle/vextendsingle/vextendsingle∂fi
∂xj(X)/vextendsingle/vextendsingle/vextendsingle≤γfor allX∈Dand for all i= 1,...,m ,
andj= 1,...,n , then
/bardblF(X2)−F(X1)/bardbl ≤c/bardblX2−X1/bardbl
for allX1andX2inD, whereC=√nmγ .
Proof: The idea is to use the components of Fand to appeal to the similar theorem (p.
597-8) for the function from En→E1. By that theorem, if X1andX2are inD, then
there is a point Z1on the line segment joining X1toX2such that
f1(X2) =f1(X1) +f/prime
1(Z1)(X2−X1),
and similarly for the other components f2,f3,...,f m. Thus
f1(X2)
·
·
·
fm(X2)
=
f1(X1)
·
·
·
fm(X1)
=
f/prime
1(Z1)
·
·
·
f/prime
m(Zm)
(X2−X1),
whereZ1,...,Z mare all on the segment joining
a figure goes here
X1toX2. Observe that the f/prime
j(Zj) ’s are all vectors. Let Lbe the matrix of derivatives
in the last term above, that is
L=
f/prime
1(Z1)
·
·
·
f/prime
m(Zm)
=
∂f1
∂x2(Z1)···∂f1
∂xn(Z1)
·
·
·
∂fm
∂x1(Zm)···∂fm
∂xn(Zm)
.
The above equation then reads
F(X2) =F(X1) +L(X2−X1). (9-1)
This equation itself is sometimes referred to as the mean value theorem. Note, however,
that the partial derivatives in Larenotall evaluated at the same point.
Since/vextendsingle/vextendsingle/vextendsingle∂fi
∂xj(X)/vextendsingle/vextendsingle/vextendsingle≤γfor allX, ifηis any vector in En, by Theorem 17, p. 373. we
find that
/bardblLη/bardbl ≤√nmγ/bardblη/bardbl.
Takingη=X2−X1, and using (1), we are led to the inequality
/bardblF(X2)−F(X1)/bardbl ≤√nmγ/bardblX2−X1/bardbl,
9.1. THE DERIVATIVE 369
which holds for any points X1andX2inD. WithC=√nmγ , this is the desired
inequality.
A few heuristic remarks. We have been considering mappings F:En→Em. In the
case of linear mappings, L:En→Em, it was possible to prove that the range of Lhad
dimension no greater than n, dim R(L)≤n. Although this does not remain true for an
arbitrary nonlinear map F, it is still true if Fis differentiable - after a suitable definition of
dimension for an arbitrary point set is made (for the range of Fwill not usually be a linear
space, the only sets whose dimension we have so far defined). In the case of differentiable
mapsF, it is easy to make a reasonable definition of dimension. The idea is to define
dimension of the range of Flocally, that is, in the neighborhood of every point in the
range. IfF:D⊂En→EmandFis differentiable at X∈D, then for all hsufficiently
small,
F(X+h) =F(X) +L(X)h+ remainder .
Thedimension of the range ofFatF(X) is defined to be the dimension of its affine part,
which is the same as dim bR(L(X)) . SinceL(X)is a linear operator, its range has a well
defined dimension. Geometrically, we have defined dimension of the range of FatF(X)
as the dimension of the tangent plane at F(X) . Our definition makes good physical sense
for it is exactly the number an insect on the surface would use for the dimension. The
illustration below is for a map F:DE2→E3whose range has dimension 2,
a figure goes here
Some special remarks should be made about maps from one space into another of the
same dimension,
F:DEn→En.
Let us assume Fis differentiable throughout D. Then the dimension of the range of
FatF(X), X∈D, is the dimension of the range of L(X)=F/prime(X) . IfFis to preserve
dimension at every point, then we must have dim R(L(X)) =nfor allX∈D. For maps
Fgiven in terms of coordinates, this means the determinant of the n×nmatrixL(X)
does not vanish,
detL(X)= detF/prime(X)/negationslash= 0
for allx∈D. In more conceptual terms, this states that a map F:D⊂En→Enis
dimension preserving at X0∈Dif its “affine part” Φ( X0+h) =F(X0) +F/prime(X0)his
dimension preserving at X0(there is no trouble with the constant vector F(X0) since it
only represents a translation of the origin - which does not affect dimensionality).
From the geometric interpretation of determinants as volume, we see that the condition
detF/prime(X0)/negationslash= 0 means that if a small set S⊂Dhas non-zero volume, then its image F(X)
also has non-zero volume. In fact, we expect that if Sis a small set about X, then
Vol (F(S)) =/vextendsingle/vextendsingledetF/prime(X0)/vextendsingle/vextendsingleVol (S).
Our expectation is based upon the realization that if the points of Sare all near X0, then
Fwill behave like its affine part, ( X0+h) =F(X0) +F/prime(X0)h, on the points X0+h∈S.
The above formula is a restatement of the effect of affine maps on volume (Corollary to
Theorem 30, page 426). We shall return to this later (Chapter 10, Section 4).
370 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S.
Because of its frequent appearance, det F/prime(X) has a name of its own. It is called the
Jacobian determinant or just the Jacobian ofF. IfFis given in terms of coordinates,
y1=f1(x1,...,x n)
·
·
·
yn=fn(x1,...,x n),
then another common notation for the Jacobian is
detF/prime(X) =∂(f1,f2,...,f n)
∂(x1,x2,...,x n).
For these maps Ffrom a space into one of the same dimension, F:D⊂En→En,
there is a very special derivative which appears often. It is the sum of the diagonal elements
of the derivative matrix F/prime(X) . One writes this expression as ∇.For÷F, the divergence
ofF,
∇ ·F(X) = divF(X) =∂f1(X)
∂x1+∂f2(X)
∂x2+···+∂fn(X)
∂xn
For example, if Y=F(X) is defined by
y1=x1+ 2x1x2
y2=x2
1−3x2,
then
F/prime(X) =/parenleftbigg1 + 2x22x1
2x1 −3/parenrightbigg
and
∇ ·F(X) = divF(X) = (1 + 2x2) + (−3) =−2 + 2x2.
The significance of the divergence will become clear later (Chapter 10, Section 2). You will
probably find it helpful to think of ∇as the operator
∇= (∂
∂x1,···,∂
∂xn).
Then ∇ ·Fis the “scalar product” of the operator ∇with the vector F.
EXERCISES.
(1) (a) Find the derivative matrix at the given point for the following mappings Y=
F(X) .
(i)y1=x2
1+ sinx1x2
y2=x2
2+ cosx1x2atX0= (0,0)
(ii)y1=x2
1+x3ex2−x3
2
y2=x1−3x2+x1logx3
y3=x2+x3
y4= 5x1x2x3atX0= (2,0,1)
(b) Find the equation of the tangent plane to the above surfaces at the given point.
9.1. THE DERIVATIVE 371
(2) Consider the following map from E2→E2,
/braceleftbiggu=excosy
v=exsiny
(a) Find the image of the following regions
i)x≥0,0≤y≤π
4
ii)x≥0,0≤y≤π
iii)x≤0,0≤y≤2π
iv) 1<x< 2,π
6≤y≤π
3.
(b) Compute the derivative matrix and its determinant.
(3) IfF:D⊂En→Emis differentiable at X0∈D, prove it is then also continuous at
X0.
(4) LetFandGboth mapD⊂En→Em, so the function f(X) =/angbracketleftF(X), G(X)/angbracketrightis
defined for all X∈Dandf:D→E1.
(a) IfFandGare differentiable in D, provefis also, and that
f/prime=F/primeG+G/primeF
(b) Apply this result to the function
f(X) =/angbracketleftX, AX /angbracketright −2/angbracketleftX, Y/angbracketright,
whereAis a constant linear operator from En→EnandYis a constant vector
inEn. How does the result simplify if Ais self adjoint?
(5) Ifϕ:D⊂En→E1andF:D⊂En→Em, then the function G(X) :=ϕ(X)F(X)
is defined for all x∈DandG:En→Em.
(a) Letϕ(x2,x2) =ax1+bx2andF(x1,x2) = (αx1+βx2,γx 1+δx2) . LetG=ϕF
and compute G/prime(X) .
(b) More generally, prove that if ϕandFare differentiable in D, thenG:=ϕF
is also differentiable and find a formula for G/prime. IfFis expressed in terms of
coordinate functions, F= (f1,f2,...,f m) , how does your formula read? Check
the result with that of part (a).
(6) (a) If F:D⊂En→Emis differentiable in the open connected set D, and if
F/prime(X)≡0 for allx∈D, prove that Fis a constant vector.
(b) IfFandGmapD⊂En→Emare differentiable in the open connected set
D, and ifF/prime(X)≡G/prime(X) for allx∈D, what can you conclude?
(7) Consider the map F:Q→R3defined by
F:x= (a+bcosϕ) cosθ
y= (a+bcosϕ) sinθ
z=bsinϕ
372 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S.
a figure goes here
(a) Compute F/prime.
(b) Find the equation of the tangent map at (0 ,0) and at ( π/2,π/2) .
(c) Determine the range of the tangent map at the above two points and indicate
your findings in a sketch.
9.2. THE DERIVATIVE OF COMPOSITE MAPS (“THE CHAIN RULE”). 373
9.2 The Derivative of Composite Maps (“The Chain Rule”).
Consider the two mappings
F:A⊂En→EmandG:B⊂Em→Er.
Then the composite map H:=G◦F:A⊂En→Eris defined if Bcontains the image of
all the points from A, F (A)⊂B.
a figure goes here
The mapH=G◦Ftakes points from A⊂Enand sends them into Er. From
knowledge of the derivatives of FandG, it is possible to compute the derivative of the
composite map G◦F.
Theorem 9.6 . LetF:A⊂En→EmandG:B⊂Em→Erbe differentiable maps
defined in the open sets AandB, respectively, with F(A)⊂B(so the composite map
H(X) := (G◦F)(X)is defined for all X∈A). IfX0∈A, letY0=F(X0)∈B. Then
the composite map His differentiable at X0and
H/prime(X0) =G/prime(Y0)◦F/prime(X0).
Remark: The multiplication G/prime◦F/primeis the multiplication of the linear operators G/primeand
F/prime. IfFandGare given in terms of coordinates, then the formula is just the product of
two matrices G/primeandF/prime.
Before proving this theorem, we shall illustrate its meaning.
Example: LetF:E2→E2andG:E2→E3be defined by Y=F(X) andZ=G(Y)
as follows/braceleftbiggy1=x1−x2
2
y2=x2sinπx1
z1=y1y2
z2= 1 +y2
1+y2
z3= 5−y3
2.
Then
F/prime(X) =/parenleftbigg1 −2x2
πx2cosπx1sinπx1/parenrightbigg
, G/prime(X) =
y2y1
2y1 1
0−3y2
2
.
AtX0= (3,2) , we find Y0=F(X0) = (−1,0) . Thus
F/prime(X0) =/parenleftbigg1−4
−2π 0/parenrightbigg
, G/prime(Y0) =
0−1
−2 1
0 0
.
IfH(X) = (G◦F)(X) =G(F(X)) , then the derivative of HatX0is
H/prime(X0) =G/prime(Y0)◦F/prime(X0) =
0−1
−2 1
0 0
/parenleftbigg1−4
−2π 0/parenrightbigg
=
2π 0
−2−2π8
0 0
.
374 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S.
The derivative could also have been found in a longer way by explicitly finding Z=H(X)
from the formulas for FandG
z1=y1y2= (x1−x2
2)(x2sinπx1)
z2= 1 +y2
1+y2= 1 + (x1−x2
2)2+x2sinπx1
z3= 5−y3
2= 5−(x2sinπx1)3
and now directly computing H/prime(X0) .
Proof of Theorem . SinceFis differentiable at X0∈A⊂EnandGis differentiable
atY0∈B⊂Er, for all sufficiently small vectors h∈Enandk∈Em, we can write
F(X0+h) =F(X0) +F/prime(X0)h+R1(X0,h)/bardblh/bardbl
G(Y0+k) =G(Y0) +G/prime(Y0)k+R2(Y0,k)/bardblk/bardbl
where
lim
/bardblh/bardbl→0/bardblR1(X0;h)/bardbl= 0 and lim
/bardblk/bardbl→0/bardblR2(Y0,k)/bardbl= 0.
Consequently, since H(X) := (G◦F)(X) =G(F(X)) ,
H(X0+h) =G(F(X0+h))
=G(F(X0) +F/prime(X0)h+R1(X0;h)/bardblh/bardbl
=G(F(X0)) +G/prime(Y0)F/prime(X0)h+R3(X0,h)/bardblh/bardbl,
where
R3(X0;h) =G/prime(Y0)R1(X0;h) +R2(Y0,k)/bardblk/bardbl
/bardblh/bardbl,
and
k=F/prime(X0)h+R1(X0;h)/bardblh/bardbl.
Thus, for all sufficiently small h,
H(X0+h) =H(X0) +G/prime(Y0)F/prime(X0)h+R3(X/prime
0,h)/bardblh/bardbl.
BecauseG/prime(Y0) andF/prime(X0) are linear maps, so is their product. Therefore we are done if
we prove lim
/bardblh/bardbl→0/bardblR3(X0;h)/bardbl= 0 .
By the triangle inequality
/bardblR3(X0;h)/bardbl ≤ /bardblG/prime(Y0)R1(X0;h)/bardbl+/bardblR2(Y0,k)/bardbl/bardblk/bardbl
/bardblh/bardbl.
Since for fixed X0, the operators F/prime(X0) andG/prime(Y0) are constant operators, by Theorem
17, p. 373, there exist constants αandβsuch that for any vectors ξ∈Enandη∈Em,
/bardblF/prime(X0)ξ/bardbl ≤α/bardblξ/bardbland /bardblG/prime(Y0)η/bardbl ≤β/bardblη/bardbl.
This means
/bardblk/bardbl ≤ /bardblF/prime(X0)h/bardbl+/bardblR1(X0;h)/bardbl/bardblh/bardbl ≤(α+/bardblR1(X0;h)/bardbl)/bardblh/bardbl
9.2. THE DERIVATIVE OF COMPOSITE MAPS (“THE CHAIN RULE”). 375
and
/bardblG/prime(Y0)R1(X0;h)/bardbl ≤β/bardblR1(X0;h)/bardbl.
Thus,
/bardblR3(X0;h)/bardbl ≤β/bardblR1(X0;h)/bardbl+ (α+/bardblR1(X0;h)/bardbl)/bardblR2(Y0,k)/bardbl
Now, as /bardblh/bardbl → 0 , so does /bardblk/bardbl ≤(α+/bardblR1(X0;h)/bardbl)/bardblh/bardbl. From the definition of R1and
R2, this implies /bardblR3(X0;h)/bardbl →0 as/bardblh/bardbl →0 and completes the proof.
Incidentally, if one writes Y=F(X) andZ=G(Y) , then the chain rule can be
written in the form
d
dx(G◦F) =dG
dY◦dY
dX,
which could hardly be more simple to remember.
For the balance of this section, we shall work out a few more illustrations showing how
the chain rule is applied in different concrete situations. We isolate the next example as an
important
Corollary 9.7 . LetF:D⊂En→Emand the scalar valued function g:Em→E1
both satisfy the hypotheses of Theorem 1. If we write Y=F(X)in coordinates F=
(f1,f2,...,f m), and leth=g◦F, then
∂h
∂x1=∂g
∂y1∂f1
∂x1+∂g
∂y2∂f2
∂x1+···+∂g
∂ym∂fm
∂x1
·
·
·
∂h
∂xn=∂g
∂y1∂f1
∂xn+∂g
∂y2∂f2
∂xn+···+∂g
∂ym∂fm
∂xn
Remark: This is the chain rule for scalar-valued functions.
Proof: By Theorem 3,
dh
dX=dq
dYdF
dX
Since
dq
dY= (∂g
∂y1,···,∂g
∂ym)
and
dF
dX=
∂f1
∂x1···∂f1
∂xn
·
·
·
∂fm
∂x1···∂fm
∂xn
,
we find upon multiplying the matrices that
dh
dX= (m/summationdisplay
j=1∂g
∂yj∂fj
∂x1,m/summationdisplay
j=1∂g
∂yj∂fj
∂x2,···,m/summationdisplay
j=1∂g
∂yj∂fj
∂xn).
But we also know
dh
dX= (∂h
∂x1,∂h
∂x2,···∂h
∂xn).
376 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S.
Comparison of the last two formulas gives the stated result.
EXAMPLE. Let F:E2→E2andg:E2→E1be defined by
/braceleftbiggf1(x1,x2) =x1−ex2, g(y1,y2) =y2
1+y1y2.
f2(x1,x2) =ex1+x2
Then
F/prime(X) =/parenleftbigg1−ex2
ex1 1/parenrightbigg
, g/prime(Y) = (2y1+y2,y1).
Ifh=g◦F=g(F(x1,x2)) , then
dh
dX= (2y1+y2,y1)/parenleftbigg1−ex2
ex1 1/parenrightbigg
= (2y1+y2+y1ex1,−(2y1+y2)ex2+y1).
In particular, we find
∂h
∂x1= 2y1+y2+y1ex1
and
∂h
∂x2=−(2y1+y2)ex2+y1.
These formulas could also have been found by directly applying the corollary, viz.
∂h
∂x1=∂g
∂y1∂f1
∂x1+∂g
∂y2∂f2
∂x1= (2y1+y2)1 +y1(ex1),
and similarly for ∂h/∂x 2.
Many applications of the chain rule are more complicated. Consider a real valued
functiong(x1,x2,x3,t) , which depends on the point ˜X= (x1,x2,x3) as well as t. The
functiongcould be an expression of the temperature at a point ˜Xat timet. If the point
˜Xrepresents your position in the room, then since you move around the room, ˜Xis itself
a function of t. Thus, if your position is specified by ˜X=˜F(t) ,
x1=f1(t), x 2=f2(t), x 3=f3(t),
the temperature where you stand is h(t) =g(f1(t),f2(t),f3(t),t) . Since ˜F:E1→E3while
g:E4→E1, the chain rule is not directly applicable because gis defined on E4, while the
image of ˜Fis inE3.
A simple - if artificial - device clears up the difficulty. Introduce another variable x4
and letX= (x1,x2,x3,x4) . Then write g(x1,x2,x3,x4) , as well as X=F(t) , with
x1=f1(t), x 2=f2(t), x 3=f3(t), x 4=f4(t)≡t.
Now, as before, h(t) =g(f1(t),f2(t),f3(t),t) , butF:E1→E4andg:E4→E1. The
chain rule is thus applicable and gives
dh
dt=dg
dXdF
dt
9.2. THE DERIVATIVE OF COMPOSITE MAPS (“THE CHAIN RULE”). 377
= (∂g
∂x1,∂g
∂x2,∂g
∂x3,∂g
∂x4
d f1
dtd f2
dtd f3
dt
1
,
so thatdh
dt=∂g
∂x1∂f1
∂t+∂g
∂x2∂f2
∂t+∂g
∂x3∂f3
∂t+∂g
∂x4.
Sincex4≡t, the last equation can also be written as
dh
dt=∂g
∂x1df1
dt+∂g
∂x2∂f2
∂t+∂g
∂x3df3
dt+∂g
dt.
From a less formal viewpoint, this could have been obtained directly from the equation
h(t) =g(f1(t),f2(t),f3(t),t) without dragging in the artificial auxiliary variable x4. The
variablex4has been introduced to show how the chain rule applies. Once the process is
understood, the variable x4can (and should) be omitted.
EXAMPLE. Let g(x1,x2,x3,t) =x1t+ 3x2
2−x1x3+4
1+t2, and letx1= 3t−1, x2=
et−1, x3=t2−1 . Ifh(t) =g(x1(t), x2(t), x3(t),t) , we find
dh
dt=∂g
∂x1∂x1
dt+∂g
∂x2dx2
∂t+∂g
∂x3dx3
dt+∂g
∂t.
= (t−x3)3 + (6x2)et−1−(x1)2t+x1−8t
(1 +t2)2.
In particular, at t= 1 , we have x1= 2, x2= 1, x3= 0 so that
dh(1)
dt= (1−0)3 + (6)1 −(2)2 + 2 −8
4= 5.
It is straightforward to compute the second derivative d2h/dt2from the formula for
the first derivative.
d2h
dt2=∂
∂x1(dg
dt)dx1
dt+∂
∂x2(dg
dt)dx2
dt+∂
∂x3(dg
dt)dx3
dt+∂
∂t(dg
dt).
For this example, this gives
d2h
dt2= (−2t+ 1)3 + (6et−1)et−1+ (−3)2t+
(3 + 6x2et−1−2x1−81−3t2
(1 +t2)3).
Att= 1 , we have
∂2h
∂t2(1) = ( −2 + 1)3 + 6 −6 + (3 + 6 −4−8−2
8) = 4.
The next example brings to the surface an ambiguity in the notation∂
∂xfor partial
derivatives. This ambiguity is often a source of great confusion. Consider a scalar valued
functiong(x1,x2,t,s) . Ifx1=f1(t) andx2=f2(t) , then
h(t,s) =g(f1(t), f2(t),t,s)
378 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S.
depends on the two variables tands. In order to see how hchanges with respect to t,
we regardsas being held fixed and use the previous example to find
∂h
dt=∂g
∂x1∂f1
dt+∂g
∂x2∂f2
∂t+∂g
∂t.
We were careful and realized that the functions g(x1,x2,t,s) , a function with four
independent variables, and h(t,s) :=g(f1(t), f2(t), t,s) , a function with only two indepen-
dent variables, were different functions. The usual (occasionally confusing) approach is to
be less careful and write∂g
dt=∂g
∂x1∂f1
dt+∂g
∂x2∂f2
∂t+∂g
∂t.
In the above equation, the term ∂g/∂t on the right is the partial derivative of g(x1,x2,t,s)
with respect to twhile thinking of all four variables x1,x2,tandsas being independent.
On the other hand, the term ∂g/∂t on the left is the partial derivative of g(f1(t),f2(t),t,s)
as a function of two variables. After being spelled out like this, the formula does have a
clear meaning - but this is not at all obvious from a glance. One might even be mistakenly
tempted to cancel the terms ∂g/∂t from both sides of the equation.
It is often awkward to introduce a new name, as h(t,s) , forg(f1(t),f2(t),t,s) . Another
unambiguous procedure is available: use the numerical subscript notation for the partial
derivatives. Then g,1always refers to the partial derivative of gwith respect to its first
variable,g,2with respect to the second variable, etc. Thus, for the above example of
g(x1,x2,t,s) wherex1=f1(t) andx2=f2(t) , we have
∂g
∂t=g,1df1
dt+g,2df2
dt+g,3.
This clearly distinguishes the two time derivatives g,3and∂g/∂t .
The seemingly unnecessary comma in the notation is to take care of the possibility
of vector valued functions G(x1,x2,t,s) whose coordinate functions are indicated by sub-
scripts. For example, if G=/parenleftbiggg1
g2/parenrightbigg
is a map into E2, where the coordinate functions are
g1(x1,x2,t,s) andg2(x1,x2,t,s) , then ifx1=f1(t) andx2=f2(t) , we have
∂G
∂t=/parenleftbigg∂g1
∂t∂g2
∂t/parenrightbigg
=/parenleftbiggg1,1f/prime
1+g1,2f/prime
2+g1,3
g2,1f/prime
1+g2,2f/prime
2+g1,3/parenrightbigg
.
Hereg1,1=∂g1/∂x 1, etc. The notation f/prime
1fordf1(t)/dtcould also have been replaced by
f1,1—but this is unnecessary here since the fjare functions of one variable.
In applications, one commonly meets a problem of the following type. Let u(x,y) be
a scalar valued function which satisfies the wave equation uxx−uyy= 0 . IfF:E2→E2
is defined by
x=f1(ξ,η) =1
2(ξ,+η)
y=f2(ξ,η) =1
2(ξ−η)
and ifh=u◦F, that is,h(ξ,η) =u(f1(ξ,η),f2(ξ,η)) , what differential equation does h
satisfy? First, we compute hξandhη
∂h
∂ξ=∂u
∂x∂f1
∂ξ+∂u
∂y∂f2
∂ξ=ux(1
2) +uy(1
2) =1
2(ux+uy)
9.2. THE DERIVATIVE OF COMPOSITE MAPS (“THE CHAIN RULE”). 379
∂h
∂η=∂u
∂x∂f1
∂η+∂u
∂y∂f2
∂η=ux(1
2) +uy(−1
2) =1
2(ux−uy)
In a similar way the second derivatives hξξ,hξηandhηηare found,
∂2h
∂ξ2=∂(hξ)
∂x∂f1
∂ξ+∂(hξ)
∂y∂f2
∂ξ
=1
2∂
∂x(ux+uy)1
2+1
2∂
∂y(ux+uy)·1
2=1
4[uxx+ 2uxy+uyy]
∂2h
∂ξ∂η=∂(hξ)
∂η=∂(hξ)
∂x∂f1
∂η+∂(hξ)
∂y∂f2
∂η
=1
2∂
∂x(ux+uy)·1
2+1
2∂
∂y(ux+uy)·−1
2=1
4[uxx−uyy]
∂2h
∂η2=∂(hη)
∂x∂f1
∂η+∂(hη)
∂y∂f2
∂η
=1
2∂
∂x(ux−uy)·1
2+1
2∂
∂y(ux−uy)(−1
2) =1
4[uxx−2uxy+uyy]
Sincehξη=1
4[uxx−uyy] , andusatisfies the wave equation, we see that hsatisfies the
equation
hξη= 0,
so, in fact, the equations for hxiξandhηηare superfluous to obtain the desired result.
From this, it is easy to give another procedure for solving the wave equation, indepen-
dent of Fourier series. Because hξη= 0 , we know that h(ξ,η) =ϕ(ξ) +ψ(η) , where the
functionsϕandψare any twice differentiable functions. However, h(ξ,η) =u(ξ+η
2,ξ−η
2) .
Since the equations x=ξ+η
2, y =ξ−η
2may be solved for ξandηin terms of xandy,
viz.ξ=x+yandη=x−y, we haveh(x+y,x−y) =u(x,y) . Buth(ξ,η) =ϕ(ξ)+ψ(η) .
Consequently
u(x,y) =ϕ(x+y) +ψ(x−y).
This formula is the general solution of the one space dimensional wave equation. It expresses
uin terms of two arbitrary functions ϕandψ.
These functions ϕandψcan be chosen so that the function u(x,y) , a solution
of the wave equation, has any given initial position u(x,0) =f(x) and initial velocity
uy(x,0) =g(x) . Let us do this.
From the initial conditions we find
f(x) =u(x,0) =ϕ(x) +ψ(x)
g(x) =uy(x,0) =ϕ/prime(x)−ψ/prime(x).
After differentiating the first expression, one can solve for ϕ/primeandψ/prime,
ϕ/prime(x) =f/prime(x) +g(x)
2, ψ/prime(x) =f/prime(x)−g(x)
2.
Integrate these:
ϕ(x) =ϕ(0) +/integraldisplayx
0f/prime(s) +g(s)
2ds=ϕ(0) +f(x)−f(0)
2+1
2/integraldisplayx
0g(s)ds.
380 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S.
ψ(x) =ψ(0) +/integraldisplayx
0f/prime(s) +g(s)
2ds=ψ(0) +f(x)−f(0)
2+1
2/integraldisplayx
0g(s)ds.
Thus,
u(x,y) =ϕ(x+y) +ψ(x−y) =ϕ(0) +f(x+y)−f(0)
2+1
2/integraldisplayx+y
0g(s)ds+
ψ(0) +f(x−y)−f(0)
2−1
2intx−y
0g(s)ds.
Becausef(0) =ϕ(0) +ψ(0) , this simplifies to
u(x,y) =f(x+y)−f(x−y)
2s+1
2/integraldisplayx+y
x−yg(s)ds,
the famous d’Alembert formula for the solution of the one space dimensional wave equa-
tion in terms of the initial position f(x) and initial velocity g(x) . Unfortunately, simple
formulas like this are exceedingly rare. That is why a different, more generally applicable,
procedure was used earlier to solve the wave equation. As was seen in Exercise 6, p. 645,
the d’Alembert formula is recoverable from the Fourier series.
Exercises
(1) For the following function gandf, computed
dX(g◦F) and evaluate∂
∂x1(g◦F) at
the pointX0= (2,2) .
(a)g(y1,y2) =y1y2−y2e2y1,
F:yz= 2x1−x1x2, y 2=x2
1+x2
2
(b)g(y1,y2) = 7 +ey1siny2
F:y1= 2x1x2, y 2=x2
1−x2
2
(c)g(y1,y2,y3) =y2
1−y2
2−3y1y3+y2
F:y1= 2x1−x2, y 2= 2x1+x2, y 3=x2
1
(2) Letϕ(x1,x2,t) :=x2x2−te2x1. IfX=F(t) is defined by x1= 1−t2, x 2= 3t+1 ,
findd
dt(ϕ◦F) att= 1 .
(3) Letϕ(x,s,t ) :=xs+xt+st. Ifx=f(t) =t3−7 , compute∂
∂t(ϕ◦f) att= 3 .
Also compute∂2
∂t2(ϕ◦f) att= 3 .
(4) Ifu(x,y) =x2−y2, whileF:= (f1,fx) is given by x=f1(r,θ) =rcosθ, y =
f2(r,θ) =rsinθfindhrandhθ, whereh:=u◦F. Also compute, hrr, h rθand
hθθ.
9.2. THE DERIVATIVE OF COMPOSITE MAPS (“THE CHAIN RULE”). 381
(5) (a) Let u(x,y) be a scalar valued function and F:E2→E2be defined by the polar
coordinate transformation
f1(r,0) =rcosθ, f 2(rθ) =rsinθ,
Takeh:=u◦F. Findhr,hθ,hrr,hrθ, andhθ,θ. [Answer: h4=uxcosθ+
uysinθ, h rr=−uxxrsinθ+uyy(rcosθ−rsinθ)+uyyrcosθ−uxsinθ+uycosθ]
(b) Show that
uxx+uyy=hrr+1
r2hθθ+1
rhr.
(6) The two space dimensional wave equation is
utt=uxx+uyy
(a) If the space variables x,y are changed to polar coordinates (ex. 5) while the
time variable is not changed, the wave equation reads
htt=?
whereh(r,θ,t ) =u(rcosθ,rsinθ,t).
(b) If a given wave form depends only on the distance rfrom the origin and time
t, but not on the angle ∂, how does the wave equation for hsimplify?
(c) Consider the equation you found in b. Use the method of separation of variables
and seek a solution in the form h(r,t) =R(r)T(t) . What are the resulting
ordinary differential equations? Compare the equation for R(r) with Bessel’s
differential equation.
(7) Ifw=f(x,y,s ) , whilex=ϕ(y,s,t ) andy=ψ(s,t) , find expressions for the partial
derivative of the composite function g(ϕ(ψ,s,t ),ψs) with respect to sandt.
(8) (a) Let u(x,y) =f(x−y) . Show that usatisfies the partial differential equation
ux+uy= 0.
(b) Letu(x,y) =f(xy) . Show that usatisfies the equation xux−yuy= 0 .
(c) Letu(x,y) =f(x
y) . Show that usatisfies the equation
xux+yuy= 0.
(d) Letu(x,y) =f(x2+y2) , souonly depends on the distance from the origin.
Show that usatisfies
yux−xuy= 0.
(9) Letu(x,y) satisfy the equation xux+yuy= 0 .
(a) Change the equation to polar coordinates [Answer: if h(r,θ) :=u(rcosθ,rsinθ) ,
thenrhr= 0 ].
(b) Solve the equation for h(r,θ) and use it to deduce that u(x,y) =f(x
y) for some
functionf. (cf. Ex. 8c)
382 CHAPTER 9. DIFFERENTIAL CALCULUS OF MAPS FROM ENTOEM, S.
(10) Assume u(x,y) satisfies the equation
uxx−2uxy−3uyy= 0.
(a) Choose the constants α,β,γ , andδso that after the change of variables x=
αξ+βη, y =γξ+δη, the equation for h(ξ,η) =u(αξ+βη,γξ +δη) ishξη= 0 .
(b) Use the result of part (a) to find the general solution of the equation for u.
[Answer:u(x,y) =ϕ(3x−y) +ψ(x+y) ].
(11) Iff(x,y) is a known scalar valued function, find both partial derivatives of the
functionf(f(x,y),y) .
(12) IfW=G(Y) andY=F(X) are defined by
G:/braceleftbiggw1=ey1−y2
w2=ey1+y2, F :/braceleftbiggy1=x2
1−3x2−x3
y2=x1+x2
2+ 3x3,
findd
dX(G◦F) .
(13) Letu(x,y) be a solution of the two dimensional Laplace’s equation uxx+uyy= 0 .
(a) Ifudepends only on the distance from the origin u(x,y) =h(r) , wherer=
x2+y2, what ordinary differential equation does hsatisfy? Compare your
answer with that found in Exercise 5.
(b) Solve the resulting equation for hand deduce that all the solutions of the two
dimensional Laplace equation which depend only on the distance from the origin
are of the form
u(x,y) =A+Blog(x2+y2),
whereAandBare constants.
(c) Now do the same thing all over again for a solution u(x1,x2,...,x n) of then
dimensional Laplace equation ux1x1+...+uxnxn= 0 , i.e. find the form of
all solutions which only depend on r=/radicalbig
x2
1+...+x2n,u(x1,...,x n) =h(r) .
[Answer:u(x1,...,x n) =A+B
(x2
1+...+x2n)n−2
2=A+B
rn−2, n≥3 ].
(14) Iff(t) is a differentiable scalar valued function with the property that f(x+y) =
f(x) +f(y) for allx,y∈E1, prove that f(x)≡kxwherek=f(1) .
(15) (a) Find the general solution of the partial differential equation ux−2uy= 0 . [Hint:
Introduce new variables as in Ex. 10]
(b) What is the solution if one requires that u(x,0) =x2? [Answer: u(x,y) =
(x+1
2y)2].
Chapter 10
Miscellaneous Supplementary
Problems
1. (a)Sn, n= 1,2,..., be a given sequence. Find another sequence ansuch that
SN=N/summationdisplay
n=1an. In other words, given the partial sums Sn, find a series whose
partial sums are Sn. To what extent are the anuniquely determined?
(b) Apply part (a) to find an infinite series/summationtextanwhosenth partial sum Snis
given by
(i)Sn=1
n, (ii)Sn=e−n
2. LetS={x∈R:x∈(−1,1)}. Define addition on Sby the formula x⊕y=
x+y
1+xy, x,y∈S, where the operations on the right are the usual ones of arithmetic.
Show that the elements of Sform a commutative group with the operation ⊕.
3. (a) Ifan→a, prove thata1+a2+···+an
n→aalso.
(b) Assume that fis continuous on the interval [0 ,∞] and lim
x→∞f(x) =A. Define
HN=1
N/integraldisplayN
0f(x)dx. Prove that lim
x→∞HNexists and find its value. [Hint:
InterpretHNas the average height of the function f].
4. (a) Suppose that allthe zeroes of a polynomial P(x) are real. Does this imply that
all the zeroes of its derivative P/prime(x) are also real? (Proof or counterexample).
What can you say about higher derivatives P(k)(x) ?
(b) Define the nth Laguerre polynomial by
Ln(x) =exdn
dxn[xne−x].
Show that Lnis a polynomial of degree n. Prove that the zeroes of Ln(x) are
all positive real numbers, and that there are exactly nof them.
383
384 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
5. Iff(x) has a Taylor series: f(x) =∞/summationdisplay
n=0anxn(which converges to ffor|x|< ρ so
the remainder does go to zero there) prove that f(cxk) , wherecis a constant and
ka positive integer, has the Taylor series
f(cxk) =∞/summationdisplay
n=0ancnxnk
which converges to f(cxk) for |x|<(ρ
|c|)1/k. You must show that i) the Taylor
coefficients for f(cxk) areancn, that ii) the power series for f(cxk) converges for
|x|<(ρ
|c|)1/k, and that iii) the remainder tends to zero. Apply the result to obtain
the Taylor series for cos(2 x2) from that of cos x.
6. Yet another proof of Taylor’s Theorem. Beginning with equation 9 on p. 98, define
the function K(s) by
K(s) =f(s)−N/summationdisplay
n=0f(n)(x0)
n!(s−x0)n−A(s−x0)N+1
(N+ 1)!,
whereAis picked so that K(ˆx) = 0 .
(a) Verify that K(x0) =K/prime(x0) =...K(N)(x0) = 0 .
(b) Use Rolle’s Theorem to prove that if a function K(s) satisfies the properties of a),
and ifK(ˆx) = 0 , then there is a ξbetween ˆxandx0such thatK(N+1)(ξ) = 0 .
(c) Apply parts a) and b) to prove Taylor’s Theorem.
7. Assume/summationtextanconverges. You are to investigate the convergence of/summationtexta2
nand/summationtext/radicalbig
|an|under various hypotheses.
(a)anarbitrary complex number
(b)an≥0 .
(c) lim
n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglean+ 1
an/vextendsingle/vextendsingle/vextendsingle/vextendsingle<1 (not= 1).
8. The harmonic series 1+1
2+1
3+···has been said to diverge with “infuriating slowness”.
Find a number Nsuch that 1 +1
2+1
3+···+1
Nis at least 100. Compare this with
Avogadro’s number ∼6×1023.
9. Consider the series/summationtext∞
n=1an, where the an’s are real.
(a) Letb1,b2,b3,...andc1,c2,c3,...denote the positive and negative terms respec-
tively from a1,a2,.... If/summationtext∞
n=1anconverges conditionally but not absolutely,
prove that both series/summationtext∞
n=1bnand/summationtext∞
n=1cndiverge .
(b) Letd1,d2,d3,..., denote the terms a1,a2,a3,... rearranged in any way. Prove
Riemann’s theorem, which states that if/summationtext∞
n=1anconverges conditionally but
not absolutely, then by picking some suitable rearrangement, the series/summationtext∞
n=1dn
can be made to converge to any real number, while using other rearrangements,
it can be made to diverge to plus or minus infinity.
385
10. IfAandBare subsets of a linear space V, a) show that span {A∩B} ⊂span{A}∩
span{B}. Give an example showing that span {A∩B}may be smaller than
span{A} ∩span{B}.
b). Show that if A⊂B⊂span{A}, then span {A} ⊃span{B}.
11. LetA={X1,...,X k}be a set of vectors in a linear space V. Denote by cs A
(coset ofA) the set
csA={X∈V:X=k/summationdisplay
j=1ajXj,wherek/summationdisplay
j=1aj= 1}.
Prove that cs Ais a coset of V, in fact, the smallest coset of Vwhich contains the
vectorsX1,...,X k.
12. (a) Consider the set of real numbers of the form a+b√
2 , whereaandbare rational
numbers. Prove that this set is a vector space over the field of rational numbers.
What is the dimension of this vector space?
(b) Consider the set of numbers of the form a+bi, whereaandbare real numbers
andi=√−1 . Prove that this set is a vector space over the field of realnumbers
and find its dimension.
13. IfF1andF2are fields with F1⊂F2, we callF2anextension field ofF1– such
asR⊂C. As such, we may think of F2as a vector space over the field F1(see
exercise 1l). In other words, take F2as an additive group and take the scalars from
F1. If this vector space is finite dimensional, the field F2is called a finite extension
ofF1, and the dimension nof this vector space is called the degree of the extension
and written n= [F2:F1] .
(a) Prove that every element ξ∈F2satisfies an equation
anξn+an−1ξn−1+···+a0= 0,
where theak∈F1andn= [F2:F1] . [Hint: look at the examples of exercise
1l].
(b) IfF1⊂F2⊂F3are fields with
[F2:F1] =n<∞and [F3:F2] =m<∞,
prove that [ F3:F1]<∞, in fact, prove
[F3:F1] = [F3:F2]]F2:F1] =nm.
(c) LetF1be the field of rationals, F2the field whose elements have the form
a+b√
3 , whereaandbare rational, and let F3be the field whose elements
have the form c+d√
5 , wherecanddare inF2. Compute [ F2:F1] and find
the polynomial of part a) satisfied by (1 −√
3)∈F)2. Compute [ F3:F2] and
[F3:F1] . Find a basis for F3as a vector space whose scalars are elements of
F1. [The ideas in this problem are basic to modern algebra, particularly Galois’
theory of equations.]
386 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
14. LetPj= (αj,βj), j= 1,...,n,α j/negationslash=αkbe anyndistinct points in the plane R2.
One often wants to find a polynomial p(x) =a0+a1x+···+aNxNwhich passes
through these npoints,p(αj) =βj, j= 1,...,n . Thus,p(x) is an interpolating
polynomial . Given any points P1,...,P n, prove that a unique interpolating polyno-
mialp(x) degreen−1(=N) can be found. (More about this is in Exercises 17-18
below).
15. LetL1andL2be linear operators mapping V→V. Then they can be both
multiplied and added (or subtracted). The bracket product orcommutator
[L1,L2]≡L1L2−L2L1
“measures the non-commutativity”. It is important in mathematics and physics. [In
quantum mechanics, the observables - like energy, momentum, and position - are
represented by self-adjoint operators. Two observables can be measured at the same
time if and only if their associated operators commute]. Prove the identities
(a) [L1,L1] = 0,[L1,I] = 0
(b) [L1,L2] =−[L2,L1]
(c) [aL1,L2] =a[L1,L2] , a scalar
(d) [L1+L2,L3] = [L1,L3] + [L2,L3]
(e) [L1,L2,L3] = [L1,L2]L3+L2[L1,L3]
(f) [L1,[L2,L3]] + [L2,[L3,L1]] + [L3,[L1,L2]] = 0
(Part f is the Jacobi identity . It has been said that everyone should verify it once in
her lifetime.)
16. * Consider the normalized Legendre Polynomials,
en(x) =/radicalbigg
2
2n+ 11
2nn!dn
dxn(x2−1)n, n = 0,1,2,...
which are an orthonormal set of polynomials in L2[−1,1], enbeing of degree n. If
f∈C[−1,1] , prove that
PNf=N/summationdisplay
n=0/angbracketleftf, en/angbracketrighten
converges to fin the norm of L2[−1,1] . [Hint: Use the form of the Weierstrass
Approximation Theorem (p. 255) and the method of Theorem (p. 241)].
17. * We again take up the interpolation problem begun in Exercise 13 above. Let Pj=
(αj,βj), j= 1,2,...,n benpoints in the plane, αi/negationslash=αj. Although we proved there
is a unique polynomial p(x) =a0+a1x+···+an−1xn−1of degreen−1 passing
through the npoints, the proof was entirely non-constructive. Here we (or you)
explicitly construct the polynomial.
(a) Show that the polynomial of degree n−1
˜pj(x) = Πn
k=/lscript
k/negationslash=j(x−αk)
is zero ifx=αk, k/negationslash=j, but ˜pj(αj)/negationslash= 0 .
387
(b) Construct a polynomial pj(x) with the property pj(αk) =δjk.
(c) Show that
p(x) =n/summationdisplay
j=1βjpj(x)
is the desired (unique by Ex. 13) interpolating polynomial.
(d) LetP1= (1,1), P2= (2,1), P3= (4,−1), P4= (−1,−2) .
Find the interpolating polynomial using the above construction.
18. * Iffis some complicated function, it is often useful to use an interpolating polyno-
mial instead of the function. Then the polynomial p(x) will pass through the points
Pj= (αm,f(αj)), j = 1,...,n , so by Exercise 16,
p(x) =n/summationdisplay
j=1f(αj)pj(x).
] How much will pdiffer from fin an interval [ a,b] containing the αj? You must
estimate the remainder R=f−p.
(a) Assume f∈Cn[a,b] . SinceR(x) =f(x)−p(x) vanishes at x=αj, j=
1,...,n , it is reasonable to write
R(x) = (x−α1)···(x−αn)·(?)
Fixˆxand define the constant Aby
f(ˆx)−p(ˆx) =A(ˆx−α1)···(ˆx−αn).
By a trick similar to that used in Taylor’s Theorem (cf. P. 104j Ex. 12), prove
thatA=f(n)(ξ)/n! whereξis some point in ( a,b) . Thus,
f(ˆx) =p(ˆx) +(ˆx−α1)···(ˆx−αn)
n!f(n−1)(ξ), ξ∈(a,b).
(b) Letf(x) = 2x, andα1=−1, α2= 0, α3= 1, α4= 2.
Find the approximating polynomial and find an upper bound for the error in the
interval [ −2,2] .
19. Ifxis irrational and a,b,c, and d are rational (with ad−bc/negationslash= 0) , prove thatax+b
cx+d
is irrational.
20. Prove by induction that
1 + 3 + 5 + ···+ (2n−1) =n2.
21. (a) If x≥0 , use the mean value theorem to prove
ex≥1 +x.
388 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
(b) Ifak≥0 , prove that
n/summationdisplay
k=1ak≤Πn
k=1(1 +ak)≤ePn
k=1ak,
(where Πn
k=1bk=b1b2···bn).
(c) Ifak≥0 , prove that the infinite product Π∞
k=1(1 +ak) := lim n→∞Πn
k=1(1 +ak)
converges if and only if the infinite series/summationtext∞
k=1converges.
22. Letan+1=2
1+an, wherea1>1 . Prove that
(a) the sequence a2n+1is monotone decreasing and bounded from below.
(b) the sequence a2nis monotone increasing and bounded from above.
(c) does lim
n→∞anexist?
23. Letak, k= 1,...,n + 1 be arbitrary real numbers which satisfy a1+a2
2+···+
an
n+an+1
n+1= 0 . Show that P(x) =a1+a2x+···+anxn−1has at least one zero for
x∈(0,1) .
24. Suppose f∈C2in some neighborhood of x0. Prove that
lim
h→0f(x0+h)−2f(x0) +f(x0−h)
h2=f/prime/prime(x0).
25. Lets(x) andc(x) be continuously differentiable functions defined for all x, and
having the properties
s/prime(x) =c(x), c/prime(x) =s(x)
s(0) = 0, c (0) = 1.
(a) Prove that c2(x)−s2(x) = 1 .
(b) Show that c(s) ands(x) are uniquely determined by these properties.
26. Consider/summationtextanand/summationtextbn.
(a) If lim
n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglebn
an/vextendsingle/vextendsingle/vextendsingle/vextendsingle=K, K /negationslash= 0,∞, then the series both converge or diverge together.
(b) If/summationtextanconverges and lim
n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglebn
an/vextendsingle/vextendsingle/vextendsingle/vextendsingle= 0 , then/summationtextbnconverges.
(c) If/summationtextanconverges and lim
n→∞/vextendsingle/vextendsingle/vextendsingle/vextendsinglebn
an/vextendsingle/vextendsingle/vextendsingle/vextendsingle=∞, then the series/summationtextbnmay converge or
diverge (give examples).
(d) Apply these to:
(i)∞/summationdisplay
n=21
n−√n
389
(ii)∞/summationdisplay
n=11
n3−2√n
(iii)∞/summationdisplay
n=1(−1)nsinπ
n. (Hint: as x→0,sinx
x→1 ).
27. The following (a weak form of Stirling’s formula ) is an improvement of the result on
page 64, Ex. 6.
nlogn−(n−1)<logn!<(n+ 1) log(n+ 1)−2 log 2 −(n−1),
from which one finds
nne−n+1<n!<1
4(n+ 1)(n+1)e−n+1.
Prove these.
28. (a) Find the Taylor series expansion for f(x) =e−xaboutx= 0 .
(b) Show that the series found in (a) converges to e−xfor allxin the interval
[−r,r] , wherer>0 is an arbitrary but fixed real number.
29. Consider the sequence
SN=/integraldisplayN
2sinπx
xdx.
Does lim
N→∞SNexist? [Hint: observe that SNcan be written as
SN=N−1/summationdisplay
2an,
where
an=/integraldisplayn+1
nsinπx
xdx.
Sketch a graph ofsinπx
x,x≥2 , to deduce - by inspection - the needed properties of
thean’s. Please do not attempt to evaluate the integrals for an].
30. LetA={p∈P9:p(x) =p(−x)}.
(a) Prove that Ais a subspace of P9.
(b) Compute the dimension of A.
31. LetXandYbe elements in a real linear space. Prove that /bardblX/bardbl=/bardblY/bardblif and only
if (X+Y)⊥(X−Y) .
32. In the space R2, introduce the new scalar product
<X,Y > =x1y1+ 4x2y2,
whereX= (x1,x2) andY= (y1,y2) .
390 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
(a) Verify that this indeed is a scalar product and define the associated norm /bardblX/bardbl.
(b) LetX1= (0,1) andX2= (4,−2) . Using thisnorm and scalar product, find an
orthonormal set of vectors e1ande2such thate1is in the subspace spanned
byX1.
33. LetHbe a scalar product space with XandYinH. Find a scalar αwhich
makes /bardblX−αY/bardbla minimum. For this α, how areX−αYandYrelated? [Hint:
Draw a picture in E2].
34. If∞/summationdisplay
n=1anconverges, where an≥0 , does the series∞/summationdisplay
1√a
n2also converge? Proof or
counterexample.
35. Use the Taylor series about x0= 0 to calculate sin.2 making an error less than
.005 . Justify your statements.
36. LetA= span {(1,1,1,1),(1,0,1,0)}be a subspace of E4. Find the orthogonal
complement, A⊥, ofAby giving a basis for A⊥.
37. Prove that
(a) 1 +1
8<∞/summationdisplay
k=11
k3<1 +1
2.
(b) 1 +1
2∞/summationdisplay
k=11
k2<1 +3
4.
38. Letakbe a sequence of positive numbers decreasing to zero, ak→0 , and letSN=
a1+a2+···+aN.
(a) Prove that SN≥NaN.
(b) Use this to estimate the number, N, of terms needed to make
N/summationdisplay
k=1k−1/4>1000.
39. Prove or give a counterexample:
(a) If∞/summationdisplay
n=1bnconverges, then∞/summationdisplay
n=1b2nmust converge.
(b) If∞/summationdisplay
n=1|bn|converges, then∞/summationdisplay
n=1|b2n|must converge.
40. LetX1andX2be elements of a scalar product space.
(a) IfX1⊥X2, prove that /bardblX1−aX2/bardbl ≤ /bardblX1/bardblfor any real number a.
391
(b) Prove the converse, that is, if /bardblX1−aX2/bardbl ≤ /bardblX1/bardblfor every real number a,
thenX1⊥X2. [Hint: After your first approach has failed, try looking at the
problem geometrically. How would you pick ato minimize the left side of the
inequality?].
41. LetSn=a1+a2+···+an, wherean→0 asn→ ∞ . Prove that Snconverges if
and only if S2n=a1+a2+···+a2n−1+a2nconverges (one could also use S3netc.).
42. Show that the error in approximating the series∞/summationdisplay
n=11
nnby the first Nterms is less
thanN−N−1.
43. A sample “multiplication” for points X= (x1,x2,x3) andY= (y1,y2,y3) inR3is
to define
X⊙Y≡(x1y2,x2y2,x3y3).
Define a multiplicative identity by yourself. Using these definitions for the multiplica-
tive structure and the usual rules for the additive structure, show that the resulting
algebraic object is not a field.
44. (a) Assume an≥0 andbn≥0 . Prove that ∠(an+bn) converges if and only if the
series ∠anand∠bnbothconverge.
(b) What if you allow the bn’s to be negative?
45. (a) Show that the vectors e1= (1√
2,1√
2), e2= (1√
2,−1√
2) form an orthonormal basis
forE2.
(b) Write the vector X= (7,−3) in the form X=a1e1+a2e2, using the scalar
product to find a1anda2(don’t solve linear equations).
46. Consider the linear space P2as a subspace of L2[0,1] .
(a) Ifp(x) = 1−x2, compute /bardblp/bardbl.
(b) Find the orthonormal basis for A⊥, whereA= span {2 +x}.
(c) Find the polynomial ϕ∈P2such that
/angbracketleftp, ϕ/angbracketright=p(1) for all p∈P2,
that is, the same ϕshould work for all p’s.
47. Give formal proofs for the following (trivial) properties of a norm on a linear space.
Only the axioms may be used.
(a)/bardbl −X/bardbl=/bardblX/bardbl
(b)/bardblX−Y/bardbl=/bardblY−X/bardbl
(c)/bardblX+Y/bardbl ≥ /bardblX/bardbl − /bardblY/bardbl
(d)/bardblX1+X2+···+Xn/bardbl ≤ /bardblX1/bardbl+/bardblX2/bardbl+···+/bardblXn/bardbl(I suggest induction here).
392 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
48. Consider R2with the norms /bardbl /bardbl 1,/bardbl /bardbl 2, and /bardbl /bardbl ∞.
(a) Draw a sketch of R2indicating the unit ball for each of these three norms. (The
ball may not turn out to be “round”).
(b) Which of these three linear spaces have the following property: “given any sub-
spaceMand a point X0not inM, then there is a unique point onMwhich
is closest to M.”
49. Are the following scalar products the set of functions continuous on [ a,b] ? Proof or
counterexample.
(a) [f,g] = (/integraldisplayb
af(x)dx)(/integraldisplayb
ag(x)dx)
(b) [f,g] = (/integraldisplayb
a|f(x)|dx)(/integraldisplayb
a|g(x)|dx)
50. (a) Let dim V=nand{X1,...,X n} ∈V. Prove that {X1,...,X n}are linearly
independent if and only if they span V(so in either case, they form a basis for
V).
(b) Let {e1,...,e n}be an orthonormal set of vectors for an inner product space
H. Prove this set of vectors is a complete orthonormal set for Hif and only if
n= dimH.
(c) Prove that dim V= largest possible number of linearly independent vectors in
V.
51. (a) Let XandYbe any two elements in an inner product space. Prove that the
parallelogram law holds
/bardblX+Y/bardbl2+/bardblX−Y/bardbl2= 2/bardblX/bardbl2+ 2/bardblY/bardbl2
(cf. page 192, Ex. 9).
(b) Consider the set of continuous functions on [0 ,1] with the uniform norm, /bardblf/bardbl∞=
max 0≤x≤1|f(x)|. Show that this norm cannot arise from an inner product, i.e.
there is no inner product such that for all f,/bardblf/bardbl∞=/radicalbig
/angbracketleftf, f/angbracketright. [Hint: If there
were, the relationship of part a would hold between the norms of various ele-
ments. Show that relationship does not, in fact, hold for the function f(x) = 1
andg(x) =x].
52. (a) Let Hbe a finite dimensional inner product space and /lscript(X) a linear functional
defined for all X∈H. Show that there is a fixed vector X0∈Hsuch that
/lscript(X) =/angbracketleftX, X 0/angbracketright for allX∈H.
This shows that every linear functional can be represented simply as the result
of taking the inner product with some vector X0. [Hint: First pick a basis
{e1,...,e n}forHand letcj=/lscript(en) . Now use the fact that the ej’s are a
basis and that /lscriptis linear].
393
(b) Consider the linear space P2with theL2[0,1] inner product. This gives an
inner product space H.
(i) Show that /lscript(p) =p(1
3) is a linear functional.
(ii) Find a polynomial p0such that/lscript(p) =/angbracketleftp, p 0/angbracketrightfor allp∈H.
53. Consider the set Sof pairs of real numbers X= (x1,x2) . Define
X+Y= (x1+y1,x2+y2), aX = (ax1,x2).
IsS, with this definition of vector addition and multiplication by scalars, a vector
space?
54. By inspection, place suitable restrictions on the contents a,b,c, ···in order to make
the following operator linear:
Tu=a[d3u
dx3]2+bx2d2u
dx2+cudu
dx+eu+fsinu+g.
55. Consider the operator D=d
dxon the linear space Pnof all polynomials of degree
less than or equal to n. Find R(D) and N(D) as well as dim R(D) and dim N(D) .
56. Let
A=
1−2
2 0
3 1
, B =/parenleftbigg−1 0 −2
2 1 0/parenrightbigg
,andC=
3 0 2
1 4 −1
0−2 0
.
Compute all of the following products which make sense:
AB, BA, AC, CA, BC, CB, A2, B2, C2, ABC,CAB.
57. Consider the mapping A:R4→R3which is defined by the matrix
A=
1−1 1 1
2 1 1 4
0−3 1 −2
(a) Find bases for N(A) and R(A) .
(b) Compute dim N(A) and dim R(A) .
58. LetAbe a square matrix. Consider the system of linear algebraic equations
AX=Y0,
whereY0is a fixed vector. Assume these equations have two distinct solutionsX1
andX2,
AX 1=Y0, AX 2=Y0, X 1/negationslash=X2.
(a) Find a third solution X3.
394 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
(b) Does there exist a vector Y1such that the equations
AX=Y1
have nosolutions? Why?
(c) detA=?
59. LetQbe a parallelepiped in Enwhose vertices Xkare at points with integer coor-
dinates,
Xk= (a1k,a2k,···ank), a ikintegers.
Prove that the volume of Qis an integer.
60. LetAandBbe self-adjoint matrices. Prove that their product ABis self-adjoint
if and only if AB=BA.
61. Solve the following initial value problems.
(a)u/prime/prime+ 8u/prime+ 16u= 0, u (0) =1
2, u/prime(0) = 0
(b)u/prime/prime+ 10u/prime+ 16u= 0, u (0) = 1, u/prime(0) = 2
(c)u/prime/prime+ 64u= 0, u (0) =1
4, u/prime(0) = 1
(d)u/prime/prime+ 4u/prime+ 5u= 0, u (0) = 2, u/prime(0) = −1
(e) 2u/prime/prime+ 6u/prime+ 5u= 0, u (0) = 0, u/prime(0) = −2
(f) 4u/prime/prime−4u/prime+u= 0, u (1) = −1, u/prime(1) = 0
(g)u/prime/prime+ 8u/prime+ 16u= 2, u (0) =1
2, u/prime(0) = 0
(h)u/prime/prime+ 8u/prime+ 16u=t, u (0) =1
2, u/prime(0) = 0
(i)u/prime/prime+ 8u/prime+ 16u=t−2, u (0) = 0, u/prime(0) = 0
(j)u/prime/prime+ 8u/prime+ 16u=t−2, u (0) =1
2, u/prime(0) = 0
(k)u/prime/prime+ 10u/prime+ 16u=t, u (0) = 1, u/prime(0) = 2
(l)u/prime/prime+ 64u= 64, u (0) =1
4, u/prime(0) = 2
(m)u/prime/prime+ 64u=t−64, u (0) =3
4, u(0) = 0
(n) 2u/prime/prime+ 6u/prime+ 5u=t2, u (0) = 0, u/prime(0) = −2
62. (The complex numbers as matrices).
(a) Show that the set of matrices
C={/parenleftbigga−b
b a/parenrightbigg
:aandbare real numbers }
is a field.
(b) Find a map ϕ:C→complex numbers such that ϕis bijective and such that
for allA,B∈C
(i)ϕ(A+B) =ϕ(A) +ϕ(B)
(ii)ϕ(AB) =ϕ(A)ϕ(B).
395
63. (Quaternions as matrices). A definition: A division ring is an algebraic object which
satisfies all of the field axioms except commutativity of multiplication.
(a) Show that the set of matrices
Q={/parenleftbiggz−¯w
w ¯z/parenrightbigg
:z,w are complex numbers }
form a division ring with the usual definitions of additions and multiplication for
matrices.
(b) If we write z=x+iy, w =u+ivwherei=√−1 andx,y,u , andvare real
numbers, then Qcan be considered as a vector space over the reals with basis
1=/parenleftbigg1 0
0 1/parenrightbigg
i=/parenleftbiggi0
0−i/parenrightbigg
j=/parenleftbigg0−1
1 0/parenrightbigg
k=/parenleftbigg0i
i0/parenrightbigg
.
Compute i2,j2,k2,ij,jk,ki,ji,kj, and ik. (The set Qis called the quater-
nions ).
64. Let
A=
2−3 1 0
0 2 −3 1
0 0 2 −3
0 0 0 2
.
(a) Find det A.
(b) FindA−1.
(c) SolveAX=Y, whereY=
2
8
8
−16
.
(d) LetL:P3→P3be the linear operator defined by
Lp=p/prime/prime−3p/prime+ 2p,(u/prime=du
dx).
Find the matrix eLeforLwith respect to the following basis for P3
e1(x) = 1, e 2(x) =x, e 3(x) =x2
2, e 4(x) =x3
3!.
(e) Use the above results to find a solution of
Lu= 2 + 8x+ 4x2−8
3x3.
[Hint: Express the right side in the basis of part d.].
65. LetHbe an inner product space, and suppose that Ais a symmetric operator,
A∗=A, with the additional property that A2=A. Show that there exist two
subspacesV1andV2ofHwith all of the following properties
396 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
(i)V1⊥V2
(ii) IfX∈V1, thenAX=X
(iii) IfY∈V2, thenAY= 0
(iv) IfZ∈H, thenZcan be written uniquely as Z=X+YwhereX∈V1and
Y∈V2.
66. (a) Find the inverse of the matrix
A=
2 1 0
−1 0 1
0−1−1
.
(b) Use the result of a) to solve AX=bforXwhereb= (7,−3,2) .
67. LetAandBbe 2×2 positive definite matrices with det A= detB. Prove that
det(A−B)<0 .
68. LetL:V1→V2be a linear operator with LX 1=Y1andLX 2=Y2. Give a proof
or counterexample to each of the following assertions:
(a) IfX1andX2are linearly independent, then Y1andY2must be linearly
independent.
(b) IfY1andY2are linearly independent, then X1andX2must be linearly
independent.
69. Letp0,p1,p2,,... be an orthogonal set of polynomials in [ a,b] wherepnhas degree
n.
(a) Prove that pnis orthogonal to 1 ,x,x2,...,xn−1.
(b) Prove that pnis orthogonal to any polynomial qof degree less than n.
(c) Prove that pnhas exactly ndistinct real zeros in ( a,b) . [Hint: Let α1,...,α k
be the places in ( a,b) wherepn(x) changes sign, so p(x) =r(x)(x−α1)(x−
α2)...(x−αk) wherer(x) is a polynomial of degree n−kwhich does not
change sign for xin (a,b) , sayr(x)≥0 . Show that
/integraldisplayb
ap(x)(x−α1)···(x−αk)dx> 0.
Ifk<n , show that this contradicts the result of part b).].
70. Consider the system of inhomogeneous equations
a11x1+···+a1nxn=b,
...
ak1x1+···+aknxn=bn.
397
LetA= ((aij)) and letAbdenote the augmented matrix
Ab=
a11···a1nb1
...
ak1···aknbn
formed by adding the bj’s as an extra column to A. Prove that the given system of
equations has a solution if and only if dim R(A) = dim R(Ab) .
71. LetAbe ann×nmatrix.
(a) Show that you can not solve the equation
A2=−I
ifnis odd.
(b) Find a 2 ×2 matrixAsuch thatA2=−I.
(c) Ifnis even, find an n×nmatrixAsuch thatA2=−I.
72. LetAbe ann×nmatrix such that A2=I. Prove that dim R(A+I)+dim R(A−I) =
n.
73. Letf(x,y) = (y−2x2)(y−x2) . Show that the origin is a critical point. Then
show that if you approach the origin along a straight line, the origin appears to be
a minimum. On the other hand, show that if curved paths are also used, then the
origin is a saddle point of f. [The point of this exercise is to illustrate the fact that
the nature of a critical point cannot be determined by merely approaching it along
straight lines].
74. (a) Let Abe a diagonal matrix, no two of whose diagonal elements are the same.
IfBis another matrix and AB=BA, prove that Bis also diagonal.
(b) LetAbe a diagonal matrix, Ba matrix with at least one zero-free column
and with the further property that AB=BA. Prove that all of the diagonal
elements of Aare equal.
75. (a) If/summationdisplay
anconverges, where an≥0 , prove that/summationdisplay√an
npconverges if p >1
2.
[Hint: Schwarz].
(b) Find an example showing that the series may diverge if p=1
2.
76. Let [X,Y] be an inner product on R3with basis vectors e1,e2,e3, not necessarily
orthonormal. Let aij= [ei,ej] . Prove that the quadratic form
Q(X) =3/summationdisplay
i=13/summationdisplay
j=1aijxixj
is positive definite.
77. IfAis self-adjoint and AX=λ1X, AY =λ2Ywithλ1/negationslash=λ2, prove that X⊥Y.
398 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
78. LetSbe a positive definite matrix. Prove that det S > 0 . [Hint: Consider the
matrixA(t)≡tS+ (1−t)I, where 0 ≤t≤1 . Show that A(t) is positive definite,
so detA(t)/negationslash= 0 . Then use the fact that A(0) =IandA(1) =Sto obtain the
conclusion].
79. Consider the linear space of infinite sequences
X= (x1,x2,x3,···)
with the usual addition. Define the linear operator S(the right shift operator) by
SX= (0,x1,x2,x3,···)
(a) DoesShave a left inverse? If so, what is it?
(b) DoesShave a right inverse? If so, what is it?
80. Find a right inverse for the matrix
A=/parenleftbigg1 0 1
0 1 0/parenrightbigg
.
CanAhave a left inverse? Why?
81. Which of the following statements are true for allsquare matrices A? Proof or
counterexample.
(a) IfA2=I, then detA=I.
(b) IfA2=A, then detA= 1.
(c) IfA2= 0 , then det A= 0
(d) IfA2=I−A, then detA2= 1−detA.
82. LetLbe a linear operator on an inner product space Hwith inner product <,> .
Define
[X,Y] =/angbracketleftLX, LY /angbracketright.
Under what further condition(s) on Lis [X,Y] an inner product too?
83. LetL:H→Hbe an invertible transformation on the inner product space H. If
L“preserves orthogonality” in the sense that X⊥YimpliesLX⊥LY, prove that
there is a constant αsuch thatR≡αLis an orthogonal transformation.
84. LetHbe an inner product space. If the vectors X1andX2are at opposite ends of
a diameter of the sphere of radius rabout the origin, and if Yis any other point on
that sphere, prove that Y−X1is perpendicular to Y−X2, proving that an angle
inscribed in a hemisphere is a right angle.
85. IfLis skew-adjoint, L∗=−L, prove that
/angbracketleftX, LX /angbracketright= 0 for all X.
399
86. LetDnbe an×nmatrix with xon the main diagonal and 1/primeson both the sub-
and super-diagonals, so
D2=/parenleftbiggx1
1x/parenrightbigg
, D 3=
x1 0
1x1
0 1x
, D 4=
x1 0 0
1x1 0
0 1x1
0 0 1x
, D 5=···.
Ifx= 2 cosθ, prove that det Dn=sin(n+1)θ
sinθ.
87. LetAandBbe square matrices of the same size. If I−AB is invertible, prove
thatI−BAis also invertible by exhibiting a formula for its inverse.
88. Assume/summationtextanconverges, where an≥0 . Does the series
/summationdisplay√anan+1
also converge? Proof or counterexample.
89. LetAbe a square matrix.
(a) Prove that AA∗is self-adjoint.
(b) IsAA∗always equal to A∗A? Proof or counterexample.
90. Show that C[0,1] is a direct sum of the space V1spanned by e1(x) =xande2(x) =
x4, and the subspace V2of all functions ϕ(x) such that
0 =/integraldisplay1
0xϕ(x)dx, 0 =/integraldisplay1
0x4ϕ(x)dx.
[Hint: Show that if f∈[0,1] , there are unique constants aandbsuch thatg(x)≡
f(x)−[ax+bx4] belongs to V2].
91. LetV1be the linear space of all complex-valued analytic functions in the open unit
disc, that is, V1consists of all complex-valued functions fof the complex variable
zwhich have convergent power series expansions
f(z) =∞/summationdisplay
0anzn
in the open disc, |z|<1 .
LetV2be the linear space of all sequences of complex numbers ( a0,a1,a2,···) with
the natural definition of addition and multiplication by constants.
DefineL:V1→V2by the rule
Lf= (a0,a1,a2,···),
where theaj’s are the Taylor series coefficients of f. Answer the following questions
with a proof or counterexample.
400 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
(a) IsLinjective?
(b) IsLsurjective?
(c) Is/lscript2contained in R(L) ? (Note:/lscript2is the subspace of V2such that
∞/summationdisplay
k=0|ak|2<∞).
92. Do the following series converge or diverge?
(a)∞/summationdisplay
n=1/radicalbig
1 + 1/n, (b)∞/summationdisplay
n=1(/radicalbig
1 + 1/n2−1).
93. Consider the set of four operators {T1,T2,T3,T4}defined as follows on the set of
square invertible matrices.
T1A=A, T 2A=A−1
T3A=A∗, T 4A= (A−1)∗.
Show that this set of four operators forms a commutative group with the group oper-
ation being ordinary operator multiplication.
94. LetSn=a1−a2+a3−a4+a5− ··· . If 0<akand theak’s are increasing, prove
that|SN| ≤aN.
95. The Monge-Ampere equation is uxxuyy−u2
xy= 0 . Show that it is satisfied by any
u(x,y)∈C2of the form u(x,y) =ϕ(ax+by) , whereaandbare constants.
96. (a) Consider the differential operator
Lu=u/prime/prime−4u
(i) Find a basis for the nullspace of L.
(ii) Find a particular solution of Lu=e2x+1.
(iii) Find the general solution of Lu=e2x+1.
(b) Consider the differential operator
Lu=u/prime/prime+ 4u
Repeat part (a), only here use Lu=f, wheref(x) = sec 2x.
97. Find the general solution for each of the following
(a) 2u/prime/prime+ 5u/prime−3u= 0
(b)u/prime/prime−6u/prime+ 9u= 0
(c)u/prime/prime−4u/prime+ 5u= 0
401
98. Find the first four non-zero terms in the series solution of
4x2u/prime/prime−4xu/prime+ (3−4x2)u= 0
corresponding to the largest root of the indicial equation. Where does the series
converge?
99. Find the complete solution of each of the following equations valid near x= 0 by
using power series.
(a)x2u/prime/prime+xu/prime−(x2−1
4)u= 0
(b)u/prime/prime+xu/prime−u= 0 (only first five non-zero terms)
[Answers:
(a)u(x) =Ax−1/2∞/summationdisplay
k=0x2k
(2k)!+Bx1/2∞/summationdisplay
k=0x2k
(2k+ 1)!,
(b)u(x) =Ax+B(1 +x2
2!−x4
4!+3x6
6!−15x8
8!+···) ].
100. Consider the matrix
A=
−1−4−12 0
1 3 6 0
0 0 −1 0
0−4−12 1
.
(a) Compute det A.
(b) Compute A−1.
(c) SolveAX=bwhereb= (1,2,3,−1) .
101. True or false. Justify your response if you believe the statement is false (a counterex-
ample is adequate).
(a) The set A={X∈R3:x1= 2}is a linear subspace ofR3.
(b) The vectors X1= (2,4) andX2= (−2,4)span R2.
(c) The vectors X1= (1,2,3), X 2= (−7,3,2), X 3= (2,−1,1) , andX4= (π,e,5)
are linearly independent .
(d) The set A={u∈C[0,1]:u(x) =a1x+a2ex}is an infinite dimensional
subspace of C[0,1] .
(e) The functions f1(x) =xandf2(x) =exarelinearly dependent functions in
C[0,1] .
(f) If {e1,e2,...,e n}are an orthonormal set of vectors in E8, thenn≤7 .
(g) The vector Y= (1,2,3) is orthogonal to the subspace of E3spanned by e1=
(0,3,−2) ande2= (−1,−1,1) .
402 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
(h) The elements of the set
A={u∈C2[0,10]:u/prime/prime+xu/prime−3u= 6x}
can be represented as u(x) = ˜u(x) +x3, where
˜u∈S={u∈C2[0,10]:u/prime/prime+xu−3x= 0}.
(i) The set of vectors e1= (1
3,0,2
3,−2
3), e 2= (0,0,1√
2,1√
2),ande3= (8
9,3
9,−2
9,2
9)
constitute a complete orthonormal basis for E4.
(j) In the vector space of bounded functions f(x), x∈[0,1] , the functions
f1(x) = 1, f 2(x) =/braceleftbigg1,0≤x≤1
2,
0,1
2<x≤1f3(x) =/braceleftbigg0,0≤x≤1
2
1,1
2<x≤1
are linearly independent .
(k) The function f(x) =|x|can be represented by a convergent Taylor series about
the pointx0= 0 .
(l) The function f(x) =x2−x73can be represented by a convergent Taylor series
about the point x0=−1 .
(m) The function f(x) =|x|can be represented by a convergent Taylor series about
the pointx0=−1 .
(n) The plane of all points ( x1,x2,x3,x4)∈E4such that
2x1−4x2+ 6x3−5x4= 7
is perpendicular to the vector (2 ,−4,6,−5) .
(o) Ife1= (3
5,4
5) ande2= (4
5,−3
5) , thenX= (−1,2) can be written as X=
2e1−e2.
(p) The set of all integers (positive, negative, and zero) is a field.
(q) Consider the infinite series
∞/summationdisplay
k=0ak.
If lim
k→0|ak|= 0 , then the series must converge .
(r) Let {an}be a sequence of rational numbers. If this sequence converges to a,
then the limiting value, a, must be a rational number too.
(s) The equation x6+ 3 = 0 , where xis an element of an ordered field, has no
solutions .
(t) It is possible to write√
iin the form a+ib, whereaandbare real numbers.
(Herei=√−1 , of course).
(u) Letanbe a sequence of complex numbers. If the sequence of absolute values,
|an|, converges, then the sequence anmust converge .
(v) If
∞/summationdisplay
k=0akzk
converges at the point z= 3 , then it must converge atz= 1 +i.
403
(w) The linear subspace A={p∈P7:p(x) =a1x+a2x5}is afive dimensional
subspace of P7.
(x) The linear subspace A={u∈C[−1,1]:u(x) =a1x+a2x5}is an infinite
dimensional subspace of C[−1,1] .
(y) There is a number αsuch that the vectors X= (1,1,1) andY= (1,α,α2)
form a basis forR3.
(z) The operator T:C2→C1defined for u∈C2byTu=u/prime−7uis alinear
operator.
102. (a) The operator T:C[0,1]→Rdefined for u∈C[0,1] by
Tu=/integraldisplay1
0|u(x)|dx
is alinear operator.
(b) The sequence (1 + i)nconverges to√
2 .
(c) The series
∞/summationdisplay
k=1k+ 1
2k+ 1=2
3+3
5+4
7+5
9+···
converges .
(d) Iftis real, then/vextendsingle/vextendsingleeit/vextendsingle/vextendsingle= 1 .
(e) LetV1andV2be linear spaces and let the operator TmapV1intoV2. If
T0 = 0 , then Tis alinear operator.
(f) The operator T:C∞[−7,13]→C∞[−7,13] defined by
Tu=udu
dx
islinear .
(g) The operator T:C[0,13]→C[0,13] defined by
(Tu)(x) =/integraldisplayx
0u(t) sintdt, x ∈[0,13]
islinear
(h) In the scalar product space L2[0,1] , the functions fandgwhose graphs are
a figure goes here
areorthogonal .
(i) LetLbe a linear operator. IF LX 1=YandLX 2=Y, whereX1/negationslash=X2, then
the solution of the homogeneous equation LX= 0 is not unique .
(j) LetLbe a linear operator. If X1andX2are solutions of LX= 0 , then
3X1−7X2isalsoa solution of LX= 0 .
404 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
(k) Lete1= (1,1) ande2= (0,1) , and let the linear operator Lwhich maps R2
intoR3satisfy
Le1= (1,2,3), Le 2= (1,−2,−1).
ThenL(2,3) = (1,1,1).
(l) In the space L2[0,1] , iffisorthogonal to the function x2, then either f≡0
orelsefmust be positive somewhere in [0,1] .
(m) IfF/prime(X) = (2,3,4) for allX∈E3, thenFis an affine mapping.
(n) Iff:E3→E1is such that f: (1,0,0)→1 andf: (0,4,0)→2 , there is a
pointZ∈E3such that /bardblf/prime(Z)/bardbl ≥1
5.
(o) LetAandBbe square matrices with det A= 7 and det B= 3 . Then
detAB= 10. det(A+B) = 10.
(p) IfA:R3→R3is given by
A=/parenleftbigg2 3 1
1 9 2/parenrightbigg
,
then dim N(A) = 2 .
(q) The function f(x,y,z ) = 9 + 3x+ 4y−7zdoes not take on its maximum value.
(r) If the function u(x) has two derivatives in some neighborhood of x= 0 , and
satisfies the differential equation
9x2u/prime/prime−28u= 0,
thenu(0) = 0 .
(s) There are constantsa,bandcsuch that the function u(x) =ex+ 2e2x−e−x
is a solution of
au/prime/prime+bu/prime+cu= 0.
(t) The vector ( xy,x)isthe derivative of some real-valued function f(x,y) .
(u) The vector ( y,x) isnotthe derivative of some real-valued function f(x,y) .
(v) Given any q×pmatrixA= ((aij(X))) , where X= (x1,···,xp) and where
the elements aij(X) are sufficiently differentiable functions, then there is a map
F:Rp→Rqsuch thatF/prime(X) =A.
(w) IfAis a square matrix and A2=A, thenA=I.
(x) IfAis a square matrix and A2= 0 , thenA= 0 .
(y) IfAis a square matrix and det A/negationslash= 0 , thenA2=Aif and only if A=I.
(z) IfX,Y , andZare three linearly independent vectors, then X+Y, Y +Z,
andX+Zare also linearly independent .
103. Define L:P2→P2as follows: if p∈P2
Lp= (x+ 1)dp
dx
(a) Find the matrix eLerepresenting the operator Lwith respect to the bases
e1= 1, e2=x1, e3=x2forP2.
405
(b) IsLan invertible operator? Why?
(c) Find dim R(L) and dim N(L) .
104. Let
A=/parenleftBigg
1
2−√
3
2 √
3
21
2/parenrightBigg
, B =/parenleftbigg5√
3√
3 3/parenrightbigg
.
(a) Compute AA∗,ABA∗, and (ABA∗)100.
(b) How could you use the result of part (a) to compute B100?
105. Consider the following system of three equations as a linear map L:R2→R3
x1+x2=y1
4x1+x2=y2
x1−2x2=y3
(a) Find a basis for N(L∗) .
(b) Use the result of part a) to determine the value(s) of αsuch thatY= (1,2,α)
is inR(L) .
106. Find the unique solution to each of the following initial value problems.
(a)u/prime/prime+u/prime−2u= 0, u (0) = 3, u/prime(0) = 0
(b)u/prime/prime+ 4u/prime+ 4u= 0, u (0) = 1 u/prime(0) = −1
(c)u/prime/prime−2u/prime+ 5u= 0, u (0) = 2, u/prime(0) = 2
107. Consider the special second order inhomogeneous constant coefficient O.D.E. Lu=f,
where
Lu≡u/prime/prime−4u,
and where fis assumed to be a suitably differentiable function which is periodic with
period 2π, f(x+ 2π) =f(x) .
(a) Expand fin its Fourier series and seek a candidate, u, for a solution of Lu=f
as a Fourier series, showing how the Fourier coefficients of uare determined by
the Fourier coefficients of f.
(b) Apply the above procedure to the trivial example where
f(x) = sin 3x−4 cos 17x+ 3 sin 36x.
108. (a) Find the directional derivative of the function
f(x,y) = 2−x+xy
at the point (0 ,6) in the direction (3 ,−4) by using the definition of the direc-
tional derivative as a limit. Check your answer by using the short method.
406 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
(b) Repeat part (a) for f(x,y) = 1−3y+xy.
109. Find and classify the critical points of the following functions.
(a)f(x,y) =x3+y2−3x−2y+ 2
(b)f(x,y) =x2−4x+y2−2y+ 6
(c)f(x,y) = (x2+y2)2−8y2
(d)f(x,y) = (x2−y2)2−8y2
(e)f(x,y) = (x2−y2)2
(f)f(x,y) =x2−2xy+1
3y3−3y
110. Consider the function x3+y2−3x−2y+ 2 . At the point (2 ,1) find the direction
in which the directional derivative is greatest. Find the direction where it is least.
111. Letf:E2→Ebe a suitably differentiable function and let X(t) be the equation
of a smooth curve CinE2on whichfis identically constant, say, f(X(t))≡4 .
Show that on this curve, f/primeis perpendicular to the velocity vector X/prime(t) . [Hint: Do
something to ϕ(t) =f(X(t)) . The proof takes but one line.].
112. Consider the following statements concerning a function f:En→E.
(A)fis continuous.
(B)fhas a total derivative everywhere.
(C)fhas first order partial derivatives everywhere.
(D)fhas a total derivative everywhere which is continuous everywhere.
(E)fhas first order partial derivatives everywhere and they are continuous functions
everywhere.
(F)fis an affine function.
(G)f/prime≡0 .
(a) Which of these statements always imply which others. A sample (possibly incor-
rect) answer might look like
(A)⇒B,F,···
(B)⇒A,···
(b) Find examples illustrating each case where a given statement does not imply
another (the Exercises, pp. 588-95, contain the required examples).
113. Solve the following ordinary differential equations subject to the given auxiliary con-
ditions
(a)u/prime/prime−u/prime−6u= 0, u (0) = 0, u/prime(0) = 5
(b)xu/prime+u=ex−1, u (1) = 2
(c)u/prime/prime−6u/prime+ 10u= 0 , general solution.
407
114. (a) If u(x,y,t ) =xexy+t2, whilex= 1−t3andy= logt2, then letw(t) =
u(x(t), y(t), t) . Finddw
dtatt= 1 .
(b) IfF:E3→E2andG:E2→E2are defined by
F(X) =/parenleftbigg2x1−x2
2+x2x3+ 1
x2
1−x2
3+x2/parenrightbigg
, G (Y) =/parenleftbiggy1+y2siny1
−3y1y2+y2
2/parenrightbigg
,
(i) Why doesn’t F◦Gmake sense?
(ii) Compute [ G◦F]/primeat the point X0= (0,1,0) .
115. LetF=E2→E2andG:E3→E2be defined by
F(w,z) =/parenleftBigg
ew+z2
ez+w2/parenrightBigg
G(r,s,t ) =/parenleftbiggr+s2+t3
s+t2+r3/parenrightbigg
.
(a) FindF/primeandG/prime.
(b) Which of F◦GorG◦Fmakes sense?
(c) IfG◦Fmakes sense, compute ( G◦F)/primeat (−1,−1) .
(d) IfF◦Gmakes sense, compute ( F◦G)/primeat (−1,0,0) .
116. LetF:X→YandG:Y→Zbe defined by
F:/braceleftbiggy1=x2−ex1+2x2
y2=x1x2, G :/braceleftbiggw1=y2+y2siny1
w2= (y1+y2)2
(a) Compute F/primeatX0= (−2,1) andG/primeatY0=F(X0) .
(b) LetH=G◦F. Compute H/primeatX0= (−2,1) .
117. Consider the map F:E2→E3defined by
F:
f1(x,y) =y+ex−y
f2(x,y) = sin(x−2y+ 1)
f3(x,y) =x−3x2+y2
(a) Find the tangent map at the point X0= (1,1) .
(b) Use the result of part (a) to evaluate approximately FatX1= (1.1,.9) .
118. Consider the system of O.D.E.’s
u/prime=αu
v/prime=αu−βv,
whereαandβare constants. If u(0) =Aandv(0) =B,
(a) Findu(t) .
(b) Findv(t) (remember to consider the case α=βseparately).
408 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
119. (a) Consider the homogeneous equation
u/prime/prime+a(t)u= 0,
wherea(t) is continuous and periodic with period P, soa(t+P) =a(t) .
(i) Ifa(t)≡1 , show that there is no non-trivial periodic solution by merely
solving the equation.
(ii) Ifa(t) = cost, show (again by solving the equation) that there is a periodic
solutionu(t) with period 2 π.
(iii) In general, if u(t) is a solution, not necessarily periodic, show that v(t)≡
u(t+P) is also a solution.
(iv) Show that the homogeneous equation has a non-trivial periodic solution of
periodPif and only if/integraldisplayP
0a(t)dt= 0
(b) Consider the inhomogeneous equation
u+a(t)u=f(t),
where both a(t) andf(t) are continuous and periodic with period P.
(i) If/integraldisplayP
0a(t)dt=K/negationslash= 0 , show that the inhomogeneous equation has one and
only one periodic solution with period P.
(ii) If/integraldisplayP
0a(t)dt= 0 , find a necessary condition on fthat the inhomogeneous
equation have a periodic solution with period P.
120. Letf:En→Ebe a differentiable function and denote the directional derivative in
the direction of the unit vector ebyDef. Prove that D−ef=−Def.
121. Letf:En→Ebe of the form f(a1x1+...+anxn) . Writeα= (a1,···,an) and
β= (b1,···,bn) . Ifβis perpendicular to α, prove that β⊥f/prime.
122. LetRdenote the rectangle 0 ≤x1<2π,0≤x2<2π, and define the map
f:R→E1by
f(x1,x2) = (3 + 2 cos x2) sinx1
Find and classify the critical points of f. (This function is the height function of a
torus with major radius 3 and minor radius 2).
123. Consider the constant coefficient differential operator
Lu≡au/prime/prime+bu/prime+cu, (a,b,c real, a/negationslash= 0.)
Letλ1andλ2denote the roots of the characteristic polynomial p(λ) =aλ2+bλ+c.
(a) Ifλ1/negationslash=λ2, find a formula for a particular solution of Lu=f.
[Answer:up(x) =1
λ1−λ2/integraldisplayx
[eλ1(x−t)−e−λ2(x−t)]f(t)dt.
409
(b) Ifλ1is complex, say, λ1=α+iβ, thenλ2=¯λ1=α−iβ. Show that in this
case, the above formula simplifies to
up(x) =1
β/integraldisplayx
eα(x−t)sinβ(x−t)f(t)dt.
(c) Ifλ1=λ2, find a formula for a particular solution of Lu=f.
[Answer:up(x) =/integraldisplayx
(x−t)eλ1(x−t)f(t)dt].
124. Consider/integraldisplay/integraldisplay
DfdA whereDis the triangle with vertices at ( −1,1),(0,0) , and
(3,1) .
(a) Set up the iterated integrals in two ways.
(b) Evaluate one of the integrals in (a) for the integrand
f(x,y) = (x+y)2.
125. When a double integral was set up for the mass Mof a certain plate with density
f(x,y) , the following sum of iterated integrals was obtained
M=/integraldisplay2
1(/integraldisplayx3
xf(x,y)dy)dx+/integraldisplay8
2(/integraldisplay8
xf(x,y)dy)dx.
(a) Sketch the domain of integration and express Mas an iterated integral in which
the order of integration is reversed.
(b) Evaluate Mif
f(x,y) =/radicalbiggx
y.
126. Evaluate/integraldisplay1
0/integraldisplay1
0xydxdy .
127. It is difficult to evaluate the integral I=/integraldisplay/integraldisplay
DfdA , wheref(x,y) =1
1+x+y2andD
is the indicated rectangle. However, you can show that (trivially)
1
3<I <3
2,
and, with a bit more effort but the same method, that
1
2<I <3
2.
Please do so.
410 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
128. Consider the integral I=/integraltext/integraltext
DfdA , where
f(x,h) =3
8 +/radicalbig
x4+y4
andDis the domain inside the curve x4+y4= 16 . Show that
2√
2<I < 6.
[Hint: Show that1
4<f <3
8inD. Then approximate the area of Dby an inscribed
and circumscribed square. For the record, it turns out that I=3A
4ln(3
2) , whereA
is the
area =2
πΓ(1
4)2].
129. (a) Find the derivative matrix for the following mappings Y=F(X) at the given
pointX0.
(i)F:/braceleftbiggy1=x2
1+ sinx1x2
y2=x2
2+ cosx1x2atX0= (0,0)
(ii)F:
y1=x2
1+x3ex2−x3
2
y2=x1−3x2+x1logx3
y3=x2+x3
y4=fx1x2x3atX0= (2,0,1)
(b) Find the equation of the tangent plane to the above surfaces at the given point.
130. Consider the following map Ffrom E2→E2, the familiar change of variables from
polar to rectangular coordinates.
F:/braceleftbiggy1=x1cosx2
y2=x1sinx2
(a) Find the images of
(i) the semi-infinite strip 1 ≤x1<∞,0≤x2≤π
2.
(ii) the semi-infinite strip 0 ≤x1<∞,0≤x2≤3π
2.
(b) Compute F/primeand detF/prime.
131. Given that up(x) =e3x+e−2x−2ex/2is a solution of
au/prime/prime+bu/prime+cu=e3x,
find the constants a,b, andc.
132. Evaluate the determinants of the following matrices.
(a)
1 1 1 1
1−1 1 −1
1−1−1 1
1 1 −1−1
(b)
1 1y z t
2x z t y
w2x20 0 0
w3x30 0 0
w4x40 0 0
.
411
133. For what value(s) of xis the following matrix invertible?
1 1 1 1
1 2 2223
1 3 3233
1x x2x3
(Hint: Observe that the determinant is a cubic polynomial all of whose roots are
obvious).
134. Letf(x) =n/summationdisplay
k=1aksinkx√πandg(x) =n/summationdisplay
r=1bsinrx√π.
Bydirect integration prove that
/integraldisplayπ
−πf(x)g(x)dx=n/summationdisplay
j=1ajbj.
After you are done, compare with Theorem 15, page 206-7 and its proof.
135. Let
A=
1 1 0 0
1 0 0 0
0 0 0 −1
0 0 1 0
(a) Find det A.
(b) FindA−1.
(c) SolveAX=Y, whereY=
2
2
1
3
.
(d) LetS={u:u(x) =aex+bex+csinx+dcosx}, wherea,b,c , anddare any
real numbers, and define a linear operator L:S→Sby the rule
Lu≡u/prime/prime−u/prime+u.
Find the matrix eLeforLwith respect to the following basis for S:
e1(x) =xex, e2(x) =x, e 3(x) = sinx, e 4(x) = cosx.
(e) Use the above results to find a solution of
Lu= 2xex+ 2ex+ sinx+ 3 cosx.
136. Letu1andu2be solutions of the homogeneous equation
Lu≡a2(x)u/prime/prime+a1(x)u/prime+a0(x)u= 0.
412 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
(a) Show that W(x)≡W(u1,u2)(x) , the Wronskian of u1andu2satisfies the
differential equation
W/prime=−a1(x)
a2(x)W.
(b) Find the equation of (a) for the particular operator
Lu≡x2u/prime/prime−2xu/prime+ 2u
and solve it for Wunder the condition that W(1) = 1 .
(c) Given that u1(x) =xis a solution of Lu= 0 for the operator of part (b), use
the result of (b) to show that if u2is another solution of Lu= 0 , then u2
satisfies the equation
u/prime
2−1
xu2=x,
provided that W(x,u 2)(1) = 1 .
(d) Solve the equation of part (c) under the assumption that u2(1) = 1 , and thus
find a second independent solution of the equation Lu= 0 for the operator of
part (b).
(e) Generalize the idea of parts (c) - (d) by stating and proving some theorem.
137. Here are some linear transformations defined in terms of matrices. In each case,
describe geometrically what the transformation does, by computing the images of the
three parallelograms
Q1: with vertices at (0 ,0),(2,0),(3,1),(1,1).
Q2: with vertices at (1 ,2),(3,2),(4,3),(2,3).
Q3: with vertices at (1 ,0),(0,2),(−1,0),(0,−2).
(a)Diagonal Maps (Stretchings)
L1=/parenleftbigg3 0
0 1/parenrightbigg
, L 2=/parenleftbigga0
0 1/parenrightbigg
, L 3=/parenleftbigg−4 0
0 6/parenrightbigg
,
L4=/parenleftbigg1 0
0−1/parenrightbigg
, L 5=/parenleftbigg−1 0
0−1/parenrightbigg
, L 6=/parenleftbigg−2 0
0 0/parenrightbigg
,
L7=/parenleftbigga0
0a/parenrightbigg
, L 8=/parenleftbigg1 0
0b/parenrightbigg
, L 9=/parenleftbigga0
0b/parenrightbigg
,
(Remember to consider negative values of aandb).
(b)Maps with 0 on the diagonal .
L1=/parenleftbigg0 1
0 0/parenrightbigg
, L 2=/parenleftbigg0a
0 0/parenrightbigg
, L 3=/parenleftbigg0 0
−1 0/parenrightbigg
,
L4=/parenleftbigg0 1
1 0/parenrightbigg
, L 5=/parenleftbigg0a
1 0/parenrightbigg
, L 6=/parenleftbigg0a
b0/parenrightbigg
.
413
(c)Upper Triangular Matrices .
L1=/parenleftbigg1 1
0 1/parenrightbigg
, L 2=/parenleftbigg1−1
0 1/parenrightbigg
, L 3=/parenleftbigg1 1
0−1/parenrightbigg
,
L4=/parenleftbigg1a
0 1/parenrightbigg
, L 5=/parenleftbigg−1−1
0 0/parenrightbigg
, L 6=/parenleftbigga1
0b/parenrightbigg
.
(d)Orthogonal Matrices (Rotations and Reflections).
L1=/parenleftbigg0−1
1 0/parenrightbigg
L3=/parenleftbigg3
54
5
L2=4
5−3
5/parenrightbigg
L4=/parenleftBigg
−1√
21√
21√
21√
2/parenrightBigg
.
138. Letaandbbe real numbers such that a2+b2= 1 . Let
S=/parenleftbigga2−b22ab
2ab b2−a2/parenrightbigg
, P =/parenleftbigga2ab
ab b2/parenrightbigg
,
and lete1= (a,b), e 2= (−b,a) , soe1⊥e2. Show that
(a)Se1=e2, Pe 1=e1
(b)Se2=−e2, Pe 2= 0
(c)S2=I, P2=P
(d) Show that Scan be interpreted as the reflection which leaves the line through
e1fixed, and that Pcan be interpreted as the projection onto the line through
e1parallel to e2.
139. (a) Consider the following relation defined on the set of allintegers:nRmifnand
mare both even integers. Verify that this relation is symmetric and transitive -
but not reflexive (since, for example, 1 R/1 ).
(b) Let Rbe a symmetric and transitive relation defined on a set A. If, given any
elementxinA, there is some element yrelated to it, xRy, prove that the
relation Ris also reflexive. (The example in part (a) shows that the assertion
will be false if some element is related to no others).
140. Letanbe a decreasing sequence of positive real numbers which satisfy an−1an+1≤
a2
n. If/summationtexta1/n
nconverges, prove that/summationtextan
an−1converges too. [Hint: Show that
(an/an−1)1/n≤an].
141. (a) Prove that the series/summationtextanznand/summationtexta2
nznhave the same radii of convergence.
(b) Prove that the series/summationtextanznand/summationtext(an)kzn, wherek>0 , have the same radii
of convergence.
142. LetVbe a linear space and Lan invertible linear map, L:V→V. If{e1,...,e n}
is a basis for V, prove that its image {Le1,Le 2,...,Le n}is also a basis for V.
143. LetHbe an inner product space and Ran orthogonal transformation, R:H→
H. If{e1,...,e n}is a complete orthonormal set for H, prove that its image
{Re1,...,Re n}is also a complete orthonormal set for H.
414 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
144. (a) Let Rbe an orthogonal matrix and let ρ1andρ2be any two of its column
vectors. Prove that ρ1⊥ρ2. Prove that any two rows of an orthogonal matrix
are also orthogonal to each other.
(b) Conversely, let Abe a square matrix whose column vectors are orthogonal. Must
Abe an orthogonal matrix? Proof or counterexample.
145. LetHbe an inner product space and Athe subspace of Hspanned by the vectors
X1,...,X n. The Gram determinant of those vectors is defined as
G(X1,...,X n) =/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/angbracketleftX1, X 1/angbracketright ··· /angbracketleftXn, X 1/angbracketright
/angbracketleftX1, X 2/angbracketright ·
· ·
· ·
· ·
/angbracketleftX1, Xn/angbracketright ··· /angbracketleftXn, Xn/angbracketright/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle/vextendsingle
(a) Prove that X1,···,Xnare linearly dependent if and only if G(X1,···,Xn) = 0 .
[Suggestion: If Z∈A, thenZ=a1X1+···anXn, where the scalars a1,···,an
are to be found. This can be done in two ways, by Theorem 31, page 428, or by
solving the nequations
/angbracketleftZ, X 1/angbracketright=a1/angbracketleftX1, X 1/angbracketright+···+an/angbracketleftXn, X 1/angbracketright
·
·
·
·
/angbracketleftZ, X n/angbracketright=a1/angbracketleftX1, Xn/angbracketright+···+an/angbracketleftXn, Xn/angbracketright
which are obtained from /angbracketleftZ, X j/angbracketright=/angbracketleftaiX1+···+anXn, Xj/angbracketright. Couple both
methods to prove the result].
(b) IfX1,···,Xnare an orthogonal set of vectors, compute G(X1,···,Xn) .
(c) IfY∈H, prove that the distance of Yfrom the subspace A,/bardblY−PAY/bardbl=δ,
is given by the formula
δ2=/bardblY−PAY/bardbl2=G(Y,X 1,...,X n)
G(X1,...,X n).
[Suggestion: Observe that δ2=/bardblY−PAY/bardbl2=/angbracketleftY−PAY, Y/angbracketrightand that /angbracketleftPAY, Y/angbracketright=
an/angbracketleftX1, Y/angbracketright+···+an/angbracketleftXn, Y/angbracketright. Now write PAYasZ, use thenequations in
a) and the one equation δ2=/angbracketleftY, Y/angbracketright −a1/angbracketleftX1, Y/angbracketright − ··· −an/angbracketleftXn, Y/angbracketrightto solve for
δ2by using Cramer’s rule].
(d) Use the fact that G(X1) =/angbracketleftX1, X 1/angbracketrightto prove the Gram determinant of lin-
early independent vectors is always positive . In particular, deduce the Cauchy -
Schwarz inequality from G(X1,X2)≥0 .
(e) InL2[0,1] , letX1= 1+x, andX2=x3. Compute G(X1,X2) . LetY= 2−x4
and compute /bardblY−PAY/bardbl, whereAis the subspace spanned by X1andX2.
415
(f) (Muntz) In L2[0,1] , letAn= span {xj1,xj2,···,xjn}wherej1,···,jnare
distinct positive integers. Let Y=xk, wherekis a positive integer by not one
of thej’s. Prove that lim
n→∞/bardblY−PAnY/bardbl= 0 if and only if/summationdisplay1
jndiverges.
146. (a) Use Theorem 17, page 217 to find linear polynomials PandQsuch that,
respectively,
(i)/integraldisplay1
−1[x2−P(x)]2dxis minimized,
(ii)/integraldisplay1
0[x2−Q(x)]2dxis minimized.
(b) WriteP(x) =a+bxand use calculus to again find the values of aandbsuch
that/integraldisplay1
−1[x2−P(x)]2dx
is minimized.
147. LetZ= (1,1,1,1,1)∈E5and letAbe the subspace of E5spanned by X1=
(1,0,1,0,0), X 2= (1,0,0,−1,0) , andX3= (0,1,0,0,1) . Find /bardblZ−PAZ/bardbl.
148. Let Γ 0be a closed planar curve which encloses a convex region, and let Γ rbe the
“parallel” curve obtained by moving out a distance of ralong the outer normal.
(a) Discover a formula relating the arc length of Γ rto that of Γ 0. [Advise: Examine
the special cases of a circle, rectangle, and convex polygon].
(b) Prove the result you conjectured in part a).
149. The hypergeometric function F(a,b;c;x) is defined by the power series
F(a,b;c;x) = 1 +a·b
1·cx+a(a+ 1)b(b+ 1)
1·2·c(c+ 1)x2+a(a+ 1)(a+ 2)b(b+ 1)(b+ 2)
1·2·3c(c+ 1)(c+ 2)x2+···
(a) Show that the series converges for all |x|<1 .
(b) Show thatd
dxF(a,b;c;x) =ab
cF(a+ 1,b+ 1;c+ 1;x) .
(c) Show that
(i) (1 −x)n=F(−n,b;b;x)
(ii) (1 +x)n=F(−n,b;b;−x)
(iii) log(1 −x) =−xF(1,1; 2,x)
(iv) log(1+x
1−x) = 2xF(1
2,1;3
2;x2)
(v)ex= lim
b→∞F(1,b; 1;x/b)
(vi) cosx=F(1
2,−1
2;1
2,sin2x)
(vii) sin−1x=xF(1
2,1
2;3
2;x2)
(viii) tan−1x=xF(1
2,1;3
2;−x2)
416 CHAPTER 10. MISCELLANEOUS SUPPLEMENTARY PROBLEMS
(d) Show that Fsatisfies the hypergeometric differential equation
x(1−x)d2F
dx2+ [c−(a+b+ 1)x]dF
dx−abF = 0.
[This equation is essentially the most general one with three regular singular
points - in this case located at 0 ,1 , and ∞].
150. Let {e1,···,en}be a complete orthonormal set of Enand let {X1,···,Xn}be a
set of vectors which are close to the ej’s in the sense that
n/summationdisplay
j=1/bardblXj−ej/bardbl2<1.
Prove that the Xj’s are linearly independent. Give an example in E3of linearly
dependent vectors {X1,X2,X3}which satisfy
n/summationdisplay
j=1/bardblXj−ej/bardbl2= 1.
[In fact, one can prove that
dimA⊥≤n/summationdisplay
j=1/bardblXj−ej/bardbl2,
] whereA= span {X1,···,Xn}].
151. (a) Show that the function f(z) =ez, z∈C, is never zero.
(b) Scrutinize the proof of the Fundamental Theorem of Algebra (pp. 544-548) and
find where it breaks down if one attempts to extend it to prove that ezhas at
least one zero.