Linear Algebra Done Right, 2nd Ed - Sheldon Axler
PDF · 261 pages · 1.8 MB
Open PDF file
A published undergraduate textbook by Sheldon Axler (Springer, 1997), kept in the tensor products folder of the archive. It covers vector spaces, linear maps, polynomials, eigenvalues, inner-product spaces, the spectral theorem, operators on complex and real vector spaces, Jordan form, and trace and determinant. Determinants are deliberately left to the end. This is a reference copy, not Phil's own work.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
Linear Algebra
Done Right,
Second Edition
Sheldon Axler
Springer
Undergraduate Texts inMathematics
Editors
S.Axler
FW. Gehring
K.A. Ribet
Springer
New York
Berlin
Heidelberg
Barcelona
Hong Kong
London
Milan
Paris
Singapore
Tokyo
Undergraduate Texts inMathematics
Abbott: Understanding Analysis Childs: AConerete Introduction to
‘Anglin: Mathematics: AConcise History Higher Algebra. Second edition.
andPhilosophy. ‘Chung: Elementary Probability Theory
Readings inMathematics. with Stochastic Processes. Third
Anglin/Lambek: The Heritage of edition.Thales. Cox/Little/O'Shea: Ideals,Varieties,Readings inMathematics. andAlgorithms. Second edition.
Apostol: Introduction toAnalytic Croom: Basic Concepts ofAlgebraic
Number Theory. Second edition. Topology.
Armstrong: Basic Topology. Curtis: Linear Algebra: AnIntroductory
‘Armstrong: Groups andSymmetry Approach, Fourth edition.‘Axler:LinearAlgebraDoneRight. Devlin:TheJoyofSets:FundamentalsSecond edition. ofContemporary SetTheory.
Beardon: Limits: ANew Approach to Second edition.
Real Analysis. Dixmier: General Topology.
Bak/Newman: Complex Analysis. Driver: Why Math?
Second edition Ebbinghaus/Flum/Thomas:
Banchoff/Wermer: Linear Algebra Mathematical Logic. Second edition
‘Through Geometry. Second edition. Edgar: Measure, Topology, andFractal
Berberian: AFirst Course inReal GeometryAnalysis, Elaydi:AnIntroduction toDifferenceBix: Conics andCubies: A Equations. Second edition.
Concrete Introduction toAlgebraic Exner: AnAccompaniment toHigherCurves. Mathematics.Brémaud: An Introduction to Exner: Inside Calculus.
Probabilistic Modeling, Fine/Rosenberger: The Fundamental
Bressoud: Factorization and Primality Theory ofAlgebraTesting Fischer:Intermediate RealAnalysis.Bressoud: Second Year Calculus Flanigan/Kazdan: Calculus Two: Linear
Readings inMathematics. andNonlinear Functions. Second
Brickman: Mathematical Introduction edition
toLinear Programming andGame Fleming: Functions ofSeveral Variables‘Theory. Secondedition.Browder: Mathematical Analysis: Foulds: Combinatorial Optimization for
‘AnIntroduction Undergraduates.
Buchmann: Introduction to Foulds: Optimization Techniques: AnCryptography. Introduction.Buskes/van Rooij: Topological Spaces: Franklin: Methods ofMathematical
From Distance toNeighborhood. Economies,
Callahan: TheGeometry ofSpacetime: Frazier: AnIntroduction toWavelets
‘AnIntroduction toSpecial andGeneral ‘Through Linear Algebra,Relavitity Gamelin:ComplexAnalysis.Carter/van Brunt: The Lebesgue Gordon: Discrete Probability
Stieltjes Integral: APractical Hairer/Wanner: Analysis byItsHistoryIntroduction ReadingsinMathematics.Cederberg: ACourse inModem Halmos: Finite-Dimensional Vector
‘Geometries. Second edition. ‘Spaces. Second edition
Sheldon Axler
Linear Algebra
Done Right
Second Edition
Springer
Sheldon Axter
Mathematics Department
SanFrancisco State University
‘San Francisco, CA 94132
USA
Editorial Board
S.Axler F.W. Gehring
Mathematics Department Mathematics Department
SanFrancisco State University East Hall
SanFrancisco, CA94132 University ofMichiganUSA ‘AnnArbor,MI48109-1109 USA
K.A. Ribet
Mathematics Department
University ofCalifornia atBerkeley
Berkeley, CA94720-3840
USA
Mathematics SubjectClassification (1991).15.01
LibraryofCongress Cataloging-in-Publication Data ‘Axler, Sheldon Jay
Linear algebra done right Sheldon Axler.~ 2nded.
.ci. (Undergraduate texts inmathematics)
Includes index.ISBN637-98250-0 (ak,paper.—ISBN0-387-98258-2 (pbk"Algebra Linea. 1,Title. I.Series.
QAiss.A96 1997
312'5--de20 97-1664
©1997, 1996 Springer-Verlag. New York, Ine
Allrights reserved. This work may notbetranslated orcopied iwhole orinpart without the
‘written permission ofthepublisher (Springer-Verlag New York, Inc. 175 Fifth Avenue, New
York, NY10010, USA), except forbrief excerpis inconnection with reviews orscholarly
analysis, Use inconnection with any form ofinformation storage and retrieval, electronic ad-
sptation, computer software, orbysimilar ordissimilar methodology now known orhereafter
Geveloped isforbidden.‘Theuseofgeneraldescriptive names,tradenames,trademarks, etc,inthispublication, eveniftheformer arenotespecially identified, isnot tobetaken asasign that such names, asunder-
stood bytheTrade Marks andMerchandise Marks Act, may accordingly beused freely byany-
ISBN 0-387-98259-0 (hardcover) SPIN 10629393,
ISBN 0:387.98258-2 (softcover) SPIN 10794473,
Springer-Verlag New York Berlin Heidelberg‘AmemberofBertelsmannSpringer ScienceBusinessMediaGmbH
Contents
Preface to the Instructor ix
Preface to the Student xiii
Acknowledgments xv
Chapter 1
Vector Spaces 1
Complex Numbers .......................... 2
Definition of Vector Space ...................... 4
Properties of Vector Spaces ..................... 11
Subspaces............................... 13
Sums and Direct Sums ........................ 14
Exercises................................ 19
Chapter 2
Finite-Dimensional Vector Spaces 21
Span and Linear Independence ................... 22
Bases.................................. 27
Dimension............................... 31
Exercises................................ 35
Chapter 3
Linear Maps 37
Definitions and Examples ...................... 38
Null Spaces and Ranges ....................... 41
The Matrix of a Linear Map ..................... 48
Invertibility .............................. 53
Exercises................................ 59
v
vi Contents
Chapter 4
Polynomials 63
Degree................................. 64
Complex Coefficients ........................ 67
Real Coefficients ........................... 69
Exercises................................ 73
Chapter 5
Eigenvalues and Eigenvectors 75
Invariant Subspaces ......................... 76
Polynomials Applied to Operators ................. 80
Upper-Triangular Matrices ..................... 81
Diagonal Matrices ........................... 87
Invariant Subspaces on Real Vector Spaces ........... 91
Exercises................................ 94
Chapter 6
Inner-Product Spaces 97
Inner Products ............................. 98
Norms................................. 102
Orthonormal Bases .......................... 106
Orthogonal Projections and Minimization Problems ...... 111
Linear Functionals and Adjoints .................. 117
Exercises................................ 122
Chapter 7
Operators on Inner-Product Spaces 127
Self-Adjoint and Normal Operators ................ 128
The Spectral Theorem ........................ 132
Normal Operators on Real Inner-Product Spaces ........ 138
Positive Operators .......................... 144
Isometries............................... 147
Polar and Singular-Value Decompositions ............ 152
Exercises................................ 158
Chapter 8
Operators on Complex Vector Spaces 163
Generalized Eigenvectors ...................... 164
The Characteristic Polynomial ................... 168
Decomposition of an Operator ................... 173
Contents vii
Square Roots .............................. 177
The Minimal Polynomial ....................... 179
Jordan Form .............................. 183
Exercises................................ 188
Chapter 9
Operators on Real Vector Spaces 193
Eigenvalues of Square Matrices ................... 194
Block Upper-Triangular Matrices .................. 195
The Characteristic Polynomial ................... 198
Exercises................................ 210
Chapter 10
Trace and Determinant 213
Change of Basis ............................ 214
Trace.................................. 216
Determinant of an Operator .................... 222
Determinant of a Matrix ....................... 225
Volume................................. 236
Exercises................................ 244
Symbol Index 247
Index 249
Preface to the Instructor
You are probably about to teach a course that will give students
their second exposure to linear algebra. During their first brush with
the subject, your students probably worked with Euclidean spaces andmatrices. In contrast, this course will emphasize abstract vector spacesand linear maps.
The audacious title of this book deserves an explanation. Almost
all linear algebra books use determinants to prove that every linear op-erator on a finite-dimensional complex vector space has an eigenvalue.Determinants are difficult, nonintuitive, and often defined without mo-tivation. To prove the theorem about existence of eigenvalues on com-
plex vector spaces, most books must define determinants, prove that a
linear map is not invertible if and only if its determinant equals 0, andthen define the characteristic polynomial. This tortuous (torturous?)path gives students little feeling for why eigenvalues must exist.
In contrast, the simple determinant-free proofs presented here of-
fer more insight. Once determinants have been banished to the endof the book, a new route opens to the main goal of linear algebra—understanding the structure of linear operators.
This book starts at the beginning of the subject, with no prerequi-
sites other than the usual demand for suitable mathematical maturity.Even if your students have already seen some of the material in the
first few chapters, they may be unaccustomed to working exercises ofthe type presented here, most of which require an understanding ofproofs.
•Vector spaces are defined in Chapter 1, and their basic properties
are developed.
•Linear independence, span, basis, and dimension are defined in
Chapter 2, which presents the basic theory of finite-dimensionalvector spaces.
ix
x Preface to the Instructor
•Linear maps are introduced in Chapter 3. The key result here
is that for a linear map T, the dimension of the null space of T
plus the dimension of the range of Tequals the dimension of the
domain ofT.
•The part of the theory of polynomials that will be needed to un-
derstand linear operators is presented in Chapter 4. If you takeclass time going through the proofs in this chapter (which con-tains no linear algebra), then you probably will not have time to
cover some important aspects of linear algebra. Your studentswill already be familiar with the theorems about polynomials inthis chapter, so you can ask them to read the statements of theresults but not the proofs. The curious students will read someof the proofs anyway, which is why they are included in the text.
•The idea of studying a linear operator by restricting it to small
subspaces leads in Chapter 5 to eigenvectors. The highlight of the
chapter is a simple proof that on complex vector spaces, eigenval-ues always exist. This result is then used to show that each linearoperator on a complex vector space has an upper-triangular ma-trix with respect to some basis. Similar techniques are used toshow that every linear operator on a real vector space has an in-variant subspace of dimension 1 or 2. This result is used to provethat every linear operator on an odd-dimensional real vector space
has an eigenvalue. All this is done without defining determinants
or characteristic polynomials!
•Inner-product spaces are defined in Chapter 6, and their basic
properties are developed along with standard tools such as ortho-normal bases, the Gram-Schmidt procedure, and adjoints. Thischapter also shows how orthogonal projections can be used tosolve certain minimization problems.
•The spectral theorem, which characterizes the linear operators for
which there exists an orthonormal basis consisting of eigenvec-tors, is the highlight of Chapter 7. The work in earlier chapterspays off here with especially simple proofs. This chapter also
deals with positive operators, linear isometries, the polar decom-
position, and the singular-value decomposition.
Preface to the Instructor xi
•The minimal polynomial, characteristic polynomial, and general-
ized eigenvectors are introduced in Chapter 8. The main achieve-ment of this chapter is the description of a linear operator on
a complex vector space in terms of its generalized eigenvectors.This description enables one to prove almost all the results usu-ally proved using Jordan form. For example, these tools are used
to prove that every invertible linear operator on a complex vectorspace has a square root. The chapter concludes with a proof thatevery linear operator on a complex vector space can be put into
Jordan form.
•Linear operators on real vector spaces occupy center stage in
Chapter 9. Here two-dimensional invariant subspaces make upfor the possible lack of eigenvalues, leading to results analogousto those obtained on complex vector spaces.
•The trace and determinant are defined in Chapter 10 in terms
of the characteristic polynomial (defined earlier without determi-
nants). On complex vector spaces, these definitions can be re-stated: the trace is the sum of the eigenvalues and the determi-nant is the product of the eigenvalues (both counting multiplic-ity). These easy-to-remember definitions would not be possiblewith the traditional approach to eigenvalues because that methoduses determinants to prove that eigenvalues exist. The standardtheorems about determinants now become much clearer. The po-
lar decomposition and the characterization of self-adjoint opera-
tors are used to derive the change of variables formula for multi-variable integrals in a fashion that makes the appearance of thedeterminant there seem natural.
This book usually develops linear algebra simultaneously for real
and complex vector spaces by letting Fdenote either the real or the
complex numbers. Abstract fields could be used instead, but to do sowould introduce extra abstraction without leading to any new linear al-gebra. Another reason for restricting attention to the real and complexnumbers is that polynomials can then be thought of as genuine func-tions instead of the more formal objects needed for polynomials with
coefficients in finite fields. Finally, even if the beginning part of the the-
ory were developed with arbitrary fields, inner-product spaces wouldpush consideration back to just real and complex vector spaces.
xii Preface to the Instructor
Even in a book as short as this one, you cannot expect to cover every-
thing. Going through the first eight chapters is an ambitious goal for aone-semester course. If you must reach Chapter 10, then I suggest cov-
ering Chapters 1, 2, and 4 quickly (students may have seen this materialin earlier courses) and skipping Chapter 9 (in which case you shoulddiscuss trace and determinants only on complex vector spaces).
A goal more important than teaching any particular set of theorems
is to develop in students the ability to understand and manipulate theobjects of linear algebra. Mathematics can be learned only by doing;
fortunately, linear algebra has many good homework problems. Whenteaching this course, I usually assign two or three of the exercises each
class, due the next class. Going over the homework might take up a
third or even half of a typical class.
A solutions manual for all the exercises is available (without charge)
only to instructors who are using this book as a textbook. To obtainthe solutions manual, instructors should send an e-mail request to me(or contact Springer if I am no longer around).
Please check my web site for a list of errata (which I hope will be
empty or almost empty) and other information about this book.
I would greatly appreciate hearing about any errors in this book,
even minor ones. I welcome your suggestions for improvements, eventiny ones. Please feel free to contact me.
Have fun!
Sheldon Axler
Mathematics Department
San Francisco State University
San Francisco, CA 94132, USA
e-mail: [email protected]
www home page: http://math.sfsu.edu/axler
Preface to the Student
You are probably about to begin your second exposure to linear al-
gebra. Unlike your first brush with the subject, which probably empha-sized Euclidean spaces and matrices, we will focus on abstract vectorspaces and linear maps. These terms will be defined later, so don’tworry if you don’t know what they mean. This book starts from the be-ginning of the subject, assuming no knowledge of linear algebra. The
key point is that you are about to immerse yourself in serious math-
ematics, with an emphasis on your attaining a deep understanding ofthe definitions, theorems, and proofs.
You cannot expect to read mathematics the way you read a novel. If
you zip through a page in less than an hour, you are probably going toofast. When you encounter the phrase “as you should verify”, you shouldindeed do the verification, which will usually require some writing onyour part. When steps are left out, you need to supply the missing
pieces. You should ponder and internalize each definition. For each
theorem, you should seek examples to show why each hypothesis isnecessary.
Please check my web site for a list of errata (which I hope will be
empty or almost empty) and other information about this book.
I would greatly appreciate hearing about any errors in this book,
even minor ones. I welcome your suggestions for improvements, even
tiny ones.
Have fun!
Sheldon Axler
Mathematics DepartmentSan Francisco State UniversitySan Francisco, CA 94132, USA
e-mail: [email protected]
www home page: http://math.sfsu.edu/axler
xiii
Acknowledgments
I owe a huge intellectual debt to the many mathematicians who cre-
ated linear algebra during the last two centuries. In writing this book I
tried to think about the best way to present linear algebra and to proveits theorems, without regard to the standard methods and proofs used
in most textbooks. Thus I did not consult other books while writing
this one, though the memory of many books I had studied in the pastsurely influenced me. Most of the results in this book belong to thecommon heritage of mathematics. A special case of a theorem mayfirst have been proved in antiquity (which for linear algebra means the
nineteenth century), then slowly sharpened and improved over decades
by many mathematicians. Bestowing proper credit on all the contrib-utors would be a difficult task that I have not undertaken. In no caseshould the reader assume that any theorem presented here representsmy original contribution.
Many people helped make this a better book. For useful sugges-
tions and corrections, I am grateful to William Arveson (for suggestingthe proof of 5.13), Marilyn Brouwer, William Brown, Robert Burckel,
Paul Cohn, James Dudziak, David Feldman (for suggesting the proof of
8.40), Pamela Gorkin, Aram Harrow, Pan Fong Ho, Dan Kalman, RobertKantrowitz, Ramana Kappagantu, Mizan Khan, Mikael Lindstr ¨om, Ja-
cob Plotkin, Elena Poletaeva, Mihaela Poplicher, Richard Potter, WadeRamey, Marian Robbins, Jonathan Rosenberg, Joan Stamm, ThomasStarbird, Jay Valanju, and Thomas von Foerster.
Finally, I thank Springer for providing me with help when I needed
it and for allowing me the freedom to make the final decisions about
the content and appearance of this book.
xv
Chapter 1
Vector Spaces
Linear algebra is the study of linear maps on finite-dimensional vec-
tor spaces. Eventually we will learn what all these terms mean. In thischapter we will define vector spaces and discuss their elementary prop-erties.
In some areas of mathematics, including linear algebra, better the-
orems and more insight emerge if complex numbers are investigatedalong with real numbers. Thus we begin by introducing the complex
numbers and their basic properties.
✽
1
2 Chapter 1.Vector Spaces
Complex Numbers
You should already be familiar with the basic properties of the set R
of real numbers. Complex numbers were invented so that we can take
square roots of negative numbers. The key idea is to assume we have
a square root of −1, denoted i, and manipulate it using the usual rules The symbol iwas first
used to denote√
−1by
the Swiss
mathematician
Leonhard Euler in 1777.of arithmetic. Formally, a complex number is an ordered pair (a,b),
wherea,b∈R, but we will write this as a+bi. The set of all complex
numbers is denoted by C:
C={a+bi:a,b∈R}.
Ifa∈R, we identify a+0iwith the real number a. Thus we can think
ofRas a subset of C.
Addition and multiplication on Care defined by
(a+bi)+(c+di)=(a+c)+(b+d)i,
(a+bi)(c+di)=(ac−bd)+(ad+bc)i;
herea,b,c,d∈R. Using multiplication as defined above, you should
verify thati2=−1. Do not memorize the formula for the product
of two complex numbers; you can always rederive it by recalling thati
2=−1 and then using the usual rules of arithmetic.
You should verify, using the familiar properties of the real num-
bers, that addition and multiplication on Csatisfy the following prop-
erties:
commutativity
w+z=z+wandwz=zwfor allw,z∈C;
associativity
(z1+z2)+z3=z1+(z2+z3)and(z1z2)z3=z1(z2z3)for all
z1,z2,z3∈C;
identities
z+0=zandz1=zfor allz∈C;
additive inverse
for everyz∈C, there exists a unique w∈Csuch thatz+w=0;
multiplicative inverse
for everyz∈Cwithz/negationslash=0, there exists a unique w∈Csuch that
zw=1;
Complex Numbers 3
distributive property
λ(w+z)=λw+λzfor allλ,w,z∈C.
Forz∈C, we let−zdenote the additive inverse of z. Thus−zis
the unique complex number such that
z+(−z)=0.
Subtraction on Cis defined by
w−z=w+(−z)
forw,z∈C.
Forz∈Cwithz/negationslash=0, we let 1/zdenote the multiplicative inverse
ofz. Thus 1/z is the unique complex number such that
z(1/z)=1.
Division on Cis defined by
w/z=w(1/z)
forw,z∈Cwithz/negationslash=0.
So that we can conveniently make definitions and prove theorems
that apply to both real and complex numbers, we adopt the followingnotation:
The letter Fis used
because Rand Care
examples of what arecalled fields. In this
book we will not need
to deal with fields other
than RorC. Many of
the definitions,theorems, and proofs
in linear algebra that
work for both Rand C
also work withoutchange if an arbitrary
field replaces RorC.Throughout this book,
Fstands for either RorC.
Thus if we prove a theorem involving F, we will know that it holds when
Fis replaced with Rand when Fis replaced with C. Elements of Fare
called scalars . The word “scalar”, which means number, is often used
when we want to emphasize that an object is a number, as opposed toa vector (vectors will be defined soon).
Forz∈Fandma positive integer, we define z
mto denote the
product ofzwith itselfmtimes:
zm=z·····z/bracehtipupleft/bracehtipdownright/bracehtipdownleft/bracehtipupright
mtimes.
Clearly(zm)n=zmnand(wz)m=wmzmfor allw,z∈Fand all
positive integers m,n .
4 Chapter 1.Vector Spaces
Definition of Vector Space
Before defining what a vector space is, let’s look at two important
examples. The vector space R2, which you can think of as a plane,
consists of all ordered pairs of real numbers:
R2={(x,y) :x,y∈R}.
The vector space R3, which you can think of as ordinary space, consists
of all ordered triples of real numbers:
R3={(x,y,z) :x,y,z∈R}.
To generalize R2and R3to higher dimensions, we first need to dis-
cuss the concept of lists. Suppose nis a nonnegative integer. A listof
lengthnis an ordered collection of nobjects (which might be num-
bers, other lists, or more abstract entities) separated by commas and
surrounded by parentheses. A list of length nlooks like this: Many mathematicians
call a list of length nan
n-tuple. (x1,...,xn).
Thus a list of length 2 is an ordered pair and a list of length 3 is an
ordered triple. For j∈{1,...,n}, we say that xjis thejthcoordinate
of the list above. Thus x1is called the first coordinate, x2is called the
second coordinate, and so on.
Sometimes we will use the word listwithout specifying its length.
Remember, however, that by definition each list has a finite length thatis a nonnegative integer, so that an object that looks like
(x
1,x2,...),
which might be said to have infinite length, is not a list. A list of length
0 looks like this: (). We consider such an object to be a list so that
some of our theorems will not have trivial exceptions.
Two lists are equal if and only if they have the same length and
the same coordinates in the same order. In other words, (x1,...,xm)
equals(y1,...,yn)if and only if m=nandx1=y1,...,xm=ym.
Lists differ from sets in two ways: in lists, order matters and repeti-
tions are allowed, whereas in sets, order and repetitions are irrelevant.
For example, the lists (3,5)and(5,3)are not equal, but the sets {3,5}
and{5,3}are equal. The lists (4,4)and(4,4,4)are not equal (they
Definition of Vector Space 5
do not have the same length), though the sets {4,4}and{4,4,4}both
equal the set {4}.
To define the higher-dimensional analogues of R2and R3, we will
simply replace Rwith F(which equals RorC) and replace the 2 or 3
with an arbitrary positive integer. Specifically, fix a positive integer n
for the rest of this section. We define Fnto be the set of all lists of
lengthnconsisting of elements of F:
Fn={(x1,...,xn):xj∈Fforj=1,...,n}.
For example, if F=Randnequals 2 or 3, then this definition of Fn
agrees with our previous notions of R2and R3. As another example,
C4is the set of all lists of four complex numbers:
C4={(z1,z2,z3,z4):z1,z2,z3,z4∈C}.
Ifn≥4, we cannot easily visualize Rnas a physical object. The same For an amusing
account of how R3
would be perceived by
a creature living in R2,
read Flatland: A
Romance of Many
Dimensions, by Edwin
A. Abbott. This novel,published in 1884, can
help creatures living inthree-dimensional
space, such as
ourselves, imagine a
physical space of fouror more dimensions.problem arises if we work with complex numbers: C1can be thought
of as a plane, but for n≥2, the human brain cannot provide geometric
models of Cn. However, even if nis large, we can perform algebraic
manipulations in Fnas easily as in R2orR3. For example, addition is
defined on Fnby adding corresponding coordinates:
1.1(x1,...,xn)+(y1,...,yn)=(x1+y1,...,xn+yn).
Often the mathematics of Fnbecomes cleaner if we use a single
entity to denote an list of nnumbers, without explicitly writing the
coordinates. Thus the commutative property of addition on Fnshould
be expressed as
x+y=y+x
for allx,y∈Fn, rather than the more cumbersome
(x1,...,xn)+(y1,...,yn)=(y1,...,yn)+(x1,...,xn)
for allx1,...,xn,y1,...,yn∈F(even though the latter formulation
is needed to prove commutativity). If a single letter is used to denotean element of F
n, then the same letter, with appropriate subscripts,
is often used when coordinates must be displayed. For example, if
x∈Fn, then letting xequal(x1,...,xn)is good notation. Even better,
work with just xand avoid explicit coordinates, if possible.
6 Chapter 1.Vector Spaces
We let 0 denote the list of length nall of whose coordinates are 0:
0=(0,...,0).
Note that we are using the symbol 0 in two different ways—on the
left side of the equation above, 0 denotes a list of length n, whereas
on the right side, each 0 denotes a number. This potentially confusing
practice actually causes no problems because the context always makesclear what is intended. For example, consider the statement that 0 is
an additive identity for F
n:
x+0=x
for allx∈Fn. Here 0 must be a list because we have not defined the
sum of an element of Fn(namely,x) and the number 0.
A picture can often aid our intuition. We will draw pictures de-
picting R2because we can easily sketch this space on two-dimensional
surfaces such as paper and blackboards. A typical element of R2is a
pointx=(x1,x2). Sometimes we think of xnot as a point but as an
arrow starting at the origin and ending at (x1,x2), as in the picture
below. When we think of xas an arrow, we refer to it as a vector .
x -axis1x -axis2
(x , x )2 1
x
Elements of R2can be thought of as points or as vectors.
The coordinate axes and the explicit coordinates unnecessarily clut-
ter the picture above, and often you will gain better understanding bydispensing with them and just thinking of the vector, as in the nextpicture.
Definition of Vector Space 7
x
0
A vector
Whenever we use pictures in R2or use the somewhat vague lan-
guage of points and vectors, remember that these are just aids to our
understanding, not substitutes for the actual mathematics that we will
develop. Though we cannot draw good pictures in high-dimensionalspaces, the elements of these spaces are as rigorously defined as ele-ments of R
2. For example, (2,−3,17,π,√
2)is an element of R5, and we
may casually refer to it as a point in R5or a vector in R5without wor-
rying about whether the geometry of R5has any physical meaning.
Recall that we defined the sum of two elements of Fnto be the ele- Mathematical models
of the economy often
have thousands of
variables, say
x1,...,x 5000 , which
means that we must
operate in R5000. Such
a space cannot be dealt
with geometrically, but
the algebraic approachworks well. That’s why
our subject is called
linear algebra.ment of Fnobtained by adding corresponding coordinates; see 1.1. In
the special case of R2, addition has a simple geometric interpretation.
Suppose we have two vectors xandyinR2that we want to add, as in
the left side of the picture below. Move the vector yparallel to itself so
that its initial point coincides with the end point of the vector x. The
sumx+ythen equals the vector whose initial point equals the ini-
tial point of xand whose end point equals the end point of the moved
vectory, as in the right side of the picture below.
y
x+y
yx
0x
0
The sum of two vectors
Our treatment of the vector yin the picture above illustrates a standard
philosophy when we think of vectors in R2as arrows: we can move an
arrow parallel to itself (not changing its length or direction) and still
think of it as the same vector.
8 Chapter 1.Vector Spaces
Having dealt with addition in Fn, we now turn to multiplication. We
could define a multiplication on Fnin a similar fashion, starting with
two elements of Fnand getting another element of Fnby multiplying
corresponding coordinates. Experience shows that this definition is notuseful for our purposes. Another type of multiplication, called scalarmultiplication, will be central to our subject. Specifically, we need todefine what it means to multiply an element of F
nby an element of F.
We make the obvious definition, performing the multiplication in eachcoordinate:
a(x
1,...,xn)=(ax 1,...,axn);
herea∈Fand(x1,...,xn)∈Fn.
Scalar multiplication has a nice geometric interpretation in R2.I f In scalar multiplication,
we multiply together a
scalar and a vector,
getting a vector. You
may be familiar with
the dot product in R2
orR3, in which we
multiply together two
vectors and obtain a
scalar. Generalizations
of the dot product will
become important
when we study inner
products in Chapter 6.
You may also be
familiar with the cross
product in R3, in which
we multiply together
two vectors and obtain
another vector. No
useful generalization of
this type of
multiplication exists in
higher dimensions.ais a positive number and xis a vector in R2, thenaxis the vector
that points in the same direction as xand whose length is atimes the
length ofx. In other words, to get ax, we shrink or stretch xby a
factor ofa, depending upon whether a<1o ra>1. The next picture
illustrates this point.
x
(1/2)x(3/2)x
Multiplication by positive scalars
Ifais a negative number and xis a vector in R2, thenaxis the vector
that points in the opposite direction as xand whose length is |a|times
the length of x, as illustrated in the next picture.
x
(−1/2)x
(−3/2)x
Multiplication by negative scalars
Definition of Vector Space 9
The motivation for the definition of a vector space comes from the
important properties possessed by addition and scalar multiplicationonF
n. Specifically, addition on Fnis commutative and associative and
has an identity, namely, 0. Every element has an additive inverse. Scalarmultiplication on F
nis associative, and scalar multiplication by 1 acts
as a multiplicative identity should. Finally, addition and scalar multi-plication on F
nare connected by distributive properties.
We will define a vector space to be a set Valong with an addition
and a scalar multiplication on Vthat satisfy the properties discussed
in the previous paragraph. By an addition onVwe mean a function
that assigns an element u+v∈Vto each pair of elements u,v∈V.
By a scalar multiplication onVwe mean a function that assigns an
elementav∈Vto eacha∈Fand eachv∈V.
Now we are ready to give the formal definition of a vector space.
Avector space is a setValong with an addition on Vand a scalar
multiplication on Vsuch that the following properties hold:
commutativity
u+v=v+ufor allu,v∈V;
associativity
(u+v)+w=u+(v+w)and(ab)v=a(bv) for allu,v,w∈V
and alla,b∈F;
additive identity
there exists an element 0 ∈Vsuch thatv+0=vfor allv∈V;
additive inverse
for everyv∈V, there exists w∈Vsuch thatv+w=0;
multiplicative identity
1v=vfor allv∈V;
distributive properties
a(u+v)=au+avand(a+b)u=au+bufor alla,b∈Fand
allu,v∈V.
The scalar multiplication in a vector space depends upon F. Thus
when we need to be precise, we will say that Vis a vector space over F
instead of saying simply that Vis a vector space. For example, Rnis
a vector space over R, and Cnis a vector space over C. Frequently, a
vector space over Ris called a real vector space and a vector space over
10 Chapter 1.Vector Spaces
Cis called a complex vector space . Usually the choice of Fis either
obvious from the context or irrelevant, and thus we often assume thatFis lurking in the background without specifically mentioning it.
Elements of a vector space are called vectors orpoints . This geo-
metric language sometimes aids our intuition.
Not surprisingly, F
nis a vector space over F, as you should verify.
Of course, this example motivated our definition of vector space.
For another example, consider F∞, which is defined to be the set of The simplest vector
space contains only
one point. In other
words,{0} is a vector
space, though not a
very interesting one.all sequences of elements of F:
F∞={(x1,x2,...) :xj∈Fforj=1,2,...}.
Addition and scalar multiplication on F∞are defined as expected:
(x1,x2,...)+(y1,y2,...)=(x1+y1,x2+y2,...),
a(x 1,x2,...)=(ax 1,ax 2,...).
With these definitions, F∞becomes a vector space over F, as you should
verify. The additive identity in this vector space is the sequence con-sisting of all 0’s.
Our next example of a vector space involves polynomials. A function
p:F→Fis called a polynomial with coefficients in Fif there exist
a
0,...,am∈Fsuch that
p(z)=a0+a1z+a2z2+···+a mzm
for allz∈F. We define P(F)to be the set of all polynomials with Though Fnis our
crucial example of a
vector space, not all
vector spaces consist
of lists. For example,
the elements of P(F)
consist of functions on
F, not lists. In general,
a vector space is an
abstract entity whose
elements might be lists,
functions, or weird
objects.coefficients in F. Addition on P(F)is defined as you would expect: if
p,q∈P(F), thenp+qis the polynomial defined by
(p+q)(z)=p(z)+q(z)
forz∈F. For example, if pis the polynomial defined by p(z)=2z+z3
andqis the polynomial defined by q(z)=7+4z, thenp+qis the
polynomial defined by (p+q)(z)=7+6z+z3. Scalar multiplication
onP(F)also has the obvious definition: if a∈Fandp∈P(F), then
apis the polynomial defined by
(ap)(z)=ap(z)
forz∈F. With these definitions of addition and scalar multiplication,
P(F)is a vector space, as you should verify. The additive identity in
this vector space is the polynomial all of whose coefficients equal 0.
Soon we will see further examples of vector spaces, but first we need
to develop some of the elementary properties of vector spaces.
Properties of Vector Spaces 11
Properties of Vector Spaces
The definition of a vector space requires that it have an additive
identity. The proposition below states that this identity is unique.
1.2 Proposition: A vector space has a unique additive identity.
Proof: Suppose 0 and 0/primeare both additive identities for some vec-
tor spaceV. Then
0/prime=0/prime+0=0,
where the first equality holds because 0 is an additive identity and the
second equality holds because 0/primeis an additive identity. Thus 0/prime=0,
proving that Vhas only one additive identity. The symbol means
“end of the proof”.
Each element vin a vector space has an additive inverse, an element
win the vector space such that v+w=0. The next proposition shows
that each element in a vector space has only one additive inverse.
1.3 Proposition: Every element in a vector space has a unique
additive inverse.
Proof: SupposeVis a vector space. Let v∈V. Suppose that w
andw/primeare additive inverses of v. Then
w=w+0=w+(v+w/prime)=(w+v)+w/prime=0+w/prime=w/prime.
Thusw=w/prime, as desired.
Because additive inverses are unique, we can let −vdenote the ad-
ditive inverse of a vector v. We definew−vto meanw+(−v) .
Almost all the results in this book will involve some vector space.
To avoid being distracted by having to restate frequently somethingsuch as “Assume that Vis a vector space”, we now make the necessary
declaration once and for all:
Let’s agree that for the rest of the book
Vwill denote a vector space over F.
12 Chapter 1.Vector Spaces
Because of associativity, we can dispense with parentheses when
dealing with additions involving more than two elements in a vectorspace. For example, we can write u+v+wwithout parentheses because
the two possible interpretations of that expression, namely, (u+v)+w
andu+(v+w), are equal. We first use this familiar convention of not
using parentheses in the next proof. In the next proposition, 0 denotesa scalar (the number 0 ∈F) on the left side of the equation and a vector
(the additive identity of V) on the right side of the equation.
1.4 Proposition: 0v=0for everyv∈V.
Note that 1.4 and 1.5
assert something about
scalar multiplication
and the additive
identity ofV. The only
part of the definition of
a vector space that
connects scalar
multiplication and
vector addition is the
distributive property.
Thus the distributive
property must be used
in the proofs.Proof: Forv∈V, we have
0v=(0+0)v=0v+0v.
Adding the additive inverse of 0 vto both sides of the equation above
gives 0=0v, as desired.
In the next proposition, 0 denotes the additive identity of V. Though
their proofs are similar, 1.4 and 1.5 are not identical. More precisely,1.4 states that the product of the scalar 0 and any vector equals thevector 0, whereas 1.5 states that the product of any scalar and thevector 0 equals the vector 0.
1.5 Proposition: a0=0for everya∈F.
Proof: Fora∈F, we have
a0=a(0+0)=a0+a0.
Adding the additive inverse of a0 to both sides of the equation above
gives 0=a0, as desired.
Now we show that if an element of Vis multiplied by the scalar −1,
then the result is the additive inverse of the element of V.
1.6 Proposition: (−1)v=−vfor everyv∈V.
Proof: Forv∈V, we have
v+(−1)v=1v+(−1)v=/parenleftbig
1+(−1)/parenrightbig
v=0v=0.
This equation says that (−1)v , when added to v, gives 0. Thus (−1)v
must be the additive inverse of v, as desired.
Subspaces 13
Subspaces
A subsetUofVis called a subspace ofVifUis also a vector space Some mathematicians
use the term linear
subspace, which means
the same as subspace.(using the same addition and scalar multiplication as on V). For exam-
ple,
{(x 1,x2,0):x1,x2∈F}
is a subspace of F3.
IfUis a subset of V, then to check that Uis a subspace of Vwe
need only check that Usatisfies the following:
additive identity
0∈U
closed under addition
u,v∈Uimpliesu+v∈U;
closed under scalar multiplication
a∈Fandu∈Uimpliesau∈U.
The first condition insures that the additive identity of Vis inU. The Clearly{0} is the
smallest subspace of V
andVitself is the
largest subspace of V.
The empty set is not a
subspace of Vbecause
a subspace must be a
vector space and a
vector space must
contain at least one
element, namely, an
additive identity.second condition insures that addition makes sense on U. The third
condition insures that scalar multiplication makes sense on U. To show
thatUis a vector space, the other parts of the definition of a vector
space do not need to be checked because they are automatically satis-
fied. For example, the associative and commutative properties of addi-
tion automatically hold on Ubecause they hold on the larger space V.
As another example, if the third condition above holds and u∈U, then
−u(which equals (−1)u by 1.6) is also in U, and hence every element
ofUhas an additive inverse in U.
The three conditions above usually enable us to determine quickly
whether a given subset of Vis a subspace of V. For example, if b∈F,
then
{(x 1,x2,x3,x4)∈F4:x3=5x4+b}
is a subspace of F4if and only if b=0, as you should verify. As another
example, you should verify that
{p∈P(F):p(3)=0}
is a subspace of P(F).
The subspaces of R2are precisely {0}, R2, and all lines in R2through
the origin. The subspaces of R3are precisely {0}, R3, all lines in R3
14 Chapter 1.Vector Spaces
through the origin, and all planes in R3through the origin. To prove
that all these objects are indeed subspaces is easy—the hard part is toshow that they are the only subspaces of R
2orR3. That task will be
easier after we introduce some additional tools in the next chapter.
Sums and Direct Sums
In later chapters, we will find that the notions of vector space sums
and direct sums are useful. We define these concepts here.
SupposeU1,...,Umare subspaces of V. The sum ofU1,...,Um, When dealing with
vector spaces, we are
usually interested only
in subspaces, as
opposed to arbitrary
subsets. The union of
subspaces is rarely a
subspace (see
Exercise 9 in this
chapter), which is why
we usually work with
sums rather than
unions.denotedU1+···+U m, is defined to be the set of all possible sums of
elements of U1,...,Um. More precisely,
U1+···+Um={u1+···+u m:u1∈U1,...,um∈Um}.
You should verify that if U1,...,Umare subspaces of V, then the sum
U1+···+Umis a subspace of V.
Let’s look at some examples of sums of subspaces. Suppose Uis the
set of all elements of F3whose second and third coordinates equal 0,
andWis the set of all elements of F3whose first and third coordinates
equal 0:
U={(x,0,0)∈F3:x∈F}andW={(0,y,0)∈F3:y∈F}.
Then
Sums of subspaces in
the theory of vector
spaces are analogous to
unions of subsets in set
theory. Given two
subspaces of a vector
space, the smallest
subspace containing
them is their sum.
Analogously, given two
subsets of a set, the
smallest subset
containing them is
their union.1.7 U+W={(x,y, 0):x,y∈F},
as you should verify.
As another example, suppose Uis as above and Wis the set of all
elements of F3whose first and second coordinates equal each other
and whose third coordinate equals 0:
W={(y,y, 0)∈F3:y∈F}.
ThenU+Wis also given by 1.7, as you should verify.
SupposeU1,...,Umare subspaces of V. ClearlyU1,...,Umare all
contained in U1+···+Um(to see this, consider sums u1+···+um
where all except one of the u’s are 0). Conversely, any subspace of V
containingU1,...,Ummust contain U1+···+Um(because subspaces
Sums and Direct Sums 15
must contain all finite sums of their elements). Thus U1+···+Umis
the smallest subspace of VcontainingU1,...,Um.
SupposeU1,...,Umare subspaces of Vsuch thatV=U1+···+Um.
Thus every element of Vcan be written in the form
u1+···+um,
where eachuj∈Uj. We will be especially interested in cases where
each vector in Vcan be uniquely represented in the form above. This
situation is so important that we give it a special name: direct sum.
Specifically, we say that Vis the direct sum of subspaces U1,...,Um,
writtenV=U1⊕···⊕Um, if each element of Vcan be written uniquely The symbol ⊕,
consisting of a plus
sign inside a circle, is
used to denote direct
sums as a reminder
that we are dealing with
a special type of sum ofsubspaces—each
element in the direct
sum can be represented
only one way as a sum
of elements from the
specified subspaces.as a sumu1+···+um, where each uj∈Uj.
Let’s look at some examples of direct sums. Suppose Uis the sub-
space of F3consisting of those vectors whose last coordinate equals 0,
andWis the subspace of F3consisting of those vectors whose first two
coordinates equal 0:
U={(x,y, 0)∈F3:x,y∈F}andW={(0,0,z)∈F3:z∈F}.
Then F3=U⊕W, as you should verify.
As another example, suppose Ujis the subspace of Fnconsisting
of those vectors whose coordinates are all 0, except possibly in the jth
slot (for example, U2={(0,x,0,...,0) ∈Fn:x∈F}). Then
Fn=U1⊕···⊕Un,
as you should verify.
As a final example, consider the vector space P(F)of all polynomials
with coefficients in F. LetUedenote the subspace of P(F)consisting
of all polynomials pof the form
p(z)=a0+a2z2+···+a2mz2m,
and letUodenote the subspace of P(F)consisting of all polynomials p
of the form
p(z)=a1z+a3z3+···+a 2m+1z2m+1;
heremis a nonnegative integer and a0,...,a 2m+1∈F(the notations
UeandUoshould remind you of even and odd powers of z). You should
verify that
16 Chapter 1.Vector Spaces
P(F)=Ue⊕Uo.
Sometimes nonexamples add to our understanding as much as ex-
amples. Consider the following three subspaces of F3:
U1={(x,y, 0)∈F3:x,y∈F};
U2={(0,0,z)∈F3:z∈F};
U3={(0,y,y)∈F3:y∈F}.
Clearly F3=U1+U2+U3because an arbitrary vector (x,y,z)∈F3can
be written as
(x,y,z)=(x,y, 0)+(0,0,z)+(0,0,0),
where the first vector on the right side is in U1, the second vector is
inU2, and the third vector is in U3. However, F3does not equal the
direct sum of U1,U2,U3because the vector (0,0,0)can be written in
two different ways as a sum u1+u2+u3, with eachuj∈Uj. Specifically,
we have
(0,0,0)=(0,1,0)+(0,0,1)+(0,−1,−1)
and, of course,
(0,0,0)=(0,0,0)+(0,0,0)+(0,0,0),
where the first vector on the right side of each equation above is in U1,
the second vector is in U2, and the third vector is in U3.
In the example above, we showed that something is not a direct sum
by showing that 0 does not have a unique representation as a sum ofappropriate vectors. The definition of direct sum requires that everyvector in the space have a unique representation as an appropriate sum.
Suppose we have a collection of subspaces whose sum equals the wholespace. The next proposition shows that when deciding whether this
collection of subspaces is a direct sum, we need only consider whether
0 can be uniquely written as an appropriate sum.
1.8 Proposition: Suppose that U
1,...,Unare subspaces of V. Then
V=U1⊕···⊕Unif and only if both the following conditions hold:
(a)V=U1+···+Un;
(b) the only way to write 0as a sumu1+···+un, where each
uj∈Uj, is by taking all the uj’s equal to 0.
Sums and Direct Sums 17
Proof: First suppose that V=U1⊕···⊕U n. Clearly (a) holds
(because of how sum and direct sum are defined). To prove (b), supposethatu
1∈U1,...,un∈Unand
0=u1+···+un.
Then eachujmust be 0 (this follows from the uniqueness part of the
definition of direct sum because 0 =0+···+ 0 and 0∈U1,...,0∈Un),
proving (b).
Now suppose that (a) and (b) hold. Let v∈V. By (a), we can write
v=u1+···+un
for someu1∈U1,...,un∈Un. To show that this representation is
unique, suppose that we also have
v=v1+···+vn,
wherev1∈U1,...,vn∈Un. Subtracting these two equations, we have
0=(u1−v1)+···+(un−vn).
Clearlyu1−v1∈U1,...,un−vn∈Un, so the equation above and (b)
imply that each uj−vj=0. Thusu1=v1,...,un=vn, as desired.
The next proposition gives a simple condition for testing which pairs Sums of subspaces are
analogous to unions of
subsets. Similarly,
direct sums of
subspaces are
analogous to disjointunions of subsets. No
two subspaces of a
vector space can be
disjoint because both
must contain 0.S o
disjointness isreplaced, at least in the
case of two subspaces,
with the requirement
that the intersection
equals{0}.of subspaces give a direct sum. Note that this proposition deals only
with the case of two subspaces. When asking about a possible direct
sum with more than two subspaces, it is not enough to test that anytwo of the subspaces intersect only at 0. To see this, consider thenonexample presented just before 1.8. In that nonexample, we had
F
3=U1+U2+U3, but F3did not equal the direct sum of U1,U2,U3.
However, in that nonexample, we have U1∩U2=U1∩U3=U2∩U3={0}
(as you should verify). The next proposition shows that with just twosubspaces we get a nice necessary and sufficient condition for a directsum.
1.9 Proposition: Suppose that UandWare subspaces of V. Then
V=U⊕Wif and only if V=U+WandU∩W={0}.
Proof: First suppose that V=U⊕W. ThenV=U+W(by the
definition of direct sum). Also, if v∈U∩W, then 0=v+(−v) , where
18 Chapter 1.Vector Spaces
v∈Uand−v∈W. By the unique representation of 0 as the sum of a
vector inUand a vector in W, we must have v=0. ThusU∩W={0},
completing the proof in one direction.
To prove the other direction, now suppose that V=U+Wand
U∩W={0}. To prove that V=U⊕W, suppose that
0=u+w,
whereu∈Uandw∈W. To complete the proof, we need only show
thatu=w=0 (by 1.8). The equation above implies that u=−w∈W.
Thusu∈U∩W, and hence u=0. This, along with equation above,
implies that w=0, completing the proof.
Exercises 19
Exercises
1. Suppose aandbare real numbers, not both 0. Find real numbers
canddsuch that
1/(a+bi)=c+di.
2. Show that
−1+√
3i
2
is a cube root of 1 (meaning that its cube equals 1).
3. Prove that −(−v)=vfor everyv∈V.
4. Prove that if a∈F,v∈V, andav=0, thena=0o rv=0.
5. For each of the following subsets of F3, determine whether it is
a subspace of F3:
(a){(x 1,x2,x3)∈F3:x1+2x2+3x3=0};
(b){(x 1,x2,x3)∈F3:x1+2x2+3x3=4};
(c){(x 1,x2,x3)∈F3:x1x2x3=0};
(d){(x 1,x2,x3)∈F3:x1=5x3}.
6. Give an example of a nonempty subset UofR2such thatUis
closed under addition and under taking additive inverses (mean-ing−u∈Uwheneveru∈U), butUis not a subspace of R
2.
7. Give an example of a nonempty subset UofR2such thatUis
closed under scalar multiplication, but Uis not a subspace of R2.
8. Prove that the intersection of any collection of subspaces of Vis
a subspace of V.
9. Prove that the union of two subspaces of Vis a subspace of Vif
and only if one of the subspaces is contained in the other.
10. Suppose that Uis a subspace of V. What isU+U?
11. Is the operation of addition on the subspaces of Vcommutative?
Associative? (In other words, if U1,U2,U3are subspaces of V,i s
U1+U2=U2+U1?I s(U1+U2)+U3=U1+(U2+U3)?)
20 Chapter 1.Vector Spaces
12. Does the operation of addition on the subspaces of Vhave an
additive identity? Which subspaces have additive inverses?
13. Prove or give a counterexample: if U1,U2,Ware subspaces of V
such that
U1+W=U2+W,
thenU1=U2.
14. Suppose Uis the subspace of P(F)consisting of all polynomials
pof the form
p(z)=az2+bz5,
wherea,b∈F. Find a subspace WofP(F)such thatP(F)=
U⊕W.
15. Prove or give a counterexample: if U1,U2,Ware subspaces of V
such that
V=U1⊕WandV=U2⊕W,
thenU1=U2.
Chapter 2
Finite-Dimensional
Vector Spaces
In the last chapter we learned about vector spaces. Linear algebra
focuses not on arbitrary vector spaces, but on finite-dimensional vector
spaces, which we introduce in this chapter. Here we will deal with thekey concepts associated with these spaces: span, linear independence,basis, and dimension.
Let’s review our standing assumptions:
Recall that Fdenotes RorC.
Recall also that Vis a vector space over F.
✽✽
21
22 Chapter 2.Finite-Dimensional Vector Spaces
Span and Linear Independence
Alinear combination of a list(v1,...,vm)of vectors in Vis a vector
of the form
2.1 a1v1+···+amvm,
wherea1,...,am∈F. The set of all linear combinations of (v1,...,vm)
is called the span of(v1,...,vm), denoted span (v1,...,vm). In other Some mathematicians
use the term linear
span, which means the
same as span.words,
span(v 1,...,vm)={a1v1+···+amvm:a1,...,am∈F}.
As an example of these concepts, suppose V=F3. The vector
(7,2,9)is a linear combination of/parenleftbig
(2,1,3),(1, 0,1)/parenrightbig
because
(7,2,9)=2(2,1,3)+3(1,0,1).
Thus(7,2,9)∈span/parenleftbig
(2,1,3),(1, 0,1)/parenrightbig
.
You should verify that the span of any list of vectors in Vis a sub-
space ofV. To be consistent, we declare that the span of the empty list
()equals{0}(recall that the empty set is not a subspace of V).
If(v1,...,vm)is a list of vectors in V, then eachvjis a linear com-
bination of(v1,...,vm)(to show this, set aj=1 and let the other a’s
in 2.1 equal 0). Thus span(v 1,...,vm)contains each vj. Conversely,
because subspaces are closed under scalar multiplication and addition,every subspace of Vcontaining each v
jmust contain span(v 1,...,vm).
Thus the span of a list of vectors in Vis the smallest subspace of V
containing all the vectors in the list.
If span(v 1,...,vm)equalsV, we say that (v1,...,vm)spansV.A
vector space is called finite dimensional if some list of vectors in it Recall that by
definition every list has
finite length.spans the space. For example, Fnis finite dimensional because
/parenleftbig
(1,0,...,0),(0, 1,0,...,0),...,( 0,...,0, 1)/parenrightbig
spans Fn, as you should verify.
Before giving the next example of a finite-dimensional vector space,
we need to define the degree of a polynomial. A polynomial p∈P(F)
is said to have degreemif there exist scalars a0,a1,...,am∈Fwith
am/negationslash=0 such that
2.2 p(z)=a0+a1z+···+amzm
Span and Linear Independence 23
for allz∈F. The polynomial that is identically 0 is said to have de-
gree−∞.
Forma nonnegative integer, let Pm(F)denote the set of all poly-
nomials with coefficients in Fand degree at most m. You should ver-
ify thatPm(F)is a subspace of P(F); hencePm(F)is a vector space.
This vector space is finite dimensional because it is spanned by the list(1,z,...,z
m); here we are slightly abusing notation by letting zkdenote
a function (so zis a dummy variable).
A vector space that is not finite dimensional is called infinite di- Infinite-dimensional
vector spaces, which
we will not mention
much anymore, are the
center of attention in
the branch of
mathematics calledfunctional analysis.
Functional analysis
uses tools from both
analysis and algebra.mensional . For example, P(F)is infinite dimensional. To prove this,
consider any list of elements of P(F). Letmdenote the highest degree
of any of the polynomials in the list under consideration (recall that bydefinition a list has finite length). Then every polynomial in the span ofthis list must have degree at most m. Thus our list cannot span P(F).
Because no list spans P(F), this vector space is infinite dimensional.
The vector space F
∞, consisting of all sequences of elements of F,
is also infinite dimensional, though this is a bit harder to prove. Youshould be able to give a proof by using some of the tools we will soondevelop.
Supposev
1,...,vm∈Vandv∈span(v 1,...,vm). By the definition
of span, there exist a1,...,am∈Fsuch that
v=a1v1+···+amvm.
Consider the question of whether the choice of a’s in the equation
above is unique. Suppose ˆa1,..., ˆamis another set of scalars such that
v=ˆa1v1+···+ ˆamvm.
Subtracting the last two equations, we have
0=(a1−ˆa1)v1+···+(am−ˆam)vm.
Thus we have written 0 as a linear combination of (v1,...,vm). If the
only way to do this is the obvious way (using 0 for all scalars), theneacha
j−ˆajequals 0, which means that each ajequals ˆaj(and thus
the choice of a’s was indeed unique). This situation is so important
that we give it a special name—linear independence—which we now
define.
A list(v1,...,vm)of vectors in Vis called linearly independent if
the only choice of a1,...,am∈Fthat makesa1v1+···+a mvmequal
0i sa 1=···=am=0. For example,
24 Chapter 2.Finite-Dimensional Vector Spaces
/parenleftbig
(1,0,0,0),(0, 1,0,0),(0, 0,1,0)/parenrightbig
is linearly independent in F4, as you should verify. The reasoning in the
previous paragraph shows that (v1,...,vm)is linearly independent if
and only if each vector in span (v1,...,vm)has only one representation
as a linear combination of (v1,...,vm).
For another example of a linearly independent list, fix a nonnegative Most linear algebra
texts define linearly
independent sets
instead of linearly
independent lists. With
that definition, the set
{(0,1),(0, 1),(1, 0)} is
linearly independent in
F2because it equals the
set{(0,1),(1,0)}. With
our definition, the list/parenleftBig
(0,1),(0, 1),(1, 0)/parenrightBig
is
not linearly
independent (because 1
times the first vector
plus−1times the
second vector plus 0
times the third vector
equals 0). By dealing
with lists instead of
sets, we will avoid
some problems
associated with the
usual approach.integerm. Then(1,z,...,zm)is linearly independent in P(F). To verify
this, suppose that a0,a1,...,am∈Fare such that
2.3 a0+a1z+···+amzm=0
for everyz∈F. If at least one of the coefficients a0,a1,...,amwere
nonzero, then 2.3 could be satisfied by at most mdistinct values of z(if
you are unfamiliar with this fact, just believe it for now; we will prove
it in Chapter 4); this contradiction shows that all the coefficients in 2.3
equal 0. Hence (1,z,...,zm)is linearly independent, as claimed.
A list of vectors in Vis called linearly dependent if it is not lin-
early independent. In other words, a list (v1,...,vm)of vectors in V
is linearly dependent if there exist a1,...,am∈F, not all 0, such that
a1v1+···+a mvm=0. For example,/parenleftbig
(2,3,1),(1,−1,2),(7, 3,8)/parenrightbig
is
linearly dependent in F3because
2(2,3,1)+3(1,−1,2)+(−1)(7, 3,8)=(0,0,0).
As another example, any list of vectors containing the 0 vector is lin-
early dependent (why?).
You should verify that a list (v)of length 1 is linearly independent if
and only ifv/negationslash=0. You should also verify that a list of length 2 is linearly
independent if and only if neither vector is a scalar multiple of the
other. Caution: a list of length three or more may be linearly dependenteven though no vector in the list is a scalar multiple of any other vectorin the list, as shown by the example in the previous paragraph.
If some vectors are removed from a linearly independent list, the
remaining list is also linearly independent, as you should verify. Toallow this to remain true even if we remove all the vectors, we declare
the empty list ()to be linearly independent.
The lemma below will often be useful. It states that given a linearly
dependent list of vectors, with the first vector not zero, one of thevectors is in the span of the previous ones and furthermore we canthrow out that vector without changing the span of the original list.
Span and Linear Independence 25
2.4 Linear Dependence Lemma: If(v1,...,vm)is linearly depen-
dent inVandv1/negationslash=0, then there exists j∈{2,...,m}such that the
following hold:
(a)vj∈span(v 1,...,vj−1);
(b) if thejthterm is removed from (v1,...,vm), the span of the
remaining list equals span(v 1,...,vm).
Proof: Suppose(v1,...,vm)is linearly dependent in Vandv1/negationslash=0.
Then there exist a1,...,am∈F, not all 0, such that
a1v1+···+amvm=0.
Not all ofa2,a3,...,amcan be 0 (because v1/negationslash=0). Letjbe the largest
element of{2,...,m}such thataj/negationslash=0. Then
2.5 vj=−a1
ajv1−···−aj−1
ajvj−1,
proving (a).
To prove (b), suppose that u∈span(v 1,...,vm). Then there exist
c1,...,cm∈Fsuch that
u=c1v1+···+cmvm.
In the equation above, we can replace vjwith the right side of 2.5,
which shows that uis in the span of the list obtained by removing the
jthterm from(v1,...,vm). Thus (b) holds.
Now we come to a key result. It says that linearly independent lists
are never longer than spanning lists.
2.6 Theorem: In a finite-dimensional vector space, the length of Suppose that for each
positive integer m,
there exists a linearly
independent list of m
vectors inV. Then this
theorem implies that V
is infinite dimensional.every linearly independent list of vectors is less than or equal to the
length of every spanning list of vectors.
Proof: Suppose that (u1,...,um)is linearly independent in Vand
that(w1,...,wn)spansV. We need to prove that m≤n. W ed os o
through the multistep process described below; note that in each step
we add one of the u’s and remove one of the w’s.
26 Chapter 2.Finite-Dimensional Vector Spaces
Step 1
The list(w1,...,wn)spansV, and thus adjoining any vector to it
produces a linearly dependent list. In particular, the list
(u1,w1,...,wn)
is linearly dependent. Thus by the linear dependence lemma (2.4),
we can remove one of the w’s so that the list B(of lengthn)
consisting of u1and the remaining w’s spansV.
Step j
The listB(of lengthn) from step j−1 spansV, and thus adjoining
any vector to it produces a linearly dependent list. In particular,the list of length (n+1)obtained by adjoining u
jtoB, placing it
just afteru1,...,uj−1, is linearly dependent. By the linear depen-
dence lemma (2.4), one of the vectors in this list is in the span ofthe previous ones, and because (u
1,...,uj)is linearly indepen-
dent, this vector must be one of the w’s, not one of the u’s. We
can remove that wfromBso that the new list B(of lengthn)
consisting of u1,...,ujand the remaining w’s spansV.
After stepm, we have added all the u’s and the process stops. If at
any step we added a uand had no more w’s to remove, then we would
have a contradiction. Thus there must be at least as many w’s asu’s.
Our intuition tells us that any vector space contained in a finite-
dimensional vector space should also be finite dimensional. We nowprove that this intuition is correct.
2.7 Proposition: Every subspace of a finite-dimensional vector
space is finite dimensional.
Proof: SupposeVis finite dimensional and Uis a subspace of V.
We need to prove that Uis finite dimensional. We do this through the
following multistep construction.
Step 1
IfU={0}, thenUis finite dimensional and we are done. If U/negationslash=
{0}, then choose a nonzero vector v
1∈U.
Step j
IfU=span(v 1,...,vj−1), thenUis finite dimensional and we are
Bases 27
done. IfU/negationslash=span(v 1,...,vj−1), then choose a vector vj∈Usuch
that
vj∉span(v 1,...,vj−1).
After each step, as long as the process continues, we have constructed
a list of vectors such that no vector in this list is in the span of theprevious vectors. Thus after each step we have constructed a linearlyindependent list, by the linear dependence lemma (2.4). This linearlyindependent list cannot be longer than any spanning list of V(by 2.6),
and thus the process must eventually terminate, which means that U
is finite dimensional.
Bases
Abasis ofVis a list of vectors in Vthat is linearly independent and
spansV. For example,
/parenleftbig
(1,0,...,0),(0, 1,0,...,0),...,( 0,...,0, 1)/parenrightbig
is a basis of Fn, called the standard basis ofFn. In addition to the
standard basis, Fnhas many other bases. For example,/parenleftbig
(1,2),(3, 5)/parenrightbig
is a basis of F2. The list/parenleftbig
(1,2)/parenrightbig
is linearly independent but is not a
basis of F2because it does not span F2. The list/parenleftbig
(1,2),(3, 5),(4, 7)/parenrightbig
spans F2but is not a basis because it is not linearly independent. As
another example, (1,z,...,zm)is a basis of Pm(F).
The next proposition helps explain why bases are useful.
2.8 Proposition: A list(v1,...,vn)of vectors in Vis a basis of V
if and only if every v∈Vcan be written uniquely in the form
2.9 v=a1v1+···+anvn,
wherea1,...,an∈F.
Proof: First suppose that (v1,...,vn)is a basis of V. Letv∈V. This proof is
essentially a repetition
of the ideas that led us
to the definition oflinear independence.Because(v1,...,vn)spansV, there exist a1,...,an∈Fsuch that 2.9
holds. To show that the representation in 2.9 is unique, suppose thatb
1,...,bnare scalars so that we also have
v=b1v1+···+bnvn.
28 Chapter 2.Finite-Dimensional Vector Spaces
Subtracting the last equation from 2.9, we get
0=(a1−b1)v1+···+(an−bn)vn.
This implies that each aj−bj=0 (because(v1,...,vn)is linearly inde-
pendent) and hence a1=b1,...,an=bn. We have the desired unique-
ness, completing the proof in one direction.
For the other direction, suppose that every v∈Vcan be written
uniquely in the form given by 2.9. Clearly this implies that (v1,...,vn)
spansV. To show that (v1,...,vn)is linearly independent, suppose
thata1,...,an∈Fare such that
0=a1v1+···+anvn.
The uniqueness of the representation 2.9 (with v=0) implies that
a1=···=a n=0. Thus(v1,...,vn)is linearly independent and
hence is a basis of V.
A spanning list in a vector space may not be a basis because it is not
linearly independent. Our next result says that given any spanning list,
some of the vectors in it can be discarded so that the remaining list islinearly independent and still spans the vector space.
2.10 Theorem: Every spanning list in a vector space can be reduced
to a basis of the vector space.
Proof: Suppose(v
1,...,vn)spansV. We want to remove some
of the vectors from (v1,...,vn)so that the remaining vectors form a
basis ofV. We do this through the multistep process described below.
Start withB=(v1,...,vn).
Step 1
Ifv1=0, deletev1fromB.I fv1/negationslash=0, leaveBunchanged.
Step j
Ifvjis in span(v 1,...,vj−1), deletevjfromB.I fvjis not in
span(v 1,...,vj−1), leaveBunchanged.
Stop the process after step n, getting a list B. This listBspansV
because our original list spanned Band we have discarded only vectors
that were already in the span of the previous vectors. The process
Bases 29
insures that no vector in Bis in the span of the previous ones. Thus B
is linearly independent, by the linear dependence lemma (2.4). HenceBis a basis of V.
Consider the list
/parenleftbig
(1,2),(3, 6),(4, 7),(5, 9)/parenrightbig
,
which spans F2. To make sure that you understand the last proof, you
should verify that the process in the proof produces/parenleftbig
(1,2),(4, 7)/parenrightbig
,a
basis of F2, when applied to the list above.
Our next result, an easy corollary of the last theorem, tells us that
every finite-dimensional vector space has a basis.
2.11 Corollary: Every finite-dimensional vector space has a basis.
Proof: By definition, a finite-dimensional vector space has a span-
ning list. The previous theorem tells us that any spanning list can bereduced to a basis.
We have crafted our definitions so that the finite-dimensional vector
space{0}is not a counterexample to the corollary above. In particular,
the empty list ()is a basis of the vector space {0}because this list has
been defined to be linearly independent and to have span {0}.
Our next theorem is in some sense a dual of 2.10, which said that
every spanning list can be reduced to a basis. Now we show that givenany linearly independent list, we can adjoin some additional vectors so
that the extended list is still linearly independent but also spans the
space.
2.12 Theorem: Every linearly independent list of vectors in a finite-
This theorem can be
used to give another
proof of the previous
corollary. Specifically,
supposeVis finite
dimensional. Thistheorem implies that
the empty list ()can be
extended to a basis
ofV. In particular, V
has a basis.dimensional vector space can be extended to a basis of the vector space.
Proof: SupposeVis finite dimensional and (v1,...,vm)is linearly
independent in V. We want to extend (v1,...,vm)to a basis of V.W e
do this through the multistep process described below. First we let
(w1,...,wn)be any list of vectors in Vthat spansV.
Step 1
Ifw1is in the span of (v1,...,vm), letB=(v1,...,vm).I fw1is
not in the span of (v1,...,vm), letB=(v1,...,vm,w1).
30 Chapter 2.Finite-Dimensional Vector Spaces
Step j
Ifwjis in the span of B, leaveBunchanged. If wjis not in the
span ofB, extendBby adjoining wjto it.
After each step, Bis still linearly independent because otherwise the
linear dependence lemma (2.4) would give a contradiction (recall that(v
1,...,vm)is linearly independent and any wjthat is adjoined to Bis
not in the span of the previous vectors in B). After step n, the span of
Bincludes all the w’s. Thus the Bobtained after step nspansVand
hence is a basis of V.
As a nice application of the theorem above, we now show that ev-
ery subspace of a finite-dimensional vector space can be paired withanother subspace to form a direct sum of the whole space.
2.13 Proposition: SupposeVis finite dimensional and Uis a sub-
Using the same basic
ideas but considerably
more advanced tools,
this proposition can be
proved without the
hypothesis that Vis
finite dimensional.space ofV. Then there is a subspace WofVsuch thatV=U⊕W.
Proof: BecauseVis finite dimensional, so is U(see 2.7). Thus
there is a basis (u1,...,um)ofU(see 2.11). Of course (u1,...,um)
is a linearly independent list of vectors in V, and thus it can be ex-
tended to a basis (u1,...,um,w1,...,wn)ofV(see 2.12). Let W=
span(w 1,...,wn).
To prove that V=U⊕W, we need to show that
V=U+WandU∩W={0};
see 1.9. To prove the first equation, suppose that v∈V. Then,
because the list (u1,...,um,w1,...,wn)spansV, there exist scalars
a1,...,am,b1,...,bn∈Fsuch that
v=a1u1+···+amum/bracehtipupleft/bracehtipdownright/bracehtipdownleft/bracehtipupright
u+b1w1+···+bnwn/bracehtipupleft/bracehtipdownright/bracehtipdownleft/bracehtipupright
w.
In other words, we have v=u+w, whereu∈Uandw∈Ware defined
as above. Thus v∈U+W, completing the proof that V=U+W.
To show that U∩W={0}, suppose v∈U∩W. Then there exist
scalarsa1,...,am,b1,...,bn∈Fsuch that
v=a1u1+···+amum=b1w1+···+bnwn.
Thus
Dimension 31
a1u1+···+amum−b1w1−···−bnwn=0.
Because(u1,...,um,w1,...,wn)is linearly independent, this implies
thata1=···=a m=b1=···=bn=0. Thusv=0, completing the
proof thatU∩W={0}.
Dimension
Though we have been discussing finite-dimensional vector spaces,
we have not yet defined the dimension of such an object. How should
dimension be defined? A reasonable definition should force the dimen-sion of F
nto equaln. Notice that the basis
/parenleftbig
(1,0,...,0),(0, 1,0,...,0),...,( 0,...,0, 1)/parenrightbig
has lengthn. Thus we are tempted to define the dimension as the
length of a basis. However, a finite-dimensional vector space in generalhas many different bases, and our attempted definition makes senseonly if all bases in a given vector space have the same length. Fortu-nately that turns out to be the case, as we now show.
2.14 Theorem: Any two bases of a finite-dimensional vector space
have the same length.
Proof: SupposeVis finite dimensional. Let B
1andB2be any two
bases ofV. ThenB1is linearly independent in VandB2spansV, so the
length ofB1is at most the length of B2(by 2.6). Interchanging the roles
ofB1andB2, we also see that the length of B2is at most the length
ofB1. Thus the length of B1must equal the length of B2, as desired.
Now that we know that any two bases of a finite-dimensional vector
space have the same length, we can formally define the dimension ofsuch spaces. The dimension of a finite-dimensional vector space is
defined to be the length of any basis of the vector space. The dimensionofV(ifVis finite dimensional) is denoted by dim V. As examples, note
that dim F
n=nand dimPm(F)=m+1.
Every subspace of a finite-dimensional vector space is finite dimen-
sional (by 2.7) and so has a dimension. The next result gives the ex-
pected inequality about the dimension of a subspace.
32 Chapter 2.Finite-Dimensional Vector Spaces
2.15 Proposition: IfVis finite dimensional and Uis a subspace
ofV, then dimU≤dimV.
Proof: Suppose that Vis finite dimensional and Uis a subspace
ofV. Any basis of Uis a linearly independent list of vectors in Vand
thus can be extended to a basis of V(by 2.12). Hence the length of a
basis ofUis less than or equal to the length of a basis of V.
To check that a list of vectors in Vis a basis ofV, we must, according The real vector space
R2has dimension 2;
the complex vector
space Chas
dimension 1. As sets,
R2can be identified
with C(and addition is
the same on both
spaces, as is scalar
multiplication by real
numbers). Thus when
we talk about the
dimension of a vector
space, the role played
by the choice of F
cannot be neglected.to the definition, show that the list in question satisfies two properties:
it must be linearly independent and it must span V. The next two
results show that if the list in question has the right length, then weneed only check that it satisfies one of the required two properties.
We begin by proving that every spanning list with the right length is a
basis.
2.16 Proposition: IfVis finite dimensional, then every spanning
list of vectors in Vwith length dimVis a basis of V.
Proof: Suppose dim V=nand(v
1,...,vn)spansV. The list
(v1,...,vn)can be reduced to a basis of V(by 2.10). However, every
basis ofVhas lengthn, so in this case the reduction must be the trivial
one, meaning that no elements are deleted from (v1,...,vn). In other
words,(v1,...,vn)is a basis of V, as desired.
Now we prove that linear independence alone is enough to ensure
that a list with the right length is a basis.
2.17 Proposition: IfVis finite dimensional, then every linearly
independent list of vectors in Vwith length dimVis a basis of V.
Proof: Suppose dim V=nand(v1,...,vn)is linearly independent
inV. The list(v1,...,vn)can be extended to a basis of V(by 2.12). How-
ever, every basis of Vhas lengthn, so in this case the extension must be
the trivial one, meaning that no elements are adjoined to (v1,...,vn).
In other words, (v1,...,vn)is a basis of V, as desired.
As an example of how the last proposition can be applied, consider
the list/parenleftbig
(5,7),(4, 3)/parenrightbig
. This list of two vectors in F2is obviously linearly
independent (because neither vector is a scalar multiple of the other).
Dimension 33
Because F2has dimension 2, the last proposition implies that this lin-
early independent list of length 2 is a basis of F2(we do not need to
bother checking that it spans F2).
The next theorem gives a formula for the dimension of the sum of
two subspaces of a finite-dimensional vector space.
2.18 Theorem: IfU1andU2are subspaces of a finite-dimensional This formula for the
dimension of the sum
of two subspaces is
analogous to a familiar
counting formula: thenumber of elements in
the union of two finite
sets equals the numberof elements in the first
set, plus the number of
elements in the second
set, minus the numberof elements in the
intersection of the two
sets.vector space, then
dim(U 1+U2)=dimU1+dimU2−dim(U 1∩U2).
Proof: Let(u1,...,um)be a basis of U1∩U2; thus dim(U 1∩U2)=
m. Because(u1,...,um)is a basis ofU1∩U2, it is linearly independent
inU1and hence can be extended to a basis (u1,...,um,v1,...,vj)ofU1
(by 2.12). Thus dim U1=m+j. Also extend (u1,...,um)to a basis
(u1,...,um,w1,...,wk)ofU2; thus dimU2=m+k.
We will show that (u1,...,um,v1,...,vj,w1,...,wk)is a basis of
U1+U2. This will complete the proof because then we will have
dim(U 1+U2)=m+j+k
=(m+j)+(m+k)−m
=dimU1+dimU2−dim(U 1∩U2).
Clearly span(u 1,...,um,v1,...,vj,w1,...,wk)containsU1andU2
and hence contains U1+U2. So to show that this list is a basis of
U1+U2we need only show that it is linearly independent. To prove
this, suppose
a1u1+···+amum+b1v1+···+bjvj+c1w1+···+ckwk=0,
where all the a’s,b’s, andc’s are scalars. We need to prove that all the
a’s,b’s, andc’s equal 0. The equation above can be rewritten as
c1w1+···+ckwk=−a1u1−···−amum−b1v1−···−bjvj,
which shows that c1w1+···+ckwk∈U1. All thew’s are inU2, so this
implies that c1w1+···+c kwk∈U1∩U2. Because(u1,...,um)is a
basis ofU1∩U2, we can write
c1w1+···+ckwk=d1u1+···+dmum
34 Chapter 2.Finite-Dimensional Vector Spaces
for some choice of scalars d1,...,dm. But(u1,...,um,w1,...,wk)
is linearly independent, so the last equation implies that all the c’s
(andd’s) equal 0. Thus our original equation involving the a’s,b’s, and
c’s becomes
a1u1+···+amum+b1v1+···+bjvj=0.
This equation implies that all the a’s andb’s are 0 because the list
(u1,...,um,v1,...,vj)is linearly independent. We now know that all
thea’s,b’s, andc’s equal 0, as desired.
The next proposition shows that dimension meshes well with direct
sums. This result will be useful in later chapters.
2.19 Proposition: SupposeVis finite dimensional and U1,...,Um Recall that direct sum
is analogous to disjoint
union. Thus 2.19 is
analogous to the
statement that if a
finite setBis written as
A1∪···∪Amand the
sum of the number of
elements in the A’s
equals the number of
elements inB, then the
union is a disjoint
union.are subspaces of Vsuch that
2.20 V=U1+···+U m
and
2.21 dimV=dimU1+···+ dimUm.
ThenV=U1⊕···⊕U m.
Proof: Choose a basis for each Uj. Put these bases together in
one list, forming a list that spans V(by 2.20) and has length dim V
(by 2.21). Thus this list is a basis of V(by 2.16), and in particular it is
linearly independent.
Now suppose that u1∈U1,...,um∈Umare such that
0=u1+···+um.
We can write each ujas a linear combination of the basis vectors (cho-
sen above) of Uj. Substituting these linear combinations into the ex-
pression above, we have written 0 as a linear combination of the basisofVconstructed above. Thus all the scalars used in this linear combina-
tion must be 0. Thus each u
j=0, which proves that V=U1⊕···⊕Um
(by 1.8).
Exercises 35
Exercises
1. Prove that if (v1,...,vn)spansV, then so does the list
(v1−v2,v2−v3,...,vn−1−vn,vn)
obtained by subtracting from each vector (except the last one)
the following vector.
2. Prove that if (v1,...,vn)is linearly independent in V, then so is
the list
(v1−v2,v2−v3,...,vn−1−vn,vn)
obtained by subtracting from each vector (except the last one)
the following vector.
3. Suppose (v1,...,vn)is linearly independent in Vandw∈V.
Prove that if (v1+w,...,vn+w)is linearly dependent, then
w∈span(v 1,...,vn).
4. Suppose mis a positive integer. Is the set consisting of 0 and all
polynomials with coefficients in Fand with degree equal to ma
subspace of P(F)?
5. Prove that F∞is infinite dimensional.
6. Prove that the real vector space consisting of all continuous real-
valued functions on the interval [0,1]is infinite dimensional.
7. Prove that Vis infinite dimensional if and only if there is a se-
quencev1,v2,...of vectors in Vsuch that(v1,...,vn)is linearly
independent for every positive integer n.
8. LetUbe the subspace of R5defined by
U={(x1,x2,x3,x4,x5)∈R5:x1=3x2andx3=7x4}.
Find a basis of U.
9. Prove or disprove: there exists a basis (p0,p1,p2,p3)ofP3(F)
such that none of the polynomials p0,p1,p2,p3has degree 2.
10. Suppose that Vis finite dimensional, with dim V=n. Prove that
there exist one-dimensional subspaces U1,...,UnofVsuch that
V=U1⊕···⊕Un.
36 Chapter 2.Finite-Dimensional Vector Spaces
11. Suppose that Vis finite dimensional and Uis a subspace of V
such that dim U=dimV. Prove thatU=V.
12. Suppose that p0,p1,...,pmare polynomials in Pm(F)such that
pj(2)=0 for eachj. Prove that (p0,p1,...,pm)is not linearly
independent in Pm(F).
13. Suppose UandWare subspaces of R8such that dim U=3,
dimW=5, andU+W=R8. Prove thatU∩W={0}.
14. Suppose that UandWare both five-dimensional subspaces of R9.
Prove thatU∩W/negationslash={0}.
15. You might guess, by analogy with the formula for the number
of elements in the union of three subsets of a finite set, that
ifU1,U2,U3are subspaces of a finite-dimensional vector space,
then
dim(U 1+U2+U3)
=dimU1+dimU2+dimU3
−dim(U 1∩U2)−dim(U 1∩U3)−dim(U 2∩U3)
+dim(U 1∩U2∩U3).
Prove this or give a counterexample.
16. Prove that if Vis finite dimensional and U1,...,Umare subspaces
ofV, then
dim(U 1+···+Um)≤dimU1+···+ dimUm.
17. Suppose Vis finite dimensional. Prove that if U1,...,Umare
subspaces of Vsuch thatV=U1⊕···⊕Um, then
dimV=dimU1+···+ dimUm.
This exercise deepens the analogy between direct sums of sub-
spaces and disjoint unions of subsets. Specifically, compare thisexercise to the following obvious statement: if a finite set is writ-ten as a disjoint union of subsets, then the number of elements in
the set equals the sum of the number of elements in the disjoint
subsets.
Chapter 3
Linear Maps
So far our attention has focused on vector spaces. No one gets ex-
cited about vector spaces. The interesting part of linear algebra is thesubject to which we now turn—linear maps.
Let’s review our standing assumptions:
Recall that Fdenotes RorC.
Recall also that Vis a vector space over F.
In this chapter we will frequently need another vector space in ad-
dition toV. We will call this additional vector space W:
Let’s agree that for the rest of this chapter
Wwill denote a vector space over F.
✽✽✽
37
38 Chapter 3.Linear Maps
Definitions and Examples
Alinear map fromVtoWis a function T:V→Wwith the following Some mathematicians
use the term linear
transformation, which
means the same as
linear map.properties:
additivity
T(u+v)=Tu+Tvfor allu,v∈V;
homogeneity
T(av)=a(Tv) for alla∈Fand allv∈V.
Note that for linear maps we often use the notation Tvas well as the
more standard functional notation T(v) .
The set of all linear maps from VtoWis denotedL(V,W). Let’s
look at some examples of linear maps. Make sure you verify that each
of the functions defined below is indeed a linear map:
zero
In addition to its other uses, we let the symbol 0 denote the func-
tion that takes each element of some vector space to the additive
identity of another vector space. To be specific, 0 ∈L(V,W) is
defined by
0v=0.
Note that the 0 on the left side of the equation above is a function
fromVtoW, whereas the 0 on the right side is the additive iden-
tity inW. As usual, the context should allow you to distinguish
between the many uses of the symbol 0.
identity
The identity map , denotedI, is the function on some vector space
that takes each element to itself. To be specific, I∈L(V,V) is
defined by
Iv=v.
differentiation
DefineT∈L(P(R),P(R))by
Tp=p/prime.
The assertion that this function is a linear map is another way of
stating a basic result about differentiation: (f+g)/prime=f/prime+g/primeand
(af)/prime=af/primewheneverf,gare differentiable and ais a constant.
Definitions and Examples 39
integration
DefineT∈L(P(R),R)by
Tp=/integraldisplay1
0p(x)dx.
The assertion that this function is linear is another way of stating
a basic result about integration: the integral of the sum of twofunctions equals the sum of the integrals, and the integral of aconstant times a function equals the constant times the integral
of the function.
multiplication by x
2
DefineT∈L(P(R),P(R))by Though linear maps are
pervasive throughoutmathematics, they arenot as ubiquitous as
imagined by some
confused students whoseem to think that cos
is a linear map from R
toRwhen they write
“identities” such as
cos 2x=2 cosxand
cos(x+y)=
cosx+cosy.(Tp)(x)=x2p(x)
forx∈R.
backward shift
Recall that F∞denotes the vector space of all sequences of ele-
ments of F. DefineT∈L(F∞,F∞)by
T(x 1,x2,x3,...)=(x2,x3,...).
from Fnto Fm
DefineT∈L(R3,R2)by
T(x,y,z)=(2x−y+3z,7x+5y−6z).
More generally, let mandnbe positive integers, let aj,k∈Ffor
j=1,...,m andk=1,...,n , and define T∈L(Fn,Fm)by
T(x 1,...,xn)=(a1,1x1+···+a 1,nxn,...,am,1x1+···+am,nxn).
Later we will see that every linear map from FntoFmis of this
form.
Suppose(v1,...,vn)is a basis ofVandT:V→Wis linear. Ifv∈V,
then we can write vin the form
v=a1v1+···+anvn.
The linearity of Timplies that
40 Chapter 3.Linear Maps
Tv=a1Tv1+···+anTvn.
In particular, the values of Tv1,...,Tvndetermine the values of Ton
arbitrary vectors in V.
Linear maps can be constructed that take on arbitrary values on a
basis. Specifically, given a basis (v1,...,vn)ofVand any choice of
vectorsw1,...,wn∈W, we can construct a linear map T:V→Wsuch
thatTvj=wjforj=1,...,n . There is no choice of how to do this—we
must define Tby
T(a 1v1+···+anvn)=a1w1+···+anwn,
wherea1,...,anare arbitrary elements of F. Because(v1,...,vn)is a
basis ofV, the equation above does indeed define a function TfromV
toW. You should verify that the function Tdefined above is linear and
thatTvj=wjforj=1,...,n .
Now we will make L(V,W) into a vector space by defining addition
and scalar multiplication on it. For S,T∈L(V,W), define a function
S+T∈L(V,W) in the usual manner of adding functions:
(S+T)v=Sv+Tv
forv∈V. You should verify that S+Tis indeed a linear map from V
toWwheneverS,T∈L(V,W). For a∈FandT∈L(V,W), define a
functionaT∈L(V,W) in the usual manner of multiplying a function
by a scalar:
(aT)v=a(Tv)
forv∈V. You should verify that aTis indeed a linear map from VtoW
whenevera∈FandT∈L(V,W). With the operations we have just
defined,L(V,W) becomes a vector space (as you should verify). Note
that the additive identity of L(V,W) is the zero linear map defined
earlier in this section.
Usually it makes no sense to multiply together two elements of a
vector space, but for some pairs of linear maps a useful product exists.
We will need a third vector space, so suppose Uis a vector space over F.
IfT∈L(U,V) andS∈L(V,W), then we define ST∈L(U,W) by
(ST)(v)=S(Tv)
forv∈U. In other words, STis just the usual composition S◦Tof two
functions, but when both functions are linear, most mathematicians
Null Spaces and Ranges 41
writeSTinstead ofS◦T. You should verify that STis indeed a linear
map fromUtoWwheneverT∈L(U,V) andS∈L(V,W). Note that
STis defined only when Tmaps into the domain of S. We often call
STthe product ofSandT. You should verify that it has most of the
usual properties expected of a product:
associativity
(T1T2)T3=T1(T2T3)wheneverT1,T2, andT3are linear maps such
that the products make sense (meaning that T3must map into the
domain ofT2, andT2must map into the domain of T1).
identity
TI=TandIT=TwheneverT∈L(V,W) (note that in the first
equationIis the identity map on V, and in the second equation I
is the identity map on W).
distributive properties
(S1+S2)T=S1T+S2TandS(T 1+T2)=ST1+ST2whenever
T,T 1,T2∈L(U,V) andS,S 1,S2∈L(V,W).
Multiplication of linear maps is not commutative. In other words, it
is not necessarily true that ST=TS, even if both sides of the equation
make sense. For example, if T∈L(P(R),P(R))is the differentiation
map defined earlier in this section and S∈L(P(R),P(R))is the mul-
tiplication by x2map defined earlier in this section, then
((ST)p)(x)=x2p/prime(x) but((TS)p)(x)=x2p/prime(x)+2xp(x).
In other words, multiplying by x2and then differentiating is not the
same as differentiating and then multiplying by x2.
Null Spaces and Ranges
ForT∈L(V,W), the null space ofT, denoted null T, is the subset Some mathematicians
use the term kernel
instead of null space.ofVconsisting of those vectors that Tmaps to 0:
nullT={v∈V:Tv=0}.
Let’s look at a few examples from the previous section. In the dif-
ferentiation example, we defined T∈L(P(R),P(R))byTp=p/prime. The
42 Chapter 3.Linear Maps
only functions whose derivative equals the zero function are the con-
stant functions, so in this case the null space of Tequals the set of
constant functions.
In the multiplication by x2example, we defined T∈L(P(R),P(R))
by(Tp)(x)=x2p(x) . The only polynomial psuch thatx2p(x)=0
for allx∈Ris the 0 polynomial. Thus in this case we have
nullT={0}.
In the backward shift example, we defined T∈L(F∞,F∞)by
T(x 1,x2,x3,...)=(x2,x3,...).
ClearlyT(x 1,x2,x3,...) equals 0 if and only if x2,x3,...are all 0. Thus
in this case we have
nullT={(a,0,0,...) :a∈F}.
The next proposition shows that the null space of any linear map is
a subspace of the domain. In particular, 0 is in the null space of every
linear map.
3.1 Proposition: IfT∈L(V,W), then nullTis a subspace of V.
Proof: SupposeT∈L(V,W). By additivity, we have
T(0)=T(0+0)=T(0)+T(0),
which implies that T(0)=0. Thus 0∈nullT.
Ifu,v∈nullT, then
T(u+v)=Tu+Tv=0+0=0,
and henceu+v∈nullT. Thus nullTis closed under addition.
Ifu∈nullTanda∈F, then
T(au)=aTu=a0=0,
and henceau∈nullT. Thus nullTis closed under scalar multiplica-
tion.
We have shown that null Tcontains 0 and is closed under addition
and scalar multiplication. Thus null Tis a subspace of V.
Null Spaces and Ranges 43
A linear map T:V→Wis called injective if whenever u,v∈V Many mathematicians
use the term
one-to-one, which
means the same asinjective.andTu=Tv, we haveu=v. The next proposition says that we
can check whether a linear map is injective by checking whether 0 isthe only vector that gets mapped to 0. As a simple application of thisproposition, we see that of the three linear maps whose null spaces wecomputed earlier in this section (differentiation, multiplication by x
2,
and backward shift), only multiplication by x2is injective.
3.2 Proposition: LetT∈L(V,W). Then Tis injective if and only
ifnullT={0}.
Proof: First suppose that Tis injective. We want to prove that
nullT={0}. We already know that {0}⊂nullT(by 3.1). To prove the
inclusion in the other direction, suppose v∈nullT. Then
T(v)=0=T(0).
BecauseTis injective, the equation above implies that v=0. Thus
nullT={0}, as desired.
To prove the implication in the other direction, now suppose that
nullT={0}. We want to prove that Tis injective. To do this, suppose
u,v∈VandTu=Tv. Then
0=Tu−Tv=T(u−v).
Thusu−vis in nullT, which equals {0}. Hence u−v=0, which
implies that u=v. HenceTis injective, as desired.
ForT∈L(V,W), the range ofT, denoted range T, is the subset of Some mathematicians
use the word image,
which means the same
as range.Wconsisting of those vectors that are of the form Tvfor somev∈V:
rangeT={Tv:v∈V}.
For example, if T∈L(P(R),P(R))is the differentiation map defined by
Tp=p/prime, then rangeT=P(R)because for every polynomial q∈P(R)
there exists a polynomial p∈P(R)such thatp/prime=q.
As another example, if T∈L(P(R),P(R))is the linear map of
multiplication by x2defined by(Tp)(x)=x2p(x) , then the range
ofTis the set of polynomials of the form a2x2+···+a mxm, where
a2,...,am∈R.
The next proposition shows that the range of any linear map is a
subspace of the target space.
44 Chapter 3.Linear Maps
3.3 Proposition: IfT∈L(V,W), then rangeTis a subspace of W.
Proof: SupposeT∈L(V,W). Then T(0)=0 (by 3.1), which im-
plies that 0∈rangeT.
Ifw1,w2∈rangeT, then there exist v1,v2∈Vsuch thatTv1=w1
andTv2=w2. Thus
T(v 1+v2)=Tv1+Tv2=w1+w2,
and hencew1+w2∈rangeT. Thus range Tis closed under addition.
Ifw∈rangeTanda∈F, then there exists v∈Vsuch thatTv=w.
Thus
T(av)=aTv=aw,
and henceaw∈rangeT. Thus range Tis closed under scalar multipli-
cation.
We have shown that range Tcontains 0 and is closed under addition
and scalar multiplication. Thus range Tis a subspace of W.
A linear map T:V→Wis called surjective if its range equals W. Many mathematicians
use the term onto,
which means the same
as surjective.For example, the differentiation map T∈L(P(R),P(R))defined by
Tp=p/primeis surjective because its range equals P(R). As another exam-
ple, the linear map T∈L(P(R),P(R))defined by(Tp)(x)=x2p(x) is
not surjective because its range does not equal P(R). As a final exam-
ple, you should verify that the backward shift T∈L(F∞,F∞)defined
by
T(x 1,x2,x3,...)=(x2,x3,...)
is surjective.
Whether a linear map is surjective can depend upon what we are
thinking of as the target space. For example, fix a positive integer m.
The differentiation map T∈L(Pm(R),Pm(R))defined byTp=p/prime
is not surjective because the polynomial xmis not in the range of T.
However, the differentiation map T∈L(Pm(R),Pm−1(R))defined by
Tp=p/primeis surjective because its range equals Pm−1(R), which is now
the target space.
The next theorem, which is the key result in this chapter, states that
the dimension of the null space plus the dimension of the range of a
linear map on a finite-dimensional vector space equals the dimension
of the domain.
Null Spaces and Ranges 45
3.4 Theorem: IfVis finite dimensional and T∈L(V,W), then
rangeTis a finite-dimensional subspace of Wand
dimV=dim nullT+dim rangeT.
Proof: Suppose that Vis a finite-dimensional vector space and
T∈L(V,W). Let(u1,...,um)be a basis of null T; thus dim null T=m.
The linearly independent list (u1,...,um)can be extended to a ba-
sis(u1,...,um,w1,...,wn)ofV(by 2.12). Thus dim V=m+n,
and to complete the proof, we need only show that range Tis finite
dimensional and dim range T=n. We will do this by proving that
(Tw 1,...,Twn)is a basis of range T.
Letv∈V. Because(u1,...,um,w1,...,wn)spansV, we can write
v=a1u1+···+amum+b1w1+···+bnwn,
where thea’s andb’s are in F. ApplyingTto both sides of this equation,
we get
Tv=b1Tw 1+···+bnTwn,
where the terms of the form Tujdisappeared because each uj∈nullT.
The last equation implies that (Tw 1,...,Twn)spans range T. In par-
ticular, range Tis finite dimensional.
To show that (Tw 1,...,Twn)is linearly independent, suppose that
c1,...,cn∈Fand
c1Tw 1+···+cnTwn=0.
Then
T(c 1w1+···+cnwn)=0,
and hence
c1w1+···+cnwn∈nullT.
Because(u1,...,um)spans nullT, we can write
c1w1+···+cnwn=d1u1+···+dmum,
where thed’s are in F. This equation implies that all the c’s (andd’s)
are 0 (because (u1,...,um,w1,...,wn)is linearly independent). Thus
(Tw 1,...,Twn)is linearly independent and hence is a basis for range T,
as desired.
46 Chapter 3.Linear Maps
Now we can show that no linear map from a finite-dimensional vec-
tor space to a “smaller” vector space can be injective, where “smaller”is measured by dimension.
3.5 Corollary: IfVandWare finite-dimensional vector spaces such
that dimV>dimW, then no linear map from VtoWis injective.
Proof: SupposeVandWare finite-dimensional vector spaces such
that dimV>dimW. LetT∈L(V,W). Then
dim nullT=dimV−dim rangeT
≥dimV−dimW
>0,
where the equality above comes from 3.4. We have just shown that
dim nullT> 0. This means that null Tmust contain vectors other
than 0. Thus Tis not injective (by 3.2).
The next corollary, which is in some sense dual to the previous corol-
lary, shows that no linear map from a finite-dimensional vector spaceto a “bigger” vector space can be surjective, where “bigger” is measuredby dimension.
3.6 Corollary: IfVandWare finite-dimensional vector spaces such
that dimV<dimW, then no linear map from VtoWis surjective.
Proof: SupposeVandWare finite-dimensional vector spaces such
that dimV<dimW. LetT∈L(V,W). Then
dim rangeT=dimV−dim nullT
≤dimV
<dimW,
where the equality above comes from 3.4. We have just shown that
dim rangeT<dimW. This means that range Tcannot equal W. Thus
Tis not surjective.
The last two corollaries have important consequences in the theory
of linear equations. To see this, fix positive integers mandn, and let
aj,k∈Fforj=1,...,m andk=1,...,n . DefineT:Fn→Fmby
Null Spaces and Ranges 47
T(x 1,...,xn)=/parenleftbign/summationdisplay
k=1a1,kxk,...,n/summationdisplay
k=1am,kxk/parenrightbig
.
Now consider the equation Tx=0 (wherex∈Fnand the 0 here is
the additive identity in Fm, namely, the list of length mconsisting of
all 0’s). Letting x=(x1,...,xn), we can rewrite the equation Tx=0
as a system of homogeneous equations: Homogeneous, in this
context, means that the
constant term on the
right side of each
equation equals 0.n/summationdisplay
k=1a1,kxk=0
...
n/summationdisplay
k=1am,kxk=0.
We think of the a’s as known; we are interested in solutions for the
variablesx1,...,xn. Thus we have mequations and nvariables. Obvi-
ouslyx1=···=x n=0 is a solution; the key question here is whether
any other solutions exist. In other words, we want to know if null Tis
strictly bigger than {0}. This happens precisely when Tis not injective
(by 3.2). From 3.5 we see that Tis not injective if n>m . Conclusion:
a homogeneous system of linear equations in which there are more
variables than equations must have nonzero solutions.
WithTas in the previous paragraph, now consider the equation
Tx=c, wherec=(c1,...,cm)∈Fm. We can rewrite the equation
Tx=cas a system of inhomogeneous equations:
These results about
homogeneous systems
with more variables
than equations andinhomogeneous
systems with more
equations than
variables are often
proved using Gaussian
elimination. Theabstract approach
taken here leads to
cleaner proofs.n/summationdisplay
k=1a1,kxk=c1
...
n/summationdisplay
k=1am,kxk=cm.
As before, we think of the a’s as known. The key question here is
whether for every choice of the constant terms c1,...,cm∈F, there
exists at least one solution for the variables x1,...,xn. In other words,
we want to know whether range Tequals Fm. From 3.6 we see that T
is not surjective if n<m . Conclusion: an inhomogeneous system of
linear equations in which there are more equations than variables has
no solution for some choice of the constant terms.
48 Chapter 3.Linear Maps
The Matrix of a Linear Map
We have seen that if (v1,...,vn)is a basis of VandT:V→Wis
linear, then the values of Tv1,...,Tvndetermine the values of Ton
arbitrary vectors in V. In this section we will see how matrices are used
as an efficient method of recording the values of the Tvj’s in terms of
a basis ofW.
Letmandndenote positive integers. An m-by-n matrix is a rect-
angular array with mrows andncolumns that looks like this:
3.7
a1,1... a 1,n
......
am,1... am,n
.
Note that the first index refers to the row number and the second in-
dex refers to the column number. Thus a3,2refers to the entry in the
third row, second column of the matrix above. We will usually considermatrices whose entries are elements of F.
LetT∈L(V,W). Suppose that (v
1,...,vn)is a basis of Vand
(w1,...,wm)is a basis of W. For eachk=1,...,n , we can write Tvk
uniquely as a linear combination of the w’s:
3.8 Tvk=a1,kw1+···+am,kwm,
whereaj,k∈Fforj=1,...,m . The scalars aj,kcompletely determine
the linear map Tbecause a linear map is determined by its values on
a basis. The m-by-n matrix 3.7 formed by the a’s is called the matrix
ofTwith respect to the bases (v1,...,vn)and(w1,...,wm); we denote
it by
M/parenleftbig
T,(v 1,...,vn),(w 1,...,wm)/parenrightbig
.
If the bases (v1,...,vn)and(w1,...,wm)are clear from the context
(for example, if only one set of bases is in sight), we write just M(T)
instead ofM/parenleftbig
T,(v 1,...,vn),(w 1,...,wm)/parenrightbig
.
As an aid to remembering how M(T) is constructed from T, you
might write the basis vectors v1,...,vnfor the domain across the top
and the basis vectors w1,...,wmfor the target space along the left, as
follows:
The Matrix of a Linear Map 49
v1... vk... vn
w1
...
wm
a1,k
...
am,k
Note that in the matrix above only the kthcolumn is displayed (and thus With respect to any
choice of bases, thematrix of the 0linear
map (the linear map
that takes every vector
to0) consists of all 0’s.the second index of each displayed aisk). Thekthcolumn ofM(T)
consists of the scalars needed to write Tvkas a linear combination of
thew’s. Thus the picture above should remind you that Tvkis retrieved
from the matrix M(T) by multiplying each entry in the kthcolumn by
the corresponding wfrom the left column, and then adding up the
resulting vectors.
IfTis a linear map from FntoFm, then unless stated otherwise you
should assume that the bases in question are the standard ones (wherethek
thbasis vector is 1 in the kthslot and 0 in all the other slots). If
you think of elements of Fmas columns of mnumbers, then you can
think of the kthcolumn ofM(T) asTapplied to the kthbasis vector.
For example, if T∈L(F2,F3)is defined by
T(x,y)=(x+3y,2x+5y,7x+9y),
thenT(1,0)=(1,2,7)andT(0,1)=(3,5,9), so the matrix of T(with
respect to the standard bases) is the 3-by-2 matrix
13
25
79
.
Suppose we have bases (v1,...,vn)ofVand(w1,...,wm)ofW.
Thus for each linear map from VtoW, we can talk about its matrix
(with respect to these bases, of course). Is the matrix of the sum of twolinear maps equal to the sum of the matrices of the two maps?
Right now this question does not make sense because, though we
have defined the sum of two linear maps, we have not defined the sumof two matrices. Fortunately the obvious definition of the sum of twomatrices has the right properties. Specifically, we define addition ofmatrices of the same size by adding corresponding entries in the ma-trices:
50 Chapter 3.Linear Maps
a1,1... a 1,n
......
am,1... am,n
+
b1,1... b 1,n
......
bm,1... bm,n
=
a1,1+b1,1... a 1,n+b1,n
......
am,1+bm,1... am,n+bm,n
.
You should verify that with this definition of matrix addition,
3.9 M(T+S)=M(T)+M(S)
wheneverT,S∈L(V,W).
Still assuming that we have some bases in mind, is the matrix of a
scalar times a linear map equal to the scalar times the matrix of the
linear map? Again the question does not make sense because we have
not defined scalar multiplication on matrices. Fortunately the obviousdefinition again has the right properties. Specifically, we define theproduct of a scalar and a matrix by multiplying each entry in the matrixby the scalar:
c
a
1,1... a 1,n
......
am,1... am,n
=
ca1,1... ca 1,n
......
cam,1... cam,n
.
You should verify that with this definition of scalar multiplication on
matrices,
3.10 M(cT)=cM(T)
wheneverc∈FandT∈L(V,W).
Because addition and scalar multiplication have now been defined
for matrices, you should not be surprised that a vector space is aboutto appear. We need only a bit of notation so that this new vector spacehas a name. The set of all m-by-n matrices with entries in Fis denoted
by Mat(m,n, F). You should verify that with addition and scalar mul-
tiplication defined as above, Mat (m,n, F)is a vector space. Note that
the additive identity in Mat (m,n, F)is them-by-n matrix all of whose
entries equal 0.
Suppose(v
1,...,vn)is a basis ofVand(w1,...,wm)is a basis ofW.
Suppose also that we have another vector space Uand that(u1,...,up)
The Matrix of a Linear Map 51
is a basis of U. Consider linear maps S:U→VandT:V→W. The
composition TSis a linear map from UtoW. How canM(TS) be
computed from M(T) andM(S)? The nicest solution to this question
would be to have the following pretty relationship:
3.11 M(TS)=M(T)M(S).
So far, however, the right side of this equation does not make sense
because we have not yet defined the product of two matrices. We willchoose a definition of matrix multiplication that forces the equation
above to hold. Let’s see how to do this.
Let
M(T)=
a
1,1... a 1,n
......
am,1... am,n
andM(S)=
b1,1... b 1,p
......
bn,1... bn,p
.
Fork∈{1,...,p}, we have
TSuk=T(n/summationdisplay
r=1br,kvr)
=n/summationdisplay
r=1br,kTvr
=n/summationdisplay
r=1br,km/summationdisplay
j=1aj,rwj
=m/summationdisplay
j=1(n/summationdisplay
r=1aj,rbr,k)wj.
ThusM(TS) is them-by-p matrix whose entry in row j, columnk
equals/summationtextn
r=1aj,rbr,k.
Now it’s clear how to define matrix multiplication so that 3.11 holds. You probably learned
this definition of matrix
multiplication in an
earlier course, althoughyou may not have seenthis motivation for it.Namely, ifAis anm-by-n matrix with entries aj,kandBis ann-by-p
matrix with entries bj,k, thenABis defined to be the m-by-p matrix
whose entry in row j, columnk, equals
n/summationdisplay
r=1aj,rbr,k.
In other words, the entry in row j, columnk,o fABis computed by
taking rowjofAand column kofB, multiplying together correspond-
ing entries, and then summing. Note that we define the product of two
52 Chapter 3.Linear Maps
matrices only when the number of columns of the first matrix equals
the number of rows of the second matrix.
As an example of matrix multiplication, here we multiply together You should find an
example to show that
matrix multiplication is
not commutative. In
other words, AB is not
necessarily equal to BA,
even when both are
defined.a 3-by-2 matrix and a 2-by-4 matrix, obtaining a 3-by-4 matrix:
12
3456
/bracketleftBigg
654 3
210−1/bracketrightBigg
=
10 7 4 1
26 19 12 542 31 20 9
.
Suppose(v
1,...,vn)is a basis ofV.I fv∈V, then there exist unique
scalarsb1,...,bnsuch that
3.12 v=b1v1+···+bnvn.
The matrix ofv, denotedM(v), is the n-by-1 matrix defined by
3.13 M(v)=
b1
...
bn
.
Usually the basis is obvious from the context, but when the basis needs
to be displayed explicitly use the notation M/parenleftbig
v,(v 1,...,vn)/parenrightbig
instead
ofM(v).
For example, the matrix of a vector x∈Fnwith respect to the stan-
dard basis is obtained by writing the coordinates of xas the entries in
ann-by-1 matrix. In other words, if x=(x1,...,xn)∈Fn, then
M(x)=
x1
...
xn
.
The next proposition shows how the notions of the matrix of a linear
map, the matrix of a vector, and matrix multiplication fit together. In
this proposition M(Tv) is the matrix of the vector Tvwith respect to
the basis(w1,...,wm)andM(v) is the matrix of the vector vwith re-
spect to the basis (v1,...,vn), whereasM(T) is the matrix of the linear
mapTwith respect to the bases (v1,...,vn)and(w1,...,wm).
3.14 Proposition: SupposeT∈L(V,W) and(v1,...,vn)is a basis
ofVand(w1,...,wm)is a basis of W. Then
M(Tv)=M(T)M(v)
for everyv∈V.
Invertibility 53
Proof: Let
3.15 M(T)=
a1,1... a 1,n
......
am,1... am,n
.
This means, we recall, that
3.16 Tvk=m/summationdisplay
j=1aj,kwj
for eachk. Letvbe an arbitrary vector in V, which we can write in the
form 3.12. Thus M(v) is given by 3.13. Now
Tv=b1Tv1+···+bnTvn
=b1m/summationdisplay
j=1aj,1wj+···+bnm/summationdisplay
j=1aj,nwj
=m/summationdisplay
j=1(aj,1b1+···+aj,nbn)wj,
where the first equality comes from 3.12 and the second equality comes
from 3.16. The last equation shows that M(Tv), the m-by-1 matrix of
the vectorTvwith respect to the basis (w1,...,wm), is given by the
equation
M(Tv)=
a1,1b1+···+a1,nbn
...
am,1b1+···+am,nbn
.
This formula, along with the formulas 3.15 and 3.13 and the definition
of matrix multiplication, shows that M(Tv)=M(T)M(v).
Invertibility
A linear map T∈L(V,W) is called invertible if there exists a linear
mapS∈L(W,V) such thatSTequals the identity map on VandTS
equals the identity map on W. A linear map S∈L(W,V) satisfying
ST=IandTS=Iis called an inverse ofT(note that the first Iis the
identity map on Vand the second Iis the identity map on W).
IfSandS/primeare inverses of T, then
54 Chapter 3.Linear Maps
S=SI=S(TS/prime)=(ST)S/prime=IS/prime=S/prime,
soS=S/prime. In other words, if Tis invertible, then it has a unique
inverse, which we denote by T−1. Rephrasing all this once more, if
T∈L(V,W) is invertible, then T−1is the unique element of L(W,V)
such thatT−1T=IandTT−1=I. The following proposition charac-
terizes the invertible linear maps.
3.17 Proposition: A linear map is invertible if and only if it is injec-
tive and surjective.
Proof: SupposeT∈L(V,W). We need to show that Tis invertible
if and only if it is injective and surjective.
First suppose that Tis invertible. To show that Tis injective, sup-
pose thatu,v∈VandTu=Tv. Then
u=T−1(Tu)=T−1(Tv)=v,
sou=v. HenceTis injective.
We are still assuming that Tis invertible. Now we want to prove
thatTis surjective. To do this, let w∈W. Thenw=T(T−1w), which
shows thatwis in the range of T. Thus range T=W, and henceTis
surjective, completing this direction of the proof.
Now suppose that Tis injective and surjective. We want to prove
thatTis invertible. For each w∈W, defineSwto be the unique ele-
ment ofVsuch thatT(Sw)=w(the existence and uniqueness of such
an element follow from the surjectivity and injectivity of T). Clearly
TSequals the identity map on W. To prove that STequals the identity
map onV, letv∈V. Then
T(STv)=(TS)(Tv)=I(Tv)=Tv.
This equation implies that STv=v(becauseTis injective), and thus
STequals the identity map on V. To complete the proof, we need to
show thatSis linear. To do this, let w1,w2∈W. Then
T(Sw 1+Sw2)=T(Sw 1)+T(Sw 2)=w1+w2.
ThusSw1+Sw2is the unique element of VthatTmaps tow1+w2.B y
the definition of S, this implies that S(w 1+w2)=Sw1+Sw2. Hence
Ssatisfies the additive property required for linearity. The proof of
homogeneity is similar. Specifically, if w∈Wanda∈F, then
Invertibility 55
T(aSw)=aT(Sw)=aw.
ThusaSw is the unique element of VthatTmaps toaw. By the
definition of S, this implies that S(aw)=aSw . HenceSis linear, as
desired.
Two vector spaces are called isomorphic if there is an invertible The Greek word isos
means equal; the Greekword morph means
shape. Thus
isomorphic literally
means equal shape.linear map from one vector space onto the other one. As abstract vector
spaces, two isomorphic spaces have the same properties. From this
viewpoint, you can think of an invertible linear map as a relabeling ofthe elements of a vector space.
If two vector spaces are isomorphic and one of them is finite dimen-
sional, then so is the other one. To see this, suppose that VandW
are isomorphic and that T∈L(V,W) is an invertible linear map. If V
is finite dimensional, then so is W(by 3.4). The same reasoning, with
Treplaced with T
−1∈L(W,V), shows that if Wis finite dimensional,
then so isV. Actually much more is true, as the following theorem
shows.
3.18 Theorem: Two finite-dimensional vector spaces are isomorphic
if and only if they have the same dimension.
Proof: First suppose VandWare isomorphic finite-dimensional
vector spaces. Thus there exists an invertible linear map TfromV
ontoW. BecauseTis invertible, we have null T={0}and rangeT=W.
Thus dim null T=0 and dim range T=dimW. The formula
dimV=dim nullT+dim rangeT
(see 3.4) thus becomes the equation dim V=dimW, completing the
proof in one direction.
To prove the other direction, suppose VandWare finite-dimen-
sional vector spaces with the same dimension. Let (v1,...,vn)be a
basis ofVand(w1,...,wn)be a basis of W. LetTbe the linear map
fromVtoWdefined by
T(a 1v1+···+anvn)=a1w1+···+anwn.
ThenTis surjective because (w1,...,wn)spansW, andTis injective
because(w1,...,wn)is linearly independent. Because Tis injective and
56 Chapter 3.Linear Maps
surjective, it is invertible (see 3.17), and hence VandWare isomorphic,
as desired.
The last theorem implies that every finite-dimensional vector space Because every
finite-dimensional
vector space is
isomorphic to some Fn,
why bother with
abstract vector spaces?
To answer this
question, note that an
investigation of Fn
would soon lead to
vector spaces that do
not equal Fn. For
example, we would
encounter the null
space and range of
linear maps, the set of
matrices Mat(n,n, F),
and the polynomials
Pn(F). Though each of
these vector spaces is
isomorphic to some
Fm, thinking of them
that way often adds
complexity but no new
insight.is isomorphic to some Fn. Specifically, if Vis a finite-dimensional vector
space and dim V=n, thenVand Fnare isomorphic.
If(v1,...,vn)is a basis of Vand(w1,...,wm)is a basis of W, then
for eachT∈L(V,W), we have a matrix M(T)∈Mat(m,n, F). In other
words, once bases have been fixed for VandW,Mbecomes a function
fromL(V,W) to Mat(m,n, F). Notice that 3.9 and 3.10 show that Mis
a linear map. This linear map is actually invertible, as we now show.
3.19 Proposition: Suppose that (v1,...,vn)is a basis of Vand
(w1,...,wm)is a basis of W. ThenMis an invertible linear map be-
tweenL(V,W) andMat(m,n, F).
Proof: We have already noted that Mis linear, so we need only
prove thatMis injective and surjective (by 3.17). Both are easy. Let’s
begin with injectivity. If T∈L(V,W) andM(T)=0, thenTvk=0
fork=1,...,n . Because(v1,...,vn)is a basis of V, this implies that
T=0. ThusMis injective (by 3.2).
To prove that Mis surjective, let
A=
a1,1... a 1,n
......
am,1... am,n
be a matrix in Mat (m,n, F). LetTbe the linear map from VtoWsuch
that
Tvk=m/summationdisplay
j=1aj,kwj
fork=1,...,n . Obviously M(T) equalsA, and so the range of M
equals Mat(m,n, F), as desired.
An obvious basis of Mat (m,n, F)consists of those m-by-n matrices
that have 0 in all entries except fo ra1i no n e entry. There are mnsuch
matrices, so the dimension of Mat (m,n, F)equalsmn.
Now we can determine the dimension of the vector space of linear
maps from one finite-dimensional vector space to another.
Invertibility 57
3.20 Proposition: IfVandWare finite dimensional, then L(V,W)
is finite dimensional and
dimL(V,W)=(dimV)(dimW).
Proof: This follows from the equation dim Mat (m,n, F)=mn,
3.18, and 3.19.
A linear map from a vector space to itself is called an operator .I f The deepest and most
important parts of
linear algebra, as wellas most of the rest ofthis book, deal with
operators.we want to specify the vector space, we say that a linear map T:V→V
is an operator on V. Because we are so often interested in linear maps
from a vector space into itself, we use the notation L(V) to denote the
set of all operators on V. In other words, L(V)=L(V,V).
Recall from 3.17 that a linear map is invertible if it is injective and
surjective. For a linear map of a vector space into itself, you mightwonder whether injectivity alone, or surjectivity alone, is enough toimply invertibility. On infinite-dimensional vector spaces neither con-dition alone implies invertibility. We can see this from some exampleswe have already considered. The multiplication by x
2operator (from
P(R)to itself) is injective but not surjective. The backward shift (from
F∞to itself) is surjective but not injective. In view of these examples,
the next theorem is remarkable—it states that for maps from a finite-dimensional vector space to itself, either injectivity or surjectivity aloneimplies the other condition.
3.21 Theorem: SupposeVis finite dimensional. If T∈L(V), then
the following are equivalent:
(a)Tis invertible;
(b)Tis injective;
(c)Tis surjective.
Proof: SupposeT∈L(V). Clearly (a) implies (b).
Now suppose (b) holds, so that Tis injective. Thus null T={0}
(by 3.2). From 3.4 we have
dim rangeT=dimV−dim nullT
=dimV,
which implies that range TequalsV(see Exercise 11 in Chapter 2). Thus
Tis surjective. Hence (b) implies (c).
58 Chapter 3.Linear Maps
Now suppose (c) holds, so that Tis surjective. Thus range T=V.
From 3.4 we have
dim nullT=dimV−dim rangeT
=0,
which implies that null Tequals{0}. ThusTis injective (by 3.2), and
soTis invertible (we already knew that Twas surjective). Hence (c)
implies (a), completing the proof.
Exercises 59
Exercises
1. Show that every linear map from a one-dimensional vector space
to itself is multiplication by some scalar. More precisely, provethat if dimV=1 andT∈L(V,V), then there exists a∈Fsuch
thatTv=avfor allv∈V.
2. Give an example of a function f:R
2→Rsuch that Exercise 2 shows that
homogeneity alone is
not enough to implythat a function is alinear map. Additivityalone is also not
enough to imply that a
function is a linear
map, although theproof of this involves
advanced tools that arebeyond the scope of
this book.f(av)=af(v)
for alla∈Rand allv∈R2butfis not linear.
3. Suppose that Vis finite dimensional. Prove that any linear map
on a subspace of Vcan be extended to a linear map on V.I n
other words, show that if Uis a subspace of VandS∈L(U,W),
then there exists T∈L(V,W) such thatTu=Sufor allu∈U.
4. Suppose that Tis a linear map from VtoF. Prove that if u∈V
is not in null T, then
V=nullT⊕{au:a∈F}.
5. Suppose that T∈L(V,W) is injective and (v1,...,vn)is linearly
independent in V. Prove that(Tv 1,...,Tvn)is linearly indepen-
dent inW.
6. Prove that if S1,...,Snare injective linear maps such that S1...Sn
makes sense, then S1...Snis injective.
7. Prove that if (v1,...,vn)spansVandT∈L(V,W) is surjective,
then(Tv 1,...,Tvn)spansW.
8. Suppose that Vis finite dimensional and that T∈L(V,W). Prove
that there exists a subspace UofVsuch thatU∩nullT={0}
and rangeT={Tu:u∈U}.
9. Prove that if Tis a linear map from F4toF2such that
nullT={(x1,x2,x3,x4)∈F4:x1=5x2andx3=7x4},
thenTis surjective.
60 Chapter 3.Linear Maps
10. Prove that there does not exist a linear map from F5toF2whose
null space equals
{(x 1,x2,x3,x4,x5)∈F5:x1=3x2andx3=x4=x5}.
11. Prove that if there exists a linear map on Vwhose null space and
range are both finite dimensional, then Vis finite dimensional.
12. Suppose that VandWare both finite dimensional. Prove that
there exists a surjective linear map from VontoWif and only if
dimW≤dimV.
13. Suppose that VandWare finite dimensional and that Uis a
subspace of V. Prove that there exists T∈L(V,W) such that
nullT=Uif and only if dim U≥dimV−dimW.
14. Suppose that Wis finite dimensional and T∈L(V,W). Prove
thatTis injective if and only if there exists S∈L(W,V) such
thatSTis the identity map on V.
15. Suppose that Vis finite dimensional and T∈L(V,W). Prove
thatTis surjective if and only if there exists S∈L(W,V) such
thatTSis the identity map on W.
16. Suppose that UandVare finite-dimensional vector spaces and
thatS∈L(V,W),T∈L(U,V). Prove that
dim nullST≤dim nullS+dim nullT.
17. Prove that the distributive property holds for matrix addition
and matrix multiplication. In other words, suppose A,B, andC
are matrices whose sizes are such that A(B+C)makes sense.
Prove thatAB+ACmakes sense and that A(B+C)=AB+AC.
18. Prove that matrix multiplication is associative. In other words,
supposeA,B, andCare matrices whose sizes are such that
(AB)C makes sense. Prove that A(BC) makes sense and that
(AB)C=A(BC).
Exercises 61
19. Suppose T∈L(Fn,Fm)and that This exercise shows
thatThas the form
promised on page 39.
M(T)=
a1,1... a 1,n
......
am,1... am,n
,
where we are using the standard bases. Prove that
T(x 1,...,xn)=(a1,1x1+···+a 1,nxn,...,am,1x1+···+am,nxn)
for every(x1,...,xn)∈Fn.
20. Suppose (v1,...,vn)is a basis of V. Prove that the function
T:V→Mat(n,1,F)defined by
Tv=M(v)
is an invertible linear map of Vonto Mat(n,1,F); hereM(v) is
the matrix of v∈Vwith respect to the basis (v1,...,vn).
21. Prove that every linear map from Mat(n, 1,F)to Mat(m,1,F)is
given by a matrix multiplication. In other words, prove that ifT∈L(Mat(n,1,F),Mat(m,1,F)), then there exists an m-by-n
matrixAsuch thatTB=ABfor everyB∈Mat(n,1,F).
22. Suppose that Vis finite dimensional and S,T∈L(V). Prove that
STis invertible if and only if both SandTare invertible.
23. Suppose that Vis finite dimensional and S,T∈L(V). Prove that
ST=Iif and only if TS=I.
24. Suppose that Vis finite dimensional and T∈L(V). Prove that
Tis a scalar multiple of the identity if and only if ST=TSfor
everyS∈L(V).
25. Prove that if Vis finite dimensional with dim V>1, then the set
of noninvertible operators on Vis not a subspace of L(V).
62 Chapter 3.Linear Maps
26. Suppose nis a positive integer and ai,j∈Ffori,j=1,...,n .
Prove that the following are equivalent:
(a) The trivial solution x1=···=xn=0 is the only solution
to the homogeneous system of equations
n/summationdisplay
k=1a1,kxk=0
...
n/summationdisplay
k=1an,kxk=0.
(b) For every c1,...,cn∈F, there exists a solution to the sys-
tem of equations
n/summationdisplay
k=1a1,kxk=c1
...
n/summationdisplay
k=1an,kxk=cn.
Note that here we have the same number of equations as vari-
ables.
Chapter 4
Polynomials
This short chapter contains no linear algebra. It does contain the
background material on polynomials that we will need in our studyof linear maps from a vector space to itself. Many of the results in
this chapter will already be familiar to you from other courses; they
are included here for completeness. Because this chapter is not aboutlinear algebra, your instructor may go through it rapidly. You may notbe asked to scrutinize all the proofs. Make sure, however, that youat least read and understand the statements of all the results in thischapter—they will be used in the rest of the book.
Recall that Fdenotes RorC.
✽
✽✽✽
63
64 Chapter 4.Polynomials
Degree
Recall that a function p:F→Fis called a polynomial with coeffi-
cients in Fif there exist a0,...,am∈Fsuch that
p(z)=a0+a1z+a2z2+···+a mzm
for allz∈F.I fpcan be written in the form above with am/negationslash=0, then we
say thatphas degreem. If all the coefficients a0,...,amequal 0, then
we say thatphas degree−∞. For all we know at this stage, a polynomial When necessary, use
the obvious arithmetic
with−∞. For example,
−∞<m and
−∞+m=−∞ for
every integer m. The 0
polynomial is declared
to have degree −∞ so
that exceptions are not
needed for various
reasonable results. For
example, the degree of
pqequals the degree of
pplus the degree of q
even ifp=0.may have more than one degree because we have not yet proved that
the coefficients in the equation above are uniquely determined by the
functionp.
Recall thatP(F)denotes the vector space of all polynomials with
coefficients in Fand thatPm(F)is the subspace of P(F)consisting of
the polynomials with coefficients in Fand degree at most m. A number
λ∈Fis called a root of a polynomial p∈P(F)if
p(λ)=0.
Roots play a crucial role in the study of polynomials. We begin by
showing that λis a root ofpif and only if pis a polynomial multiple
ofz−λ.
4.1 Proposition: Supposep∈P(F)is a polynomial with degree
m≥1. Letλ∈F. Thenλis a root of pif and only if there is a
polynomialq∈P(F)with degree m−1such that
4.2 p(z)=(z−λ)q(z)
for allz∈F.
Proof: One direction is obvious. Namely, suppose there is a poly-
nomialq∈P(F)such that 4.2 holds. Then
p(λ)=(λ−λ)q(λ)=0,
and henceλis a root ofp, as desired.
To prove the other direction, suppose that λ∈Fis a root ofp. Let
a0,...,am∈Fbe such that am/negationslash=0 and
p(z)=a0+a1z+a2z2+···+a mzm
Degree 65
for allz∈F. Becausep(λ)=0, we have
0=a0+a1λ+a2λ2+···+amλm.
Subtracting the last two equations, we get
p(z)=a1(z−λ)+a2(z2−λ2)+···+am(zm−λm)
for allz∈F. For eachj=2,...,m , we can write
zj−λj=(z−λ)qj−1(z)
for allz∈F, whereqj−1is a polynomial with degree j−1 (specifically,
takeqj−1(z)=zj−1+zj−2λ+···+zλj−2+λj−1). Thus
p(z)=(z−λ)(a 1+a2q2(z)+···+amqm−1(z))/bracehtipupleft /bracehtipdownright/bracehtipdownleft /bracehtipupright
q(z)
for allz∈F. Clearlyqis a polynomial with degree m−1, as desired.
Now we can prove that polynomials do not have too many roots.
4.3 Corollary: Supposep∈P(F)is a polynomial with degree m≥0.
Thenphas at most mdistinct roots in F.
Proof: Ifm=0, thenp(z)=a0/negationslash=0 and sophas no roots. If
m=1, thenp(z)=a0+a1z, witha1/negationslash=0, andphas exactly one
root, namely, −a0/a1. Now suppose m> 1. We use induction on m,
assuming that every polynomial with degree m−1 has at most m−1
distinct roots. If phas no roots in F, then we are done. If phas a root
λ∈F, then by 4.1 there is a polynomial qwith degree m−1 such that
p(z)=(z−λ)q(z)
for allz∈F. The equation above shows that if p(z)=0, then either
z=λorq(z)=0. In other words, the roots of pconsist ofλand the
roots ofq. By our induction hypothesis, qhas at most m−1 distinct
roots in F. Thusphas at most mdistinct roots in F.
The next result states that if a polynomial is identically 0, then all
its coefficients must be 0.
66 Chapter 4.Polynomials
4.4 Corollary: Supposea0,...,am∈F.I f
a0+a1z+a2z2+···+amzm=0
for allz∈F, thena0=···=a m=0.
Proof: Supposea0+a1z+a2z2+···+amzmequals 0 for all z∈F.
By 4.3, no nonnegative integer can be the degree of this polynomial.Thus all the coefficients equal 0.
The corollary above implies that (1,z,...,zm)is linearly indepen-
dent inP(F)for every nonnegative integer m. We had noted this earlier
(in Chapter 2), but now we have a complete proof. This linear indepen-dence implies that each polynomial can be represented in only one wayas a linear combination of functions of the form z
j. In particular, the
degree of a polynomial is unique.
Ifpandqare nonnegative integers, with p/negationslash=0, then there exist
nonnegative integers sandrsuch that
q=sp+r.
andr<p . Think of dividing qbyp, gettingswith remainder r. Our
next task is to prove an analogous result for polynomials.
Let degpdenote the degree of a polynomial p. The next result is
often called the division algorithm, though as stated here it is not reallyan algorithm, just a useful lemma.
4.5 Division Algorithm: Supposep,q∈P(F), withp/negationslash=0. Then
Think of 4.6 as giving
the remainder rwhen
qis divided by p.there exist polynomials s,r∈P(F)such that
4.6 q=sp+r
anddegr<degp.
Proof: Chooses∈P(F)such thatq−sphas degree as small as
possible. Let r=q−sp. Thus 4.6 holds, and all that remains is to
show that deg r<degp. Suppose that deg r≥degp.I fc∈Fandjis
a nonnegative integer, then
q−(s+czj)p=r−czjp.
Choosejandcso that the polynomial on the right side of this equation
has degree less than deg r(specifically, take j=degr−degpand then
Complex Coefficients 67
choosecso that the coefficients of zdegrinrand inczjpare equal).
This contradicts our choice of sas the polynomial that produces the
smallest degree for expressions of the form q−sp, completing the
proof.
Complex Coefficients
So far we have been handling polynomials with complex coefficients
and polynomials with real coefficients simultaneously through our con-
vention that Fdenotes RorC. Now we will see some differences be-
tween these two cases. In this section we treat polynomials with com-plex coefficients. In the next section we will use our results about poly-nomials with complex coefficients to prove corresponding results forpolynomials with real coefficients.
Though this chapter contains no linear algebra, the results so far
have nonetheless been proved using algebra. The next result, thoughcalled the fundamental theorem of algebra, requires analysis for itsproof. The short proof presented here uses tools from complex anal-ysis. If you have not had a course in complex analysis, this proof willalmost certainly be meaningless to you. In that case, just accept the
fundamental theorem of algebra as something that we need to use but
whose proof requires more advanced tools that you may learn in latercourses.
4.7 Fundamental Theorem of Algebra: Every nonconstant polyno-
This is an existence
theorem. The quadraticformula gives the roots
explicitly for
polynomials ofdegree 2. Similar but
more complicated
formulas exist for
polynomials of degree3and4. No such
formulas exist for
polynomials of degree5and above.mial with complex coefficients has a root.
Proof: Letpbe a nonconstant polynomial with complex coeffi-
cients. Suppose that phas no roots. Then 1 /pis an analytic function
onC. Furthermore, p(z)→∞ asz→∞, which implies that 1 /p→0a s
z→∞. Thus 1/p is a bounded analytic function on C. By Liouville’s the-
orem, any such function must be constant. But if 1 /pis constant, then
pis constant, contradicting our assumption that pis nonconstant.
The fundamental theorem of algebra leads to the following factor-
ization result for polynomials with complex coefficients. Note thatin this factorization, the numbers λ
1,...,λmare precisely the roots
ofp, for these are the only values of zfor which the right side of 4.9
equals 0.
68 Chapter 4.Polynomials
4.8 Corollary: Ifp∈P(C)is a nonconstant polynomial, then p
has a unique factorization (except for the order of the factors) of theform
4.9 p(z)=c(z−λ
1)...(z−λm),
wherec,λ 1,...,λm∈C.
Proof: Letp∈P(C)and letmdenote the degree of p. We will use
induction on m.I fm=1, then clearly the desired factorization exists
and is unique. So assume that m> 1 and that the desired factorization
exists and is unique for all polynomials of degree m−1.
First we will show that the desired factorization of pexists. By the
fundamental theorem of algebra (4.7), phas a rootλ. By 4.1, there is a
polynomialqwith degree m−1 such that
p(z)=(z−λ)q(z)
for allz∈C. Our induction hypothesis implies that qhas the desired
factorization, which when plugged into the equation above gives thedesired factorization of p.
Now we turn to the question of uniqueness. Clearly cis uniquely
determined by 4.9—it must equal the coefficient of z
minp. So we need
only show that except for the order, there is only one way to chooseλ
1,...,λm.I f
(z−λ1)...(z−λm)=(z−τ1)...(z−τm)
for allz∈C, then because the left side of the equation above equals 0
whenz=λ1, one of theτ’s on the right side must equal λ1. Relabeling,
we can assume that τ1=λ1. Now forz/negationslash=λ1, we can divide both sides
of the equation above by z−λ1, getting
(z−λ2)...(z−λm)=(z−τ2)...(z−τm)
for allz∈Cexcept possibly z=λ1. Actually the equation above
must hold for all z∈Cbecause otherwise by subtracting the right side
from the left side we would get a nonzero polynomial that has infinitelymany roots. The equation above and our induction hypothesis implythat except for the order, the λ’s are the same as the τ’s, completing
the proof of the uniqueness.
Real Coefficients 69
Real Coefficients
Before discussing polynomials with real coefficients, we need to
learn a bit more about the complex numbers.
Supposez=a+bi, whereaandbare real numbers. Then ais
called the real part ofz, denoted Re z, andbis called the imaginary
part ofz, denoted Im z. Thus for every complex number z, we have
z=Rez+(Imz)i.
The complex conjugate ofz∈C, denoted ¯z, is defined by Note thatz=¯zif and
only ifzis a real
number.¯z=Rez−(Imz)i.
For example, 2+3i=2−3i.
The absolute value of a complex number z, denoted|z|, is defined
by
|z|=/radicalBig
(Rez)2+(Imz)2.
For example, |1+2i|=√
5. Note that |z|is always a nonnegative
number.
You should verify that the real and imaginary parts, absolute value,
and complex conjugate have the following properties:
additivity of real part
Re(w+z)=Rew+Rezfor allw,z∈C;
additivity of imaginary part
Im(w+z)=Imw+Imzfor allw,z∈C;
sum ofzand¯z
z+¯z=2R ez for allz∈C;
difference of zand¯z
z−¯z=2(Imz)ifor allz∈C;
product ofzand¯z
z¯z=|z|2for allz∈C;
additivity of complex conjugate
w+z=¯w+¯zfor allw,z∈C;
multiplicativity of complex conjugate
wz=¯w¯zfor allw,z∈C;
70 Chapter 4.Polynomials
conjugate of conjugate
¯z=zfor allz∈C;
multiplicativity of absolute value
|wz|=|w||z|for allw,z∈C.
In the next result, we need to think of a polynomial with real coef-
ficients as an element of P(C). This makes sense because every real
number is also a complex number.
4.10 Proposition: Supposepis a polynomial with real coefficients. A polynomial with real
coefficients may have
no real roots. For
example, the
polynomial 1+x2has
no real roots. The
failure of the
fundamental theorem
of algebra for R
accounts for the
differences between
operators on real and
complex vector spaces,
as we will see in later
chapters.Ifλ∈Cis a root ofp, then so is ¯λ.
Proof: Let
p(z)=a0+a1z+···+amzm,
wherea0,...,amare real numbers. Suppose λ∈Cis a root ofp. Then
a0+a1λ+···+amλm=0.
Take the complex conjugate of both sides of this equation, obtaining
a0+a1¯λ+···+am¯λm=0,
where we have used some of the basic properties of complex conjuga-
tion listed earlier. The equation above shows that ¯λis a root ofp.
We want to prove a factorization theorem for polynomials with real
coefficients. To do this, we begin by characterizing the polynomialswith real coefficients and degree 2 that can be written as the productof two polynomials with real coefficients and degree 1.
4.11 Proposition: Letα,β∈R. Then there is a polynomial factor-
Think about the
connection between the
quadratic formula and
this proposition.ization of the form
4.12 x2+αx+β=(x−λ1)(x−λ2),
withλ1,λ2∈R, if and only if α2≥4β.
Proof: Notice that
4.13 x2+αx+β=(x+α
2)2+(β−α2
4).
Real Coefficients 71
First suppose that α2<4β. Then clearly the right side of the
equation above is positive for every x∈R, and hence the polynomial
x2+αx+βhas no real roots. Thus no factorization of the form 4.12,
withλ1,λ2∈R, can exist.
Conversely, now suppose that α2≥4β. Thus there is a real number
csuch thatc2=α2
4−β. From 4.13, we have
x2+αx+β=(x+α
2)2−c2
=(x+α
2+c)(x+α
2−c),
which gives the desired factorization.
In the following theorem, each term of the form x2+αjx+βj, with
αj2<4βj, cannot be factored into the product of two polynomials with
real coefficients and degree 1 (by 4.11). Note that in the factorizationbelow, the numbers λ
1,...,λmare precisely the real roots of p, for these
are the only real values of xfor which the right side of the equation
below equals 0.
4.14 Theorem: Ifp∈P(R)is a nonconstant polynomial, then p
has a unique factorization (except for the order of the factors) of the
form
p(x)=c(x−λ1)...(x−λm)(x2+α1x+β1)...(x2+αMx+βM),
wherec,λ 1,...,λm∈Rand(α1,β1),...,(αM,βM)∈R2withαj2<4βj Here eithermorM
may equal 0. for eachj.
Proof: Letp∈P(R)be a nonconstant polynomial. We can think
ofpas an element of P(C)(because every real number is a complex
number). The idea of the proof is to use the factorization 4.8 of pas a
polynomial with complex coefficients. Complex but nonreal roots of p
come in pairs; see 4.10. Thus if the factorization of pas an element
ofP(C)includes terms of the form (x−λ)withλa nonreal complex
number, then (x−¯λ)is also a term in the factorization. Combining
these two terms, we get a quadratic term of the required form.
The idea sketched in the paragraph above almost provides a proof
of the existence of our desired factorization. However, we need tobe careful about one point. Suppose λis a nonreal complex number
72 Chapter 4.Polynomials
and(x−λ)is a term in the factorization of pas an element of P(C).
We are guaranteed by 4.10 that (x−¯λ)also appears as a term in the
factorization, but 4.10 does not state that these two factors appear
the same number of times, as needed to make the idea above work.However, all is well. We can write
p(x)=(x−λ)(x−¯λ)q(x)
=/parenleftbig
x
2−2(Reλ)x+|λ|2/parenrightbig
q(x)
for some polynomial q∈P(C)with degree two less than the degree
ofp. If we can prove that qhas real coefficients, then, by using induc-
tion on the degree of p, we can conclude that (x−λ)appears in the
factorization of pexactly as many times as (x−¯λ).
To prove that qhas real coefficients, we solve the equation above
forq, getting Here we are not
dividing by 0because
the roots of
x2−2(Reλ)x+|λ|2
areλand¯λ, neither of
which is real.q(x)=p(x)
x2−2(Reλ)x+|λ|2
for allx∈R. The equation above implies that q(x)∈Rfor allx∈R.
Writing
q(x)=a0+a1x+···+an−2xn−2,
wherea0,...,an−2∈C, we thus have
0=Imq(x)=(Ima0)+(Ima1)x+···+(Iman−2)xn−2
for allx∈R. This implies that Im a0,...,Iman−2all equal 0 (by 4.4).
Thus all the coefficients of qare real, as desired, and hence the desired
factorization exists.
Now we turn to the question of uniqueness of our factorization. A
factor ofpof the formx2+αx+βwithα2<4βcan be uniquely written
as(x−λ)(x−¯λ)withλ∈C. A moment’s thought shows that two
different factorizations of pas an element of P(R)would lead to two
different factorizations of pas an element of P(C), contradicting 4.8.
Exercises 73
Exercises
1. Suppose mandnare positive integers with m≤n. Prove that
there exists a polynomial p∈Pn(F)with exactly mdistinct
roots.
2. Suppose that z1,...,zm+1are distinct elements of Fand that
w1,...,wm+1∈F. Prove that there exists a unique polynomial
p∈Pm(F)such that
p(zj)=wj
forj=1,...,m+1.
3. Prove that if p,q∈P(F), withp/negationslash=0, then there exist unique
polynomials s,r∈P(F)such that
q=sp+r
and degr<degp. In other words, add a uniqueness statement
to the division algorithm (4.5).
4. Suppose p∈P(C)has degreem. Prove that phasmdistinct
roots if and only if pand its derivative p/primehave no roots in com-
mon.
5. Prove that every polynomial with odd degree and real coefficients
has a real root.
Chapter 5
Eigenvalues and Eigenvectors
In Chapter 3 we studied linear maps from one vector space to an-
other vector space. Now we begin our investigation of linear maps froma vector space to itself. Their study constitutes the deepest and mostimportant part of linear algebra. Most of the key results in this areado not hold for infinite-dimensional vector spaces, so we work only onfinite-dimensional vector spaces. To avoid trivialities we also want toeliminate the vector space {0}from consideration. Thus we make the
following assumption:
Recall that Fdenotes RorC.
Let’s agree that for the rest of the book
Vwill denote a finite-dimensional, nonzero vector space over F.
✽✽✽✽✽
75
76 Chapter 5.Eigenvalues and Eigenvectors
Invariant Subspaces
In this chapter we develop the tools that will help us understand the
structure of operators. Recall that an operator is a linear map from a
vector space to itself. Recall also that we denote the set of operators
onVbyL(V); in other words, L(V)=L(V,V).
Let’s see how we might better understand what an operator looks
like. Suppose T∈L(V). If we have a direct sum decomposition
5.1 V=U1⊕···⊕Um,
where eachUjis a proper subspace of V, then to understand the be-
havior ofT, we need only understand the behavior of each T|Uj; here
T|Ujdenotes the restriction of Tto the smaller domain Uj. Dealing
withT|Ujshould be easier than dealing with TbecauseUjis a smaller
vector space than V. However, if we intend to apply tools useful in the
study of operators (such as taking powers), then we have a problem:
T|Ujmay not map Ujinto itself; in other words, T|Ujmay not be an
operator on Uj. Thus we are led to consider only decompositions of
the form 5.1 where Tmaps eachUjinto itself.
The notion of a subspace that gets mapped into itself is sufficiently
important to deserve a name. Thus, for T∈L(V)andUa subspace
ofV, we say that Uisinvariant underTifu∈UimpliesTu∈U.
In other words, Uis invariant under TifT|Uis an operator on U. For
example, ifTis the operator of differentiation on P7(R), thenP4(R)
(which is a subspace of P7(R)) is invariant under Tbecause the deriva-
tive of any polynomial of degree at most 4 is also a polynomial withdegree at most 4.
Let’s look at some easy examples of invariant subspaces. Suppose
The most famous
unsolved problem in
functional analysis is
called the invariant
subspace problem. It
deals with invariant
subspaces of operators
on infinite-dimensional
vector spaces.T∈L(V). Clearly {0}is invariant under T. Also, the whole space Vis
obviously invariant under T. MustThave any invariant subspaces other
than{0}andV? Later we will see that this question has an affirmative
answer for operators on complex vector spaces with dimension greaterthan 1 and also for operators on real vector spaces with dimension
greater than 2.
IfT∈L(V), then null Tis invariant under T(proof: ifu∈nullT,
thenTu=0, and hence Tu∈nullT). Also, range Tis invariant under T
(proof: ifu∈rangeT, thenTuis also in range T, by the definition of
range). Although null Tand rangeTare invariant under T, they do not
necessarily provide easy answers to the question about the existence
Invariant Subspaces 77
of invariant subspaces other than {0}andVbecause null Tmay equal
{0}and rangeTmay equalV(this happens when Tis invertible).
We will return later to a deeper study of invariant subspaces. Now
we turn to an investigation of the simplest possible nontrivial invariantsubspaces—invariant subspaces with dimension 1.
How does an operator behave on an invariant subspace of dimen-
sion 1? Subspaces of Vof dimension 1 are easy to describe. Take any
nonzero vector u∈Vand letUequal the set of all scalar multiples
ofu:
5.2 U={au:a∈F}.
ThenUis a one-dimensional subspace of V, and every one-dimensional
These subspaces are
loosely connected to
the subject of Herbert
Marcuse’s well-knownbook One-Dimensional
Man.subspace of Vis of this form. If u∈Vand the subspace Udefined
by 5.2 is invariant under T∈L(V), thenTumust be inU, and hence
there must be a scalar λ∈Fsuch thatTu=λu. Conversely, if u
is a nonzero vector in Vsuch thatTu=λufor someλ∈F, then the
subspaceUdefined by 5.2 is a one-dimensional subspace of Vinvariant
underT.
The equation
5.3 Tu=λu,
which we have just seen is intimately connected with one-dimensional
invariant subspaces, is important enough that the vectors uand scalars
λsatisfying it are given special names. Specifically, a scalar λ∈F
is called an eigenvalue ofT∈L(V)if there exists a nonzero vector The regrettable word
eigenvalue is
half-German,
half-English. The
German adjective eigen
means own in the sense
of characterizing someintrinsic property.
Some mathematicians
use the term
characteristic value
instead of eigenvalue.u∈Vsuch thatTu=λu. We must require uto be nonzero because
withu=0 every scalar λ∈Fsatisfies 5.3. The comments above show
thatThas a one-dimensional invariant subspace if and only if Thas
an eigenvalue.
The equation Tu=λuis equivalent to (T−λI)u=0, soλis an
eigenvalue of Tif and only if T−λIis not injective. By 3.21, λis an
eigenvalue of Tif and only if T−λIis not invertible, and this happens
if and only if T−λIis not surjective.
SupposeT∈L(V)andλ∈Fis an eigenvalue of T. A vectoru∈V
is called an eigenvector ofT(corresponding to λ)i fTu=λu. Because
5.3 is equivalent to (T−λI)u=0, we see that the set of eigenvectors
ofTcorresponding to λequals null(T−λI). In particular, the set of
eigenvectors of Tcorresponding to λis a subspace of V.
78 Chapter 5.Eigenvalues and Eigenvectors
Let’s look at some examples of eigenvalues and eigenvectors. If Some texts define
eigenvectors as we
have, except that 0is
declared not to be an
eigenvector. With the
definition used here,
the set of eigenvectors
corresponding to a
fixed eigenvalue is a
subspace.a∈F, thenaIhas only one eigenvalue, namely, a, and every vector is
an eigenvector for this eigenvalue.
For a more complicated example, consider the operator T∈L(F2)
defined by
5.4 T(w,z)=(−z,w).
IfF=R, then this operator has a nice geometric interpretation: Tis
just a counterclockwise rotation by 90◦about the origin in R2.A n
operator has an eigenvalue if and only if there exists a nonzero vectorin its domain that gets sent by the operator to a scalar multiple of itself.The rotation of a nonzero vector in R
2obviously never equals a scalar
multiple of itself. Conclusion: if F=R, the operator Tdefined by 5.4
has no eigenvalues. However, if F=C, the story changes. To find
eigenvalues of T, we must find the scalars λsuch that
T(w,z)=λ(w,z)
has some solution other than w=z=0. ForTdefined by 5.4, the
equation above is equivalent to the simultaneous equations
5.5 −z=λw, w=λz.
Substituting the value for wgiven by the second equation into the first
equation gives
−z=λ2z.
Nowzcannot equal 0 (otherwise 5.5 implies that w=0; we are looking
for solutions to 5.5 where (w,z) is not the 0 vector), so the equation
above leads to the equation
−1=λ2.
The solutions to this equation are λ=iorλ=−i. You should be
able to verify easily that iand−iare eigenvalues of T. Indeed, the
eigenvectors corresponding to the eigenvalue iare the vectors of the
form(w,−wi) , withw∈C, and the eigenvectors corresponding to the
eigenvalue−iare the vectors of the form (w,wi), with w∈C.
Now we show that nonzero eigenvectors corresponding to distinct
eigenvalues are linearly independent.
Invariant Subspaces 79
5.6 Theorem: LetT∈L(V). Suppose λ1,...,λmare distinct eigen-
values ofTandv1,...,vmare corresponding nonzero eigenvectors.
Then(v1,...,vm)is linearly independent.
Proof: Suppose(v1,...,vm)is linearly dependent. Let kbe the
smallest positive integer such that
5.7 vk∈span(v 1,...,vk−1);
the existence of kwith this property follows from the linear dependence
lemma (2.4). Thus there exist a1,...,ak−1∈Fsuch that
5.8 vk=a1v1+···+ak−1vk−1.
ApplyTto both sides of this equation, getting
λkvk=a1λ1v1+···+ak−1λk−1vk−1.
Multiply both sides of 5.8 by λkand then subtract the equation above,
getting
0=a1(λk−λ1)v1+···+ak−1(λk−λk−1)vk−1.
Because we chose kto be the smallest positive integer satisfying 5.7,
(v1,...,vk−1)is linearly independent. Thus the equation above implies
that all thea’s are 0 (recall that λkis not equal to any of λ1,...,λk−1).
However, this means that vkequals 0 (see 5.8), contradicting our hy-
pothesis that all the v’s are nonzero. Therefore our assumption that
(v1,...,vm)is linearly dependent must have been false.
The corollary below states that an operator cannot have more dis-
tinct eigenvalues than the dimension of the vector space on which itacts.
5.9 Corollary: Each operator on Vhas at most dimVdistinct eigen-
values.
Proof: LetT∈L(V). Suppose that λ
1,...,λmare distinct eigenval-
ues ofT. Letv1,...,vmbe corresponding nonzero eigenvectors. The
last theorem implies that (v1,...,vm)is linearly independent. Thus
m≤dimV(see 2.6), as desired.
80 Chapter 5.Eigenvalues and Eigenvectors
Polynomials Applied to Operators
The main reason that a richer theory exists for operators (which
map a vector space into itself) than for linear maps is that operators
can be raised to powers. In this section we define that notion and the
key concept of applying a polynomial to an operator.
IfT∈L(V), thenTTmakes sense and is also in L(V). We usually
writeT2instead ofTT. More generally, if mis a positive integer, then
Tmis defined by
Tm=T...T/bracehtipupleft/bracehtipdownright/bracehtipdownleft/bracehtipupright
mtimes.
For convenience we define T0to be the identity operator IonV.
Recall from Chapter 3 that if Tis an invertible operator, then the
inverse ofTis denoted by T−1.I fmis a positive integer, then we define
T−mto be(T−1)m.
You should verify that if Tis an operator, then
TmTn=Tm+nand(Tm)n=Tmn,
wheremandnare allowed to be arbitrary integers if Tis invertible
and nonnegative integers if Tis not invertible.
IfT∈L(V)andp∈P(F)is a polynomial given by
p(z)=a0+a1z+a2z2+···+a mzm
forz∈F, thenp(T) is the operator defined by
p(T)=a0I+a1T+a2T2+···+amTm.
For example, if pis the polynomial defined by p(z)=z2forz∈F, then
p(T)=T2. This is a new use of the symbol pbecause we are applying
it to operators, not just elements of F. If we fix an operator T∈L(V),
then the function from P(F)toL(V) given byp/arrowbarrightp(T) is linear, as
you should verify.
Ifpandqare polynomials with coefficients in F, thenpqis the
polynomial defined by
(pq)(z)=p(z)q(z)
forz∈F. You should verify that we have the following nice multiplica-
tive property: if T∈L(V), then
Upper-Triangular Matrices 81
(pq)(T)=p(T)q(T)
for all polynomials pandqwith coefficients in F. Note that any two
polynomials in Tcommute, meaning that p(T)q(T)=q(T)p(T) , be-
cause
p(T)q(T)=(pq)(T)=(qp)(T)=q(T)p(T).
Upper-Triangular Matrices
Now we come to one of the central results about operators on com-
plex vector spaces.
5.10 Theorem: Every operator on a finite-dimensional, nonzero, Compare the simple
proof of this theoremgiven here with the
standard proof using
determinants. With the
standard proof, first
the difficult concept of
determinants must bedefined, then anoperator with 0
determinant must be
shown to be not
invertible, then the
characteristic
polynomial needs to bedefined, and by the
time the proof of this
theorem is reached, no
insight remains aboutwhy it is true.complex vector space has an eigenvalue.
Proof: SupposeVis a complex vector space with dimension n>0
andT∈L(V). Choose v∈Vwithv/negationslash=0. Then
(v,Tv,T2v,...,Tnv)
cannot be linearly independent because Vhas dimension nand we have
n+1 vectors. Thus there exist complex numbers a0,...,an, not all 0,
such that
0=a0v+a1Tv+···+a nTnv.
Letmbe the largest index such that am/negationslash=0. Because v/negationslash=0, the
coefficients a1,...,amcannot all be 0, so 0 <m≤n. Make the a’s
the coefficients of a polynomial, which can be written in factored form
(see 4.8) as
a0+a1z+···+anzn=c(z−λ1)...(z−λm),
wherecis a nonzero complex number, each λj∈C, and the equation
holds for all z∈C. We then have
0=a0v+a1Tv+···+anTnv
=(a0I+a1T+···+anTn)v
=c(T−λ1I)...(T−λmI)v,
which means that T−λjIis not injective for at least one j. In other
words,Thas an eigenvalue.
82 Chapter 5.Eigenvalues and Eigenvectors
Recall that in Chapter 3 we discussed the matrix of a linear map
from one vector space to another vector space. This matrix dependedon a choice of a basis for each of the two vector spaces. Now that we are
studying operators, which map a vector space to itself, we need onlyone basis. In addition, now our matrices will be square arrays, ratherthan the more general rectangular arrays that we considered earlier.
Specifically, let T∈L(V). Suppose (v
1,...,vn)is a basis of V. For
eachk=1,...,n , we can write
Tvk=a1,kv1+···+an,kvn,
whereaj,k∈Fforj=1,...,n . Then-by-n matrix Thekthcolumn of the
matrix is formed from
the coefficients used to
writeTvkas a linear
combination of the v’s.5.11
a1,1... a 1,n
......
an,1... an,n
is called the matrix ofTwith respect to the basis (v1,...,vn); we de-
note it byM/parenleftbig
T,(v 1,...,vn)/parenrightbig
or just byM(T) if the basis(v1,...,vn)
is clear from the context (for example, if only one basis is in sight).
IfTis an operator on Fnand no basis is specified, you should assume
that the basis in question is the standard one (where the jthbasis vector
is 1 in thejthslot and 0 in all the other slots). You can then think of
thejthcolumn ofM(T) asTapplied to the jthbasis vector.
A central goal of linear algebra is to show that given an operator
T∈L(V), there exists a basis of Vwith respect to which Thas a
reasonably simple matrix. To make this vague formulation (“reasonablysimple” is not precise language) a bit more concrete, we might try tomakeM(T) have many 0’s.
IfVis a complex vector space, then we already know enough to
show that there is a basis of Vwith respect to which the matrix of T
has 0’s everywhere in the first column, except possibly the first entry.In other words, there is a basis of Vwith respect to which the matrix
ofTlooks like
We often use ∗to
denote matrix entries
that we do not know
about or that are
irrelevant to the
questions being
discussed.
λ
0∗
...
0
;
here the∗denotes the entries in all the columns other than the first
column. To prove this, let λbe an eigenvalue of T(one exists by 5.10)
Upper-Triangular Matrices 83
and letvbe a corresponding nonzero eigenvector. Extend (v)to a
basis ofV. Then the matrix of Twith respect to this basis has the form
above. Soon we will see that we can choose a basis of Vwith respect to
which the matrix of Thas even more 0’s.
The diagonal of a square matrix consists of the entries along the
straight line from the upper left corner to the bottom right corner.For example, the diagonal of the matrix 5.11 consists of the entriesa
1,1,a2,2,...,an,n.
A matrix is called upper triangular if all the entries below the di-
agonal equal 0. For example, the 4-by-4 matrix
6275
061300790008
is upper triangular. Typically we represent an upper-triangular matrix
in the form
λ
1∗
...
0λn
;
the 0 in the matrix above indicates that all entries below the diagonal
in thisn-by-n matrix equal 0. Upper-triangular matrices can be consid-
ered reasonably simple—for nlarge, ann-by-n upper-triangular matrix
has almost half its entries equal to 0.
The following proposition demonstrates a useful connection be-
tween upper-triangular matrices and invariant subspaces.
5.12 Proposition: SupposeT∈L(V) and(v1,...,vn)is a basis
ofV. Then the following are equivalent:
(a) the matrix of Twith respect to (v1,...,vn)is upper triangular;
(b)Tvk∈span(v 1,...,vk)for eachk=1,...,n ;
(c) span(v 1,...,vk)is invariant under Tfor eachk=1,...,n .
Proof: The equivalence of (a) and (b) follows easily from the def-
initions and a moment’s thought. Obviously (c) implies (b). Thus tocomplete the proof, we need only prove that (b) implies (c). So suppose
that (b) holds. Fix k∈{1,...,n}. From (b), we know that
84 Chapter 5.Eigenvalues and Eigenvectors
Tv1∈span(v 1)⊂span(v 1,...,vk);
Tv2∈span(v 1,v2)⊂span(v 1,...,vk);
...
Tvk∈span(v 1,...,vk).
Thus ifvis a linear combination of (v1,...,vk), then
Tv∈span(v 1,...,vk).
In other words, span (v1,...,vk)is invariant under T, completing the
proof.
Now we can show that for each operator on a complex vector space,
there is a basis of the vector space with respect to which the matrixof the operator has only 0’s below the diagonal. In Chapter 8 we will
improve even this result.
5.13 Theorem: SupposeVis a complex vector space and T∈L(V).
This theorem does not
hold on real vector
spaces because the first
vector in a basis with
respect to which an
operator has an
upper-triangular matrix
must be an eigenvector
of the operator. Thus if
an operator on a real
vector space has no
eigenvalues (we have
seen an example on
R2), then there is no
basis with respect to
which the operator has
an upper-triangular
matrix.ThenThas an upper-triangular matrix with respect to some basis of V.
Proof: We will use induction on the dimension of V. Clearly the
desired result holds if dim V=1.
Suppose now that dim V> 1 and the desired result holds for all
complex vector spaces whose dimension is less than the dimensionofV. Letλbe any eigenvalue of T(5.10 guarantees that Thas an
eigenvalue). Let
U=range(T−λI).
BecauseT−λIis not surjective (see 3.21), dim U<dimV. Furthermore,
Uis invariant under T. To prove this, suppose u∈U. Then
Tu=(T−λI)u+λu.
Obviously(T−λI)u∈U(from the definition of U) andλu∈U. Thus
the equation above shows that Tu∈U. HenceUis invariant under T,
as claimed.
ThusT|
Uis an operator on U. By our induction hypothesis, there
is a basis(u1,...,um)ofUwith respect to which T|Uhas an upper-
triangular matrix. Thus for each jwe have (using 5.12)
5.14 Tuj=(T|U)(uj)∈span(u 1,...,uj).
Upper-Triangular Matrices 85
Extend(u1,...,um)to a basis(u1,...,um,v1,...,vn)ofV. For
eachk, we have
Tvk=(T−λI)vk+λvk.
The definition of Ushows that(T−λI)vk∈U=span(u 1,...,um).
Thus the equation above shows that
5.15 Tvk∈span(u 1,...,um,v1,...,vk).
From 5.14 and 5.15, we conclude (using 5.12) that Thas an upper-
triangular matrix with respect to the basis (u1,...,um,v1,...,vn).
How does one determine from looking at the matrix of an operator
whether the operator is invertible? If we are fortunate enough to havea basis with respect to which the matrix of the operator is upper tri-angular, then this problem becomes easy, as the following propositionshows.
5.16 Proposition: SupposeT∈L(V) has an upper-triangular matrix
with respect to some basis of V. ThenTis invertible if and only if all
the entries on the diagonal of that upper-triangular matrix are nonzero.
Proof: Suppose(v
1,...,vn)is a basis of Vwith respect to which
Thas an upper-triangular matrix
5.17 M/parenleftbig
T,(v 1,...,vn)/parenrightbig
=
λ
1 ∗
λ2
...
0 λn
.
We need to prove that Tis not invertible if and only if one of the λ
k’s
equals 0.
First we will prove that if one of the λk’s equals 0, then Tis not
invertible. If λ1=0, thenTv1=0 (from 5.17) and hence Tis not
invertible, as desired. So suppose that 1 <k≤nandλk=0. Then,
as can be seen from 5.17, Tmaps each of the vectors v1,...,vk−1into
span(v 1,...,vk−1). Becauseλk=0, the matrix representation 5.17 also
implies that Tvk∈span(v 1,...,vk−1). Thus we can define a linear map
S: span(v 1,...,vk)→span(v 1,...,vk−1)
86 Chapter 5.Eigenvalues and Eigenvectors
bySv=Tvforv∈span(v 1,...,vk). In other words, Sis justT
restricted to span(v 1,...,vk).
Note that span(v 1,...,vk)has dimension kand span(v 1,...,vk−1)
has dimension k−1 (because(v1,...,vn)is linearly independent). Be-
cause span(v1,...,vk)has a larger dimension than span (v1,...,vk−1),
no linear map from span (v1,...,vk)to span(v 1,...,vk−1)is injective
(see 3.5). Thus there exists a nonzero vector v∈span(v 1,...,vk)such
thatSv=0. HenceTv=0, and thusTis not invertible, as desired.
To prove the other direction, now suppose that Tis not invertible.
ThusTis not injective (see 3.21), and hence there exists a nonzero
vectorv∈Vsuch thatTv=0. Because(v1,...,vn)is a basis ofV,w e
can write
v=a1v1+···+akvk,
wherea1,...,ak∈Fandak/negationslash=0 (represent vas a linear combination
of(v1,...,vn)and then choose kto be the largest index with a nonzero
coefficient). Thus
0=Tv
0=T(a 1v1+···+akvk)
=(a1Tv1+···+ak−1Tvk−1)+akTvk.
The last term in parentheses is in span (v1,...,vk−1)(because of the
upper-triangular form of 5.17). Thus the last equation shows thata
kTvk∈span(v 1,...,vk−1). Multiplying by 1/a k, which is allowed
becauseak/negationslash=0, we conclude that Tvk∈span(v 1,...,vk−1). Thus
whenTvkis written as a linear combination of the basis (v1,...,vn),
the coefficient of vkwill be 0. In other words, λkin 5.17 must be 0,
completing the proof.
Unfortunately no method exists for exactly computing the eigenval- Powerful numeric
techniques exist for
finding good
approximations to the
eigenvalues of an
operator from its
matrix.ues of a typical operator from its matrix (with respect to an arbitrary
basis). However, if we are fortunate enough to find a basis with re-spect to which the matrix of the operator is upper triangular, then theproblem of computing the eigenvalues becomes trivial, as the followingproposition shows.
5.18 Proposition: SupposeT∈L(V) has an upper-triangular matrix
with respect to some basis of V. Then the eigenvalues of Tconsist
precisely of the entries on the diagonal of that upper-triangular matrix.
Diagonal Matrices 87
Proof: Suppose(v1,...,vn)is a basis of Vwith respect to which
Thas an upper-triangular matrix
M/parenleftbig
T,(v 1,...,vn)/parenrightbig
=
λ
1 ∗
λ2
...
0 λn
.
Letλ∈F. Then
M/parenleftbig
T−λI,(v
1,...,vn)/parenrightbig
=
λ
1−λ ∗
λ2−λ
...
0 λn−λ
.
HenceT−λIis not invertible if and only if λequals one of the λ
/prime
js
(see 5.16). In other words, λis an eigenvalue of Tif and only if λ
equals one of the λ/prime
js, as desired.
Diagonal Matrices
Adiagonal matrix is a square matrix that is 0 everywhere except
possibly along the diagonal. For example,
800
020005
is a diagonal matrix. Obviously every diagonal matrix is upper triangu-
lar, although in general a diagonal matrix has many more 0’s than anupper-triangular matrix.
An operator T∈L(V)has a diagonal matrix
λ
1 0
...
0λn
with respect to a basis (v1,...,vn)ofVif and only
Tv1=λ1v1
...
Tvn=λnvn;
88 Chapter 5.Eigenvalues and Eigenvectors
this follows immediately from the definition of the matrix of an opera-
tor with respect to a basis. Thus an operator T∈L(V)has a diagonal
matrix with respect to some basis of Vif and only if Vhas a basis
consisting of eigenvectors of T.
If an operator has a diagonal matrix with respect to some basis,
then the entries along the diagonal are precisely the eigenvalues of theoperator; this follows from 5.18 (or you may want to find an easierproof that works only for diagonal matrices).
Unfortunately not every operator has a diagonal matrix with respect
to some basis. This sad state of affairs can arise even on complex vectorspaces. For example, consider T∈L(C
2)defined by
5.19 T(w,z)=(z,0).
As you should verify, 0 is the only eigenvalue of this operator and
the corresponding set of eigenvectors is the one-dimensional subspace
{(w, 0)∈C2:w∈C}. Thus there are not enough linearly independent
eigenvectors of Tto form a basis of the two-dimensional space C2.
HenceTdoes not have a diagonal matrix with respect to any basis
ofC2.
The next proposition shows that if an operator has as many distinct
eigenvalues as the dimension of its domain, then the operator has a di-agonal matrix with respect to some operator. However, some operatorswith fewer eigenvalues also have diagonal matrices (in other words, the
converse of the next proposition is not true). For example, the operatorTdefined on the three-dimensional space F
3by
T(z 1,z2,z3)=(4z 1,4z2,5z3)
has only two eigenvalues (4 and 5), but this operator has a diagonal
matrix with respect to the standard basis.
5.20 Proposition: IfT∈L(V) hasdimVdistinct eigenvalues, then Later we will find other
conditions that imply
that certain operators
have a diagonal matrix
with respect to some
basis (see 7.9 and 7.13).Thas a diagonal matrix with respect to some basis of V.
Proof: Suppose that T∈L(V)has dimVdistinct eigenvalues
λ1,...,λ dimV. For eachj, letvj∈Vbe a nonzero eigenvector cor-
responding to the eigenvalue λj. Because nonzero eigenvectors cor-
responding to distinct eigenvalues are linearly independent (see 5.6),
(v1,...,v dimV)is linearly independent. A linearly independent list of
Diagonal Matrices 89
dimVvectors inVis a basis of V(see 2.17); thus (v1,...,v dimV)is a
basis ofV. With respect to this basis consisting of eigenvectors, Thas
a diagonal matrix.
We close this section with a proposition giving several conditions
on an operator that are equivalent to its having a diagonal matrix withrespect to some basis.
5.21 Proposition: SupposeT∈L(V). Letλ
1,...,λmdenote the For complex vector
spaces, we will extend
this list of equivalences
later (see Exercises 16
and 23 in Chapter 8).distinct eigenvalues of T. Then the following are equivalent:
(a)Thas a diagonal matrix with respect to some basis of V;
(b)Vhas a basis consisting of eigenvectors of T;
(c) there exist one-dimensional subspaces U1,...,UnofV, each in-
variant under T, such that
V=U1⊕···⊕Un;
(d)V=null(T−λ1I)⊕···⊕ null(T−λmI);
(e) dimV=dim null(T−λ1I)+···+ dim null(T−λmI).
Proof: We have already shown that (a) and (b) are equivalent.
Suppose that (b) holds; thus Vhas a basis(v1,...,vn)consisting of
eigenvectors of T. For eachj, letUj=span(vj). Obviously each Uj
is a one-dimensional subspace of Vthat is invariant under T(because
eachvjis an eigenvector of T). Because(v1,...,vn)is a basis of V,
each vector in Vcan be written uniquely as a linear combination of
(v1,...,vn). In other words, each vector in Vcan be written uniquely
as a sumu1+···+u n, where each uj∈Uj. ThusV=U1⊕···⊕Un.
Hence (b) implies (c).
Suppose now that (c) holds; thus there are one-dimensional sub-
spacesU1,...,UnofV, each invariant under T, such that
V=U1⊕···⊕Un.
For eachj, letvjbe a nonzero vector in Uj. Then each vjis an eigen-
vector ofT. Because each vector in Vcan be written uniquely as a sum
u1+···+un, where each uj∈Uj(so eachujis a scalar multiple of vj),
we see that(v1,...,vn)is a basis of V. Thus (c) implies (b).
90 Chapter 5.Eigenvalues and Eigenvectors
At this stage of the proof we know that (a), (b), and (c) are all equiv-
alent. We will finish the proof by showing that (b) implies (d), that (d)implies (e), and that (e) implies (b).
Suppose that (b) holds; thus Vhas a basis consisting of eigenvectors
ofT. Thus every vector in Vis a linear combination of eigenvectors
ofT. Hence
5.22 V=null(T−λ
1I)+···+ null(T−λmI).
To show that the sum above is a direct sum, suppose that
0=u1+···+um,
where eachuj∈null(T−λjI). Because nonzero eigenvectors corre-
sponding to distinct eigenvalues are linearly independent, this implies(apply 5.6 to the sum of the nonzero vectors on the right side of theequation above) that each u
jequals 0. This implies (using 1.8) that the
sum in 5.22 is a direct sum, completing the proof that (b) implies (d).
That (d) implies (e) follows immediately from Exercise 17 in Chap-
ter 2.
Finally, suppose that (e) holds; thus
5.23 dimV=dim null(T−λ1I)+···+ dim null(T−λmI).
Choose a basis of each null (T−λjI); put all these bases together to
form a list(v1,...,vn)of eigenvectors of T, wheren=dimV(by 5.23).
To show that this list is linearly independent, suppose
a1v1+···+anvn=0,
wherea1,...,an∈F. For eachj=1,...,m , letujdenote the sum of
all the terms akvksuch thatvk∈null(T−λjI). Thus each ujis an
eigenvector of Twith eigenvalue λj, and
u1+···+um=0.
Because nonzero eigenvectors corresponding to distinct eigenvalues
are linearly independent, this implies (apply 5.6 to the sum of thenonzero vectors on the left side of the equation above) that each u
j
equals 0. Because each ujis a sum of terms akvk, where the vk’s
were chosen to be a basis of null (T−λjI), this implies that all the ak’s
equal 0. Thus (v1,...,vn)is linearly independent and hence is a basis
ofV(by 2.17). Thus (e) implies (b), completing the proof.
Invariant Subspaces on Real Vector Spaces 91
Invariant Subspaces on Real Vector Spaces
We know that every operator on a complex vector space has an eigen-
value (see 5.10 for the precise statement). We have also seen an example
showing that the analogous statement is false on real vector spaces. In
other words, an operator on a nonzero real vector space may have noinvariant subspaces of dimension 1. However, we now show that aninvariant subspace of dimension 1 or 2 always exists.
5.24 Theorem: Every operator on a finite-dimensional, nonzero, real
vector space has an invariant subspace of dimension 1or2.
Proof: SupposeVis a real vector space with dimension n>0 and
T∈L(V). Choose v∈Vwithv/negationslash=0. Then
(v,Tv,T
2v,...,Tnv)
cannot be linearly independent because Vhas dimension nand we have
n+1 vectors. Thus there exist real numbers a0,...,an, not all 0, such
that
0=a0v+a1Tv+···+a nTnv.
Make thea’s the coefficients of a polynomial, which can be written in
factored form (see 4.14) as
a0+a1x+···+anxn
=c(x−λ1)...(x−λm)(x2+α1x+β1)...(x2+αMx+βM),
wherecis a nonzero real number, each λj,αj, andβjis real,m+M≥1,Here eithermorM
might equal 0.
and the equation holds for all x∈R. We then have
0=a0v+a1Tv+···+anTnv
=(a0I+a1T+···+anTn)v
=c(T−λ1I)...(T−λmI)(T2+α1T+β1I)...(T2+αMT+βMI)v,
which means that T−λjIis not injective for at least one jor that
(T2+αjT+βjI)is not injective for at least one j.I fT−λjIis not
injective for at least one j, thenThas an eigenvalue and hence a one-
dimensional invariant subspace. Let’s consider the other possibility. Inother words, suppose that (T
2+αjT+βjI)is not injective for some j.
Thus there exists a nonzero vector u∈Vsuch that
92 Chapter 5.Eigenvalues and Eigenvectors
5.25 T2u+αjTu+βju=0.
We will complete the proof by showing that span (u,Tu), which clearly
has dimension 1 or 2, is invariant under T. To do this, consider a typical
element of span (u,Tu) of the formau+bTu, wherea,b∈R. Then
T(au+bTu)=aTu+bT2u
=aTu−bαjTu−bβju,
where the last equality comes from solving for T2uin 5.25. The equa-
tion above shows that T(au+bTu)∈span(u,Tu) . Thus span(u,Tu)
is invariant under T, as desired.
We will need one new piece of notation for the next proof. Suppose
UandWare subspaces of Vwith
V=U⊕W.
Each vectorv∈Vcan be written uniquely in the form
v=u+w,
whereu∈Uandw∈W. With this representation, define PU,W∈L(V) PU,W is often called the
projection ontoUwith
null spaceW.by
PU,Wv=u.
You should verify that PU,Wv=vif and only if v∈U. Interchanging
the roles ofUandWin the representation above, we have PW,Uv=w.
Thusv=PU,Wv+PW,Uvfor everyv∈V. You should verify that
PU,W2=PU,W; furthermore range PU,W=Uand nullPU,W=W.
We have seen an example of an operator on R2with no eigenvalues.
The following theorem shows that no such example exists on R3.
5.26 Theorem: Every operator on an odd-dimensional real vector
space has an eigenvalue.
Proof: SupposeVis a real vector space with odd dimension. We
will prove that every operator on Vhas an eigenvalue by induction (in
steps of size 2) on the dimension of V. To get started, note that the
desired result obviously holds if dim V=1.
Now suppose that dim Vis an odd number greater than 1. Using
induction, we can assume that the desired result holds for all operators
Invariant Subspaces on Real Vector Spaces 93
on all real vector spaces with dimension 2 less than dim V. Suppose
T∈L(V). We need to prove that Thas an eigenvalue. If it does, we are
done. If not, then by 5.24 there is a two-dimensional subspace UofV
that is invariant under T. LetWbe any subspace of Vsuch that
V=U⊕W;
2.13 guarantees that such a Wexists.
BecauseWhas dimension 2 less than dim V, we would like to apply
our induction hypothesis to T|W. However,Wmight not be invariant
underT, meaning that T|Wmight not be an operator on W. We will
compose with the projection PW,Uto get an operator on W. Specifically,
defineS∈L(W)by
Sw=PW,U(Tw)
forw∈W. By our induction hypothesis, Shas an eigenvalue λ.W e
will show that this λis also an eigenvalue for T.
Letw∈Wbe a nonzero eigenvector for Scorresponding to the
eigenvalueλ; thus(S−λI)w=0. We would be done if wwere an
eigenvector for Twith eigenvalue λ; unfortunately that need not be
true. So we will look for an eigenvector of TinU+span(w).T o d o
that, consider a typical vector u+awinU+span(w), where u∈U
anda∈R. We have
(T−λI)(u+aw)=Tu−λu+a(Tw−λw)
=Tu−λu+a(PU,W(Tw)+PW,U(Tw)−λw)
=Tu−λu+a(PU,W(Tw)+Sw−λw)
=Tu−λu+aPU,W(Tw).
Note that on the right side of the last equation, Tu∈U(becauseU
is invariant under T),λu∈U(becauseu∈U), andaPU,W(Tw)∈U
(from the definition of PU,W). ThusT−λImapsU+span(w) intoU.
BecauseU+span(w) has a larger dimension than U, this means that
(T−λI)|U+span(w) is not injective (see 3.5). In other words, there exists
a nonzero vector v∈U+span(w)⊂Vsuch that(T−λI)v=0. Thus
Thas an eigenvalue, as desired.
94 Chapter 5.Eigenvalues and Eigenvectors
Exercises
1. Suppose T∈L(V). Prove that if U1,...,Umare subspaces of V
invariant under T, thenU1+···+Umis invariant under T.
2. Suppose T∈L(V). Prove that the intersection of any collection
of subspaces of Vinvariant under Tis invariant under T.
3. Prove or give a counterexample: if Uis a subspace of Vthat is
invariant under every operator on V, thenU={0}orU=V.
4. Suppose that S,T∈L(V)are such that ST=TS. Prove that
null(T−λI)is invariant under Sfor everyλ∈F.
5. Define T∈L(F2)by
T(w,z)=(z,w).
Find all eigenvalues and eigenvectors of T.
6. Define T∈L(F3)by
T(z 1,z2,z3)=(2z 2,0,5z3).
Find all eigenvalues and eigenvectors of T.
7. Suppose nis a positive integer and T∈L(Fn)is defined by
T(x 1,...,xn)=(x1+···+xn,...,x 1+···+xn);
in other words, Tis the operator whose matrix (with respect to
the standard basis) consists of all 1’s. Find all eigenvalues and
eigenvectors of T.
8. Find all eigenvalues and eigenvectors of the backward shift op-
eratorT∈L(F∞)defined by
T(z 1,z2,z3,...)=(z2,z3,...).
9. Suppose T∈L(V)and dim range T=k. Prove that Thas at
mostk+1 distinct eigenvalues.
10. Suppose T∈L(V)is invertible and λ∈F\{0}. Prove that λis
an eigenvalue of Tif and only if1
λis an eigenvalue of T−1.
Exercises 95
11. Suppose S,T∈L(V). Prove that STandTShave the same eigen-
values.
12. Suppose T∈L(V)is such that every vector in Vis an eigenvector
ofT. Prove thatTis a scalar multiple of the identity operator.
13. Suppose T∈L(V)is such that every subspace of Vwith di-
mension dim V−1 is invariant under T. Prove thatTis a scalar
multiple of the identity operator.
14. Suppose S,T∈L(V)andSis invertible. Prove that if p∈P(F)
is a polynomial, then
p(STS−1)=Sp(T)S−1.
15. Suppose F=C,T∈L(V),p∈P(C), anda∈C. Prove that ais
an eigenvalue of p(T) if and only if a=p(λ) for some eigenvalue
λofT.
16. Show that the result in the previous exercise does not hold if C
is replaced with R.
17. Suppose Vis a complex vector space and T∈L(V). Prove
thatThas an invariant subspace of dimension jfor eachj=
1,...,dimV.
18. Give an example of an operator whose matrix with respect to These two exercises
show that 5.16 failswithout the hypothesisthat an upper-
triangular matrix is
under consideration.some basis contains only 0’s on the diagonal, but the operator is
invertible.
19. Give an example of an operator whose matrix with respect to
some basis contains only nonzero numbers on the diagonal, butthe operator is not invertible.
20. Suppose that T∈L(V)has dimVdistinct eigenvalues and that
S∈L(V)has the same eigenvectors as T(not necessarily with
the same eigenvalues). Prove that ST=TS.
21. Suppose P∈L(V)andP
2=P. Prove thatV=nullP⊕rangeP.
22. Suppose V=U⊕W, whereUandWare nonzero subspaces of V.
Find all eigenvalues and eigenvectors of PU,W.
96 Chapter 5.Eigenvalues and Eigenvectors
23. Give an example of an operator T∈L(R4)such thatThas no
(real) eigenvalues.
24. Suppose Vis a real vector space and T∈L(V)has no eigenval-
ues. Prove that every subspace of Vinvariant under Thas even
dimension.
Chapter 6
Inner-Product Spaces
In making the definition of a vector space, we generalized the lin-
ear structure (addition and scalar multiplication) of R2and R3.W e
ignored other important features, such as the notions of length andangle. These ideas are embedded in the concept we now investigate,inner products.
Recall that Fdenotes RorC.
Also,Vis a finite-dimensional, nonzero vector space over F.
✽✽✽✽✽✽
97
98 Chapter 6.Inner-Product Spaces
Inner Products
To motivate the concept of inner product, let’s think of vectors in R2
and R3as arrows with initial point at the origin. The length of a vec-
torxinR2orR3is called the norm ofx, denoted/bardblx/bardbl. Thus for
x=(x1,x2)∈R2, we have/bardblx/bardbl=/radicalbig
x12+x22. If we think of vectors
as points instead of
arrows, then /bardblx/bardbl
should be interpreted
as the distance from
the pointxto the
origin.
x -axis1x -axis2
(x , x )2 1
x
The length of this vector xis/radicalbig
x12+x22.
Similarly, for x=(x1,x2,x3)∈R3, we have/bardblx/bardbl=/radicalbig
x12+x22+x32.
Even though we cannot draw pictures in higher dimensions, the gener-alization to R
nis obvious: we define the norm of x=(x1,...,xn)∈Rn
by
/bardblx/bardbl=/radicalBig
x12+···+xn2.
The norm is not linear on Rn. To inject linearity into the discussion,
we introduce the dot product. For x,y∈Rn, the dot product ofx
andy, denotedx·y, is defined by
x·y=x1y1+···+xnyn,
wherex=(x1,...,xn)andy=(y1,...,yn). Note that the dot product
of two vectors in Rnis a number, not a vector. Obviously x·x=/bardblx/bardbl2
for allx∈Rn. In particular, x·x≥0 for allx∈Rn, with equality if
and only ifx=0. Also, ify∈Rnis fixed, then clearly the map from Rn
toRthat sendsx∈Rntox·yis linear. Furthermore, x·y=y·x
for allx,y∈Rn.
An inner product is a generalization of the dot product. At this
point you should be tempted to guess that an inner product is defined
Inner Products 99
by abstracting the properties of the dot product discussed in the para-
graph above. For real vector spaces, that guess is correct. However,so that we can make a definition that will be useful for both real and
complex vector spaces, we need to examine the complex case beforemaking the definition.
Recall that if λ=a+bi, wherea,b∈R, then the absolute value
ofλis defined by
|λ|=/radicalbig
a2+b2,
the complex conjugate of λis defined by
¯λ=a−bi,
and the equation
|λ|2=λ¯λ
connects these two concepts (see page 69 for the definitions and the
basic properties of the absolute value and complex conjugate). Forz=(z
1,...,zn)∈Cn, we define the norm of zby
/bardblz/bardbl=/radicalBig
|z1|2+···+|zn|2.
The absolute values are needed because we want /bardblz/bardblto be a nonnega-
tive number. Note that
/bardblz/bardbl2=z1z1+···+znzn.
We want to think of /bardblz/bardbl2as the inner product of zwith itself, as we
did in Rn. The equation above thus suggests that the inner product of
w=(w1,...,wn)∈Cnwithzshould equal
w1z1+···+wnzn.
If the roles of the wandzwere interchanged, the expression above
would be replaced with its complex conjugate. In other words, we
should expect that the inner product of wwithzequals the complex
conjugate of the inner product of zwithw. With that motivation, we
are now ready to define an inner product on V, which may be a real or
a complex vector space.
An inner product onVis a function that takes each ordered pair
(u,v) of elements of Vto a number /angbracketleftu,v/angbracketright∈F and has the following
properties:
100 Chapter 6.Inner-Product Spaces
positivity
/angbracketleftv,v/angbracketright≥0 for all v∈V; Ifzis a complex
number, then the
statementz≥0means
thatzis real and
nonnegative.definiteness
/angbracketleftv,v/angbracketright=0 if and only if v=0;
additivity in first slot
/angbracketleftu+v,w/angbracketright=/angbracketleftu,w/angbracketright+/angbracketleftv,w/angbracketrightfor allu,v,w∈V;
homogeneity in first slot
/angbracketleftav,w/angbracketright=a/angbracketleftv,w/angbracketrightfor alla∈Fand allv,w∈V;
conjugate symmetry
/angbracketleftv,w/angbracketright=/angbracketleftw,v/angbracketrightfor allv,w∈V.
Recall that every real number equals its complex conjugate. Thus
if we are dealing with a real vector space, then in the last condition
above we can dispense with the complex conjugate and simply statethat/angbracketleftv,w/angbracketright=/angbracketleftw,v/angbracketrightfor allv,w∈V.
An inner-product space is a vector space Valong with an inner
product onV.
The most important example of an inner-product space is F
n.W e
can define an inner product on Fnby If we are dealing with
Rnrather than Cn, then
again the complex
conjugate can be
ignored.6.1/angbracketleft(w 1,...,wn),(z 1,...,zn)/angbracketright=w 1z1+···+wnzn,
as you should verify. This inner product, which provided our motiva-
tion for the definition of an inner product, is called the Euclidean inner
product onFn. When Fnis referred to as an inner-product space, you
should assume that the inner product is the Euclidean inner productunless explicitly told otherwise.
There are other inner products on F
nin addition to the Euclidean
inner product. For example, if c1,...,cnare positive numbers, then we
can define an inner product on Fnby
/angbracketleft(w 1,...,wn),(z 1,...,zn)/angbracketright=c 1w1z1+···+cnwnzn,
as you should verify. Of course, if all the c’s equal 1, then we get the
Euclidean inner product.
As another example of an inner-product space, consider the vector
spacePm(F)of all polynomials with coefficients in Fand degree at
mostm. We can define an inner product on Pm(F)by
Inner Products 101
6.2 /angbracketleftp,q/angbracketright=/integraldisplay1
0p(x)q(x)dx,
as you should verify. Once again, if F=R, then the complex conjugate
is not needed.
Let’s agree for the rest of this chapter that
Vis a finite-dimensional inner-product space over F.
In the definition of an inner product, the conditions of additivity
and homogeneity in the first slot can be combined into a requirement
of linearity in the first slot. More precisely, for each fixed w∈V, the
function that takes vto/angbracketleftv,w/angbracketrightis a linear map from VtoF. Because
every linear map takes 0 to 0, we must have
/angbracketleft0,w/angbracketright=0
for everyw∈V. Thus we also have
/angbracketleftw,0/angbracketright=0
for everyw∈V(by the conjugate symmetry property).
In an inner-product space, we have additivity in the second slot as
well as the first slot. Proof:
/angbracketleftu,v+w/angbracketright=/angbracketleftv+w,u/angbracketright
=/angbracketleftv,u/angbracketright+/angbracketleftw,u/angbracketright
=/angbracketleftv,u/angbracketright+/angbracketleftw,u/angbracketright
=/angbracketleftu,v/angbracketright+/angbracketleftu,w/angbracketright;
hereu,v,w∈V.
In an inner-product space, we have conjugate homogeneity in the
second slot, meaning that /angbracketleftu,av/angbracketright= ¯a/angbracketleftu,v/angbracketrightfor all scalars a∈F.
Proof:
/angbracketleftu,av/angbracketright=/angbracketleftav,u/angbracketright
=a/angbracketleftv,u/angbracketright
=¯a/angbracketleftv,u/angbracketright
=¯a/angbracketleftu,v/angbracketright;
herea∈Fandu,v∈V. Note that in a real vector space, conjugate
homogeneity is the same as homogeneity.
102 Chapter 6.Inner-Product Spaces
Norms
Forv∈V, we define the norm ofv, denoted/bardblv/bardbl,b y
/bardblv/bardbl=/radicalBig
/angbracketleftv,v/angbracketright.
For example, if (z1,...,zn)∈Fn(with the Euclidean inner product),
then
/bardbl(z1,...,zn)/bardbl=/radicalBig
|z1|2+···+|zn|2.
As another example, if p∈Pm(F)(with inner product given by 6.2),
then
/bardblp/bardbl=/radicalBigg/integraldisplay1
0|p(x)|2dx.
Note that/bardblv/bardbl=0 if and only if v=0 (because/angbracketleftv,v/angbracketright=0 if and only
ifv=0). Another easy property of the norm is that /bardblav/bardbl=|a|/bardblv/bardbl
for alla∈Fand allv∈V. Here’s the proof:
/bardblav/bardbl2=/angbracketleftav,av/angbracketright
=a/angbracketleftv,av/angbracketright
=a¯a/angbracketleftv,v/angbracketright
=|a|2/bardblv/bardbl2;
taking square roots now gives the desired equality. This proof illus-
trates a general principle: working with norms squared is usually easierthan working directly with norms.
Two vectors u,v∈Vare said to be orthogonal if/angbracketleftu,v/angbracketright=0. Note
Some mathematicians
use the term
perpendicular, which
means the same as
orthogonal.that the order of the vectors does not matter because /angbracketleftu,v/angbracketright=0i f
and only if/angbracketleftv,u/angbracketright=0. Instead of saying that uandvare orthogonal,
sometimes we say that uis orthogonal to v. Clearly 0 is orthogonal
to every vector. Furthermore, 0 is the only vector that is orthogonal toitself.
For the special case where V=R
2, the next theorem is over 2,500 The word orthogonal
comes from the Greek
word orthogonios,
which means
right-angled.years old.
6.3 Pythagorean Theorem: Ifu,vare orthogonal vectors in V, then
6.4 /bardblu+v/bardbl2=/bardblu/bardbl2+/bardblv/bardbl2.
Norms 103
Proof: Suppose that u,v are orthogonal vectors in V. Then The proof of the
Pythagorean theoremshows that 6.4 holds ifand only if/angbracketleftu,v/angbracketright+/angbracketleftv,u/angbracketright, which
equals 2R e/angbracketleftu,v/angbracketright,i s0.
Thus the converse of
the Pythagoreantheorem holds in real
inner-product spaces./bardblu+v/bardbl2=/angbracketleftu+v,u+v/angbracketright
=/bardblu/bardbl2+/bardblv/bardbl2+/angbracketleftu,v/angbracketright+/angbracketleftv,u/angbracketright
=/bardblu/bardbl2+/bardblv/bardbl2,
as desired.
Supposeu,v∈V. We would like to write uas a scalar multiple of v
plus a vector worthogonal to v, as suggested in the next picture.
0u
vλvw
An orthogonal decomposition
To discover how to write uas a scalar multiple of vplus a vector or-
thogonal tov, leta∈Fdenote a scalar. Then
u=av+(u−av).
Thus we need to choose aso thatvis orthogonal to (u−av). In other
words, we want
0=/angbracketleftu−av,v/angbracketright=/angbracketleftu,v/angbracketright−a/bardblv/bardbl2.
The equation above shows that we should choose ato be/angbracketleftu,v/angbracketright//bardblv/bardbl2
(assume that v/negationslash=0 to avoid division by 0). Making this choice of a,w e
can write
6.5 u=/angbracketleftu,v/angbracketright
/bardblv/bardbl2v+/parenleftbigg
u−/angbracketleftu,v/angbracketright
/bardblv/bardbl2v/parenrightbigg
.
As you should verify, if v/negationslash=0 then the equation above writes uas a
scalar multiple of vplus a vector orthogonal to v.
The equation above will be used in the proof of the next theorem,
which gives one of the most important inequalities in mathematics.
104 Chapter 6.Inner-Product Spaces
6.6 Cauchy-Schwarz Inequality: Ifu,v∈V, then In 1821 the French
mathematician
Augustin-Louis Cauchy
showed that this
inequality holds for the
inner product defined
by 6.1. In 1886 the
German mathematician
Herman Schwarz
showed that this
inequality holds for the
inner product defined
by 6.2.6.7 |/angbracketleftu,v/angbracketright|≤/bardblu/bardbl/bardblv/bardbl.
This inequality is an equality if and only if one of u,v is a scalar mul-
tiple of the other.
Proof: Letu,v∈V.I fv=0, then both sides of 6.7 equal 0 and
the desired inequality holds. Thus we can assume that v/negationslash=0. Consider
the orthogonal decomposition
u=/angbracketleftu,v/angbracketright
/bardblv/bardbl2v+w,
wherewis orthogonal to v(herewequals the second term on the right
side of 6.5). By the Pythagorean theorem,
/bardblu/bardbl2=/vextenddouble/vextenddouble/vextenddouble/vextenddouble/angbracketleftu,v/angbracketright
/bardblv/bardbl2v/vextenddouble/vextenddouble/vextenddouble/vextenddouble2
+/bardblw/bardbl2
=|/angbracketleftu,v/angbracketright|2
/bardblv/bardbl2+/bardblw/bardbl2
≥|/angbracketleftu,v/angbracketright|2
/bardblv/bardbl2. 6.8
Multiplying both sides of this inequality by /bardblv/bardbl2and then taking square
roots gives the Cauchy-Schwarz inequality 6.7.
Looking at the proof of the Cauchy-Schwarz inequality, note that 6.7
is an equality if and only if 6.8 is an equality. Obviously this happens ifand only ifw=0. Butw=0 if and only if uis a multiple of v(see 6.5).
Thus the Cauchy-Schwarz inequality is an equality if and only if uis a
scalar multiple of vorvis a scalar multiple of u(or both; the phrasing
has been chosen to cover cases in which either uorvequals 0).
The next result is called the triangle inequality because of its geo-
metric interpretation that the length of any side of a triangle is lessthan the sum of the lengths of the other two sides.
v
uu+v
The triangle inequality
Norms 105
6.9 Triangle Inequality: Ifu,v∈V, then The triangle inequality
can be used to showthat the shortest pathbetween two points is a
straight line segment.6.10 /bardblu+v/bardbl≤/bardblu/bardbl+/bardblv/bardbl.
This inequality is an equality if and only if one of u,v is a nonnegative
multiple of the other.
Proof: Letu,v∈V. Then
/bardblu+v/bardbl2=/angbracketleftu+v,u+v/angbracketright
=/angbracketleftu,u/angbracketright+/angbracketleftv,v/angbracketright+/angbracketleftu,v/angbracketright+/angbracketleftv,u/angbracketright
=/angbracketleftu,u/angbracketright+/angbracketleftv,v/angbracketright+/angbracketleftu,v/angbracketright+/angbracketleftu,v/angbracketright
=/bardblu/bardbl2+/bardblv/bardbl2+2R e/angbracketleftu,v/angbracketright
≤/bardblu/bardbl2+/bardblv/bardbl2+2|/angbracketleftu,v/angbracketright| 6.11
≤/bardblu/bardbl2+/bardblv/bardbl2+2/bardblu/bardbl/bardblv/bardbl 6.12
=(/bardblu/bardbl+/bardblv/bardbl)2,
where 6.12 follows from the Cauchy-Schwarz inequality (6.6). Taking
square roots of both sides of the inequality above gives the triangle
inequality 6.10.
The proof above shows that the triangle inequality 6.10 is an equality
if and only if we have equality in 6.11 and 6.12. Thus we have equalityin the triangle inequality 6.10 if and only if
6.13 /angbracketleftu,v/angbracketright=/bardblu/bardbl/bardblv/bardbl.
If one ofu,vis a nonnegative multiple of the other, then 6.13 holds, as
you should verify. Conversely, suppose 6.13 holds. Then the condition
for equality in the Cauchy-Schwarz inequality (6.6) implies that one ofu,v must be a scalar multiple of the other. Clearly 6.13 forces the
scalar in question to be nonnegative, as desired.
The next result is called the parallelogram equality because of its
geometric interpretation: in any parallelogram, the sum of the squares
of the lengths of the diagonals equals the sum of the squares of the
lengths of the four sides.
106 Chapter 6.Inner-Product Spaces
u+v
uu−vu
v v
The parallelogram equality
6.14 Parallelogram Equality: Ifu,v∈V, then
/bardblu+v/bardbl2+/bardblu−v/bardbl2=2(/bardblu/bardbl2+/bardblv/bardbl2).
Proof: Letu,v∈V. Then
/bardblu+v/bardbl2+/bardblu−v/bardbl2=/angbracketleftu+v,u+v/angbracketright+/angbracketleftu−v,u−v/angbracketright
=/bardblu/bardbl2+/bardblv/bardbl2+/angbracketleftu,v/angbracketright+/angbracketleftv,u/angbracketright
+/bardblu/bardbl2+/bardblv/bardbl2−/angbracketleftu,v/angbracketright−/angbracketleftv,u/angbracketright
=2(/bardblu/bardbl2+/bardblv/bardbl2),
as desired.
Orthonormal Bases
A list of vectors is called orthonormal if the vectors in it are pair-
wise orthogonal and each vector has norm 1. In other words, a list
(e1,...,em)of vectors in Vis orthonormal if /angbracketleftej,ek/angbracketrightequals 0 when
j/negationslash=kand equals 1 when j=k(forj,k=1,...,m ). For example, the
standard basis in Fnis orthonormal. Orthonormal lists are particularly
easy to work with, as illustrated by the next proposition.
6.15 Proposition: If(e1,...,em)is an orthonormal list of vectors
inV, then
/bardbla1e1+···+amem/bardbl2=|a1|2+···+|am|2
for alla1,...,am∈F.
Proof: Because each ejhas norm 1, this follows easily from re-
peated applications of the Pythagorean theorem (6.3).
Now we have the following easy but important corollary.
Orthonormal Bases 107
6.16 Corollary: Every orthonormal list of vectors is linearly inde-
pendent.
Proof: Suppose(e1,...,em)is an orthonormal list of vectors in V
anda1,...,am∈Fare such that
a1e1+···+amem=0.
Then|a1|2+···+|am|2=0 (by 6.15), which means that all the aj’s
are 0, as desired.
An orthonormal basis ofVis an orthonormal list of vectors in V
that is also a basis of V. For example, the standard basis is an ortho-
normal basis of Fn. Every orthonormal list of vectors in Vwith length
dimVis automatically an orthonormal basis of V(proof: by the pre-
vious corollary, any such list must be linearly independent; because ithas the right length, it must be a basis—see 2.17). To illustrate thisprinciple, consider the following list of four vectors in R
4:
/parenleftbig
(1
2,1
2,1
2,1
2),(1
2,1
2,−1
2,−1
2),(1
2,−1
2,−1
2,1
2),(−1
2,1
2,−1
2,1
2)/parenrightbig
.
The verification that this list is orthonormal is easy (do it!); because we
have an orthonormal list of length four in a four-dimensional vector
space, it must be an orthonormal basis.
In general, given a basis (e1,...,en)ofVand a vector v∈V,w e
know that there is some choice of scalars a1,...,amsuch that
v=a1e1+···+anen,
but finding the aj’s can be difficult. The next theorem shows, however,
that this is easy for an orthonormal basis.
6.17 Theorem: Suppose(e1,...,en)is an orthonormal basis of V. The importance of
orthonormal bases
stems mainly from thistheorem.Then
6.18 v=/angbracketleftv,e 1/angbracketrighte1+···+/angbracketleftv,en/angbracketrighten
and
6.19 /bardblv/bardbl2=|/angbracketleftv,e 1/angbracketright|2+···+|/angbracketleftv,en/angbracketright|2
for everyv∈V.
108 Chapter 6.Inner-Product Spaces
Proof: Letv∈V. Because(e1,...,en)is a basis of V, there exist
scalarsa1,...,ansuch that
v=a1e1+···+anen.
Take the inner product of both sides of this equation with ej, get-
ting/angbracketleftv,ej/angbracketright=aj. Thus 6.18 holds. Clearly 6.19 follows from 6.18
and 6.15.
Now that we understand the usefulness of orthonormal bases, how
do we go about finding them? For example, does Pm(F), with inner
product given by integration on [0,1](see 6.2), have an orthonormal
basis? As we will see, the next result will lead to answers to these ques-tions. The algorithm used in the next proof is called the Gram-Schmidt
procedure. It gives a method for turning a linearly independent list into
The Danish
mathematician Jorgen
Gram (1850–1916) and
the German
mathematician Erhard
Schmidt (1876–1959)
popularized this
algorithm for
constructing
orthonormal lists.an orthonormal list with the same span as the original list.
6.20 Gram-Schmidt: If(v1,...,vm)is a linearly independent list
of vectors in V, then there exists an orthonormal list (e1,...,em)of
vectors inVsuch that
6.21 span(v 1,...,vj)=span(e 1,...,ej)
forj=1,...,m .
Proof: Suppose(v1,...,vm)is a linearly independent list of vec-
tors inV. To construct the e’s, start by setting e1=v1//bardblv 1/bardbl. This
satisfies 6.21 for j=1. We will choose e2,...,eminductively, as fol-
lows. Suppose j>1 and an orthornormal list (e1,...,ej−1)has been
chosen so that
6.22 span(v 1,...,vj−1)=span(e 1,...,ej−1).
Let
6.23 ej=vj−/angbracketleftvj,e1/angbracketrighte1−···−/angbracketleftvj,ej−1/angbracketrightej−1
/bardblvj−/angbracketleftvj,e1/angbracketrighte1−···−/angbracketleftvj,ej−1/angbracketrightej−1/bardbl.
Note thatvj∉span(v 1,...,vj−1)(because(v1,...,vm)is linearly inde-
pendent) and thus vj∉span(e 1,...,ej−1). Hence we are not dividing
by 0 in the equation above, and so ejis well defined. Dividing a vector
by its norm produces a new vector with norm 1; thus /bardblej/bardbl=1.
Orthonormal Bases 109
Let 1≤k<j . Then
/angbracketleftej,ek/angbracketright=/angbracketleftBigg
vj−/angbracketleftvj,e1/angbracketrighte1−···−/angbracketleftvj,ej−1/angbracketrightej−1
/bardblvj−/angbracketleftvj,e1/angbracketrighte1−···−/angbracketleftvj,ej−1/angbracketrightej−1/bardbl,ek/angbracketrightBigg
=/angbracketleftvj,ek/angbracketright−/angbracketleftvj,ek/angbracketright
/bardblvj−/angbracketleftvj,e1/angbracketrighte1−···−/angbracketleftvj,ej−1/angbracketrightej−1/bardbl
=0.
Thus(e1,...,ej)is an orthonormal list.
From 6.23, we see that vj∈span(e 1,...,ej). Combining this infor-
mation with 6.22 shows that
span(v 1,...,vj)⊂span(e 1,...,ej).
Both lists above are linearly independent (the v’s by hypothesis, the e’s
by orthonormality and 6.16). Thus both subspaces above have dimen-sionj, and hence they must be equal, completing the proof.
Now we can settle the question of the existence of orthonormal
bases.
6.24 Corollary: Every finite-dimensional inner-product space has an Until this corollary,
nothing we had done
with inner-product
spaces required our
standing assumption
thatVis finite
dimensional.orthonormal basis.
Proof: Choose a basis of V. Apply the Gram-Schmidt procedure
(6.20) to it, producing an orthonormal list. This orthonormal list islinearly independent (by 6.16) and its span equals V. Thus it is an
orthonormal basis of V.
As we will soon see, sometimes we need to know not only that an
orthonormal basis exists, but also that any orthonormal list can beextended to an orthonormal basis. In the next corollary, the Gram-Schmidt procedure shows that such an extension is always possible.
6.25 Corollary: Every orthonormal list of vectors in Vcan be ex-
tended to an orthonormal basis of V.
Proof: Suppose(e
1,...,em)is an orthonormal list of vectors in V.
Then(e1,...,em)is linearly independent (by 6.16), and hence it can be
extended to a basis (e1,...,em,v1,...,vn)ofV(see 2.12). Now apply
110 Chapter 6.Inner-Product Spaces
the Gram-Schmidt procedure (6.20) to (e1,...,em,v1,...,vn), produc-
ing an orthonormal list
6.26 (e1,...,em,f1,...,fn);
here the Gram-Schmidt procedure leaves the first mvectors unchanged
because they are already orthonormal. Clearly 6.26 is an orthonormal
basis ofVbecause it is linearly independent (by 6.16) and its span
equalsV. Hence we have our extension of (e1,...,em)to an orthonor-
mal basis of V.
Recall that a matrix is called upper triangular if all entries below the
diagonal equal 0. In other words, an upper-triangular matrix looks likethis:
∗∗
...
0∗
.
In the last chapter we showed that if Vis a complex vector space, then
for each operator on Vthere is a basis with respect to which the matrix
of the operator is upper triangular (see 5.13). Now that we are dealingwith inner-product spaces, we would like to know when there exists an
orthonormal basis with respect to which we have an upper-triangular
matrix. The next corollary shows that the existence of any basis withrespect to which Thas an upper-triangular matrix implies the existence
of an orthonormal basis with this property. This result is true on bothreal and complex vector spaces (though on a real vector space, the hy-pothesis holds only for some operators).
6.27 Corollary: SupposeT∈L(V).I fThas an upper-triangular
matrix with respect to some basis of V, thenThas an upper-triangular
matrix with respect to some orthonormal basis of V.
Proof: SupposeThas an upper-triangular matrix with respect to
some basis(v
1,...,vn)ofV. Thus span(v 1,...,vj)is invariant under
Tfor eachj=1,...,n (see 5.12).
Apply the Gram-Schmidt procedure to (v1,...,vn), producing an
orthonormal basis (e1,...,en)ofV. Because
span(e 1,...,ej)=span(v 1,...,vj)
Orthogonal Projections and Minimization Problems 111
for eachj(see 6.21), we conclude that span(e 1,...,ej)is invariant un-
derTfor eachj=1,...,n . Thus, by 5.12, Thas an upper-triangular
matrix with respect to the orthonormal basis (e1,...,en).
The next result is an important application of the corollary above.
6.28 Corollary: SupposeVis a complex vector space and T∈L(V). This result is
sometimes called
Schur’s theorem. The
German mathematician
Issai Schur published
the first proof of this
result in 1909.ThenThas an upper-triangular matrix with respect to some orthonor-
mal basis of V.
Proof: This follows immediately from 5.13 and 6.27.
Orthogonal Projections and
Minimization Problems
IfUis a subset of V, then the orthogonal complement ofU, de-
notedU⊥, is the set of all vectors in Vthat are orthogonal to every
vector inU:
U⊥={v∈V:/angbracketleftv,u/angbracketright=0 for allu∈U}.
You should verify that U⊥is always a subspace of V, thatV⊥={0},
and that{0}⊥=V. Also note that if U1⊂U2, thenU⊥
1⊃U⊥
2.
Recall that if U1,U2are subspaces of V, thenVis the direct sum of
U1andU2(writtenV=U1⊕U2) if each element of Vcan be written in
exactly one way as a vector in U1plus a vector in U2. The next theorem
shows that every subspace of an inner-product space leads to a naturaldirect sum decomposition of the whole space.
6.29 Theorem: IfUis a subspace of V, then
V=U⊕U
⊥.
Proof: Suppose that Uis a subspace of V. First we will show that
6.30 V=U+U⊥.
To do this, suppose v∈V. Let(e1,...,em)be an orthonormal basis
ofU. Obviously
112 Chapter 6.Inner-Product Spaces
6.31
v=/angbracketleftv,e 1/angbracketrighte1+···+/angbracketleftv,em/angbracketrightem/bracehtipupleft /bracehtipdownright/bracehtipdownleft /bracehtipupright
u+v−/angbracketleftv,e 1/angbracketrighte1−···−/angbracketleftv,em/angbracketrightem/bracehtipupleft /bracehtipdownright/bracehtipdownleft /bracehtipupright
w.
Clearlyu∈U. Because(e1,...,em)is an orthonormal list, for each j
we have
/angbracketleftw,ej/angbracketright=/angbracketleftv,ej/angbracketright−/angbracketleftv,ej/angbracketright
=0.
Thuswis orthogonal to every vector in span (e1,...,em). In other
words,w∈U⊥. Thus we have written v=u+w, whereu∈U
andw∈U⊥, completing the proof of 6.30.
Ifv∈U∩U⊥, thenv(which is inU) is orthogonal to every vector
inU(includingvitself), which implies that /angbracketleftv,v/angbracketright=0, which implies
thatv=0. Thus
6.32 U∩U⊥={0}.
Now 6.30 and 6.32 imply that V=U⊕U⊥(see 1.9).
The next corollary is an important consequence of the last theorem.
6.33 Corollary: IfUis a subspace of V, then
U=(U⊥)⊥.
Proof: Suppose that Uis a subspace of V. First we will show that
6.34 U⊂(U⊥)⊥.
To do this, suppose that u∈U. Then/angbracketleftu,v/angbracketright=0 for every v∈U⊥(by
the definition of U⊥). Becauseuis orthogonal to every vector in U⊥,
we haveu∈(U⊥)⊥, completing the proof of 6.34.
To prove the inclusion in the other direction, suppose v∈(U⊥)⊥.
By 6.29, we can write v=u+w, whereu∈Uandw∈U⊥. We have
v−u=w∈U⊥. Becausev∈(U⊥)⊥andu∈(U⊥)⊥(from 6.34), we
havev−u∈(U⊥)⊥. Thusv−u∈U⊥∩(U⊥)⊥, which implies that v−u
is orthogonal to itself, which implies that v−u=0, which implies that
v=u, which implies that v∈U. Thus(U⊥)⊥⊂U, which along with
6.34 completes the proof.
Orthogonal Projections and Minimization Problems 113
SupposeUis a subspace of V. The decomposition V=U⊕U⊥given
by 6.29 means that each vector v∈Vcan be written uniquely in the
form
v=u+w,
whereu∈Uandw∈U⊥. We use this decomposition to define an op-
erator onV, denotedPU, called the orthogonal projection ofVontoU.
Forv∈V, we definePUvto be the vector uin the decomposition above.
In the notation introduced in the last chapter, we have PU=PU,U⊥. You
should verify that PU∈L(V)and that it has the following proper-
ties:
•rangePU=U;
•nullPU=U⊥;
•v−PUv∈U⊥for everyv∈V;
•PU2=PU;
•/bardblPUv/bardbl≤/bardblv/bardblfor everyv∈V.
Furthermore, from the decomposition 6.31 used in the proof of 6.29
we see that if (e1,...,em)is an orthonormal basis of U, then
6.35 PUv=/angbracketleftv,e 1/angbracketrighte1+···+/angbracketleftv,em/angbracketrightem
for everyv∈V.
The following problem often arises: given a subspace UofVand
a pointv∈V, find a point u∈Usuch that/bardblv−u/bardblis as small as
possible. The next proposition shows that this minimization problemis solved by taking u=P
Uv.
6.36 Proposition: SupposeUis a subspace of Vandv∈V. Then The remarkable
simplicity of the
solution to thisminimization problemhas led to many
applications ofinner-product spaces
outside of pure
mathematics./bardblv−PUv/bardbl≤/bardblv−u/bardbl
for everyu∈U. Furthermore, if u∈Uand the inequality above is an
equality, then u=PUv.
Proof: Supposeu∈U. Then
/bardblv−PUv/bardbl2≤/bardblv−PUv/bardbl2+/bardblPUv−u/bardbl26.37
=/bardbl(v−PUv)+(PUv−u)/bardbl26.38
=/bardblv−u/bardbl2,
114 Chapter 6.Inner-Product Spaces
where 6.38 comes from the Pythagorean theorem (6.3), which applies
becausev−PUv∈U⊥andPUv−u∈U. Taking square roots gives the
desired inequality.
Our inequality is an equality if and only if 6.37 is an equality, which
happens if and only if /bardblPUv−u/bardbl=0, which happens if and only if
u=PUv.
0v
U
PUv
PUvis the closest point in Utov.
The last proposition is often combined with the formula 6.35 to
compute explicit solutions to minimization problems. As an illustra-tion of this procedure, consider the problem of finding a polynomial u
with real coefficients and degree at most 5 that on the interval [−π,π]
approximates sin xas well as possible, in the sense that
/integraldisplay
π
−π|sinx−u(x)|2dx
is as small as possible. To solve this problem, let C[−π,π] denote the
real vector space of continuous real-valued functions on [−π,π] with
inner product
6.39 /angbracketleftf,g/angbracketright=/integraldisplayπ
−πf(x)g(x)dx.
Letv∈C[−π,π] be the function defined by v(x)=sinx. LetU
denote the subspace of C[−π,π] consisting of the polynomials with
real coefficients and degree at most 5. Our problem can now be re-formulated as follows: find u∈Usuch that/bardblv−u/bardblis as small as
possible.
To compute the solution to our approximation problem, first apply
the Gram-Schmidt procedure (using the inner product given by 6.39)
Orthogonal Projections and Minimization Problems 115
to the basis(1,x,x2,x3,x4,x5)ofU, producing an orthonormal basis
(e1,e2,e3,e4,e5,e6)ofU. Then, again using the inner product given A machine that can
perform integrations isuseful here. by 6.39, compute PUvusing 6.35 (with m=6). Doing this computation
shows thatPUvis the function
6.40 0.987862x−0.155271x3+0.00564312x5,
where theπ’s that appear in the exact answer have been replaced with
a good decimal approximation.
By 6.36, the polynomial above should be about as good an approxi-
mation to sin xon[−π,π] as is possible using polynomials of degree
at most 5. To see how good this approximation is, the picture belowshows the graphs of both sin xand our approximation 6.40 over the
interval[−π,π] .
-3 -2 -1 1 2 3
-1-0.50.51
Graphs of sinxand its approximation 6.40
Our approximation 6.40 is so accurate that the two graphs are almost
identical—our eyes may see only one graph!
Another well-known approximation to sin xby a polynomial of de-
gree 5 is given by the Taylor polynomial
6.41 x−x3
3!+x5
5!.
To see how good this approximation is, the next picture shows the
graphs of both sin xand the Taylor polynomial 6.41 over the interval
[−π,π] .
116 Chapter 6.Inner-Product Spaces
-3 -2 -1 1 2 3
-1-0.50.51
Graphs of sinxand the Taylor polynomial 6.41
The Taylor polynomial is an excellent approximation to sin xforx
near 0. But the picture above shows that for |x|>2, the Taylor poly-
nomial is not so accurate, especially compared to 6.40. For example,takingx=3, our approximation 6.40 estimates sin 3 with an error of
about 0.001, but the Taylor series 6.41 estimates sin 3 with an error of
about 0.4. Thus at x=3, the error in the Taylor series is hundreds of
times larger than the error given by 6.40. Linear algebra has helped usdiscover an approximation to sin xthat improves upon what we learned
in calculus!
We derived our approximation 6.40 by using 6.35 and 6.36. Our
standing assumption that Vis finite dimensional fails when Vequals
C[−π,π] , so we need to justify our use of those results in this case.
First, reread the proof of 6.29, which states that if Uis a subspace of V,
then
6.42 V=U⊕U
⊥.
Note that the proof uses the finite dimensionality of U(to get a basis If we allowVto be
infinite dimensional
and allowUto be an
infinite-dimensional
subspace of V, then
6.42 is not necessarily
true without additional
hypotheses.ofU) but that it works fine regardless of whether or not Vis finite
dimensional. Second, note that the definition and properties of PU(in-
cluding 6.35) require only 6.29 and thus require only that U(but not
necessarilyV) be finite dimensional. Finally, note that the proof of 6.36
does not require the finite dimensionality of V. Conclusion: for v∈V
andUa subspace of V, the procedure discussed above for finding the
vectoru∈Uthat makes/bardblv−u/bardblas small as possible works if Uis finite
dimensional, regardless of whether or not Vis finite dimensional. In
the example above Uwas indeed finite dimensional (we had dim U=6),
so everything works as expected.
Linear Functionals and Adjoints 117
Linear Functionals and Adjoints
Alinear functional onVis a linear map from Vto the scalars F.
For example, the function ϕ:F3→Fdefined by
6.43 ϕ(z 1,z2,z3)=2z1−5z2+z3
is a linear functional on F3. As another example, consider the inner-
product space P6(R)(here the inner product is multiplication followed
by integration on [0,1]; see 6.2). The function ϕ:P6(R)→Rdefined
by
6.44 ϕ(p)=/integraldisplay1
0p(x)( cosx)dx
is a linear functional on P6(R).
Ifv∈V, then the map that sends uto/angbracketleftu,v/angbracketrightis a linear functional
onV. The next result shows that every linear functional on Vis of this
form. To illustrate this theorem, note that for the linear functional ϕ
defined by 6.43, we can take v=(2,−5,1)∈F3. The linear functional
ϕdefined by 6.44 better illustrates the power of the theorem below be-
cause for this linear functional, there is no obvious candidate for v(the
function cos xis not eligible because it is not an element of P6(R)).
6.45 Theorem: Supposeϕis a linear functional on V. Then there is
a unique vector v∈Vsuch that
ϕ(u)=/angbracketleftu,v/angbracketright
for everyu∈V.
Proof: First we show that there exists a vector v∈Vsuch that
ϕ(u)=/angbracketleftu,v/angbracketrightfor everyu∈V. Let(e1,...,en)be an orthonormal
basis ofV. Then
ϕ(u)=ϕ(/angbracketleftu,e 1/angbracketrighte1+···+/angbracketleftu,en/angbracketrighten)
=/angbracketleftu,e 1/angbracketrightϕ(e 1)+···+/angbracketleftu,en/angbracketrightϕ(en)
=/angbracketleftu,ϕ(e 1)e1+···+ϕ(en)en/angbracketright
for everyu∈V, where the first equality comes from 6.17. Thus setting
v=ϕ(e 1)e1+···+ϕ(en)en, we haveϕ(u)=/angbracketleftu,v/angbracketrightfor everyu∈V,
as desired.
118 Chapter 6.Inner-Product Spaces
Now we prove that only one vector v∈Vhas the desired behavior.
Supposev1,v2∈Vare such that
ϕ(u)=/angbracketleftu,v 1/angbracketright=/angbracketleftu,v 2/angbracketright
for everyu∈V. Then
0=/angbracketleftu,v 1/angbracketright−/angbracketleftu,v 2/angbracketright=/angbracketleftu,v 1−v2/angbracketright
for everyu∈V. Takingu=v1−v2shows thatv1−v2=0. In other
words,v1=v2, completing the proof of the uniqueness part of the
theorem.
In addition to V, we need another finite-dimensional inner-product
space.
Let’s agree that for the rest of this chapter
Wis a finite-dimensional, nonzero, inner-product space over F.
LetT∈L(V,W). The adjoint ofT, denotedT∗, is the function from The word adjoint has
another meaning in
linear algebra. We will
not need the second
meaning, related to
inverses, in this book.
Just in case you
encountered the
second meaning for
adjoint elsewhere, be
warned that the two
meanings for adjoint
are unrelated to one
another.WtoVdefined as follows. Fix w∈W. Consider the linear functional
onVthat mapsv∈Vto/angbracketleftTv,w/angbracketright. LetT∗wbe the unique vector in V
such that this linear functional is given by taking inner products withT
∗w(6.45 guarantees the existence and uniqueness of a vector in V
with this property). In other words, T∗wis the unique vector in V
such that
/angbracketleftTv,w/angbracketright=/angbracketleftv,T∗w/angbracketright
for allv∈V.
Let’s work out an example of how the adjoint is computed. Define
T:R3→R2by
T(x 1,x2,x3)=(x2+3x3,2x1).
ThusT∗will be a function from R2toR3. To compute T∗, fix a point
(y1,y2)∈R2. Then
/angbracketleft(x1,x2,x3),T∗(y1,y2)/angbracketright=/angbracketleftT(x 1,x2,x3),(y 1,y2)/angbracketright
=/angbracketleft(x2+3x3,2x1),(y 1,y2)/angbracketright
=x2y1+3x3y1+2x1y2
=/angbracketleft(x1,x2,x3),(2y 2,y1,3y1)/angbracketright
for all(x1,x2,x3)∈R3. This shows that
Linear Functionals and Adjoints 119
T∗(y1,y2)=(2y 2,y1,3y1).
Note that in the example above, T∗turned out to be not just a func- Adjoints play a crucial
role in the important
results in the nextchapter.tion from R2toR3, but a linear map. That is true in general. Specif-
ically, ifT∈L(V,W), then T∗∈L(W,V). To prove this, suppose
T∈L(V,W). Let’s begin by checking additivity. Fix w1,w2∈W.
Then
/angbracketleftTv,w 1+w2/angbracketright=/angbracketleftTv,w 1/angbracketright+/angbracketleftTv,w 2/angbracketright
=/angbracketleftv,T∗w1/angbracketright+/angbracketleftv,T∗w2/angbracketright
=/angbracketleftv,T∗w1+T∗w2/angbracketright,
which shows that T∗w1+T∗w2plays the role required of T∗(w1+w2).
Because only one vector can behave that way, we must have
T∗w1+T∗w2=T∗(w1+w2).
Now let’s check the homogeneity of T∗.I fa∈F, then
/angbracketleftTv,aw/angbracketright=¯a/angbracketleftTv,w/angbracketright
=¯a/angbracketleftv,T∗w/angbracketright
=/angbracketleftv,aT∗w/angbracketright,
which shows that aT∗wplays the role required of T∗(aw). Because
only one vector can behave that way, we must have
aT∗w=T∗(aw).
ThusT∗is a linear map, as claimed.
You should verify that the function T/arrowbarrightT∗has the following prop-
erties:
additivity
(S+T)∗=S∗+T∗for allS,T∈L(V,W);
conjugate homogeneity
(aT)∗=¯aT∗for alla∈FandT∈L(V,W);
adjoint of adjoint
(T∗)∗=Tfor allT∈L(V,W);
identity
I∗=I, whereIis the identity operator on V;
120 Chapter 6.Inner-Product Spaces
products
(ST)∗=T∗S∗for allT∈L(V,W) andS∈L(W,U) (hereUis an
inner-product space over F).
The next result shows the relationship between the null space and
the range of a linear map and its adjoint. The symbol ⇐⇒means “if and
only if”; this symbol could also be read to mean “is equivalent to”.
6.46 Proposition: SupposeT∈L(V,W). Then
(a) nullT∗=(rangeT)⊥;
(b) rangeT∗=(nullT)⊥;
(c) nullT=(rangeT∗)⊥;
(d) rangeT=(nullT∗)⊥.
Proof: Let’s begin by proving (a). Let w∈W. Then
w∈nullT∗⇐⇒T∗w=0
⇐⇒ /angbracketleftv,T∗w/angbracketright=0 for allv∈V
⇐⇒ /angbracketleftTv,w/angbracketright=0 for allv∈V
⇐⇒w∈(rangeT)⊥.
Thus nullT∗=(rangeT)⊥, proving (a).
If we take the orthogonal complement of both sides of (a), we get (d),
where we have used 6.33. Finally, replacing TwithT∗in (a) and (d) gives
(c) and (b).
The conjugate transpose of anm-by-n matrix is the n-by-m matrix IfF=R, then the
conjugate transpose of
a matrix is the same as
itstranspose, which is
the matrix obtained by
interchanging the rows
and columns.obtained by interchanging the rows and columns and then taking the
complex conjugate of each entry. For example, the conjugate transpose
of /bracketleftBigg
23+4i 7
658 i/bracketrightBigg
is the matrix
26
3−4i 5
7−8i
.
The next proposition shows how to compute the matrix of T∗from
the matrix of T. Caution: the proposition below applies only when
Linear Functionals and Adjoints 121
we are dealing with orthonormal bases—with respect to nonorthonor-
mal bases, the matrix of T∗does not necessarily equal the conjugate
transpose of the matrix of T.
6.47 Proposition: SupposeT∈L(V,W).I f(e1,...,en)is an or- The adjoint of a linear
map does not dependon a choice of basis.
This explains why we
will emphasize adjoints
of linear maps instead
of conjugatetransposes of matrices.thonormal basis of Vand(f1,...,fm)is an orthonormal basis of W,
then
M/parenleftbig
T∗,(f1,...,fm),(e 1,...,en)/parenrightbig
is the conjugate transpose of
M/parenleftbig
T,(e 1,...,en),(f 1,...,fm)/parenrightbig
.
Proof: Suppose that (e1,...,en)is an orthonormal basis of Vand
(f1,...,fm)is an orthonormal basis of W. We writeM(T) instead of the
longer expression M/parenleftbig
T,(e 1,...,en),(f 1,...,fm)/parenrightbig
; we also write M(T∗)
instead ofM/parenleftbig
T∗,(f1,...,fm),(e 1,...,en)/parenrightbig
.
Recall that we obtain the kthcolumn ofM(T) by writingTekas a lin-
ear combination of the fj’s; the scalars used in this linear combination
then become the kthcolumn ofM(T). Because (f1,...,fm)is an ortho-
normal basis of W, we know how to write Tekas a linear combination
of thefj’s (see 6.17):
Tek=/angbracketleftTek,f1/angbracketrightf1+···+/angbracketleftTek,fm/angbracketrightfm.
Thus the entry in row j, columnk,o fM(T) is/angbracketleftTek,fj/angbracketright. Replacing T
withT∗and interchanging the roles played by the e’s andf’s, we see
that the entry in row j, columnk,o fM(T∗)is/angbracketleftT∗fk,ej/angbracketright, which equals
/angbracketleftfk,Tej/angbracketright, which equals /angbracketleftTej,fk/angbracketright, which equals the complex conjugate
of the entry in row k, columnj,o fM(T). In other words, M(T∗)equals
the conjugate transpose of M(T).
122 Chapter 6.Inner-Product Spaces
Exercises
1. Prove that if x,y are nonzero vectors in R2, then
/angbracketleftx,y/angbracketright=/bardblx/bardbl/bardbly/bardblcosθ,
whereθis the angle between xandy(thinking of xandyas
arrows with initial point at the origin). Hint: draw the triangle
formed byx,y, andx−y; then use the law of cosines.
2. Suppose u,v∈V. Prove that/angbracketleftu,v/angbracketright=0 if and only if
/bardblu/bardbl≤/bardblu+av/bardbl
for alla∈F.
3. Prove that
/parenleftBign/summationdisplay
j=1ajbj/parenrightBig2
≤/parenleftBign/summationdisplay
j=1jaj2/parenrightBig/parenleftBign/summationdisplay
j=1bj2
j/parenrightBig
for all real numbers a1,...,anandb1,...,bn.
4. Suppose u,v∈Vare such that
/bardblu/bardbl=3,/bardblu+v/bardbl=4,/bardblu−v/bardbl=6.
What number must /bardblv/bardblequal?
5. Prove or disprove: there is an inner product on R2such that the
associated norm is given by
/bardbl(x 1,x2)/bardbl=|x1|+|x2|
for all(x1,x2)∈R2.
6. Prove that if Vis a real inner-product space, then
/angbracketleftu,v/angbracketright=/bardblu+v/bardbl2−/bardblu−v/bardbl2
4
for allu,v∈V.
7. Prove that if Vis a complex inner-product space, then
/angbracketleftu,v/angbracketright=/bardblu+v/bardbl2−/bardblu−v/bardbl2+/bardblu+iv/bardbl2i−/bardblu−iv/bardbl2i
4
for allu,v∈V.
Exercises 123
8. A norm on a vector space Uis a function /bardbl/bardbl:U→[0,∞)such
that/bardblu/bardbl=0 if and only if u=0,/bardblαu/bardbl=|α|/bardblu/bardbl for allα∈F
and allu∈U, and/bardblu+v/bardbl≤/bardblu/bardbl+/bardblv/bardblfor allu,v∈U. Prove
that a norm satisfying the parallelogram equality comes froman inner product (in other words, show that if /bardbl/bardblis a norm
onUsatisfying the parallelogram equality, then there is an inner
product/angbracketleft,/angbracketrightonUsuch that/bardblu/bardbl=/angbracketleftu,u/angbracketright
1/2for allu∈U).
9. Suppose nis a positive integer. Prove that This orthonormal list is
often used for
modeling periodic
phenomena such as
tides./parenleftBig1√
2π,sinx√π,sin 2x√π,...,sinnx√π,cosx√π,cos 2x√π,...,cosnx√π/parenrightBig
is an orthonormal list of vectors in C[−π,π] , the vector space of
continuous real-valued functions on [−π,π] with inner product
/angbracketleftf,g/angbracketright=/integraldisplayπ
−πf(x)g(x)dx.
10. OnP2(R), consider the inner product given by
/angbracketleftp,q/angbracketright=/integraldisplay1
0p(x)q(x)dx.
Apply the Gram-Schmidt procedure to the basis (1,x,x2)to pro-
duce an orthonormal basis of P2(R).
11. What happens if the Gram-Schmidt procedure is applied to a list
of vectors that is not linearly independent?
12. Suppose Vis a real inner-product space and (v1,...,vm)is a
linearly independent list of vectors in V. Prove that there exist
exactly 2morthonormal lists (e1,...,em)of vectors in Vsuch
that
span(v 1,...,vj)=span(e 1,...,ej)
for allj∈{1,...,m}.
13. Suppose (e1,...,em)is an orthonormal list of vectors in V. Let
v∈V. Prove that
/bardblv/bardbl2=|/angbracketleftv,e 1/angbracketright|2+···+|/angbracketleftv,em/angbracketright|2
if and only if v∈span(e 1,...,em).
124 Chapter 6.Inner-Product Spaces
14. Find an orthonormal basis of P2(R)(with inner product as in
Exercise 10) such that the differentiation operator (the operatorthat takesptop
/prime)o nP 2(R)has an upper-triangular matrix with
respect to this basis.
15. Suppose Uis a subspace of V. Prove that
dimU⊥=dimV−dimU.
16. Suppose Uis a subspace of V. Prove thatU⊥={0}if and only if
U=V.
17. Prove that if P∈L(V)is such that P2=Pand every vector
in nullPis orthogonal to every vector in range P, thenPis an
orthogonal projection.
18. Prove that if P∈L(V)is such thatP2=Pand
/bardblPv/bardbl≤/bardblv/bardbl
for everyv∈V, thenPis an orthogonal projection.
19. Suppose T∈L(V)andUis a subspace of V. Prove that Uis
invariant under Tif and only if PUTPU=TPU.
20. Suppose T∈L(V)andUis a subspace of V. Prove that Uand
U⊥are both invariant under Tif and only if PUT=TPU.
21. In R4, let
U=span/parenleftbig
(1,1,0,0),(1, 1,1,2)/parenrightbig
.
Findu∈Usuch that/bardblu−(1,2,3,4)/bardblis as small as possible.
22. Findp∈P 3(R)such thatp(0)=0,p/prime(0)=0, and
/integraldisplay1
0|2+3x−p(x)|2dx
is as small as possible.
23. Findp∈P 5(R)that makes
/integraldisplayπ
−π|sinx−p(x)|2dx
as small as possible. (The polynomial 6.40 is an excellent approx-
imation to the answer to this exercise, but here you are asked tofind the exact solution, which involves powers of π. A computer
that can perform symbolic integration will be useful.)
Exercises 125
24. Find a polynomial q∈P 2(R)such that
p(1
2)=/integraldisplay1
0p(x)q(x)dx
for everyp∈P 2(R).
25. Find a polynomial q∈P 2(R)such that
/integraldisplay1
0p(x)( cosπx)dx=/integraldisplay1
0p(x)q(x)dx
for everyp∈P 2(R).
26. Fix a vector v∈Vand defineT∈L(V,F)byTu=/angbracketleftu,v/angbracketright. For
a∈F, find a formula for T∗a.
27. Suppose nis a positive integer. Define T∈L(Fn)by
T(z 1,...,zn)=(0,z 1,...,zn−1).
Find a formula for T∗(z1,...,zn).
28. Suppose T∈L(V)andλ∈F. Prove thatλis an eigenvalue of T
if and only if ¯λis an eigenvalue of T∗.
29. Suppose T∈L(V)andUis a subspace of V. Prove that Uis
invariant under Tif and only if U⊥is invariant under T∗.
30. Suppose T∈L(V,W). Prove that
(a)Tis injective if and only if T∗is surjective;
(b)Tis surjective if and only if T∗is injective.
31. Prove that
dim nullT∗=dim nullT+dimW−dimV
and
dim rangeT∗=dim rangeT
for everyT∈L(V,W).
32. Suppose Ais anm-by-n matrix of real numbers. Prove that the
dimension of the span of the columns of A(inRm) equals the
dimension of the span of the rows of A(inRn).
Chapter 7
Operators on
Inner-Product Spaces
The deepest results related to inner-product spaces deal with the
subject to which we now turn—operators on inner-product spaces. By
exploiting properties of the adjoint, we will develop a detailed descrip-tion of several important classes of operators on inner-product spaces.
Recall that Fdenotes RorC.
Let’s agree that for this chapter
Vis a finite-dimensional, nonzero, inner-product space over F.
✽✽✽✽✽✽✽
127
128 Chapter 7.Operators on Inner-Product Spaces
Self-Adjoint and Normal Operators
An operator T∈L(V)is called self-adjoint ifT=T∗. For example, Instead of self-adjoint,
some mathematicians
use the term Hermitian
(in honor of the French
mathematician Charles
Hermite, who in 1873
published the first
proof thateis not the
root of any polynomial
with integer
coefficients).ifTis the operator on F2whose matrix (with respect to the standard
basis) is/bracketleftBigg
2b
37/bracketrightBigg
,
thenTis self-adjoint if and only if b=3 (becauseM(T)=M(T∗)if and
only ifb=3; recall that M(T∗)is the conjugate transpose of M(T)—
see 6.47).
You should verify that the sum of two self-adjoint operators is self-
adjoint and that the product of a real scalar and a self-adjoint operatoris self-adjoint.
A good analogy to keep in mind (especially when F=C) is that
the adjoint on L(V) plays a role similar to complex conjugation on C.
A complex number zis real if and only if z=¯z; thus a self-adjoint
operator (T =T
∗) is analogous to a real number. We will see that
this analogy is reflected in some important properties of self-adjointoperators, beginning with eigenvalues.
7.1 Proposition: Every eigenvalue of a self-adjoint operator is real.
IfF=R, then by
definition every
eigenvalue is real, so
this proposition is
interesting only when
F=C.Proof: SupposeTis a self-adjoint operator on V. Letλbe an
eigenvalue of T, and letvbe a nonzero vector in Vsuch thatTv=λv.
Then
λ/bardblv/bardbl2=/angbracketleftλv,v/angbracketright
=/angbracketleftTv,v/angbracketright
=/angbracketleftv,Tv/angbracketright
=/angbracketleftv,λv/angbracketright
=¯λ/bardblv/bardbl2.
Thusλ=¯λ, which means that λis real, as desired.
The next proposition is false for real inner-product spaces. As an
example, consider the operator T∈L(R2)that is a counterclockwise
rotation of 90◦around the origin; thus T(x,y)=(−y,x) . Obviously
Tvis orthogonal to vfor everyv∈R2, even though Tis not 0.
Self-Adjoint and Normal Operators 129
7.2 Proposition: IfVis a complex inner-product space and Tis an
operator on Vsuch that
/angbracketleftTv,v/angbracketright=0
for allv∈V, thenT=0.
Proof: SupposeVis a complex inner-product space and T∈L(V).
Then
/angbracketleftTu,w/angbracketright=/angbracketleftT(u+w),u+w/angbracketright−/angbracketleftT(u−w),u−w/angbracketright
4
+/angbracketleftT(u+iw),u+iw/angbracketright−/angbracketleftT(u−iw),u−iw/angbracketright
4i
for allu,w∈V, as can be verified by computing the right side. Note
that each term on the right side is of the form /angbracketleftTv,v/angbracketrightfor appropriate
v∈V.I f/angbracketleftTv,v/angbracketright=0 for allv∈V, then the equation above implies that
/angbracketleftTu,w/angbracketright=0 for allu,w∈V. This implies that T=0 (takew=Tu).
The following corollary is false for real inner-product spaces, as
shown by considering any operator on a real inner-product space thatis not self-adjoint.
7.3 Corollary: LetVbe a complex inner-product space and let
This corollary provides
another example ofhow self-adjoint
operators behave like
real numbers.T∈L(V). ThenTis self-adjoint if and only if
/angbracketleftTv,v/angbracketright∈R
for everyv∈V.
Proof: Letv∈V. Then
/angbracketleftTv,v/angbracketright−/angbracketleftTv,v/angbracketright=/angbracketleftTv,v/angbracketright−/angbracketleftv,Tv/angbracketright
=/angbracketleftTv,v/angbracketright−/angbracketleftT∗v,v/angbracketright
=/angbracketleft(T−T∗)v,v/angbracketright.
If/angbracketleftTv,v/angbracketright∈R for everyv∈V, then the left side of the equation above
equals 0, so /angbracketleft(T−T∗)v,v/angbracketright=0 for every v∈V. This implies that
T−T∗=0 (by 7.2), and hence Tis self-adjoint.
Conversely, if Tis self-adjoint, then the right side of the equation
above equals 0, so /angbracketleftTv,v/angbracketright=/angbracketleftTv,v/angbracketrightfor everyv∈V. This implies that
/angbracketleftTv,v/angbracketright∈R for everyv∈V, as desired.
130 Chapter 7.Operators on Inner-Product Spaces
On a real inner-product space V, a nonzero operator Tmay satisfy
/angbracketleftTv,v/angbracketright=0 for all v∈V. However, the next proposition shows that
this cannot happen for a self-adjoint operator.
7.4 Proposition: IfTis a self-adjoint operator on Vsuch that
/angbracketleftTv,v/angbracketright=0
for allv∈V, thenT=0.
Proof: We have already proved this (without the hypothesis that
Tis self-adjoint) when Vis a complex inner-product space (see 7.2).
Thus we can assume that Vis a real inner-product space and that Tis
a self-adjoint operator on V. Foru,w∈V, we have
7.5/angbracketleftTu,w/angbracketright=/angbracketleftT(u+w),u+w/angbracketright−/angbracketleftT(u−w),u−w/angbracketright
4;
this is proved by computing the right side, using
/angbracketleftTw,u/angbracketright=/angbracketleftw,Tu/angbracketright
=/angbracketleftTu,w/angbracketright,
where the first equality holds because Tis self-adjoint and the second
equality holds because we are working on a real inner-product space.If/angbracketleftTv,v/angbracketright=0 for all v∈V, then 7.5 implies that /angbracketleftTu,w/angbracketright=0 for all
u,w∈V. This implies that T=0 (takew=Tu).
An operator on an inner-product space is called normal if it com-
mutes with its adjoint; in other words, T∈L(V)is normal if
TT∗=T∗T.
Obviously every self-adjoint operator is normal. For an example of a
normal operator that is not self-adjoint, consider the operator on F2
whose matrix (with respect to the standard basis) is
/bracketleftBigg
2−3
32/bracketrightBigg
.
Clearly this operator is not self-adjoint, but an easy calculation (which
you should do) shows that it is normal.
We will soon see why normal operators are worthy of special at-
tention. The next proposition provides a simple characterization ofnormal operators.
Self-Adjoint and Normal Operators 131
7.6 Proposition: An operator T∈L(V) is normal if and only if Note that this
proposition implies
that nullT=nullT∗
for every normal
operatorT./bardblTv/bardbl=/bardblT∗v/bardbl
for allv∈V.
Proof: LetT∈L(V). We will prove both directions of this result
at the same time. Note that
Tis normal⇐⇒T∗T−TT∗=0
⇐⇒ /angbracketleft(T∗T−TT∗)v,v/angbracketright=0 for all v∈V
⇐⇒ /angbracketleftT∗Tv,v/angbracketright=/angbracketleftTT∗v,v/angbracketright for allv∈V
⇐⇒ /bardblTv/bardbl2=/bardblT∗v/bardbl2for allv∈V,
where we used 7.4 to establish the second equivalence (note that the
operatorT∗T−TT∗is self-adjoint). The equivalence of the first and
last conditions above gives the desired result.
Compare the next corollary to Exercise 28 in the previous chapter.
That exercise implies that the eigenvalues of the adjoint of any operatorare equal (as a set) to the complex conjugates of the eigenvalues of theoperator. The exercise says nothing about eigenvectors because anoperator and its adjoint may have different eigenvectors. However, thenext corollary implies that a normal operator and its adjoint have the
same eigenvectors.
7.7 Corollary: SupposeT∈L(V) is normal. If v∈Vis an eigen-
vector ofTwith eigenvalue λ∈F, thenvis also an eigenvector of T
∗
with eigenvalue ¯λ.
Proof: Supposev∈Vis an eigenvector of Twith eigenvalue λ.
Thus(T−λI)v=0. BecauseTis normal, so is T−λI, as you should
verify. Using 7.6, we have
0=/bardbl(T−λI)v/bardbl=/bardbl(T−λI)∗v/bardbl=/bardbl(T∗−¯λI)v/bardbl,
and hencevis an eigenvector of T∗with eigenvalue ¯λ, as desired.
Because every self-adjoint operator is normal, the next result applies
in particular to self-adjoint operators.
132 Chapter 7.Operators on Inner-Product Spaces
7.8 Corollary: IfT∈L(V) is normal, then eigenvectors of T
corresponding to distinct eigenvalues are orthogonal.
Proof: SupposeT∈L(V)is normal and α,β are distinct eigen-
values ofT, with corresponding eigenvectors u,v. ThusTu=αuand
Tv=βv. From 7.7 we have T∗v=¯βv. Thus
(α−β)/angbracketleftu,v/angbracketright=/angbracketleftαu,v/angbracketright−/angbracketleftu,¯βv/angbracketright
=/angbracketleftTu,v/angbracketright−/angbracketleftu,T∗v/angbracketright
=0.
Becauseα/negationslash=β, the equation above implies that /angbracketleftu,v/angbracketright=0. Thusuand
vare orthogonal, as desired.
The Spectral Theorem
Recall that a diagonal matrix is a square matrix that is 0 everywhere
except possibly along the diagonal. Recall also that an operator on V
has a diagonal matrix with respect to some basis if and only if there isa basis ofVconsisting of eigenvectors of the operator (see 5.21).
The nicest operators on Vare those for which there is an ortho-
normal basis ofVwith respect to which the operator has a diagonal
matrix. These are precisely the operators T∈L(V)such that there is
an orthonormal basis of Vconsisting of eigenvectors of T. Our goal
in this section is to prove the spectral theorem, which characterizesthese operators as the normal operators when F=Cand as the self-
adjoint operators when F=R. The spectral theorem is probably the
most useful tool in the study of operators on inner-product spaces.
Because the conclusion of the spectral theorem depends on F,w e
will break the spectral theorem into two pieces, called the complexspectral theorem and the real spectral theorem. As is often the case inlinear algebra, complex vector spaces are easier to deal with than realvector spaces, so we present the complex spectral theorem first.
As an illustration of the complex spectral theorem, consider the
normal operator T∈L(C
2)whose matrix (with respect to the standard
basis) is/bracketleftBigg
2−3
32/bracketrightBigg
.
You should verify that
The Spectral Theorem 133
/parenleftbigg(i,1)√
2,(−i,1)√
2/parenrightbigg
is an orthonormal basis of C2consisting of eigenvectors of Tand that
with respect to this basis, the matrix of Tis the diagonal matrix
/bracketleftBigg
2+3i 0
02−3i/bracketrightBigg
.
7.9 Complex Spectral Theorem: Suppose that Vis a complex Because every
self-adjoint operator is
normal, the complex
spectral theorem
implies that every
self-adjoint operator on
a finite-dimensionalcomplex inner-product
space has a diagonal
matrix with respect to
some orthonormal
basis.inner-product space and T∈L(V). ThenVhas an orthonormal basis
consisting of eigenvectors of Tif and only if Tis normal.
Proof: First suppose that Vhas an orthonormal basis consisting of
eigenvectors of T. With respect to this basis, Thas a diagonal matrix.
The matrix of T∗(with respect to the same basis) is obtained by taking
the conjugate transpose of the matrix of T; henceT∗also has a diag-
onal matrix. Any two diagonal matrices commute; thus Tcommutes
withT∗, which means that Tmust be normal, as desired.
To prove the other direction, now suppose that Tis normal. There
is an orthonormal basis (e1,...,en)ofVwith respect to which Thas
an upper-triangular matrix (by 6.28). Thus we can write
7.10 M/parenleftbig
T,(e 1,...,en)/parenrightbig
=
a1,1... a 1,n
......
0an,n
.
We will show that this matrix is actually a diagonal matrix, which means
that(e1,...,en)is an orthonormal basis of Vconsisting of eigenvectors
ofT.
We see from the matrix above that
/bardblTe 1/bardbl2=|a1,1|2
and
/bardblT∗e1/bardbl2=|a1,1|2+|a1,2|2+···+|a1,n|2.
BecauseTis normal,/bardblTe 1/bardbl=/bardblT∗e1/bardbl(see 7.6). Thus the two equations
above imply that all entries in the first row of the matrix in 7.10, exceptpossibly the first entry a
1,1, equal 0.
Now from 7.10 we see that
/bardblTe 2/bardbl2=|a2,2|2
134 Chapter 7.Operators on Inner-Product Spaces
(becausea1,2=0, as we showed in the paragraph above) and
/bardblT∗e2/bardbl2=|a2,2|2+|a2,3|2+···+|a2,n|2.
BecauseTis normal,/bardblTe 2/bardbl=/bardblT∗e2/bardbl. Thus the two equations above
imply that all entries in the second row of the matrix in 7.10, except
possibly the diagonal entry a2,2, equal 0.
Continuing in this fashion, we see that all the nondiagonal entries
in the matrix 7.10 equal 0, as desired.
We will need two lemmas for our proof of the real spectral theo-
rem. You could guess that the next lemma is true and even discover its
proof by thinking about quadratic polynomials with real coefficients.
Specifically, suppose α,β∈Randα2<4β. Letxbe a real number.
Then This technique of
completing the square
can be used to derive
the quadratic formula.x2+αx+β=/parenleftbig
x+α
2/parenrightbig2+/parenleftbig
β−α2
4/parenrightbig
>0.
In particular, x2+αx+βis an invertible real number (a convoluted
way of saying that it is not 0). Replacing the real number xwith a
self-adjoint operator (recall the analogy between real numbers and self-adjoint operators), we are led to the lemma below.
7.11 Lemma: SupposeT∈L(V) is self-adjoint. If α,β∈Rare such
thatα
2<4β, then
T2+αT+βI
is invertible.
Proof: Supposeα,β∈Rare such that α2<4β. Letvbe a nonzero
vector inV. Then
/angbracketleft(T2+αT+βI)v,v/angbracketright=/angbracketleftT2v,v/angbracketright+α/angbracketleftTv,v /angbracketright+β/angbracketleftv,v/angbracketright
=/angbracketleftTv,Tv/angbracketright+α/angbracketleftTv,v/angbracketright+β/bardblv/bardbl2
≥/bardblTv/bardbl2−|α|/bardblTv/bardbl/bardblv/bardbl+β/bardblv/bardbl2
=/parenleftbig
/bardblTv/bardbl−|α|/bardblv/bardbl
2/parenrightbig2+/parenleftbig
β−α2
4/parenrightbig
/bardblv/bardbl2
>0,
The Spectral Theorem 135
where the first inequality holds by the Cauchy-Schwarz inequality (6.6).
The last inequality implies that (T2+αT+βI)v/negationslash=0. ThusT2+αT+βI
is injective, which implies that it is invertible (see 3.21).
We have proved that every operator, self-adjoint or not, on a finite-
dimensional complex vector space has an eigenvalue (see 5.10), so thenext lemma tells us something new only for real inner-product spaces.
7.12 Lemma: SupposeT∈L(V) is self-adjoint. Then Thas an
eigenvalue.
Proof: As noted above, we can assume that Vis a real inner-
product space. Let n=dimVand choosev∈Vwithv/negationslash=0. Then
Here we are imitating
the proof that Thas an
invariant subspace of
dimension 1or2
(see 5.24).(v,Tv,T2v,...,Tnv)
cannot be linearly independent because Vhas dimension nand we have
n+1 vectors. Thus there exist real numbers a0,...,an, not all 0, such
that
0=a0v+a1Tv+···+a nTnv.
Make thea’s the coefficients of a polynomial, which can be written in
factored form (see 4.14) as
a0+a1x+···+anxn
=c(x2+α1x+β1)...(x2+αMx+βM)(x−λ1)...(x−λm),
wherecis a nonzero real number, each αj,βj, andλjis real, each
αj2<4βj,m+M≥1, and the equation holds for all real x. We then
have
0=a0v+a1Tv+···+anTnv
=(a0I+a1T+···+anTn)v
=c(T2+α1T+β1I)...(T2+αMT+βMI)(T−λ1I)...(T−λmI)v.
EachT2+αjT+βjIis invertible because Tis self-adjoint and each
αj2<4βj(see 7.11). Recall also that c/negationslash=0. Thus the equation above
implies that
0=(T−λ1I)...(T−λmI)v.
HenceT−λjIis not injective for at least one j. In other words, Thas
an eigenvalue.
136 Chapter 7.Operators on Inner-Product Spaces
As an illustration of the real spectral theorem, consider the self-
adjoint operator TonR3whose matrix (with respect to the standard
basis) is
14−13 8
−13 14 8
88−7
.
You should verify that
/parenleftbigg(1,−1,0)√
2,(1,1,1)√
3,(1,1,−2)√
6/parenrightbigg
is an orthonormal basis of R3consisting of eigenvectors of Tand that
with respect to this basis, the matrix of Tis the diagonal matrix
27 0 0
09 000−15
.
Combining the complex spectral theorem and the real spectral the-
orem, we conclude that every self-adjoint operator on Vhas a diagonal
matrix with respect to some orthonormal basis. This statement, which
is the most useful part of the spectral theorem, holds regardless of
whether F=CorF=R.
7.13 Real Spectral Theorem: Suppose that Vis a real inner-product
space andT∈L(V). ThenVhas an orthonormal basis consisting of
eigenvectors of Tif and only if Tis self-adjoint.
Proof: First suppose that Vhas an orthonormal basis consisting of
eigenvectors of T. With respect to this basis, Thas a diagonal matrix.
This matrix equals its conjugate transpose. Hence T=T
∗and soTis
self-adjoint, as desired.
To prove the other direction, now suppose that Tis self-adjoint. We
will prove that Vhas an orthonormal basis consisting of eigenvectors
ofTby induction on the dimension of V. To get started, note that our
desired result clearly holds if dim V=1. Now assume that dim V>1
and that the desired result holds on vector spaces of smaller dimen-
sion.
The idea of the proof is to take any eigenvector uofTwith norm 1,
then adjoin to it an orthonormal basis of eigenvectors of T|{u}⊥. Now
The Spectral Theorem 137
for the details, the most important of which is verifying that T|{u}⊥is
self-adjoint (this allows us to apply our induction hypothesis).
Letλbe any eigenvalue of T(becauseTis self-adjoint, we know
from the previous lemma that it has an eigenvalue) and let u∈V
denote a corresponding eigenvector with /bardblu/bardbl=1. Let Udenote the To get an eigenvector
of norm 1, take any
nonzero eigenvector
and divide it by its
norm.one-dimensional subspace of Vconsisting of all scalar multiples of u.
Note that a vector v∈Vis inU⊥if and only if /angbracketleftu,v/angbracketright=0.
Supposev∈U⊥. Then because Tis self-adjoint, we have
/angbracketleftu,Tv/angbracketright=/angbracketleftTu,v/angbracketright=/angbracketleftλu,v/angbracketright=λ/angbracketleftu,v/angbracketright=0,
and henceTv∈U⊥. ThusTv∈U⊥wheneverv∈U⊥. In other words,
U⊥is invariant under T. Thus we can define an operator S∈L(U⊥)by
S=T|U⊥.I fv,w∈U⊥, then
/angbracketleftSv,w/angbracketright=/angbracketleftTv,w/angbracketright=/angbracketleftv,Tw/angbracketright=/angbracketleftv,Sw/angbracketright,
which shows that Sis self-adjoint (note that in the middle equality
above we used the self-adjointness of T). Thus, by our induction hy-
pothesis, there is an orthonormal basis of U⊥consisting of eigenvec-
tors ofS. Clearly every eigenvector of Sis an eigenvector of T(because
Sv=Tvfor everyv∈U⊥). Thus adjoining uto an orthonormal basis
ofU⊥consisting of eigenvectors of Sgives an orthonormal basis of V
consisting of eigenvectors of T, as desired.
ForT∈L(V)self-adjoint (or, more generally, T∈L(V)normal
when F=C), the corollary below provides the nicest possible decom-
position ofVinto subspaces invariant under T. On each null (T−λjI),
the operator Tis just multiplication by λj.
7.14 Corollary: Suppose that T∈L(V) is self-adjoint (or that F=C
and thatT∈L(V) is normal). Let λ1,...,λmdenote the distinct eigen-
values ofT. Then
V=null(T−λ1I)⊕···⊕ null(T−λmI).
Furthermore, each vector in each null(T−λjI)is orthogonal to all vec-
tors in the other subspaces of this decomposition.
Proof: The spectral theorem (7.9 and 7.13) implies that Vhas a
basis consisting of eigenvectors of T. The desired decomposition of V
now follows from 5.21.
The orthogonality statement follows from 7.8.
138 Chapter 7.Operators on Inner-Product Spaces
Normal Operators on Real
Inner-Product Spaces
The complex spectral theorem (7.9) gives a complete description
of normal operators on complex inner-product spaces. In this sectionwe will give a complete description of normal operators on real inner-
product spaces. Along the way, we will encounter a proposition (7.18)and a technique (block diagonal matrices) that are useful for both realand complex inner-product spaces.
We begin with a description of the operators on a two-dimensional
real inner-product space that are normal but not self-adjoint.
7.15 Lemma: SupposeVis a two-dimensional real inner-product
space andT∈L(V). Then the following are equivalent:
(a)Tis normal but not self-adjoint;
(b) the matrix of Twith respect to every orthonormal basis of V
has the form /bracketleftBigg
a−b
ba/bracketrightBigg
,
withb/negationslash=0;
(c) the matrix of Twith respect to some orthonormal basis of Vhas
the form /bracketleftBigg
a−b
ba/bracketrightBigg
,
withb>0.
Proof: First suppose that (a) holds, so that Tis normal but not
self-adjoint. Let (e
1,e2)be an orthonormal basis of V. Suppose
7.16 M/parenleftbig
T,(e 1,e2)/parenrightbig
=/bracketleftBigg
ac
bd/bracketrightBigg
.
Then/bardblTe 1/bardbl2=a2+b2and/bardblT∗e1/bardbl2=a2+c2. BecauseTis normal,
/bardblTe 1/bardbl=/bardblT∗e1/bardbl(see 7.6); thus these equations imply that b2=c2.
Thusc=borc=−b. Butc/negationslash=bbecause otherwise Twould be self-
adjoint, as can be seen from the matrix in 7.16. Hence c=−b,s o
7.17 M/parenleftbig
T,(e 1,e2)/parenrightbig
=/bracketleftBigg
a−b
bd/bracketrightBigg
.
Normal Operators on Real Inner-Product Spaces 139
Of course, the matrix of T∗is the transpose of the matrix above. Use
matrix multiplication to compute the matrices of TT∗andT∗T(do it
now). Because Tis normal, these two matrices must be equal. Equating
the entries in the upper-right corner of the two matrices you computed,you will discover that bd=ab. Nowb/negationslash=0 because otherwise Twould
be self-adjoint, as can be seen from the matrix in 7.17. Thus d=a,
completing the proof that (a) implies (b).
Now suppose that (b) holds. We want to prove that (c) holds. Choose
any orthonormal basis (e
1,e2)ofV. We know that the matrix of Twith
respect to this basis has the form given by (b), with b/negationslash=0. Ifb> 0,
then (c) holds and we have proved that (b) implies (c). If b<0, then,
as you should verify, the matrix of Twith respect to the orthonormal
basis(e1,−e2)equals/bracketleftBig
ab
−ba/bracketrightBig
, where−b> 0; thus in this case we also
see that (b) implies (c).
Now suppose that (c) holds, so that the matrix of Twith respect to
some orthonormal basis has the form given in (c) with b>0. Clearly
the matrix of Tis not equal to its transpose (because b/negationslash=0), and hence
Tis not self-adjoint. Now use matrix multiplication to verify that the
matrices ofTT∗andT∗Tare equal. We conclude that TT∗=T∗T, and
henceTis normal. Thus (c) implies (a), completing the proof.
As an example of the notation we will use to write a matrix as a
matrix of smaller matrices, consider the matrix
D=
11222
11222003330033300333
.
We can write this matrix in the form
Often we can
understand a matrix
better by thinking of it
as composed of smallermatrices. We will use
this technique in the
next proposition and in
later chapters.D=/bracketleftBigg
AB
0C/bracketrightBigg
,
where
A=/bracketleftBigg
11
11/bracketrightBigg
,B=/bracketleftBigg
222
222/bracketrightBigg
,C=
333
333333
,
and 0 denotes the 3-by-2 matrix consisting of all 0’s.
140 Chapter 7.Operators on Inner-Product Spaces
The next result will play a key role in our characterization of the
normal operators on a real inner-product space.
7.18 Proposition: SupposeT∈L(V) is normal and Uis a subspace Without normality, an
easier result also holds:
ifT∈L(V) andU
invariant under T, then
U⊥is invariant under
T∗; see Exercise 29 in
Chapter 6.ofVthat is invariant under T. Then
(a)U⊥is invariant under T;
(b)Uis invariant under T∗;
(c)(T|U)∗=(T∗)|U;
(d)T|Uis a normal operator on U;
(e)T|U⊥is a normal operator on U⊥.
Proof: First we will prove (a). Let (e1,...,em)be an orthonormal
basis ofU. Extend to an orthonormal basis (e1,...,em,f1,...,fn)ofV
(this is possible by 6.25). Because Uis invariant under T, eachTejis
a linear combination of (e1,...,em). Thus the matrix of Twith respect
to the basis(e1,...,em,f1,...,fn)is of the form
e1... emf1... fn
M(T)=e1
...
em
f1
...
fn
AB
0C
;
hereAdenotes anm-by-m matrix, 0 denotes the n-by-m matrix con-
sisting of all 0’s, Bdenotes anm-by-n matrix,Cdenotes ann-by-n
matrix, and for convenience the basis has been listed along the top andleft sides of the matrix.
For eachj∈{1,...,m},/bardblTe
j/bardbl2equals the sum of the squares of the
absolute values of the entries in the jthcolumn ofA(see 6.17). Hence
7.19m/summationdisplay
j=1/bardblTej/bardbl2=the sum of the squares of the absolute
values of the entries of A.
For eachj∈{1,...,m},/bardblT∗ej/bardbl2equals the sum of the squares of the
absolute values of the entries in the jthrows ofAandB. Hence
Normal Operators on Real Inner-Product Spaces 141
7.20m/summationdisplay
j=1/bardblT∗ej/bardbl2=the sum of the squares of the absolute
values of the entries of AandB.
BecauseTis normal,/bardblTej/bardbl=/bardblT∗ej/bardblfor eachj(see 7.6); thus
m/summationdisplay
j=1/bardblTej/bardbl2=m/summationdisplay
j=1/bardblT∗ej/bardbl2.
This equation, along with 7.19 and 7.20, implies that the sum of the
squares of the absolute values of the entries of Bmust equal 0. In
other words, Bmust be the matrix consisting of all 0’s. Thus
e1... emf1... fn
M(T)=e1
...
em
f1
...
fn
A 0
0C
. 7.21
This representation shows that Tf
kis in the span of (f1,...,fn)for
eachk. Because(f1,...,fn)is a basis ofU⊥, this implies that Tv∈U⊥
wheneverv∈U⊥. In other words, U⊥is invariant under T, completing
the proof of (a).
To prove (b), note that M(T∗)has a block of 0’s in the lower left
corner (because M(T), as given above, has a block of 0’s in the upper
right corner). In other words, each T∗ejcan be written as a linear
combination of (e1,...,em). ThusUis invariant under T∗, completing
the proof of (b).
To prove (c), let S=T|U. Fixv∈U. Then
/angbracketleftSu,v/angbracketright=/angbracketleftTu,v/angbracketright
=/angbracketleftu,T∗v/angbracketright
for allu∈U. BecauseT∗v∈U(by (b)), the equation above shows that
S∗v=T∗v. In other words, (T|U)∗=(T∗)|U, completing the proof
of (c).
To prove (d), note that Tcommutes with T∗(becauseTis normal)
and that(T|U)∗=(T∗)|U(by (c)). Thus T|Ucommutes with its adjoint
and hence is normal, completing the proof of (d).
142 Chapter 7.Operators on Inner-Product Spaces
To prove (e), note that in (d) we showed that the restriction of Tto
any invariant subspace is normal. However, U⊥is invariant under T
(by (a)), and hence T|U⊥is normal.
In proving 7.18 we thought of a matrix as composed of smaller ma-
trices. Now we need to make additional use of that idea. A block diag-
onal matrix is a square matrix of the form The key step in the
proof of the last
proposition was
showing that M(T) is
an appropriate block
diagonal matrix;
see 7.21.
A1 0
...
0Am
,
whereA1,...,Amare square matrices lying along the diagonal and all
the other entries of the matrix equal 0. For example, the matrix
7.22 A=
40 0 0 0
02−30 0
03 2 0 000 0 1−7
00 0 7 1
is a block diagonal matrix with
A=
A
1 0
A2
0A3
,
where
7.23A1=/bracketleftBig
4/bracketrightBig
,A 2=/bracketleftBigg
2−3
32/bracketrightBigg
,A 3=/bracketleftBigg
1−7
71/bracketrightBigg
.
IfAandBare block diagonal matrices of the form
A=
A1 0
...
0Am
,B=
B1 0
...
0Bm
,
whereAjhas the same size as Bjforj=1,...,m , thenABis a block
diagonal matrix of the form
7.24 AB=
A1B1 0
...
0AmBm
,
Normal Operators on Real Inner-Product Spaces 143
as you should verify. In other words, to multiply together two block
diagonal matrices (with the same size blocks), just multiply together thecorresponding entries on the diagonal, as with diagonal matrices.
A diagonal matrix is a special case of a block diagonal matrix where
each block has size 1-by-1. At the other extreme, every square matrix is
Note that if an operator
Thas a block diagonal
matrix with respect to
some basis, then the
entry in any 1-by-1
block on the diagonal
of this matrix must bean eigenvalue of T.a block diagonal matrix because we can take the first (and only) block
to be the entire matrix. Thus to say that an operator has a block di-agonal matrix with respect to some basis tells us nothing unless weknow something about the size of the blocks. The smaller the blocks,
the nicer the operator (in the vague sense that the matrix then containsmore 0’s). The nicest situation is to have an orthonormal basis that
gives a diagonal matrix. We have shown that this happens on a com-
plex inner-product space precisely for the normal operators (see 7.9)and on a real inner-product space precisely for the self-adjoint opera-tors (see 7.13).
Our next result states that each normal operator on a real inner-
product space comes close to having a diagonal matrix—specifically,we get a block diagonal matrix with respect to some orthonormal basis,with each block having size at most 2-by-2. We cannot expect to do bet-ter than that because on a real inner-product space there exist normal
operators that do not have a diagonal matrix with respect to any basis.
For example, the operator T∈L(R
2)defined byT(x,y)=(−y,x) is
normal (as you should verify) but has no eigenvalues; thus this partic-ularTdoes not have even an upper-triangular matrix with respect to
any basis of R
2.
Note that the matrix in 7.22 is the type of matrix promised by the
theorem below. In particular, each block of 7.22 (see 7.23) has sizeat most 2-by-2 and each of the 2-by-2 blocks has the required form(upper left entry equals lower right entry, lower left entry is positive,
and upper right entry equals the negative of lower left entry).
7.25 Theorem: Suppose that Vis a real inner-product space and
T∈L(V). ThenTis normal if and only if there is an orthonormal
basis ofVwith respect to which Thas a block diagonal matrix where
each block is a 1-by-1 matrix or a 2-by-2 matrix of the form
7.26/bracketleftBigg
a−b
ba/bracketrightBigg
,
withb>0.
144 Chapter 7.Operators on Inner-Product Spaces
Proof: To prove the easy direction, first suppose that there is an
orthonormal basis of Vsuch that the matrix of Tis a block diagonal
matrix where each block is a 1-by-1 matrix or a 2-by-2 matrix of the
form 7.26. With respect to this basis, the matrix of Tcommutes with
the matrix of T∗(which is the conjugate of the matrix of T), as you
should verify (use formula 7.24 for the product of two block diagonalmatrices). Thus Tcommutes with T
∗, which means that Tis normal.
To prove the other direction, now suppose that Tis normal. We will
prove our desired result by induction on the dimension of V. To get
started, note that our desired result clearly holds if dim V=1 (trivially)
or if dimV=2 (ifTis self-adjoint, use the real spectral theorem 7.13;
ifTis not self-adjoint, use 7.15).
Now assume that dim V> 2 and that the desired result holds on
vector spaces of smaller dimension. Let Ube a subspace of Vof di-
mension 1 that is invariant under Tif such a subspace exists (in other
words, ifThas a nonzero eigenvector, let Ube the span of this eigen-
vector). If no such subspace exists, let Ube a subspace of Vof dimen-
sion 2 that is invariant under T(an invariant subspace of dimension 1
or 2 always exists by 5.24).
If dimU=1, choose a vector in Uwith norm 1; this vector will In a real vector space
with dimension 1, there
are precisely two
vectors with norm 1.be an orthonormal basis of U, and of course the matrix of T|Uis a
1-by-1 matrix. If dim U=2, thenT|Uis normal (by 7.18) but not self-
adjoint (otherwise T|U, and henceT, would have a nonzero eigenvector;
see 7.12), and thus we can choose an orthonormal basis of Uwith re-
spect to which the matrix of T|Uhas the form 7.26 (see 7.15).
NowU⊥is invariant under TandT|U⊥is a normal operator on U⊥
(see 7.18). Thus by our induction hypothesis, there is an orthonormal
basis ofU⊥with respect to which the matrix of T|U⊥has the desired
form. Adjoining this basis to the basis of Ugives an orthonormal basis
ofVwith respect to which the matrix of Thas the desired form.
Positive Operators
An operator T∈L(V)is called positive ifTis self-adjoint and Many mathematicians
also use the term
positive semidefinite
operator, which means
the same as positive
operator./angbracketleftTv,v/angbracketright≥0
for allv∈V. Note that if Vis a complex vector space, then the
condition that Tbe self-adjoint can be dropped from this definition
(by 7.3).
Positive Operators 145
You should verify that every orthogonal projection is positive. For
another set of examples, look at the proof of 7.11, where we showedthat ifT∈L(V)is self-adjoint and α,β∈Rare such that α
2<4β,
thenT2+αT+βIis positive.
An operator Sis called a square root of an operator TifS2=T.
For example, if T∈L(F3)is defined by T(z 1,z2,z3)=(z3,0,0), then
the operator S∈L(F3)defined byS(z 1,z2,z3)=(z2,z3,0)is a square
root ofT.
The following theorem is the main result about positive operators. The positive operators
correspond, in somesense, to the numbers
[0,∞), so better
terminology would callthese nonnegative
instead of positive.
However, operator
theorists consistently
call these the positive
operators, so we willfollow that custom.Note that its characterizations of the positive operators correspond to
characterizations of the nonnegative numbers among C. Specifically,
a complex number zis nonnegative if and only if it has a nonnegative
square root, corresponding to condition (c) below. Also, zis nonnega-
tive if and only if it has a real square root, corresponding to condition
(d) below. Finally, zis nonnegative if and only if there exists a complex
numberwsuch thatz=¯ww, corresponding to condition (e) below.
7.27 Theorem: LetT∈L(V). Then the following are equivalent:
(a)Tis positive;
(b)Tis self-adjoint and all the eigenvalues of Tare nonnegative;
(c)Thas a positive square root;
(d)Thas a self-adjoint square root;
(e) there exists an operator S∈L(V) such thatT=S∗S.
Proof: We will prove that (a) ⇒(b)⇒(c)⇒(d)⇒(e)⇒(a).
First suppose that (a) holds, so that Tis positive. Obviously Tis
self-adjoint (by the definition of a positive operator). To prove the othercondition in (b), suppose that λis an eigenvalue of T. Letvbe a nonzero
eigenvector of Tcorresponding to λ. Then
0≤/angbracketleftTv,v/angbracketright
=/angbracketleftλv,v/angbracketright
=λ/angbracketleftv,v/angbracketright,
and thusλis a nonnegative number. Hence (b) holds.
Now suppose that (b) holds, so that Tis self-adjoint and all the eigen-
values ofTare nonnegative. By the spectral theorem (7.9 and 7.13),
146 Chapter 7.Operators on Inner-Product Spaces
there is an orthonormal basis (e1,...,en)ofVconsisting of eigen-
vectors ofT. Letλ1,...,λnbe the eigenvalues of Tcorresponding to
e1,...,en, so that each λjis a nonnegative number. Define S∈L(V)
by
Sej=/radicalBig
λjej
forj=1,...,n . ThenSis a positive operator, as you should verify.
Furthermore, S2ej=λjej=Tejfor eachj, which implies that S2=T.
ThusSis a positive square root of T, and hence (c) holds.
Clearly (c) implies (d) (because, by definition, every positive operator
is self-adjoint).
Now suppose that (d) holds, meaning that there exists a self-adjoint
operatorSonVsuch thatT=S2. ThenT=S∗S(becauseS∗=S), and
hence (e) holds.
Finally, suppose that (e) holds. Let S∈L(V)be such that T=S∗S.
ThenT∗=(S∗S)∗=S∗(S∗)∗=S∗S=T, and henceTis self-adjoint.
To complete the proof that (a) holds, note that
/angbracketleftTv,v/angbracketright=/angbracketleftS∗Sv,v/angbracketright
=/angbracketleftSv,Sv/angbracketright
≥0
for everyv∈V. ThusTis positive.
Each nonnegative number has a unique nonnegative square root.
The next proposition shows that positive operators enjoy a similarproperty. Because of this proposition, we can use the notation√
T
to denote the unique positive square root of a positive operator T, just
as√
λdenotes the unique nonnegative square root of a nonnegative
numberλ.
7.28 Proposition: Every positive operator on Vhas a unique positive A positive operator can
have infinitely many
square roots (though
only one of them can
be positive). For
example, the identity
operator on Vhas
infinitely many square
roots if dimV>1.square root.
Proof: SupposeT∈L(V)is positive. Let λ1,...,λmdenote the
distinct eigenvalues of T; becauseTis positive, all these numbers are
nonnegative (by 7.27). Because Tis self-adjoint, we have
7.29 V=null(T−λ1I)⊕···⊕ null(T−λmI);
see 7.14.
Isometries 147
Now suppose S∈L(V)is a positive square root of T. Supposeαis
an eigenvalue of S.I fv∈null(S−αI), thenSv=αv, which implies
that
7.30 Tv=S2v=α2v,
sov∈null(T−α2I). Thusα2is an eigenvalue of T, which means
thatα2must equal some λj. In other words, α=/radicalBig
λjfor somej.
Furthermore, 7.30 implies that
7.31 null(S−/radicalBig
λjI)⊂null(T−λjI).
In the paragraph above, we showed that the only possible eigenval-
ues forSare/radicalbig
λ1,...,/radicalbig
λm. BecauseSis self-adjoint, this implies that
7.32 V=null(S−/radicalBig
λ1I)⊕···⊕ null(S−/radicalBig
λmI);
see 7.14. Now 7.29, 7.32, and 7.31 imply that
null(S−/radicalBig
λjI)=null(T−λjI)
for eachj. In other words, on null (T−λjI), the operator Sis just
multiplication by/radicalBig
λj. ThusS, the positive square root of T, is uniquely
determined by T.
Isometries
An operator S∈L(V)is called an isometry if The Greek word isos
means equal; the Greekword metron means
measure. Thus
isometry literally
means equal measure./bardblSv/bardbl=/bardblv/bardbl
for allv∈V. In other words, an operator is an isometry if it preserves
norms. For example, λIis an isometry whenever λ∈Fsatisfies|λ|=1.
More generally, suppose λ1,...,λnare scalars with absolute value 1 and
S∈L(V)satisfiesS(ej)=λjejfor some orthonormal basis (e1,...,en)
ofV. Supposev∈V. Then
7.33 v=/angbracketleftv,e 1/angbracketrighte1+···+/angbracketleftv,en/angbracketrighten
and
7.34 /bardblv/bardbl2=|/angbracketleftv,e 1/angbracketright|2+···+|/angbracketleftv,en/angbracketright|2,
148 Chapter 7.Operators on Inner-Product Spaces
where we have used 6.17. Applying Sto both sides of 7.33 gives
Sv=/angbracketleftv,e 1/angbracketrightSe 1+···+/angbracketleftv,en/angbracketrightSen
=λ1/angbracketleftv,e 1/angbracketrighte1+···+λn/angbracketleftv,en/angbracketrighten.
The last equation, along with the equation |λj|=1, shows that
7.35 /bardblSv/bardbl2=|/angbracketleftv,e 1/angbracketright|2+···+|/angbracketleftv,en/angbracketright|2.
Comparing 7.34 and 7.35 shows that /bardblv/bardbl=/bardblSv/bardbl. In other words, Sis
an isometry.
For another example, let θ∈R. Then the operator on R2of coun- An isometry on a real
inner-product space is
often called an
orthogonal operator.
An isometry on a
complex inner-product
space is often called a
unitary operator. We
will use the term
isometry so that our
results can apply to
both real and complex
inner-product spaces.terclockwise rotation (centered at the origin) by an angle of θis an
isometry (you should find the matrix of this operator with respect tothe standard basis of R
2).
IfS∈L(V)is an isometry, then Sis injective (because if Sv=0,
then/bardblv/bardbl=/bardblSv/bardbl=0, and hence v=0). Thus every isometry is
invertible (by 3.21).
The next theorem provides several conditions that are equivalent
to being an isometry. These equivalences have several important in-terpretations. In particular, the equivalence of (a) and (b) shows thatan isometry preserves inner products. Because (a) implies (d), we seethat ifSis an isometry and (e
1,...,en)is an orthonormal basis of V,
then the columns of the matrix of S(with respect to this basis) are or-
thonormal; because (e) implies (a), we see that the converse also holds.Because (a) is equivalent to conditions (i) and (j), we see that in the lastsentence we can replace “columns” with “rows”.
7.36 Theorem: SupposeS∈L(V). Then the following are equiva-
lent:
(a)Sis an isometry;
(b)/angbracketleftSu,Sv/angbracketright=/angbracketleftu,v/angbracketrightfor allu,v∈V;
(c)S
∗S=I;
(d)(Se1,...,Sen)is orthonormal whenever (e1,...,en)is an ortho-
normal list of vectors in V;
(e) there exists an orthonormal basis (e1,...,en)ofVsuch that
(Se1,...,Sen)is orthonormal;
(f)S∗is an isometry;
Isometries 149
(g)/angbracketleftS∗u,S∗v/angbracketright=/angbracketleftu,v/angbracketrightfor allu,v∈V;
(h)SS∗=I;
(i)(S∗e1,...,S∗en)is orthonormal whenever (e1,...,en)is an or-
thonormal list of vectors in V;
(j) there exists an orthonormal basis (e1,...,en)ofVsuch that
(S∗e1,...,S∗en)is orthonormal.
Proof: First suppose that (a) holds. If Vis a real inner-product
space, then for every u,v∈Vwe have
/angbracketleftSu,Sv/angbracketright=(/bardblSu+Sv/bardbl2−/bardblSu−Sv/bardbl2)/4
=(/bardblS(u+v)/bardbl2−/bardblS(u−v)/bardbl2)/4
=(/bardblu+v/bardbl2−/bardblu−v/bardbl2)/4
=/angbracketleftu,v/angbracketright,
where the first equality comes from Exercise 6 in Chapter 6, the second
equality comes from the linearity of S, the third equality holds because
Sis an isometry, and the last equality again comes from Exercise 6 in
Chapter 6. If Vis a complex inner-product space, then use Exercise 7
in Chapter 6 instead of Exercise 6 to obtain the same conclusion. Ineither case, we see that (a) implies (b).
Now suppose that (b) holds. Then
/angbracketleft(S
∗S−I)u,v/angbracketright=/angbracketleftSu,Sv/angbracketright−/angbracketleftu,v/angbracketright
=0
for everyu,v∈V. Takingv=(S∗S−I)u, we see that S∗S−I=0.
HenceS∗S=I, proving that (b) implies (c).
Now suppose that (c) holds. Suppose (e1,...,en)is an orthonormal
list of vectors in V. Then
/angbracketleftSej,Sek/angbracketright=/angbracketleftS∗Sej,ek/angbracketright
=/angbracketleftej,ek/angbracketright.
Hence(Se1,...,Sen)is orthonormal, proving that (c) implies (d).
Obviously (d) implies (e).
Now suppose (e) holds. Let (e1,...,en)be an orthonormal basis of V
such that(Se1,...,Sen)is orthonormal. If v∈V, then
150 Chapter 7.Operators on Inner-Product Spaces
/bardblSv/bardbl2=/bardblS/parenleftbig
/angbracketleftv,e 1/angbracketrighte1+···+/angbracketleftv,en/angbracketrighten/parenrightbig
/bardbl2
=/bardbl/angbracketleftv,e 1/angbracketrightSe 1+···+/angbracketleftv,en/angbracketrightSen/bardbl2
=|/angbracketleftv,e 1/angbracketright|2+···+|/angbracketleftv,en/angbracketright|2
=/bardblv/bardbl2,
where the first and last equalities come from 6.17. Taking square roots,
we see thatSis an isometry, proving that (e) implies (a).
Having shown that (a) ⇒(b)⇒(c)⇒(d)⇒(e)⇒(a), we know at this
stage that (a) through (e) are all equivalent to each other. Replacing S
withS∗, we see that (f) through (j) are all equivalent to each other. Thus
to complete the proof, we need only show that one of the conditions
in the group (a) through (e) is equivalent to one of the conditions inthe group (f) through (j). The easiest way to connect the two groups ofconditions is to show that (c) is equivalent to (h). In general, of course,Sneed not commute with S
∗. However,S∗S=Iif and only if SS∗=I;
this is a special case of Exercise 23 in Chapter 3. Thus (c) is equivalentto (h), completing the proof.
The last theorem shows that every isometry is normal (see (a), (c),
and (h) of 7.36). Thus the characterizations of normal operators canbe used to give complete descriptions of isometries. We do this in thenext two theorems.
7.37 Theorem: SupposeVis a complex inner-product space and
S∈L(V). ThenSis an isometry if and only if there is an orthonormal
basis ofVconsisting of eigenvectors of Sall of whose corresponding
eigenvalues have absolute value 1.
Proof: We already proved (see the first paragraph of this section)
that if there is an orthonormal basis of Vconsisting of eigenvectors of S
all of whose eigenvalues have absolute value 1, then Sis an isometry.
To prove the other direction, suppose Sis an isometry. By the com-
plex spectral theorem (7.9), there is an orthonormal basis (e
1,...,en)
ofVconsisting of eigenvectors of S. Forj∈{1,...,n}, letλjbe the
eigenvalue corresponding to ej. Then
|λj|=/bardblλjej/bardbl=/bardblSej/bardbl=/bardblej/bardbl=1.
Thus each eigenvalue of Shas absolute value 1, completing the proof.
Isometries 151
Ifθ∈R, then the operator on R2of counterclockwise rotation (cen-
tered at the origin) by an angle of θhas matrix 7.39 with respect to
the standard basis, as you should verify. The next result states that ev-
ery isometry on a real inner-product space is composed of pieces thatlook like rotations on two-dimensional subspaces, pieces that equal theidentity operator, and pieces that equal multiplication by −1.
7.38 Theorem: Suppose that Vis a real inner-product space and
This theorem implies
that an isometry on an
odd-dimensional real
inner-product spacemust have 1or−1as
an eigenvalue.S∈L(V). ThenSis an isometry if and only if there is an orthonormal
basis ofVwith respect to which Shas a block diagonal matrix where
each block on the diagonal is a 1-by-1 matrix containing 1or−1or a
2-by-2 matrix of the form
7.39/bracketleftBigg
cosθ−sinθ
sinθcosθ/bracketrightBigg
,
withθ∈(0,π) .
Proof: First suppose that Sis an isometry. Because Sis normal,
there is an orthonormal basis of Vsuch that with respect to this basis
Shas a block diagonal matrix, where each block is a 1-by-1 matrix or a
2-by-2 matrix of the form
7.40/bracketleftBigg
a−b
ba/bracketrightBigg
,
withb>0 (see 7.25).
Ifλis an entry in a 1-by-1 along the diagonal of the matrix of S(with
respect to the basis mentioned above), then there is a basis vector ej
such thatSej=λej. BecauseSis an isometry, this implies that |λ|=1.
Thusλ=1o rλ=−1 because these are the only real numbers with
absolute value 1.
Now consider a 2-by-2 matrix of the form 7.40 along the diagonal of
the matrix of S. There are basis vectors ej,ej+1such that
Sej=aej+bej+1.
Thus
1=/bardblej/bardbl2=/bardblSej/bardbl2=a2+b2.
The equation above, along with the condition b>0, implies that there
exists a number θ∈(0,π) such thata=cosθandb=sinθ. Thus the
152 Chapter 7.Operators on Inner-Product Spaces
matrix 7.40 has the required form 7.39, completing the proof in this
direction.
Conversely, now suppose that there is an orthonormal basis of V
with respect to which the matrix of Shas the form required by the
theorem. Thus there is a direct sum decomposition
V=U1⊕···⊕Um,
where eachUjis a subspace of Vof dimension 1 or 2. Furthermore,
any two vectors belonging to distinct U’s are orthogonal, and each S|Uj
is an isometry mapping UjintoUj.I fv∈V, we can write
v=u1+···+um,
where eachuj∈Uj. ApplyingSto the equation above and then taking
norms gives
/bardblSv/bardbl2=/bardblSu1+···+Sum/bardbl2
=/bardblSu1/bardbl2+···+/bardblSum/bardbl2
=/bardblu1/bardbl2+···+/bardblum/bardbl2
=/bardblv/bardbl2.
ThusSis an isometry, as desired.
Polar and Singular-Value Decompositions
Recall our analogy between CandL(V). Under this analogy, a com-
plex number zcorresponds to an operator T, and ¯zcorresponds to T∗.
The real numbers correspond to the self-adjoint operators, and the non-negative numbers correspond to the (badly named) positive operators.
Another distinguished subset of Cis the unit circle, which consists of
the complex numbers zsuch that|z|=1. The condition |z|=1i s
equivalent to the condition ¯zz=1. Under our analogy, this would cor-
respond to the condition T
∗T=I, which is equivalent to Tbeing an
isometry (see 7.36). In other words, the unit circle in Ccorresponds to
the isometries.
Continuing with our analogy, note that each complex number zex-
cept 0 can be written in the form
z=/parenleftbiggz
|z|/parenrightbigg
|z|=/parenleftbiggz
|z|/parenrightbigg/radicalbig
¯zz,
Polar and Singular-Value Decompositions 153
where the first factor, namely, z/|z|, is an element of the unit circle. Our
analogy leads us to guess that any operator T∈L(V)can be written
as an isometry times√
T∗T. That guess is indeed correct, as we now
prove.
7.41 Polar Decomposition: IfT∈L(V), then there exists an isom- If you know a bit of
complex analysis, you
will recognize the
analogy to polar
coordinates for
complex numbers:
every complex numbercan be written in the
forme
θir, where
θ∈[0,2π) andr≥0.
Note thateθiis in the
unit circle,
corresponding to S
being an isometry, andris nonnegative,
corresponding to√
T∗Tbeing a positive
operator.etryS∈L(V) such that
T=S√
T∗T.
Proof: SupposeT∈L(V).I fv∈V, then
/bardblTv/bardbl2=/angbracketleftTv,Tv/angbracketright
=/angbracketleftT∗Tv,v/angbracketright
=/angbracketleft√
T∗T√
T∗Tv,v/angbracketright
=/angbracketleft√
T∗Tv,√
T∗Tv/angbracketright
=/bardbl√
T∗Tv/bardbl2.
Thus
7.42 /bardblTv/bardbl=/bardbl√
T∗Tv/bardbl
for allv∈V.
Define a linear map S1: range√
T∗T→rangeTby
7.43 S1(√
T∗Tv)=Tv.
The idea of the proof is to extend S1to an isometry S∈L(V)such that
T=S√
T∗T. Now for the details.
First we must check that S1is well defined. To do this, suppose
v1,v2∈Vare such that√
T∗Tv1=√
T∗Tv2. For the definition given
by 7.43 to make sense, we must show that Tv1=Tv2. However,
/bardblTv 1−Tv2/bardbl=/bardblT(v 1−v2)/bardbl
=/bardbl√
T∗T(v 1−v2)/bardbl
=/bardbl√
T∗Tv1−√
T∗Tv2/bardbl
=0,
where the second equality holds by 7.42. The equation above shows
thatTv1=Tv2,s oS1is indeed well defined. You should verify that S1
is a linear map.
154 Chapter 7.Operators on Inner-Product Spaces
We see from 7.43 that S1maps range√
T∗Tonto rangeT. Clearly In the rest of the proof
all we are doing is
extendingS1to an
isometrySon all ofV.7.42 and 7.43 imply that /bardblS1u/bardbl=/bardblu/bardblfor allu∈range√
T∗T.I n
particular,S1is injective. Thus from 3.4, applied to S1, we have
dim range√
T∗T=dim rangeT.
This implies that dim(range√
T∗T)⊥=dim(rangeT)⊥(see Exercise 15
in Chapter 6). Thus orthonormal bases (e1,...,em)of(range√
T∗T)⊥
and(f1,...,fm)of(rangeT)⊥can be chosen; the key point here is that
these two orthonormal bases have the same length. Define a linear map
S2:(range√
T∗T)⊥→(rangeT)⊥by
S2(a1e1+···+amem)=a1f1+···+a mfm.
Obviously/bardblS2w/bardbl=/bardblw/bardblfor allw∈(range√
T∗T)⊥.
Now letSbe the operator on Vthat equalsS1on range√
T∗Tand
equalsS2on(range√
T∗T)⊥. More precisely, recall that each v∈V
can be written uniquely in the form
7.44 v=u+w,
whereu∈range√
T∗Tandw∈(range√
T∗T)⊥(see 6.29). For v∈V
with decomposition as above, define Svby
Sv=S1u+S2w.
For eachv∈Vwe have
S(√
T∗Tv)=S1(√
T∗Tv)=Tv,
soT=S√
T∗T, as desired. All that remains is to show that Sis an isom-
etry. However, this follows easily from the two uses of the Pythagoreantheorem: ifv∈Vhas decomposition as in 7.44, then
/bardblSv/bardbl
2=/bardblS1u+S2w/bardbl2
=/bardblS1u/bardbl2+/bardblS2w/bardbl2
=/bardblu/bardbl2+/bardblw/bardbl2
=/bardblv/bardbl2,
where the second equality above holds because S1u∈rangeTand
S2u∈(rangeT)⊥.
Polar and Singular-Value Decompositions 155
The polar decomposition (7.41) states that each operator on Vis the
product of an isometry and a positive operator. Thus we can write eachoperator on Vas the product of two operators, each of which comes
from a class that we have completely described and that we under-stand reasonably well. The isometries are described by 7.37 and 7.38;the positive operators (which are all self-adjoint) are described by the
spectral theorem (7.9 and 7.13).
Specifically, suppose T=S√
T∗Tis the polar decomposition of
T∈L(V), whereSis an isometry. Then there is an orthonormal basis
ofVwith respect to which Shas a diagonal matrix (if F=C) or a block
diagonal matrix with blocks of size at most 2-by-2 (if F=R), and there
is an orthonormal basis of Vwith respect to which√
T∗Thas a diag-
onal matrix. Warning: there may not exist an orthonormal basis thatsimultaneously puts the matrices of both Sand√
T∗Tinto these nice
forms (diagonal or block diagonal with small blocks). In other words, S
may require one orthonormal basis and√
T∗Tmay require a different
orthonormal basis.
SupposeT∈L(V). The singular values ofTare the eigenvalues
of√
T∗T, with each eigenvalue λrepeated dim null (√
T∗T−λI)times.
The singular values of Tare all nonnegative because they are the eigen-
values of the positive operator√
T∗T.
For example, if T∈L(F4)is defined by
7.45 T(z 1,z2,z3,z4)=(0,3z1,2z2,−3z 4),
thenT∗T(z 1,z2,z3,z4)=(9z 1,4z2,0,9z4), as you should verify. Thus
√
T∗T(z 1,z2,z3,z4)=(3z 1,2z2,0,3z4),
and we see that the eigenvalues of√
T∗Tare 3, 2,0. Clearly
dim null(√
T∗T−3I)=2,dim null(√
T∗T−2I)=1,dim null√
T∗T=1.
Hence the singular values of Tare 3, 3,2,0. In this example −3 and 0
are the only eigenvalues of T, as you should verify.
EachT∈L(V)has dimVsingular values, as can be seen by applying
the spectral theorem and 5.21 (see especially part (e)) to the positive(hence self-adjoint) operator√
T∗T. For example, the operator Tde-
fined by 7.45 on the four-dimensional vector space F4has four singular
values (they are 3, 3,2,0), as we saw in the previous paragraph.
The next result shows that every operator on Vhas a nice descrip-
tion in terms of its singular values and two orthonormal bases of V.
156 Chapter 7.Operators on Inner-Product Spaces
7.46 Singular-Value Decomposition: SupposeT∈L(V) has sin-
gular values s1,...,sn. Then there exist orthonormal bases (e1,...,en)
and(f1,...,fn)ofVsuch that
7.47 Tv=s1/angbracketleftv,e 1/angbracketrightf1+···+sn/angbracketleftv,en/angbracketrightfn
for everyv∈V.
Proof: By the spectral theorem (also see 7.14) applied to√
T∗T,
there is an orthonormal basis (e1,...,en)ofVsuch that√
T∗Tej=sjej
forj=1,...,n . We have
v=/angbracketleftv,e 1/angbracketrighte1+···+/angbracketleftv,en/angbracketrighten
for everyv∈V(see 6.17). Apply√
T∗Tto both sides of this equation,
getting√
T∗Tv=s1/angbracketleftv,e 1/angbracketrighte1+···+sn/angbracketleftv,en/angbracketrighten
for everyv∈V. By the polar decomposition (see 7.41), there is an This proof illustrates
the usefulness of the
polar decomposition.isometryS∈L(V)such thatT=S√
T∗T. ApplySto both sides of the
equation above, getting
Tv=s1/angbracketleftv,e 1/angbracketrightSe 1+···+sn/angbracketleftv,en/angbracketrightSen
for everyv∈V. For eachj, letfj=Sej. BecauseSis an isometry,
(f1,...,fn)is an orthonormal basis of V(see 7.36). The equation above
now becomes
Tv=s1/angbracketleftv,e 1/angbracketrightf1+···+sn/angbracketleftv,en/angbracketrightfn
for everyv∈V, completing the proof.
When we worked with linear maps from one vector space to a second
vector space, we considered the matrix of a linear map with respectto a basis for the first vector space and a basis for the second vectorspace. When dealing with operators, which are linear maps from avector space to itself, we almost always use only one basis, making it
play both roles.
The singular-value decomposition allows us a rare opportunity to
use two different bases for the matrix of an operator. To do this, sup-poseT∈L(V). Lets
1,...,sndenote the singular values of T, and let
(e1,...,en)and(f1,...,fn)be orthonormal bases of Vsuch that the
singular-value decomposition 7.47 holds. Then clearly
Polar and Singular-Value Decompositions 157
M/parenleftbig
T,(e 1,...,en),(f 1,...,fn)/parenrightbig
=
s1 0
...
0sn
.
In other words, every operator on Vhas a diagonal matrix with respect
to some orthonormal bases of V, provided that we are permitted to
use two different bases rather than a single basis as customary whenworking with operators.
Singular values and the singular-value decomposition have many ap-
plications (some are given in the exercises), including applications incomputational linear algebra. To compute numeric approximations tothe singular values of an operator T, first compute T
∗Tand then com-
pute approximations to the eigenvalues of T∗T(good techniques exist
for approximating eigenvalues of positive operators). The nonnegativesquare roots of these (approximate) eigenvalues of T
∗Twill be the (ap-
proximate) singular values of T(as can be seen from the proof of 7.28).
In other words, the singular values of Tcan be approximated without
computing the square root of T∗T.
158 Chapter 7.Operators on Inner-Product Spaces
Exercises
1. MakeP2(R)into an inner-product space by defining
/angbracketleftp,q/angbracketright=/integraldisplay1
0p(x)q(x)dx.
DefineT∈L(P2(R))byT(a 0+a1x+a2x2)=a1x.
(a) Show that Tis not self-adjoint.
(b) The matrix of Twith respect to the basis (1,x,x2)is
000
010000
.
This matrix equals its conjugate transpose, even though T
is not self-adjoint. Explain why this is not a contradiction.
2. Prove or give a counterexample: the product of any two self-
adjoint operators on a finite-dimensional inner-product space isself-adjoint.
3. (a) Show that if Vis a real inner-product space, then the set
of self-adjoint operators on Vis a subspace of L(V).
(b) Show that if Vis a complex inner-product space, then the
set of self-adjoint operators on Vis not a subspace of
L(V).
4. Suppose P∈L(V)is such thatP
2=P. Prove thatPis an orthog-
onal projection if and only if Pis self-adjoint.
5. Show that if dim V≥2, then the set of normal operators on Vis
not a subspace of L(V).
6. Prove that if T∈L(V)is normal, then
rangeT=rangeT∗.
7. Prove that if T∈L(V)is normal, then
nullTk=nullTand rangeTk=rangeT
for every positive integer k.
Exercises 159
8. Prove that there does not exist a self-adjoint operator T∈L(R3)
such thatT(1,2,3)=(0,0,0)andT(2,5,7)=(2,5,7).
9. Prove that a normal operator on a complex inner-product space Exercise 9 strengthens
the analogy (for normal
operators) betweenself-adjoint operators
and real numbers.is self-adjoint if and only if all its eigenvalues are real.
10. Suppose Vis a complex inner-product space and T∈L(V)is a
normal operator such that T9=T8. Prove that Tis self-adjoint
andT2=T.
11. Suppose Vis a complex inner-product space. Prove that every
normal operator on Vhas a square root. (An operator S∈L(V)
is called a square root ofT∈L(V)ifS2=T.)
12. Give an example of a real inner-product space VandT∈L(V) This exercise shows
that the hypothesisthatTis self-adjoint is
needed in 7.11, evenfor real vector spaces.and real numbers α,βwithα2<4βsuch thatT2+αT+βIis
not invertible.
13. Prove or give a counterexample: every self-adjoint operator on
Vhas a cube root. (An operator S∈L(V)is called a cube root
ofT∈L(V)ifS3=T.)
14. Suppose T∈L(V)is self-adjoint, λ∈F, and/epsilon1>0. Prove that if
there existsv∈Vsuch that/bardblv/bardbl=1 and
/bardblTv−λv/bardbl</epsilon1,
thenThas an eigenvalue λ/primesuch that|λ−λ/prime|</epsilon1.
15. Suppose Uis a finite-dimensional real vector space and T∈
L(U). Prove that Uhas a basis consisting of eigenvectors of Tif
and only if there is an inner product on Uthat makesTinto a
self-adjoint operator.
16. Give an example of an operator Ton an inner product space such This exercise shows
that 7.18 can fail
without the hypothesisthatTis normal.thatThas an invariant subspace whose orthogonal complement
is not invariant under T.
17. Prove that the sum of any two positive operators on Vis positive.
18. Prove that if T∈L(V)is positive, then so is Tkfor every positive
integerk.
160 Chapter 7.Operators on Inner-Product Spaces
19. Suppose that Tis a positive operator on V. Prove that Tis in-
vertible if and only if
/angbracketleftTv,v/angbracketright>0
for everyv∈V\{0}.
20. Prove or disprove: the identity operator on F2has infinitely many
self-adjoint square roots.
21. Prove or give a counterexample: if S∈L(V)and there exists
an orthonormal basis (e1,...,en)ofVsuch that/bardblSej/bardbl=1 for
eachej, thenSis an isometry.
22. Prove that if S∈L(R3)is an isometry, then there exists a nonzero
vectorx∈R3such thatS2x=x.
23. Define T∈L(F3)by
T(z 1,z2,z3)=(z3,2z1,3z2).
Find (explicitly) an isometry S∈L(F3)such thatT=S√
T∗T.
24. Suppose T∈L(V),S∈L(V)is an isometry, and R∈L(V)is a Exercise 24 shows that
if we writeTas the
product of an isometry
and a positive operator
(as in the polar
decomposition), then
the positive operator
must equal√
T∗T.positive operator such that T=SR. Prove thatR=√
T∗T.
25. Suppose T∈L(V). Prove that Tis invertible if and only if there
exists a unique isometry S∈L(V)such thatT=S√
T∗T.
26. Prove that if T∈L(V)is self-adjoint, then the singular values
ofTequal the absolute values of the eigenvalues of T(repeated
appropriately).
27. Prove or give a counterexample: if T∈L(V), then the singular
values ofT2equal the squares of the singular values of T.
28. Suppose T∈L(V). Prove that Tis invertible if and only if 0 is
not a singular value of T.
29. Suppose T∈L(V). Prove that dim range Tequals the number of
nonzero singular values of T.
30. Suppose S∈L(V). Prove that Sis an isometry if and only if all
the singular values of Sequal 1.
Exercises 161
31. Suppose T1,T2∈L(V). Prove that T1andT2have the same
singular values if and only if there exist isometries S1,S2∈L(V)
such thatT1=S1T2S2.
32. Suppose T∈L(V)has singular-value decomposition given by
Tv=s1/angbracketleftv,e 1/angbracketrightf1+···+sn/angbracketleftv,en/angbracketrightfn
for everyv∈V, wheres1,...,snare the singular values of Tand
(e1,...,en)and(f1,...,fn)are orthonormal bases of V.
(a) Prove that
T∗v=s1/angbracketleftv,f 1/angbracketrighte1+···+s n/angbracketleftv,fn/angbracketrighten
for everyv∈V.
(b) Prove that if Tis invertible, then
T−1v=/angbracketleftv,f 1/angbracketrighte1
s1+···+/angbracketleftv,fn/angbracketrighten
sn
for everyv∈V.
33. Suppose T∈L(V). Let ˆsdenote the smallest singular value of T,
and letsdenote the largest singular value of T. Prove that
ˆs/bardblv/bardbl≤/bardblTv/bardbl≤s/bardblv/bardbl
for everyv∈V.
34. Suppose T/prime,T/prime/prime∈L(V). Lets/primedenote the largest singular value
ofT/prime, lets/prime/primedenote the largest singular value of T/prime/prime, and lets
denote the largest singular value of T/prime+T/prime/prime. Prove thats≤s/prime+s/prime/prime.
Chapter 8
Operators on
Complex Vector Spaces
In this chapter we delve deeper into the structure of operators on
complex vector spaces. An inner product does not help with this ma-terial, so we return to the general setting of a finite-dimensional vectorspace (as opposed to the more specialized context of an inner-productspace). Thus our assumptions for this chapter are as follows:
Recall that Fdenotes RorC.
Also,Vis a finite-dimensional, nonzero vector space over F.
Some of the results in this chapter are valid on real vector spaces,
so we have not assumed that Vis a complex vector space. Most of the
results in this chapter that are proved only for complex vector spaceshave analogous results on real vector spaces that are proved in the nextchapter. We deal with complex vector spaces first because the proofson complex vector spaces are often simpler than the analogous proofs
on real vector spaces.
✽✽✽
✽✽✽✽✽
163
164 Chapter 8.Operators on Complex Vector Spaces
Generalized Eigenvectors
Unfortunately some operators do not have enough eigenvectors to
lead to a good description. Thus in this section we introduce the con-
cept of generalized eigenvectors, which will play a major role in our
description of the structure of an operator.
To understand why we need more than eigenvectors, let’s examine
the question of describing an operator by decomposing its domain intoinvariant subspaces. Fix T∈L(V). We seek to describe Tby finding a
“nice” direct sum decomposition
8.1 V=U
1⊕···⊕Um,
where eachUjis a subspace of Vinvariant under T. The simplest pos-
sible nonzero invariant subspaces are one-dimensional. A decompo-
sition 8.1 where each Ujis a one-dimensional subspace of Vinvariant
underTis possible if and only if Vhas a basis consisting of eigenvectors
ofT(see 5.21). This happens if and only if Vhas the decomposition
8.2 V=null(T−λ1I)⊕···⊕ null(T−λmI),
whereλ1,...,λmare the distinct eigenvalues of T(see 5.21).
In the last chapter we showed that a decomposition of the form
8.2 holds for every self-adjoint operator on an inner-product space(see 7.14). Sadly, a decomposition of the form 8.2 may not hold formore general operators, even on a complex vector space. An exam-
ple was given by the operator in 5.19, which does not have enough
eigenvectors for 8.2 to hold. Generalized eigenvectors, which we nowintroduce, will remedy this situation. Our main goal in this chapter isto show that if Vis a complex vector space and T∈L(V), then
V=null(T−λ
1I)dimV⊕···⊕ null(T−λmI)dimV,
whereλ1,...,λmare the distinct eigenvalues of T(see 8.23).
SupposeT∈L(V)andλis an eigenvalue of T. A vectorv∈Vis
called a generalized eigenvector ofTcorresponding to λif
8.3 (T−λI)jv=0
for some positive integer j. Note that every eigenvector of Tis a gen-
eralized eigenvector of T(takej=1 in the equation above), but the
converse is not true. For example, if T∈L(C3)is defined by
Generalized Eigenvectors 165
T(z 1,z2,z3)=(z2,0,z 3),
thenT2(z1,z2,0)=0 for allz1,z2∈C. Hence every element of C3
whose last coordinate equals 0 is a generalized eigenvector of T.A s
you should verify,
C3={(z1,z2,0):z1,z2∈C}⊕{(0,0,z 3):z3∈C},
where the first subspace on the right equals the set of generalized eigen-
vectors for this operator corresponding to the eigenvalue 0 and the sec-
ond subspace on the right equals the set of generalized eigenvectorscorresponding to the eigenvalue 1. Later in this chapter we will provethat a decomposition using generalized eigenvectors exists for everyoperator on a complex vector space (see 8.23).
Thoughjis allowed to be an arbitrary integer in the definition of a
Note that we do not
define the concept of a
generalized eigenvalue
because this would notlead to anything new.
Reason: if(T−λI)
jis
not injective for some
positive integer j, then
T−λIis not injective,
and henceλis an
eigenvalue of T.generalized eigenvector, we will soon see that every generalized eigen-
vector satisfies an equation of the form 8.3 with jequal to the dimen-
sion ofV. To prove this, we now turn to a study of null spaces of
powers of an operator.
SupposeT∈L(V)andkis a nonnegative integer. If Tkv=0, then
Tk+1v=T(Tkv)=T(0)=0. Thus null Tk⊂nullTk+1. In other words,
we have
8.4{0}=nullT0⊂nullT1⊂···⊂null Tk⊂nullTk+1⊂···.
The next proposition says that once two consecutive terms in this se-
quence of subspaces are equal, then all later terms in the sequence areequal.
8.5 Proposition: IfT∈L(V) andmis a nonnegative integer such
that nullT
m=nullTm+1, then
nullT0⊂nullT1⊂···⊂ nullTm=nullTm+1=nullTm+2=···.
Proof: SupposeT∈L(V)andmis a nonnegative integer such
that nullTm=nullTm+1. Letkbe a positive integer. We want to prove
that
nullTm+k=nullTm+k+1.
We already know that null Tm+k⊂nullTm+k+1. To prove the inclusion
in the other direction, suppose that v∈nullTm+k+1. Then
166 Chapter 8.Operators on Complex Vector Spaces
0=Tm+k+1v=Tm+1(Tkv).
Hence
Tkv∈nullTm+1=nullTm.
Thus 0=Tm(Tkv)=Tm+kv, which means that v∈nullTm+k. This
implies that null Tm+k+1⊂nullTm+k, completing the proof.
The proposition above raises the question of whether there must ex-
ist a nonnegative integer msuch that null Tm=nullTm+1. The propo-
sition below shows that this equality holds at least when mequals the
dimension of the vector space on which Toperates.
8.6 Proposition: IfT∈L(V), then
nullTdimV=nullTdimV+1=nullTdimV+2=···.
Proof: SupposeT∈L(V). To get our desired conclusion, we need
only prove that null TdimV=nullTdimV+1(by 8.5). Suppose this is not
true. Then, by 8.5, we have
{0}=nullT0⊊nullT1⊊···⊊nullTdimV⊊nullTdimV+1,
where the symbol ⊊means “contained in but not equal to”. At each of
the strict inclusions in the chain above, the dimension must increase byat least 1. Thus dim null T
dimV+1≥dimV+1, a contradiction because
a subspace of Vcannot have a larger dimension than dim V.
Now we have the promised description of generalized eigenvectors.
8.7 Corollary: SupposeT∈L(V) andλis an eigenvalue of T. Then This corollary implies
that the set of
generalized
eigenvectors of
T∈L(V)
corresponding to an
eigenvalueλis a
subspace of V.the set of generalized eigenvectors of Tcorresponding to λequals
null(T−λI)dimV.
Proof: Ifv∈null(T−λI)dimV, then clearly vis a generalized
eigenvector of Tcorresponding to λ(by the definition of generalized
eigenvector).
Conversely, suppose that v∈Vis a generalized eigenvector of T
corresponding to λ. Thus there is a positive integer jsuch that
v∈null(T−λI)j.
From 8.5 and 8.6 (with T−λIreplacingT), we getv∈null(T−λI)dimV,
as desired.
Generalized Eigenvectors 167
An operator is called nilpotent if some power of it equals 0. For The Latin word nil
means nothing or zero;the Latin word potent
means power. Thus
nilpotent literally
means zero power.example, the operator N∈L(F4)defined by
N(z 1,z2,z3,z4)=(z3,z4,0,0)
is nilpotent because N2=0. As another example, the operator of dif-
ferentiation on Pm(R)is nilpotent because the (m+1)stderivative of
any polynomial of degree at most mequals 0. Note that on this space of
dimensionm+1, we need to raise the nilpotent operator to the power
m+1 to get 0. The next corollary shows that we never need to use a
power higher than the dimension of the space.
8.8 Corollary: SupposeN∈L(V) is nilpotent. Then NdimV=0.
Proof: BecauseNis nilpotent, every vector in Vis a generalized
eigenvector corresponding to the eigenvalue 0. Thus from 8.7 we seethat nullN
dimV=V, as desired.
Having dealt with null spaces of powers of operators, we now turn
our attention to ranges. Suppose T∈L(V)andkis a nonnegative
integer. Ifw∈rangeTk+1, then there exists v∈Vwith
w=Tk+1v=Tk(Tv)∈rangeTk.
Thus rangeTk+1⊂rangeTk. In other words, we have
These inclusions go in
the opposite directionfrom the corresponding
inclusions for null
spaces (8.4).V=rangeT0⊃rangeT1⊃···⊃range Tk⊃rangeTk+1⊃···.
The proposition below shows that the inclusions above become equal-
ities once the power reaches the dimension of V.
8.9 Proposition: IfT∈L(V), then
rangeTdimV=rangeTdimV+1=rangeTdimV+2=···.
Proof: We could prove this from scratch, but instead let’s make use
of the corresponding result already proved for null spaces. Supposem> dimV. Then
dim rangeT
m=dimV−dim nullTm
=dimV−dim nullTdimV
=dim rangeTdimV,
168 Chapter 8.Operators on Complex Vector Spaces
where the first and third equalities come from 3.4 and the second equal-
ity comes from 8.6. We already know that range TdimV⊃rangeTm.W e
just showed that dim range TdimV=dim rangeTm, so this implies that
rangeTdimV=rangeTm, as desired.
The Characteristic Polynomial
SupposeVis a complex vector space and T∈L(V). We know that
Vhas a basis with respect to which Thas an upper-triangular matrix
(see 5.13). In general, this matrix is not unique— Vmay have many
different bases with respect to which Thas an upper-triangular matrix,
and with respect to these different bases we may get different upper-triangular matrices. However, the diagonal of any such matrix mustcontain precisely the eigenvalues of T(see 5.18). Thus if Thas dimV
distinct eigenvalues, then each one must appear exactly once on thediagonal of any upper-triangular matrix of T.
What ifThas fewer than dim Vdistinct eigenvalues, as can easily
happen? Then each eigenvalue must appear at least once on the diag-onal of any upper-triangular matrix of T, but some of them must be
repeated. Could the number of times that a particular eigenvalue isrepeated depend on which basis of Vwe choose?
You might guess that a number λappears on the diagonal of an
IfThappens to have a
diagonal matrix Awith
respect to some basis,
thenλappears on the
diagonal ofAprecisely
dim null(T−λI) times,
as you should verify.upper-triangular matrix of Tprecisely dim null (T−λI)times. In gen-
eral, this is false. For example, consider the operator on C2whose
matrix with respect to the standard basis is the upper-triangular matrix
/bracketleftBigg
51
05/bracketrightBigg
.
For this operator, dim null (T−5I)=1 but 5 appears on the diago-
nal twice. Note, however, that dim null (T−5I)2=2 for this oper-
ator. This example illustrates the general situation—a number λap-
pears on the diagonal of an upper-triangular matrix of Tprecisely
dim null(T−λI)dimVtimes, as we will show in the following theorem.
Because null (T−λI)dimVdepends only on Tandλand not on a choice
of basis, this implies that the number of times an eigenvalue is repeatedon the diagonal of an upper-triangular matrix of Tis independent of
which particular basis we choose. This result will be our key tool inanalyzing the structure of an operator on a complex vector space.
The Characteristic Polynomial 169
8.10 Theorem: LetT∈L(V) andλ∈F. Then for every basis of V
with respect to which Thas an upper-triangular matrix, λappears on
the diagonal of the matrix of Tprecisely dim null(T−λI)dimVtimes.
Proof: We will assume, without loss of generality, that λ=0 (once
the theorem is proved in this case, the general case is obtained by re-
placingTwithT−λI).
For convenience let n=dimV. We will prove this theorem by induc-
tion onn. Clearly the desired result holds if n=1. Thus we can assume
thatn>1 and that the desired result holds on spaces of dimension
n−1.
Suppose(v1,...,vn)is a basis of Vwith respect to which Thas an
upper-triangular matrix Recall that an asterisk
is often used in
matrices to denoteentries that we do not
know or care about.8.11
λ
1 ∗
...
λn−1
0 λn
.
LetU=span(v
1,...,vn−1). ClearlyUis invariant under T(see 5.12),
and the matrix of T|Uwith respect to the basis (v1,...,vn−1)is
8.12
λ1∗
...
0λn−1
.
Thus, by our induction hypothesis, 0 appears on the diagonal of 8.12
dim null(T|U)n−1times. We know that null (T|U)n−1=null(T|U)n(be-
causeUhas dimension n−1; see 8.6). Hence
8.13 0 appears on the diagonal of 8.12 dim null (T|U)ntimes.
The proof breaks into two cases, depending on whether λn=0. First
consider the case where λn/negationslash=0. We will show that in this case
8.14 nullTn⊂U.
Once this has been verified, we will know that null Tn=null(T|U)n, and
hence 8.13 will tell us that 0 appears on the diagonal of 8.11 exactly
dim nullTntimes, completing the proof in the case where λn/negationslash=0.
BecauseM(T) is given by 8.11, we have
170 Chapter 8.Operators on Complex Vector Spaces
M(Tn)=M(T)n=
λ
1n∗
...
λn−1n
0 λnn
.
This shows that
T
nvn=u+λnnvn
for someu∈U. To prove 8.14 (still assuming that λn/negationslash=0), suppose
v∈nullTn. We can write vin the form
v=˜u+avn,
where ˜u∈Uanda∈F. Thus
0=Tnv=Tn˜u+aTnvn=Tn˜u+au+aλnnvn.
BecauseTn˜uandauare inUandvn∉U, this implies that aλnn=0.
However,λn/negationslash=0, soa=0. Thusv=˜u∈U, completing the proof
of 8.14.
Now consider the case where λn=0. In this case we will show that
8.15 dim nullTn=dim null(T|U)n+1,
which along with 8.13 will complete the proof when λn=0.
Using the formula for the dimension of the sum of two subspaces
(2.18), we have
dim nullTn=dim(U∩nullTn)+dim(U+nullTn)−dimU
=dim null(T|U)n+dim(U+nullTn)−(n−1).
Suppose we can prove that null Tncontains a vector not in U. Then
n=dimV≥dim(U+nullTn)>dimU=n−1,
which implies that dim (U+nullTn)=n, which when combined with
the formula above for dim null Tngives 8.15, as desired. Thus to com-
plete the proof, we need only show that null Tncontains a vector not
inU.
Let’s think about how we might find a vector in null Tnthat is not
inU. We might try a vector of the form
u−vn,
The Characteristic Polynomial 171
whereu∈U. At least we are guaranteed that any such vector is not
inU. Can we choose u∈Usuch that the vector above is in null Tn?
Let’s compute:
Tn(u−vn)=Tnu−Tnvn.
To make the above vector equal 0, we must choose (if possible) u∈U
such thatTnu=Tnvn. We can do this if Tnvn∈range(T|U)n. Because
8.11 is the matrix of Twith respect to (v1,...,vn), we see that Tvn∈U
(recall that we are considering the case where λn=0). Thus
Tnvn=Tn−1(Tvn)∈range(T|U)n−1=range(T|U)n,
where the last equality comes from 8.9. In other words, we can indeed
chooseu∈Usuch thatu−vn∈nullTn, completing the proof.
SupposeT∈L(V). The multiplicity of an eigenvalue λofTis de- Our definition of
multiplicity has a clear
connection with the
geometric behavior
ofT. Most texts define
multiplicity in terms ofthe multiplicity of the
roots of a certain
polynomial defined by
determinants. These
two definitions turn
out to be equivalent.fined to be the dimension of the subspace of generalized eigenvectors
corresponding to λ. In other words, the multiplicity of an eigenvalue λ
ofTequals dim null (T−λI)dimV.I fThas an upper-triangular matrix
with respect to some basis of V(as always happens when F=C), then
the multiplicity of λis simply the number of times λappears on the
diagonal of this matrix (by the last theorem).
As an example of multiplicity, consider the operator T∈L(F3)de-
fined by
8.16 T(z 1,z2,z3)=(0,z 1,5z3).
You should verify that 0 is an eigenvalue of Twith multiplicity 2, that
5 is an eigenvalue of Twith multiplicity 1, and that Thas no additional
eigenvalues. As another example, if T∈L(F3)is the operator whose
matrix is
8.17
677
067007
,
then 6 is an eigenvalue of Twith multiplicity 2 and 7 is an eigenvalue
ofTwith multiplicity 1 (this follows from the last theorem).
In each of the examples above, the sum of the multiplicities of the
eigenvalues of Tequals 3, which is the dimension of the domain of T.
The next proposition shows that this always happens on a complex
vector space.
172 Chapter 8.Operators on Complex Vector Spaces
8.18 Proposition: IfVis a complex vector space and T∈L(V), then
the sum of the multiplicities of all the eigenvalues of Tequals dimV.
Proof: SupposeVis a complex vector space and T∈L(V). Then
there is a basis of Vwith respect to which the matrix of Tis upper
triangular (by 5.13). The multiplicity of λequals the number of times λ
appears on the diagonal of this matrix (from 8.10). Because the diagonalof this matrix has length dim V, the sum of the multiplicities of all the
eigenvalues of Tmust equal dim V.
SupposeVis a complex vector space and T∈L(V). Letλ1,...,λm
denote the distinct eigenvalues of T. Letdjdenote the multiplicity
ofλjas an eigenvalue of T. The polynomial
(z−λ1)d1...(z−λm)dm
is called the characteristic polynomial ofT. Note that the degree of Most texts define the
characteristic
polynomial using
determinants. The
approach taken here,
which is considerably
simpler, leads to an
easy proof of the
Cayley-Hamilton
theorem.the characteristic polynomial of Tequals dimV(from 8.18). Obviously
the roots of the characteristic polynomial of Tequal the eigenvalues
ofT. As an example, the characteristic polynomial of the operator
T∈L(C3)defined by 8.16 equals z2(z−5).
Here is another description of the characteristic polynomial of an
operator on a complex vector space. Suppose Vis a complex vector
space andT∈L(V). Consider any basis of Vwith respect to which T
has an upper-triangular matrix of the form
8.19 M(T)=
λ1∗
...
0λn
.
Then the characteristic polynomial of Tis given by
(z−λ1)...(z−λn);
this follows immediately from 8.10. As an example of this procedure,
ifT∈L(C3)is the operator whose matrix is given by 8.17, then the
characteristic polynomial of Tequals(z−6)2(z−7).
In the next chapter we will define the characteristic polynomial of
an operator on a real vector space and prove that the next result also
holds for real vector spaces.
Decomposition of an Operator 173
8.20 Cayley-Hamilton Theorem: Suppose that Vis a complex vector The English
mathematician ArthurCayley published threemathematics papers
before he completed
his undergraduatedegree in 1842. TheIrish mathematicianWilliam Hamilton was
made a professor in
1827 when he was 22
years old and still anundergraduate!space andT∈L(V). Letqdenote the characteristic polynomial of T.
Thenq(T)=0.
Proof: Suppose(v1,...,vn)is a basis of Vwith respect to which
the matrix of Thas the upper-triangular form 8.19. To prove that
q(T)=0, we need only show that q(T)vj=0 forj=1,...,n .T o
do this, it suffices to show that
8.21 (T−λ1I)...(T−λjI)vj=0
forj=1,...,n .
We will prove 8.21 by induction on j. To get started, suppose j=1.
BecauseM/parenleftbig
T,(v 1,...,vn)/parenrightbig
is given by 8.19, we have Tv1=λ1v1, giving
8.21 whenj=1.
Now suppose that 1 <j≤nand that
0=(T−λ1I)v1
=(T−λ1I)(T−λ2I)v2
...
=(T−λ1I)...(T−λj−1I)vj−1.
BecauseM/parenleftbig
T,(v 1,...,vn)/parenrightbig
is given by 8.19, we see that
(T−λjI)vj∈span(v 1,...,vj−1).
Thus, by our induction hypothesis, (T−λ1I)...(T−λj−1I)applied to
(T−λjI)vjgives 0. In other words, 8.21 holds, completing the proof.
Decomposition of an Operator
We saw earlier that the domain of an operator might not decompose
into invariant subspaces consisting of eigenvectors of the operator,even on a complex vector space. In this section we will see that everyoperator on a complex vector space has enough generalized eigenvec-tors to provide a decomposition.
We observed earlier that if T∈L(V), then null Tis invariant un-
derT. Now we show that the null space of any polynomial of Tis also
invariant under T.
174 Chapter 8.Operators on Complex Vector Spaces
8.22 Proposition: IfT∈L(V) andp∈P(F), then nullp(T) is
invariant under T.
Proof: SupposeT∈L(V)andp∈P(F). Letv∈nullp(T) . Then
p(T)v=0. Thus
(p(T))(Tv) =T(p(T)v)=T(0)=0,
and henceTv∈nullp(T) . Thus null p(T) is invariant under T,a s
desired.
The following major structure theorem shows that every operator on
a complex vector space can be thought of as composed of pieces, eachof which is a nilpotent operator plus a scalar multiple of the identity.
Actually we have already done all the hard work, so at this point the
proof is easy.
8.23 Theorem: SupposeVis a complex vector space and T∈L(V).
Letλ
1,...,λmbe the distinct eigenvalues of T, and letU1,...,Umbe
the corresponding subspaces of generalized eigenvectors. Then
(a)V=U1⊕···⊕Um;
(b) eachUjis invariant under T;
(c) each(T−λjI)|Ujis nilpotent.
Proof: Note thatUj=null(T−λjI)dimVfor eachj(by 8.7). From
8.22 (withp(z)=(z−λj)dimV), we get (b). Obviously (c) follows from
the definitions.
To prove (a), recall that the multiplicity of λjas an eigenvalue of T
is defined to be dim Uj. The sum of these multiplicities equals dim V
(see 8.18); thus
8.24 dimV=dimU1+···+ dimUm.
LetU=U1+···+U m. ClearlyUis invariant under T. Thus we can
defineS∈L(U)by
S=T|U.
Note thatShas the same eigenvalues, with the same multiplicities, as T
because all the generalized eigenvectors of Tare inU, the domain of S.
Thus applying 8.18 to S, we get
Decomposition of an Operator 175
dimU=dimU1+···+dim Um.
This equation, along with 8.24, shows that dim V=dimU. BecauseU
is a subspace of V, this implies that V=U. In other words,
V=U1+···+Um.
This equation, along with 8.24, allows us to use 2.19 to conclude that
(a) holds, completing the proof.
As we know, an operator on a complex vector space may not have
enough eigenvectors to form a basis for the domain. The next resultshows that on a complex vector space there are enough generalizedeigenvectors to do this.
8.25 Corollary: SupposeVis a complex vector space and T∈L(V).
Then there is a basis of Vconsisting of generalized eigenvectors of T.
Proof: Choose a basis for each U
jin 8.23. Put all these bases
together to form a basis of Vconsisting of generalized eigenvectors
ofT.
Given an operator TonV, we want to find a basis of Vso that the
matrix ofTwith respect to this basis is as simple as possible, meaning
that the matrix contains many 0’s. We begin by showing that if Nis
nilpotent, we can choose a basis of Vsuch that the matrix of Nwith
respect to this basis has more than half of its entries equal to 0.
8.26 Lemma: SupposeNis a nilpotent operator on V. Then there is IfVis complex vector
space, a proof of this
lemma follows easily
from Exercise 6 in this
chapter, 5.13, and 5.18.
But the proof givenhere uses simpler ideasthan needed to prove
5.13, and it works for
both real and complex
vector spaces.a basis ofVwith respect to which the matrix of Nhas the form
8.27
0∗
...
00
;
here all entries on and below the diagonal are 0’s.
Proof: First choose a basis of null N. Then extend this to a basis
of nullN2. Then extend to a basis of null N3. Continue in this fashion,
eventually getting a basis of V(because null Nm=Vformsufficiently
large).
176 Chapter 8.Operators on Complex Vector Spaces
Now let’s think about the matrix of Nwith respect to this basis. The
first column, and perhaps additional columns at the beginning, consistsof all 0’s because the corresponding basis vectors are in null N. The
next set of columns comes from basis vectors in null N
2. ApplyingN
to any such vector, we get a vector in null N; in other words, we get a
vector that is a linear combination of the previous basis vectors. Thusall nonzero entries in these columns must lie above the diagonal. Thenext set of columns come from basis vectors in null N
3. ApplyingN
to any such vector, we get a vector in null N2; in other words, we get a
vector that is a linear combination of the previous basis vectors. Thus,once again, all nonzero entries in these columns must lie above thediagonal. Continue in this fashion to complete the proof.
Note that in the next theorem we get many more zeros in the matrix
ofTthan are needed to make it upper triangular.
8.28 Theorem: SupposeVis a complex vector space and T∈L(V).
Letλ1,...,λmbe the distinct eigenvalues of T. Then there is a basis
ofVwith respect to which Thas a block diagonal matrix of the form
A1 0
...
0Am
,
where eachAjis an upper-triangular matrix of the form
8.29 Aj=
λj∗
...
0λj
.
Proof: Forj=1,...,m , letUjdenote the subspace of generalized
eigenvectors of Tcorresponding to λj. Thus(T−λjI)|Ujis nilpotent
(see 8.23(c)). For each j, choose a basis of Ujsuch that the matrix of
(T−λjI)|Ujwith respect to this basis is as in 8.26. Thus the matrix of
T|Ujwith respect to this basis will look like 8.29. Putting the bases for
theUj’s together gives a basis for V(by 8.23(a)). The matrix of Twith
respect to this basis has the desired form.
Square Roots 177
Square Roots
Recall that a square root of an operator T∈L(V)is an operator
S∈L(V)such thatS2=T. As an application of the main structure
theorem from the last section, in this section we will show that everyinvertible operator on a complex vector space has a square root.
Every complex number has a square root, but not every operator on
a complex vector space has a square root. An example of an operatoronC
3that has no square root is given in Exercise 4 in this chapter.
The noninvertibility of that particular operator is no accident, as wewill soon see. We begin by showing that the identity plus a nilpotentoperator always has a square root.
8.30 Lemma: SupposeN∈L(V) is nilpotent. Then I+Nhas a
square root.
Proof: Consider the Taylor series for the function√
1+x:
Becausea1=1/2, this
formula shows that
1+x/2is a good
estimate for√
1+x
whenxis small.8.31/radicalbig
1+x=1+a1x+a2x2+···.
We will not find an explicit formula for all the coefficients or worry
about whether the infinite sum converges because we are using thisequation only as motivation, not as a formal part of the proof.
BecauseNis nilpotent, N
m=0 for some positive integer m. In 8.31,
suppose we replace xwithNand 1 withI. Then the infinite sum on
the right side becomes a finite sum (because Nj=0 for allj≥m). In
other words, we guess that there is a square root of I+Nof the form
I+a1N+a2N2+···+am−1Nm−1.
Having made this guess, we can try to choose a1,a2,...,am−1so that
the operator above has its square equal to I+N. Now
(I+a1N+a2N2+a3N3+···+am−1Nm−1)2
=I+2a1N+(2a 2+a12)N2+(2a 3+2a1a2)N3+···
+(2am−1+terms involving a1,...,am−2)Nm−1.
We want the right side of the equation above to equal I+N. Hence
choosea1so that 2a1=1 (thusa1=1/2). Next, choose a2so that
2a2+a12=0 (thusa2=−1/8). Then choose a3so that the coefficient
ofN3on the right side of the equation above equals 0 (thus a3=1/16).
178 Chapter 8.Operators on Complex Vector Spaces
Continue in this fashion for j=4,...,m−1, at each step solving for
ajso that the coefficient of Njon the right side of the equation above
equals 0. Actually we do not care about the explicit formula for thea
j’s. We need only know that some choice of the aj’s gives a square
root ofI+N.
The previous lemma is valid on real and complex vector spaces.
However, the next result holds only on complex vector spaces.
8.32 Theorem: SupposeVis a complex vector space. If T∈L(V) On real vector spaces
there exist invertible
operators that have no
square roots. For
example, the operator
of multiplication by −1
onRhas no square
root because no real
number has its square
equal to−1.is invertible, then Thas a square root.
Proof: SupposeT∈L(V)is invertible. Let λ1,...,λmbe the dis-
tinct eigenvalues of T, and letU1,...,Umbe the corresponding sub-
spaces of generalized eigenvectors. For each j, there exists a nilpotent
operatorNj∈L(Uj)such thatT|Uj=λjI+Nj(see 8.23(c)). Because T
is invertible, none of the λj’s equals 0, so we can write
T|Uj=λj/parenleftbig
I+Nj
λj/parenrightbig
for eachj. ClearlyNj/λjis nilpotent, and so I+Nj/λjhas a square
root (by 8.30). Multiplying a square root of the complex number λjby
a square root of I+Nj/λj, we obtain a square root SjofT|Uj.
A typical vector v∈Vcan be written uniquely in the form
v=u1+···+um,
where eachuj∈Uj(see 8.23). Using this decomposition, define an
operatorS∈L(V)by
Sv=S1u1+···+Smum.
You should verify that this operator Sis a square root of T, completing
the proof.
By imitating the techniques in this section, you should be able to
prove that if Vis a complex vector space and T∈L(V)is invertible,
thenThas akth-root for every positive integer k.
The Minimal Polynomial 179
The Minimal Polynomial
As we will soon see, given an operator on a finite-dimensional vec-
tor space, there is a unique monic polynomial of smallest degree that Amonic polynomial is
a polynomial whose
highest degree
coefficient equals 1.
For example,
2+3z2+z8is a monic
polynomial.when applied to the operator gives 0. This polynomial is called the
minimal polynomial of the operator and is the focus of attention inthis section.
SupposeT∈L(V), where dim V=n. Then
(I,T,T
2,...,Tn2)
cannot be linearly independent in L(V) becauseL(V) has dimension n2
(see 3.20) and we have n2+1 operators. Let mbe the smallest positive
integer such that
8.33 (I,T,T2,...,Tm)
is linearly dependent. The linear dependence lemma (2.4) implies that
one of the operators in the list above is a linear combination of theprevious ones. Because mwas chosen to be the smallest positive in-
teger such that 8.33 is linearly dependent, we conclude that T
mis
a linear combination of (I,T,T2,...,Tm−1). Thus there exist scalars
a0,a1,a2,...,am−1∈Fsuch that
a0I+a1T+a2T2+···+am−1Tm−1+Tm=0.
The choice of scalars a0,a1,a2,...,am−1∈Fabove is unique because
two different such choices would contradict our choice of m(subtract-
ing two different equations of the form above, we would have a linearlydependent list shorter than 8.33). The polynomial
a
0+a1z+a2z2+···+am−1zm−1+zm
is called the minimal polynomial ofT. It is the monic polynomial
p∈P(F)of smallest degree such that p(T)=0.
For example, the minimal polynomial of the identity operator Iis
z−1. The minimal polynomial of the operator on F2whose matrix
equals/bracketleftBig
41
05/bracketrightBig
is 20−9z+z2, as you should verify.
Clearly the degree of the minimal polynomial of each operator on V
is at most(dimV)2. The Cayley-Hamilton theorem (8.20) tells us that
ifVis a complex vector space, then the minimal polynomial of each
operator onVhas degree at most dim V. This remarkable improvement
also holds on real vector spaces, as we will see in the next chapter.
180 Chapter 8.Operators on Complex Vector Spaces
A polynomial p∈P(F)is said to divide a polynomial q∈P(F)if
there exists a polynomial s∈P(F)such thatq=sp. In other words,
pdividesqif we can take the remainder rin 4.6 to be 0. For exam- Note that(z−λ)
divides a polynomial q
if and only if λis a
root ofq. This follows
immediately from 4.1.ple, the polynomial (1+3z)2divides 5+32z+57z2+18z3because
5+32z+57z2+18z3=(2z+5)(1+3z)2. Obviously every nonzero
constant polynomial divides every polynomial.
The next result completely characterizes the polynomials that when
applied to an operator give the 0 operator.
8.34 Theorem: LetT∈L(V) and letq∈P(F). Thenq(T)=0if
and only if the minimal polynomial of Tdividesq.
Proof: Letpdenote the minimal polynomial of T.
First we prove the easy direction. Suppose that pdividesq. Thus
there exists a polynomial s∈P(F)such thatq=sp. We have
q(T)=s(T)p(T)=s(T)0=0,
as desired.
To prove the other direction, suppose that q(T)=0. By the division
algorithm (4.5), there exist polynomials s,r∈P(F)such that
8.35 q=sp+r
and degr<degp. We have
0=q(T)=s(T)p(T)+r(T)=r(T).
Becausepis the minimal polynomial of Tand degr<degp, the equa-
tion above implies that r=0. Thus 8.35 becomes the equation q=sp,
and hencepdividesq, as desired.
Now we describe the eigenvalues of an operator in terms of its min-
imal polynomial.
8.36 Theorem: LetT∈L(V). Then the roots of the minimal poly-
nomial ofTare precisely the eigenvalues of T.
The Minimal Polynomial 181
Proof: Let
p(z)=a0+a1z+a2z2+···+am−1zm−1+zm
be the minimal polynomial of T.
First suppose that λ∈Fis a root ofp. Then the minimal polynomial
ofTcan be written in the form
p(z)=(z−λ)q(z),
whereqis a monic polynomial with coefficients in F(see 4.1). Because
p(T)=0, we have
0=(T−λI)(q(T)v)
for allv∈V. Because the degree of qis less than the degree of the
minimal polynomial p, there must exist at least one vector v∈Vsuch
thatq(T)v/negationslash=0. The equation above thus implies that λis an eigenvalue
ofT, as desired.
To prove the other direction, now suppose that λ∈Fis an eigen-
value ofT. Letvbe a nonzero vector in Vsuch thatTv=λv. Repeated
applications of Tto both sides of this equation show that Tjv=λjv
for every nonnegative integer j. Thus
0=p(T)v=(a0+a1T+a2T2+···+am−1Tm−1+Tm)v
=(a0+a1λ+a2λ2+···+am−1λm−1+λm)v
=p(λ)v.
Becausev/negationslash=0, the equation above implies that p(λ)=0, as desired.
Suppose we are given, in concrete form, the matrix (with respect to
some basis) of some operator T∈L(V). To find the minimal polyno-
mial ofT, consider
(M(I),M(T),M(T)2,...,M(T)m)
form=1,2,... until this list is linearly dependent. Then find the
scalarsa0,a1,a2,...,am−1∈Fsuch that You can think of this as
a system of (dimV)2
equations in m
variables
a0,a1,...,am−1.a0M(I)+a1M(T)+a2M(T)2+···+am−1M(T)m−1+M(T)m=0.
The scalars a0,a1,a2,...,am−1,1 will then be the coefficients of the
minimal polynomial of T. All this can be computed using a familiar
process such as Gaussian elimination.
182 Chapter 8.Operators on Complex Vector Spaces
For example, consider the operator TonC5whose matrix is given
by
8.37
0000−3
1000 60100 00010 0
0001 0
.
Because of the large number of 0’s in this matrix, Gaussian elimination
is not needed here. Simply compute powers of M(T) and notice that
there is no linear dependence until the fifth power. Do the computa-tions and you will see that the minimal polynomial of Tequals
8.38 z
5−6z+3.
Now what about the eigenvalues of this particular operator? From 8.36,
we see that the eigenvalues of Tequal the solutions to the equation
z5−6z+3=0.
Unfortunately no solution to this equation can be computed using ra-
tional numbers, arbitrary roots of rational numbers, and the usual rules
of arithmetic (a proof of this would take us considerably beyond linear
algebra). Thus we cannot find an exact expression for any eigenvaluesofTin any familiar form, though numeric techniques can give good ap-
proximations for the eigenvalues of T. The numeric techniques, which
we will not discuss here, show that the eigenvalues for this particularoperator are approximately
−1.67, 0.51, 1.40,−0.12+1.59i,−0.12−1.59i.
Note that the nonreal eigenvalues occur as a pair, with each the complex
conjugate of the other, as expected for the roots of a polynomial with
real coefficients (see 4.10).
SupposeVis a complex vector space and T∈L(V). The Cayley-
Hamilton theorem (8.20) and 8.34 imply that the minimal polynomialofTdivides the characteristic polynomial of T. Both these polynomials
are monic. Thus if the minimal polynomial of Thas degree dim V, then
it must equal the characteristic polynomial of T. For example, if Tis
the operator on C
5whose matrix is given by 8.37, then the character-
istic polynomial of T, as well as the minimal polynomial of T, is given
by 8.38.
Jordan Form 183
Jordan Form
We know that if Vis a complex vector space, then for every T∈L(V)
there is a basis of Vwith respect to which Thas a nice upper-triangular
matrix (see 8.28). In this section we will see that we can do even better—there is a basis of Vwith respect to which the matrix of Tcontains zeros
everywhere except possibly on the diagonal and the line directly abovethe diagonal.
We begin by describing the nilpotent operators. Consider, for ex-
ample, the nilpotent operator N∈L(F
n)defined by
N(z 1,...,zn)=(0,z 1,...,zn−1).
Ifv=(1,0,...,0), then clearly (v,Nv,...,Nn−1v)is a basis of Fnand
(Nn−1v)is a basis of null N, which has dimension 1.
As another example, consider the nilpotent operator N∈L(F5)de-
fined by
8.39 N(z 1,z2,z3,z4,z5)=(0,z 1,z2,0,z 4).
Unlike the nilpotent operator discussed in the previous paragraph, for
this nilpotent operator there does not exist a vector v∈F5such that
(v,Nv,N2v,N3v,N4v)is a basis of F5. However, if v1=(1,0,0,0,0)
andv2=(0,0,0,1,0), then(v1,Nv 1,N2v1,v2,Nv 2)is a basis of F5
and(N2v1,Nv 2)is a basis of null N, which has dimension 2.
SupposeN∈L(V)is nilpotent. For each nonzero vector v∈V, let
m(v) denote the largest nonnegative integer such that Nm(v)v/negationslash=0. For Obviouslym(v)
depends on Nas well
as onv, but the choice
ofNwill be clear from
the context.example, ifN∈L(F5)is defined by 8.39, then m(1,0,0,0,0)=2.
The lemma below shows that every nilpotent operator N∈L(V)
behaves similarly to the example defined by 8.39, in the sense that thereis a finite collection of vectors v
1,...,vk∈Vsuch that the nonzero
vectors of the form Njvrform a basis of V; herervaries from 1 to k
andjvaries from 0 to m(vr).
8.40 Lemma: IfN∈L(V) is nilpotent, then there exist vectors
v1,...,vk∈Vsuch that
(a)(v1,Nv 1,...,Nm(v 1)v1,...,vk,Nvk,...,Nm(vk)vk)is a basis ofV;
(b)(Nm(v 1)v1,...,Nm(vk)vk)is a basis of nullN.
184 Chapter 8.Operators on Complex Vector Spaces
Proof: SupposeNis nilpotent. Then Nis not injective and thus
dim rangeN< dimV(see 3.21). By induction on the dimension of V,
we can assume that the lemma holds on all vector spaces of smaller
dimension. Using range Nin place ofVandN|rangeNin place ofN,w e
thus have vectors u1,...,uj∈rangeNsuch that
(i)(u1,Nu 1,...,Nm(u 1)u1,...,uj,Nuj,...,Nm(uj)uj)is a basis of
rangeN;
(ii)(Nm(u 1)u1,...,Nm(uj)uj)is a basis of null N∩rangeN.
Because each ur∈rangeN, we can choose v1,...,vj∈Vsuch that
Nvr=urfor eachr. Note thatm(vr)=m(ur)+1 for eachr.
LetWbe a subspace of null Nsuch that The existence of a
subspaceWwith this
property follows from
2.13.8.41 nullN=(nullN∩rangeN)⊕W
and choose a basis of W, which we will label (vj+1,...,vk). Because
vj+1,...,vk∈nullN, we havem(vj+1)=···=m(v k)=0.
Having constructed v1,...,vk, we now need to show that (a) and
(b) hold. We begin by showing that the alleged basis in (a) is linearly
independent. To do this, suppose
8.42 0=k/summationdisplay
r=1m(vr)/summationdisplay
s=0ar,sNs(vr),
where eachar,s∈F. ApplyingNto both sides of the equation above,
we get
0=k/summationdisplay
r=1m(vr)/summationdisplay
s=0ar,sNs+1(vr)
=j/summationdisplay
r=1m(ur)/summationdisplay
s=0ar,sNs(ur).
The last equation, along with (i), implies that ar,s=0 for 1≤r≤j,
0≤s≤m(vr)−1. Thus 8.42 reduces to the equation
0=a1,m(v 1)Nm(v 1)v1+···+aj,m(vj)Nm(vj)vj
+aj+1,0vj+1+···+ak,0vk.
Jordan Form 185
The terms on the first line on the right are all in null N∩rangeN; the
terms on the second line are all in W. Thus the last equation and 8.41
imply that
0=a1,m(v 1)Nm(v 1)v1+···+aj,m(vj)Nm(vj)vj
=a1,m(v 1)Nm(u 1)u1+···+aj,m(vj)Nm(uj)uj 8.43
and8.44 0=a
j+1,0vj+1+···+ak,0vk.
Now 8.43 and (ii) imply that a1,m(v 1)=···=aj,m(vj)=0. Because
(vj+1,...,vk)is a basis ofW, 8.44 implies that aj+1,0=···=a k,0=0.
Thus all the a’s equal 0, and hence the list of vectors in (a) is linearly
independent.
Clearly (ii) implies that dim(null N∩rangeN)=j. Along with 8.41,
this implies that
8.45 dim nullN=k.
Clearly (i) implies that
dim rangeN=j/summationdisplay
r=0(m(ur)+1)
=j/summationdisplay
r=0m(vr). 8.46
The list of vectors in (a) has length
k/summationdisplay
r=0(m(vr)+1)=k+j/summationdisplay
r=0m(vr)
=dim nullN+dim rangeN
=dimV,
where the second equality comes from 8.45 and 8.46, and the third
equality comes from 3.4. The last equation shows that the list of vectors
in (a) has length dim V; because this list is linearly independent, it is a
basis ofV(see 2.17), completing the proof of (a).
Finally, note that
(Nm(v 1)v1,...,Nm(vk)vk)=(Nm(u 1)u1,...,Nm(uj)uj,vj+1,...,vk).
186 Chapter 8.Operators on Complex Vector Spaces
Now (ii) and 8.41 show that the last list above is a basis of null N, com-
pleting the proof of (b).
SupposeT∈L(V). A basis of Vis called a Jordan basis forTif
with respect to this basis Thas a block diagonal matrix
A1 0
...
0Am
,
where eachAjis an upper-triangular matrix of the form
Aj=
λ
j10
......
...1
0 λj
.
In eachA
j, the diagonal is filled with some eigenvalue λjofT, the line To understand why
eachλjmust be an
eigenvalue of T,
see 5.18.directly above the diagonal is filled with 1’s, and all other entries are 0
(Ajmay be just a 1-by-1 block consisting of just some eigenvalue).
Because there exist operators on real vector spaces that have no
eigenvalues, there exist operators on real vector spaces for which there
is no corresponding Jordan basis. Thus the hypothesis that Vis a com-
plex vector space is required for the next result, even though the pre-vious lemma holds on both real and complex vector spaces.
8.47 Theorem: SupposeVis a complex vector space. If T∈L(V),
The French
mathematician Camille
Jordan first published a
proof of this theorem
in 1870.then there is a basis of Vthat is a Jordan basis for T.
Proof: First consider a nilpotent operator N∈L(V)and the vec-
torsv1,...,vk∈Vgiven by 8.40. For each j, note thatNsends the first
vector in the list (Nm(vj)vj,...,Nvj,vj)to 0 and that Nsends each vec-
tor in this list other than the first vector to the previous vector. In otherwords, if we reverse the order of the basis given by 8.40(a), then we ob-tain a basis of Vwith respect to which Nhas a block diagonal matrix,
where each matrix on the diagonal has the form
01 0
......
...1
00
.
Jordan Form 187
Thus the theorem holds for nilpotent operators.
Now suppose T∈L(V). Letλ1,...,λmbe the distinct eigenval-
ues ofT, withU1,...,Umthe corresponding subspaces of generalized
eigenvectors. We have
V=U1⊕···⊕Um,
where each(T−λjI)|Ujis nilpotent (see 8.23). By the previous para-
graph, there is a basis of each Ujthat is a Jordan basis for (T−λjI)|Uj.
Putting these bases together gives a basis of Vthat is a Jordan basis
forT.
188 Chapter 8.Operators on Complex Vector Spaces
Exercises
1. Define T∈L(C2)by
T(w,z)=(z,0).
Find all generalized eigenvectors of T.
2. Define T∈L(C2)by
T(w,z)=(−z,w).
Find all generalized eigenvectors of T.
3. Suppose T∈L(V),mis a positive integer, and v∈Vis such
thatTm−1v/negationslash=0 butTmv=0. Prove that
(v,Tv,T2v,...,Tm−1v)
is linearly independent.
4. Suppose T∈L(C3)is defined by T(z 1,z2,z3)=(z2,z3,0). Prove
thatThas no square root. More precisely, prove that there does
not existS∈L(C3)such thatS2=T.
5. Suppose S,T∈L(V). Prove that if STis nilpotent, then TSis
nilpotent.
6. Suppose N∈L(V)is nilpotent. Prove (without using 8.26) that
0 is the only eigenvalue of N.
7. Suppose Vis an inner-product space. Prove that if N∈L(V)is
self-adjoint and nilpotent, then N=0.
8. Suppose N∈L(V)is such that null NdimV−1/negationslash=nullNdimV. Prove
thatNis nilpotent and that
dim nullNj=j
for every integer jwith 0≤j≤dimV.
9. Suppose T∈L(V)andmis a nonnegative integer such that
rangeTm=rangeTm+1.
Prove that range Tk=rangeTmfor allk>m .
Exercises 189
10. Prove or give a counterexample: if T∈L(V), then
V=nullT⊕rangeT.
11. Prove that if T∈L(V), then
V=nullTn⊕rangeTn,
wheren=dimV.
12. Suppose Vis a complex vector space, N∈L(V), and 0 is the only
eigenvalue of N. Prove that Nis nilpotent. Give an example to
show that this is not necessarily true on a real vector space.
13. Suppose that Vis a complex vector space with dim V=nand
T∈L(V)is such that
nullTn−2/negationslash=nullTn−1.
Prove thatThas at most two distinct eigenvalues.
14. Give an example of an operator on C4whose characteristic poly-
nomial equals (z−7)2(z−8)2.
15. Suppose Vis a complex vector space. Suppose T∈L(V)is such
that 5 and 6 are eigenvalues of Tand thatThas no other eigen-
values. Prove that
(T−5I)n−1(T−6I)n−1=0,
wheren=dimV.
16. Suppose Vis a complex vector space and T∈L(V). Prove that For complex vector
spaces, this exerciseadds another
equivalence to the list
given by 5.21.Vhas a basis consisting of eigenvectors of Tif and only if every
generalized eigenvector of Tis an eigenvector of T.
17. Suppose Vis an inner-product space and N∈L(V)is nilpotent.
Prove that there exists an orthonormal basis of Vwith respect to
whichNhas an upper-triangular matrix.
18. Define N∈L(F5)by
N(x 1,x2,x3,x4,x5)=(2x 2,3x3,−x4,4x5,0).
Find a square root of I+N.
190 Chapter 8.Operators on Complex Vector Spaces
19. Prove that if Vis a complex vector space, then every invertible
operator on Vhas a cube root.
20. Suppose T∈L(V)is invertible. Prove that there exists a polyno-
mialp∈P(F)such thatT−1=p(T) .
21. Give an example of an operator on C3whose minimal polynomial
equalsz2.
22. Give an example of an operator on C4whose minimal polynomial
equalsz(z−1)2.
23. Suppose Vis a complex vector space and T∈L(V). Prove that For complex vector
spaces, this exercise
adds another
equivalence to the list
given by 5.21.Vhas a basis consisting of eigenvectors of Tif and only if the
minimal polynomial of Thas no repeated roots.
24. Suppose Vis an inner-product space. Prove that if T∈L(V)is
normal, then the minimal polynomial of Thas no repeated roots.
25. Suppose T∈L(V)andv∈V. Letpbe the monic polynomial of
smallest degree such that
p(T)v=0.
Prove thatpdivides the minimal polynomial of T.
26. Give an example of an operator on C4whose characteristic and
minimal polynomials both equal z(z−1)2(z−3).
27. Give an example of an operator on C4whose characteristic poly-
nomial equals z(z−1)2(z−3)and whose minimal polynomial
equalsz(z−1)(z−3).
28. Suppose a0,...,an−1∈C. Find the minimal and characteristic This exercise shows
that every monic
polynomial is the
characteristic
polynomial of some
operator.polynomials of the operator on Cnwhose matrix (with respect to
the standard basis) is
0 −a
0
10 −a1
1...−a2
......
0−an−2
1−an−1
.
Exercises 191
29. Suppose N∈L(V)is nilpotent. Prove that the minimal poly-
nomial ofNiszm+1, wheremis the length of the longest con-
secutive string of 1’s that appears on the line directly above thediagonal in the matrix of Nwith respect to any Jordan basis for N.
30. Suppose Vis a complex vector space and T∈L(V). Prove that
there does not exist a direct sum decomposition of Vinto two
proper subspaces invariant under Tif and only if the minimal
polynomial of Tis of the form (z−λ)
dimVfor someλ∈C.
31. Suppose T∈L(V)and(v1,...,vn)is a basis ofVthat is a Jordan
basis forT. Describe the matrix of Twith respect to the basis
(vn,...,v 1)obtained by reversing the order of the v’s.
Chapter 9
Operators on
Real Vector Spaces
In this chapter we delve deeper into the structure of operators on
real vector spaces. The important results here are somewhat more com-plex than the analogous results from the last chapter on complex vectorspaces.
Recall that Fdenotes RorC.
Also,Vis a finite-dimensional, nonzero vector space over F.
Some of the new results in this chapter are valid on complex vector
spaces, so we have not assumed that Vis a real vector space.
✽✽✽✽✽✽✽✽✽
193
194 Chapter 9.Operators on Real Vector Spaces
Eigenvalues of Square Matrices
We have defined eigenvalues of operators; now we need to extend
that notion to square matrices. Suppose Ais ann-by-n matrix with
entries in F. A number λ∈Fis called an eigenvalue ofAif there
exists a nonzero n-by-1 matrix xsuch that
Ax=λx.
For example, 3 is an eigenvalue of/bracketleftBig
78
15/bracketrightBig
because
/bracketleftBigg
78
15/bracketrightBigg/bracketleftBigg
2
−1/bracketrightBigg
=/bracketleftBigg
6
−3/bracketrightBigg
=3/bracketleftBigg
2
−1/bracketrightBigg
.
As another example, you should verify that the matrix/bracketleftBig
0−1
10/bracketrightBig
has no
eigenvalues if we are thinking of Fas the real numbers (by definition,
an eigenvalue must be in F) and has eigenvalues iand−iif we are
thinking of Fas the complex numbers.
We now have two notions of eigenvalue—one for operators and one
for square matrices. As you might expect, these two notions are closelyconnected, as we now show.
9.1 Proposition: SupposeT∈L(V) andAis the matrix of Twith
respect to some basis of V. Then the eigenvalues of Tare the same as
the eigenvalues of A.
Proof: Let(v
1,...,vn)be the basis of Vwith respect to which T
has matrixA. Letλ∈F. We need to show that λis an eigenvalue of T
if and only if λis an eigenvalue of A.
First suppose λis an eigenvalue of T. Letv∈Vbe a nonzero vector
such thatTv=λv. We can write
9.2 v=a1v1+···+anvn,
wherea1,...,an∈F. Letxbe the matrix of the vector vwith respect
to the basis(v1,...,vn). Recall from Chapter 3 that this means
9.3 x=
a1
...
an
.
Block Upper-Triangular Matrices 195
We have
Ax=M(T)M(v)=M(Tv)=M(λv)=λM(v)=λx,
where the second equality comes from 3.14. The equation above shows
thatλis an eigenvalue of A, as desired.
To prove the implication in the other direction, now suppose λis an
eigenvalue of A. Letxbe a nonzero n-by-1 matrix such that Ax=λx.
We can write xin the form 9.3 for some scalars a1,...,an∈F. Define
v∈Vby 9.2. Then
M(Tv)=M(T)M(v)=Ax=λx=M(λv).
where the first equality comes from 3.14. The equation above implies
thatTv=λv, and thusλis an eigenvalue of T, completing the proof.
Because every square matrix is the matrix of some operator, the
proposition above allows us to translate results about eigenvalues ofoperators into the language of eigenvalues of square matrices. Forexample, every square matrix of complex numbers has an eigenvalue(from 5.10). As another example, every n-by-n matrix has at most n
distinct eigenvalues (from 5.9).
Block Upper-Triangular Matrices
Earlier we proved that each operator on a complex vector space has
an upper-triangular matrix with respect to some basis (see 5.13). Inthis section we will see that we can almost do as well on real vectorspaces.
In the last two chapters we used block diagonal matrices, which
extend the notion of diagonal matrices. Now we will need to use thecorresponding extension of upper-triangular matrices. A block upper-
triangular matrix is a square matrix of the form
As usual, we use an
asterisk to denote
entries of the matrix
that play no importantrole in the topics under
consideration.
A1∗
...
0Am
,
whereA1,...,Amare square matrices lying along the diagonal, all en-
tries belowA1,...,Amequal 0, and the ∗denotes arbitrary entries. For
example, the matrix
196 Chapter 9.Operators on Real Vector Spaces
A=
41 01 11 21 3
0−3−31 42 5
0−3−31 61 7
00 0 5 500 0 5 5
is a block upper-triangular matrix with
A=
A
1∗
A2
0A3
,
where
A1=/bracketleftBig
4/bracketrightBig
,A 2=/bracketleftBigg
−3−3
−3−3/bracketrightBigg
,A 3=/bracketleftBigg
55
55/bracketrightBigg
.
Now we prove that for each operator on a real vector space, we can Every upper-triangular
matrix is also a block
upper-triangular matrix
with blocks of size
1-by-1 along the
diagonal. At the other
extreme, every square
matrix is a block
upper-triangular matrix
because we can take
the first (and only)
block to be the entire
matrix. Smaller blocks
are better in the sense
that the matrix then
has more 0’s.find a basis that gives a block upper-triangular matrix with blocks of
size at most 2-by-2 on the diagonal.
9.4 Theorem: SupposeVis a real vector space and T∈L(V).
Then there is a basis of Vwith respect to which Thas a block upper-
triangular matrix
9.5
A1∗
...
0Am
,
where eachAjis a1-by-1 matrix or a 2-by-2 matrix with no eigenvalues.
Proof: Clearly the desired result holds if dim V=1.
Next, consider the case where dim V=2. IfThas an eigenvalue λ,
then letv1∈Vbe any nonzero eigenvector. Extend (v1)to a basis
(v1,v2)ofV. With respect to this basis, Thas an upper-triangular
matrix of the form /bracketleftBigg
λa
0b/bracketrightBigg
.
In particular, if Thas an eigenvalue, then there is a basis of Vwith
respect to which Thas an upper-triangular matrix. If Thas no eigen-
values, then choose any basis (v1,v2)ofV. With respect to this basis,
Block Upper-Triangular Matrices 197
the matrix of Thas no eigenvalues (by 9.1). Thus regardless of whether
Thas eigenvalues, we have the desired conclusion when dim V=2.
Suppose now that dim V>2 and the desired result holds for all real
vector spaces with smaller dimension. If Thas an eigenvalue, let Ube a
one-dimensional subspace of Vthat is invariant under T; otherwise let
Ube a two-dimensional subspace of Vthat is invariant under T(5.24
guarantees that we can choose Uin this fashion). Choose any basis
ofUand letA1denote the matrix of T|Uwith respect to this basis. If
A1is a 2-by-2 matrix, then Thas no eigenvalues (otherwise we would
have chosen Uto be one-dimensional) and thus T|Uhas no eigenvalues.
Hence ifA1is a 2-by-2 matrix, then A1has no eigenvalues (see 9.1).
LetWbe any subspace of Vsuch that
V=U⊕W;
2.13 guarantees that such a Wexists. Because Whas dimension less
than the dimension of V, we would like to apply our induction hypoth-
esis toT|W. However,Wmight not be invariant under T, meaning that
T|Wmight not be an operator on W. We will compose with the pro-
jectionPW,Uto get an operator on W. Specifically, define S∈L(W) Recall that if
v=w+u, where
w∈Wandu∈U,
thenPW,Uv=w.by
Sw=PW,U(Tw)
forw∈W. Note that
Tw=PU,W(Tw)+PW,U(Tw)
=PU,W(Tw)+Sw 9.6
for everyw∈W.
By our induction hypothesis, there is a basis of Wwith respect to
whichShas a block upper-triangular matrix of the form
A2∗
...
0Am
,
where eachAjis a 1-by-1 matrix or a 2-by-2 matrix with no eigenvalues.
Adjoin this basis of Wto the basis of Uchosen above, getting a basis
ofV. A minute’s thought should convince you (use 9.6) that the matrix
ofTwith respect to this basis is a block upper-triangular matrix of the
form 9.5, completing the proof.
198 Chapter 9.Operators on Real Vector Spaces
The Characteristic Polynomial
For operators on complex vector spaces, we defined characteristic
polynomials and developed their properties by making use of upper-
triangular matrices. In this section we will carry out a similar procedure
for operators on real vector spaces. Instead of upper-triangular matri-ces, we will have to use the block upper-triangular matrices furnishedby the last theorem.
In the last chapter, we did not define the characteristic polynomial
of a square matrix with complex entries because our emphasis is onoperators rather than on matrices. However, to understand operatorson real vector spaces, we will need to define characteristic polynomialsof 1-by-1 and 2-by-2 matrices with real entries. Then, using block-upper
triangular matrices with blocks of size at most 2-by-2 on the diagonal,
we will be able to define the characteristic polynomial of an operatoron a real vector space.
To motivate the definition of characteristic polynomials of square
matrices, we would like the following to be true (think about the Cayley-Hamilton theorem; see 8.20): if T∈L(V)has matrixAwith respect
to some basis of Vandqis the characteristic polynomial of A, then
q(T)=0.
Let’s begin with the trivial case of 1-by-1 matrices. Suppose Vis a
real vector space with dimension 1 and T∈L(V).I f[λ]equals the
matrix ofTwith respect to some basis of V, thenTequalsλI. Thus
if we letqbe the degree 1 polynomial defined by q(x)=x−λ, then
q(T)=0. Hence we define the characteristic polynomial of [λ]to be
x−λ.
Now let’s look at 2-by-2 matrices with real entries. Suppose Vis a
real vector space with dimension 2 and T∈L(V). Suppose
/bracketleftBigg
ac
bd/bracketrightBigg
is the matrix of Twith respect to some basis (v
1,v2)ofV. We seek
a monic polynomial qof degree 2 such that q(T)=0. Ifb=0, then
the matrix above is upper triangular. If in addition we were dealing
with a complex vector space, then we would know that Thas charac-
teristic polynomial (z−a)(z−d). Thus a reasonable candidate might
be(x−a)(x−d), where we use xinstead ofzto emphasize that
now we are working on a real vector space. Let’s see if the polynomial
The Characteristic Polynomial 199
(x−a)(x−d), when applied to T, gives 0 even when b/negationslash=0. We have
(T−aI)(T−dI)v 1=(T−dI)(T−aI)v 1=(T−dI)(bv 2)=bcv 1
and
(T−aI)(T−dI)v 2=(T−aI)(cv 1)=bcv 2.
Thus(T−aI)(T−dI)is not equal to 0 unless bc=0. However, the
equations above show that (T−aI)(T−dI)−bcI=0 (because this
operator equals 0 on a basis, it must equal 0 on V). Thus ifq(x)=
(x−a)(x−d)−bc, thenq(T)=0.
Motivated by the previous paragraph, we define the characteristic
polynomial of a 2-by-2 matrix/bracketleftbigac
bd/bracketrightbig
to be(x−a)(x−d)−bc. Here
we are concerned only with matrices with real entries. The next re-sult shows that we have found the only reasonable definition for thecharacteristic polynomial of a 2-by-2 matrix.
9.7 Proposition: SupposeVis a real vector space with dimension 2
Part (b) of this
proposition would befalse without the
hypothesis that Thas
no eigenvalues. For
example, defineT∈L(R
2)by
T(x 1,x2)=(0,x 2).
Takep(x)=x(x−2).
Thenpis not the
characteristic
polynomial of the
matrix ofTwith
respect to the standard
basis, butp(T) is not
invertible.andT∈L(V) has no eigenvalues. Let p∈P(R)be a monic polynomial
with degree 2. SupposeAis the matrix of Twith respect to some basis
ofV.
(a) Ifpequals the characteristic polynomial of A, thenp(T)=0.
(b) Ifpdoes not equal the characteristic polynomial of A, thenp(T)
is invertible.
Proof: We already proved (a) in our discussion above. To prove (b),
letqdenote the characteristic polynomial of Aand suppose that p/negationslash=q.
We can write p(x)=x2+α1x+β1andq(x)=x2+α2x+β2for some
α1,β1,α2,β2∈R. Now
p(T)=p(T)−q(T)=(α1−α2)T+(β1−β2)I.
Ifα1=α2, thenβ1/negationslash=β2(otherwise we would have p=q). Thus if
α1=α2, thenp(T) is a nonzero multiple of the identity and hence is
invertible, as desired. If α1/negationslash=α2, then
p(T)=(α1−α2)(T−β2−β1
α1−α2I),
which is an invertible operator because Thas no eigenvalues. Thus (b)
holds.
200 Chapter 9.Operators on Real Vector Spaces
SupposeVis a real vector space with dimension 2 and T∈L(V)has
no eigenvalues. The last proposition shows that there is precisely onemonic polynomial with degree 2 that when applied to Tgives 0. Thus,
thoughTmay have different matrices with respect to different bases,
each of these matrices must have the same characteristic polynomial.For example, consider T∈L(R
2)defined by
9.8 T(x 1,x2)=(3x 1+5x2,−2x 1−x2).
The matrix of Twith respect to the standard basis of R2is
/bracketleftBigg
35
−2−1/bracketrightBigg
.
The characteristic polynomial of this matrix is (x−3)(x+1)+2·5,
which equals x2−2x+7. As you should verify, the matrix of Twith
respect to the basis/parenleftbig
(−2, 1),(1, 2)/parenrightbig
equals
/bracketleftBigg
1−6
11/bracketrightBigg
.
The characteristic polynomial of this matrix is (x−1)(x−1)+1·6,
which equals x2−2x+7, the same result we obtained by using the
standard basis.
When analyzing upper-triangular matrices of an operator Ton a
complex vector space V, we found that subspaces of the form
null(T−λI)dimV
played a key role (see 8.10). Those spaces will also play a role in study-
ing operators on real vector spaces, but because we must now consider
block upper-triangular matrices with 2-by-2 blocks, subspaces of the
form
null(T2+αT+βI)dimV
will also play a key role. To get started, let’s look at one- and two-
dimensional real vector spaces.
First suppose that Vis a one-dimensional real vector space and that
T∈L(V).I fλ∈R, then null(T−λI)equalsVifλis an eigenvalue
ofTand{0}otherwise. If α,β∈Rwithα2<4β, then
The Characteristic Polynomial 201
null(T2+αT+βI)={0}.
(Proof: Because Vis one-dimensional, there is a constant λ∈Rsuch Recall thatα2<4β
implies that
x2+αx+βhas no real
roots; see 4.11.thatTv=λvfor allv∈V. Thus(T2+αT+βI)v=(λ2+αλ+β)v.
However, the inequality α2<4βimplies that λ2+αλ+β/negationslash=0, and thus
null(T2+αT+βI)={0}.)
Now suppose Vis a two-dimensional real vector space and T∈L(V)
has no eigenvalues. If λ∈R, then null(T−λI)equals{0}(becauseT
has no eigenvalues). If α,β∈Rwithα2<4β, then null (T2+αT+βI)
equalsVifx2+αx+βis the characteristic polynomial of the matrix
ofTwith respect to some (or equivalently, every) basis of Vand equals
{0}otherwise (by 9.7). Note that for this operator, there is no middle
ground—the null space of T2+αT+βIis either{0}or the whole space;
it cannot be one-dimensional.
Now suppose that Vis a real vector space of any dimension and
T∈L(V). We know that Vhas a basis with respect to which Thas
a block upper-triangular matrix with blocks on the diagonal of size at
most 2-by-2 (see 9.4). In general, this matrix is not unique— Vmay
have many different bases with respect to which Thas a block upper-
triangular matrix of this form, and with respect to these different baseswe may get different block upper-triangular matrices.
We encountered a similar situation when dealing with complex vec-
tor spaces and upper-triangular matrices. In that case, though we mightget different upper-triangular matrices with respect to the differentbases, the entries on the diagonal were always the same (though possi-bly in a different order). Might a similar property hold for real vectorspaces and block upper-triangular matrices? Specifically, is the num-ber of times a given 2-by-2 matrix appears on the diagonal of a block
upper-triangular matrix of Tindependent of which basis is chosen?
Unfortunately this question has a negative answer. For example, theoperatorT∈L(R
2)defined by 9.8 has two different 2-by-2 matrices,
as we saw above.
Though the number of times a particular 2-by-2 matrix might appear
on the diagonal of a block upper-triangular matrix of Tcan depend on
the choice of basis, if we look at characteristic polynomials insteadof the actual matrices, we find that the number of times a particularcharacteristic polynomial appears is independent of the choice of basis.This is the content of the following theorem, which will be our key tool
in analyzing the structure of an operator on a real vector space.
202 Chapter 9.Operators on Real Vector Spaces
9.9 Theorem: SupposeVis a real vector space and T∈L(V).
Suppose that with respect to some basis of V, the matrix of Tis
9.10
A1∗
...
0Am
,
where eachAjis a1-by-1 matrix or a 2-by-2 matrix with no eigenvalues.
(a) Ifλ∈R, then precisely dim null(T−λI)dimVof the matrices
A1,...,Amequal the 1-by-1 matrix[λ].
(b) Ifα,β∈Rsatisfyα2<4β, then precisely This result implies that
null(T2+αT+βI)dimV
must have even
dimension.dim null(T2+αT+βI)dimV
2
of the matrices A1,...,Amhave characteristic polynomial equal
tox2+αx+β.
Proof: We will construct one proof that can be used to prove both This proof uses the
same ideas as the proof
of the analogous result
on complex vector
spaces (8.10). As usual,
the real case is slightly
more complicated but
requires no new
creativity.(a) and (b). To do this, let λ,α,β∈Rwithα2<4β. Definep∈P(R)
by
p(x)=/braceleftBigg
x−λ if we are trying to prove (a);
x2+αx+βif we are trying to prove (b).
Letddenote the degree of p. Thusd=1 if we are trying to prove (a)
andd=2 if we are trying to prove (b).
We will prove this theorem by induction on m, the number of blocks
along the diagonal of 9.10. If m=1, then dimV=1 or dimV=2; the
discussion preceding this theorem then implies that the desired resultholds. Thus we can assume that m> 1 and that the desired result
holds whenmis replaced with m−1.
For convenience let n=dimV. Consider a basis of Vwith respect
to whichThas the block upper-triangular matrix 9.10. Let U
jdenote
the span of the basis vectors corresponding to Aj. Thus dimUj=1
ifAjis a 1-by-1 matrix and dim Uj=2i fAjis a 2-by-2 matrix. Let
U=U1+···+Um−1. ClearlyUis invariant under Tand the matrix
ofT|Uwith respect to the obvious basis (obtained from the basis vec-
tors corresponding to A1,...,Am−1)i s
9.11
A1∗
...
0Am−1
.
The Characteristic Polynomial 203
Thus, by our induction hypothesis,
9.12precisely(1/d) dim nullp(T|U)nof the matrices
A1,...,Am−1have characteristic polynomial p.
Actually the induction hypothesis gives 9.12 with exponent dim Uin-
stead ofn, but then we can replace dim Uwithn(by 8.6) to get the
statement above.
Supposeum∈Um. LetS∈L(Um)be the operator whose matrix
(with respect to the basis corresponding to Um) equalsAm. In particu-
lar,Sum=PUm,UTum. Now
Tum=PU,UmTum+PUm,UTum
=∗U+Sum,
where∗Udenotes a vector in U. Note thatSum∈Um; thus applying
Tto both sides of the equation above gives
T2um=∗U+S2um,
where again∗Udenotes a vector in U, though perhaps a different vector
than the previous usage of ∗U(the notation ∗Uis used when we want
to emphasize that we have a vector in Ubut we do not care which
particular vector—each time the notation ∗Uis used, it may denote a
different vector in U). The last two equations show that
9.13 p(T)um=∗U+p(S)um
for some∗U∈U. Note that p(S)um∈Um; thus iterating the last
equation gives
9.14 p(T)num=∗U+p(S)num
for some∗U∈U.
The proof now breaks into two cases. First consider the case where
the characteristic polynomial of Amdoes not equal p. We will show
that in this case
9.15 nullp(T)n⊂U.
Once this has been verified, we will know that
nullp(T)n=nullp(T|U)n,
204 Chapter 9.Operators on Real Vector Spaces
and hence 9.12 will tell us that precisely (1/d) dim nullp(T)nof the
matricesA1,...,Amhave characteristic polynomial p, completing the
proof in the case where the characteristic polynomial of Amdoes not
equalp.
To prove 9.15 (still assuming that the characteristic polynomial of
Amdoes not equal p), supposev∈nullp(T)n. We can write vin the
formv=u+um, whereu∈Uandum∈Um. Using 9.14, we have
0=p(T)nv=p(T)nu+p(T)num=p(T)nu+∗U+p(S)num
for some∗U∈U. Because the vectors p(T)nuand∗Uare inUand
p(S)num∈Um, this implies that p(S)num=0. However, p(S) is in-
vertible (see the discussion preceding this theorem about one- and two-dimensional subspaces and note that dim U
m≤2), soum=0. Thus
v=u∈U, completing the proof of 9.15.
Now consider the case where the characteristic polynomial of Am
equalsp. Note that this implies dim Um=d. We will show that
9.16 dim nullp(T)n=dim nullp(T|U)n+d,
which along with 9.12 will complete the proof.
Using the formula for the dimension of the sum of two subspaces
(2.18), we have
dim nullp(T)n=dim(U∩nullp(T)n)+dim(U+nullp(T)n)−dimU
=dim nullp(T|U)n+dim(U+nullp(T)n)−(n−d).
IfU+nullp(T)n=V, then dim(U +nullp(T)n)=n, which when com-
bined with the last formula above for dim null p(T)nwould give 9.16,
as desired. Thus we will finish by showing that U+nullp(T)n=V.
To prove that U+nullp(T)n=V, supposeum∈Um. Because the
characteristic polynomial of the matrix of S(namely,Am) equalsp,w e
havep(S)=0. Thusp(T)um∈U(from 9.13). Now
p(T)num=p(T)n−1(p(T)um)∈rangep(T|U)n−1=rangep(T|U)n,
where the last equality comes from 8.9. Thus we can choose u∈U
such thatp(T)num=p(T|U)nu. Now
p(T)n(um−u)=p(T)num−p(T)nu
=p(T)num−p(T|U)nu
=0.
The Characteristic Polynomial 205
Thusum−u∈nullp(T)n, and henceum, which equals u+(um−u),
is inU+nullp(T)n. In other words, Um⊂U+nullp(T)n. Therefore
V=U+Um⊂U+nullp(T)n, and henceU+nullp(T)n=V, completing
the proof.
As we saw in the last chapter, the eigenvalues of an operator on a
complex vector space provide the key to analyzing the structure of theoperator. On a real vector space, an operator may have fewer eigen-values, counting multiplicity, than the dimension of the vector space.
The previous theorem suggests a definition that makes up for this defi-
ciency. We will see that the definition given in the next paragraph helps
make operator theory on real vector spaces resemble operator theoryon complex vector spaces.
SupposeVis a real vector space and T∈L(V). An ordered pair
(α,β) of real numbers is called an eigenpair ofTifα
2<4βand Though the word
eigenpair was chosen
to be consistent with
the word eigenvalue,
this terminology is not
in widespread use.T2+αT+βI
is not injective. The previous theorem shows that Tcan have only
finitely many eigenpairs because each eigenpair corresponds to thecharacteristic polynomial of a 2-by-2 matrix on the diagonal of 9.10and there is room for only finitely many such matrices along that diag-onal. Guided by 9.9, we define the multiplicity of an eigenpair (α,β)
ofTto be
dim null(T
2+αT+βI)dimV
2.
From 9.9, we see that the multiplicity of (α,β) equals the number of
times thatx2+αx+βis the characteristic polynomial of a 2-by-2 matrix
on the diagonal of 9.10.
As an example, consider the operator T∈L(R3)whose matrix (with
respect to the standard basis) equals
3−1−2
32−3
12 0
.
You should verify that (−4, 13)is an eigenpair of Twith multiplicity 1;
note thatT2−4T+13Iis not injective because (−1, 0,1)and(1,1,0)
are in its null space. Without doing any calculations, you should verifythatThas no other eigenpairs (use 9.9). You should also verify that 1 is
an eigenvalue of Twith multiplicity 1, with corresponding eigenvector
(1,0,1), and that Thas no other eigenvalues.
206 Chapter 9.Operators on Real Vector Spaces
In the example above, the sum of the multiplicities of the eigenval-
ues ofTplus twice the multiplicities of the eigenpairs of Tequals 3,
which is the dimension of the domain of T. The next proposition shows
that this always happens on a real vector space.
9.17 Proposition: IfVis a real vector space and T∈L(V), then This proposition shows
that though an
operator on a real
vector space may have
no eigenvalues, or it
may have no
eigenpairs, it cannot be
lacking in both these
useful objects. It also
shows that an operator
on a real vector space
Vcan have at most
(dimV)/2distinct
eigenpairs.the sum of the multiplicities of all the eigenvalues of Tplus the sum
of twice the multiplicities of all the eigenpairs of Tequals dimV.
Proof: SupposeVis a real vector space and T∈L(V). Then there
is a basis of Vwith respect to which the matrix of Tis as in 9.9. The
multiplicity of an eigenvalue λequals the number of times the 1-by-1
matrix[λ]appears on the diagonal of this matrix (from 9.9). The multi-
plicity of an eigenpair (α,β) equals the number of times x2+αx+βis
the characteristic polynomial of a 2-by-2 matrix on the diagonal of thismatrix (from 9.9). Because the diagonal of this matrix has length dim V,
the sum of the multiplicities of all the eigenvalues of Tplus the sum of
twice the multiplicities of all the eigenpairs of Tmust equal dim V.
SupposeVis a real vector space and T∈L(V). With respect to
some basis of V,Thas a block upper-triangular matrix of the form
9.18
A1∗
...
0Am
,
where eachAjis a 1-by-1 matrix or a 2-by-2 matrix with no eigenval-
ues (see 9.4). We define the characteristic polynomial ofTto be the
product of the characteristic polynomials of A1,...,Am. Explicitly, for
eachj, defineqj∈P(R)by
9.19qj(x)=/braceleftBigg
x−λ ifAjequals[λ];
(x−a)(x−d)−bc ifAjequals/bracketleftbigac
bd/bracketrightbig
.
Then the characteristic polynomial of Tis Note that the roots of
the characteristic
polynomial of Tequal
the eigenvalues of T,a s
was true on complex
vector spaces.q1(x)...qm(x).
Clearly the characteristic polynomial of Thas degree dim V. Fur-
thermore, 9.9 insures that the characteristic polynomial of Tdepends
only onTand not on the choice of a particular basis.
The Characteristic Polynomial 207
Now we can prove a result that was promised in the last chapter,
where we proved the analogous theorem (8.20) for operators on com-plex vector spaces.
9.20 Cayley-Hamilton Theorem: SupposeVis a real vector space
andT∈L(V). Letqdenote the characteristic polynomial of T. Then
q(T)=0.
Proof: Choose a basis of Vwith respect to which Thas a block
This proof uses the
same ideas as the proofof the analogous resulton complex vectorspaces (8.20).upper-triangular matrix of the form 9.18, where each Ajis a 1-by-1
matrix or a 2-by-2 matrix with no eigenvalues. Suppose Ujis the one- or
two-dimensional subspace spanned by the basis vectors correspondingtoA
j. Defineqjas in 9.19. To prove that q(T)=0, we need only show
thatq(T)|Uj=0 forj=1,...,m . To do this, it suffices to show that
9.21 q1(T)...qj(T)|Uj=0
forj=1,...,m .
We will prove 9.21 by induction on j. To get started, suppose that
j=1. BecauseM(T) is given by 9.18, we have q1(T)|U1=0 (obvious if
dimU1=1; from 9.7(a) if dim U1=2), giving 9.21 when j=1.
Now suppose that 1 <j≤nand that
0=q1(T)|U1
0=q1(T)q 2(T)|U2
...
0=q1(T)...qj−1(T)|Uj−1.
Ifv∈Uj, then from 9.18 we see that
qj(T)v=u+qj(S)v,
whereu∈U1+···+Uj−1andS∈L(Uj)has characteristic poly-
nomialqj. Becauseqj(S)=0 (obvious if dim Uj=1; from 9.7(a) if
dimUj=2), the equation above shows that
qj(T)v∈U1+···+Uj−1
wheneverv∈Uj. Thus, by our induction hypothesis, q1(T)...qj−1(T)
applied toqj(T)v gives 0 whenever v∈Uj. In other words, 9.21 holds,
completing the proof.
208 Chapter 9.Operators on Real Vector Spaces
SupposeVis a real vector space and T∈L(V). Clearly the Cayley-
Hamilton theorem (9.20) implies that the minimal polynomial of Thas
degree at most dim V, as was the case on complex vector spaces. If
the degree of the minimal polynomial of Tequals dimV, then, as was
also the case on complex vector spaces, the minimal polynomial of T
must equal the characteristic polynomial of T. This follows from the
Cayley-Hamilton theorem (9.20) and 8.34.
Finally, we can now prove a major structure theorem about oper-
ators on real vector spaces. The theorem below should be compared
to 8.23, the corresponding result on complex vector spaces.
9.22 Theorem: SupposeVis a real vector space and T∈L(V). Let
λ1,...,λmbe the distinct eigenvalues of T, withU1,...,Umthe corre-
sponding sets of generalized eigenvectors. Let (α1,β1),...,(αM,βM) EithermorM
might be 0. be the distinct eigenpairs of Tand letVj=null(T2+αjT+βjI)dimV.
Then
(a)V=U1⊕···⊕Um⊕V1⊕···⊕VM;
(b) eachUjand eachVjis invariant under T;
(c) each(T−λjI)|Ujand each(T2+αjT+βjI)|Vjis nilpotent.
Proof: From 8.22, we get (b). Clearly (c) follows from the defini- This proof uses the
same ideas as the proof
of the analogous result
on complex vector
spaces (8.23).tions.
To prove (a), recall that dim Ujequals the multiplicity of λjas an
eigenvalue of Tand dimVjequals twice the multiplicity of (αj,βj)as
an eigenpair of T. Thus
9.23 dimV=dimU1+···+ dimUm+dimV1+···+VM;
this follows from 9.17. Let U=U1+···+U m+V1+···+VM. Note
thatUis invariant under T. Thus we can define S∈L(U)by
S=T|U.
Note thatShas the same eigenvalues, with the same multiplicities, as T
because all the generalized eigenvectors of Tare inU, the domain of S.
Similarly,Shas the same eigenpairs, with the same multiplicities, as T.
Thus applying 9.17 to S, we get
dimU=dimU1+···+ dimUm+dimV1+···+V M.
The Characteristic Polynomial 209
This equation, along with 9.23, shows that dim V=dimU. BecauseU
is a subspace of V, this implies that V=U. In other words,
V=U1+···+Um+V1+···+VM.
This equation, along with 9.23, allows us to use 2.19 to conclude that
(a) holds, completing the proof.
210 Chapter 9.Operators on Real Vector Spaces
Exercises
1. Prove that 1 is an eigenvalue of every square matrix with the
property that the sum of the entries in each row equals 1.
2. Consider a 2-by-2 matrix of real numbers
A=/bracketleftBigg
ac
bd/bracketrightBigg
.
Prove thatAhas an eigenvalue (in R) if and only if
(a−d)2+4bc≥0.
3. Suppose Ais a block diagonal matrix
A=
A1 0
...
0Am
,
where eachAjis a square matrix. Prove that the set of eigenval-
ues ofAequals the union of the eigenvalues of A1,...,Am.
4. Suppose Ais a block upper-triangular matrix Clearly Exercise 4 is a
stronger statement
than Exercise 3. Even
so, you may want to do
Exercise 3 first because
it is easier than
Exercise 4.A=
A1∗
...
0Am
,
where eachAjis a square matrix. Prove that the set of eigenval-
ues ofAequals the union of the eigenvalues of A1,...,Am.
5. Suppose Vis a real vector space and T∈L(V). Suppose α,β∈R
are such that T2+αT+βI=0. Prove that Thas an eigenvalue
if and only if α2≥4β.
6. Suppose Vis a real inner-product space and T∈L(V). Prove
that there is an orthonormal basis of Vwith respect to which T
has a block upper-triangular matrix
A1∗
...
0Am
,
where eachAjis a 1-by-1 matrix or a 2-by-2 matrix with no eigen-
values.
Exercises 211
7. Prove that if T∈L(V)andjis a positive integer such that
j≤dimV, thenThas an invariant subspace whose dimension
equalsj−1o rj .
8. Prove that there does not exist an operator T∈L(R7)such that
T2+T+Iis nilpotent.
9. Give an example of an operator T∈L(C7)such thatT2+T+I
is nilpotent.
10. Suppose Vis a real vector space and T∈L(V). Suppose α,β∈R
are such that α2<4β. Prove that
null(T2+αT+βI)k
has even dimension for every positive integer k.
11. Suppose Vis a real vector space and T∈L(V). Suppose α,β∈R
are such that α2<4βandT2+αT+βIis nilpotent. Prove that
dimVis even and
(T2+αT+βI)dimV/2=0.
12. Prove that if T∈L(R3)and 5, 7 are eigenvalues of T, thenThas
no eigenpairs.
13. Suppose Vis a real vector space with dim V=nandT∈L(V)
is such that
nullTn−2/negationslash=nullTn−1.
Prove thatThas at most two distinct eigenvalues and that Thas
no eigenpairs.
14. Suppose Vis a vector space with dimension 2 and T∈L(V). You do not need to find
the eigenvalues of Tto
do this exercise. As
usual unless otherwise
specified, here Vmay
be a real or complex
vector space.Prove that if /bracketleftBigg
ac
bd/bracketrightBigg
is the matrix of Twith respect to some basis of V, then the char-
acteristic polynomial of Tequals(z−a)(z−d)−bc.
15. Suppose Vis a real inner-product space and S∈L(V)is an isom-
etry. Prove that if (α,β) is an eigenpair of S, thenβ=1.
Chapter 10
Trace and Determinant
Throughout this book our emphasis has been on linear maps and op-
erators rather than on matrices. In this chapter we pay more attentionto matrices as we define and discuss traces and determinants. Deter-
minants appear only at the end of this book because we replaced their
usual applications in linear algebra (the definition of the characteris-tic polynomial and the proof that operators on complex vector spaceshave eigenvalues) with more natural techniques. The book concludeswith an explanation of the important role played by determinants inthe theory of volume and integration.
Recall that Fdenotes RorC.
Also,Vis a finite-dimensional, nonzero vector space over F.
✽✽✽✽✽✽✽✽✽✽
213
214 Chapter 10. Trace and Determinant
Change of Basis
The matrix of an operator T∈L(V)depends on a choice of basis
ofV. Two different bases of Vmay give different matrices of T. In this
section we will learn how these matrices are related. This informationwill help us find formulas for the trace and determinant of Tlater in
this chapter.
With respect to any basis of V, the identity operator I∈L(V)has a
diagonal matrix
10
...
01
.
This matrix is called the identity matrix and is denoted I. Note that we
use the symbol Ito denote the identity operator (on all vector spaces)
and the identity matrix (of all possible sizes). You should always beable to tell from the context which particular meaning of Iis intended.
For example, consider the equation
M(I)=I;
on the left side Idenotes the identity operator and on the right side I
denotes the identity matrix.
IfAis a square matrix (with entries in F, as usual) with the same
size asI, thenAI=IA=A, as you should verify. A square matrix A
is called invertible if there is a square matrix Bof the same size such
Some mathematicians
use the terms
nonsingular, which
means the same as
invertible, and
singular, which means
the same as
noninvertible.thatAB=BA=I, and we call Baninverse ofA. To prove that Ahas
at most one inverse, suppose BandB/primeare inverses of A. Then
B=BI=B(AB/prime)=(BA)B/prime=IB/prime=B/prime,
and henceB=B/prime, as desired. Because an inverse is unique, we can use
the notation A−1to denote the inverse of A(ifAis invertible). In other
words, ifAis invertible, then A−1is the unique matrix of the same size
such thatAA−1=A−1A=I.
Recall that when discussing linear maps from one vector space to
another in Chapter 3, we defined the matrix of a linear map with respectto two bases—one basis for the first vector space and another basis forthe second vector space. When we study operators, which are linearmaps from a vector space to itself, we almost always use the same basis
Change of Basis 215
for both vector spaces (after all, the two vector spaces in question are
equal). Thus we usually refer to the matrix of an operator with respectto a basis, meaning that we are using one basis in two capacities. The
next proposition is one of the rare cases where we need to use twodifferent bases even though we have an operator from a vector spaceto itself.
Let’s review how matrix multiplication interacts with multiplication
of linear maps. Suppose that along with Vwe have two other finite-
dimensional vector spaces, say UandW. Let(u
1,...,up)be a basis
ofU, let(v1,...,vn)be a basis of V, and let(w1,...,wm)be a basis
ofW.I fT∈L(U,V) andS∈L(V,W), then ST∈L(U,W) and
10.1M/parenleftbig
ST,(u 1,...,up),(w 1,...,wm)/parenrightbig
=
M/parenleftbig
S,(v 1,...,vn),(w 1,...,wm)/parenrightbig
M/parenleftbig
T,(u 1,...,up),(v 1,...,vn)/parenrightbig
.
The equation above holds because we defined matrix multiplication to
make it true—see 3.11 and the material following it.
The following proposition deals with the matrix of the identity op-
erator when we use two different bases. Note that the kthcolumn of
M/parenleftbig
I,(u 1,...,un),(v 1,...,vn)/parenrightbig
consists of the scalars needed to write
ukas a linear combination of the v’s. As an example of the proposi-
tion below, consider the bases/parenleftbig
(4,2),(5, 3)/parenrightbig
and/parenleftbig
(1,0),(0, 1)/parenrightbig
ofF2.
Obviously
M/parenleftBig
I,/parenleftbig
(4,2),(5,3)/parenrightbig
,/parenleftbig
(1,0),(0, 1)/parenrightbig/parenrightBig
=/bracketleftBigg
45
23/bracketrightBigg
.
The inverse of the matrix above is/bracketleftBig
3/2−5/2
−12/bracketrightBig
, as you should verify. Thus
the proposition below implies that
M/parenleftBig
I,/parenleftbig
(1,0),(0, 1)/parenrightbig
,/parenleftbig
(4,2),(5, 3)/parenrightbig/parenrightBig
=/bracketleftBigg
3/2−5/2
−12/bracketrightBigg
.
10.2 Proposition: If(u1,...,un)and(v1,...,vn)are bases of V,
thenM/parenleftbig
I,(u 1,...,un),(v 1,...,vn)/parenrightbig
is invertible and
M/parenleftbig
I,(u 1,...,un),(v 1,...,vn)/parenrightbig−1=M/parenleftbig
I,(v 1,...,vn),(u 1,...,un)/parenrightbig
.
Proof: In 10.1, replace UandWwithV, replacewjwithuj, and
replaceSandTwithI, getting
216 Chapter 10. Trace and Determinant
I=M/parenleftbig
I,(v 1,...,vn),(u 1,...,un)/parenrightbig
M/parenleftbig
I,(u 1,...,un),(v 1,...,vn)/parenrightbig
.
Now interchange the roles of the u’s andv’s, getting
I=M/parenleftbig
I,(u 1,...,un),(v 1,...,vn)/parenrightbig
M/parenleftbig
I,(v 1,...,vn),(u 1,...,un)/parenrightbig
.
These two equations give the desired result.
Now we can see how the matrix of Tchanges when we change
bases.
10.3 Theorem: SupposeT∈L(V). Let(u1,...,un)and(v1,...,vn)
be bases ofV. LetA=M/parenleftbig
I,(u 1,...,un),(v 1,...,vn)/parenrightbig
. Then
10.4M/parenleftbig
T,(u 1,...,un)/parenrightbig
=A−1M/parenleftbig
T,(v 1,...,vn)/parenrightbig
A.
Proof: In 10.1, replace UandWwithV, replacewjwithvj, replace
TwithI, and replace SwithT, getting
10.5M/parenleftbig
T,(u 1,...,un),(v 1,...,vn)/parenrightbig
=M/parenleftbig
T,(v 1,...,vn)/parenrightbig
A.
Again use 10.1, this time replacing UandWwithV, replacingwj
withuj, and replacing SwithI, getting
M/parenleftbig
T,(u 1,...,un)/parenrightbig
=A−1M/parenleftbig
T,(u 1,...,un),(v 1,...,vn)/parenrightbig
,
where we have used 10.2. Substituting 10.5 into the equation above
gives 10.4, completing the proof.
Trace
Let’s examine the characteristic polynomial more closely than we
did in the last two chapters. If Vis ann-dimensional complex vector
space andT∈L(V), then the characteristic polynomial of Tequals
(z−λ1)...(z−λn),
whereλ1,...,λnare the eigenvalues of T, repeated according to multi-
plicity. Expanding the polynomial above, we can write the characteristicpolynomial of Tin the form
10.6z
n−(λ1+···+λn)zn−1+···+(−1)n(λ1...λn).
Trace 217
IfVis ann-dimensional real vector space and T∈L(V), then the
characteristic polynomial of Tequals
HeremorMmight
equal 0.(x−λ1)...(x−λm)(x2+α1x+β1)...(x2+αMx+βM),
whereλ1,...,λmare the eigenvalues of Tand(α1,β1),...,(αM,βM)are
the eigenpairs of T, each repeated according to multiplicity. Expanding Recall that a pair (α,β)
of real numbers is aneigenpair of Tif
α
2<4βand
T2+αT+βIis not
injective.the polynomial above, we can write the characteristic polynomial of T
in the form
10.7xn−(λ1+···+λm−α1−···−αm)xn−1+...
+(−1)m(λ1...λmβ1...βM).
In this section we will study the coefficient of zn−1(usually denoted
xn−1when we are dealing with a real vector space) in the characteristic
polynomial. In the next section we will study the constant term in thecharacteristic polynomial.
ForT∈L(V), the negative of the coefficient of z
n−1(orxn−1for real
vector spaces) in the characteristic polynomial of Tis called the trace Note that traceT
depends only on Tand
not on a basis of V
because the
characteristic
polynomial of Tdoes
not depend on a choice
of basis.ofT, denoted trace T.I fVis a complex vector space, then 10.6 shows
that traceTequals the sum of the eigenvalues of T, counting multiplic-
ity. IfVis a real vector space, then 10.7 shows that trace Tequals the
sum of the eigenvalues of Tminus the sum of the first coordinates of
the eigenpairs of T, each repeated according to multiplicity.
For example, suppose T∈L(C3)is the operator whose matrix is
10.8
3−1−2
32−3
12 0
.
Then the eigenvalues of Tare 1, 2+3i, and 2−3i, each with multi-
plicity 1, as you can verify. Computing the sum of the eigenvalues, we
have traceT=1+(2+3i)+(2−3i); in other words, trace T=5.
As another example, suppose T∈L(R3)is the operator whose ma-
trix is also given by 10.8 (note that in the previous paragraph we wereworking on a complex vector space; now we are working on a real vec-tor space). Then 1 is the only eigenvalue of T(it has multiplicity 1)
and(−4, 13)is the only eigenpair of T(it has multiplicity 1), as you
should have verified in the last chapter (see page 205). Computing the
sum of the eigenvalues minus the sum of the first coordinates of theeigenpairs, we have trace T=1−(−4); in other words, trace T=5.
218 Chapter 10. Trace and Determinant
The reason that the operators in the two previous examples have
the same trace will become clear after we find a formula (valid on bothcomplex and real vector spaces) for computing the trace of an operator
from its matrix.
Most of the rest of this section is devoted to discovering how to cal-
culate traceTfrom the matrix of T(with respect to an arbitrary basis).
Let’s start with the easiest situation. Suppose Vis a complex vector
space,T∈L(V), and we choose a basis of Vwith respect to which
Thas an upper-triangular matrix A. Then the eigenvalues of Tare
precisely the diagonal entries of A, repeated according to multiplicity
(see 8.10). Thus trace Tequals the sum of the diagonal entries of A.
The same formula works for the operator T∈L(F
3)whose matrix is
given by 10.8 and whose trace equals 5. Could such a simple formulabe true in general?
We begin our investigation by considering T∈L(V)whereVis a
real vector space. Choose a basis of Vwith respect to which Thas a
block upper-triangular matrix M(T), where each block on the diagonal
is a 1-by-1 matrix containing an eigenvalue of Tor a 2-by-2 block with
no eigenvalues (see 9.4 and 9.9). Each entry in a 1-by-1 block on thediagonal ofM(T) is an eigenvalue of Tand thus makes a contribution
to traceT.I fM(T) has any 2-by-2 blocks on the diagonal, consider a
typical one/bracketleftBigg
ac
bd/bracketrightBigg
.
The characteristic polynomial of this 2-by-2 matrix is (x−a)(x−d)−bc,
which equals
x
2−(a+d)x+(ad−bc).
Thus(−a−d,ad−bc)is an eigenpair of T. The negative of the first You should carefully
review 9.9 to
understand the
relationship between
eigenpairs and
characteristic
polynomials of 2-by-2
blocks.coordinate of this eigenpair, namely, a+d, is the contribution of this
block to trace T. Note thata+dis the sum of the entries on the di-
agonal of this block. Thus for any basis of Vwith respect to which
the matrix of Thas the block upper-triangular form required by 9.4
and 9.9, trace Tequals the sum of the entries on the diagonal.
At this point you should suspect that trace Tequals the sum of
the diagonal entries of the matrix of Twith respect to an arbitrary
basis. Remarkably, this turns out to be true. To prove it, let’s de-fine the trace of a square matrix A, denoted trace A, to be the sum
of the diagonal entries. With this notation, we want to prove that
Trace 219
traceT=traceM/parenleftbig
T,(v 1,...,vn)/parenrightbig
, where(v1,...,vn)is an arbitrary
basis ofV. We already know this is true if (v1,...,vn)is a basis with
respect to which Thas an upper-triangular matrix (if Vis complex) or
an appropriate block upper-triangular matrix (if Vis real). We will need
the following proposition to prove our trace formula for an arbitrarybasis.
10.9 Proposition: IfAandBare square matrices of the same size,
then
trace(AB)=trace(BA).
Proof: Suppose
A=
a
1,1... a 1,n
......
a
n,1... an,n
,B=
b1,1... b 1,n
......
b
n,1... bn,n
.
Thejthterm on the diagonal of ABequals
n/summationdisplay
k=1aj,kbk,j.
Thus
trace(AB)=n/summationdisplay
j=1n/summationdisplay
k=1aj,kbk,j
=n/summationdisplay
k=1n/summationdisplay
j=1bk,jaj,k
=n/summationdisplay
k=1kthterm on the diagonal of BA
=trace(BA),
as desired.
Now we can prove that the sum of the diagonal entries of the matrix
of an operator is independent of the basis with respect to which thematrix is computed.
10.10 Corollary: SupposeT∈L(V).I f(u
1,...,un)and(v1,...,vn)
are bases of V, then
traceM/parenleftbig
T,(u 1,...,un)/parenrightbig
=traceM/parenleftbig
T,(v 1,...,vn)/parenrightbig
.
220 Chapter 10. Trace and Determinant
Proof: Suppose(u1,...,un)and(v1,...,vn)are bases of V. Let
A=M/parenleftbig
I,(u 1,...,un),(v 1,...,vn)/parenrightbig
. Then
The third equality here
depends on the
associative property of
matrix multiplication.traceM/parenleftbig
T,(u 1,...,un)/parenrightbig
=trace/parenleftBig
A−1/parenleftbig
M/parenleftbig
T,(v 1,...,vn)/parenrightbig
A/parenrightbig/parenrightBig
=trace/parenleftBig/parenleftbig
M/parenleftbig
T,(v 1,...,vn)/parenrightbig
A/parenrightbig
A−1/parenrightBig
=traceM/parenleftbig
T,(v 1,...,vn)/parenrightbig
,
where the first equality follows from 10.3 and the second equality fol-
lows from 10.9. The third equality completes the proof.
The theorem below states that the trace of an operator equals the
sum of the diagonal entries of the matrix of the operator. This theoremdoes not specify a basis because, by the corollary above, the sum ofthe diagonal entries of the matrix of an operator is the same for every
choice of basis.
10.11 Theorem: IfT∈L(V), then traceT=traceM(T).
Proof: LetT∈L(V). As noted above, trace M(T) is independent
of which basis of Vwe choose (by 10.10). Thus to show that
traceT=traceM(T)
for every basis of V, we need only show that the equation above holds
for some basis of V. We already did this (on page 218), choosing a basis
ofVwith respect to which M(T) is an upper-triangular matrix (if Vis a
complex vector space) or an appropriate block upper-triangular matrix(ifVis a real vector space).
If we know the matrix of an operator on a complex vector space, the
theorem above allows us to find the sum of all the eigenvalues withoutfinding any of the eigenvalues. For example, consider the operatoronC
5whose matrix is
0000−3
1000 60100 00010 0
0001 0
.
Trace 221
No one knows an exact formula for any of the eigenvalues of this op-
erator. However, we do know that the sum of the eigenvalues equals 0because the sum of the diagonal entries of the matrix above equals 0.
The theorem above also allows us easily to prove some useful prop-
erties about traces of operators by shifting to the language of tracesof matrices, where certain properties have already been proved or areobvious. We carry out this procedure in the next corollary.
10.12 Corollary: IfS,T∈L(V), then
trace(ST)=trace(TS) and trace(S+T)=traceS+traceT.
Proof: SupposeS,T∈L(V). Choose any basis of V. Then
trace(ST)=traceM(ST)
=trace/parenleftbig
M(S)M(T)/parenrightbig
=trace/parenleftbig
M(T)M(S)/parenrightbig
=traceM(TS)
=trace(TS),
where the first and last equalities come from 10.11 and the middle
equality comes from 10.9. This completes the proof of the first asser-tion in the corollary.
To prove the second assertion in the corollary, note that
trace(S+T)=traceM(S+T)
=trace/parenleftbig
M(S)+M(T)/parenrightbig
=traceM(S)+traceM(T)
=traceS+traceT,
where again the first and last equalities come from 10.11; the third
equality is obvious from the definition of the trace of a matrix. Thiscompletes the proof of the second assertion in the corollary.
The techniques we have developed have the following curious corol-
lary. The generalization of this result to infinite-dimensional vector
spaces has important consequences in quantum theory.
222 Chapter 10. Trace and Determinant
10.13 Corollary: There do not exist operators S,T∈L(V) such that The statement of this
corollary does not
involve traces, though
the short proof uses
traces. Whenever
something like this
happens in
mathematics, we can be
sure that a good
definition lurks in the
background.ST−TS=I.
Proof: SupposeS,T∈L(V). Then
trace(ST−TS)=trace(ST)−trace(TS)
=0,
where the second equality comes from 10.12. Clearly the trace of I
equals dimV, which is not 0. Because ST−TSandIhave different
traces, they cannot be equal.
Determinant of an Operator
ForT∈L(V), we define the determinant ofT, denoted det T,t o Note that detT
depends only on Tand
not on a basis of V
because the
characteristic
polynomial of Tdoes
not depend on a choice
of basis.be(−1)dimVtimes the constant term in the characteristic polynomial
ofT. The motivation for the factor (−1)dimVin this definition comes
from 10.6.
IfVis a complex vector space, then det Tequals the product of
the eigenvalues of T, counting multiplicity; this follows immediately
from 10.6. Recall that if Vis a complex vector space, then there is
a basis ofVwith respect to which Thas an upper-triangular matrix
(see 5.13); thus det Tequals the product of the diagonal entries of this
matrix (see 8.10).
IfVis a real vector space, then det Tequals the product of the
eigenvalues of Ttimes the product of the second coordinates of the
eigenpairs of T, each repeated according to multiplicity—this follows
from 10.7 and the observation that m=dimV−2M(in the notation
of 10.7), and hence (−1)m=(−1)dimV.
For example, suppose T∈L(C3)is the operator whose matrix is
given by 10.8. As we noted in the last section, the eigenvalues of Tare
1, 2+3i, and 2−3i, each with multiplicity 1. Computing the product
of the eigenvalues, we have det T=(1)(2+3i)(2−3i); in other words,
detT=13.
As another example, suppose T∈L(R3)is the operator whose ma-
trix is also given by 10.8 (note that in the previous paragraph we were
working on a complex vector space; now we are working on a real vec-
tor space). Then, as we noted earlier, 1 is the only eigenvalue of T(it
Determinant of an Operator 223
has multiplicity 1) and (−4, 13)is the only eigenpair of T(it has multi-
plicity 1). Computing the product of the eigenvalues times the productof the second coordinates of the eigenpairs, we have det T=(1)(13);
in other words, det T=13.
The reason that the operators in the two previous examples have the
same determinant will become clear after we find a formula (valid onboth complex and real vector spaces) for computing the determinantof an operator from its matrix.
In this section, we will prove some simple but important properties
of determinants. In the next section, we will discover how to calculatedetTfrom the matrix of T(with respect to an arbitrary basis). We begin
with a crucial result that has an easy proof with our approach.
10.14 Proposition: An operator is invertible if and only if its deter-
minant is nonzero.
Proof: First suppose Vis a complex vector space and T∈L(V).
The operator Tis invertible if and only if 0 is not an eigenvalue of T.
Clearly this happens if and only if the product of the eigenvalues of T
is not 0. Thus Tis invertible if and only if det T/negationslash=0, as desired.
Now suppose Vis a real vector space and T∈L(V). Again, Tis
invertible if and only if 0 is not an eigenvalue of T. Using the notation
of 10.7, we have
10.15 detT=λ
1...λmβ1...βM,
where theλ’s are the eigenvalues of Tand theβ’s are the second coor-
dinates of the eigenpairs of T, each repeated according to multiplicity.
For each eigenpair (αj,βj), we haveαj2<4βj. In particular, each βj
is positive. This implies (see 10.15) that λ1...λm/negationslash=0 if and only if
detT/negationslash=0. ThusTis invertible if and only if det T/negationslash=0, as desired.
IfT∈L(V)andλ,z∈F, thenλis an eigenvalue of Tif and only if
z−λis an eigenvalue of zI−T. This follows from
−(T−λI)=(zI−T)−(z−λ)I.
Raising both sides of this equation to the dim Vpower and then taking
null spaces of both sides shows that the multiplicity of λas an eigen-
value ofTequals the multiplicity of z−λas an eigenvalue of zI−T.
224 Chapter 10. Trace and Determinant
The next lemma gives the analogous result for eigenpairs. We will use
this lemma to show that the characteristic polynomial can be expressedas a certain determinant.
10.16 Lemma: SupposeVis a real vector space, T∈L(V), and
Real vector spaces are
harder to deal with
than complex vector
spaces. The first time
you read this chapter,
you may want to
concentrate on the
basic ideas by
considering only
complex vector spaces
and ignoring the
special procedures
needed to deal with
real vector spaces.α,β,x∈Rwithα2<4β. Then(α,β) is an eigenpair of Tif and only
if(−2x−α,x2+αx+β)is an eigenpair of xI−T. Furthermore, these
eigenpairs have the same multiplicities.
Proof: First we need to check that (−2x−α,x2+αx+β)satisfies
the inequality required of an eigenpair. We have
(−2x−α)2=4x2+4αx+α2
<4x2+4αx+4β
=4(x2+αx+β).
Thus(−2x−α,x2+αx+β)satisfies the required inequality.
Now
T2+αT+βI=(xI−T)2−(2x+α)(xI−T)+(x2+αx+β)I,
as you should verify. Thus (α,β) is an eigenpair of Tif and only if
(−2x−α,x2+αx+β)is an eigenpair of xI−T. Furthermore, raising
both sides of the equation above to the dim Vpower and then taking
null spaces of both sides shows that the multiplicities are equal.
Most textbooks take the theorem below as the definition of the char-
acteristic polynomial. Texts using that approach must spend consider-ably more time developing the theory of determinants before they getto interesting linear algebra.
10.17 Theorem: SupposeT∈L(V). Then the characteristic poly-
nomial ofTequals det(zI−T).
Proof: First suppose Vis a complex vector space. Let λ
1,...,λn
denote the eigenvalues of T, repeated according to multiplicity. Thus
forz∈C, the eigenvalues of zI−Tarez−λ1,...,z−λn, repeated
according to multiplicity. The determinant of zI−Tis the product of
these eigenvalues. In other words,
det(zI−T)=(z−λ1)...(z−λn).
Determinant of a Matrix 225
The right side of the equation above is, by definition, the characteristic
polynomial of T, completing the proof when Vis a complex vector
space.
Now suppose Vis a real vector space. Let λ1,...,λmdenote the
eigenvalues of Tand let(α1,β1),...,(αM,βM)denote the eigenpairs
ofT, each repeated according to multiplicity. Thus for x∈R, the
eigenvalues of xI−Tarex−λ1,...,x−λmand, by 10.16, the eigenpairs
ofxI−Tare
(−2x−α1,x2+α1x+β1),...,(−2x−αM,x2+αMx+βM),
each repeated according to multiplicity. Hencedet(xI−T)=(x−λ
1)...(x−λm)(x2+α1x+β1)...(x2+αMx+βM).
The right side of the equation above is, by definition, the characteristic
polynomial of T, completing the proof when Vis a real vector space.
Determinant of a Matrix
Most of this section is devoted to discovering how to calculate det T
from the matrix of T(with respect to an arbitrary basis). Let’s start with
the easiest situation. Suppose Vis a complex vector space, T∈L(V),
and we choose a basis of Vwith respect to which Thas an upper-
triangular matrix. Then, as we noted in the last section, det Tequals
the product of the diagonal entries of this matrix. Could such a simple
formula be true in general?
Unfortunately the determinant is more complicated than the trace.
In particular, det Tneed not equal the product of the diagonal entries
ofM(T) with respect to an arbitrary basis. For example, the operator
onF3whose matrix equals 10.8 has determinant 13, as we saw in the
last section. However, the product of the diagonal entries of that matrixequals 0.
For each square matrix A, we want to define the determinant of A,
denoted detA, in such a way that det T=detM(T) regardless of which
basis is used to compute M(T). We begin our search for the correct def-
inition of the determinant of a matrix by calculating the determinantsof some special operators.
Letc
1,...,cn∈Fbe nonzero scalars and let (v1,...,vn)be a basis
ofV. Consider the operator T∈L(V)such thatM/parenleftbig
T,(v 1,...,vn)/parenrightbig
equals
226 Chapter 10. Trace and Determinant
10.18
0 c
n
c10
c20
......
cn−1 0
;
here all entries of the matrix are 0 except for the upper-right corner
and along the line just below the diagonal. Let’s find the determinantofT. Note that
(v
1,Tv 1,T2v1,...,Tn−1v1)=(v1,c1v2,c1c2v3,...,c 1...cn−1vn).
Thus(v1,Tv 1,...,Tn−1v1)is linearly independent (the c’s are all non-
zero). Hence if pis a nonzero polynomial with degree at most n−1,
thenp(T)v 1/negationslash=0. In other words, the minimal polynomial of Tcannot
have degree less than n. As you should verify, Tnvj=c1...cnvjfor
eachj, and hence Tn=c1...cnI. Thuszn−c1...cnis the minimal
polynomial of T. Becausen=dimV, we see that zn−c1...cnis also Recall that if the
minimal polynomial of
an operator T∈L(V)
has degree dimV, then
the characteristic
polynomial of Tequals
the minimal polynomial
ofT. Computing the
minimal polynomial is
often an efficient
method of finding the
characteristic
polynomial.the characteristic polynomial of T. Multiplying the constant term of
this polynomial by (−1)n, we get
10.19 detT=(−1)n−1c1...cn.
If somecjequals 0, then clearly Tis not invertible, so det T=0 and
the same formula holds. Thus in order to have det T=detM(T),w e
will have to make the determinant of 10.18 equal to (−1)n−1c1...cn.
However, we do not yet have enough evidence to make a reasonableguess about the proper definition of the determinant of an arbitrary
square matrix.
To compute the determinants of a more complicated class of op-
erators, we introduce the notion of permutation. A permutation of
(1,...,n) is a list(m
1,...,mn)that contains each of the numbers
1,...,n exactly once. The set of all permutations of (1,...,n) is de-
noted permn. For example, (2,3,...,n,1)∈permn. You should think
of an element of perm nas a rearrangement of the first nintegers.
For simplicity we will work with matrices with complex entries (at
this stage we are providing only motivation—formal proofs will comelater). Letc
1,...,cn∈Cand let(v1,...,vn)be a basis of V, which
we are assuming is a complex vector space. Consider a permutation
(p1,...,pn)∈permnthat can be obtained as follows: break (1,...,n)
Determinant of a Matrix 227
into lists of consecutive integers and in each list move the first term to
the end of that list. For example, taking n=9, the permutation
10.20 (2,3,1,5,6,7,4,9,8)
is obtained from (1,2,3),(4, 5,6,7),(8, 9)by moving the first term of
each of these lists to the end, producing (2,3,1),(5, 6,7,4),(9, 8), and
then putting these together to form 10.20. Let T∈L(V)be the operator
such that
10.21 Tvk=ckvpk
fork=1,...,n . We want to find a formula for det T. This generalizes
our earlier example because if (p1,...,pn)happens to be the permuta-
tion(2,3,...,n,1), then the operator Twhose matrix equals 10.18 is
the same as the operator Tdefined by 10.21.
With respect to the basis (v1,...,vn), the matrix of the operator T
defined by 10.21 is a block diagonal matrix
A=
A1 0
...
0AM
,
where each block is a square matrix of the form 10.18. The eigenvalues
ofTequal the union of the eigenvalues of A1,...,AM(see Exercise 3 in
Chapter 9). Recalling that the determinant of an operator on a complex
vector space is the product of the eigenvalues, we see that our definition
of the determinant of a square matrix should force
detA=(detA1)...( detAM).
However, we already know how to compute the determinant of each Aj,
which has the same form as 10.18 (of course with a different value of n).
Putting all this together, we see that we should have
detA=(−1)n1−1...(−1)nM−1c1...cn,
whereAjhas sizenj-by-nj. The number (−1)n1−1...(−1)nM−1is called
the sign of the permutation (p1,...,pn), denoted sign(p 1,...,pn)(this
is a temporary definition that we will change to an equivalent definition
later, when we define the sign of an arbitrary permutation).
228 Chapter 10. Trace and Determinant
To put this into a form that does not depend on the particular per-
mutation(p1,...,pn), letaj,kdenote the entry in row j, columnk,o fA;
thus
aj,k=/braceleftBigg
0i fj/negationslash=pk;
ckifj=pk.
Then
10.22 detA=/summationdisplay
(m1,...,mn)∈permn/parenleftbig
sign(m 1,...,mn)/parenrightbig
am1,1...amn,n,
because each summand is 0 except the one corresponding to the per-
mutation(p1,...,pn).
Consider now an arbitrary matrix Awith entryaj,kin rowj, col-
umnk. Using the paragraph above as motivation, we guess that det A
should be defined by 10.22. This will turn out to be correct. We can
now dispense with the motivation and begin the more formal approach.
First we will need to define the sign of an arbitrary permutation.
The sign of a permutation (m1,...,mn)is defined to be 1 if the Some texts use the
unnecessarily fancy
term signum, which
means the same
as sign.number of pairs of integers (j,k) with 1≤j<k≤nsuch thatjap-
pears afterkin the list(m1,...,mn)is even and−1 if the number of
such pairs is odd. In other words, the sign of a permutation equals 1 ifthe natural order has been changed an even number of times and equals−1 if the natural order has been changed an odd number of times. Forexample, in the permutation (2,3,...,n,1) the only pairs (j,k) with
j<k that appear with changed order are (1,2),(1, 3),...,( 1,n) ; be-
cause we have n−1 such pairs, the sign of this permutation equals
(−1)
n−1(note that the same quantity appeared in 10.19).
The permutation (2,1,3,4), which is obtained from the permutation
(1,2,3,4)by interchanging the first two entries, has sign −1. The next
lemma shows that interchanging any two entries of any permutationchanges the sign of the permutation.
10.23 Lemma: Interchanging two entries in a permutation multiplies
the sign of the permutation by −1.
Proof: Suppose we have two permutations, where the second per-
mutation is obtained from the first by interchanging two entries. If thetwo entries that we interchanged were in their natural order in the firstpermutation, then they no longer are in the second permutation, and
Determinant of a Matrix 229
vice versa, for a net change (so far) of 1 or −1 (both odd numbers) in
the number of pairs not in their natural order.
Consider each entry between the two interchanged entries. If an in-
termediate entry was originally in the natural order with respect to thefirst interchanged entry, then it no longer is, and vice versa. Similarly,if an intermediate entry was originally in the natural order with respect
to the second interchanged entry, then it no longer is, and vice versa.Thus the net change for each intermediate entry in the number of pairsnot in their natural order is 2, 0, or −2 (all even numbers).
For all the other entries, there is no change in the number of pairs
not in their natural order. Thus the total net change in the number of
pairs not in their natural order is an odd number. Thus the sign of the
second permutation equals −1 times the sign of the first permutation.
IfAis ann-by-n matrix
10.24 A=
a1,1... a 1,n
......
an,1... an,n
,
then the determinant ofA, denoted det A, is defined by Our motivation for this
definition comes
from 10.22.10.25 detA=/summationdisplay
(m1,...,mn)∈permn/parenleftbig
sign(m 1,...,mn)/parenrightbig
am1,1...amn,n.
For example, if Ais the 1-by-1 matrix [a1,1], then detA=a1,1be-
cause perm 1 has only one element, namely, (1), which has sign 1. For
a more interesting example, consider a typical 2-by-2 matrix. Clearlyperm 2 has only two elements, namely, (1,2), which has sign 1, and
(2,1), which has sign −1. Thus
10.26 det/bracketleftBigg
a
1,1a1,2
a2,1a2,2/bracketrightBigg
=a1,1a2,2−a2,1a1,2.
To make sure you understand this process, you should now find the
formula for the determinant of the 3-by-3 matrix The set perm 3
contains 6elements. In
general, permn
containsn!elements.
Note thatn!rapidly
grows large as n
increases.
a1,1a1,2a1,3
a2,1a2,2a2,3
a3,1a3,2a3,3
using just the definition given above (do this even if you already know
the answer).
230 Chapter 10. Trace and Determinant
Let’s compute the determinant of an upper-triangular matrix
A=
a1,1∗
...
0an,n
.
The permutation (1,2,...,n) has sign 1 and thus contributes a term
ofa1,1...an,nto the sum 10.25 defining det A. Any other permutation
(m1,...,mn)∈permncontains at least one entry mjwithmj>j,
which means that amj,j=0 (becauseAis upper triangular). Thus all
the other terms in the sum 10.25 defining det Amake no contribu-
tion. Hence det A=a1,1...an,n. In other words, the determinant of an
upper-triangular matrix equals the product of the diagonal entries. Inparticular, this means that if Vis a complex vector space, T∈L(V),
and we choose a basis of Vwith respect to which M(T) is upper trian-
gular, then det T=detM(T). Our goal is to prove that this holds for
every basis of V, not just bases that give upper-triangular matrices.
Generalizing the computation from the paragraph above, next we
will show that if Ais a block upper-triangular matrix
A=
A
1∗
...
0Am
,
where eachAjis a 1-by-1 or 2-by-2 matrix, then
10.27 detA=(detA1)...( detAm).
To prove this, consider an element of perm n. If this permutation
moves an index corresponding to a 1-by-1 block on the diagonal any-
place else, then the permutation makes no contribution to the sum10.25 defining det A(becauseAis block upper triangular). For a pair
of indices corresponding to a 2-by-2 block on the diagonal, the permu-
tation must either leave these indices fixed or interchange them; oth-erwise again the permutation makes no contribution to the sum 10.25defining det A(becauseAis block upper triangular). These observa-
tions, along with the formula 10.26 for the determinant of a 2-by-2 ma-trix, lead to 10.27. In particular, if Vis a real vector space, T∈L(V),
and we choose a basis of Vwith respect to which M(T) is a block
upper-triangular matrix with 1-by-1 and 2-by-2 blocks on the diagonal
as in 9.9, then det T=detM(T).
Determinant of a Matrix 231
Our goal is to prove that det T=detM(T) for everyT∈L(V)and An entire book could
be devoted just to
deriving properties ofdeterminants.Fortunately we needonly a few of the basic
properties.every basis of V. To do this, we will need to develop some proper-
ties of determinants of matrices. The lemma below is the first of theproperties we will need.
10.28 Lemma: SupposeAis a square matrix. If Bis the matrix
obtained from Aby interchanging two columns, then
detA=−detB.
Proof: SupposeAis given by 10.24 and Bis obtained from Aby
interchanging two columns. Think of the sum 10.25 defining det Aand
the corresponding sum defining det B. The same products of a’s appear
in both sums, though they correspond to different permutations. The
permutation corresponding to a given product of a’s when computing
detBis obtained by interchanging two entries in the corresponding
permutation when computing det A, thus multiplying the sign of the
permutation by −1 (see 10.23). Hence det A=−detB.
IfT∈L(V)and the matrix of T(with respect to some basis) has two
equal columns, then Tis not injective and hence det T=0. Though
this comment makes the next lemma plausible, it cannot be used in theproof because we do not yet know that det T=detM(T).
10.29 Lemma: IfAis a square matrix that has two equal columns,
then detA=0.
Proof: SupposeAis a square matrix that has two equal columns.
Interchanging the two equal columns of Agives the original matrix A.
Thus from 10.28 (with B=A), we have det A=−detA, which implies
that detA=0.
This section is long, so let’s pause for a paragraph. The symbols ✽
that appear on the first page of each chapter are decorations intended
to take up space so that the first section of the chapter can start on thenext page. Chapter 1 has one of these symbols, Chapter 2 has two ofthem, and so on. The symbols get smaller with each chapter. What youmay not have noticed is that the sum of the areas of the symbols at the
beginning of each chapter is the same for all chapters. For example, the
diameter of each symbol at the beginning of Chapter 10 equals 1 /√
10
times the diameter of the symbol in Chapter 1.
232 Chapter 10. Trace and Determinant
We need to introduce notation that will allow us to represent a ma-
trix in terms of its columns. If Ais ann-by-n matrix
A=
a1,1... a 1,n
......
an,1... an,n
,
then we can think of the kthcolumn ofAas ann-by-1 matrix
ak=
a1,k
...
an,k
.
We will write Ain the form
[a1... an],
with the understanding that akdenotes thekthcolumn ofA. With this
notation, note that aj,k, with two subscripts, denotes an entry of A,
whereasak, with one subscript, denotes a column of A.
The next lemma shows that a permutation of the columns of a matrix
changes the determinant by a factor of the sign of the permutation.
10.30 Lemma: SupposeA=[a1... an]is ann-by-n matrix. Some texts define the
determinant to be the
function defined on the
square matrices that is
linear as a function of
each column separately
and that satisfies 10.30
anddetI=1. To prove
that such a function
exists and that it is
unique takes a
nontrivial amount of
work.If(m1,...,mn)is a permutation, then
det[am1... amn]=/parenleftbig
sign(m 1,...,mn)/parenrightbig
detA.
Proof: Suppose(m1,...,mn)∈permn. We can transform the
matrix[am1... amn]intoAthrough a series of steps. In each
step, we interchange two columns and hence multiply the determinant
by−1 (see 10.28). The number of steps needed equals the number
of steps needed to transform the permutation (m1,...,mn)into the
permutation (1,...,n) by interchanging two entries in each step. The
proof is completed by noting that the number of such steps is even if
(m1,...,mn)has sign 1, odd if (m1,...,mn)has sign−1 (this follows
from 10.23, along with the observation that the permutation (1,...,n)
has sign 1).
LetA=[a1... an]. For 1≤k≤n, think of all columns of A
except thekthcolumn as fixed. We have
Determinant of a Matrix 233
detA=det[a1... ak... an],
and we can think of det Aas a function of the kthcolumnak. This
function, which takes akto the determinant above, is a linear map
from the vector space of n-by-1 matrices with entries in FtoF. The
linearity follows easily from 10.25, where each term in the sum containsprecisely one entry from the k
thcolumn ofA.
Now we are ready to prove one of the key properties about determi-
nants of square matrices. This property will enable us to connect the
determinant of an operator with the determinant of its matrix. Notethat this proof is considerably more complicated than the proof of thecorresponding result about the trace (see 10.9).
10.31 Theorem: IfAandBare square matrices of the same size,
This theorem was first
proved in 1812 by the
French mathematiciansJacques Binet and
Augustin-Louis Cauchy.then
det(AB)=det(BA)=(detA)(detB).
Proof: LetA=[a1... an], where each akis ann-by-1 column
ofA. Let
B=
b1,1... b 1,n
......
bn,1... bn,n
=[b1... bn],
where eachbkis ann-by-1 column of B. Letekdenote then-by-1 matrix
that equals 1 in the kthrow and 0 elsewhere. Note that Aek=akand
Bek=bk. Furthermore, bk=/summationtextn
m=1bm,kem.
First we will prove that det (AB)=(detA)(detB). A moment’s
thought about the definition of matrix multiplication shows that AB=
[Ab1... Abn]. Thus
det(AB)=det[Ab1... Abn]
=det[A(/summationtextn
m1=1bm1,1em1). . .A (/summationtextn
mn=1bmn,nemn)]
=det[/summationtextn
m1=1bm1,1Aem1.../summationtextn
mn=1bmn,nAemn]
=n/summationdisplay
m1=1···n/summationdisplay
mn=1bm1,1...bmn,ndet[Aem1... Aemn],
where the last equality comes from repeated applications of the linear-
ity of det as a function of one column at a time. In the last sum above,
234 Chapter 10. Trace and Determinant
all terms in which mj=mkfor somej/negationslash=kcan be ignored because the
determinant of a matrix with two equal columns is 0 (by 10.29). Thusinstead of summing over all m
1,...,mnwith eachmjtaking on values
1,...,n , we can sum just over the permutations, where the mj’s have
distinct values. In other words,
det(AB)=/summationdisplay
(m1,...,mn)∈permnbm1,1...bmn,ndet[Aem1... Aemn]
=/summationdisplay
(m1,...,mn)∈permnbm1,1...bmn,n/parenleftbig
sign(m 1,...,mn)/parenrightbig
detA
=(detA)/summationdisplay
(m1,...,mn)∈permn/parenleftbig
sign(m 1,...,mn)/parenrightbig
bm1,1...bmn,n
=(detA)(detB),
where the second equality comes from 10.30.
In the paragraph above, we proved that det (AB)=(detA)(detB).
Interchanging the roles of AandB, we have det (BA)=(detB)(detA).
The last equation can be rewritten as det (BA)=(detA)(detB), com-
pleting the proof.
Now we can prove that the determinant of the matrix of an oper-
ator is independent of the basis with respect to which the matrix iscomputed.
10.32 Corollary: SupposeT∈L(V).I f(u
1,...,un)and(v1,...,vn)
are bases of V, then
detM/parenleftbig
T,(u 1,...,un)/parenrightbig
=detM/parenleftbig
T,(v 1,...,vn)/parenrightbig
.
Proof: Suppose(u1,...,un)and(v1,...,vn)are bases of V. Let Note the similarity of
this proof to the proof
of the analogous result
about the trace
(see 10.10).A=M/parenleftbig
I,(u 1,...,un),(v 1,...,vn)/parenrightbig
. Then
detM/parenleftbig
T,(u 1,...,un)/parenrightbig
=det/parenleftBig
A−1/parenleftbig
M/parenleftbig
T,(v 1,...,vn)/parenrightbig
A/parenrightbig/parenrightBig
=det/parenleftBig/parenleftbig
M/parenleftbig
T,(v 1,...,vn)/parenrightbig
A/parenrightbig
A−1/parenrightBig
=detM/parenleftbig
T,(v 1,...,vn)/parenrightbig
,
where the first equality follows from 10.3 and the second equality fol-
lows from 10.31. The third equality completes the proof.
Determinant of a Matrix 235
The theorem below states that the determinant of an operator equals
the determinant of the matrix of the operator. This theorem does notspecify a basis because, by the corollary above, the determinant of the
matrix of an operator is the same for every choice of basis.
10.33 Theorem: IfT∈L(V), then detT=detM(T).
Proof: LetT∈L(V). As noted above, 10.32 implies that det M(T)
is independent of which basis of Vwe choose. Thus to show that
detT=detM(T)
for every basis of V, we need only show that the equation above holds
for some basis of V. We already did this (on page 230), choosing a basis
ofVwith respect to which M(T) is an upper-triangular matrix (if Vis a
complex vector space) or an appropriate block upper-triangular matrix(ifVis a real vector space).
If we know the matrix of an operator on a complex vector space, the
theorem above allows us to find the product of all the eigenvalues with-out finding any of the eigenvalues. For example, consider the operatoronC
5whose matrix is
0000−3
1000 60100 0
0010 0
0001 0
.
No one knows an exact formula for any of the eigenvalues of this opera-
tor. However, we do know that the product of the eigenvalues equals −3
because the determinant of the matrix above equals −3.
The theorem above also allows us easily to prove some useful prop-
erties about determinants of operators by shifting to the language ofdeterminants of matrices, where certain properties have already beenproved or are obvious. We carry out this procedure in the next corol-lary.
10.34 Corollary: IfS,T∈L(V), then
det(ST)=det(TS)=(detS)(detT).
236 Chapter 10. Trace and Determinant
Proof: SupposeS,T∈L(V). Choose any basis of V. Then
det(ST)=detM(ST)
=det/parenleftbig
M(S)M(T)/parenrightbig
=/parenleftbig
detM(S)/parenrightbig/parenleftbig
detM(T)/parenrightbig
=(detS)(detT),
where the first and last equalities come from 10.33 and the third equal-
ity comes from 10.31.
In the paragraph above, we proved that det (ST)=(detS)(detT). In-
terchanging the roles of SandT, we have det (TS)=(detT)(detS). Be-
cause multiplication of elements of Fis commutative, the last equation
can be rewritten as det (TS)=(detS)(detT), completing the proof.
Volume
We proved the basic results of linear algebra before introducing de-
terminants in this final chapter. Though determinants have value as aresearch tool in more advanced subjects, they play little role in basiclinear algebra (when the subject is done right). Determinants do have
Most applied
mathematicians agree
that determinants
should rarely be used
in serious numeric
calculations.one important application in undergraduate mathematics, namely, in
computing certain volumes and integrals. In this final section we willuse the linear algebra we have learned to make clear the connectionbetween determinants and these applications. Thus we will be dealing
with a part of analysis that uses linear algebra.
We begin with some purely linear algebra results that will be use-
ful when investigating volumes. Recall that an isometry on an inner-
product space is an operator that preserves norms. The next resultshows that every isometry has determinant with absolute value 1.
10.35 Proposition: Suppose that Vis an inner-product space. If
S∈L(V) is an isometry, then |detS|=1.
Proof: SupposeS∈L(V)is an isometry. First consider the case
whereVis a complex inner-product space. Then all the eigenvalues of S
have absolute value 1 (by 7.37). Thus the product of the eigenvaluesofS, counting multiplicity, has absolute value one. In other words,
|detS|=1, as desired.
Volume 237
Now suppose Vis a real inner-product space. Then there is an ortho-
normal basis of Vwith respect to which Shas a block diagonal matrix,
where each block on the diagonal is a 1-by-1 matrix containing 1 or −1
or a 2-by-2 matrix of the form
10.36/bracketleftBigg
cosθ−sinθ
sinθcosθ/bracketrightBigg
,
withθ∈(0,π) (see 7.38). Note that the constant term of the charac-
teristic polynomial of each matrix of the form 10.36 equals 1 (because
cos2θ+sin2θ=1). Hence the second coordinate of every eigenpair
ofSequals 1. Thus the determinant of Sis the product of 1’s and −1’s.
In particular, |detS|=1, as desired.
SupposeVis a real inner-product space and S∈L(V)is an isometry.
By the proposition above, the determinant of Sequals 1 or−1. Note
that
{v∈V:Sv=−v}
is the subspace of Vconsisting of all eigenvectors of Scorresponding
to the eigenvalue −1 (or is the subspace {0}if−1 is not an eigenvalue
ofS). Thinking geometrically, we could say that this is the subspace
on whichSreverses direction. A careful examination of the proof of
the last proposition shows that det S=1 if this subspace has even
dimension and det S=−1 if this subspace has odd dimension.
A self-adjoint operator on a real inner-product space has no eigen-
pairs (by 7.11). Thus the determinant of a self-adjoint operator on areal inner-product space equals the product of its eigenvalues, count-ing multiplicity (of course, this holds for any operator, self-adjoint ornot, on a complex vector space).
Recall that if Vis an inner-product space and T∈L(V), thenT
∗T
is a positive operator and hence has a unique positive square root, de-noted√
T∗T(see 7.27 and 7.28). Because√
T∗Tis positive, all its eigen-
values are nonnegative (again, see 7.27), and hence its determinant isnonnegative. Thus in the corollary below, taking the absolute value ofdet√
T∗Twould be superfluous.
10.37 Corollary: SupposeVis an inner-product space. If T∈L(V),
then
|detT|=det√
T∗T.
238 Chapter 10. Trace and Determinant
Proof: SupposeT∈L(V). By the polar decomposition (7.41), there Another proof of this
corollary is suggested
in Exercise 24 in this
chapter.is an isometry S∈L(V)such that
T=S√
T∗T.
Thus
|detT|=|detS|det√
T∗T
=det√
T∗T,
where the first equality follows from 10.34 and the second equality
follows from 10.35.
SupposeVis a real inner-product space and T∈L(V)is invertible.
The detTis either positive or negative. A careful examination of the
proof of the corollary above can help us attach a geometric meaningto whichever of these possibilities holds. To see this, first apply thereal spectral theorem (7.13) to the positive operator√
T∗T, getting an
orthonormal basis (e1,...,en)ofVsuch that√
T∗Tej=λjej, where
λ1,...,λnare the eigenvalues of√
T∗T, repeated according to multi-
plicity. Because each λjis positive,√
T∗Tnever reverses direction. We are not formally
defining the phrase
“reverses direction”
because these
comments are meant to
be an intuitive aid to
our understanding, not
rigorous mathematics.Now consider the polar decomposition
T=S√
T∗T,
whereS∈L(V)is an isometry. Then det T=(detS)(det√
T∗T). Thus
whether det Tis positive or negative depends on whether det Sis pos-
itive or negative. As we saw earlier, this depends on whether the spaceon whichSreverses direction has even or odd dimension. Because
Tis the product of Sand an operator that never reverses direction
(namely,√
T∗T), we can reasonably say that whether det Tis positive
or negative depends on whether Treverses vectors an even or an odd
number of times.
Now we turn to the question of volume, where we will consider only
the real inner-product space Rn(with its standard inner product). We
would like to assign to each subset ΩofRnitsn-dimensional volume,
denoted volume Ω(whenn=2, this is usually called area instead of
volume). We begin with cubes, where we have a good intuitive notion ofvolume. The cube inR
nwith side length rand vertex(x1,...,xn)∈Rn
is the set
Volume 239
{(y 1,...,yn)∈Rn:xj<yj<xj+rforj=1,...,n};
you should verify that when n=2, this gives a square, and that when
n=3, it gives a familiar three-dimensional cube. The volume of a cube
inRnwith side length ris defined to be rn. To define the volume of
an arbitrary set Ω⊂Rn, the idea is to write Ωas a subset of a union of Readers familiar with
outer measure will
recognize that concepthere.many small cubes, then add up the volumes of these small cubes. As
we approximate Ωmore accurately by unions (perhaps infinite unions)
of small cubes, we get a better estimate of volume Ω.
Rather than take the trouble to make precise this definition of vol-
ume, we will work only with an intuitive notion of volume. Our purposein this book is to understand linear algebra, whereas notions of volumebelong to analysis (though as we will soon see, volume is intimately con-nected with determinants). Thus for the rest of this section we will rely
on intuitive notions of volume rather than on a rigorous development,
though we shall maintain our usual rigor in the linear algebra partsof what follows. Everything said here about volume will be correct—the intuitive reasons given here can be converted into formally correctproofs using the machinery of analysis.
ForT∈L(V)andΩ⊂R
n, defineT(Ω)by
T(Ω)={Tx:x∈Ω}.
Our goal is to find a formula for the volume of T(Ω)in terms of T
and the volume of Ω. First let’s consider a simple example. Suppose
λ1,...,λnare positive numbers. Define T∈L(Rn)byT(x 1,...,xn)=
(λ1x1,...,λnxn).I fΩis a cube in Rnwith side length r, thenT(Ω)
is a box in Rnwith sides of length λ1r,...,λnr. This box has volume
λ1...λnrn, whereas the cube Ωhas volumern. Thus this particular T,
when applied to a cube, multiplies volumes by a factor of λ1...λn,
which happens to equal det T.
As above, assume that λ1,...,λnare positive numbers. Now sup-
pose that(e1,...,en)is an orthonormal basis of RnandTis the op-
erator on Rnthat satisfies Tej=λjejforj=1,...,n . In the special
case where(e1,...,en)is the standard basis of Rn, this operator is the
same one as defined in the paragraph above. Even for an arbitrary or-thonormal basis (e
1,...,en), this operator has the same behavior as
the one in the paragraph above—it multiplies the jthbasis vector by
a factor ofλj. Thus we can reasonably assume that this operator also
multiplies volumes by a factor of λ1...λn, which again equals det T.
240 Chapter 10. Trace and Determinant
We need one more ingredient before getting to the main result in
this section. Suppose S∈L(Rn)is an isometry. For x,y∈Rn,w e
have
/bardblSx−Sy/bardbl=/bardblS(x−y)/bardbl
=/bardblx−y/bardbl.
In other words, Sdoes not change the distance between points. As you
can imagine, this means that Sdoes not change volumes. Specifically,
ifΩ⊂Rn, then volume S(Ω)=volumeΩ.
Now we can give our pseudoproof that an operator T∈L(Rn)
changes volumes by a factor of |detT|.
10.38 Theorem: IfT∈L(Rn), then
volumeT(Ω)=|detT|(volumeΩ)
forΩ⊂Rn.
Proof: First consider the case where T∈L(Rn)is a positive
operator. Let λ1,...,λnbe the eigenvalues of T, repeated according
to multiplicity. Each of these eigenvalues is a nonnegative number(see 7.27). By the real spectral theorem (7.13), there is an orthonormalbasis(e
1,...,en)ofVsuch thatTej=λjejfor eachj. As discussed
above, this implies that Tchanges volumes by a factor of det T.
Now suppose T∈L(Rn)is an arbitrary operator. By the polar de-
composition (7.41), there is an isometry S∈L(V)such that
T=S√
T∗T.
IfΩ⊂Rn, thenT(Ω)=S/parenleftbig√
T∗T(Ω)/parenrightbig
. Thus
volumeT(Ω)=volumeS/parenleftbig√
T∗T(Ω)/parenrightbig
=volume√
T∗T(Ω)
=(det√
T∗T)(volumeΩ)
=|detT|(volumeΩ),
where the second equality holds because volumes are not changed by
the isometry S(as discussed above), the third equality holds by the
previous paragraph (applied to the positive operator√
T∗T), and the
fourth equality holds by 10.37.
Volume 241
The theorem above leads to the appearance of determinants in the
formula for change of variables in multivariable integration. To de-scribe this, we will again be vague and intuitive. If Ω⊂R
nandfis
a real-valued function (not necessarily linear) on Ω, then the integral
offoverΩ, denoted/integraltext
Ωfor/integraltext
Ωf(x)dx , is defined by breaking Ωinto
pieces small enough so that fis almost constant on each piece. On
each piece, multiply the (almost constant) value of fby the volume of
the piece, then add up these numbers for all the pieces, getting an ap-proximation to the integral that becomes more accurate as we divide
Ωinto finer pieces. Actually Ωneeds to be a reasonable set (for ex-
ample, open or measurable) and fneeds to be a reasonable function
(for example, continuous or measurable), but we will not worry aboutthose technicalities. Also, notice that the xin/integraltext
Ωf(x)dx is a dummy
variable and could be replaced with any other symbol.
Fix a setΩ⊂Rnand a function (not necessarily linear) σ:Ω→Rn.
We will useσto make a change of variables in an integral. Before we
can get to that, we need to define the derivative of σ, a concept that
uses linear algebra. For x∈Ω, the derivative ofσatxis an operator Ifn=1, then the
derivative in this sense
is the operator on Rof
multiplication by the
derivative in the usual
sense of one-variable
calculus.T∈L(Rn)such that
lim
y→0/bardblσ(x+y)−σ(x)−Ty/bardbl
/bardbly/bardbl=0.
If an operator T∈L(Rn)exists satisfying the equation above, then
σis said to be differentiable atx.I fσis differentiable at x, then
there is a unique operator T∈L(Rn)satisfying the equation above
(we will not prove this). This operator Tis denotedσ/prime(x). Intuitively,
the idea is that for xfixed and/bardbly/bardblsmall, a good approximation to
σ(x+y)isσ(x)+/parenleftbig
σ/prime(x)/parenrightbig
(y)(note thatσ/prime(x)∈L(Rn), so this makes
sense). Note that for xfixed the addition of the term σ(x) does not
change volumes. Thus if Γis a small subset of Ωcontainingx, then
volumeσ(Γ)is approximately equal to volume/parenleftbig
σ/prime(x)/parenrightbig
(Γ).
Becauseσis a function from ΩtoRn, we can write
σ(x)=/parenleftbig
σ1(x),...,σ n(x)/parenrightbig
,
where eachσjis a function from ΩtoR. The partial derivative of σj
with respect to the kthcoordinate is denoted Dkσj. Evaluating this
partial derivative at a point x∈ΩgivesDkσj(x).I fσis differentiable
atx, then the matrix of σ/prime(x)with respect to the standard basis of Rn
242 Chapter 10. Trace and Determinant
containsDkσj(x)in rowj, columnk(we will not prove this). In other
words,
10.39 M(σ/prime(x))=
D1σ1(x) ... D nσ1(x)
......
D1σn(x) ... D nσn(x)
.
Suppose that σis differentiable at each point of Ωand thatσis
injective on Ω. Letfbe a real-valued function defined on σ(Ω). Let
x∈Ωand letΓbe a small subset of Ωcontainingx. As we noted above,
volumeσ(Γ)≈volume/parenleftbig
σ/prime(x)/parenrightbig
(Γ),
where the symbol ≈means “approximately equal to”. Using 10.38, this
becomes
volumeσ(Γ)≈|detσ/prime(x)|(volume Γ).
Lety=σ(x) . Multiply the left side of the equation above by f(y) and
the right side by f/parenleftbig
σ(x)/parenrightbig
(becausey=σ(x) , these two quantities are
equal), getting
10.40f(y) volumeσ(Γ)≈f/parenleftbig
σ(x)/parenrightbig
|detσ/prime(x)|(volume Γ).
Now divide Ωinto many small pieces and add the corresponding ver-
sions of 10.40, getting
10.41/integraldisplay
σ(Ω)f(y)dy=/integraldisplay
Ωf/parenleftbig
σ(x)/parenrightbig
|detσ/prime(x)|dx.
This formula was our goal. It is called a change of variables formula
because you can think of y=σ(x) as a change of variables.
The key point when making a change of variables is that the factor
of|detσ/prime(x)| must be included, as in the right side of 10.41. We finish
up by illustrating this point with two important examples. When n=2,
we can use the change of variables induced by polar coordinates. In this If you are not familiar
with polar and
spherical coordinates,
skip the remainder of
this section.caseσis defined by
σ(r,θ)=(rcosθ,rsinθ),
where we have used r,θas the coordinates instead of x1,x2for reasons
that will be obvious to everyone familiar with polar coordinates (and
will be a mystery to everyone else). For this choice of σ, the matrix of
partial derivatives corresponding to 10.39 is
Volume 243
/bracketleftBigg
cosθ−rsinθ
sinθr cosθ/bracketrightBigg
,
as you should verify. The determinant of the matrix above equals r,
thus explaining why a factor of ris needed when computing an integral
in polar coordinates.
Finally, when n=3, we can use the change of variables induced by
spherical coordinates. In this case σis defined by
σ(ρ,ϕ,θ)=(ρsinϕcosθ,ρsinϕsinθ,ρcosϕ),
where we have used ρ,θ,ϕ as the coordinates instead of x1,x2,x3
for reasons that will be obvious to everyone familiar with spherical
coordinates (and will be a mystery to everyone else). For this choiceofσ, the matrix of partial derivatives corresponding to 10.39 is
sinϕcosθρ cosϕcosθ−ρsinϕsinθ
sinϕsinθρ cosϕsinθρ sinϕcosθ
cosϕ−ρsinϕ 0
,
as you should verify. You should also verify that the determinant of the
matrix above equals ρ
2sinϕ, thus explaining why a factor of ρ2sinϕ
is needed when computing an integral in spherical coordinates.
244 Chapter 10. Trace and Determinant
Exercises
1. Suppose T∈L(V)and(v1,...,vn)is a basis of V. Prove that
M/parenleftbig
T,(v 1,...,vn)/parenrightbig
is invertible if and only if Tis invertible.
2. Prove that if AandBare square matrices of the same size and
AB=I, thenBA=I.
3. Suppose T∈L(V)has the same matrix with respect to every ba-
sis ofV. Prove thatTis a scalar multiple of the identity operator.
4. Suppose that (u1,...,un)and(v1,...,vn)are bases of V. Let
T∈L(V)be the operator such that Tvk=ukfork=1,...,n .
Prove that
M/parenleftbig
T,(v 1,...,vn)/parenrightbig
=M/parenleftbig
I,(u 1,...,un),(v 1,...,vn)/parenrightbig
.
5. Prove that if Bis a square matrix with complex entries, then there
exists an invertible square matrix Awith complex entries such
thatA−1BAis an upper-triangular matrix.
6. Give an example of a real vector space VandT∈L(V)such that
trace(T2)<0.
7. Suppose Vis a real vector space, T∈L(V), andVhas a basis
consisting of eigenvectors of T. Prove that trace(T2)≥0.
8. Suppose Vis an inner-product space and v,w∈L(V). Define
T∈L(V)byTu=/angbracketleftu,v/angbracketrightw. Find a formula for trace T.
9. Prove that if P∈L(V)satisfiesP2=P, then tracePis a nonneg-
ative integer.
10. Prove that if Vis an inner-product space and T∈L(V), then
traceT∗=traceT.
11. Suppose Vis an inner-product space. Prove that if T∈L(V)is
a positive operator and trace T=0, thenT=0.
Exercises 245
12. Suppose T∈L(C3)is the operator whose matrix is
51−12−21
60−40−28
57−68 1
.
Someone tells you (accurately) that −48 and 24 are eigenvalues
ofT. Without using a computer or writing anything down, find
the third eigenvalue of T.
13. Prove or give a counterexample: if T∈L(V)andc∈F, then
trace(cT)=ctraceT.
14. Prove or give a counterexample: if S,T∈L(V), then trace(ST) =
(traceS)(traceT).
15. Suppose T∈L(V). Prove that if trace (ST)=0 for allS∈L(V),
thenT=0.
16. Suppose Vis an inner-product space and T∈L(V). Prove that
if(e1,...,en)is an orthonormal basis of V, then
trace(T∗T)=/bardblTe1/bardbl2+···+/bardblTen/bardbl2.
Conclude that the right side of the equation above is independent
of which orthonormal basis (e1,...,en)is chosen for V.
17. Suppose Vis a complex inner-product space and T∈L(V). Let
λ1,...,λnbe the eigenvalues of T, repeated according to multi-
plicity. Suppose
a1,1... a 1,n
......
an,1... an,n
is the matrix of Twith respect to some orthonormal basis of V.
Prove that
|λ1|2+···+|λn|2≤n/summationdisplay
k=1n/summationdisplay
j=1|aj,k|2.
18. Suppose Vis an inner-product space. Prove that
/angbracketleftS,T/angbracketright=trace(ST∗)
defines an inner product on L(V).
246 Chapter 10. Trace and Determinant
19. Suppose Vis an inner-product space and T∈L(V). Prove that Exercise 19 fails on
infinite-dimensional
inner-product spaces,
leading to what are
called hyponormal
operators, which have a
well-developed theory.if
/bardblT∗v/bardbl≤/bardblTv/bardbl
for everyv∈V, thenTis normal.
20. Prove or give a counterexample: if T∈L(V)andc∈F, then
det(cT)=cdimVdetT.
21. Prove or give a counterexample: if S,T∈L(V), then det (S+T)=
detS+detT.
22. Suppose Ais a block upper-triangular matrix
A=
A1∗
...
0Am
,
where eachAjalong the diagonal is a square matrix. Prove that
detA=(detA1)...( detAm).
23. Suppose Ais ann-by-n matrix with real entries. Let S∈L(Cn)
denote the operator on Cnwhose matrix equals A, and letT∈
L(Rn)denote the operator on Rnwhose matrix equals A. Prove
that traceS=traceTand detS=detT.
24. Suppose Vis an inner-product space and T∈L(V). Prove that
detT∗=detT.
Use this to prove that |detT|=det√
T∗T, giving a different
proof than was given in 10.37.
25. Leta,b,c be positive numbers. Find the volume of the ellipsoid
/braceleftbig
(x,y,z)∈R3:x2
a2+y2
b2+z2
c2<1/bracerightbig
by finding a set Ω⊂R3whose volume you know and an operator
T∈L(R3)such thatT(Ω)equals the ellipsoid above.
Symbol Index
R,2
C,2
F,3
Fn,5
F∞,1 0
P(F),1 0
−∞,2 3
Pm(F),2 3
dimV,3 1
L(V,W),3 8
I,3 8
nullT,4 1
rangeT,4 3
M/parenleftbig
T,(v 1,...,vn),(w 1,...,wm)/parenrightbig
,
48
M(T),4 8Mat(m,n, F),5 0
M(v),5 2
M/parenleftbig
v,(v
1,...,vn)/parenrightbig
,5 2
T−1,5 4
L(V),5 7
degp,6 6Re, 69
Im, 69
¯z,6 9
|z|,6 9
T|U,7 6
M/parenleftbig
T,(v 1,...,vn)/parenrightbig
,8 2
M(T),8 2
PU,W,9 2
/angbracketleftu,v/angbracketright,9 9
/bardblv/bardbl, 102
U⊥, 111
PU, 113
T∗, 118
⇐⇒, 120√
T, 146
⊊, 166
✽, 231
T(Ω), 239/integraltext
Ωf, 241
Dk, 241
≈, 242
247
Index
absolute value, 69
addition, 9adjoint, 118
basis, 27
block diagonal matrix, 142block upper-triangular matrix,
195
Cauchy-Schwarz inequality,
104
Cayley-Hamilton theorem for
complex vectorspaces, 173
Cayley-Hamilton theorem for
real vector spaces,207
change of basis, 216characteristic polynomial of a
2-by-2 matrix, 199
characteristic polynomial of an
operator on a complex
vector space, 172
characteristic polynomial of an
operator on a realvector space, 206
characteristic value, 77
closed under addition, 13closed under scalar
multiplication, 13
complex conjugate, 69complex number, 2
complex spectral theorem,
133
complex vector space, 10conjugate transpose, 120
coordinate, 4
cube, 238
cube root of an operator,
159
degree, 22
derivative, 241
determinant of a matrix,
229
determinant of an operator,
222
diagonal matrix, 87
diagonal of a matrix, 83
differentiable, 241
dimension, 31
direct sum, 15
divide, 180
division algorithm, 66
dot product, 98
eigenpair, 205
eigenvalue of a matrix, 194
eigenvalue of an operator,
77
eigenvector, 77
249
250 Index
Euclidean inner product,
100
field, 3
finite dimensional, 22
functional analysis, 23
fundamental theorem of
algebra, 67
generalized eigenvector,
164
Gram-Schmidt procedure,
108
Hermitian, 128
homogeneous, 47
identity map, 38
identity matrix, 214
image, 43
imaginary part, 69infinite dimensional, 23
injective, 43
inner product, 99
inner-product space, 100
integral, 241
invariant, 76
inverse of a linear map, 53inverse of a matrix, 214
invertible, 53
invertible matrix, 214
isometry, 147
isomorphic, 55
Jordan basis, 186
kernel, 41
length, 4
linear combination, 22linear dependence lemma,
25
linear functional, 117linear map, 38linear span, 22linear subspace, 13linear transformation, 38linearly dependent, 24linearly independent, 23
list, 4
matrix, 48
matrix of a linear map, 48matrix of a vector, 52matrix of an operator, 82
minimal polynomial, 179
monic polynomial, 179multiplicity of an eigenpair,
205
multiplicity of an eigenvalue,
171
nilpotent, 167
nonsingular matrix, 214norm, 102normal, 130
null space, 41
one-to-one, 43
onto, 44operator, 57
orthogonal, 102
orthogonal complement,
111
orthogonal operator, 148
orthogonal projection, 113orthonormal, 106orthonormal basis, 107
parallelogram equality, 106
Index 251
permutation, 226
perpendicular, 102point, 10
polynomial, 10positive operator, 144positive semidefinite operator,
144
product, 41projection, 92
Pythagorean theorem, 102
range, 43
real part, 69real spectral theorem, 136
real vector space, 9
root, 64
scalar, 3
scalar multiplication, 9
self-adjoint, 128sign of a permutation, 228signum, 228singular matrix, 214singular values, 155span, 22
spans, 22spectral theorem, 133, 136square root of an operator,
145, 159
standard basis, 27subspace, 13sum of subspaces, 14surjective, 44
trace of a square matrix,
218
trace of an operator, 217transpose, 120
triangle inequality, 105
tuple, 4
unitary operator, 148
upper triangular, 83
vector, 6, 10
vector space, 9
volume, 238
Undergraduate Texts inMathematics
‘Contemp
Halmos: Naive SetTheory. Malitz: Introduction toMathematical
‘Hammerlin/Hoffmann: Numerical Logic.Mathematics. Marsden/Weinstein: Calculus1,I,IlReadings inMathematics. Second edition.Harris/Hirst/Mossinghoff: Martin:TheFoundations ofGeometry‘Combinatorics andGraph Theory. andtheNon-Euclidean Plane.
Hartshorne: Geometry: Euclid and Martin: Geometric Constructions.Beyond, Martin:Transformation Geometry: AnHijab: Introduction toCalculus and Introduction toSymmetry.
‘Ciassical Analysis. Millman/Parker: Geometry: AMetric
Hitton/Holton/Pedersen: Mathematical ‘Approach with Models. Second
Reflections: InaRoom with Many editionMirrors Moschovakis: NotesonSetTheorylooss/Joseph: Elementary Stability Owen: AFirst Course inthe
andBifurcation Theory. Second Mathematical Foundations ofedition. ‘Thermodynamics.Isaac: The Pleasures ofProbability Palka: AnIntroduction toComplex
Readings inMathematics Function Theory.
James: Topological andUniform Pedriek: AFirst Course inAnalysisSpaces. Peressini/Sullivan/Uhl: TheMathematicsJanich:LinearAlgebra ‘ofNonlinearProgramming.Sanich: Topology. Prenowitz/Jantoseiak: Join Geometries.
‘anieh: Vector Analysis. Priestley: Calculus: ALiberal Art.
Kemeny/Snell: Finite Markov Chains Second edition.
Kinsey: Topology ofSurfaces Protter/Morrey: AFirst Course inReal
Klambauer: AspectsofCalculus. ‘Analysis.Secondedition Lang: AFirst Course inCalculus. Fifth _Protter/Morrey: Intermediate Calculus.dition. Secondedition.Lang: Calculus ofSeveral Variables. Roman: AnIntroduction toCoding and
Third edition. Information Theory.
Lang: Introduction toLinear Algebra Ross: Elementary Analysis: TheTheory
‘Second edition ofCalculus
Lang: Linear Algebra. Third edition ‘Samuel: Projective Geometry
Lang: Undergraduate Algebra. Second Readings inMathematics.edition. Scharlaw/Opotka: FromFermattoLang: Undergraduate Analysis. Minkowski.
Lax/Burstein/Lax: Calculus with ‘Schiff: TheLaplace Transform: Theory
Applications andComputing. andApplications
Volume | ‘Sethuraman: Rings, Fields, andVector
LeCuyer: College Mathematics with ‘Spaces: AnApproachtoGeometric APL.Constructability.
LidvPitz: Applied Abstract Algebra. Sigler: Algebra.
‘Second edition Sllverman/Tate: Rational Points on
Logan: Applied Partial Differential Elliptic Curves.Equations, Simmonds: ABriefonTensorAnalysis. ‘Macki-Strauss: Introduction toOptimal ‘Second edition.
Control Theory,