allen matrix notes
PDF · 253 pages · 2.0 MB
Open PDF file
Textbook-style notes on linear algebra and matrices, apparently by an author named Allen and stored among downloaded math books rather than as Phil's own work. The opening chapter defines vector spaces and gives examples (Rn, polynomial spaces, sequence spaces). It then covers subspaces, spans, linear dependence and independence, bases and the extension-to-a-basis theorem, with proofs. The text shown ends partway through the section on dimension.
AI-written summary; may contain errors. This description is approximate.
Extracted text (machine-read; may contain errors)
1
Chapter 1
Vectors and Vector Spaces
1.1 Vector Spaces
Underlying every vector space (to be de fined shortly) is a scalar fieldF.
Examples of scalar fields are the real and the complex numbers
R:= real numbers
C:= complex numbers.
These are the only fields we use here.
Definition 1.1.1. Avector space Vis a collection of objects with a (vector)
addition and scalar multiplication de fined that closed under both operations
and which in addition satis fies the following axioms:
(i) (α+β)x=αx+βxfor all x∈Vandα,β∈F
(ii)α(βx)=(αβ)x
(iii)x+y=y+xfor all x, y∈V
(iv)x+(y+z)=(x+y)+zfor all x, y, z∈V
(v)α(x+y)=αx+αy
(vi)∃O∈Vz0+x=x; 0 is usually called the origin
(vii) 0 x=0
(viii) ex=xwhere eis the multiplicative unit in F.
7
8 CHAPTER 1. VECTORS AND VECTOR SPACES
The “closed” property mentioned above means that for all α,β∈Fand
x, y∈V
αx+βy∈V
(i.e. you can’t leave Vusing vector addition and scalar multiplication). Also,
when we write for α,β∈Fandx∈V
(α+β)x
the ‘+’ is in the field, whereas when we write x+yforx, y∈V,t h e‘ + ’i s
in the vector space. There is a multiple usage of this symbol.
Examples.
(1)R2={(a1,a2)|a1,a2∈R}two dimensional space.
(2)Rn={(a1,a2,... ,a n)|a1,a2,... ,a n∈R},n dimensional space.
(a1,a2,... ,a n)i sc a l l e da n n-tuple.
(3)C2andCnrespectively to R2andRnwhere the underlying field isC,
the complex numbers.
(4)Pn=l
n
j=0ajxj|a0,a1,... ,a n∈RM
is called the polynomial space of
all polynomials of degree n. Note this includes not just the polynomials
of exactly degree nbut also those of lesser degree.
(5)fp={(ai,...)|ai∈R,Σ|ai|p<∞}. This space is comprised of
vectors in the form of in finite-tuples of numbers. Properly we would
write
fp(R)o rfp(C)
to designate the field.
(6)TN=FN
n=1ansinnπx|a1,... ,a n∈Rk
, trigonometric polynomials.
Standard vectors in Rn
e1=( 1,0,... , 0)
e2=( 0,1,0,... , 0)
e3=( 0,0,1,0,... , 0)
...
en=( 0,0,... , 0,1)These are the
unit∗vec-
tors whichpoint in thenorthogonal
∗
directions.
1.1. VECTOR SPACES 9
∗Precise de finitions will be given later.
ForR2, the standard vectors are
10 CHAPTER 1. VECTORS AND VECTOR SPACES
e1=( 1,0)
e2=( 0,1) (1,0)(0,1)
(0,0)12
e
Graphical representa-
tion of e1ande2in the
usual two dimensional
plane.
Recall the usual vector addition in the plane uses the parallelogram rule
yx+y
ForR3, the standard vectors are
e1=( 1,0,0)
e2=( 0,1,0)
e3=( 0,0,1)(0,0,1)
(1,0,0)ee
e(0,1,0)23
1
Graphical representa-
tion of e1,e2,a n d e3in
the usual
Linear algebra is the mathematics of v ector spaces and their subspaces. We
will see that many questions about vector spaces can be reformulated asquestions about arrays of numbers.
1.1.1 Subspaces
LetVbe a vector space and U⊂V.W e w i l l c a l l Uasubspace ofVifU
is closed under vector addition, scalar multiplication and satis fies all of the
vector space axioms. We also use the term linear subspace synonymously.
1.1. VECTOR SPACES 11
Examples. Proofs will be given later
letV=R3={(a, b, c )|a, b, c∈R} (1.1)
U={(a, b,0)|a, b∈R}.
Clearly U⊂Vand also Uis a subspace of V.
let v1,v2∈R3(1.2)
W={av1+bv2|a, b∈R}
Wis a subspace of R3.
In this case we say Wis “spanned” by {v1,v2}. In general, let S⊂V,a
vector space, have the form
S={v1,v2,... ,v k}.
Thespan ofSis the set
U=
k3
j=1ajvj|a1,... ,a k∈R
.
We will use the notion
S(v1,v2,... ,v k)
for the span of a set of vectors.
Definition 1.1.2. We say that
u=a1v1+···+akvk
is alinear combination of the vectors v1,v2,... ,v k.
Theorem 1.1.1. LetVbe a vector space and U⊂V.I fUis closed under
vector addition and scalar multiplication, then Uis a subspace of V.
Proof. We remark that this result provides a “short cut” to proving that a
particular subset of a vector space is in fact a subspace. The actual proofof this result is simple. To show (i), note that if x∈Uthen x∈Vand so
(ab)x=ax+bx.
Nowax, bx, ax +bxand ( a+b)xareallinUby the closure hypothesis. The
equality is due to vector space properties of V.T h u s( i )h o l d sf o r U.E a c h
of the other axioms is proved similarly.
12 CHAPTER 1. VECTORS AND VECTOR SPACES
A very important corollary follows about spans.
Corollary 1.1.1. LetVbe a vector space and S={v1,v2,... ,v k}⊂V.
ThenS(v1,... ,v k)is a linear subspace of V.
Proof. We merely observe that
S(v1,... ,v k)=lk3
1ajvj|a1,... ,a k∈RorCM
.
This means that the closure is built right into the de finition of span. Thus,
if
v=a1v1+···+akvk
w=b1v1+···+bkvk
then both
v+w=(a1+b1)v1+···+(ak+bk)vk
and
cv=ca1v+ca2v+···+cakv
are in U.T h u s Uis closed under both operations; therefore Uis a subspace
ofV.
Example 1.1.1. (Product spaces.) Let VandWbe vector spaces de fined
over the same field. We de fine the new vector space Z=V×Wby
Z={(v, w)|u∈V, w∈W}
We de fine vector addition as ( v1,w1)+(v2,w2)=( v1+v2,w1+w2)a n d
scalar multiplication by α(v, w)=(αv,αw). With these operations, Zis a
vector space, sometimes called the product ofVandW.
Example 1.1.2. Using set-builder notation, de fineV13={(a,0,b)|a, b,∈
R}.Then Uis a subspace of R3.It can also be realized as the subspace of
the standard vectors e1=( 1,0,0) and e3=( 0,0,1), that is to say V13=
S(e1,e3).
1.2. LINEAR INDEPENDENCE AND LINEAR DEPENDENCE 13
Example 1.1.3. More subspaces of R3.There are two other important
methods to construct subspaces of R3. Besides the set builder notation
used above, we have just considered the method of spanning sets. Forexample, let S={v
1,v2}⊂R3.ThenS(S) is a subspace of R3.Simi-
larly, if T={v1}⊂R3.ThenS(T) is a subspace of R3.At h i r dw a y
to construct subspaces is by using inner products. Let x, w∈R3.Ex-
p r e s s e di nc o o r d i n a t e s x=(x1,x2,x3)a n d w=(w1,w2,w3).Define
the inrner product of xandwbyx·w=x1w1+x2w2+x3w3.Then
Uw={x∈R3|x·w=0}is a subpace of R3. To prove this it is neces-
sary to prove closure under vector addition and scalar multiplication. Thelatter is easy to see because the inner product is homogeneous in α,that is,
(αx)·w=αx
1w1+αx2w2+αx3w3=α(x·w).Therefore if x·w=0s o
also is (αx)·w.The additivity is also straightforward. Let x, y∈U.T h e n
the sum
(x+y)·w=(x1+y1)w1+(x2+y2)w2+(x3+y3)w3
=(x1w1+x2w2+x3w3)+(y1w1+y2w2+y3w3)
=0 + 0 = 0
However, by choosing two vectors v,w,∈R3we can de fineUv,w={x∈
R3|x·y=0a n d x·w=0}.E s t a b l i s h i n g Uv,wis a subspace of R3is proved
similarly. In fact, what is that both these sets of subspaces, those formedby spanning sets and those formed from the inner products are the same set
of subspaces. For example, referring to the previous example, it follows that
V
13=S(e1,e3)=Ue2. Can you see how to correspond the others?
1.2 Linear independence and linear dependence
One of the most important problems in vector spaces is to determine if
a given subspace is the span of a collection of vectors and if so, to deter-
mine a spanning set. Given the importance of spanning sets, we intend to
examine the notion in more detail. In particular, we consider the conceptof uniqueness of representation.
LetS={v
1,... ,v k}⊂V, a vector space, and let U=S(v1,... ,v k)( o r
S(S) for simpler notation). Certainly we know that any vector v∈Uhas
the representation
v=a1v1+···+akvk
for some set of scalars a1,... ,a k. Is this representation unique? Or,c a nw e
14 CHAPTER 1. VECTORS AND VECTOR SPACES
find another set of scalars b1,... ,b kn o ta l lt h es a m ea s a1,... ,a krespec-
tively for which
v=b1v1+···+bkvk.
We need more information about Sto answer this question either way.
Definition 1.2.1. LetS={v1,... ,v k}⊂V, a vector space. We say that
Sislinearly dependent (l.d.) if there are scalars a1,... ,a knot all zero for
which
a1v1+a2v2+···+akvk=0. (T)
O t h e r w i s ew es a y Sislinearly independent (l.i.).
Note. If we allow all the scalars to be zero we can always arrange for ( T)
to hold, making the concept vacuous.
Proposition 1.2.1. IfS={v1,... ,v k}⊂V, a vector space, is linearly
dependent, then one member of this set can be expressed as a linear combi-
nation of the others.
Proof. We know that there are scalars a1,... ,a ksuch that
a1v1+a2v2+···+akvk=0
Since not all of the coe fficients are zero, we can solve for one of the vectors
as a linear combination of the other vectors.
Remark 1.2.1. Actually we have shown that there is novector with a
unique representation in S(S).
Corollary 1.2.1. If0∈S={v1,... ,v k},t h e n Sis linearly dependent.
Proof. Trivial.
Corollary 1.2.2. IfS={v1,... ,v k}is linearly independent then every
subset of Sis linearly independent.
1.3. BASES 15
1.3 Bases
The idea of a basis is that of finding a minimal generating set for a vector
space. Through basis, unicity of representation and a number of other usefulproperties, both theoretical and com putational, can be concluded. Thinking
of the concept in operations research ideas, a basis will be a redundancy freeand complete generating set for a vector space
Definition 1.3.1. LetVbe a vector space and S={v
1,... ,v k}⊂V.W e
callSaspanning set for the subspace U=S(S).
Suppose that Vis a vector space, and S={v1,... ,v k}is a linearly
independent spanning set for V.T h e n Sis called a basis ofV.M o d i f yt h i s
definition correspondingly for subspaces.
Proposition 1.3.1. IfSis a basis of V, then every vector has a unique
representation.
Proof. LetS={v1,... ,v k}andv∈V.T h e n
v=a1v1+···+akvk
for some choice of scalars. If there is a second choice of scalars b1,... ,b k
not all the same, respectively, as a1,... ,a k,w eh a v e
v=b1v1+···+bkvk
and
0=(a1−b1)v1+···+(ak−bk)vk.
Since not all of the di fferences a1−b1,... ,a k−bkare zero we must have
thatSis linearly dependent. This is a contradiction to our hypothesis, and
the result is proved.
Example. LetV=R3andS={e1,e2,e3}.T h e n Sis a basis for V.
Proof. Clearly Vis spanned by S. Now suppose that
0=a1e1+a2e2+a3e3
or
(0,0,0) =a1(1,0,0) +a2(0,1,0) +a3(0,0,1)
=(a1,a2,a3).
Hence a1=a2=a3=0 . T h u st h es e t {e1,e2,e3}is linearly independent.
16 CHAPTER 1. VECTORS AND VECTOR SPACES
Remark 1.3.1. Note how we resolved the linearly dependent/linearly in-
dependent issue by converting a vector problem to a numbers problem. Thisis at the heart of linear algebra.
Exercise. LetS={v
1,v2}={(1,0,1),(1,−1,0)}⊂R3. Show that Sis
linearly independent and therefore a basis of S(S).
1.4 Extension to a basis
In this section, we show that given a linearly independent set of vectors
from a vector space with a finite spanning set, it is possible add to this set
more vectors until it becomes a basis. Thus any set of linearly independent
vectors can be a part (subset) of a basis.
Theorem 1.4.1 (Extension to a basis). Assume that the given vector space
Vhas a finite spanning set S1, i.e. V=S(S1).L e t S0={x1,... ,x f}be a
linearly independent subset of Vso that S(S0)V. Then, there is a subset
SI
1ofS1,s u c ht h a t S0∪SIis a basis for V.
Proof. Our intention is to add vectors to S0keeping it linearly independent
and eventually becoming a basis. There are a couple of steps.Steps.
1. Since S(S
1)SS(S0), there is a vector y1∈S1such that S0,1=
{S0,y1}is linearly independent and thus S(S0,1)SS(S0).
2. Continue this process generating sets
S0,1={S0,y1}
S0,2={S0,1,y2}
...
S0,j={S0,j−1,yj−1}
...
At each step S0,1,S0,2,... are linearly independent sets. Since S1is
finite we must eventually have that
S(S0,m)=S(S1)=V.
3. Since S0,mis linearly independent and spans V, it must be a basis.
1.5. DIMENSION 17
Remark 1.4.1. In the proof it was important to begin with anyspanning
set for Vand to extract vectors from it as we did. Assuming merely that
there exists a finite spanning set and extracting vectors directly from V
leads to a problem of terminus. That is, when can we say that the new
linearly independent set being generated in Step 2 above is a spanning set
forV? What we would need is a theorem that says something to the e ffect
that if Vhas a finite basis, then every linearly independent set having the
same number of vectors is also a basis. This result is the content of the nextsection. However, to prove it we need the Extension theorem.
Corollary 1.4.1. IfS={v
1,... ,v k}is linearly dependent then the repre-
sentation of vectors in S(S)isnotunique.
Proof. We know there are scalars a1,... ,a knot all zero, for which
a1v1+···+akvk=0
letv∈S(S) have the representation
v=b1v1+b2v2+···+bkvk.
Then we also have the representation
v=(a1+b1)v1+(a2+b2)v2+···+(ak+bk)vk
establishing the result.
Remark 1.4.2. The upshot of this construction is that we can always con-
struct a basis from a spanning set. In actual practice this process may bequite difficult to carry out. In fact, we will spend some time achieving this
goal. The main tool will be matrix theory.
1.5 Dimension
One of the most remarkable features of vector spaces is the notion of
dimension. We need one simple result that makes this happen, the basis
theorem.
Theorem 1.5.1 (Basis Theorem). LetS={v1,... ,v k}⊂Vbe a basis
forV. Then every basis of Vhaskelements.
18 CHAPTER 1. VECTORS AND VECTOR SPACES
Proof. We proceed by induction. Suppose S={v1}andT={w1,w2}are
both bases of V.T h e ns i n c e Sis a basis
w1=α1v1w2=α2v1
and therefore
1
α1w1−1
α2w2=0
which implies that Tis linearly dependent (we tacitly assumed that both
α1andα2were nonzero. Why can we do this?)
The next step is to assume the result holds for bases having up to k
elements. Suppose that S={v1,... ,v k+1}andT={w1,... ,w k+2}are
both bases of V. Now consider SI={v1,... ,v k}. We know that S(SI)
S(S)=S(T)=V. By our extension of bases result, there is a vector
wf1∈Tsuch that
SI
1={v1,... ,v k,wf1}
is linearly independent and S(SI
1)⊂S(S)=V.I fS(SI
1)V, our extension
result applies again to give a vector vf1such that
SI
11={v1,... ,v k,wf1,vf1}
is linearly independent The only possible selection is vf1=vk+1. But in this
casewfiwill depend on v1,... ,v k,vk+1, and that is a contradiction. Hence
S(v1,... ,v k,wf1)=V.
The next step is to remove the vector vkfrom SI
1and apply the extension
to conclude that the span of the set
SI
2={v1,... ,v k−1,wf1,wf2}
isV. We continue in this way eventually concluding that
SI
k+1={wf1,wf2,... ,w fk+1}
has span V.B u t SI
k+1T, whence Tis linearly dependent.
Proposition 1.5.1 (Reduced spanning sets). (a) Suppose that S=
{v1,... ,v k}spans Vandvjdepends (linearly) on
Sj={v1,... ,v j−1,vj+1...v k}.
Then Sjalso spans V.
1.5. DIMENSION 19
( b )I fa tl e a s to n ev e c t o ri n Sis nonzero (that is VW={0}, the smallest
vector space), then there is a subset S0⊂Sthat is linearly independent
and spans V.
Proof. (Left to reader.)
Definition 1.5.1. Thedimension of a vector space Vis the (unique) num-
b e ro fv e c t o r si nab a s i so f V.W ew r i t ed i m ( V) for the dimension.
Remark 1.5.1. This de finition make sense possible only because of our
basis theorem from which we are assured all every linearly independentspanning sets of V, that is all bases, have the same number of elements.
Examples.
(1) dim( R
n)=n,
(2) dim( Pn)=n+1 .
Exercise. LetM= all rectangular arrays of two rows and three columns
with real entries. Find a basis for M,a n d find the dimension of M.N o t e
M=F}abc
def]eeeea, b, c, d, e, f ∈Rk
Example 1.5.1. P
n={anxn+an−1xn−1+···+a1x+a0=0 }is the
vector space of polynomials of degree n. We claim that the powers, x0=1 ,
x, x2,... ,xnare linearly independent, and since
Pn=S(1,x ,... ,xn)
they form a basis of Pn.
Proof. There are several ways we can prove this fact. Here is the most
direct and it requires essentially no machinery. Suppose they are linearly
dependent, which means that there are coe fficients a0,a1,... ,a nso that
anxn+an−1xn−1+···+a1x+a0=0, (T)
the function . (This functional view is critically important because every
polynomial has roots.) There must be a coe fficient which is nonzero and
which corresponds to the highest power. Let us assume that anW=0 ,f o r
convenience, and with no loss in generality.
20 CHAPTER 1. VECTORS AND VECTOR SPACES
Solve for xnto get
xn=−an−1
anxn−1+···+−a1
anx−a0
an(TT)
Now compute the ratio of this expression divided by xnon both sides, and
letx→∞ . The left side of course will be 1. Again for convenience we take
n= 2. So, condensing terms we will have
b1x+b0
x2=b1w1
xW
+b0w1
x2W
=1
where bj=−aj/a2.B u t a s x→∞ the expression b1D1
xi
+b0D1
x2i
→0.
This is a contradiction. It cannot be that the functions 1 ,x,a n d x2are
linearly dependent.
In the general case for nwe have
bn−1w1
xW
+bn−2w1
x2W
+···+b0w1
xnW
=1,
where bj=−aj/an. Apply the same limiting argument to obtain the con-
tradiction. Thus
T={1,x ,... ,xn}
is a basis of Pn.
A calculus proof is available. It is also based on the fact that if the
powers are linearly independent and ( T) holds, then we can assume that the
same relation ( TT)i st r u e . N o wt a k et h e nthderivative of both sides. We
obtain
n!=0
a contraction, and the result if proved
Finally, one more technique used to prove this result is by using the
Fundamental Theorem of Algebra.
Theorem 1.5.2. Every polynomial ( T)o f exactly nthdegree (i.e. with
anW=0) has exactly nroots counted with mu ltiplicity (i.e. if q(x)=qnxn+
qn−1xn−1+···+q1x+q0∈Pn(C),qnW=0 t h e nt h en u m b e ro fs o l u t i o n so f
q(x)=0 is exactly n).
From (T) above we have an nthdegree polynomial that is zero for every
x. Thus the polynomial is zero, and this means allthe coefficients are
zero. This is a contradiction to the hypothesis, and therefore the theorem
is proved.
1.5. DIMENSION 21
Remark 1.5.2.
P0P1P2···Pn···.
On the other hand this is not true for the Euclidean spaces R1,R2,... .
However, we may say that there is a subspace of R3which is “like” R2in
every possible way. Do you see this? We have
R2={(a, b)|a, b∈R}
R3={(a, b, c )|a, b, c∈R}.
No element in R2, an ordered pair,c a nb ei n R3, a set of ordered triples.
However
U={(a, b,0)|a, b∈R}
is “like” R2is just about every way. Later on we will give a precise mathe-
matical meaning to this comparison.
Example 1.5.2. Find a basis for the subspace V0ofR3of all solutions to
x1+x2+x3=0 ( T)
where x=(x1,x2,x3)∈R3.
Solution. First show that the set V0={(x1,x2,x3)∈R3|x1+x2+x3=0}
is in fact a subspace of R3. Clearly if x=(x1,x2,x3)∈V0andy=
(y1,y2,y3)∈V0then x+y=(x1+y1,x2+y2,x3+y3)∈V0,p r o v i n g
closure under vector addition. Similarly V0is closed under scalar multi-
plication. Next, we seek a collection of vectors v1,v2,... ,v k∈V0so that
S(v1,... ,v k)=V0.L e t x3=αandx2=βbe free parameters. Then
x1=−(α+β).
Hence all solutions of ( T)h a v et h ef o r m
x=(−(α+β),β,α)
x=α(−1,0,1) +β(−1,1,0).
Obviously the vectors v1=(−1,0,1) and v2=(−1,1,0) are linearly inde-
pendent, and xis expressed as being in the span of them. So, V0=S(v1,v2).
V0has dimension 2.
22 CHAPTER 1. VECTORS AND VECTOR SPACES
Theorem 1.5.3 (Uniqueness). LetS={v1,... ,v k}be a basis of V.
Then each vector v∈Vhas a unique representation with respect to S.
Proof. SinceS(S)=Vwe have that
v=a1v1+a2v2+···+akvk
for some coe fficients a1,a2,... ,a kin the given field. (This is the represen-
tation of vwith respect to S.) If it is notunique there is another
v=b1v1+b2v2+···+bkvk.
So, subtracting we have
(a1−b1)v1+(a2−b2)v2+···+(ak−bk)vk=0
where the di fferences aj−bjare not all zero. This implies that Sis a linearly
dependent set.
Theorem 1.5.4. Suppose that S={v1,... ,v k}is a basis of the vector
space V. Suppose that T={w1,... ,w m}is a linearly independent subset
ofV.T h e n m≤k.
Proof. We know that Sis a linearly independent spanning set. This means
that every linearly independent set of kvectors is also a spanning set. There-
fore,m>k renders a contradiction as T0={w1,... ,w k}is a spanning set
andwk+1∈S(T0).
Definition 1.5.2. IfAis any set we de fine
|A|:= cardinality of A,
that is to say |A|is the number of elements of A.
Example 1.5.3. LetT={1,x ,x2,x3}.T h e n |T|=4 .
Theorem 1.5.5. Both RkandCkarek-dimensional and Sk={e1,e2,... ,e k}
is a basis of both.
Proof. It is easy to see that e1,... ,e kare linearly independent, and any
vector xinRkhas the form
x=a1e1+a2e2+···+akek
fora1,... ,a k∈R.T h u s Skis a linearly independent spanning set and
hence a basis of Rk.
1.6. NORMS 23
Question: What single change to the proof above gives the theorem for Ck?
The following results follow easily from previous results.
Theorem 1.5.6. LetVbe ak-dimensional vector space.
(i) Every set Twith |T|>k is linearly dependent.
(ii) If D={v1,... ,v j}is linearly independent and j<k , then there are
vectors vf1,... ,v fk−j∈Vsuch that
D∪{vf1,... ,v fk−j}
is a basis of V.
(iii) If D⊂V,|D|=k,a n d Dis either a spanning set for Vor linearly
independent, then Dis a basis for V.
1.6 Norms
Norms are a way of putting a measure of distance on vector spaces. The
purpose is for the re fined analysis of vector spaces from the viewpoint of
many applications. It is also to all the comparison of various vectors on
the basis of their length. Ultimately , we wish to discuss vector spaces as
representatives of points. Naturally, we are all accustomed to the “shortestdistance” distance from the Pythagorean theorem. This is an example of anorm, but we shall consider them as real valued functions with very special
properties.
Definition 1.6.1. Norms on vector spaces over C,orR.L e t Vbe a vec-
tor space and suppose that ,·,:V→R
+is a function from Vto the
nonnegative reals for which
(i),x,≥0 for all x∈Vand
,x,=0i fa n do n l yi f x=0
(ii),αx,=|α|,x,for allα∈C,Randx∈V
(iii),x+y,≤, x,+,y,for all x, y∈V “The Triangle inequality”.
Then,·,is called a norm onV. The second condition is often termed the
(positive) homogeneity property.
Remark 1.6.1. The notation is a substitute function notation. The ex-
pression ,·,, without the vector, is just the way a norm is expressed.
24 CHAPTER 1. VECTORS AND VECTOR SPACES
Examples. LetV=Rn(orCn). De fine for x=(x1,... ,x n)
(i),x,2=(|x1|2+|x2|2+···+|xn|2)1/2Euclidean norm
(ii),x,1=(|x1|+|x2|+···+|xn|)f1norm
(iii),x,∞=m a x
1≤i≤n|xi|f∞norm
Norm (ii) is read as: ell one norm. Norm (iii) is read as: ell in finity norm.
Proof that (ii) is a norm. Clearly (i) holds. Next
,αx,1=(|αx1|+|αx2|+···+|αxn|)
=(|α||x1|+|α||x2|+···+|α||xn|)
=|α|(|x1|+|x2|+···+|xn|)=|α|,x,1
which is what we needed to prove. Also,
,x+y,1=(|x1+y1|+|x2+y2|+···+|xn+yn|)
≤(|x1|+|y1|+|x2|+|y2|+···+|xn|+|yn|)
=(|x1|+|x2|+···+|xn|)+( |y1|+|y2|+···+|yn|)
=,x,1+,y,1.
Here we used the fact that |α+β|≤|α|+|β|for numbers.
To prove that (i) is a norm we need a very famous inequality.
Lemma 1.6.1 (Cauchy—Schwartz). Given that a1,... ,a nandb1,... ,b n
are in C.T h e n
n3
1|aibi|≤Xn3
1a2
i~1/2Xn3
1b2
i~1/2
. (T)
Proof. We consider for the variable t
Σ(ai+tbi)2=Σa2
i+2tΣaibi+t2Σb2
i.
Note that ( T) is obvious ifn
1aibi=0 . I fn o tt a k e
t=−n
1a2
i
n
1aibi.
1.6. NORMS 25
Then
Σ(ai+tbi)2=Σa2
i−2Σa2
i
ΣaibiΣaibi+D
Σa2
ii2
(Σaibi)2Σb2
i
=−Σa2
i+D
Σa2
ii2Σb2
i
(Σaibi)2
=D
Σa2
iiw
−1+Σa2
iΣb2
i
(Σaibi)2W
.
Since the left side is ≥0 and since Σa2
i≥0, we must have that
w
−1+Σa2
iΣb2
i
(Σaibi)2W
≥0.
Solving this inequality we have
(Σaibi)2≤Σa2
iΣb2
i.
Now that square roots to get the result.
To prove that (i) is a norm, we note that conditions (i) and (ii) are
straightforward. The truth of condition (iii) is a consequence of anotherfamous result.
Theorem 1.6.1 (Minkowski). ,x+y,
2≤,x,2+,y,2.
Proof.
Σ(ai+bi)2=Σa2
i+2Σaibi+Σb2
i
≤Σa2
i+2D
Σa2
ii1/2D
Σb2
ii1/2+Σb2
i
=pD
Σa2
ii1/2+D
Σb2
ii1/2Q2
.
Taking square roots gives the result.
Continuity and Equivalence of Norms
Lemma 1.6.2. Every vector norm on Cnis continuous in the vector com-
ponents.
26 CHAPTER 1. VECTORS AND VECTOR SPACES
Proof. Letx∈Cnand,·,some norm on Cn. We need to show that if the
vectorδ→0, in components, then ,x+δ,→, x,. First, by the triangle
inequality
,x+δ,≤, x,+,δ,or
,x+δ,−,x,≤,δ,
Similarly
,x,≤, x+δ−δ,
≤,x+δ,+,δ,or
−,δ,≤, x+δ,−,x,
Therefore
|,x+δ,−,x,|≤,δ,
Now expressing δin components and standard bases vectors, we write δ=
δ1e1+···+δnenand
,δ,≤ |δ1|,e1,+···+|δn|,en,
≤max
1≤i≤n|δi|(,e1,+···+,en,)
≤Mmax
1≤i≤n|δi|
where M=,e1,+···+,en,.We know that if δ→0i nc o m p o n e n t s ,t h e n
max 1≤i≤n|δi|→0.Therefore |,x+δ,−,x,|→0, as well.
Definition 1.6.2. Let,·,aand,·,bbe two vector norms on Cn.We say
that these norms are equivalent if there are postive constants m, M such
that for all x∈Cn
m,x,a≤,x,b≤M,x,a
The remarkable fact about vector norms on Cnis that they are allequiv-
alent. The only tool we need to prove this is the following result: Every
continuous function on a compact set of Cnassumes its maximum (and
minimum) on that set. The term “compact” refers to a particular kind of
setK, one which is both bounded and closed. Bounded means that for
max x∈K,x,≤B<∞a n dc l o s e dm e a n st h a ti fl i m n→∞xn=x,t h e n
x∈K.
1.6. NORMS 27
Theorem 1.6.2. All norms on Cnare equivalent.
Proof. Since equivalence of norms is an equivalence condition, we can take
one of the norms to be the in finity norm ,·,∞.Denote the other norm by
,·,.N o w d e fineK={x|,x,∞=1 }.This set, called the unit ball in the
infinity norm, is compact. Now we de fine
m=m i n
x∈K,x,and
M=m a x
x∈K,x,
Since,·,is a continuous function on K(from the lemma above) and since
Kis compact, we have that both the minimum and maximum are attained
by speci fic vectors in K. Since these vectors are nonzero (they’re in K)a n d
because ,x,is positive for nonzero vectors, it must follow that 0 <m<
M<∞. Hence, on K,t h er e l a t i o n
m,x,≤, x,∞≤M,x,
holds true. For any vector x∈Cnwe can write x=wx
,x,∞W
,x,∞and
x
,x,∞∈K.Thus
mEEEEx
,x,∞EEEE≤EEEEx
,x,∞EEEE
∞≤MEEEEx
,x,∞EEEE
mEEEEx
,x,∞EEEE,x,
∞≤EEEEx
,x,∞EEEE
∞,x,∞≤MEEEEx
,x,∞EEEE,x,
∞
m,x,≤, x,∞≤M,x,
and the theorem is proved.
Example 1.6.1. Example. Find the estimates for the equivalence of ,·,2
and,·,∞
Solution. Letx∈Cn. Then, because we know for any finite sequences
28 CHAPTER 1. VECTORS AND VECTOR SPACES
thatn
i=1|aibi|≤max 1≤i≤n|ai|n
i=1|bi|
,x,2=Xn3
i=1|xi|2~1
2
≤Xn3
i=11·|xi|2~1
2
≤w
max
1≤i≤n|xi|2W1/2Xn3
i=11~1
2
=n1
2,x,∞
On the other hand, by the Cauchy-Schwartz inequality
,x,∞=m a x
1≤i≤n|xi|
≤n3
i=1|xi|
≤Xn3
i=11~1
2Xn3
i=1|xi|2~1/2
=n1
2,x,2
Putting these inequalities together we have
n−1
2,x,2≤,x,∞≤n1
2,x,2
This makes m=n−1/2andM=n1
2.
Remark 1.6.2. Note that the constants mandMdepend on the dimension
of the vector space. Though not the rule in all cases, it is mostly thesituation.
Norms on polynomial spaces
Polynomial spaces, as we have considered earlier, can be given norms as well.Since they are function spaces, our norms usually need to consider all the
values of the independent variable. In many, though not all, cases we need
1.6. NORMS 29
to restrict the domains of the polynomials. With that in mind we introduce
the notation
Pk(a, b)=Pkwith domain restricted to the interval [ a, b]
We now de fine the function versions of the same three norms we have just
studied. For functions p(x)i nPk(a, b)w ed e fine
1.,p(x),2=D$b
a|p(x)|2dxi1
2
2.,p(x),1=$b
a|p(x)|dx
3.,p(x),∞=m a x
a≤x≤b|p(x)|
The positivity and homogeneity properties are fairly easy to prove. The
triangle property is a little more involved. However, it has essentially beenproved for the earlier norms. In the present case, one merely “integrates”
over the inequality. Sometimes ,·,
2is called the energy norm.
T h ei n t e g r a ln o r m sa r er e a l l yt h e norm for polynomial spaces. Alternate
norms use pointwise evaluation or even derivatives depending on the appli-cation. Here is a common type of norm that features the first derivative.
Forp∈P
n(a, b)d efine
,p,=m a x
a≤x≤b|p(x)|+m a x
a≤x≤beepI(x)ee
As is evident this norm becomes large not only when the polynomial is large
but also when its derivative is large. If we remove the term max a≤x≤b|p(x)|
from the norm above and de fine
N(p)= m a x
a≤x≤beepI(x)ee
This function satis fies all the norm properties except one and thus is not a
norm. (See the exercises.) Point evaluation-type norms take us too far a field
of our goals partly because making poi nt evaluations into norms requires
some knowledge of interpolation and related topics. Leave it said that the
obvious point evaluation functions such as p(a)a n dt h el i k ew i l ln o tp r o v i d e
us with norms.
30 CHAPTER 1. VECTORS AND VECTOR SPACES
1.7 Ordered Bases
Given a vector space Vwith a basis S={v1,v2,... ,v k}we now know
that every vector v∈Vhas a representation with respect to the basis
v=a1v1+a2v2+···+akvk.
But no order is implied. For example, for R2we have S={e1,e2}={e2,e1}
shows us that there is no particular order convey through the de finition of
a basis. When we place an order on a basis we will notice an underlyingalgebraic structure of all k-dimensional vector spaces.
Definition 1.7.1. LetVbe a k-dimensional vector space with basis S=
{v
1;v2;...;vk}is speci fied with a fixed and well de fined order as indicated
by their relevant positions. Then Sis called an ordered basis. With or-
dered bases we obtain coordinates .L e t Vbe a vector space of dimension
kwith ordered basis S, and suppose v∈V.F o r 1 ≤i≤k,w ed e fine
theithcoordinate ofvwith respect to Sto be the ithcoefficient aiin the
representation
v=a1v1+a2v2+···+aivi+···+akvk.
In this way we can associate each v∈Vwith a k-tuple of numbers
(a1,a2,... ,a k)∈Rkthat are the coe fficients of vwith respect to S.T h e
k-tuple is unique, owing to the fixed ordering of S. Conversely, for each
ordered k-tuple ( a1,a2,... ,a k)∈Rkthere is associated a unique vector
v∈Vgiven by v=a1v1+a2v2+···+aivi+···+akvk.
We will express this association as
v∼(a1,a2,... ,a k)
The following properties are each simple propositions:
•Ifv∼(a1,a2,... ,a k)a n d w∼(b1,b2,... ,b k)t h e n
v+w∼(a1+b1,a2+b2,... ,a k+bk)
•Ifv∼(a1,a2,... ,a k)a n dα∈R(orC), then
αv∼α(a1,a2,... ,a k)=(αa1,αa2,... ,αak)
•Ifv∼(a1,a2,... ,a k)=( 0 ,0,...0), then v=0 .
1.7. ORDERED BASES 31
We now de fine a special type of linear function from one linear space to
another. The special condition is linearity of the map.
Definition 1.7.2. LetVandWbe two vector spaces. We say that Vand
Warehomomorphic if there is a mapping Φbetween VandWfor which
(1.) For vandwinV
Φ(v+w)=Φ(v)+Φ(w)
(2.) For vandα∈R(orC)
Φ(αv)=αΦ(v)
In this case we call Φahomomorphism from VtoW.F u r t h e r m o r e ,w es a y
thatVandWare isomorphic if they are homomorphic and if
(3.) For each w∈Wthere exists a unique v∈Vsuch that
Φ(v)=w
In this case we call Φaisomorphism from VtoW.
We put all this together to show that finite dimensional vector spaces
over the reals (resp. complex numers) and the standard Euclidean spaces
Rk(resp. Ck) are very, very closely related. Indeed from the point of view
of isometry, they are identical.
Theorem 1.7.1. IfVis ak-dimensional vector space over R(respC), then
Vis isomorphic to Rk(resp. Ck).
This constitutes the beginning of the su fficiency of matrix theory as a
tool to study finite dimentsional vector spaces.
Definition 1.7.3. The mapping cs:V→Rkdefined by
cs(v)=(a1,a2,... ,a k)
where v∼(a1,a2,... ,a k)i st h es o - c a l l e d coordinate map.
Example 1.7.1. We have shown that in R3the solutions to the equation
x1+x2+x3=0f o r x=(x1,x2,x3)∈R3is a subspace V0with basis
S={v1,v2}={(−1,0,1),(−1,1,0)}. With respect to this basis v0∈V0if
there are constants α0,β0so that
v0=α0(−1,0,1) +β0(−1,1,0)
With respect to this basis the coordinate map has the form
cs(v0)=(α0,β0)
Therefore, we have established that V0is isomorphic to R2.
32 CHAPTER 1. VECTORS AND VECTOR SPACES
1.8 Exercises.
1. Show that {(a, b,0)|a, b∈R}is a subspace of R3by proving that it
is spanned by vectors in R3. Find at least two sets of spanning sets.
2. Show that {(a, b,1)|a, b∈R}cannot be a subspace of R3.
3. Show that {(a−b,2b−a, a−b)|a, b∈R}is a subspace of R3by
proving that it is spanned by vectors in R3.
4. For any w∈R3, show that Uw={x∈R3|x·w=dW=0}isnot
subpace of R3.
5. Find a set of vectors in R3that spans the subspace Uw={x∈R3|x·
w=0},w h e r e w=( 1,1,1).
6. Why can {(a−b, a2,a b)|a, b∈R}never be a subspace of R3?
7. Let Q={x1,...,x k}be a set of distinct points on the real line
with k<n . Show that the subset PQof the polynomial space Pnof
polynomials zero on the set Qis in fact a subspace of Pn. Characterize
PQifk>n andk=n.
8. In the product space de fined above prove that de finitions given the
result is a vector space.
9. What is the product space R2×R3?
10. Find a basis for Q={ax+bx3|a, b∈R}.
11. Let T⊂Pnbe those polynomials of exactly degree n. Show that Tis
not a subspace of Pn.
12. What is the dimension of
Q={ax+ax2+bx3|a, b∈Q}.
What is a basis for Q?
13. Given that S={x1,x2, ..., x 2k}andT={y1,y2, ..., y 2k}are
both bases of a vector space V.(Note, the space Vhas dimension 2 k.)
Consider the set of any kintegers L={l1,..., l k}⊂{1,2,..., 2k}.(i)
Show that associated with P={xl1,xl2, ..., x lk}there are exactly
kvectors from T,s a yQ={ym1,ym2,. . . ,y mk}so that P∪Qis also
ab a s i sf o r V.(ii) Is the set of vectors from Tunique? Why or why
not?
1.8. EXERCISES. 33
14. Given Pn.D efineZn={pI(x)|p(x)∈Pn}.( T h e n o t a t i o n pI(x)i st h e
standard notation for the derivative of the function p(x)w i t hr e s p e c t
to the variable x.) What is another way to express Znin terms of
previously de fined spaces?
15. Show that ,x,∞is a norm. (Hint. The condition (iii) should be the
focal point of your e ffort.)
16. Let S={x1,x2,. . . , x n}⊂Rn.F o r e a c h j=1,2,...,n suppose
xj∈Shas the property that its firstj−1 entries equal zero and the
jthentry is nonzero. Show that Sis a basis of Rn.
17. Let w=(w1,w2,w3)∈R3,where all the components of ware strictly
positive. De fine,·,wonR3by,x,w=p
w1|x1|2+w2|x3|2+w2|x3|2Q1/2
.
Show that ,·,wis a norm on R3.
18. Show that equivalence of norms is an equivalence relation.19. De fine for p∈P
n(a, b) the function N(p)=m a x a≤x≤b|pI(x)|.Show
that this is not a norm on Pn(a, b).
20. For p∈Pn(a, b)d efine,p,=m a x a≤x≤b|p(x)|+m a x a≤x≤b|pII(x)|.
Show this is a norm on Pn(a, b).
21. Suppose that Vis a vector space with dimension k. Find two (linearly
independent) spanning sets S={v1,v2,...,v k}andW={w1,w2,...,w k}
ofVsuch that if any m<k vectors are chosen from Sand any k−m
vectors are chosen from T,the resulting set will be a basis for V.
22. For p∈Pn(a, b)d e fineN(p)=eepDa+b
2iee.Show this is not a norm
onPn(a, b).
23. Find the estimates for the equivalence of ,·,1and,·,∞.
24. Find the estimates for the equivalence of ,·,1and,·,2.
25. Show that Pndefined over the reals is isomorphic to Rn+1.
26. Show that Tn, the space of trigonometric polynomials, de fined over
the reals is isomorphic to Rn.
27. Show that the product space Ck×Cmis isomorphic to Ck+m.
28. What is the relation between the product space Pn×PnandP2n?
Find the polynomial space that is isomorphic to Pn×Pn.
34 CHAPTER 1. VECTORS AND VECTOR SPACES
Terms.
Field
Vector space
scalar multiplication
Closed spaceOriginPolynomial spaceSubspace, linear subspaceSpan
Spanning set
RepresentationUniqueness of representationlinear dependencelinear independencelinear combination
Basis
Extension to a basisDimensionNormf
2norm;f1norm;f2∞norm
Cauchy-Schwartz inequality
Fundamental Theorem of AlgebraCardinalityTriangle inquality
Chapter 2
Matrices and Linear Algebra
2.1 Basics
Definition 2.1.1. Amatrix is an m×narray of scalars from a given field
F. The individual values in the matrix are called entries .
Examples.
A=^
21 3
−124
B=^
12
34
Thesizeof the array is–written as m×n,w h e r e
m×n
cA
number of rows number of columns
Notation
A=
a
11a12... a 1n
a21a22... a 2n
an1an2... a mn
A
←− rows
t
AAc
columns
A:= uppercase denotes a matrix
a:= lower case denotes an entry of a matrix a∈F.
Special matrices
33
34 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
(1) If m=n, the matrix is called square .I nt h i sc a s ew eh a v e
(1a) A matrix Ais said to be diagonal if
aij=0 iW=j.
(1b) A diagonal matrix Amay be denoted by diag( d1,d2,... ,d n)
where
aii=diaij=0 jW=i.
The diagonal matrix diag(1 ,1,... , 1) is called the identity matrix
and is usually denoted by
In=
10 ... 0
01
......
01
or simply I,w h e n nis assumed to be known. 0 = diag(0 ,... , 0)
is called the zero matrix .
(1c) A square matrix Lis said to be lower triangular if
fij=0 i<j .
(1d) A square matrix Uis said to be upper triangular if
uij=0 i>j .
(1e) A square matrix Ais called symmetric if
aij=aji.
(1f) A square matrix Ais called Hermitian if
aij=¯aji(¯z:= complex conjugate of z).
(1g) Eijhas a 1 in the ( i, j) position and zeros in all other positions.
(2) A rectangular matrix Ais called nonnegative if
aij≥0a l l i, j.
It is called positive if
aij>0a l l i, j.
Each of these matrices has some speci al properties, which we will study
during this course.
2.1. BASICS 35
Definition 2.1.2. The set of all m×nmatrices is denoted by Mm,n(F),
where Fis the underlying field (usually RorC). In the case where m=n
we write Mn(F) to denote the matrices of size n×n.
Theorem 2.1.1. Mm,nis a vector space with basis given by Eij,1≤i≤
m,1≤j≤n.
Equality, Additi on, Multiplication
Definition 2.1.3. Two matrices AandBare equal if and only if they have
t h es a m es i z ea n d
aij=bijalli, j.
Definition 2.1.4. IfAis any matrix and α∈Fthen the scalar multipli-
cation B=αAis defined by
bij=αaijalli, j.
Definition 2.1.5. IfAandBare matrices of the same size then the sum
AandBis defined by C=A+B,w h e r e
cij=aij+bijalli, j
We can also compute the difference D=A−Bby summing Aand (−1)B
D=A−B=A+(−1)B.
matrix subtraction.
Matrix addition “inherits” many properties from the fieldF.
Theorem 2.1.2. IfA, B, C∈Mm,n(F)andα,β∈F,t h e n
(1)A+B=B+A commutivity
(2)A+(B+C)=(A+B)+C associativity
(3)α(A+B)=αA+αB distributivity of a scalar
(4) If B=0 (a matrix of all zeros) then
A+B=A+0= A
(4)(α+β)A=αA+βA
36 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
(5)α(βA)=αβA
(6)0A=0
(7)α0=0 .
Definition 2.1.6. Ifxandy∈Rn,
x=(x1...x n)
y=(y1...y n).
Then the scalar or dot product of xandyis given by
x, yX=n3
i=1xiyi.
Remark 2.1.1. (i) Alternate notation for the scalar product: x, yX=x·y.
(ii) The dot product is de fined only for vectors of the same length.
Example 2.1.1. Letx=( 1,0,3,−1) and y=( 0,2,−1,2) thenx, yX=
1(0) + 0(2) + 3( −1)−1(2) =−5.
Definition 2.1.7. IfAism×nandBisn×p.L e t ri(A) denote the vector
with entries given by the ithrow of A,a n dl e t cj(B) denote the vector with
entries given by the jthrow of B. The product C=ABis the m×pmatrix
defined by
cij=ri(A),cj(B)X
where ri(A) is the vector in Rnconsisting of the ithrow of Aand similarly
cj(B) is the vector formed from the jthcolumn of B. Other notation for
C=AB
cij=n
k=1aikbkj1≤i≤m
1≤j≤p.
Example 2.1.2. Let
A=}101
321]
and B=
21
30
−11
.
Then
AB=}12
11 4]
.
2.1. BASICS 37
Properties of matrix multiplication
(1) If ABexists, does it happen that BAexists and AB=BA?T h e
answer is usually no. First AB andBA exist if and only if A∈
Mm,n(F)a n d B∈Mn,m(F). Even if this is so the sizes of ABand
BAare different ( ABism×mandBAisn×n) unless m=n.
However even if m=nwe may have ABW=BA.S e e t h e e x a m p l e s
below. They may be di fferent sizes and if they are the same size (i.e.
AandBa r es q u a r e )t h ee n t r i e sm a yb ed i fferent
A=[ 1,2]B=}−1
1]
AB=[ 1 ]
BA=}−1−2
12]
A=}12
34]
B=}−11
01]
AB=}−13
−37]
BA=}22
34]
(2) If Ais square we de fine
A1=A, A2=AA, A3=A2A=AAA
An=An−1A=A···A(nfactors) .
(3)I= diag(1 ,... , 1). If A∈Mm,n(F)t h e n
AIn=Aand
ImA=A.
Theorem 2.1.3 (Matrix Multiplication Rules). Assume A, B ,a n d C
are matrices for which all products below make sense. Then
(1)A(BC)=(AB)C
(2)A(B±C)=AB±ACand(A±B)C=AC±BC
(3)AI=AandIA=A
(4)c(AB)=(cA)B
(5)A0=0 and0B=0
38 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
(6) For Asquare
ArAs=AsArfor all integers r, s≥1.
Fact: IfACandBCare equal, it does not follow that A=B. See Exercise
60.
Remark 2.1.2. We use an alternate notation for matrix entries. For any
matrix Bdenote the ( i, j)-entry by ( B)ij.
Definition 2.1.8. LetA∈Mm,n(F).
(i) De fine the transpose ofA, denoted by AT,t ob et h e n×mmatrix
with entries
(AT)ij=aji.
(ii) De fine the adjoint ofA, denoted by A∗,t ob et h e n×mmatrix with
entries
(A∗)ij=¯ajicomplex conjugate
Example 2.1.3.
A=}123
541]
AT=
15
24
31
In words ...“The rows of Abecome the columns of AT, taken in the same
order.” The following results are easy to prove.
Theorem 2.1.4 (Laws of transposes). (1)(AT)T=Aand(A∗)∗=A
(2)(A±B)T=AT±BT(and for ∗)
(3)(cA)T=cAT(cA)∗=¯cA∗
(4)(AB)T=BTAT
(5) If Ais symmetric
A=AT
2.1. BASICS 39
(6) If Ais Hermitian
A=A∗.
More facts about symmetry.
Proof. (1) We know ( AT)ij=aji.S o( ( AT)T)ij=aij.T h u s( AT)T=A.
(2) (A±B)T=aji±bji.S o( A±B)T=AT±BT.
Proposition 2.1.1. (1)Ais symmetric if and only if ATis symmetric.
(1)∗Ais Hermitian if and only if A∗is Hermitian.
(2) If Ais symmetric, then A2is also symmetric.
(3) If Ais symmetric, then Anis also symmetric for all n.
Definition 2.1.9. A matrix is called skew-symmetric if
AT=−A.
Example 2.1.4. The matrix
A=
012
−10−3
−23 0
is skew-symmetric.
Theorem 2.1.5. (1) If Ais skew symmetric, then Ais a square matrix
andaii=0,i=1,... ,n .
(2) For any matrix A∈Mn(F)
A−AT
is skew-symmetric while A+ATis symmetric.
(3) Every matrix A∈Mn(F)can be uniquely written as the sum of a
skew-symmetric and symmetric matrix.
Proof. (1) If A∈Mm,n(F), then AT∈Mn,m(F). So, if AT=−Awe
must have m=n.A l s o
aii=−aii
fori=1,... ,n .S oaii=0f o ra l l i.
40 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
(2) Since ( A−AT)T=AT−A=−(A−AT), it follows that A−ATis
skew-symmetric.
(3) Let A=B+Cbe a second such decomposition. Subtraction gives
1
2(A+AT)−B=C−1
2(A−AT).
The left matrix is symmetric while the right matrix is skew-symmetric.
Hence both are the zero matrix.
A=1
2(A+AT)+1
2(A−AT).
Examples. A=J0−1
10o
is skew-symmetric. Let
B=}12
−14]
BT=}1−1
24]
B−BT=}03
−30]
B+BT=}21
18]
.
Then
B=1
2(B−BT)+1
2(B+BT).
An important observation about matri x multiplication is related to ideas
from vector spaces. Indeed, two very important vector spaces are associatedwith matrices.
Definition 2.1.10. LetA∈M
m,n(C).
(i)Denote by
cj(A): =jthcolumn of A
cj(A)∈Cm. We call the subspace of Cmspanned by the columns of Athe
column space ofA.W i t h c1(A),...,c n(A) denoting the columns of A
2.1. BASICS 41
the column space is S(c1(A),...,c n(A)).
(ii) Similarly, we call the subspace of Cnspanned by the rows of Atherow
space ofA.W i t h r1(A),...,r m(A) denoting the rows of Athe row space
is therefore S(r1(A),...,r m(A)).
Letx∈Cn,w h i c hw ev i e wa st h e n×1m a t r i x x=[x1...x n]T.T h e
product Axis defined and
Ax=n3
j=1xjcj(A).
That is to say, Ax∈S(c1(A),... ,c n(A)) = column space of A.
Definition 2.1.11. LetA∈Mn(F). The matrix Ais said to be invertible
if there is a matrix B∈Mn(F) such that
AB=BA=I.
In this case Bis called the inverse ofA, and the notation for the inverse is
A−1.
Examples.
(i) Let
A=}13
−12]
Then
A−1=1
5}2−3
11]
.
(ii) For n=3w eh a v e
A=
12−1
−13−1
−23−1
A−1=
01−1
−13−2
−37−5
A square matrix need not have an inverse, as will be discussed in the
next section. As examples, the two matrices below do not have inverses
A=}1−2
−12]
B=
101
021122
42 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
2.2 Linear Systems
The solutions of linear systems is likely the single largest application of ma-
trix theory. Indeed, most reasonable problems of the sciences and economicsthat have the need to solve problems of several variable almost without ex-ception are reduced to component parts where one of them is the solutionof a linear system. Of course the entire solution process may have the linear
system solver as a relatively small component, but an essential one. Even
the solution of nonlinear problems, esp ecially, employ linear systems to great
and crucial advantage.
To be precise, we suppose that the coe fficients a
ij,1≤i≤mand 1≤
j≤nand the data bj,1≤j≤ma r ek n o w n . W ed e fine the linear system
for the nunknowns x1,...,x nto be
a11x1+a12x2+···+a1nxn=b1
a21x1+a22x2+···+a2nxn=b2 (∗)
am1x1+am2x2+···+amnxn=bm
The solution set is defined to be the subset of Rnof vectors ( x1,...,x n)t h a t
satisfy each of the mequations of the system. The question of how to solve
a linear system includes a vast literature of theoretical and computation
methods. Certain systems form the model of what to do. In the systemsbelow we note that the first one has three highly coupled (interrelated)
variables.
3x
1−2x2+4x3=7
x1−6x2−2x3=0
−x1+3x2+6x3=−2
The second system is more tractable because there appears even to the
untrained eye a clear and direct method of solution.
3x1−2x2−x3=7
x2−2x3=1
2x3=−2
I n d e e d ,w ec a ns e er i g h to ffthatx3=−1.Substituting this value into the
second equation we obtain x2=1−2=−1.Substituting both x2andx3
into the first equation, we obtain 2 x1−2(−1)−(−1) = 7 ,gives x1=2.The
2.2. LINEAR SYSTEMS 43
solution set is the vector (2 ,−1,−1).The virtue of the second system is
that the unknowns can be determined on e-by-one, back substituting those
already found into the next equation until all unknowns are determined. Soif we can convert the given system of the first kind to one of the second kind,
we can determine the solution.
This procedure for solving linear systems is therefore the applications of
operations to e ffect the gradual elimination of unknowns from the equations
until a new system results that can be solved by direct means. The oper-ations allowed in this process must have precisely one important property:They must not change the solution set by either adding to it or subtracting
from it. There are exactly three such operations needed to reduce any set
of linear equations so that it can be solved directly.
(E1) Interchange two equations.(E2) Multiply any equation by a nonzero constant.
(E3) Add a multiple of one equation to another.
This can be summarized in the following theorem
Theorem 2.2.1. Given the linear system (*). The set of equation opera-
tions E1, E2, and E3 on the equations of (*) does not alter the solution setof the system (*).
We leave this result to the exercises. Our main intent is to convert these
operations into corresponding operations for matrices. Before we do this
we clarify which linear systems can have a soltution. First, the system can
be converted to matrix form by setting Aequal to the m×nmatrix of
coefficients, bequal to the m×1 vector of data, and xequal to the n×1
vector of unknowns. Then the system (*) can be written as
Ax=b
In this way we see that with c
i(A)d e n o t i n gt h e ithcolumn of A,the system
is expressible as
x1c1(A)+···+xncn(A)=b
From this equation it is clear that the system has a solution if and only if
the vector bis inS(c1(A),···,cn(A)). This is summarized in the following
theorem.
44 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
Theorem 2.2.2. An e c e s s a r ya n ds u fficient condition that Ax=bhas a
solution is that b∈S(c1(A)...c n(A)).
In the general matrix product C=AB, we note that the column space of
C⊂column space of A.I nt h ef o l l o w i n gd e finition we regard the matrix A
as a function acting upon vectors in one vector space with range in anothervector space. This is entirely similar to the domain-range idea of function
theory.
Definition 2.2.1. Therange ofA={Ax|x∈R
n(o rCn)}.
It follows directly from our discussion above that the range of Aequals
S(c1(A),... ,c n(A)).
Row operations: To solve Ax=bwe use a process called Gaussian
elimination , which is based on row operations.
Type 1: Interchange two rows. (Notation: Ri←→Rj)
Type 2: Multiply a row by a nonzero constant. (Notation: cRi→Ri)
Type 3: Add a multiple of one row to another row. (Notation: cRi+Rj→
Rj)
Gaussian elimination is the process of reducing a matrix to its RREF using
these row operations. Each of these operations is the respective analogue of
the equation operations described above, and each can be realized by leftmatrix multiplication. We have the following.Type 1
E
1=
1......
1......
... ... 0... ... ... 1... ... ...
...1...
.........
...1...
... ... 1... ... ... 0... ... ...
......1
.........
......1
rowi
row j
column
icolumn
j
2.2. LINEAR SYSTEMS 45
Notation: Ri↔Rj
Type 2
E2=
1...
......
1...
... ... ... c ... ... ...
...1
......
...1
rowi
column i
Notation: cR
i
Type 3
E3=
1...
1...
...
......
... ... c ... ... ... ...
...1
...1
rowj
column
i
Notation: cR
i+Rj, the abbreviated form of cRi+Rj→Rj
Example 2.2.1. The operations
21 0
02 1
−102
R1←→R2
→
02 1
21 0
−102
4R3
→
02 1
21 0
−408
46 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
can also be realized as
R1←→ R2:
010
100001
21 0
02 1
−102
=
02 1
21 0
−102
4R
3 :
100
010004
02 1
21 0
−102
=
02 1
21 0
−408
The operations
21 0
02 1
−102
−3R
1+R2
→
2R1+R3
21 0
−6−11
32 2
can be realized by the left matrix multiplications
100
010
201
10 0
−310
00 1
21 0
02 1
−102
=
21 0
−6−11
32 2
Note there are two matrix multiplications them, one for each Type 3 ele-
mentary operation.
Row-reduced echelon form. To each A∈Mm,n(E) there is a canonical
form also in Mm,n(E) which may be obtained by row operations. Called the
RREF, it has the following properties.
(a) Each nonzero row has a 1 as the first nonzero entry (:= leading one ).
(b) All column entries above and below a leading one are zero.
(c) All zero rows are at the bottom.
(d) The leading one of one row is to the left of leading ones of all lower
rows.
Example 2.2.2.
B=
1200 −1
0010 30001 00000 0
is in RREF.
2.2. LINEAR SYSTEMS 47
Theorem 2.2.3. LetA∈Mm,n(F). Then the RREF is necessarily unique.
We defer the proof of this result. Let A∈Mm,n(F). Recall that the
row space ofAis the subspace of Rn(orCn) spanned by the rows of A.I n
symbols the row space is
S(r1(A),... ,r m(A)).
Proposition 2.2.1. ForA∈Mm,n(F)the rows of its RREF span the rows
space of A.
Proof. First, we know the nonzero rows of the RREF are linearly indepen-
dent. And all row operations are linear combinations of the rows. Thereforet h er o ws p a c eg e n e r a t e df r o mt h eR R E Fi sc o n t a i n e di nt h er o ws p a c eo fA. If the containment is proper. That is there is a row of Athat is lin-
early independent from the row space of the RREF, this is a contradiction
because every row of Acan be obtained by the inverse row operations from
the RREF.
Proposition 2.2.2. IfA∈Mm,n(F)and a row operation is applied to A,
then linearly dependent columns of Aremain linearly dependent and linearly
independent columns of Aremain linearly independent.
Proposition 2.2.3. The number of linearly independent columns of A∈
Mm,n(F)is the same as the number of leading ones in the RREF of A.
Proof. LetS={i1...i k}be the columns of the RREF of Ahaving a lead-
ing one. These columns of the RREF are linearly independent Thus these
columns were originally linearly independent. If another column is linearlyindependent, this column of the RREF is linearly dependent on the columnswith a leading one. This is a contradiction to the above proposition.
Proof of Theorem 2.2.3. By the way the RREF is constructed, left-to-right,
and top-to-bottom, it should be apparent that if the right most row of the
RREF is removed, there results the RREF of the m×(n−1) matrix formed
from Aby deleting the nthcolumn. Similarly, if the bottom row of the
RREF is removed there results a new matrix in RREF form, though notsimply related to the original matrix A.
To prove that the RREF is unique, we proceed by a double induction,
first on the number of columns. We take it as given that for an m×1m a t r i x
the RREF is unique. It is either the zero m×1 matrix, which would be
t h ec a s ei f Awas zero or the matrix with a 1 in the first row and zeros in
48 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
the other rows. Assume therefore that the RREF is unique if the number
o fc o l u m n si sl e s st h a n n. Assume there are two RREF forms, B1and
B2forA.N o w t h e R R E F o f Ais therefore unique through the ( n−1)st
columns. The only di fference between the RREF’s B1andB2must occur
in the nthcolumn. Now proceed by induction on the number of nonzero
rows. Assume that AW=0 . I f Ahas just one row, the RREF of Ais simply
thescalar multiple of Athat makes the first nonzero column entry a one.
Thus it is unique. If A= 0, the RREF is also zero. Assume now that
the RREF is unique for matrices with less than mrows. By the comments
above that the only di fference between the RREF’s B1andB2can occur at
the (m, n)-entry. That is ( B1)m,nW=(B2)m,n. They are therefore not leading
ones. (Why?) There is a leading one in the mthrow, however, because it
is a non zero row. Because the row spaces of B1andB2are identical, this
results in a contradiction, and therefore the ( m, n)-entries must be equal.
Finally, B1=B2.This completes the induction. (Alternatively, the two
systems pertaining to the RREF’s must have the same solution set to thesystem Ax=0 . W i t h( B
1)m,nW=(B2)m,n, it is easy to see that the solution
sets to B1x=0a n d B2x=0m u s td i ffer.) ¤
Definition 2.2.2. LetA∈Mm,nandb∈Rm(orCn). De fine
[A|b]=
a11... a 1nb1
a21... a 2nb2
am1... a mnbm
[A|b]i sc a l l e dt h e augmented matrix ofAbyb.[A|b]∈Mm,n+1(F). The
augmented matrix is a useful notation for finding the solution of systems
using row operations.
Identical to other de finitions for solutions of equations, the equivalence
of two systems is de fined via the idea of equality of the solution set.
Definition 2.2.3. Two linear systems Ax=bandBx=care called equiv-
alent if one can be converted to the other by elementary equation opera-
tions.
It is easy to see that this implies the followingTheorem 2.2.4. Two linear systems Ax=bandBx=care equivalent if
and only if both [A|b]and[B|c]have the same row reduced echelon form.
We leave the prove to the reader. (See Exercise 23.) Note that the solution
set need not be a single vector; it can be null or in finite.
2.3. RANK 49
2.3 Rank
Definition 2.3.1. Therank of any matrix A,d e n o t eb y r(A), is the di-
mension of its column space.
Proposition 2.3.1. (i) The rank of Aequals the number of nonzero rows
of the RREF of A, i.e. the number of leading ones.
(ii)r(A)=r(AT).
Proof. (i) Follows from previous results.
(ii) The number of linearly independent rows equals the number of lin-
early independent columns. The number of linearly independent rows is
the number of linearly independent columns of AT–by de finition. Hence
r(A)=r(AT).
Proposition 2.3.2. LetA∈Mm,n(C)andb∈Cm.T h e n Ax=bhas a
solution if and only if r(A)=r([A|b]),w h e r e [A|b]is the augmented matrix.
Remark 2.3.1. Solutions may exist and may not. However, even if a so-
lution exists, it may not be unique. Indeed if it is not unique, there is an
infinity of solutions.
Definition 2.3.2. When Ax=bhas a solution we say the system is con-
sistent .
Naturally, in practical applications we want our systems to be consistent.
When they are not, this can be an indicator that something is wrong withthe underlying physical model. In mathematics, we also want consistentsystems; they are usually far more interesting and o ffer richer environments
for study.
In addition to the column and row spaces, another space of great impor-
tance is the so-called null space, the set of vectors x∈R
nfor which Ax=0 .
In contrast, when solving the simple single variable linear equation ax=b
with aW= 0 we know there is always a unique solution x=b/a.I n s o l v i n g
even the simplest higher dimensional systems, the picture is not as clear.
Definition 2.3.3. LetA∈Mm,n(F). The null space ofAis defined to be
Null( A)={x∈Rn|Ax=0}.
It is a simple consequence of the linearity of matrix multiplication that
Null( A) is a linear subspace of Rn. That is to say, Null( A) is closed under
vector addition and scalar multiplication. In fact, A(x+y)=Ax+Ay=
0+0=0 , i f x, y∈Null( A). Also, A(αx)=αAx=0 ,i f x∈Null( A). We
state this formally as
50 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
Theorem 2.3.1. LetA∈Mm,n(F).T h e n N u l l (A)is a subspace of Rn
w h i l et h er a n g eo f Ais in Rm.
Having such solutions gives valuable information about the solution set
of the linear system Ax=b. For, if we have found asolution, x, and have
any vector z∈Null( A), then x+zis a solution of the same linear system.
Indeed, what is easy to see is that if uandvare both solutions to Ax=b,
then A(u−v)=Au−Av= 0, or what is the same x−y∈Null( A). This
means that to findallsolutions to Ax=b, we need only find a single solution
and the null space. We summarize this as the following theorem.
Theorem 2.3.2. LetA∈Mm,n(F)with null space Null (A).L e t xbe any
nonzero solution to Ax=b. Then the set x+Null(A)is the entire solution
set to Ax=b.
Example 2.3.1. Find the null space of A=}13
−3−9]
.
Solution. Solve Ax=0.The RREF for Ais}13
00]
.S o l v i n g x1+3x2=0,
take x2=t, a “free” parameter and solve for x1to get x1=−3t.Thus
every solution to Ax= 0 can be written in the form
x=}−3t
t]
=t}−3
1]
t∈R
Expressed this way we see that Null( A)=F
t}−3
1]
|t∈Rk
,a subspace
ofR2of dimension 1.
Theorem 2.3.3 (Fundamental theorem on rank). A∈Mm,n(F).T h e
following are equivalent
(a)r(A)=k.
(b) There exist exactly klinearly independent columns of A.
(c) There exist exactly klinearly independent rows of A.
(d) The dimension of the column space of Aisk(i.e. dim( Range A)=k).
(e) There exists a set Sof exactly kvectors in Rmfor which Ax=bhas
a solution for each b∈S(S).
(f) The null space of Ahas dimension n−k.
2.3. RANK 51
Proof. The equivalence of (a), (b), (c) and (d) follow from previous con-
siderations. To establish (e), let S={cf1,cf2,... ,c fk}denote the linearly
independent column vectors of A.L e t T={ef1,ef2,... ,e fk}⊂Rnbe the
standard vectors. Then Aefj=cfj.I fb∈S(S), then b=a1cf1+a2cf2+
···+akcfk.As o l u t i o nt o Ax=bis given by x=a1ef1+a2ef2+···+akefk.
Conversely, if (e) holds, then the set Smust be linearly independent for
otherwise Scould be reduced to k−1 or fewer vectors. Similarly if Ahas
k+ 1 linearly independent columns then set Scan be expanded. Therefore,
the column space of Amust have exactly kvectors.
To prove (f) we assume that S={v1,... ,v k}is a basis for the column
space of A.L e t T={w1,... ,w k}⊂Rnfor which Awi=vi,i=1,... ,k .
By our extension theorem, we select n−kvectors wk+1,... ,w nsuch that
U={w1,... ,w k,wk+1,... ,w n}is a basis of Rn. We must have that
Awk+1∈S(S). Hence there are scalars b1,... ,b ksuch that
Awk+1=A(b1w1+···+bkwk)
and thus wI
k+1=wk+1−(b1w1+···+bkwk)i si nt h en u l ls p a c eo f A.
Repeat this process for each wk+j,j=1,... ,n−k. We generate a total
ofn−kvectors {wI
k+1,... ,wI
n}in this manner. This set must be linearly
independent. (Why?) Therefore, the dimension of the null space must beat least n−k. Now we consider a new basis which consists of the original
vectors and the n−kvectors {w
I
k+1,wI
k+2,... ,wI
n}for which Aw=0 . W e
assert that the dimension of the null space is exactly n−k.F o ri f z∈Rnis
av e c t o rf o rw h i c h Az=0 ,t h e n zcan be uniquely written as a component
z1fromS(T) and a component z2fromS({wI
k+1,... ,wI
n}). But Az1W=0
andAz2= 0. Therefore Az= 0 is impossible unless the component z1=0 .
Conversely, if (f) holds we take a basis for the null space T={u1,u2,... ,u n−k}
a n de x t e n dt h eb a s i s
TI=T∪{un−k+1,. . . ,u n}
toRn.N e x ta r g u es i m i l a r l yt oa b o v et h a t
Aun−k+1,A u n−k+2,... ,A u n
must be linearly independent, for otherwise there is yet another linearly
independent vector that can be added to its basis, a contradiction. Thereforethe column space must have dimension at least, and hence equal to k.
The following corollary assembles many consequences of this theorem.
52 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
Corollary 2.3.1. (1)r(A)≤min(m, n).
(2)r(AB)≤min(r(A),r(B)).
(3)r(A+B)≤r(A)+r(B).
(4)r(A)=r(AT)=r(A∗)=r(¯A).
(5) If A∈Mm(F)andB∈Mm,n(F),a n di f Ais invertible, then
r(AB)=r(B).
Similarly, if C∈Mn(F)is invertible and B∈Mm,n(F)
r(BC)=r(B).
(6)r(A)=r(ATA)=r(A∗A).
(7) Let A∈Mm,n(F),w i t h r(A)=k.T h e n A=XBY where X∈Mm,k,
Y∈Mk,nandB∈Mkis invertible.
(8) In particular, every rank 1 matrix has the form A=xyT,w h e r e x∈
Rmandy∈Rn.H e r e
xyT=
x1y1x1y2... x 1yn
.........
xmy1xmy2... x myn
.
Proof. (1) The rank of any matrix is the number of linearly independent
rows, which is the same as the number of linearly independent columns.
The maximum this value can be is therefore the maximum of theminimum of the dimensions of the matrix, or r(A)≤min ( m, n).
(2) The product ABc a nb ev i e w e di nt w ow a y s . T h e fir s ti sa sas e t
of linear combinations of the rows of B,and the other is as a set of
linear combinations of the columns of A.In either case the number
of linear independent rows (or columns as the case may be) In otherwords, the rank of the product ABcannot be greater than the number
of linearly independent columns of Anor greater than the number of
linearly independent rows of B.Another way to express this is as
r(AB)≤min(r(A),r(B))
2.3. RANK 53
(3) Now let S={v1,...v r(A)}andT={w1,...,w r(B)}be basis of the
column spaces of AandBrespectively. Then, the dimension of the
union S∪T={v1,...v r(A),w1,...,w r(B)}cannot exceed r(A)+r(B).
Also, every vector in the column space of A+Bis clearly in the span
ofS∪T.The result follows.
(4) The rank of Ais the number of linearly independent rows (and columns)
ofA,which in turn is the number of linearly independent columns of
AT,which in turn is the rank of AT.That is, r(A)=rD
ATi
.Similar
proofs hold for A∗and ¯A.
(5) Now suppose that A∈Mm(F) is invertible and B∈Mm,n.As
we have emphasized many times the rows of the product ABcan be
viewed as a set of linear combinations of the rows of B.Since Ahas
rank m any set of linearly independent rows of Bremains linearly
independent. To see why, let ri(AB)d e n o t et h e ithrow of the product
AB. Then it is easy to see that
ri(AB)=m3
j=1aijrj(B)
Suppose we can determine constants c1,...,c mnot all zero so that
0=m3
j=1ciri(AB)=m3
i=1cim3
j=1aijrj(B)
=m3
j=1rj(B)m3
i=1ciaij
This linear combination of the rows of Bhas coefficient given by ATc,
where c=[c1,...,c k]T.Because the rank of A(and AT)i sm,we
can solve this system for any vector d∈Rm.Suppose that the row
vectors rjl(B),f=1,...,r (B),are linearly independent. Arrange
that the components of dto be zero for indices not included in the set
jl,f=1,...,r (B) and not all zero otherwise. Then the conclusion
0=m
j=1rj(B)m
i=1ciaij=r(B)
l=1rjl(B)djlis impossible. Indeed,
the same basis of the row space of Bwill be a basis of the row space
ofAB.T h i sp r o v e st h er e s u l t .
(6) We postpone the proof of this result until we discuss orthogonality.
54 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
(7) Place Ain RREF, say ARREF .S i n c e r(A)=kwe know the top k
rows of ARREF are linearly independent and the remaining rows are
zero. De fineYto be the k×nmatrix consisting of these top krows.
DefineB=Ik.Now the rows of Aare linear combinations of these
rows. So, de fine the m×kmatrix Xto have rows as follows: The
first row of consists of the coe fficients so thatx1jrj(Y)=r1(A).
In general, the ithrow of Xis selected so that
3
xijrj(Y)=ri(A)
(8) This is an application of (7) noting in this special case that Xis an
m×1 matrix that can be interpretted as a vector x∈Rm. Similarly,
Yis an 1 ×nmatrix that can be interpretted as a vector y∈Rn.
Thus, with I=[ 1 ] ,w eh a v e
A=xyT
Example 2.3.2. Here is the decomposition of the form given in Lemma
2.3.1 (7). The 3 ×4m a t r i x Ahas rank 2.
A=
12 −1
002
−1−23
240
=
1−1
02
−13
20
}10
01]}120
001]
=XBY
The matrix Yis the RREF of A.
Example 2.3.3. Letx=[x1,x2,...,x m]T∈Rmandy=[y1,y2,...y n]T∈
Rn. Then the rank one m×nmatrix xyThas the form
xyT=
x
1y1x1y2··· x1yn
x2y1x2y2 x2yn
.........
xmy1xmy2···xmyn
In particular, with x=[ 1,3,5]T,a n d y=[−2,7]T,the rank one 3 ×2m a t r i x
xyTis given by
xyT=
1
35
[−2,7] =
−27
−62 1
−10 35
2.3. RANK 55
Invertible Matrices
A subclass matrices A∈Mn(F) that have only the zero kernel is very
important in applications and theoretical developments.
Definition 2.3.4. A∈Mnis called nonsingular ifAx= 0 implies that
x=0 .
In many texts such matrices are introduced though an equivalent alter-
nate de finition involving rank.
Definition 2.3.5. A∈Mnisnonsingular ifr(A)=n.
We also say that nonsingular matrices have fullrank. That nonsingular
matrices are invertible and conversely together with many other equivalencesis the content of the next theorem.
Theorem 2.3.4. [Fundamental theorem on inverses] Let A∈M
n(F).T h e n
the following statements are equivalent.
(a)Ais nonsingular.
(b)Ais invertible.
(c)r(A)=n.
(d) The rows and columns of Aare linearly independent.
(e)dim(Range( A)) =n.
(f)dim(Null( A)) = 0 .
(g)Ax=bis consistent for all b∈Rn(orCn).
(h)Ax=bhas a unique solution for every x∈Rn(orCn).
(i)Ax=0 has only the zero solution.
(j)* 0 is not an eigenvalue of A.
(k)* detAW=0.
* The statements about eigenvalues and the determinant (det A)o fam a -
trix will be clari fied later after they have been properly de fined. They are
included now for completeness.
56 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
Definition 2.3.6. Two linear systems Ax=bandBx=care called equiv-
alent if one can be converted to the other by elementary equation opera-
tions. Equivalently, the systems are equivalent if [ A|b] can be converted to
[B|c] by elementary row operations.
Alternatively, the systems are equivalent if they have the same solution
set which means of course that both can be reduced to the same RREF.
Theorem 2.3.5. IfA∈Mn(F)andB∈Mn(F)with AB=I,t h e n Bis
unique.
Proof. IfAB=Ithen for every e1...e nthere is a solution to the system
Abi=eifor all 1 = 1 ,2,... ,n .T h u st h es e t {bi}n
i=1is linearly independent
(because the set {ei}is) and moreover a basis. Similarly if AC=Ithen
A(C−B) = 0, and there are ci∈Rn(orCn),i=1, ..., n such that
Aci=ei. Suppose for example that c1−b1W=0 . S i n c et h e {bi}n
i=1is a basis
it follows that c1−b1=Σαjbj, where not all αjare zero. Therefore,
A(c1−b1)=ΣαjAbj=ΣαjejW=0.
and this is a contradiction.
Theorem 2.3.6. LetA∈Mn(F).I fBis a right inverse, AB=I,t h e n B
is a left inverse.
Proof. DefineC=BA−I+B,a n da s s u m e CW=Bor what is the same
thing that Bis not a left inverse. Then
AC=ABA−A+AB=(AB)(A)−A+AB
=A−A+AB=I
This implies that Cis another right inverse of A, contradicting Theorem
2.3.5.
2.4 Orthogonality
Let V be a vector space over C.W ed e fine an inner product ·,·XonV×V
to be a function from VtoCthat satis fies the following properties:
1.av, wX=av,wXandv,awX=av,wX(ais the complex conjugate
ofa)
2.v,wX=w,vX
2.4. ORTHOGONALITY 57
3.u+v,wX=u, wX+v,wX(linearity)
4.u, v+wX=u, vX+u, wX
5.v,vX≥ 0w i t hv,vX=0i fa n do n l yi f v=0.
For inner products over real vector spaces, we neglect the complex con-
jugate operation. In addition, we want our inner products to de fine anorm
as follows:
6. For any v∈V,,v,2=v,vX
We assume thoughout the text that all vector spaces with inner products
have norms de fined exactly in this way. With the norm and vector vcan
benormalized by dilating it to have length 1, say vn=v1
,v,. The simplest
type of inner product on Cnis given by
v,wX=n3
i=1xi¯yi
We call this the standard inner product.
Using any inner product, we can de fine an angle between vectors.
Definition 2.4.1. Theangleθxybetween vectors xandyinRnis defined
by
cosθxy=x, yX
,x,,y,
=x, yX
(x, xX)1/2(y,yX)1/2.
This comes from the well known result in R2
x·y=,x,,y,cosθ
which can be proved using the law of cosines. With angle comes the notion
of orthogonality.
Definition 2.4.2. Two vectors uandvare said to be orthogonal if the
angle between them ifπ
2or what is the same thing u, vX=0 . I nt h i sc a s e
we commonly write x⊥y. W ee x t e n dt h i sn o t a t i o nt os e t s Uwriting x⊥U
to mean that x⊥ufor every u∈U. Similarly two sets UandVare called
orthogonal if u⊥vfor every u∈Uandv∈V.
58 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
Remark 2.4.1. It is important to note that the notion of orthogonality
depends completely on the inner product. For example, the weighted inner
product defined byv,wX=n
i=1wixi¯yiwhere the wi>0g i v e sv e r yd i fferent
orthogonal vectors from the standard inner product.
Example 2.4.1. InRnorCnthe standard unit vectors are orthogonal with
respect to the standard inner product.
Example 2.4.2. In the R3the vectors u=( 1,2,−1) and v=( 1,1,3) are
orthogonal because
x, yX=1( 1 )+2( 2 ) −1( 3 )=0
Note that in R3the complex conjugate is not written. The set of vectors
(x1,x2,x3)∈R3orthogonal to u=( 1,2,−1) satis fies the equation x1+
2x2−x3= 0 is recognizable as the plane with normal vector u.
Definition 2.4.3. We de fine the projection Puvof one vector vin the di-
rection of an other vector uto be
Puv=u, vX
,u,2u
A sy o uc a ns e e ,w eh a v em e r e l yw r i t t e na ne x p r e s s i o nf o rt h em o r ei n -
tuitive version of the projection in question given by ,v,cosθuvu
,u,.I n t h e
figure below, we show the fundamental diagram for the projection of one
vector in the direction of another.
/c113uv
Pvu
If the vectors uandvare orthogonal, it is easy to see that Puv=0.
(Why?)
Example 2.4.3. Find the projection of the vector v=( 1,2,1) on the vector
u=(−2,1,3)
2.4. ORTHOGONALITY 59
Solution. We have
Puv=u, vX
,u,2u=(1,2,1),(−2,1,3)X
,(−2,1,3),2(−2,1,3)
=1(−2) + 2 (1) + 1 (3)
14(−2,1,3)
=5
14(−2,1,3)
We are now ready to findorthogonal sets of vectors and orthogonal bases.
First we make an important de finition.
Definition 2.4.4. LetVbe a vector space with an inner product. A set
of vectors S={x1,... ,x n}inVis said to be orthogonal ifxi,xjX=0f o r
iW=j.I ti sc a l l e d orthonormal if alsoxi,xiX= 1. If, in addition, Sis a
basis it is called an orthogonal basis or orthonomal basis.
Note: Sometimes the conditions for orthonormality are written as
xi,xjX=δij
whereδijis the “Dirac” delta: δij=0 ,iW=j,δii=1 .
Theorem 2.4.1. Suppose Uis a subspace of the (inner product) vector
space Vand that Uhas the basis S={x1...x k},t h e n Uhas an orthogonal
basis.
Proof. Definey1=x1
,|x,|.T h u s y1is the “normalized” x1.N o w d e fine the
new orthonormal basis recursively by
yI
j+1=xj+1−j3
i=1yi,xj+1Xyi
yj+1=yI
j+1
,yI
j+1,
forj=1,2,...,k−1. Then
(1)yj+1is orthogonal to y1,. . .,y j
(2)yj+1W=0 .
In the language above we have yi,yjX=δij.
60 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
Basically, what the proof accomplishes is to take the di fferences of the
vector from the projections to the others. Referring to the figure above we
compute v−Puva sn o t e di nt h e figure below. The process of orthogonal-
ization described above is called the Gram—Schmidt process.
/c113uv
v
PvPv
uu-
Representation of vectors
One of the great advantages of orthonormal bases is that they make the
representation of vectors particularl y easy. It is as simple as computing an
inner product. Let Vbe a vector space with inner product ·,·Xand with
subspace Uhaving basis S={u1,u2,...,u k}. Then for every u∈Uwe
know there are constants a1,a2,...,a ksuch that
x=a1u1+a2u2+···+akuk.
Taking the inner product of both sides with ujand applying the orthogo-
nality relations
x, u jX=a1u1+a2u2+···+akuk.,ujX
=k3
j=1aiui.,ujX=aj
Thus aj=x, u jX,j=1,2, ..., k ,a n d
x=k3
j=1u.,ujXuj
Example 2.4.4. One basis of R2is given by the orthonormal vectors S=
{u1,u2},w h e r e u1=
1√
2,1√
2=T
and u2=
1√
2,−1√
2=T
. The representa-
tion of x=[ 3,2]Tis given by
x=23
j=1u.,ujXuj=5
2√
2}1√
2,1√
2]T
+1
2√
2}1√
2,−1√
2]T
2.4. ORTHOGONALITY 61
Orthogonal subspaces
Definition 2.4.5. For any set of vectors Swe de fine
S⊥={v∈V|v⊥S}
That is, S⊥is the set of vectors orthogonal to S.O f t e n , S⊥is called the
orthogonal complement ororthocomplement ofS.
For example the orthocomplement of any vector v=[v1,v2,v3]T∈R3is the
(unique) plane passing through the origin that is orthogonal to v.I ti se a s y
to see that the equation of the plane is x1v1+x2v2+x3v3=0 .
For any set of vectors Sthe orthocomplement S⊥has the remarkable
property of being a subspace of V, and therefore it is must have an orthog-
onal basis.
Proposition 2.4.1. Suppose that Vis a vector space with an inner product,
andS⊂V.T h e n S⊥is a subspace of V.
Proof. Ify1,. . .,y m∈S⊥thenΣaiyi∈S⊥for every set of coe fficients
a1,. . .,a minR(orC).
Corollary 2.4.1. Suppose that Vi sav e c t o rs p a c ew i t ha ni n n e rp r o d u c t ,
andS⊂V.
(i) If Sis a basis of V,S⊥={0}.
(ii) If U=S(S),t h e n U⊥=S⊥.
The proofs of these facts are elementary consequences of the proposition.
An important decomposition result is based on orthogonality of subspaces.For example, suppose that Vis afinite dimensional inner product space and
that Uis a subspace of V.L e t U
⊥be the orthocomplement of U,and let
S={u1,u2,...,u k}be an orthonormal basis of U.L e t x∈V.D e fine
x1=k
j=1x, u jXuj,a n d x2=x−x1. Then it follows that x1∈Uand
x2∈U⊥.M o r e o v e r , x=x1+x2. We summarize this in the following.
Proposition 2.4.2. LetVis a vector space with inner product ·,·Xand
with subspace U. Then every vector x∈Vc a nb ew r i t t e na sas u mo ft w o
orthogonal vectors x=x1+x2,w h e r e x1∈Uandx2∈U⊥.
62 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
Geometrically what this results asserts is that for a given subspace of
an inner product space, every vector has an orthogonal decomposition astwo unique sum of a vector from the subspace and its orthocomplement.We write the vector components as the respective projections of the givenvector to the orthogonal subspaces
x
1=PUx
x2=PU⊥x
Such decompositions are important in the analysis of vector spaces and
matrices. In the case of vector spaces, of course, the representation ofvectors is of great value. In the case of matrices, this type of decompositionserves to allow reductions of the matrices while preserving the informationthey carry.
2.4.1 An important equality for matrix multiplication and
the inner product
LetA∈Mmn(C). Then we know that both A∗AandAA∗(Alternatively,
ATAandAATexist) exist, and we can surely inquire about the rank of these
matrices. The main result of this section is on the rank of ATA,n a m e l y
that r(A)=r(A∗A)=r(AA∗). The proof is quite simple but requires an
important equality. Let A∈Mmn(C)a n d v∈Cnandw∈Cm.Then
Av, wX=rm3
i=1(Av)i,wiS
=m3
i=1n3
j=1aijvj¯wi
=n3
j=1vjm3
i=1aij¯wi
=n3
j=1vjm3
i=1¯aijwi
=n3
j=1vj(A∗w)j
=v,A∗wX
2.4. ORTHOGONALITY 63
As a consequence we have A∗Av, wX=Av, AwXif both v, w∈Cn.T h i s
important equality allows the adjoint or transpose matrices to be used oneither side of the inner product, as needed. Indeed we shall use this below.
Proposition 2.4.3. LetA∈M
mn(C)have rank r(A).Then
r(A)=r(A∗A)=r(AA∗)
Proof. Assume that r(A)=k.Then there are kstandard vectors ej1,..., e jk
such that for each l=1,2,...k, the vectors Aejlis one of the linearly
independent columns of A.Moreover, it also follows that for every set of
constants a1,...,a kthe vector ADalelji
W=0.Now A∗ADalelji
W=0
follows because
?
A∗Ap3
aleljQ
,p3
aleljQ#
=?
Ap3
aleljQ
,Ap3
aleljQ#
=EEEAp3
a
leljQEEE2
W=0
This in turn establishes that A∗Acannot be zero on a linear space of di-
mension kexcept for the zero element of course, and since the rank of A∗A
cannot be larger than kthe result is proved.
Remark 2.4.2. This establishes (6) of the Corollary 2.3.1 above. Also, it
is easy to see that the result is also true for real matrices.
2.4.2 The Legendre Polynomials
When a vector space has an inner product, it is possible to construct anorthogonal basis from any given basis. We do this now for the polynomial
space P
n(−1,1) and a particular basis.
Consider the space the polynomials of degree ndefin e do nt h ei n t e r v a l
[−1,1] over the reals .Recall that this is a vector space and has as a basis
themonomials\
1,x ,x2,...,xn
.We can de fine an assortment of inner
products on this space, but the most common inner product is given by
p, qX=81
−1p(x)q(x)dx
Verifying the inner product properties is fairly straight forward and we leave
it as an exercise. This inner product also de fines a norm
,p,2=81
−1|p(x)|2dx
64 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
This norm satis fies the triangle inequality requires an integral version of the
Cauchy-Schwartz inequality.
Now that we have an inner product and norm, we could proceed to find
an othogonal basis of Pn(−1,1) by applying the Gram-Schmidt procedure
to the basis\
1,x ,x2,...,xn
. This procedure can be clumsy and tedious.
It is easier to build an orthogonal basis from scratch. Following traditionwe will use capital letters P
0,P1,... to denote our orthogonal polynomials.
Toward this end take P0= 1. Note we are numbering from 0 onwards so
that the polynomial degree will agree with the index. Now let P1=ax+b.
For orthogonality, we need
81
−1P0(x)P1(x)dx=81
−11·(ax+b)dx=2b=0
Thus b=0a n d acan be arbitrary. We take a=1.This gives y1=x.Now
we assume the model for the next orthogonal function to be y2=ax2+bx+c.
This time there are two orthogonality conditions to satisfy.
81
−1P0(x)P2(x)dx=81
−11·D
ax2+bx+ci
dx=2
3a+2c=0
81
−1P1(x)P2(x)dx=81
−1x·D
ax2+bx+ci
dx=2
3b=0
We conclude that b=0.From the equation2
3a+2c= 0, we can assign one
of the variables and solve for the other one. Following tradition we takec=−
1
2and solve for ato get a=3
2.
The next polynomial will be modeled as P3(x)=ax3+bx2+cx+d.
Three orthogonality relations need to be satis fied.
81
−1P0(x)P3(x)dx=81
−11·D
ax3+bx2+cx+di
dx=2
3b+2d=0
81
−1P1(x)P3(x)dx=81
−1x·D
ax3+bx2+cx+di
dx=2
5a+2
3c=0
81
−1P2(x)P3(x)dx=81
−11
2(3x−1)D
ax3+bx2+cx+di
dx
=3
5a−1
3b+c−d=0
It is easy to see that b=d=0( w h y ? ) a n df r o m2
5a+2
3c=0,we select
2.4. ORTHOGONALITY 65
c=−3
2anda=5
2.Our table of orthogonal polynomials so far is
k Pk(x)
0 1
1 x
21
2(3x−1)
31
2D
5x3−3xi
Continue in this fashion, generating polynomials of increasing order each
orthogonal to all of the lower order ones.
P0(x)=1
P1(x)= x
P2(x)=3 /2x2−1/2
P3,x)=5 /2x3−3/2x
P4(x)=35
8x4−15
4x2+3/8
P5(x)=63
8x5−35
4x3+15
8x
P6(x)=231
16x6−315
16x4+105
16x2−5
16
P7(x)=429
16x7−693
16x5+315
16x3−35
16x
P8(x)=6435
128x8−3003
32x6+3465
64x4−315
32x2+35
128
P9(x)=12155
128x9−6435
32x7+9009
64x5−1155
32x3+315
128x
P10(x)=46189
256x10−109395
256x8+45045
128x6−15015
128x4+3465
256x2−63
256
2.4.3 Orthogonal matrices
Besides sets of vectors being orthogonal, there is also a de finition of orthog-
onal matrices. The two notions are closely linked.
Definition 2.4.6. We say a matrix A∈Mn(C)i sorthogonal ifA∗A=I.
The same de finition applies to matrices A∈Mn(R)w i t h A∗replaced by
AT.
66 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
For example, the rotation matrices (Exercise ??)Bθ=}cosθ−sinθ
sinθcosθ]
are all orthogonal.
A simple consequence of this de finition is that the rows and the columns
ofAareorthonormal . We see for example that when Ais orthogonal then
(A∗)2=(A−1)2=A−1A−1=(A2)−1. Such a de finition applies, as well to
higher powers. For instance, if Ais orthogonal then Amis orthogonal for
every positive integer m.
One way to generate orthogonal matrices in Cn(orRn)i st ob e g i nw i t h
an orthonormal basis and arrange it into an n×nmatrix either as its columns
or rows.
Theorem 2.4.2. (i) Let {xi},i=1,...,n be an orthonormal basis of Cn
or(Rn). Then the matrices
U=
x1···xn
↓···↓
··
and V=
x1−→ ·
......
xn−→ ·
formed by arranging the vectors xias its respective columns or rows are
orthogonal.
(ii) Conversely, Uis an orthogonal matrix, the sets of its rows and
columns are each orthonormal, an d moreover each forms a basis of Cnor
(Rn).
The proofs are entirely trivial. We shall consider these types of results
in more detail later in Chapter 4. In the meantime there are a few moreinteresting results that are direct consequences of the de finition and facts
about the transpose (adjoint).
Theorem 2.4.3. LetA, B∈M
n(C)(orMn(R))be orthogonal matrices.
Then(a)Ais invertible and A
−1=A∗.
(b) For each integer k=0,±1,±2,...,b o t h Akand−Akare orthogonal.
(c)AB is orthogonal.
2.5 Determinants
This section is about determinants that can be regarded as a measure of
singularity of a matrix. More generally, in many applied situations that
deal with complex objects, a single number is sought that will in some way
2.5. DETERMINANTS 67
classify an aspect of those objects. The determinant is such a measure for
singularity of the matrix. The determinant is di fficult to calculate and of
not much practical use. However, it has considerable theoretical value andcertainly has a place of historical interest.
Definition 2.5.1. LetA∈M
n(F). De fine the determinant ofAto be
the value in F
detA=3
σXn
i=1aiσ(i)~
·sgnσ
whereσis a permutation of the integers {1,2,... ,n }and
(1)
σdenotes the sum over all permutations
(2) sgnσ=s i g no f σ=±1
Atransposition is the exchange of two elements of an ordered list with
all others staying the same. With respect to permutations, a transposition
of one permutation is another permutation formed by the exchange of twovalues. For example a transposition of {1,4,3,2}is{1,3,4,2}.T h e sign of
a given permutation σis
(a) +1 ,if the number of transpositions required to bring σto{1,2,... ,n }
is even.
(b)−1,if the number of transpositions required to bring σto{1,2,... ,n }
is odd.
Alternatively, and what is the same thing, we may count the number mof
transpositions required to bring σto{1,2,... ,n }and to compute the sign
is (−1)
m.
Example 2.5.1.
σ1={2,1,3}1↔2−−→ {1,2,3} odd
σ2={2,3,1}3↔1−−→ {2,1,3}1↔2−−→ {1,2,3}even
sgnσ1=−1s g n σ2=+ 1
Proposition 2.5.1. LetA∈Mn.
68 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
(i) If two rows of Aare interchanged to obtain B,t h e n
detB=−detA.
(ii) Given A∈Mn(F). If any row is multiplied by a scalar c,t h er e s u l t i n g
matrix Bhas determinant
detB=cdetA.
(iii) If any two rows of A∈Mn(F)are equal,
detA=0.
Proof. (i) Suppose rows i1andi2are interchanged. Now for the given
permutations σapply the transposition i1↔i2to getσ1.T h e n
n
i=1ai1σ(i)=n
i=1bi2σ1(i)
because
ai1σ(i1)=bi2σ1(i2)
as
bi2j=ai1jandσ1(i2)=σ2(i1)
and similarly ai2σ(i2)=bi1σ1(i1). All other terms are equal. In the
computation of the full determinant with signs of the permutations,
we see that the change is caused only by the fact sgn( σ1)=−sgn(σ).
Thus,
detB=−detA.
(ii) Is trivial.
(iii) If two rows are equal then by part (i)
detA=−detA
and this implies det A=0 .
2.5. DETERMINANTS 69
Corollary 2.5.1. LetA∈Mn.I fAhas two rows equal up to a multiplica-
tive constant, it has has determinant zero.
What happens to the determinant when two matrices are added. The
result is too complicated to write down is not very important. However,when a single vector is added to a row or a column of a matrix, then theresult can be simply stated.
Proposition 2.5.2. Suppose A∈M
n(F).S u p p o s e Bis obtained from A
by adding a vector vto a given row (resp. column) and Cis obtained from
Aby replacing the given row (resp. column) by the vector v.T h e n
detB=d e t A+d e t C.
Proof. Assume the jthrow is altered. Using the de finition of the determi-
nant,
detA=3
σsgn(σ)
ibiσ(i)=3
σsgn(σ)
iW=jbiσ(i)
bjσ(j)
=3
σsgn(σ)
iW=jaiσ(i)
(a+v)jσ(j)
=3
σsgn(σ)
iW=jaiσ(i)
ajσ(j)+3
σsgn(σ)
iW=jaiσ(i)
vjσ(j)
detA+d e t C
For column replacement the proof is similar, particularly using the alternate
representation of the determinant given in Exercise 15.
Corollary 2.5.2. Suppose A∈Mn(F)andBis obtained by multiplying a
given row (resp. column) of Aby a scalar and adding it to another row
(resp. column), then
detB=d e t A.
Proof. First note that in applying Proposition 2.5.2 Chas two rows equal
up to a multiplicative constant. Thus det C=0 .
Computing determinants is usually di fficult and many techniques have
been devised to compute them out. As is evident from counting, computing
70 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
the determinant of an n×nmatrix using the de finition above would require
the expression of all n! permutations of the integers {1,2,..., n }and the
determination of their signs together with all the concommitant productsand summation. This method is prohibitively costly. Using elementaryrow operations and Gaussian elimination, the evaluation of the determinant
becomes more manageable. First we need the result below.
Theorem 2.5.1. For the elementary matrices the following results hold.
(a) for Type 1 (row interchange) E
1
detE1=−1
(b) for Type 2 (multiply a row by a constant c)E2
detE2=c
(c) for Type 3 (add a multiple of one row to another row) E3
detE3=1.
Note that (c) is a consequence of Corollary 2.5.2. Proof of parts (a) and (b)
are left as exercises. Thus for any matrix A∈Mn(F)w eh a v e
det(E1A)=−detA=d e t E1detA
det(E2A)=cdetA =d e t E2detA
det(E3A)=d e t A =d e t E3detA.
Suppose F1...F kis a sequence of row operations to reduce Ato its RREF.
Then
FkFk−1...F 1A=B.
Now we see that
detB=d e t ( FkFk−1...F 1A)
=d e t ( Fk)d e t ( Fk−1...F 1A)
=...
=d e t ( Fk)d e t ( Fk−1)...det(F1)d e tA.
ForBin RREF andB∈Mn(F), we have that Bis upper triangular. The
next result establishes Theorem 2.3.4( k) about the determinant of singular
and non singular matrices. Moreover, the determinant of triangular matrices
is computed simply as the product of its diagonal elements.
2.5. DETERMINANTS 71
Proposition 2.5.3. LetA∈Mn.T h e n
(i) If r(A)<n,t h e n detA=0.
(ii) If Ais triangular then
detA=
aii
(iii) If r(A)=n,t h e n detAW=0.
Proof. (i) If r(A)<n, then its RREF has a row of zeros, and det A=0b y
Theorem 2.5.1. (ii) If Ais triangular the only product without possible zero
entries isaii.H e n c e d e t A=aii. (iii) If If r(A)=n, then its RREF has
no nonzero rows. Since it is square and has a leading one in each column,
it follows that the RREF is the identity matrix. Therefore det AW=0 .
Now let A, B∈Mn.I fAis singular the RREF must have a zero row.
It follows that det A=0 . I f Ais singular it follows that ABis singular.
Therefore
0=d e t AB=d e t AdetB.
The same reasoning applies if Bis singular. If AandBare not singular
both AandBcan be row reduced to the identity. Let F1...F k1be the row
operations that reduce AtoI,a n d G1...G kBbe the row operations that
reduce BtoI.T h e n
detA=[ d e t ( F1)...det(FkA)]−1
detB=[ d e t ( G1)...det(GkB)]−1.
Also
I=(GkB...G 1)(FkB...F 1)AB
and we have
detI=( d e t A)−1(detB)−1detAB.
This proves the
Theorem 2.5.2. IfA, B∈Mn(F),detAB=d e t AdetB.
72 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
2.5.1 Minors and Determinants
The method of row reduction is one of the simplest methods to compute
the determinant of a matrix. Indeed, it is not necessary to use Type 2 el-
ementary transformation. This results in the computing the determinantas the product of the diagonal elements of the resulting triangular matrixpossibly multiplied by a minus sign. An alternate approach to computingdeterminants using minors is both interesting and useful. However, unless
the matrix has some special form, it does not provide a computational al-
ternative to row reduction.
Definition 2.5.2. LetA∈M
n(C).For any row iand column jdefine the
(ij)-minor ofAby
Mij=d e t A
ithrow removed
jthcolumn removed
The notation
A
ithrow removed
jthcolumn removed
denotes the ( n−1)×(n−1) matrix formed from Aby removing the ithrow
andjthcolumn. With minors an alternative formulation of the determinant
can be given. This method, while not of great value computationally, hassome theoretical importance. For example, the inverse of a matrix can beexpressed using minors. We begin by consideration the determinant.
Theorem 2.5.3. LetA∈M
n(C).(i) Fix any row, say row k.The de-
terminant of Ais given by
detA=n3
j=1akj(−1)k+jMkj
(ii) Fix any column, say column m. The determinant of Ais given by
detA=n3
j=1ajm(−1)m+jMjm
Proof. (i) Suppose that k=1.Consider the quantity
a11M11
2.5. DETERMINANTS 73
We observe that this is equivalent to all the products of the form
sgn(σ)a11·a2σ(2)····· anσ(n)
where only permutations that fix the integer (i.e. position) 1 are taken.
Thusσ(1) = 1 .Since this position is fixed the signs taken in the determi-
nant M11for permutations of n−1 integers are respectively the same as the
signs for the new permutation of nintegers.
Now consider all permutations that fix the integer 2 in the sense that
σ(1) = 2. The quantity a12(−1)1+2M12consists of all the products of the
form
sgn(σ)a12·a2σ(1)a3σ(3)····· anσ(n)
We need here the extra sign change because if the part of the permutation
σof the integers {1,3,4,...,n }is of one sign, which is the sign used in
the computation of det Mij, then the permutation of σof the integers
{1,2,3,4,...,n }is of the other sign, and that sign is sgn(σ).
When we proceed to the kthcomponent, we consider permutations that
fixt h ei n t e g e r k.T h a t i s , σ(1) = k. In this case the quantity a1k(−1)1+kM1k
consists of all products of the form
sgn(σ)a1ka2σ(1)····ak−1σ(k−1)ak+1σ(k+1)····· anσ(n)
Continuing in this way we exhaust all possible products a1σ(1)·a2σ(2)·····
anσ(n)over all possible permutations of the integers {1,2,...,n }.T h i s
proves the assertion. The proof for expanding from any row is similar, with
only a possible change of sign needed, which is a prescribed.
(ii) The proof is similar.
Example 2.5.2. Find the determinant of
A=
32−1
01 3
12−1
expanding across the first row and then expanding down the second column.
Solution. Expanding across the first row gives
detA=a11M11−a12M12+a13M12
=3 d e t}13
2−1]
−2d e t}03
1−1]
−1d e t}01
12]
=3 (−7)−2(−3)−(−1)
=−14
74 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
Expanding across the second column gives
detA=−2d e t}03
1−1]
+1d e t}3−1
1−1]
−2d e t}3−1
03]
=−2(−3) + (−2)−2( 9 )=−14
The inverse of the matrix can be formulated in terms of minors, which
is formulated below.
Definition 2.5.3. LetA∈Mn(C)( o r Mn(R)). De fine the adjugate (or
adjoint )m a t r i x ˆAby
ˆAij=(−1)i+jMji
where Mjiis the jiminor.
The adjugate has traditionally been call ed the “adjoint”, but that terminol-
ogy is somewhat ambiguous in light of the previous de finition as complex
conjugate transpose. Note that it is de fined for all square matrices; when
restricted to invertible matrices the inverse appears.
Theorem 2.5.4. LetA∈Mn(C)(orMn(R)) be invertible. Then A−1=
1
detAˆA
Proof. A quick examination of the ij-entry of the product AˆAyields the
following sum
n3
j=1aijˆAjk=1
det (A)n3
j=1aij(−1)k+jMkj
There are two possibilities. (1) If i=k,then the summation above is the
summation to form the determinant as described in Theorem 2.5.3. (2) IfiW=k,the summation is the computation of the determinant of the matrix
Awith the k
throw replaced by the ithrow. Thus the determinant of a
matrix with two identical rows is represented above and this must be zero.
We conclude thatp
AˆAQ
ij=δij,the usual Kronecker ‘delta,’ and the result
is proved.
Example 2.5.1. Find the adjugate and inverse of
A=}24
21]
2.5. DETERMINANTS 75
It is easy to see that
ˆA=}1−4
−22]
Also det A=−6. Therefore, the inverse
A−1=−1
6}1−4
−22]
Remark 2.5.1. The notation for cofactors of a square matrix Ais often
used
ˆaij=(−1)i+jMji
Note the reversed order of the subscripts ijand then jiabove.
Cramer’s Rule
We know now that the solution to the system Ax=bis given by x=A−1b.
Moreover, the inverse A−1is given by A−1=ˆA
detA,w h e r e ˆAis the adjugate
matrix. The the ithcomponent of the solution vector is therefore
xi=1
detAn3
j=1ˆaijbj
=1
detAn3
j=1(−1)i+jMjibj
=detAi
detA
w h e r ew ed e fine the matrix Aito be the modi fication to Aby replacing its
ithcolumn by the vector b. In this way we obtain a very compact formula
for the solution of a linear system. Called Cramer’s rule we state thisconclusion as
Theorem 2.5.1. (Cramer’s Rule.) Let A∈M
n(C)be invertible and b∈
Cn.F o re a c h i=1,, n , define the matrix Aito be the modi fication of A
by replacing its ithcolumn by the vector b. Then the solution to the linear
system Ax=bis given by components xi=detAi
detA,i=1,, n.
76 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
Example 2.5.2. Given the matrix A=}24
21]
, and the vector b=
}2
−1]
.Solve the system Ax=bby Cramer’s rule.
We have
A1=}24
−11]
and A2=}22
2−1]
and det A1=6,detA2=−6,detA=−6. Therefore
x1=−1a n d x2=1
A curious formula
The useful formula using cofactors given below will have some consequence
when we study positive de finite operators in Chapter ??.
Proposition 2.5.1. Consider the matrix
B=
0x
1x2··· xn
x1a11a12···a1n
x2a21a22···a2n
...............
x
nan1an2 ann
Then
detB=−3
ˆa
ijxixj
where ˆaijis the ij-cofactor of A.
Proof. Expand by minors along the top row to get
detB=3
(−1)jxjM1j(B)
Now expand the matrix of M1j(B)d o w nt h e first column. This gives
M1j(B)=3
(−1)i−1xiMij(A)
2.6. PARTITIONED MATRICES 77
Combining we obtain
detB=3
(−1)jxjM1j(B)
=3
(−1)jxj3
(−1)i−1xiMij(A)=
=33
(−1)i+j−1xjxiMij(A)
=−33
ˆaijxjxi
The reader may note that in the last line of the equation above, we should
have used ˆ aij. However, the formulation given is correct, as well. (Why?)
2.6 Partitioned Matrices
It is convenient to study partitioned or “blocked” matrices, or more graph-
ically said, matrices whose entries are themselves matrices. For example,with I
2denoting the 2 ×2i d e n t i t ym a t r i xw ec a nc r e a t et h e4 ×4m a t r i x
w r i t t e ni np a r t i t i o n e df o r ma n de x p a n d e df o r m .
A=}aI2cI2
cI2dI2]
=
a0b0
0a0b
c0d0
0c0d
Partitioning matrices allows our attention to focus on certain structural
properties. In many applications part ititioned matrices appear in a natural
way, with the particular blocks having some system context. Many similarsubclasses and processes apply to partitioned matrices. In speci fics i t u a t i o n s
they can be added, multiplied, and inverted, just like regular matrices. It iseven possible to perform “blocked” version of Gaussian elimination. In the
few results here, we touch on some of these possibilities.
Definition 2.6.1. For each 1 ≤i≤mand 1≤j≤n,letA
ijbe an mi×nj
matrices where . Then the matrix
A=
A
11A12··· A1n
A21A2n··· A2n
............
Am1Am2···Amn
78 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
is a partitioned matrix of order (m)i×(nj).
The usual operations of addition and multiplication of partitioned ma-
trices can be performed provided each of the operations makes sense. Foraddition of two partitioned matrices AandBit is necessary to have the
same numbers of blocks of the respective same sizes. Then
A+B=
A
11A12··· A1n
A21A2n··· A2n
............
Am1Am2···Amn
+
B11B12··· B1n
B21B2n··· B2n
............
Bm1Bm2···Bmn
=
A
11+B11 A12+B12··· A1n+B1n
A21+B21 A2n+B2n··· A2n+B2n
............
Am1+Bm1Am2+Bm2···Amn+Bmn
For multiplication, the situation is a bit more complicated. For de finiteness,
suppose that Bis a partitioned matrix with block sizes s
i×tj,w h e r e1 ≤
i≤pand 1≤j≤qThe usual operations to construct C=AB,
n3
j=1AijBjk
then make sense provided p=nandnj=sj,1≤j≤n.
A special category of partitioned matrices are the so-called quasi-triangular
matrices, wherein Aij=0i f i>j for the “lower” triangular version. The
special subclass of quasi-triangular matrices wherein Aij=0i f iW=jare
called quasi-diagonal. In the case of the multiplication of partitioned
matrices ( C=AB) with the left multiplicand Aa quasi-diagonal matrix,
we have Cik=AiiBik. Thus the multiplication is similar in form to the usual
multiplication of matrices where the left multiplicand is a diagonal matrix.In the case of the multiplication of partitioned matrices ( C=AB)w i t ht h e
right multiplicand Ba quasi-diagonal matrix, we have C
ik=AikBkk.F o r
quasi-triangular matrices with square diagonal blocks, there is an interestingresult about the determinant.
Theorem 2.6.1. LetAbe a quasi-triangular matrix, where the diagonal
blocks A
iiare square. Then
detA=
idetAii
2.7. LINEAR TRANSFORMATIONS 79
Proof. Apply row operations on each vertical block without row interchanges
between blocks, without any Type 2 operations. The resulting matrix ineach diagonal block position ( i, i) is triangular. Be sure to multiply one
of the diagonal entries by ±1,reflecting the number of row interchanges
within a block. The resulting matrix c an still be regarded as partitioned,
though the diagonal blocks are now actually upper triangular. Now apply
Proposition 2.5.3, noting that the product of each of the diagonal entriespertaining to the i
thblock is in fact det Aii.
A simple consequence of this result, proved al´ a Gaussian elimination, is
contained in the following corollary.
Corollary 2.6.1. Consider the partitioned matrix
A=}A11A12
A21A22]
with square diagonal blocks and with A11invertible. Then the rank of Ais
the same as the rank of A11if and only if A22=A21A−1
11A12.
Proof. Multiplication of Aby the elementary partitioned matrix
E=}I 0
−A21A−1
11I]
yields
EA =}I 0
−A21A−1
11I]}A11A12
A21A22]
=}A11 A12
0A22−A21A−1
11A12]
Since Ehas full rank, it follows that rank( EA)=r a n k A.S i n c e EAis
quasi-triangular, it follows that the rank of Ai st h es a m ea st h er a n ko f A11
if and only if A22−A21A−1
11A12=0.
2.7 Linear Transformations
Definition 2.7.1. A mapping Tfrom RntoRmis called a linear trans-
formation if
T(x+y)=Tx+Ty∀x, y∈Rn
T(ax)=aTx ∀a∈F.
80 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
Note: We normally write Txinstead of T(x).
Example 2.7.1. T:Rn→Rm.L e t a∈Rmandy∈Rn.T h e n f o r e a c h
x∈Rn,Tx=x, yXais a linear transformation. Let S={v1...v n}be
a basis of Rn, and de fine the m×nmatrix with columns given by the
coordinates of Tv1,Tv 2,... ,Tv n. Then this matrix
A=^
Tv1Tv2 Tvn
↓↓ ···↓
is the matrix representation of Twith respect to the basis S.T h u s , i f
x=Σaivi, whence [ x]S=(a1...a n), we have
[Tx]S=A[x]S
TThere is a duality between all linear transformations from RntoRm
and the set Mm,n(F).
Note that Mm,n(F) is itself a vector space over F. Hence L(Fn,Fm),
the set of linear transformations from FntoFmis likewise. As such it has
subspaces.
Example 2.7.2. (1) Let ¯ x∈Fn.D efiniteJ={T∈L|T¯x=0}.T h e n
Jis a subspace of L(Fn,Fm).
(2) Let U={T∈L(Rn,Rn)|Tx≥0i fx≥0},w h e r e {x≥0}means
the positive orthant of Rn.Uisnota linear subspace of L(Rn,Rn),
though it is a convex set.
(3) De fineT:Pn→PnbyTp=d
dxp.Tis a linear transformation.
Example 2.7.3. Express the linear transformation D:P3→P3given by
Dp=d
dxp(x) as a matrix with respect to the basis. S={1,x ,x2,x3}.W e
have D1=0=0+0 x+0x2+0x3.A l s o
[D1]S=[ 0,0,0,0]T
similarly
[Dx]S=[ 1,0,0,0]T
[Dx2]S=[ 0,2,0,0]T
[Dx3]S=[ 0,0,3,0]T.
2.7. LINEAR TRANSFORMATIONS 81
Hence
[D]S=
0100
0020
00030000
.
In this context the di fferentiation operator is rather simple.
Example 2.7.4. Consider the linear transformation Tdefined by Tq=
3x
d
dxq+x2qforq∈P2.Find the matrix representation of T.
Solution. First o ffwe notice that this transformation has range in
P4.Let’s use the standard bases for this problem. We then determine
the coordinates of Tfor vectors in the P2basis {1,x ,x2}in the P4basis
{1,x ,x2,x3,x4}.Compute
T(1) = x2
T(x)=3 x+x3
TD
x2i
=6 x2+x4
The coordinates of the input vectors we know are [1 ,0,0]T,[0,1,0]T,and
[0,0,1]T.For the output vectors the coordinates are [0 ,0,1,0,0]T,[0,3,0,1,0]T,
and [0 ,0,6,0,1]T.So, with respect to these two bases, the matrix of the
transformation is
A=
000
030
106
010001
Observe that the dimensionality corresponds with the dimentionality of the
respective spaces.
Example 2.7.5. LetV=R
2,w i t h S0={v1,v2}={[1
0],[1
1]},S1=
{w1,w2}=\
[1
2],J−2
1o
,a n d T=Ithe identity. The vectors above are
expressed in the standard E={e1,e2},Tvj=Ivj=vj.T o find [vj]S1we
solve
vj=ajw1+βjw2
82 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
v1:}1−2
21]}α1
β1]
=}1
0]
−→}α1
β1]
=}1
5
−2
5]
v2:}1−2
21]}α2
β2]
=}1
1]
−→}α2
β2]
=}3
5
−1
5]A
tsolve linear
systems
Therefore
S1[I]S0=}1
53
5
−2
5−1
5]
←change of
basis
matrix
If
[x]S0=}−1
2]
[x]S1=S1[I]S0}−1
2]
=1
5}13
−2−1]}−1
2]
=1
5}5
0]
=}1
0]
.
Note the necessity of using the standard basis to express the vectors in both
bases S0andS1.
2.8 Change of Basis
LetVbe a vector space with bases S0={v1...v n}andS1={w1...w n},
and suppose T:V→Vis a linear transformation. We want to find the
representation of Tas a matrix that takes a vector xg i v e ni nt e r m so fi t s S0
coordinates and produces the vector Txg i v e ni nt e r m so fi t s S1coordinates.
We know that x→[x]S0is well de fined. The action of Tis known if the
nvector [ x]S0=}c1...cn]
and the vectors Tv1,Tv 2,... ,Tv nare known, for if
x=Σcjvj,t h e n Tx=ΣcjTvj,b yl i n e a r i t y .
To determine [ Tx]S1we need to convert the Tvj,j=1,...,n to coor-
dinates in the other S1basis, This is done as follows. Find
[Tvj]S1=
t
1j
t2j
...
tnj
j=1,2,... ,n .
2.8. CHANGE OF BASIS 83
Then if x∈V
[Tx]S1=[ΣcjTvj]S1=Σcj[Tvj]S1
=
3
jtijcj
=
t11... t 1n
tn1 tnn
c1
...
cn
.
This n×narray [ tij] depends on T,S 0andS1but not on x.W ed e fine the
S0→S1basis representation of Tto be [ tij], and we write this as
S1[T]S0=
t11... t 1n
.........
tn1... t nn
.
In the special case that Tis the identity operator the matrix S1[I]S0converts
the coordinates of a vector in the basis S0to coordinates in the basis S1.I t
is easy to see that S0[I]S1must be the inverse of S1[I]S0and thus
S0[I]S1·S1[I]S0=I.
We can also establish the equality
S1[T]S1=S1[I]S0S0[T]S0S0[I]S1.
In this way we see that the matrix representation of Tdepends on the bases
involved. If Xi sa n yi n v e r t i b l em a t r i xi n Mn(F)w ec a nw r i t e
B=X−1AX.
The interpretation in this context is clear
X: change of coordinate from one basis to another S0→S1
X−1: change of coordinate S1→S0
A: matrix of the linear transformation in the basis S0
B: matrix of the same linear transformation in the basis S1.
With this in mind it seems prudent to study linear transformations in the
basis that makes their matrix representation as simple as possible.
84 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
Example 2.8.1. LetA=}13
−11]
be the matrix representation of a lin-
ear transformation given with respect to the standard basis S0={e1,e2}=
{(1,0),(0,1)}Find the matrix representation of this transformation with
resepect to the basis S1={v1,v2}={(2,1),(1,1)}.
Solution. According to the analysis above we need to determine S0[I]S1and
S1[I]S0.Of course S1[I]S0=S0[I]−1
S1.Since the coordinates of the vectors
inS1are expressed in terms of the basis vectors S0we obtain directly
S0[I]S1=}21
11]
Its inverse is given by
S0[I]−1
S1=}1−1
−12]
Assembling these matrices we have the final matrix converted to the new
basis.
S1[A]S1= S1[I]S0AS0[I]S1
=}1−1
−12]}13
−11]}21
11]
=}64
−7−4]
Example 2.8.2. Consider the same problem as above except that the ma-
trixAi sg i v e ni nt h eb a s i s S1. Find matrix representation of this transfor-
mation with resepect to the basis S0.
Solution. To solve this problem we need to determine S0[A]S0=S0[I]S1AS1[I]S0.
As we already have these matrices, we determine that
S0[A]S0= S0[I]S1AS1[I]S0.
=}21
11]}13
−11]}1−1
−12]
=}−61 3
−48]
2.9. APPENDIX A – SOLVING LINEAR SYSTEMS 85
2.9 Appendix A – Solving linear systems
The key to solving linear systems is to reduce the augmented system to
RREF and solve the resulting equations. While this may be so, there is anintermediate step that occurs about half way through the computation ofthe RREF where the reduced matrix achieves an upper triangular form. Atthis point the solution can be determined directly by back substitution. To
clarify the rules on back substitution, suppose that we have the triangular
form
a
11a12···a1n
0a22···a2n
.........
0··· 0anneeeeeeeeeb
1
b2
...
bn
Assuming that the diagonal part consists of all nonzero terms, we can solve
this system by back substitution. First solve for x
n=bn
ann. Now inductively
solve for the remaining solution coordinates using the formula
xn−j=1
bn−j,n−j^j−13
k=0an−j,n−kxn−k
,j =1,2,..., n−1
This inconvenient looking formula can be replaced by
xj=1
bjj
n3
k=j+1ajkxk
,j =n−1,n−2,..., 1
where the index runs from j=n−1u pt o j=1.The upshot is that the
row reduction process can be halted when a triangular-like form has beenattained. The applies as well to nonsingular and non square systems, wherethe the process is stopped when all the leading ones have been identi fied,
entries below them have been zeroed out, and all the zero rows are present.
The principle reason for using back substitution is to reduce the number
of computations required, an important consideration in numerical linearalgebra. In the example below we solve a 3 ×3 nonsingular system.
Example 2.9.1. Solve Ax=bwhere
A=
120
22−1
−13 2
b=
3
6
−2
86 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
Solution. Find the RREF of [ A|b]. Then solve Ax=b.
120
22−1
−13 2eeeeee3
6
−2
−2R
1+R2
→
R1+R3
12 0
0−2−1
05 2eeeeee3
0
1
−
1
2R2
→
120
011
2
052eeeeee3
0
1
−5R
2+R3
→
12 0
011
2
00−1
2eeeeee3
01
(∗)
−
1
2R3+R2
→
120
010001eeeeee3
1
−2
−2R
3
→
120
011
2
001eeeeee3
0
−2
−2R
2+R1
→
100
010001eeeeee1
1
−2
Hence solving we obtain x
3=−2,x 2=1,andx1=1 . T h i si s fine,
but there is a faster way to solve this system. Stop the reduction when
the system attains a triangular form at ( ∗).From this point solve to obtain
x3=−2.Now back substitute x3= 2 into the second row (equation) to
solve for x2.T h u s x2=−1
2(−2) = 1 .Finally, back substitute x3=2 a n d
x2= 1 into the first row (equation) to solve for x1.T h u s x1=3−2( 1 )=1 .
Sometimes the form ( ∗) is called the row reduced form.
Example 2.9.2. Given the augmented system for Ax=bis in RREF.
12000 0
00120 −1
00001 300000 0eeeeeeee4
110
2.9. APPENDIX A – SOLVING LINEAR SYSTEMS 87
Find the solution.
Solution. The leading ones occur in columns 1, 3, and 5. The values in
columns 2, 4, and 6 can be taken as free parameters. So, take x2=r, x 4=s,
andx6=t. Now solving for the other varables we have
x1=4−2r
x3=1−2s+t
x5=1−3t
The solution set is comprised of the vectorx=[ 4−2r, r,1−2s+t, s, 1−t, t]
T
=[ 4 ,0,1,0,1,0]T+r[−2,1,0,0,0,0]T+s[0,0,−2,1,0,0]T+t[0,0,1,0−3,1]T
for all r, s, andt.We can rewrite this as the set
S=
4
0
1010
+r
−2
1
0000
+s
0
0
−2
100
+t
0
0
10
−3
1
eeeeeeeeeeeer, s, t∈RorC
This representation shows better the connection between the free constants
and the component vectors that make up the solution. Note this expressionalso reveals the solution of the homogeneous solution Ax=0a st h es e t
+
r[−2,1,0,0,0,0]
T+s[0,0,−2,1,0,0]T+t[0,0,1,0−3,1]Teeer, s, t∈RorC
Indeed, this is a full subspace.
Example 2.9.3. The RREF can be used to determine the inverse, as well.
Given the matrix A∈M
n,the inverse is given by the matrix Xfor
which AX =I. I nt u r nw i t h x1, ..., x nrepresenting the columns of
Xande1, ..., e nrepresenting the standard vectors we see that Axj=
ej,j=1,2,...,n . To solve for these vectors, form the augmented matrix
[A|ej],j=1,2,..., n and row reduce as above. A massive short cut to
this process is to augment all the standard vectors at one and row reduce the
88 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
resulting n×2nmatrix [ A|I]. If Ais invertible, its RREF is the identity.
Therefore,
[A|I]row
→
operations[I|X]
and, of course, A−1=X.T h u s , f o r
A=
−21 0
1123−2−1
we row reduce [ A|I]a sf o l l o w s
[A|I]=
−21 0
1123−2−1eeeeee100
010001
row
→
operations
100
010001eeeeee312
724
−5−1−3
2.10 Exercises
1. Consider the di fferential operator T=2xd
dx(·)−4 acting on the vector
space of cubic polynomials, P3.Show that Tis a linear transformation
andfind a matrix representation of it. Assume the basis is given by
{1,x ,x2,x3}.
2. (i) Find matrices AandB, each with positive rank, for which r(A+
B)=r(A)+r(B). (ii) Find matrices AandB,e a c hw i t hp o s i t i v e
rank, for which r(A+B) = 0. (iii) Give a method to find two nonzero
matrices AandBfor which the sum has any preassigned rank. Of
course, the matrix sizes may depend on this value.
3. Find square matrices AandBfor which r(A)=r(B)=2a n df o r
which r(AB)=0 .
4. Suppose that A is an m×nmatrix and that xis a solution of Ax=b
over the prescribed field. Show that every solution of Ax=bhave the
form x+x0,w h e r e x0is a solution of Ax0=0 .
2.10. EXERCISES 89
5. Find a matrix A∈Mnof rank n−1f o rw h i c h r(Ak)=n−kfork≤n.
Is it possible to begin this process with a matrix A∈Mnof rank n
and for which r(Ak)=n−k+1 f o r k≤n?
6. Show that Ax=bhas a solution if and only if yTb=0i fa n do n l yi f
yTA= 0 for some column vector.
7. In R2the linear transformation that rotates any vector by θradians
counter clockwise Thas matrix representation with respect to the
standard basis given by
A=}cosθ−sinθ
sinθcosθ]
What is the matrix representation with respect to the standard basis
of the transformation that rotates any vector by θradians clockwise?
What is the relation between the matrices?
8. Show that if B,C∈Mn(F), where Bis symmetric and Cis skew-
symmetric, then B=Cimplies that B=C=0 .
9. Prove the general formula for the inverse of the 2 ×2m a t r i x A=}ab
cd]
isA−1=1
detA}d−b
−ca]
.
10. Prove Theorem 2.5.1(a).
11. Prove Theorem 2.5.1(b).
12. Prove that every permutation σmust have an inverse σ−1(i . e .σ−1(σ(j)) =
j), and the signs of σ−1andσare the same.
13. Show that the sign of every transposition is −1.
14. Prove that det A=
σ(−1)sgn(σ)w
iaσ(i)iW
15. Prove Proposition 2.5.2 using minors.
16. Suppose that the n×nmatrix Ais singular. Show that each column
of the adjugate matrix ˆAis a solution of Ax=0 . ( M c D u ffee, Chapter
3, Theorem 29.)
90 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
17. Suppose that Ais an ( n−1)×nmatrix, and consider the homogeneous
system Ax=0f o r x∈Rn.D e finehito be the determinant of the
(n−1)×(n−1) matrix formed by removing the ithcolumn of A. Show
that the vector h=(h1, ..., h n)Tis a solution to Ax=0 . ( M c D u ffee,
Chapter 3, Corollary 29.)
18. Show that A=J1−1
−11o
has no inverse by trying to solve AB=IThat
is, assume the form
B=}ab
cd]
multiply the matrices ( AandB) together, and then solve for the un-
knowns a, b, c, andd. (This is not a very e fficient way to determine
inverses of matrices. Try the same thing for any 3 ×3m a t r i x . )
19. Prove that the elementary equation operations do not change the so-
lution set of a linear system.
20. Find the inverses of E1,E2,a n d E3.
21. Find the matrix representation of linear transformation Tthat rotates
any vector by θradians counter clockwise (ccw) with respect to the
basis S={(2,1),(1,1)}.
22. Consider R3.Suppose that we have angles {θi}3
i=1and pairs of co-
ordinate vectors {(e1,e2),(e1,e3),(e2,e3)}.LetTbe the linear trans-
formation that successively rotates a vector in the respective planes
{(ei1,ei2)}k
i=1through the respective angles {θi}k
i=1.Find the matrix
representation of Twith respect to the standard basis .Prove that it
is invertible.
23. Prove Theorem 2.2.4.24. Prove or disprove the equivalence of the linear systems.
2x−3y=−1
x+4y=5−x+4y=3
x+2y=3
25. Find basis for the orthoc omplement of the subspace of R
3spanned by
the vectors {[2,1,1]T,[1,1,2]T}.
26. Find basis for the orthoc omplement of the subspace of R3spanned by
the vector [1 ,1,1]T.
2.10. EXERCISES 91
27. Consider planar rotations in Rnwith respect to the standard bases
elements .Prove that there must ben(n−1)
2of them – discounting
the particular angle. Display the general representation of any of
them. Prove or disprove that any two of them are commutative. Thatis for two angles {θ
i}2
i=1and pairs of coordinate vectors {(ei1,ei2)}k
i=1
the respective counter clockwise rotations are commutative.
28. Suppose that A∈Mmk,B∈Mknand both have rank k.Show that
the rank of ABisk.
29. Suppose that A∈Mmkhas rank k.P r o v e t h a t ARREF =}Ik
0]
where Ikis the identity matrix of size kand 0 is the m−k×kzero
matrix.
30. Suppose that B∈Mknhas rank k.P r o v e t h a t BRREF =J
Ik0o
where Ikis the identity matrix of size kand 0 is the k×n−kzero
matrix.
31. Determine and prove a version of Corollary 2.6.1 for 3 ×3b l o c k e d
matrices, where we assume the diagonal blocks A11is invertible and
wish to conclude the result that the rank of Ais the equal to the rank
ofA11.
32. Suppose that we have angles {θi}k
i=1and pairs of coordinate vectors
{(ei1,ei2)}k
i=1.LetTbe the linear transformation that successively
rotates a vector in the respective planes {(ei1,ei2)}k
i=1through the
respective angles {θi}k
i=1.Prove that the matrix representation of the
linear transformation with respect to any basis must be invertible.
33. The super-diagonal of a matrix is the set of elements ai,i+1.The subdi-
agonal of a matrix is the set of elements ai−1,i.A tri-banded matrix is
one for which the entries are zero above the super-diagonal and belowthe subdiagonal. Suppose that for an n×ntri-banded matrix T,w e
have a
i−1,i=a, a ii=0,andai,i+1=c.Prove the following facts:
(a) If nis odd det A=0.
(b) If n=2mis even det T=(−1)mamcm.
34. For the banded matrix of the previous example, prove the following
for the powers TpofT.
92 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
(a) If pis odd, prove that ( Tp)ij=0i f i+jis even.
(b) If pis even, prove that ( Tp)ij=0i f i+jis odd.
35. Consider the vector space P2(1,2) with inner product de fined byp, qX=$2
1p(x)q(x)dx.Find an orthogonal basis of P2(1,2).(Hint. Begin
with the standard basis {1,x ,x2}.Apply the Gram-Schmidt procedure.)
36. For what values of aandbis the matrix below singular
A=
a21
21 b
1a−2
37. The Vandermonde matrix, de fined for a sequence of numbers {x1,...x n},
is given by the n×nmatrix
Vn=
1x
1x2
1···xn−1
1
1x2x2
2···xn−1
2
...............
1xnx2
n···xn−1
n
Prove that the determinant is given by
detV
n=n
i>j=1(xi−xj)
38. In the case the x-values are the integers {1,...,n },p r o v et h a td e t Vn
is divisible byn
i=1(i−1)!.(These numbers are called superfactorials.)
39. Prove that for the weighted functional de fined in Remark 2.4.1, it is
necessary and su fficient that the weights be strictly positive for it to
be an inner product.
40. For what values of aandbis the matrix below singular
A=
b0a0
00 ba
0a0b
ba 00
2.10. EXERCISES 93
41. Find an orthogonal basis for R2from the vectors {(1,2),(2,1)}.
42. Find an orthogonal basis of the subspace of R3spanned by {(1,0,1),(0,1,−1)}
43. Suppose that Vis a vector space with an inner product, and S⊂V.
Show that if Sis a basis of V,S⊥={0}.
44. Suppose that Vis a vector space with an inner product, and S⊂V.
Show that if U=S(S), then U⊥=S⊥.
45. Let A∈Mmn(F). Show it may not be true that r(A)=r(ATA)=
r(AAT) unless F=R, in which case it is true.
46. If A∈Mn(C) is orthogonal, show that the rows and columns of Aare
orthogonal.
47. If Ais orthogonal then Amis orthogonal for every positive integer m.
(This is a part of Theorem 2.4.3(b).)
48. Consider the polynomial space Pn[−1,1] with the inner product p, qX=$1
−1p(t)q(t)dt.Show that every polynomial p∈Pnfor which p(1) =
p(−1) = 0 is orthogonal to its derivative.
49. Consider the polynomial space Pn[−1,1] with the inner product p, qX=$1
−1p(t)q(t)dt.Show that the subspace of polynomials in even pow-
ers (e.g. p(t)=t2−5t6) is orthogonal to the subspace of polynomials
in odd powers.
50. Let A=}13
−11]
be the matrix representation of a linear trans-
formation given with respect to the standard basis S0={e1,e2}=
{(1,0),(0,1)}Find the matrix representation of this transformation
with resepect to the basis S1={v1,v2}={(2,−3),(1,−2)}.
51. Show that the sign of every transposition from the set {1,2, ..., n }is
−1.
52. What are the signs of the permutations {7, 6, 5, 4, 3, 2, 1 }and
{7,1, 6, 4, 3, 5, 2 }of the integers {1, 2, 3, 4, 5, 6, 7 }?
53. Prove that the sign of the permutation {m, m−1,..., 2,1}is (−1)m.
54. Suppose A, B∈Mn(F). IfAis singular, use a row space argument to
show that det AB=0 .
94 CHAPTER 2. MATRICES AND LINEAR ALGEBRA
55. Show that if A∈Mm,nandB∈Mm,nand both r(A)=r(B)=m.I f
r(AB)=m−k, what can be said about n?
56. Prove that
deteeeeeeeex
1x2x3x4
−x2x1−x4x3
−x3x4x1−x2
−x4−x3−x2x1eeeeeeee=D
x
2
1+x2
2+x2
3+x2
4i2
57. Prove that there is no invertible 3 ×3 matrix that has all the same
cofactors. What similar statement can be made for n×nmatrices?
58. Show by example that there are matrices AandBfor which lim n→∞An
and lim n→∞Bnboth exist, but for which lim n→∞(AB)ndoes not ex-
ist.
59. Let A∈M2(C). Show that there is no matrix solution B∈M2(C)
toAB−BA=I. What can you say about the same problem with
A, B∈Mn(C)?
60. Show by example that if AC=BCthen it does not follow that A=B.
However, show that if Cis inveritble the conclusion A=Bis valid.
Chapter 3
Eigenvalues and Eigenvectors
In this chapter we begin our study of the most important, and certainly the
most dominant aspect, of matrix theory. Called spectral theory, it allows usto give fundamental structure theorems for matrices and to develop power
tools for comparing and computing w i t hm a t r i c e s . W eb e g i nw i t has t u d y
of norms on matrices.
3.1 Matrix Norms
We know Mnis a vector space. It is most useful to apply a metric on this
vector space. The reasons are manifold, ranging from general information ofa metrized system to perturbation theory where the “smallness” of a matrixmust be measured. For that reason we de fine metrics called matrix norms
that are regular norms with one additional property pertaining to the matrix
product.
Definition 3.1.1. LetA∈M
n.R e c a l l t h a t a norm ,,·,,o na n yv e c t o r
space sati fies the properties:
(i),A,≥0a n d |A,=0i fa n do n l yi f A=0
(ii),cA,=|c|,A,forc∈R
(iii),A+B,≤, A,+,B,.
There is a true vector product on Mndefined by matrix multiplication. In
this connection we say that the norm is submultiplicative if
(iv),AB,≤, A,,B,
95
96 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
In the case that the norm ,·,satifies all four properties (i) - (iv) we call it
amatrix norm .
Here are a few simple consequences for matrix norms. The proofs are
straightforward.
Proposition 3.1.1. Let,·,be a matrix norm on Mn, and suppose that
A∈Mn.T h e n
(a),A,2≤,A,2,,Ap,≤, A,p,p=2,3,. . .
(b) If A2=Athen,A,≥1
(c) If Ais invertible, then ,A−1,≥,I,
,A,
(d),I,≥1.
Proof. The proof of (a) is a consequence of induction. Supposing that
A2=A,we have by the submultiplicativity property that ,A,=EEA2EE≤
,A,2. Hence ,A,≥ 1, and therefore (b) follows. If Ais invertible, we
apply the submultiplicativity again to obtain ,I,=EEAA−1EE≤,A,EEA−1EE,
whence (c) follows. Finally, (d) follows because I2=Iand (b) applies.
Matrices for which A2=Aare called idempotent . Idempotent matrices
turn up in most unlikely places and are useful for applications.
Examples. We can easily apply standard vector space type norms, i.e. f1,
f2,a n df∞to matrices. Indeed, an n×nmatrix can clearly be viewed as
an element of Cn2w i t ht h ec o o r d i n a t es t a c k e di nr o w so f nnumbers each.
The trick is usually to verify the submultiplicativity condition (iv).1.f
1.D efine
,A,1=3
i,j|aij|
The usual norm conditions (i)—(iii) hold. To show submultiplicativity we
write
,AB,1=3
ijeeeee3
kaikbkjeeeee
≤3
ij3
k|aik|bkj|
≤3
ijkm|aik||bmj|=3
i,k|aik|3
mj|bmj|
=,A,1,B,1.
3.1. MATRIX NORMS 97
Thus,A,1is a matrix norm.
2.f2.D efine
,A,2=
3
i,ja2
ij
1/2
conditions (i)—(iii) clearly hold. ,A,2is also a matrix norm as we see by
application of the Cauchy—Schwartz inequality. We have
,AB,2
2=3
ijX3
kaikbkj~2
≤3
i,jX3
ka2
ik~X3
mb2
jm~
=
3
i,k|aik|2
3
j,m|bjm|2
=,A,2
2,B,2
2.
This norm has three common names: The (a) Frobenius norm, (b) Schur
norm, and (c) Hilbert—Schmidt norm. It has considerable importance inmatrix theory.
3.f
∞.D efine for A∈Mn(R)
,A,∞=s u p
i,j|aij|=m a x
i,j|aij|.
Note that if J=[11
11],,J,∞=1 . A l s o J2=2J.T h u s,J2,=2,J,=1W≤
,J,2.S o,A,∞is not a matrix norm, though it is a vector space norm. We
can make it into a matrix norm by
,A,=n,A,∞.
Note
|||AB|||=nmax
i,jeeeee3
kaikbkjeeeee
≤nmax
ijnmax
k|aik||bkj|
≤n2max
i,k|aik|max
k,j|bkj|
=|||A||| ||| B|||.
98 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
In the inequalities above we use the fundamental inequality
3
k|ckdk|≤max
k|dk|3
k|ck|
(See Exercise 4.) While these norms have some use in general matrix theory,
most of the widely applicable norms are those that are subordinate to vectornorms in the manner de fined below.
Definition 3.1.2. Let,·,be a vector norm on R
n(orCn). For A∈Mn(R)
(orMn(C)) we de fine the norm ,A,onMnby
,A,=m a x
,x,=1,Ax,. (T)
and call ,A,the norm subordinate to the vector norm. Note the use of
the same notation for both the vector and subordinate norms.
Theorem 3.1.1. The subordinate norm is a matrix norm and ,Ax,≤
,A,,x,.
Proof. We need to verify conditions (i)—(iv). Conditions (i) and (ii) are
obvious and are left to the reader . To show (iii), we have
,A+B,=m a x
,x,=1,(A+B)x,≤max
,x,=1(,Ax,+,Bx,)
≤max
,x,=1,Ax,+m a x
,x,=1,Bx,
=,A,+,B,.
Note that
,Ax,=,x,Awx
,x,W
≤,A,,x,
sinceEEEx
,x,EEE= 1. Finally, it follows that for any x∈R
n
,ABx,≤, A,,Bx,≤, A,,B,,x,
and therefore ,AB,≤, A,,B,.
Corollary 3.1.1. (i),I,=1.
3.2. CONVERGENCE AND PERTURBATION THEORY 99
(ii) If Ais invertible, then
,A−1,≥(,A,)−1.
Proof. For (i) we have
,I,=m a x
,x,=1,Ix,=m a x
,x,=1,x,=1.
To prove (ii) begin with A−1A=I. Then by the submultiplicativity and (i)
1=,I,≤, A−1,,A,
and so,A−1,≥1/,A,.
There are many results connected with matrix norms and eigenvectors that
we shall explore before long. The relation between the norm of the matrixand its inverse is important in computational linear algebra. The quantity
,A
−1,,A,thecondition number of the matrix A. When it is very large,
the solution of the linear system Ax=bby general methods such as Gaussian
elimination may produce results with considerable error. The conditionn u m b e r ,t h e r e f o r e ,t i p so ffinvestigators to this possibility. Naturally enough
t h ec o n d i t i o nn u m b e rm a yb ed i fficult to compute accurately in exactly these
circumstances. Alternative and very approximate methods are often used
as reliable substitutes for the condition number.
A special type of matrix, one for which ,Ax,=,x,for every x∈C,
is called an isometry . Such matrices which do not “stretch” any vectors
have remarkable spectral properties and play an important roll in spectraltheory.
3.2 Convergence and perturbation theory
It will often be necessary to compare o ne matrix with another matrix that
isnearby in some sense. When a matrix norm at is hand it is possible
to measure the proximity of two matrices by computing the norm of their
difference. This is just as we do for numbers. We begin with this study by
showing that if the norm of a matrix is less than one, then its di fference
with the identity is invertible. Again, this is just as with numbers; that is,if|r|<1,
1
1−ris defined. Let us assume R∈Mnand,, is some norm on
Mn. We want to show that if ,R,<1t h e n( I−R)−1exists. Toward this
end we prove the following lemma.
100 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
Lemma 3.2.1. For every R∈Mn
(I−R)(I+R+R2+···+Rn)=I−Rn+1.
Proof. This result for matrices is the direct analog of the result for numbers
(1−r)(1+r+r2+···+rn)=1−rn+1, also often written as 1+ r+r2+···+rn=
1−rn+1
1−r. We prove the result inductively. If n=1t h er e s u l tf o l l o w sf r o m
direct computation, ( I−R)(I+R)=I−R2. Assume the result holds up
ton−1. Then
(I−R)(I+R+R2+···+Rn−1+Rn)=(I−R)(I+R+···+Rn−1)
+(I−R)Rn
=(I−Rn)+Rn−Rn+1=I−Rn+1
by our inductive hypothesis. This calculation completes the induction, and
hence the proof.
Remark 3.2.1. Sometimes the proof is presented in a “quasi-inductive”
manner. That is, you will see
(I−R)(I+R+R2+···+Rn)=(I+R+R2+···+Rn)
−(R+R2+···+Rn+1)(∗)
=I−Rn+1
This is usually considered acceptable b ecause the correct induction is trans-
parent in the calculation.
Below we will show that if ,R,=λ<1, then ( I+R+R2+···)=
(I−R)−1.I t w o u l d b e incorrect to apply the obvious fact that ,Rn+1,<
λn+1→∞ to draw the conclusion from the equality ( ∗)a b o v ew i t h o u t first
establishing convergence of the series∞
0Rk. A crucial step in showing that
an in finite series is convergent is showing that its partial sums satisfy the
Cauchy criterion:∞
k=1akconverges if and only if for each ε>0,there
exists an integer Nsuch that if m, n > N, theneen
k=m+1akee<ε.(See
Appendix A.) There is just one more aspect of this problem. While it is easyto establishe the Cauchy criterion for our present situation, we still need toresolve the situation between norm convergence andpointwise convergence .
We need to conclude that if ,R,<1t h e nl i m
n→∞Rn=0,and by this
expression we mean that ( Rn)ij→0 for all 1 ≤i, j≤n.
3.2. CONVERGENCE AND PERTURBATION THEORY 101
Lemma 3.2.2. Suppose that the norm ,·,is a subordinate norm on Mn
andR∈Mn.
(i) If,R,<ε, then there is a constant Msuch that |pij|<Mε.
(ii) If limn→∞,Rn,=0,t h e n limn→∞Rn=0.
Proof. (i) If,R,<6, if follows that ,Rx,<6for each vector x,a n db y
selecting the standard vectors ejin turn, it follows that from which it follows
that,r∗j,<6,w h e r e r∗jdenotes the jthcolumn of R.By Theorem 1.6.2 all
norms are equivalent. It follows that there is a fixed constant Mindependent
of6andRsuch that |rij|<M6.
(ii) Suppose for some increasing subsequence of powers nk→∞ it happens
thateee(Rnk)ijeee≥r.Select the standard unit vector e
j.A little computation
shows that ,Rnkej,≥r,whence,Rnk,≥r, contradicting the known limit
limn→∞,Rn,= 0. The conclusion lim n→∞Rn=0f o l l o w s .
Lemma 3.2.3. Suppose that the norm ,·,is a subordinate norm on Mn.
If,R,=λ<1,t h e n I+R+R2+···+Rk+···converges.
Proof. LetPn=I+R+···+Rn. To show convergence we establish that
{Pn}is a Cauchy sequence. For n>m we have
Pn−Pm=n3
k=m+1Rk
Hence
,Pn−Pm,=EEEEEn3
k=m+1RkEEEEE
≤n3
k=m+1,Rk,
≤n3
k=m+1,R,k
=n3
m+1λk=λm+1n−m−13
j=0λj
≤λm+1(1−λ)−1→0
102 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
where in the second last step we used the inequality, which is valid for
0≤λ≤1.n−m−1
0λj≤∞
0λj<(1−λ)−1. We conclude by Lemma 3.2.2
that the individual matrix entries of the parital sums converge and thus the
series itself converges.
Note that this result is independent of the particular norm. In practice it isoften necessary to select a convenient norm to actually carry out or verifyparticular computations are valid. In the theorem below we complete theanalysis of the matrix version of the geometric series, stating that when thenorm of a matrix is less than one, the geometric series based on that matrix
converges and the inverse of the di fference with the identity exists.
Theorem 3.2.1. IfR∈M
n(F)and,R,<1for some norm, then (I−R)−1
exists and
(I−R)−1=I+R+R2+···=∞3
k=0Rk.
Proof. Apply the two previous lemmas.
T h ep e r t u r b a t i o nr e s u l ta l l u d e dt oa b o v ec a nn o wb es t a t e da n de a s i l y
proved. In words this result states that i fw eb e g i nw i t ha ni n v e r t i b l em a t r i x
and additively perturb it by a su fficiently small amount the result remains
invertible. Overall, this is the first of a series of results where what is proved
is that some property of a matrix is preserved under additive perturbations.
Corollary 3.2.1. IfA, B∈MnandAis invertible, then A+λBis invert-
ible for su fficiently small |λ|(inRorC).
Proof. A sa b o v ew ea s s u m et h a t ,·,is a norm on Mn(F). It is any easy
computation to see that
A+λB=A(I+λA−1B).
Selectλsufficiently small so that ,λA−1B,=|λ|,A−1B,<1. Then by the
theorem above, I+λA−1Bis invertible. Therefore
(A+λB)−1=(I+λA−1B)−1A−1
and the result follows.
3.3. EIGENVECTORS AND EIGENVALUES 103
Another way of stating this is to say that if A, B∈MnandAhas a nonzero
determinant, then for su fficiently small λthe matrix A+λBalso has a
nonzero determinant. This corollary can be applied directly to the identitymatrix itself being perturbed by a rank one matrix. In this case the λcan
be speci fied in terms of the two vectors comprising the matrix. (Recall
Theorem 2.3.1(8).)
Corollary 3.2.2. Letx, y∈R
nsatisfy |x, yX|=|λ|<1.T h e n I+xyTis
invertible and
(I+xyT)−1=I−xyT(1 +λ)−1.
Proof. We have that ( I+xyT)−1exists by selecting a norm ,·,consistent
with the inner product ·,·X.( F o re x a m p l e ,t a k e ,A,=s u p
,x,2=1,Ax,2,w h e r e
,·,2is the Euclidean norm.) It is easy to see that ( xyT)k=λk−1xyT.
Therefore
(I+xyT)−1=I−xyT+(xyT)2−(xyT)3+···
=I−xyT+λxyT−λ2xyT+···
=I−xyTXn3
k=0(−λ)k~
.
Thus
(I+xyT)−1=I−xyT(1 +λ)−1
and the result is proved.
In words we conclude that the perturbation of the identity by a small rank
1m a t r i xh a sa computable inverse.
3.3 Eigenvectors and Eigenvalues
Throughout this section we will consider only matrices A∈Mn(C)o rMn(R).
Furthermore, we suppress the field designation unless it is relevant.
Definition 3.3.1. IfA∈Mnandx∈CnorRn.I f t h e r e i s a c o n s t a n t
λ∈Cand a vector xW=0f o rw h i c h
Ax=λx
104 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
we callλaneigenvalue ofAandxits corresponding eigenvector .A l -
ternatively, we call xthe eigenvector pertaining to the eigenvalue λ,a n d
vice-versa.
Definition 3.3.2. ForA∈Mn,d efine
(1)σ(A)={λ|Ax=λxhas a solution for a nonzero vector x}.σ(A)i s
called the spectrum ofA.
(2)ρ(A)= s u p
λ∈σ(A)|λ|p
or equivalently max
λ∈σ(A)|λ|Q
.ρ(A)i sc a l l e dt h e
spectral radius .
Example 3.3.1. LetA=[21
12]. Then λ= 1 is an eigenvalue of Awith
eigenvector x=[−1,1]T.A l s oλ= 3 is an eigenvalue of Awith eigenvector
x=( 1,1)T. The spectrum of Aisσ(A)={1,3}and the spectral radius of
Aisρ(A)=3 .
Example 3.3.2. The 3 ×3m a t r i x B=
−30 6
−12 9 26
4−4−9
has eigenvalues:
−1,−3,1. Pertaining to the eigenvalues are the eigenvectors
3
11
↔1,
1
10
↔− 3
3
−2
2
↔− 1
The characteristic polynomial
To say that Ax=λxhas a nontrivial solution ( xW=0 )f o rs o m e λ∈Cis
the same as the assertion that ( A−λI)x= 0 has a nontrivial solution. This
means that
det(A−λI)=0
or what is more commonly written
det(λI−A)=0 .
From the original de finition (De finition 2.5.1) the determinant is sum of
products of individual matrix entries. Therefore, det( λI−A)m u s tb ea
polynomial in λ.T h i sm a k e st h ed e finition:
3.3. EIGENVECTORS AND EIGENVALUES 105
Definition 3.3.3. LetA∈Mn. The determinant
pA(λ)=d e t (λI−A)
is called the characteristic polynomial ofA. Its zeros1are the called
theeigenvalues ofA.T h e s e t σ(A) of all eigenvalues of Ais called the
spectrum ofA.
A simple consequence of the nature of the determinant of det( λI−A)i s
the following.
Proposition 3.3.1. IfA∈Mn,t h e n pA(λ)has degree exactly n.
See Appendix A for basic information on solving polynomials equations
p(λ) = 0. We may note that even though A∈Mn(C)h a s n2entries in
its de finition, its spectrum is completely determined by the ncoefficients of
pA(λ).
Procedure. The basic procedure of determining eigenvalues and eigenvec-
tors is this: (1) Solve det( λI−A) = 0 for the eigenvalues λand (2) for
any given eigenvalue λsolve the system ( A−λI)x= 0 for the pertaining
eigenvector(s). Though this procedure is not practical in general it can beeffective for small sized matrices, and for matrices with special structures.
Theorem 3.3.1. LetA∈M
n.The set of eigenvectors pertaining to any
particular eigenvalue is a subspace of the given vector space Cn.
Proof. Letλ∈σ(A).The set of eigenvectors pertaining to λis the null
space of ( A−λI).The proof is complete by application of Theorem 2.3.1
that states the null space of any matrix is a subspace of the underlying
vector space.
In light of this theorem the following de finition makes sense and is a most
important concept in the study of eigenvalues and eigenvectors.
Definition 3.3.4. LetA∈Mnand letλ∈σ(A). The null space of
(A−λI)i sc a l l e dt h e eigenspace ofApertaining to λ.
Theorem 3.3.2. LetA∈Mn.T h e n
(i) Eigenvectors pertaining to di fferent eigenvalues are linearly indepen-
dent.
1The zeros of a polynomial (or more generally a function) p(λ) are the solutions to the
equation p(λ)=0 . As o l u t i o nt o p(λ)=0i sa l s oc a l l e da root of the equation.
106 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
(ii) Suppose λ∈σ(A)with eigenvector xis different from the set of
eigenvalues {µ1,...,µ k}⊂σ(A)andVµis the span of the pertaining
eigenspaces. Then x/∈Vµ.
Proof. (i) Let µ,λ∈σ(A)w i t h µW=λpertaining eigenvectors xandy
resepectively. Suppose these vectors are linearly dependent; that is, y=cx.
Then
µy=µcx=cµx
=cAx=A(cx)
=Ay
=λy
This is a contradiction, and (i) is proved.
(ii) Suppose the contrary holds, namely that x∈Vµ.T h e n x=a1y1+···+
akymwhere {y1,···,ym}are linearly independent vectors of Vµ.E a c h o f
the vectors yiis an eigenvector, we know. Assume the pertaining eigenvalues
denoted by µji.T h a ti st os a y , Ayi=µjiyi,for each i=1,...,m . Then
λx=Ax=A(a1y1+···+amym)
=a1µj1y1+···+akµjmym
IfλW=0w eh a v e
x=a1y1+···+amym=a1µj1
λy1+···+akµjm
λym
We know by the previous part of this result that at least two of the co-
efficients aimust be nonzero. (Why?) Thus we have two di fferent rep-
resentations of the same vector by linearly independent vectors, which is
impossible. On the other hand, if λ=0t h e n a1µ1y1+···+akµkyk=0 ,
which is also impossible. Thus, (ii) is proved.
The examples below will illustrate the spectra of various matrices.
Example 3.3.1. LetA=Jab
cdo
.T h e n pA(λ)i sg i v e nb y
pA(λ)=d e t (λI−A)=d e t}λ−a−b
−cλ−d]
=(λ−a)(λ−d)−bc
=λ2−(a+d)λ+ad−bc.
3.3. EIGENVECTORS AND EIGENVALUES 107
The eigenvalues are the roots of pA(λ)=0
λ=a+d±0
(a+d)2−4(ad−bc)
2
=a+d±0
(a−d)2+4bc
2.
For this quadratic there are three possibilities:
(a) Two real rootsc
9different values
equal values
(b) Two complex roots
Here are three 2 ×2 examples that illustrate each possibility. The reader
should compute the characteristic polynomials to verify these computation.
B1=}1−1
1−1]
.Then pB1(λ)=λ2λ=0,0
B2=}0−1
10]
.Then pB2(λ)=λ2+1λ=±i
B3=}01
10]
.Then pB3(λ)=λ2−1λ=±1.
Example 3.3.2. Consider the rotation in the x, z-plane through an angle
θ
B=
cosθ0−sinθ
01 0
sinθ0c o sθ
The characteristic polynomial is given by
pB(λ)=d e t
λ
100
010001
−
cosθ0−sinθ
01 0
sinθ0c o sθ
=−1+( 1+2c o s θ)λ−(1 + 2 cos θ)λ
2+λ3
The eigenvalues are 1 ,cosθ+√
cos2θ−1,cosθ−√
cos2θ−1.Whenθ
is not equal to an even multiple of π, exactly two of the roots are complex
numbers. In fact, they are complex conjugate pairs, which can also be
written as cos θ+isinθ,cosθ−isinθ. The magnitude of each eigenvalue
108 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
is 1, which means all three eigenvalues lie on the unit circle in the complex
plane. An interesting observation is that the characteristic polynomial andhence the eigenvalues are the same regardless of which pair of axes ( x-z,
x-y,o ry-z) is selected for the rotation. Matrices of the form Bare actually
called rotations. In two dimensions the counter-clockwise rotations through
the angle θare given by
B
θ=}cosθ−sinθ
sinθcosθ]
The eigenvalues for all θis not equal to an even multiple of πare±i.( S e e
Exercise 2.)
Example 3.3.3. IfTis upper triangular with diag T=[t11,t22,... ,t nn].
ThenλI−Tis upper triangular with diag[ λ−t11,λ−t22,... ,λ−tnn]. Thus
the determinant of λI−Tgives the characteristic polynomial of Tto be
pT(λ)=n
i=1(λ−tii)
The eigenvalues of Tare the diagonal elements of T. By expanding this
product we see that
pT(λ)=λn−(Σtii)λn−1+ lower order terms.
The constant term of pT(λ)i s(−1)nn
i=1tii=(−1)ndetT.W ed e fine
trT=n3
i=1tii
and call it the trace ofT. The same statements apply to lower triangular
matrices. Moreover, the trace de finition applies to all matrices, not just to
triangular ones, and the result will be the same.
Example 3.3.4. Suppose that Ais rank 1. Then there are two vectors
w,z∈Cnfor which
A=wzT.
Tofind the spectrum of Awe consider the equation
Ax=λx
3.3. EIGENVECTORS AND EIGENVALUES 109
or
z,xXw=λx.
From this we see that x=wis an eigenvector with eigenvalue z,wX.I f
z⊥x, x is an eigenvector pertaining to the eigenvalue 0. Therefore,
σA={z,wX,0}.
T h ec h a r a c t e r i s t i cp o l y n o m i a li s
pλ(λ)=(λ−z,wX)λn−1.
Ifwandzare orthogonal then
pA(λ)=λn.
Ifwandzare not orthogonal though there are just two eigenvalues, we say
that 0 is an eigenvalue of multiplicity n−1, the order of the factor ( λ−0).
Alsoz,wXhas multiplicity 1. This is the subject of the next section.
For instance, suppose w=( 1,−1,2) and z=( 0,1,−3). Then spectrum
of the matrix A=wzTis given by σ(A)={−7,0}. The eigenvalue pertain-
ing toλ=−7i swand we may take x=( 0,3,1) and ( c,3,1), for any c∈C,
to be eigenvectors pertaining to λ=0 . N o t et h a t {w,(0,3,1),(1,3,1)}form
ab a s i sf o r R3.
To complete our discussion of characteristic polynomials we prove a re-
sult that every nthdegree polynomial with lead coe fficient one is the char-
acteristic polynomial of some matrix. You will note the similarity of thisresult and the analogous result for di fferential systems.
Theorem 3.3.3. Every polynomial of n
thdegree with lead coe fficient 1, that
is
q(λ)=λn+b1λn−1+···+bn−1λ+bn
is the characteristic polynomial of some matrix.
Proof. We consider the n×nmatrix
B=
01 0 ··· 0
00 1 0
......
−b
n−bn−1 −b1
110 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
ThenλI−Bhas the form
λI−B=
λ−10 ··· 0
0λ−10
... λ...
b
nbn−1 λ+b1
Now expand in minors across the bottom row to get
det (λI−B)= b
n(−1)n+1det
−10 ··· 0
λ−10
λ...
+bn−1(−1)n+2det
λ0··· 0
0−10
...λ...
+···
+b1(−1)n+ndet
λ−10 ···
0λ−1
... λ...
=bn(−1)n+1(−1)n−1+bn−1(−1)n+2λ(−1)n−2+···
+(λ+b1)(−1)n+nλn−1
=bn+bn−1λ+···+b1λn−1+λn
which is what we set out to prove. (The reader should check carefully the
term with bn−2to fully understand the nature of this proof.)
Multiplicity
LetA∈Mn(C). Since pA(λ) is a polynomial of degree exactly n,i tm u s t
have exactly neigenvectors (i.e. roots) λ1,λ2,... ,λncounted according to
multiplicity. Recall that the multiplicity of an eigenvalue is the number oftimes the monomial ( λ−λ
i) is repeated in the factorization of pA(λ). For
example the multiplicity of the root 2 in the polynomial ( λ−2)3(λ−5)
is 3. Suppose µ1,... ,µ kare the distinct eigenvalues with multiplicities
m1,m2,... ,m krespectively. Then the characteristic polynomial can be
3.3. EIGENVECTORS AND EIGENVALUES 111
rewritten as
pA(λ)=d e t (λI−A)=n
1(λ−λi)
=k
1(λ−µi)mi.
More precisely, the multiplicities m1,m2,... ,m kare called the algebraic
multiplicities of the respective eigenvalues. This factorization will be very
useful later.
We know that for each eigenvalue λof any multiplicity m,t h e r em u s t
be at least oneeigenvector pertaining to λ. What is desired, but not always
possible, is to findµlinearly independent eigenvectors corresponding to the
eigenvalue λof multiplicity m. This state of a ffairs makes matrix theory at
once much more challenging but also much more interesting.
Example 3.3.3. For the matrix
A=
20 0
07
41
4√
3
01
4√
35
4
the characteristic polynomial is det(λI−A)=pA(λ)=λ3−5λ2+8λ−
4,which can be factored as pA(λ)=(λ−1) (λ−2)2We see that the
eigenvalues are 1, 2, and 2. So, th e multiplicity of the eigenvalue λ=1
is 1, and the multiplicity of the eigenvalue λ=2i s2 . T h ee i g e n s p a c e
pertaining to the eigenvalue λ= 1 is generated by the vector
0
−1
3√
3
1
,
and the dimension of this eigenspace is one. The eigenspace pertaining to
λ=2i sg e n e r a t e db y
1
00
and
0
1
1
3√
3
.(That is, these two vectors form
a basis of the eigenspace.) To summarize, for the given matrix there is one
eigenvector for the eigenvalue λ= 1, and there are two linearly independent
eigenvectors for the eigenvalue λ= 2. The dimension of the eigenspace
pertaining to λ= 1 is one, and the dimension of the eigenspace pertaining
toλ=2i st w o .
Now contrast the above example where the eigenvectors span the space C3
and the next example where we have an eigenvalue of multiplicity three but
the eigenspace is of dimension one.
112 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
Example 3.3.4. Consider the matrix
A=
2−10
020
102
The characteristic polynomial is given by ( λ−2)3.Hence the eigenvalue
λ= 2 has multiplicity three. The eigenspace pertaining to λ=2i sg e n e r -
ated by the single vector [0 ,0,1]TTo see this we solve
(A−2I)x=
2−10
020102
−2
100
010001
x
=
0−10
000100
x=0
The row reduced echelon form for
0−10
000
100
is
100
010
000
.F r o m
this it is apparent that we may take x
3=t,but that x1=x2=0.Now
assign t= 1 to obtain the generating vector [0 ,0,1]T.This type of example
and its consequences seriously complexi fies the study of matrices.
Symmetric Functions
Definition 3.3.5. Letnbe a positive integer and Λ={λ1,λ2,... ,λn}be
given numbers. Suppose that kis a positive integer with 1 ≤k≤n.T h e
kthelementary symmetric function on theΛis defined by
Sk(λ1,... ,λn)=3
1≤i1<···<ik≤nk
j=1λij.
It is easy to see that
S1(λ1,... ,λn)=n3
1λi
Sn(λ1,... ,λn)=n
1λi.
3.3. EIGENVECTORS AND EIGENVALUES 113
For a given matrix A∈Mnthere are nsymmetric functions de fined with
respect to its eigenvalues Λ={λ1,λ2,... ,λn}. The symmetric functions
are sometimes called the invariants of matrices as they are invariant undersimilarity transformations that will be in Section 3.5. They also furnishdirectly the coe fficients of the characteristic polynomial. Thus specifying
thensymmetric functions of an n×nmatrix is su fficient to determine its
eigenvalues.
Theorem 3.3.4. LetA∈M
nhave symmetric functions Sk,k=1,2, ..., n .
Then
det(λI−A)=n
i=1(λ−λi)=λn+n3
k=1(−1)kSkλn−k.
Proof. The proof is a consequence of actually expanding the productn
i=1(λ−
λi). Each term in the expansion has exactly nterms multiplied together that
are combinations of the factor λand the−λI
is. For example, for the power
λn−kthe coefficient is obtained by computing the total number of products
ofk“different”−1λI
is.( T h e t e r m d i fferent is in quotes because it refers
to different indices not actual values.) Co l l e c t i n ga l lt h e s et e r m si sa c c o m -
plished by addition. Now the number of ways we can obtain products ofthese kdifferent (−1)λ
I
isis easily seen to be the number of sequences in the
set{1≤i1<···<ik≤n}. The sum of these is clearly ( −1)kSk,w i t ht h e
(−1)kfactor being the collected product of k−1’s.
Two of the symmetric functions are familiar. In the following we restate
this using familiar terms.
Theorem 3.3.5. LetA∈Mn(C).T h e n
pA(λ)=d e t (λI−A)=n3
k=0pkλn−k
where p0=1 and
(i)p1=−tr(A)=−n3
1aii
(ii) pn=−detA.
(iii) pk=(−1)kSk,for1<k<n .
114 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
Proof. Note that pA(0) =−det(A)=pn. This gives (ii). To establish (i),
we consider
det
λ·a
11−a12 ...−a1n
−a21λ−a22 −a2n
......
−an1 ... λ−ann
.
Clearly the productn
1(λ−aii) is one of the selections of products in
the calculation process. In every other product there must be no more than
n−2 diagonal terms. Hence
pA(λ)=d e t (λI−A)=n
1(λ−aii)+pn−2(λ),
where pn−2(λ) is a polynomial of degree n−2. The coe fficient ofλn−1is
−n
1aiiby the Theorem 3.3.4, and this is (i).
As a final note, observe that the characteristic polynomial is de fined by
knowing the nsymmetric functions. However, the matrix itself has n2en-
tries. Therefore, one may expect that knowing the only characteristic poly-
nomial of a matrix is insu fficient to characterize it. This is correct. Many
matrices having rather di fferent properties can have the same characteristic
polynomial.
3.4 The Hamilton-Cayley Theorem
The Hamilton-Cayley Theorem opens the doors to a finer analysis of a ma-
trix through the use of polynomials, which in turn is an important tool of
spectral analysis. The results states that any square matrix satis fies its own
charactertic polynomial, that is pA(A)=0 . T h ep r o o fi sn o td i fficult, but we
need some preliminary results about fac toring matrix-valued polynomials.
Preceding that we need to consider matrix polynomials in some detail.
Matrix polynomials
One of the very important results of matrix theory is the Hamilton-Cayleytheorem which states that a matrix satis fies its only characteristic equation.
This implies we need the notion of a matrix polynomial. It is an easy idea
3.4. THE HAMILTON-CAYLEY THEOREM 115
– just replace the coe fficients of any polynomial by matrices – but it bears
some important consequences.
Definition 3.4.1. LetA0,A1,... ,A mbe square n×nmatrices. We can
define the polynomial with matrix coe fficients
A(λ)=A0λm+A1λm−1+···+Am−1λ+Am.
Thedegree ofA(λ)i sm,p r o v i d e d A0W=0 . A(λ) is called regular if det A0W=
0. In this case we can construct an equivalent monic2polynomial.
˜A(λ)=A−1
0A(λ)=A−1
0A0λm+A−1
0A1λm−1+···+A−1
0Am
=Iλm+˜A1λm−1+···+˜Am.
The algebra of matrix polynomials mim ics the normal polynomial algebra.
Let
A(λ)=m3
0Aiλm−kB(λ)=m3
0Biλm−k.
(1)Addition:
A(λ)±B(λ)=m3
0(Ak±Bk)λm−k.
(2)Multiplication:
A(λ)B(λ)=m3
i=0λmwm3
k=0AiBm−iW
The termm
k=0AiBm−iis called the Cauchy product of the sequences.
Note that the matrices Ajalways multiply on the left of the Bk.
(3)Division: LetA(λ)a n d B(λ)b et w om a t r i xp o l y n o m i a l so fd e g r e e
m(as above) and suppose B(λ) is regular, i.e. det B0W=0 . W es a y
thatQr(λ)a n d Rr(λ)a r eright quotient andremainder ofA(λ)u p o n
division by B(λ)i f
A(λ)=Qr(λ)B(λ)+Rr(λ)( 1 )
2Recall that a monic polynomial is a polynomial where coe ffic i e n to ft h eh i g h e s tp o w e r
is one. For matrix polynomials the corresponding coe fficient is I, the identity matrix.
116 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
i ft h ed e g r e eo f Rr(λ)i slessthan that of B(λ). Similarly Qf(λ)a n d
Rf(λ) are respectively the leftquotient andremainder ofA(λ)u p o n
division by B(λ)i f
A(λ)=B(λ)Qf(λ)+Rf(λ)( 2 )
i ft h ed e g r e eo f Rf(λ)i slessthan that of B(λ).
I nt h ec a s e( 1 )w es e e
A(λ)B−1(λ)=Qr(λ)+Rr(λ)B−1(λ), (3)
which looks much likea
b=q+r
b, a way to write the quotient and remainder
of a divided by bwhen aandbare numbers. Also, the form (3) may not
properly exist for all λ.
Lemma 3.4.1. LetBi∈Mn(C),i =0,... , n with B0nonsingular.
Then the polynomial
B(λ)=B0λn+B1λn−1+···+Bn
is invertible for su fficiently large |λ|.
Proof. We factor B(λ)a s
B(λ)= B0λnD
I+B−1
0B1λ−1+···+B−1
0Bnλ−ni
=B0λnJ
I+λ−1D
B−1
0B1+···+B−1
0Bnλ1−nio
Forλ>1, the norm of the term B−1
0B1+···+B−1
0Bnλ1−nis bounded by
EEB−1
0B1+···+B−1
0Bnλ1−nEE≤EEB−1
0B1EE+EEB−1
0B2EE|λ|−1+···EEB−1
0BnEE|λ|1−n
≤p
1+|λ|−1+···+|λ|1−nQ
max
1≤i≤nEEB−1
0BiEE
≤1
1−|λ|−1max
1≤i≤nEEB−1
0BiEE
Thus the conditions of our perturbation theorem hold and for su fficiently
large |λ|,it followsEEλ−1D
B−1
0B1+···+B−1
0Bnλ1−niEE<1.Hence
I+λ−1D
B−1
0B1+···+B−1
0Bnλ1−ni
is invertible and therefore B(λ) is also invertible.
3.4. THE HAMILTON-CAYLEY THEOREM 117
Theorem 3.4.1. LetA(λ)andB(λ)be matrix polynomials in Mn(C)or
(Mn(R)). Then both left and right division of A(λ)byB(λ)is possible and
the respective quotients an d remainders are unique.
Proof. We proceed by induction on deg B, and clearly if deg B=0t h er e s u l t
holds. If deg B=p> degA(λ)=mthen the result follows simply. For,
takeQr(λ)=0a n d Rr(λ)=A(λ). The conditions of right division are met.
Now suppose that p≤m. It is easy to see that
A(λ)=A0λm+A1λm−1+···+Am−1λ+Am
=A0B−1
0λm−pB(λ)−p3
j=1A0B−1
0Bjλm−j+m3
j=1Ajλm−j
=Q1(λ)B(λ)+A1(λ)
where deg A1(λ)<degA(λ). Our inductive hypothesis assumed the division
was possible for matrix polynomials A(λ)o fd e g r e e <p.T h e r e f o r e , A1(λ)=
Q2(λ)B() + R(λ), where the degree of B(λ)<p . Finally, with Q(λ)=
Q(λ)+Q2(λ), there results A(λ)=Q(λ)B() +R(λ).
To establish uniqueness we assume two right divisors and quotients have
been determined. Thus
A=Qr1(λ)B(λ)+Rr1(λ)
A=Qr2(λ)B(λ)+Rr2(λ)
Subtract to get
0=( Qr1(λ)−Qr2(λ))B(λ)+Rr1(λ)−Rr2(λ).
IfQr1(λ)−Qr2(λ)W= 0, we know the degree of ( Qr1(λ)−Qr2(λ))B(λ)i s
greater than the degree of R1(λ)−R2(λ). This contradiction implies that
Qr1(λ)−Qr2(λ)=0 ,w h i c hi nt u r ni m p l i e st h a t Rr1(λ)−Rr2(λ)=0 . H e n c e
the decomposition is unique.
Hamilton-Cayley Theorem
Let
B(λ)=B0λm+B1λm−1+···+Bm−1λ+Bm
with B0W= 0. We can also, write B(λ)=m
i=0λm−iBi. Both versions are
the same. However, when A∈Mn(F), there are two possible evaluations of
B(A).
118 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
Definition 3.4.2. LetB(λ),A∈Mn(C)( o r Mn(R)) and B(λ)=B0λm+
B1λm−1+···+Bm−1λ+Bm.D efine
B(A)=B0Am+B1Am−1+···+Bm “right value”
B(A)=AmB0+Am−1B1+···+Bm “left value”
The generalized B´ ezout theorem gives the remainder of B(λ)d i v i d e db y
λI−A.I nf a c t ,w eh a v e
Theorem 3.4.2 (Generalized B´ ezout Theorem). The right division of
B(λ)byλI−Ahas remainder
Rr(λ)=B(A)
Similarly, the left division of B(λ)by(λI−A)has remainder
Rf(λ)=B(A).
Proof. In the case deg B(λ)=1 ,w eh a v e
B0λ+B1=B0(λI−A)+B0A+B1.
The remainder Rr(λ)=B0A+B1=B(A). Assume the result holds for all
polynomials up to degree p−1. We have
B(λ)=B0λp+B1λp−1+···+Bp
=B0λp−1(λI−A)+B0Aλp−1+B1λp−1+···
=B0λp−1(λI−A)+B1(λ)
where deg B1(λ)≤p−1. By induction
B(λ)=B0λp−1(λI−A)+Qr(λ)(λI−A)+B1(A)
B1(A)=( B0A+B1)Ap−1+B2Ap−2+···+Bp−1A+Bp=B(A). This
proves the result.
Corollary 3.4.1. (λI−A)divides B(λ)if and only if B(A)=0 (resp
B(A)=0 ).
Combining the B´ ezout result and the adjoint formulation of the matrix
inverse, we can establish the important Hamilton-Cayley theorem.
Theorem 3.4.3 (Hamilton-Cayley). LetA∈Mn(C)(orMn(R))w i t h
characteristic polynomial pA(λ).T h e n pA(A)=0 .
3.4. THE HAMILTON-CAYLEY THEOREM 119
Proof. Recall the adjoint formulation of the inverse as ˆC=1
det·CCT
ij=C−1.
Now let B=adj(A−λI). Then
B(λI−A)=d e t (λI−A)I
(λI−A)B=d e t (λI−A)I.
These equations show that p(λ)I=d e t (λI−A)Iis divisible on the right
andthe left by ( λI−A) without remainder. It follows from the generalized
B´ezout theorem that this is possible only if pA(A)=0 .
LetA∈Mn(C).Now that we know any Asatisfies its characteristic
polynomial, we might also ask if there are polynomials of lower degree that it
also satis fies. In particular, we will study the so-called minimal polynomial
that a matrix satis fies. The nature of this polynomial will shed considerable
light on the fundamental structure of A. For example, both matrices below
have the same characteristic polynomial P(λ)=(λ−2)3.
A=
200
020002
andB=
210
021002
Henceλ= 2 is an eigenvalue of multiplicity three. However, Asatifies the
much simpler first degree polynomial ( λ−2) while there is no polynomial
of degree less that three that Bsatisfies. By this time you recognize
that Ahas three linearly independent eigenvectors, while Bhas only one
eigenvector. We will take this subject up in a later chapter.
Biographies
Arthur Cayley (1821-1895), one of the most proli fic mathematicians of his
era and of all time, born in Richmond, Surrey, and studied mathematics atCambridge. For four years he taught at Cambridge having won a Fellowshipand, during this period, he published 28 papers in the Cambridge Mathe-matical Journal. A Cambridge fellowship had a limited tenure so Cayley
had to find a profession. He chose law and was admitted to the bar in 1849.
He spent 14 years as a lawyer, but Cayley always considered it as a meansto make money so that he could pursue mathematics. During these 14 yearsas a lawyer Cayley published about 250 mathematical papers! Part of thattime he worked in collaboration with James Joseph Sylvester
3(1814-
3In 1841 he went to the United States to become professor at the University of Virginia,
but just four years later resigned and returned to England. He took to teaching private
120 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
1897), another lawyer. Together, but not in collaboration, they founded the
algebraic theory of invariants 1843.
In 1863 Cayley was appointed Sadleirian professor of Pure Mathemat-
ics at Cambridge. This involved a very large decrease in income. However
Cayley was very happy to devote himself entirely to mathematics. He pub-lished over 900 papers and notes covering nearly every aspect of modernmathematics.
The most important of his work is in developing the algebra of matrices,
work in non-Euclidean geometry and n-dimensional geometry. Importantly,
he also clari fied many of the theorems of algebraic geometry that had previ-
ously been only hinted at, and he was among the first to realize how many
different areas of mathematics were linked together by group theory.
As early as 1849 Cayley wrote a paper l inking his ideas on permutations
with Cauchy’s. In 1854 Cayley wrote two papers which are remarkable forthe insight they have of abstract groups. At that time the only knowngroups were groups of permutations and even this was a radically new area,
yet Cayley de fines an abstract group and gives a table to display the group
multiplication.
Cayley developed the theory of algebraic invariance, and his develop-
ment of n-dimensional geometry has been applied in physics to the study
of the space-time continuum. His work on matrices served as a founda-
tion for quantum mechanics, which was developed by Werner Heisenberg in1925. Cayley also suggested that Euclidean and non-Euclidean geometryare special types of geometry. He united projective geometry and metricalgeometry which is dependent on sizes of angles and lengths of lines.
Heinrich Weber (1842-1913) was born and educated in Heidelberg, where
he became professor 1869. He then taught at a number of institutions inGermany and Switzerland. His main work was in algebra and number theory.
He is best known for his outstanding text Lehrbuch der Algebra published
in 1895.
Weber worked hard to connect the various theories even fundamental
concepts such as a fie l da n dag r o u p ,w h i c hw e r es e e na st o o l sa n dn o t
properly developed as theories in his Die partiellen Di fferentialgleichungen
der mathematischen Physik 1900-01, which was essentially a reworking of a
book of the same title based on lectures given by Bernhard Riemann and
pupils and had among them Florence Nightingale. By 1850 he became a barrister, and by
1855 returned to an academic life at the Royal Military Academy in Woolwich, London. Hereturned to the US again in 1877 to become prof essor at the new Johns Hopkins University,
but returned to England once again in 1877. Sylvester coined the term ‘matrix’ in 1850.
3.4. THE HAMILTON-CAYLEY THEOREM 121
written by Karl Hattendor ff.
Etienne B´ ezout (1730-1783) was a mathematician who represents a char-
acteristic aspect of the subject at that time. One of the many successful
textbook projects produced in the 18thcentury was B´ ezout’s Cours de math-
ematique ,as i xv o l u m ew o r kt h a t first appeared in 1764-1769, which was
almost immediately issued in a new e dition of 1770-1772, and which boasted
many versions in French and other languages. (The first American textbook
in analytic geometry, incidentally, was derived in 1826 from B´ ezout’s Cours .)
It was through such compilations, rather than through the original works
of the authors themselves, that the mathematical advances of Euler and
d’Alembert became widely known. B´ ezout’s name is familiar today in con-
nection with the use of determinants in algebraic elimination. In a memoirof the Paris Academy for 1764, and more extensively in a treatise of 1779entitled Theorie generale des equations algebriques ,B ´ezout gave arti ficial
rules, similar to Cramer’s, for solving nsimultaneous linear equations in n
unknowns. He is best known for an extension of these to a system of equa-
tions in one or more unknowns in which it is required to find the condition
on the coe fficients necessary for the equations to have a common solution.
To take a very simple case, one might ask for the condition that the equa-tions a
1x+b1y+c1=0 , a2x+b2y+c2=0 , a3x+bay+c3=0h a v ea
common solution. The necessary condition is that the eliminant a special
case of the “Bezoutiant,” should be 0.
Somewhat more complicated eliminants arise when conditions are sought
for two polynomial equations of unequal degree to have a common solution.B´ezout also was the first one to give a satisfactory proof of the theorem,
known to Maclaurin and Cramer, that two algebraic curves of degrees mand
nrespectively intersect in general in m·npoints; hence, this is often called
B´ezout’s theorem. Euler also had contributed to the theory of elimination,
but less extensively than did B´ ezout.
Taken from A History of Mathematics by Carl Boyer
122 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
William Rowen Hamilton
Born Aug. 3/4, 1805, Dublin, Ire. and died Sept. 2, 1865, Dublin
Irish mathematician and astronomer who developed the theory of quater-
nions, a landmark in the development of algebra, and discovered the phe-nomenon of conical refraction. His uni fication of dynamics and optics, more-
over, has had lasting in fluence on mathematical physics, even though the
full signi ficance of his work was not fully appreciated until after the rise of
quantum mechanics.
Like his English contemporaries Thomas Babington Macaulay and John
Stuart Mill, Hamilton showed unusual intellect as a child. Before the age
of three his parents sent him to live with his father’s brother, James, a
learned clergyman and schoolmaster at an Anglican school at Trim, a small
town near Dublin, where he remained until 1823, when he entered Trinity
College, Dublin. Within a few months of his arrival at his uncle’s he couldread English easily and was advanced in arithmetic; at five he could translate
Latin, Greek, and Hebrew and recite Homer, Milton and Dryden. Beforehis 12th birthday he had compiled a grammar of Syriac, and by the age of
14 he had su fficient mastery of the Persian language to compose a welcome
to the Persian ambassador on his visit to Dublin.
Hamilton became interested in mathematics after a meeting in 1820
with Zerah Colburn, an American who cou ld calculate mentally with aston-
ishing speed. Having read the El´ements d’alg` ebre of Alexis—Claude Clairaut
and Isaac Newton’s Principia ,Hamilton h a di m m e r s e dh i m s e l fi nt h e five
volumes of Pierre—Simon Laplace’s Trait´ ed em ´ ecanique c´ eleste (1798-1827;
Celestial Mechanics ) by the time he was 16. His detection of a flaw in
Laplace’s reasoning brought him to the attention of John Brinkley, pro-fessor of astronomy at Trinity College. When Hamilton was 17, he sent
Brinkley, then president of the Royal Irish Academy, an original memoir
about geometrical optics. Brinkley, in forwarding the memoir to the Acad-
emy, is said to have remarked: “This young man, I do not say will be , but
is,t h e first mathematician of his age.”
In 1823 Hamilton entered Trinity College, from which he obtained the
highest honours in both classics and m athematics. Meanwhile, he continued
his research in optics and in April 1827 submitted this “theory of Systems
of Rays” to the Academy. The paper tr ansformed geometrical optics into
a new mathematical science by establishing one uniform method for thesolution of all problems in that field.Hamilton started from the principle,
originated by the 17th-century French mathematician Pierre de Fermat, thatlight takes the shortest possible time in going from one point to another,
whether the path is straight or is bent by refraction. Hamilton ’s key idea
3.5. SIMILARITY 123
was to consider the time (or a related quantity called the “action”) as a
function of the end points between which the light passes and to show thatthis quantity varied when the coordinates of the end points varied, accordingto a law that he called the law of varying action. He showed that the entiretheory of systems of rays is reducible to the study of this characteristic
function.
Shortly after - Hamilton submitted his paper and while still an under-
graduate, Trinity College elected h im to the post of Andrews professor of
astronomy and royal astronomer of Ireland, to succeed Brinkley, who hadbeen made bishop. Thus an undergraduate (not quite 22 years old) becameex officio an examiner of graduates who were candidates for the Bishop Law
Prize in mathematics. The electors’ object was to provide Hamilton with
a research post free from heavy teachi ng duties. Accordingly, in October
1827Hamilton took up residence next to Dunsink Observatory, 5 miles (8
km) from Dublin, where he lived for the rest of his life. He proved to be anunsuccessful observer, but large audiences were attracted by the distinctlyliterary flavour of his lectures on astronomy. Throughout his life Hamilton
was attracted to literature and considered the poet William Wordsworth
among his fiends, although Wordsworth advised him to write mathematics
rather than poetry.
With eigenvalues we are able to begin spectral analysis. That part is
the derivation of the various normal forms for matrices. We begin with a
relatively weak form of the Jordan form, which is coming up.
First of all, as you have seen diagona l matrices furnish the easiest form
for matrix analysis. Also, linear systems are very simple to solve for diagonalmatrices. The next simplest class of matrices are the triangular matrices.We begin with the following result based on the idea of similarity.
3.5 Similarity
Definition 3.5.1. Am a t r i x B∈Mnis said to be similar toA∈Mnif
there exists a nonsingular matrix S∈Mnsuch that
B=S−1AS.
The transformation A→S1ASis called a similarity transformation .S o m e -
times we write A∼B. Note that similarity is an equivalence relation:
(i) A∼A reflexivity
(ii) B∼A⇒A∼B symmetry
(iii) B∼AandA∼C⇒B∼Ctransitivity
124 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
Theorem 3.5.1. Similar matrices have the same characteristic polynomial.
Proof. We suppose A, B,a n d S∈Mnwith Sinvertible and B=S−1AS.
Then
λI−B=λI−S−1AS
=λS−1IS−S−1AS
=S−1(λI−A)S.
Hence
det(λI−B)=d e t ( S−1(λI−A)S)
=d e t ( S−1)d e t (λI−A)d e tS
=d e t (λI−A)
because 1 = det I=d e t ( SS−1)=d e t SdetS−1.
A simple consequence of Theorem 3.5.1 and Corollary 2.3.1 follows.
Corollary 3.5.1. IfAandBare in Mnand if AandBare similar, then
they have the same eigenvalues counted ac cording to multiplicity, and there-
f o r et h es a m er a n k .
We remark that even though [00
00]a n d[0100] have the same eigenvalues,
0 and 0, they are not similar. Hence the converse is not true.
Another immediate corollary of Theorem 3.5.1 can be expressed in terms
of the invariance of the trace and determinant of similarity transformations ,
that is functions on Mn(F)d efined for any invertible matrix SbyTS(A)=
S−1AS. Note that such transformations are linear mappings from Mn(F)→
Mn(F). We shall see how important are those properties of matrices that
are invariant (i.e. do not change) under similarity transformations.
Corollary 3.5.2. IfA, B∈Mnare similar, then they have the same trace
and same determinant. That is, tr(A) = tr(B) anddetA=d e t B.
Theorem 3.5.2. IfA∈Mn,t h e n Ais similar to a triangular matrix.
Proof. The following sketch shows the first two steps of the proof. A formal
induction can be applied to achieve the full result.
Letλ1be an eigenvalue of Awith eigenvector u1. Select a basis , say
S1,o fCnand arrange these vectors into columns of the matrix P1,w i t h u1
3.5. SIMILARITY 125
in the first column. De fineB1=P−1
1AP1.T h e n B1is the representation of
Ain the new basis and so
B1=
λα 1...αn−1
0
... A2
0
where A 2is (n−1)×(n−1)
because Au1=λ1u1. Remembering that B1is the representation of Ain
the basis S1,t h e n[ u1]S1=e1and hence B1e1=λe1=λ1[1,0,... , 0]T.B y
similarity, the characteristic polynomial of B1i st h es a m ea st h a to f A, but
more importantly (using expansion by minors down the first column)
det(λI−B1)=(λ−λ1)d e t (λIn−1−A2).
Now select an eigenvalue λ2ofA2and pertaining eigenvector v2∈Cn;s o
A2v2=λ2v2.N o t et h a t λ2is also an eigenvalue of A. With this vector we
define
ˆu1=
1
0
...
0
ˆu2=
0
v2
.
Select a basis of Cn−1,w i t h v2selected first and create the matrix P2with
this basis as columns
P2=
10 ... 0
0
...P
2
0
.
It is an easy matter to see that P
2is invertible and
B2=P−1
2B1P2
=
λ
1∗... ... ∗
0λ2∗...∗
00 A3
......
00
.
Of course, B
2∼B1, and by the transitivity of similarity B2∼A.C o n -
tinue this process, deriving ultimately the triangular matrix Bn∼Bn−1and
hence Bn∼A. This completes the proof.
126 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
Definition 3.5.2. We say that A∈Mnisdiagonalizable ifAis similar to
a diagonal matrix.
Suppose P∈Mnis nonsingular and D∈Mnis a diagonal matrix. Let
A=PDP−1. Suppose the columns of Pare the vectors v1,v2,... ,v n.T h e n
Avj=PDP−1vj
=PDe j=λjPej=λjvj
whereλjis the jthdiagonal element of D,a n d ejis the jthstandard vector.
Similarly, if u1,... ,u narenlinearly independent vectors of Apertain-
ing to eigenvalues λ1,λ2,... ,λn,t h e nw i t h Q, the matrix with columns
u1...u n,w eh a v e
AQ=QD
where
D=
λ
1
λ2s
...
s λn
.
Therefore
Q−1AQ=D.
We have thus established the
Theorem 3.5.3. LetA∈Mn.T h e n Ais diagonalizable if and only if A
hasnlinearly independent eigenvectors.
As a practical measure, the conditions of this theorem are remarkably
difficult to verify.
Example 3.5.1. The matrix A=[01
00]i snotdiagonalizable.
Solution. F i r s tw en o t et h a tt h es p e c t r u m σ(A)={0}.S o l v i n g
Ax=0
we see that x=c(1,0)Tis the only solution. That is to say, there are not
twolinearly independent eigenvectors. hence the result.
3.5. SIMILARITY 127
Corollary 3.5.3. IfA∈Mnis diagonalizable and Bis similar to A,t h e n
Bis diagonalizable.
Proof. The proof follows directly from transitivity of similarity. However,
more directly, suppose that B∼AandSis the invertible matrix such that
B=SAS−1Then BS=SA.I fuis an eigenvector of Awith eigenvalue λ,
then SAu =λSuand therefore BSu =λSu. This is valid for each eigenvec-
tor. We see that if u1,...,u nare the eigenvalues of A,t h e n Su1,...,Su n
are the eigenvalues of B. The similarity matrix converts the eigenvectors of
one matrix to the eigenvectors of the transformed matrix.
In light of these remarks, we see that similarity transforms preserve com-pletely the dimension of eigenspaces. It is just as signi ficant to note that
if a similarity transformation diagonalizes a given matrix A,t h es i m i l a r i t y
matrix must consist of the eigenvectors of A.
Corollary 3.5.4. LetA∈M
nbe nonzero and nilpotent. Then Ais not
diagonalizable.
Proof. Suppose A∼T,w h e r e Tis triangular. Since Ais nilpotent (i.e.( Am=
0),Tis nilpotent, as well. Therefore the diagonal entries of Tare zero. The
spectrum of Tand hence Ais zero, it’s null space has dimension n.T h e r e -
fore, by Theorem 2.3.3(f), its rank is zero. Therefore T= 0, and hence
A= 0. The result is proved.
Alternatively, if the nilpotent matrix Ais similar to a diagonal matrix with
zero diagonal entries, then Ai ss i m i l a rt ot h ez e r om a t r i x . T h u s Ais itself
the zero matrix. From the obvious fact that the power of a similarity trans-
formations is the similarity transformation of the power of a matrix, that is(SAS
−1)m=SAmS−1(see Exercise 26), we have
Corollary 3.5.5. Suppose A∈Mn(C)is diagonalizable. (i) Then Amis
diagonalizable for every positive integer m. (ii) If p(·)is any polynomial,
then p(A)is diagonalizable.
Corollary 3.5.6. If all the eigenvalues of A∈Mn(C)are distinct, then A
is diagonalizable.
Proposition 3.5.1. LetA∈Mn(C)andε>0. Then for any matrix norm
,·,there is a matrix Bwith norm ,B,<εfor which A+Bis diagonalizable
and for each λ∈σ(A)there is a µ∈σ(B)for which |λ−µ|<ε.
128 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
Proof. First triangularize AtoTby a similarity transformation, T=SAS−1.
Now add to Tany diagonal matrix Dso that the resulting triangular ma-
trixT+Dhas all distinct values. Moreover this can be accomplished by a
diagonal matrix of arbitrarily small norm for any matrix norm. Then wehave
S
−1(T+D)S=S−1TS+S−1DS
=A+B
where B:=S−1DS. Now by the submultiplicative property of matrix norms
,B,≤EES−1DSEE≤EES−1EE,D,,S,. Thus to obtain the estimate ,B,<ε,
it is sufficient to take ,D,≤6
,S−1,,S,, which is possible as established
above. Since the spectrum σ(A+B)o fA+Bhasndistinct values, it
follows that A+Bis diagonalizable.
There are many, many results on diagonalizable and non-diagonalizable ma-
trices. Here is an interesting class o f nondiagonalizable matrices we will
encounter later.
Proposition 3.5.2. pro Every matrix of the form
A=λI+N
where Nis nilpotent and is not zero is not diagonalizable.
An important subclass has the form:
A=
λ1
λ1s
λ1
s...1
λ
“Jordan block”.
Eigenvectors
Once an eigenvalue is determined, it is a relatively simple matter to find the
pertaining eigenvectors. Just solve the homogeneous system ( λI−A)x=0 .
Eigenvectors have a more complex structure and their study merits ourattention.
Facts
(1)σ(A)=σ(A
T) including mutliplicities.
3.5. SIMILARITY 129
(2)σ(A)=σ(A∗), including multiplicities.
Proof. det(λI−A)=d e t ( ( λI−A)T)=d e t (λI−AT). Similarly for A∗.
Definition 3.5.3. The linear space spanned by all the eigenvectors per-
taining to an eigenvalue λis called the eigenspace corresponding to the
eigenvalue λ.
For any A∈Mnany subspace V∈Cnfor which
AV⊂V
is called an invariant subspace ofA. The determination of invariant sub-
spaces for linear transformations has been an important question for decades.
Example 3.5.2. For any upper triangular matrix Tthe spaces Vj=S(e1,e2,... ,e j),
j=1,... ,n are invariant.
Example 3.5.3. Given A∈Mn, with eigenvalue λ. The eigenspace cor-
responding to λand all of its subspaces are invariant subspaces. Corollary
to this, the null space N(A)= {x|Ax=0}is an invariant subspace of A
corresponding to the eigenvalue λ=0 .
Definition 3.5.4. LetA∈Mnwith eigenvalue λ. The dimension of the
eigenspace corresponding to λis called the geometric multiplicity ofλ.T h e
multiplicity of λas a zero of the characteristic polynomial pA(λ)i sc a l l e d
thealgebraic multiplicity .
Theorem 3.5.4. IfA∈Mnandλ0is an eigenvalue of Awith geometric
and algebraic multiplicities mgandma, respectively. Then
mg≤ma.
Proof. Letu1...u mgbe linearly independent eigenvectors pertaining to λ0.
LetSbe a matrix consisting of a basis of Cnwith u1...u mgselected among
them and placed in the firstmgcolumns. Then
B=S−1AS=
λ0I...∗............
0...B
where I=Img. It is easy to see that
pB(λ)=pA(λ)=(λ−λ0)mgp0B(λ),
whence the algebraic multiplicity ma≥mg.
130 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
Example 3.5.4. ForA=[01
00], the algebraic multiplicity of ais 2, while
the geometric multiplicity of 0 is 1. Hence, the equality mg=mais not
possible, in general.
Theorem 3.5.5. IfA∈Mnand for each eigenvalue µ∈σ(A),mg(µ)=
ma(µ),t h e n Ais diagonalizable. The converse is also true.
Proof. Extend the argument given in the theorem just above. Alternatively,
we can see that Amust have nlinearly independent eigenvalues, which
follows from the
Lemma 3.5.1. IfA∈MnandµW=λare eigenvalues with eigenspaces Eµ
andEλrespectively. Then EµandEλare linearly independent.
Proof. Suppose u∈Eµc a nb ee x p r e s s e da s
u=k3
cjvj
where, of course uW=0a n d vjW=0,j=1,... ,k where v1...v k∈Eλ.T h e n
Au=AΣcjvj
⇒ µu=λΣcjvj
⇒ u=λ
µΣcjvj,ifµW=0.
Sinceλ/µW= 1, we have a contradiction. If µ=0 ,t h e n Σcjvj=0a n dt h i s
implies u= 0. In either case we have a contradiction, the result is therefore
proved.
While both AandAThave the same eigenvalues, counted even with
multiplicities, the eigenspaces can be very much di fferent. Consider the
example where
A=}11
02]
and AT=}10
12]
.
The eigenvalue λ= 1 has eigenvector u=[1
0]f o rAandJ1
−1o
forAT.
A new concept of left eigenvector yield some interesting results.
Definition 3.5.5. We sayλis aleft eigenvector ofA∈Mnif there is a
vector y∈Cnfor which
y∗A=λy∗
3.6. EQUIVALENT NORMS AND CONVERGENT MATRICES 131
or inRnifyTA=λyT. Taking adjoints the two sides of this equality become
(y∗A)∗=A∗y
(λy∗)∗=¯λy.
Putting these lines together A∗y=¯λyand hence ¯λis an eigenvalue of A∗.
Here’s the big result.
Theorem 3.5.6. LetA∈Mn(C)with eigenvalues λand eigenvector u.
Letu, v be left eigenvectors pertaining to µ,λ, respectively. If µW=λ,v∗u=
u, vX=0.
Proof. We have
v∗u=1
µv∗Aµ=λ
µv∗u.
Assuming µW= 0 we have a contradiction. If µ=0 ,a r g u ea s v∗u=1
λv∗Au=
µ
λc∗u= 0, which was to be proved. The result is proved.
Note that left eigenvectors of Aare (right) eigenvectors of A∗(ATin the
real case).
3.6 Equivalent norms and convergent matrices
Equivalent norms
S of a rw eh a v ed e fined a number of di fferent norms. Just what “di fferent”
m e a n si ss u b j e c tt od i fferent interpretations. For example, we might agree
that different means that the two norms have a di fferent value for some
matrix A. On the other hand if we have two matrix norms ,·,aand,·,b,
we might be prepared to say that if for two positive constants mM > 0
0<m≤,A,a
,A,b≤M for all A∈Mn(F)
then these norms are not really so di fferent but rather are equivalent be-
cause “small” in one norm implies “small” in the other and the same for“large.” Indeed, when two norms satisfy the condition above we will callthem equivalent. The remarkable fact is that all subordinate norms on afinite dimensional space are equivalent. This is a direct consequence of the
similar result for vector norms. We state it below but leave the details of
the proof to the reader.
132 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
Theorem 3.6.1. Any two matrix norms, ,·,aand,·,b,o n Mn(C)are
equivalent in the sense that ther e exist two positive constants mM > 0
0<m≤,A,a
,A,b≤M for all A∈Mn(F)
We have already proved one convergence result about the invertibility of
I−Awhen,A,<1. A deeper version of this result can be proved based
on a new matrix norm. This result has important consequences in general
matrix theory and particularly in com putational matrix theory. It is most
important in applications, where having ρ(A)<1 can yield the same results
as having ,A,<1.
Lemma 3.6.1. Let,·,be a matrix norm that is subordinate to the vector
norm (on Cn),·,. Then for each A∈Mn
ρ(A)≤,A,.
Note: We use the same notation for both vector and matrix norms.
Proof. Letλ∈σ(A). Then with corresponding eigenvector xλwe have
Axλ=λxλ.N o r m a l i z i n g xλso that,xλ,=1 ,w eh a v e ,A,=m a x
,x,=1,Ax,≥
,Axλ,=|λ|,xλ,=|λ|.H e n c e
,A,≥ max
λ∈σ(A)|λ|=ρ(A).
Theorem 3.6.2. For each A∈Mnandε>0there is a vector norm ,·,
onCnfor which the subordinate matrix norm ,·,satisfies
ρ(A)≤,A,≤ρ(A)+ε.
Proof. The proof follows in a series of simple steps, the first of which is
interesting in its own right. First we know that Ais similar to a triangular
matrix B–from a previous theorem
B=SAS−1=Λ+U
whereΛis the diagonal part of BandUis strictly upper triangular part of
B,w i t hz e r o s filled in elsewhere.
3.6. EQUIVALENT NORMS AND CONVERGENT MATRICES 133
Note that the diagonal of Bisthe spectrum of Band hence that of A.
This is important. Now select a δ>0 and form the diagonal matrix
D=
1
δ−1s
...
s δ1−n
.
A brief computation reveals that
C=DBD−1=DSAS−1D−1
=D(Λ+U)D−1
=Λ+DUD−1=Λ+V
where V=DUD−1and more speci fically vij=δj−iuijforj>i and of
courseΛis a diagonal matrix with the spectrum of Afor the diagonal ele-
ments. In this way we see that for δsmall enough we have arranged that
Ais similar to the diagonal matrix of its spectral elements plus a small
triangular perturbation. We now de fine the new vector norm on Cnby
,x,A=(DS)∗(DS)x, xX1/2.
We recognise this to be a norm from a previous result. Now compute the
matrix norm
,A,A=m a x
,x,A=1,Ax,A.
Thus
,Ax,2
A=DSAx,DSAx X
=CDSx,CDSx X
≤,C,2
2,DSx,2
=,C∗C,2,x,2
A.
From C=Λ+V, it follows that
C∗C=Λ∗Λ+Λ∗V+V∗Λ+V∗V.
Because the last three terms have a δpmultiplier for various positive powers
p,w ec a nm a k et h et e r m s Λ∗V+V∗Λ+V∗Vsmall in anynorm by taking
δsufficiently small. Also the diagonal elements of Λ∗Λhave the form
|λi|2λi∈S(A).
134 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
Talkingδsufficiently small so that each of ,Λ∗V,2,,V∗Λ,2and,V∗V,2is
less than 6/3. With ,x,2
A=1a sr e q u i r e d ,w eh a v e
,Ax,A≤(ρ(A)+ε),x,A
and we know already that ,A,A≥ρ(A).
An important corollary places a lower bound on vector norms is given
below.
Corollary 3.6.1. For any A∈Mn
ρ(A)= i n f
,,(m a x
,x,=1,Ax,)
where the in fimum is taken over all vector norms.
Convergent matrices
Definition 3.6.1. We say that a matrix A∈Mn(F)f o r F=RorCis
convergent if
lim
m→∞Am=0 ←the zero matrix.
That is, for each 1 ≤i, j≤nthelimm→∞(Am)ij=0 . T h i si ss o m e t i m e s
called pointwise convergence.
Theorem 3.6.3. The following three statements are equivalent:
(a)Ais convergent.
(b) lim
n→∞,Am,=0,f o rs o m em a t r i xn o r m .
(c)ρ(A)<1.
Proof. These results follow substantially from previously proven results.
However, for completeness, assume (a) holds. Then
lim
n→∞max
ij|(An)ij|=0
or what is the same we have
lim
n→∞,An,∞=0
3.6. EQUIVALENT NORMS AND CONVERGENT MATRICES 135
which is (b). Now suppose that (b) holds. If ρ(A)≥1t h e r ei sav e c t o r
xfor which ,Anx,=,ρ(A),n,x,.T h e r e f o r e ,An,≥1, which contradicts
(b). Thus (c) holds. Next if (c) holds we can apply the above theorem toestablish that there is a norm for which ,A,<1. Hence (b) follows. Finally,
we know that
,A
m,∞≤M,Am, (T)
whence lim
n→∞Am=0 . ( T h a ti s( a ) ≡(b)). In sum we have shown that (a)
≡(b) and (b) ≡(c).
To establish ( T)w en e e dt h ef o l l o w i n g .
Theorem 3.6.4. If,, and,·,Iare two vector norms on Cnthen there
are constants mandMso that
m,x,I≤,x,<M,x,I
for all x∈Cn.
This result carries over to matrix norms subordinate to vector norms by
simple inheritance. To prove this result we use compactness ideas. Supposethere is a sequence of vectors x
jfor which
1≤,xj,<1
j,xj,I
and for which the components of xjare bounded in modulus. By compact-
ness there is a convergent subsequence of the xj,f o rw h i c hl i m xj=xW=0 .
But,x,I<∞. Hence,x,= 0, a contradiction.
Here is the new and improved version of our previous result.
Theorem 3.6.5. Ifρ(A)<1then (I−A)−1exists and
(I−A)−1=I+A+A2+···.
Proof. Select a matrix norm ,·,for which ,A,<1. Apply previous calcu-
lations. We know that every matrix can be triangularized and “almost” di-
agonalized we may ask if we can eliminate altogether the o ff-diagonal terms.
The answer is unfortunately no. But we can resolve the diagonal question
completely.
136 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
3.7 Exercises
1. Suppose that U∈Mnis orthogonal and let ,·,be a matrix norm.
Show that ,U,≥1.
2. In two dimensions the counter-clockwise rotations through the angle
θare given by
Bθ=}cosθ−sinθ
sinθcosθ]
Find the eigenvalues and eigenvectors for all θ. (Note the two special
cases,θis not equal to an even multiple of πandθ=0 .
3. Given two matrices A, B∈Mn(C). De fine the commutant ofAand
Bby [A, B]−AB−BA. Prove that tr[ A, B]=0 .
4. Given two finite sequences {ck}n
k=1and{dk}n
k=1.Prove that
k|ckdk|≤
max k|dk|
k|ck|.
5. Verify that a matrix norm which is subordinate to a vector norm sat-
isfies norm conditions (i) and (ii).
6. Let A∈Mn(C).Show that the matrix norm subordinate to the
vector norm ,·,∞is given by
,A,∞=m a x
i,ri(A),1
where as usual ri(A)d e n o t e st h e ithrow of the matrix Aand,·,1is
thef1norm.
7. The Hilbert matrix, Hnof order nis defined by
hij=1
i+j−11≤i, j≤n
8. Show that ,Hn,1<lnn.
9. Show that ,Hn,∞=1 .
10. Show that ,Hn,2∼n1
2.
11. Show that the spectral radius of Hnis bounded by 1.
12. Show that for each ε>0 there exists an integer Nsuch that if n>N
there is a vector x∈Rnwith,x,2=1 such that ,Hnx,2<ε.
3.7. EXERCISES 137
13. Same as the previous question except that you need to show that
N=!
1
ε1/2
+1w i l lw o r k .
14. Show that the matrix A=
110
031
1−12
is not diagonalizable.Let
A∈Mn(C).
15. We know that the characteristic polynomial of a matrix A∈M12(C)
is equal to pA(λ)=(λ−1)12−1.Show that Ais not similar to the
identity matrix.
16. We know that the spectrum of A∈M3(C)i sσ(A)={1,1,−2}and
the corresponding eigenvectors are {[1,2,1]T,[2,1,−1]T,[1,1,2]T}.
(i) Is it possible to determine A? Why or why not? If so, prove it.
If not show two di fferent matrices with the given spectrum and
eigenvectors.
(ii) Is Adiagonalizable?
17. Prove that if A∈Mn(C) is diagonalizable then for each λ∈σ(A),
the algebraic and geometric multiplicities are equal. That is, ma(λ)=
mg(λ).
18. Show that B−1(λ)e x i s t sf o r |λ|sufficiently large.
19. We say that a matrix A∈Mn(R) is row stochastic if all its entries
are non negative and the sum of the entries of each row is one.
(i) Prove that 1 ∈σ(A).
(ii) Prove that ρ(A)=1 .
20. We say that A, B∈Mn(C)a r e quasi -commutative if the spectrum of
AB−BAis just zero, i.e. σ(AB−BA)={0}.
21. Prove that if ABis nilpotent then so is BA.
22. A matrix A∈Mn(C) has a square root if there is a matrix B∈Mn(C)
such that B2=A.Show that if Ais diagonalizable then it has a square
root.
23. Prove that every matrix that commutes with every diagonal matrix is
itself diagonal.
138 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
24. Suppose that if A∈Mn(C) is diagonalizable and for each λ∈σ(A),
|λ|<1.Prove directly from similarity ideas that lim n→∞An=
0.(That is, do not apply the more general theorem from the lecture
notes.)
25. We say that Aisright quasi- invertible if there exists a matrix B∈
Mn(C) such that AB∼D,w h e r e Dis a diagonal matrix with diagonal
entries nonzero. Similarly, we say that Aisleft quasi- invertible if there
exists a matrix B∈Mn(C) such that BA∼D,w h e r e Dis a diagonal
matrix with diagonal entries nonzero.
(i) Show that if Aisright quasi- invertible then it is invertible.
(ii) Show that if Aisright quasi- invertible then it is left quasi- invertible.
(iii) Prove that quasi- invertibility is not an equivalence relation. (Hint.
How do we usually show that an assertion is false?)
26. Suppose that A, S∈Mn,w i t h Sinvertible, and mis a positive integer.
Show that ( SAS−1)m=SAmS−1.
27. Consider the rotation
10 0
0c o sθ−sinθ
0s i nθcosθ
Show that the eigenvalues are the same as for the rotation
cosθ0−sinθ
01 0
sinθ0c o sθ
See Example 2 of section 3.2.
28. Let A∈Mn(C). De fineeA=∞
n=0An
n!. (i) Prove that exists. (ii)
Suppose that A, B∈Mn(C). Is eA+B=eAeB?I f n o t , w h e n i s i t
true?
29. Suppose that A, B∈Mn(C). We say A`Bif [A, B]=0 ,w h e r e
[·,·] is the commutant. Prove or disprove that “ `” is an equivalence
relation. Answer the same question in the case of quasi-commutivity.(See Exercise 20
30. Let u, v∈C
n.Find ( I+uv∗)m.
3.8. APPENDIX A 139
31. De fine the function fonMm×n(C)b y f(A)=r a n k A, and suppose
that,·,is any matrix norm. Prove the following. (1) If for any
matrix Afor which f(A)=nshow that fis continuous in ,·,.T h a t
is, for every ε>0t h e n f(B)=nfor every matrix Bwith,BA,<ε.
is continuous in an ε-neighborhood of A.T h a t i s , f o r a n y Bwith
,A−B,<ε,t h e n f(B)=f(A). (This deviates slightly from the
usual de finition of continuity because the function fis integer valued.)
(2) If f(A)<n,t h e n fis not continuous in every ε-neighborhood of
A.
3.8 Appendix A
It is desired to solve the equation p(λ)=0f o rc o e fficients in p0,...,p n∈C
orR. The basic theorem on this subject is called the Fundamental Theo-
rem of Algebra (FTA), whose importance is manifest by its hundreds ofapplications. Concommitant with the FTA is the notion of reducible andirreducible factors.
Theorem 3.8.1 (Theorem Fundamental Theorem of Algebra). Given
any polynomial p(λ)=p
0λn+p1λn−1+···+pnwith coefficients in p0,...,p n∈
C. There is at least one solution λ∈Cto the equation p(λ)=0 .
Though proving this result would take us too far a field of our subject,
we remarks that the simplest proof of this result no doubt comes as a di-rect application of Liouville’s Theorem, a result in complex analysis. As a
corollary to the FTA, we have that there are exactly nsolutions to p(λ)=0
when counted with multiplicitiy. Proved originally by Gauss, the proof ofthis theorem eluded mathematicians for many years. Let us assume thatp
0= 1 to make the factorization simpler to write. Thus
p(λ)=k
i=1(λ−λi)mi(4)
.
As we know, in the case that the coe fficients p0,...,p nare real, there
may be complex solutions. As is easy to prove, complex solutions must
come in complex conjugate pairs. For if λ=r+isis a solution
p(r+is)=(r+is)n+p1(r+is)n−1+···pn−1(r+is)+pn=0
Because the coe fficients are real, the real and imaginary parts of the powers
(r+is)jremain respectively real or imaginary upon multiplication by pj.
140 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
Thus
p(r+is)=R e ( r+is)n+p1Re (r+is)n−1+···pn−1Re (r+is)+pn
+ip
Im (r+is)n+p1Im (r+is)n−1+···pn−1Im (r+is)+pnQ
=0
Hence the real and imaginary parts are each zero. Since Re ( r−is)j=
Re (r+is)jand Im ( r−is)j=−Im (r+is)jit follows that p(r−is)=0 .
In the case that the coe fficients are real it may be of interest to note
what statement of factorization analogous to (4) above. To this end we
need the de finition
Definition 3.8.1. The real polynomial x2+bx+cis called irreducible if
it has no real zeros.
In general, any polynomial that cannot be factored over the underlying
field is called irreducible. Irreducible polynomials have played a very impor-
tant role in fields such as abstract algebra and number theory. Indeed, they
were central in early attempts to prove Fermat’s last theorem and also so
but more indirectly to the acutal proof. Our attention here is restricted to
thefieldsCandR.
Theorem 3.8.2. Letp(λ)=λn+p1λn−1+···+pnwith coefficients in
p0,...,p n∈R.T h e n p(λ)can be factored as a product of linear factors
pertaining to real zeros of p(λ)=0 and irreducible quadratic factors.
Proof. The proof is an easy application of the FTA and the observation
above. If λk=r+isis a zero of p(λ), then so also is ¯λk=r−is.T h e r e f o r e
the product ( λ−λk)D
λ−¯λki
=λ2−2λr+r2+s2is an irreducible quadratic.
Combine such terms with the linear factors generated from the real zeros,and use (4).
Remark 3.8.1. It is worth noting that there are nohigher order irreducible
factors with real coe fficients. Even though there are certainly higher order
polynomials with complex roots, they can always factored as products of
either real linear or real quadratic f actors. Proving this without the FTA
may prove challenging.
3.9. APPENDIX B 141
3.9 Appendix B
3.9.1 In finite Series
Definition 3.9.1. Aninfiniteseries, denoted by
a0+a1+a2+···
is a sequence {un},w h e r e unis defined by
un=a0+a1+···+an
If the sequence {un}converges to some limit A,w es a yt h a tt h ei n finite
series converges to Aand use the notation
a0+a1+a2+···=A
We also say that the sum of the in finite series is A.I f {un}diverges, the
infinite series a0+a1+a2+···is said to be divergent.
The sequence {un}is called the sequence ofpartial sums, and the se-
quence {an}is called the sequence ofterms of the in finite series a0+a1+
a2+···.
Let us now return to the formula
1+r+r2+···+rn=1−rn+1
1−r
where rW= 1. Since the sequence {rn+1}converges (and, in fact, to 0) if and
only if−1<r< 1( o r |r|<1),rbeing different from 1, the in finite series
1+r+r2+···
converges to (1 −r)−1if and only if |r|<1. This in finite series is called the
geometric series with ratio r.
Definition 3.9.2. Geometric Series with Ratio r
1+r+r2+···=1
1−r,if|r|<1
and diverges if |r|≥1.
Multiplying both sides by ryields the following.
142 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
Definition 3.9.3. If|r|<1, then
r+r2+r3+···=r
1−r
Example 3.9.1. Investigate the convergence of each of the following in fi-
nite series. If it converges, determine its limit (sum).a. 1−
1
2+1
4−1
8+··· b. 1 +2
3+4
9+8
27+··· c.3
4+9
16+27
64+···
Solution A careful study reveals that each in finite series is a geometric
series. In fact, the ratios are, respectively, (a) −1
2,( b )2
3,a n d( c )3
4.S i n c e
they are all of absolute value less than 1, they all converge. In fact, we have:a. 1−
1
2+1
4−1
8+···=1
1−(−1
2)=2
3
b. 1 +2
3+4
9+8
27+···=1
1−2
3=3
1=3
c.3
4+9
16+27
64+···=3
4
1−3
4=3
4·4
1=3
Since an in finite series is de fined as a sequence of partial sums, the prop-
erties of sequences carry over to in finite series. The following two properties
are especially important.
1. Uniqueness of Limit Ifa0+a1+a2+···converges, then
it converges to a unique limit.
This property explains the notation
a0+a1+a2+···=A
where Ais the limit of the in finite series. Because of this notation, we often
say that the in finite series sums to A,o rt h e sum oftheinfinitely many
terms a0,a1,a2,... isA. We remark, however, that the order of the terms
cannot be changed arbitrarily in general.
Definition 3.9.4. 2. Sum of In finite Series If
a0+a1+a2+···=Aandb0+b1+b2+···=B,t h e n
(a0+b0)+(a1+b1)+(a2+b2)+··· =(a0+a1+a2+···)
+(b0+b1+b2+···)
=A+B
This property follows from the observation that
(a0+b0)+(a1+b1)+···+(an+bn)=( a0+a1+···+an)
+(b0+b1+···+bn)
converges to A+B. Another property is
3.9. APPENDIX B 143
Definition 3.9.5. 3. Constant Multiple of In finite Series Ifa0+a1+
a2+···=Aandcis a constant, then
ca0+ca1+ca2+···=cA
Example 3.9.2. Determine the sums of the following convergent in finite
series.a.
5
3·2+13
9·4+35
27·8+··· b.1
3·2+1
9·2+1
27·2+·
Solution a. By the Sum Property, the sum of the first in finite series is
w1
3+1
2W
+w1
9+1
4W
+w1
27+1
8W
+···
=w1
3+1
9+1
27+···W
+w1
2+1
4+1
8+···W
=1
3
1−1
3+1
2
1−1
2=1
2+1=3
2
b. By the Constant-Multiple Property, the sum of the second in finite series
is
1
2w1
3+1
9+1
27+···W
=1
2·1
3
1−1
3=1
2·1
2=1
4
Another type of in finite series can be illustrated by using Taylor poly-
nomial extrapolation as follows.
Example 3.9.3. In extrapolating the value of ln 2, the nth-degree Taylor
polynomial Pn(x)o fl n ( 1+ x)a tx= 0 was evaluated at x=1i nE x a m p l e
4 of Section 12.3. Interpret ln 2 as the sum of an in finite series whose nth
partial sum is Pn(1).
Solution We have seen in Example 3 of Section 12.3 that
Pn(x)=x−1
2x2+1
3x3−···+(−1)n−11
nxn
so that Pn(1) = 1−1
2+1
3−···+(−1)n−11
n,a n di ti st h e nth partial sum of
the in finite series
1−1
2+1
3−1
4+···
Since {Pn(1)}converges to ln 2 (see Example 4in Section 12.3), we have
ln 2 = 1−1
2+1
3−1
4−···
144 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
It should be noted that the terms of the in finite series in the preceding
example alternate in sign. In general, an in finite series of this type always
converges.
Alternating Series Test Leta0≥a1≥···≥0. If anap-
proaches 0, then the in finite series
a0−a1+a2−a3+···
converges to some limit A,w h e r e0 <A<a 0.
The condition that the term antends to zero is essential in the above
Alternating Series Test. In fact, if the terms do not tend to zero, the corre-sponding in finite series, alternating or not, must diverge.
Divergence Test Ifa
ndoes not approach zero as n→+∞,
then the in finite series
a0+a1+a2+···
diverges.
Example 3.9.4. Determine the convergence or divergence of the following
infinite series.
a. 1−1
3+1
5−1
7+··· b.−2
3+4
5−6
7+···
Solution a. This series is an alternating series with
1>1
3>1
5>···>0
and with the terms approaching 0. In fact, the general ( nth) term is
(−1)n1
2n+1
Hence, by the Alternating Series Test, the in finite series 1 −1
3+1
5−1
7+···
converges.
b. Let a1=−2
3,a2=4
5,a3=−6
7,... . It is clear that andoes not approach
0, and it follows from the Divergence Test that the in finite series is divergent.
[The nth term is ( −1)n2n/(2n+ 1).] Note, however, that the terms do
alternate in sign.
3.9. APPENDIX B 145
Another useful tool for testing convergence or divergence of an in finite
series is the Ratio Test, to be discussed below. Recall that the geometric
series
1+r+r2+···
where the nth term is an=rn,c o n v e r g e si f |r|<1 and diverges otherwise.
Note also that the ratio of two consecutive terms is
an+1
an=rn+1
rn=r
Hence, the geometric series converges if and only if this ratio is of absolute
value less than 1. In general, it is possible to draw a similar conclusion if thesequence of the absolute values of the ratios of consecutive terms converges.
Ratio Test Suppose that the sequence
{|a
n+1|/|an|}
converges to some limit R. Then, the in finite series a0+a1+
a2+···converges if R< 1 and diverges if R> 1.
Example 3.9.5. In each of the following, determine all values of rfor which
the in finite series converges.
a. 1 + r+r2
2!+r3
3!+··· b. 1 + 2 r+3r3+4r3+···
Solution a. The nth term is an=rn/n!, so that
an+1
an=rn+1
(n+1 ) !·n!
rn=r
n+1
For each value of r,w eh a v e
|an+1|
|an|=1
n+1|r|→0<1
Hence, by the Ratio Test, the in finite series converges for all values of r.
b. Let an=(n+1 )rn. Then,
|an+1|
|an|=(n+2 )|r|n+1
(n+1 )|r|n=n+2
n+1|r|→|r|
Hence, the Ratio Test says that the in finite series converges if |r|<1a n d
diverges if |r|>1. For |r|=1 ,w eh a v e |an|=n+1, which does not converge
to 0, so that the corresponding in finite series must be divergent as a result
of applying the Divergence Test.
Sometimes it is possible to compare the partial sums of an in finite series
with certain integrals, as illustrated in the following example.
146 CHAPTER 3. EIGENVALUES AND EIGENVECTORS
Example 3.9.6. Show that 1+1
2+···+1/nis larger than the de fine integral$n+1
1dx/x , and conclude that the in finite series 1 +1
2+1
3+···diverges.
Solution Consider the function f(x)=1 /x. Then, the de finite integral
8n+1
1dx
x=l n ( n+1 )−ln 1 = ln( n+1 )
is the area under the graph of the function y=f(x) between x=1a n d
x=n+ 1. On the other hand, the sum
1+1
2+···+1
n=[f(1) + f(2) + ···+f(n)]∆x
where∆x= 1, is the sum of the areas of nrectangles with base ∆x=1a n d
heights f(1),... ,f (n), consecutively, as shown in Figure 12.5. Since the
union of these rectangles covers the region bounded by the curve y=f(x),
thex-axis, and the vertical lines x=1a n d x=n+1 ,w eh a v e :
1+1
2+···+1
n>8n+1
1dx
x=l n ( n+1 )
Now recall that ln( n+ 1) approaches ∞asnapproaches ∞. Hence, the
sequence of partial sums of the in finite series
1+1
2+1
3+···
must be divergent.
The above in finite series is called the harmonic series. It diverges “to
infinity” in the sense that its sequence of partial sums becomes arbitrarily
large for all large values of n.I f a n i n finite series diverges to in finity, we also
say that it sums toinfinity and use the notation “= ∞” accordingly. We
have
TheHarmonic Series is defined by
1+1
2+1
3+···=∞
Chapter 4
Unitary Matrices
4.1 Basics
This chapter considers a very important class of matrices that are quite use-
ful in proving a number of structure theorems about all matrices. Calledunitary matrices, they comprise a class of matrices that have the remarkable
properties that as transformations they preserve length, and preserve the an-
gle between vectors. This is of course true for the identity transformation.Therefore it is helpful to regard unitary matrices as “generalized identities,”though we will see that they form quite a large class. An important exam-ple of these matrices, the rotations, have already been considered. In this
chapter, the underlying field is usually C, the underlying vector space is C
n,
and almost without exception the underlying norm is k·k 2.W e b e g i n b y
recalling a few important facts.
Recall that a set of vectors x1,...,x k∈Cnis called orthogonal if
x∗
jxm=hxm,xji=0f o r1 ≤j6=m≤k.T h es e ti s orthonormal if
x∗
jxm=δmj=½1j=m
0j6=m.
An orthogonal set of vectors can be made orthonormal by scaling:
xj−→1
(x∗
jxj)1/2xj.
Theorem 4.1.1. Every set of orthonormal vecto rs is linearly independent.
Proof. The proof is routine, using a common technique. Suppose S=
{xj}k
j=1is orthonormal and linearly dependent. Then, without loss of gen-
157
158 CHAPTER 4. UNITARY MATRICES
erality (by relabeling if needed), we can assume
uk=k−1X
j=1cjuj.
Compute u∗
kukas
1=u∗
kuk=u∗
k
k−1X
j=1cjuj
=k−1X
j=1cju∗
kuj=0.
This contradiction proves the result.
Corollary 4.1.1. IfS={u1...u k}⊂Cnis orthonormal then k≤n.
Corollary 4.1.2. Every k-dimensional subspace of Cnhas an orthonormal
basis.
Proof. Apply the Gram—Schmidt process to any basis to orthonormalize
it.
Definition 4.1.1. Am a t r i x U∈Mnis said to be unitary ifU∗U=I.[ I f
U∈Mn(R)a n d UTU=I,t h e n Uis called real orthogonal .]
Note: A linear transformation T:Cn→Cnis called an isometry if
kTxk=kxkfor all x∈Cn.
Proposition 4.1.1. Suppose that U∈Mnis unitary. (i) Then the columns
ofUform an orthonormal basis of Cn,o rRn,i fUis real. (ii) The spectrum
σ(u)⊂{z||z|=1}. (iii) |detU|=1.
Proof. The proof of (i) is a consequence of the de finition. To prove (ii), first
denote the columns of Ubyui,i=1,...,n .I fλis an eigenvalue of Uwith
pertaining eigenvector x,t h e n kUxk=kPxiuik=(P|xi|2)1/2=kxk=
|λ|kxk.H e n c e |λ|= 1. Finally, (iii) follows directly because det U=Qλi.
Thus |detU|=Q|λi|=1 .
This important result is just one of many equivalent results about unitary
matrices. In the result below, a number of equivalences are established.
Theorem 4.1.2. LetU∈Mn. The following are equivalent.
4.1. BASICS 159
(a)Uis unitary.
(b)Uis nonsingular and U∗=U−1.
(c)UU∗=I.
(d)U∗is unitary.
(e) The columns of Uform an orthonormal set.
(f) The rows of Uform an orthonormal set.
(g)Uis an isometry.
(h)Ucarries every set of orthonormal vectors to a set of orthonormal
vectors.
Proof. (a)⇒(b) follows from the de finition of unitary and the fact that
the inverse is unique.
(b)⇒(c) follows from the fact that a left inverse is also a right inverse.
(a)⇒(d)UU∗=(U∗)∗U∗=I.
(d)⇒(e) (e) ≡(b)⇒ u∗
jukδjk,w h e r e u1...u nare the columns of U.
Similarly (b) ⇒(e).
(d)≡(f) same reasoning.
(e)⇒(g) We know the columns of Uare orthonormal. Denoting the columns
byu1...u n,w eh a v e
Ux=nX
1xiui
where x=(x1,... ,x n)T. It is an easy matter to see that
kUxk2=nX
1|xi|2=kxk2.
(g)⇒(e). Consider x=ej.T h e n Ux=uj.H e n c e 1 = kejk=kUejk=
kujk.T h e c o l u m n s o f Uhave norm one. Now let x=αei+βejbe chosen
160 CHAPTER 4. UNITARY MATRICES
such that kxk=kαei+βejk=q
|α|2+|β|2=1 . T h e n
1= kUxk2
=kU(αei+βej)k2
=hU(αei+βej),U(αei+βej)i
=|α|2hUei,Ue ii+|β|2hUej,Ue ji+α¯βhUei,Ue ji+¯αβhUej,Ue ii
=|α|2hui,uii+|β|2huj,uji+α¯βhui,uji+¯αβhuj,uii
=|α|2+|β|2+2<¡
α¯βhui,uji¢
=1 + 2 <¡
α¯βhui,uji¢
Thus <¡
α¯βhui,uji¢
=0.Now suppose hui,uji=s+it. Selecting α=
β=1√
2we obtain that <hui,uji= 0, and selecting α=iβ=1√
2we obtain
that =hui,uji=0 . T h u s hui,uji= 0. Since the coordinates iandjare
arbitrary, it follows that the columns of Uare orthogonal.
(g)⇒(h) Suppose {v1,...,v n}is orthogonal. For any two of them kU(vj+
vk)k2=kvj+vkk2. Hence hUvj,U v ki=0 .
(h)⇒(e) The orthormal set of the standard unit vectors ej,j=1,...,n
is carried to the columns of U.T h a t i s Uej=uj,t h e jthcolumn of U.
Therefore the columns of Uare orthonormal.
Corollary 4.1.3. IfU∈Mn(C)is unitary, then the transformation de fined
byUpreserves angles.
Proof. We have for any vectors x, y∈Cnthat the angle θis completely
determined from the inner product via cos θ=hx,yi
kxkkyk.S i n c e Uis unitary
(and thus an isometry) it follows that
hUx,Uy i=hU∗Ux,y i=hx, yi
This proves the result.
Example 4.1.1. LetT(θ)=£cosθ−sinθ
sinθcosθ¤
whereθis any real. Then T(θ)i s
realorthogonal.
Proposition 4.1.1. IfU∈M2(R)is real orthogonal, then Uhas the form
T(θ)for some θor the form
U=·10
0−1¸
T(θ)=·cosθsinθ
sinθ−cosθ¸
Finally, we can easily establish the di agonalizability of unitary matrices.
4.1. BASICS 161
Theorem 4.1.3. IfU∈Mnis unitary, then it is diagonalizable.
Proof. To prove this we need to revisit the proof of Theorem 3.5.2. As
before, select the first vector to be a normalized eigenvector u1pertaining
toλ1.Now choose the remaining vectors to be orthonormal to u1.T h i s
makes the matrix P1with all these vectors as columns a unitary matrix.
Therefore B1=P−1UPis also unitary. However it has the form
B1=
λα
1...αn−1
0
... A2
0
where A
2is (n−1)×(n−1)
Since it is unitary, it must have orthogonal columns by Theorem 4.1.2. It
follows then that α1=α2=···=αn=0 a n d
B1=
λ0... 0
0
... A2
0
At this point one may apply and inductive hypothesis to conclude that A2
is similar to a diagonal matrix. Thus by the manner in which the full
similarity was constructed, we see that Amust also be similar to a diagonal
matrix.
Corollary 4.1.1. LetU∈Mnbe unitary. Then
(i) Then Uhas a set of northogonal eigenvectors.
(ii) Let {λ1,...,λn}and{v1,...,v n}denote respectively the eigenvalues
and their pertaining orthonormal eigenvectors of U.Then Uhas the
representation as the sum of rank one matrices given by
U=nX
j=1λjvjvT
j
This representation is often called the spectral respresentation or spectral
decomposition of U.
162 CHAPTER 4. UNITARY MATRICES
4.1.1 Groups of matrices
Invertible and unitary matrices have a fundamental structure that makes
possible a great many general statements about their nature and the waythey act upon vectors other vectors matrices. A group is a set with a math-ematical operation, product, that obeys some minimal set of properties so
as to resemble the nonzero numbers under multiplication.
Definition 4.1.2. Agroup Gi sas e tw i t hab i n a r yo p e r a t i o n G×G→G
which assigns to every pair a, bof elements of Ga unique element abinG.
The operation, called the product ,s a t i s fies four properties:
1. Closure. If a, b∈G,t h e n ab∈G.
2. Associativity. If a, b, c∈G,t h e n a(bc)=(ab)c.
3. Identity. There exists an element e∈Gsuch that ae=ea=afor
every a∈G.eis called the identity ofG.
4. Inverse. For each a∈G, there exists an element ˆ a∈Gsuch that
aˆa=ˆaa=e.ˆais called the inverse of aand is often denoted by a
−1.
A subset of Gthat is itself a group under the same product is called a
subgroup ofG.
It may be interesting to note that removal of any of the properties 2-4 leads
to other categories of sets that have interest, and in fact applications, in
their own right. Moreover, many groups have additional properties such ascommutativity, i.e. ab=bafor all a, b∈G.B e l o w a r e a f e w e x a m p l e s o f
matrix groups. Note matrix addition is not involved in these de finitions.
Example 4.1.2. As usual M
nis the vector space of n×nmatrices. The
product in these examples is the usual matrix product.
•The group GL(n, F) is the group of invertible n×nmatrices. This is
the so-called general linear group. The subset of Mnof invertible
lower (resp. upper) triangular matrices is a subgroup of GL(n, F).
•Theunitary group Unof unitary matrices in Mn(C).
•Theorthogonal group Onorthogonal matrices in Mn(R). The sub-
group of Ondenoted by SOnconsists of orthogonal matrices with
determinant 1.
4.1. BASICS 163
Because element inverses are required, it is obvious that the only subsets
of invertible matrices in Mnwill be groups. Clearly, GL(n, F) is a group
because the properties follow from those matrix of multiplication. We havealready established that invertible lo wer triangular matrices have lower tri-
angular inverses. Therefore, they form a subgroup of GL(n, F). We consider
the unitary and orthogonal groups below.
Proposition 4.1.2. For any integer n=1,2,... the set of unitary matrices
U
n(resp. real orthogonal) forms a group. Similarly Onis a group, with
subgroup SOn.
Proof. The result follows if we can show that unitary matrices are closed
under multiplication. Let UandVbe unitary. Then
(UV)∗(UV)=V∗U∗UV
=V∗V=I
For orthogonal matrices the proof is essentially identical. That SOnis a
group follows from the determinant equality det( AB)=d e t AdetB.T h e r e -
fore it is a subgroup of On.
4.1.2 Permutation matrices
Another example of matrix groups comes from the idea of permutations of
integers.
Definition 4.1.3. The matrix P∈Mn(C)i sc a l l e da permutation matrix
if each row and each column has exactly one 1, the rest of the entries being
zero.
Example 4.1.3. Let
P=
100
001010
Q=
0001
01001000
0010
PandQare permutation matrices.
Another way to view a permutation matrix is with the game of chess.
On an n×nchess board place nrooks in positions where none of them
attack one another. Viewing the board as an n×nmatrix with ones where
the rooks are and zeros elsewhere, this matrix will be a permutation matrix.
164 CHAPTER 4. UNITARY MATRICES
Of course there are n! such placements, exactly the number of permutations
of the integers {1,2,...,n }.
Permutation matrices are closely linked with permutations as discussed
in Chapter 2.5. Let σbe a permutation of the integers {1,2,...,n }.Define
the matrix Aby
aij=½1i f j=σ(i)
0i fo t h e r w i s e
Then Ais a permutation matrix. We could also use the Dirac notation
to express the same matrix, that is to say aij=δiσ(j).The product of
permutation matrices is again a permutation matrix. This is apparent bystraight multiplication. Let PandQbe two n×npermutation matrices
with pertaining permutations σ
PandσQof the integers {1,2,...,n }.Then
theithrow of PiseσP(i)and the ithrow of QiseσQ(i). (Recall the eiare
the usual standard vectors.) Now the ithrow of the product PQcan be
computed by
nX
j=1pijeσQ(j)=nX
j=1δiσ(j)eσQ(j)=eσQ(σP(i))
Thus the multiplication ithrow of PQis a standard vector. Since the σP(i)
ranges over the integers {1,2,...,n },i ti st r u ea l s ot h a t σQ(σP(i)) does
likewise. Therefore the “product” σQ(σP(i)) is also a permutation. We
conclude that the standard vectors constitute the rows of PQ.T h u s p e r m u -
tation matrices are orthogonal under multiplication. Moreover the inverse ofevery permutation is permutation matrix, with the inverse describe throughthe inverse of the pertaining permutation of {1,2,...,n }. Therefore, we
have the following result.
Proposition 4.1.3. Permutation matrices are orthogonal. Permutation
matrices form a subgroup of O
n.
4.1.3 Unitary equivalence
Definition 4.1.4. Am a t r i x B∈Mnis said to be unitarily equivalent toA
if there is a unitary matrix U∈Mnsuch that
B=U∗AU.
( I nt h er e a lc a s ew es a y Bis orthogonally equivalent to A.)
4.1. BASICS 165
Theorem 4.1.4. IfBandAare unitarily equivalent. Then
nX
i,j=1|bij|2=nX
i,j=1|aij|2.
Proof. We havenP
i,j=1|aij|2=t rA∗A,a n d
Σ|bij|2=t rB∗B=t r( U∗AU)∗U∗AU=t rU∗A∗AU
=t rA∗A,
since the trace is invariant under similarity transformations.
Alternatively, we have BU∗=U∗A.S i n c e U(and U∗) are isometries
we have each column of U∗Ahas the same norm as the norm of the same
column of A. The same holds for the rows of BU∗, whence the result.
Example 4.1.4. B=£31
−20¤
andA=[11
02] are similar but not unitarily
equivalent. AandBare similar because (1) they have the same spectrum,
σ(A)=σ(B)={1,2}and (2) they have two linearly independent eigenvec-
tors. They are not unitarily equivalent because the conditions of the above
theorem are not met.
Remark 4.1.1. Unitary equivalence is a finer classi fication than similarity.
Indeed, consider the two sets S(A)={B|Bis similar to A}andU(A)=
{B|Bis unitarily equivalent to A},t h e n
U(A)⊂S(A).
(Can you show that U(A)$S(A) for some large class of A∈Mn?)
4.1.4 Householder transformations
An important class of unitary transformations are elementary re flections.
These can be realized as transformations that re flect one vector to its neg-
a t i v ea n dl e a v ei n v a r i a n tt h eo r t h o c o m p l e m e n to fv e c t o r s .
Definition 4.1.5. Simple Householder transformation.) Suppose w∈Cn,
kwk=1 . D e fine the Householder transformation Hwby
Hw=I−2ww∗
166 CHAPTER 4. UNITARY MATRICES
Computing
HwH∗
w=(I−2ww∗)(I−2ww∗)∗
=I−2ww∗−2ww∗+4ww∗ww∗
=I−4ww∗+4hw,wiww∗=I
it follows that Hwis unitary.
Example 4.1.5. Consider the vector w=·cosθ
sinθ¸
and the Householder
transformation
Hw=I−2wwT=·1−2c o s2θ−2c o sθsinθ
−2c o sθsinθ1−2s i n2θ¸
=·−cos 2θ−sin 2θ
−sin 2θcos 2θ¸
The transformation properties for the standard vectors are
Hwe1=·−cos 2θ
−sin 2θ¸
and
Hwe2=·cos 2θ
−sin 2θ¸
This is shown below. It is evident that this unitary transformation is not a
rotation. Though, it can be imagined as a “rotation with a one dimensionalreflection.”
ww
wH e
H eee2
1
θ2θ2θ
Householder transformation
Example 4.1.6. Letθbe real. For any n≥2a n d1≤i, j≤nwith i6=j
4.1. BASICS 167
define
Un(θ,i ,j)i
j
10 ......0
.........
01......
........... c o s θ............ −sinθ...........
...1
...
1...
........... s i n θ............ c o s θ.................1...
......0...0
0.........1
ij
Then U
n(θ;i, j) is a rotation and is unitary.
Proposition 4.1.4 (Limit theorems). (i) Show that the unitary matri-
ces are closed with respect to any norm. That is, if the sequence {Un}⊂
Mn(C)are all unitary and the limn→∞Un=U in the k·k2norm, then U
is also unitary.
(ii) The unitary matrices are closed under pointwise convergence. That
is, if the sequence {Un}⊂Mn(C)are all unitary and limn→∞Un=Ufor
each ( ij)entry, then Uis also unitary.
Householder transformations can als ob eu s e dt ot r i a n g u l a r i z eam a t r i x .
The procedure successively removes the lower triangular portion of a matrixcolumn by column in a way similar to Gaussian elimination. The elemen-
tary row operations are replaced by elementary re flectors. The result is a
triangular matrix T=H
vn···Hv2Hv1A.Since these re flectors are unitary,
the factorization yields an e ffective method for solving linear systems. The
actual process is rather straightforward. Construct the vector vsuch that
¡
I−2vvT¢
A=¡
I−2vvT¢
a
11a12···a1n
a21a22···a2n
............
an1an2···ann
=
ˆa
11ˆa12··· ˆa1n
0ˆa22··· ˆa2n
............
0ˆan2··· ˆann
168 CHAPTER 4. UNITARY MATRICES
This is accomplished as follows. This means we wish to find a vector v
such that
¡
I−2vvT¢
[a11,a21,···,an1]T=[ ˆa11,0,···,0]T
For notational convenience and to emphasize the construction is vector
based, relabel the column vector [ a11,a21,···,an1]Tas [x1,x2,...,x n]T
vj=xj
2hv,xi
forj=2,3,...,n .D e fine
v1=x1+α
2hv,xi
Then
hv,xi=hx, xi
2hv,xi+αx1
2hv,xi=kxk2
2hv,xi+αx1
2hv,xi
4hv,xi2=2 kxk2+2αx1
where k·kdenotes the Euclidean norm. Also, for Hvto be unitary we need
1= hv,vi=1
4hv,xi2³
kxk2+2αx1+α2´
4hv,xi2=kxk2+2αx1+α2
Equating the two expressions for 4 hv,xi2gives
2kxk2+2αx1=kxk2+2αx1+α2
α2=kxk2
α=±kxk
We now have that
4hv,xi2=2 kxk2±2kxkx1
hv,xi=µ1
2³
kxk2±kxkx1´¶1
2
This makes v1=x1±kxk
(1
2(kxk2±kxkx1))1
2.With the construction of Hvto “elim-
inate” the first column of A, we relabel the vector vasv1(with the small
4.1. BASICS 169
possibility of notational confusion) and move to describe Hv2using a trans-
formation of the same kind with a vector of the type v=( 0,x2,x3,...,x n).
Such a selection will not a ffect the structure of the first column. Continue
this until the matrix is triangularized. The upshot is that every matrix canbe factored as A=UT, where Uis unitary and Tis upper triangular.
The main result for this section is the factorization theorem. As it turns out,
every Unitary matrix can be written as a product of elementary re flectors.
The proof requires a slightly more general notion of re flector.
Definition 4.1.6. The general form of the Householder matrix, also called
anelementary re flector , has the form
H
v=I−τvv∗
where the vector v∈Cn.
In order for Hvto be unitary, it must be true that
HvH∗
v=(I−τvv∗)(I−τvv∗)∗
=(I−τvv∗)(I−¯τvv∗)
=I−τvv∗−¯τvv∗+|τ|2(v∗v)vv∗
=I−2Re (τ)vv∗+|τ|2|v|2vv∗
=I
Therefore, for v6=0w em u s th a v e
−2Re (τ)vv∗+|τ|2|v|2vv∗=³
−2Re (τ)+|τ|2|v|2´
vv∗
=0
or
−2Re (τ)+|τ|2|v|2=0
Now suppose that Q∈Mn(R) is orthogonal and that the spectrum σ(Q)⊂
{−1,1}.Suppose Qhas a complete set of normalized orthogonal eigen-
vectors1, it can be expressed as Q=nP
i=1λivivT
iwhere the set v1,...,v nare
the eigenvectors and λi⊂σ(Q). Now assume the eigenvectors have been
arranged so that this simpli fies to
Q=−kX
i=1vivT
i+nX
i=k+1vivT
i
1This is in fact a theorem that will be established in Chapter 4.2. It follows as a
consequence of Schur’s theorem
170 CHAPTER 4. UNITARY MATRICES
Here we have just arrange the eigenvectors with eigenvalue −1t oc o m e first.
Define
Hj=I−2vjvT
j,j =1,...,k
It follows that
Q=kY
j=1Hj=kY
j=1¡
I−2vjvT
j¢
for it is easy to check that
Uvm=kY
j=1Hj=½−vmifm≤k
vm ifm>k
Since we have agreement with Qon a basis, the equality follows. In words we
may say that an orthogonal matrix Uwith spectrum σ(Q)⊂{−1,1}with
can be written as a product of re flectors. With this simple case out of
t h ew a y ,w ec o n s i d e rm o r eg e n e r a lc a s ew ew r i t e Q=nP
i=1λivivT
i.where the
setv1,...,v nare the eigenvectors and λi⊂σ(Q). We wish to represent Q
similar to the above formula as a product of elementary re flectors.
Q=kY
j=1Hj=kY
j=1¡
I−τjwjw∗
j¢
where wi=αivifor some scalars αi. On the one hand it must be true
that−2Re (τi)+|τi|2|wi|2=0f o re a c h i=1,...n , and on the other hand
it must follow that
(I−τiwiw∗
i)vm=½λiviifm=i
vmifm6=i
The second of these relations is automatically satis fied by the orthogonality
of the eigenvectors. The second relation can needs to be solved. This
simpli fies to ( I−τiwiw∗
i)vi=³
1−τi|αi|2´
vi=λivi.T h e r e f o r e , i t i s
necessary to solve the system
−2Re (τi)+|τi|2|vi|2=0
1−τi|αi|2=λi
Having done so there results the factorization.
4.2. SCHUR’S THEOREM 171
Theorem 4.1.5. LetQ∈Mnbe real orthogonal or unitary. Then Qcan
be factored as the product of elementary re flectors Q=nQ
j=1³
I−τjwjw∗
j´
,
where the wjare the eigenvectors of Q.
Note that the product written here is up to nwhereas the earlier product
was just up to k. The difference here is slight, for by taking the scalar ( α)
equal zero when necessary, it is possible to equate the second form to thefirst form when the spectrum is contained in the set {−1,1}.
4.2 Schur’s theorem
It has already been established in Theorem 3.5.2 that every matrix is similar
to a triangular matrix. A far stronger result is possible. Called Schur’stheorem, this result proves that the similarity is actually unitary similarity.
Theorem 4.2.1 (Schur’s Theorem). Every matrix A∈M
n(C)is uni-
tarily equivalent to a triangular matrix.
Proof. We proceed by induction on the size of the matrix n. First suppose
theAis 2×2.Then for a given eigenvalue λand normalized eigenvector
vform the matrix Pwith its first column vand second column any vector
orthogonal to vwith norm one. Then Pis an unitary matrix and P∗AP=·λ∗
0∗¸
. This is the desired triangular form. Now assume the Schur
factorization is possible for matrices up to size ( n−1)×(n−1). For the
given n×nmatrix Aselect any eigenvalue λ. With its pertaining normalized
eigenvector vconstruct the matrix Pwith vin the first column and an
orthonormal complementary basis in the remaining n−1 columns. Then
P∗AP=
λˆa
12··· ˆa1n
0ˆa22··· ˆa2n
............
0ˆan2··· ˆann
=
λˆa
12··· ˆa1n
0
... A2
0
The eigenvalues of A
2together with λconstitute the eigenvalues of A.B y
the inductive hypothesis there is an ( n−1)×(n−1) unitary matrix ˆQsuch
that ˆQ∗A2ˆQ=T2,where T2is triangular. Now embed ˆQin an n×nmatrix
172 CHAPTER 4. UNITARY MATRICES
Qas shown below
Q=
λ0··· 0
0
... ˆQ
0
It follows that
Q
∗P∗APQ =
λ0··· 0
0
... T
2
0
=T
is triangular, as indicated. Therefore the factorization is complete upon
defining the similarity transformation U=PQ. Of course, the eigenvalues
ofA
2together with λconstitute the eigenvalues of A, and therefore by
similarity the diagonal of Tcontains only eigenvalues of A.
The triangular matrix above is not unique as is easy to see by mixing or
permuting the eigenvalues. An import ant consequence of Schur’s theorem
pertains to unitary matrices. Suppose that Bis unitary and we apply the
Schur factorization to write B=U∗TU.T h e n T=UBU∗. It follows that
the upper triangular matrix is itself unitary, which is to say T∗T=I.It is a
simple fact to prove this implies that Tis in fact a diagonal matrix. Thus,
the following result is proved.
Corollary 4.2.1. Every unitary matrix is diagon alizable. Moreover, every
unitary matrix has northogonal eigenvectors.
Proposition 4.2.1. IfB,A∈M2are similar and tr (B∗B)=tr(A∗A),t h e n
BandAare unitarily equivalent.
This result higher dimensions.
Theorem 4.2.2. Suppose that A∈Mn(C)has distinct eigenvalues λ1...λk
with multiplicities n1,n2,... ,n krespectively. Then Ais similar to a matrix
of the form
T
1
0
T2
0...
Tk
4.2. SCHUR’S THEOREM 173
where Ti∈Mniis upper triangular with diagonal entries λi.
Proof. LetE1be eigenspace of λ1. Assume dim E1=m1.L e t u1...u m1be
an orthogonal basis of E1, and suppose v1...v n−m1is an orthonormal basis
ofCn−E1.T h e n w i t h P=[u1...u m1,v1...v n−m1]w eh a v e P−1APhas
the block structure
·T10
0A2¸
.
Continue this way, as we have done before. However, if dim Ei<m iwe
must proceed a di fferent way. First apply Schur’s Theorem to triangularize.
Arrange the firstm1eigenvalues to the firstm1diagonal entries of T,t h e
similar triangular matrix. Let Er,sb et h em a t r i xw i t ha1i nt h e r, sposition
and 0’s elsewhere. We assume throughout that r6=s. Then, it is easy to
see that I+αEr,sis invertible and that ( I+αErs)−1=I−αErs.N o w i f
we de fine
(I+αErs)−1T(I+αErs)
we see that trs→trs+α(trr−tss). We know trr−tss6=0i f randspertain
to values with di fferent eigenvalues. Now this value is changed, but so also
are the values above and to the right of trs.This is illustrated below.
columns → s
row r
↑
∗↑
∗→ →
To see how to use these similarity transformation to zero out the upper
blocks, consider the special case with just two blocks
A=
T
1 T12
0 T2
We will give a procedure to use elementary matrices αE
ijto zero-out each
of the entries in T12.The matrix T12hasm1·m2entries and we will need
174 CHAPTER 4. UNITARY MATRICES
to use exactly that many of the matrices αEij.In the diagram below we
illustrate the order of removal (zeroing-out) of upper entries
T
1m12m1···m1·m2
.........
2...
1m1+1 ··· ···
0 T2
W h e np e r f o r m e di nt h i sw a yw ed e v e l o p m
1·m2constants αij.We proceed in
the upper right block ( T12) in the left-most column and bottom entry, that is
position ( m1,m1+1 ).For the ( m1,m1+ 1) position we use the elementary
matrixαEm1,m1+1to zero out this position. This is possible because the
diagonal entries of T1andT2are different. Now use the elementary matrix
αEm1−1,m1+1to zero-out the position ( m1−1,m1+ 1). Proceed up this
column to the first row each time using a new α.W e a r e finished when
all the entries in column m1+1 f r o m r o w m1to row 1 are zero. Now
focus on the next column to the right, column m1+ 2. Proceed in the
same way from the ( m1,m1+ 2) position to the (1 ,m1+ 1) position, using
elementary matrices αEk,m 1+2zeroing out the entries ( k,m 1+ 2) positions,
k=m1,..., 1. Ultimately, we can zero out every entry in the upper right
block with this marching left-to-right, bottom-to-top procedure.
In the general scheme with k×kblocks, next use the matrices T2and
T3to zero-out the block matrix T23.(See below.)
T=
T
10T13
T2T23
T3... ...
0...
Tk−1Tk−1,k
Tk
After that it is possible using blocks T
1andT3to zero-out the block T13.
This is the general scheme for the entire triangular structure, moving downthe super-diagonal ( j.j+ 1) blocks and then up the (block) columns. In this
way we can zero-out any value not in the square diagonal blocks pertaining
to the eigenvalues λ
1...λk.
4.2. SCHUR’S THEOREM 175
Remark 4.2.1. IfA∈Mn(R)a n dσ(A)⊂R, then all operations can be
carried out with real numbers.
Lemma 4.2.1. LetJ⊂Mnbe a commuting family. Then there is a vector
x∈Cnfor which xis an eigenvector for every A∈J.
Proof. LetW⊂Cnbe a subspace of minimal dimension that is invariant
under J.S i n c e Cnis invariant, we know Wexists. Since Cnis invariant,
we know Wexists. Suppose there is an A∈Jfor which there is a vector
inWwhich is not an eigenvector of A.D e fineW0={y∈W|Ay=
λyfor some λ}.T h a t i s , W0is a set of eigenvectors of A.S i n c e Wis
invariant under A, it follows that W06=φ. Also, by assumption. W0$W.
For any x∈W0
ABx =(AB)x=B(Ax)=λBx
and so Bx∈W0. It follows that W0isJinvariant, and W0has lower
positive dimension than W.
As a consequence we have the following result.
Theorem 4.2.3. LetJbe a commuting family in Mn.I fA∈Jis diago-
nalizable, then Jis simultaneously diagonalizable.
Proof. Since Ais diagonalizable there are nlinearly independent eigenvec-
tors of Aand if Sis in Mnand consists of those eigenvectors the matrix
S−1ASis diagonal. Since eigenvectors of Aare the same as eigenvectors of
B∈J, it follows that S−1BSis also diagonal. Thus S−1JS=D:= all
diagonal matrices.
Theorem 4.2.4. LetJ⊂Mnbe a commuting family. There is a unitary
matrix U∈Mnsuch that U∗AU is upper triangular for each A∈J.
Proof. From the proof of Schur’s Theorem we have that the eigenvectors
chosen for Uare the same for all A∈J. This follows because after the first
step we have reduced AandBto
·A11A12
0A22¸
and·B11B12
0B22¸
respectively. Commutativity is preserved under simultaneous similarity and
therefore A22andB22commutes. Therefore at the second step the same
eigenvector can be selected for allB22⊂J2, a commuting family.
176 CHAPTER 4. UNITARY MATRICES
Theorem 4.2.5. IfA∈Mn(R), there is a real orthogonal matrix Q∈
Mn(R)such that
(?) QTAQ=
A1
A2?
0...
Ak
1≤k≤n
where each Aiis a real 1×1matrix or a real 2×2matrix with a non real
pair of complex conjugate eigenvalues.
Theorem 4.2.6. LetJ⊂Mn(R)be a commuting family. There is a real
orthogonal matrix Q∈Mn(R)for which QTAQ has the form (?)for every
A∈J.
Theorem 4.2.7. Suppose A, B∈Mn(C)have eigenvalues α1,... ,αnand
β1,... ,βnrespectively. If AandBcommute then there is a permutation
i1...i nof the integers 1,... ,n for which the eigenvalues of A+Bareαj+
βij,j=1,... ,n .T h u sσ(A+B)⊂σ(A)+σ(B).
Proof. Since J={A, B}f o r m sac o m m u t i n gf a m i l yw eh a v et h a te v e r y
eigenvector of Ais an eigenvector of B, and conversely. Thus if Ax=αjx
we must have that Bx=Bijx. But how do we get the permutation? The
answer is to simultaneously triangularize with U∈Mn.W eh a v e
U∗AU=Tand U∗BU=R.
Since
U∗(A+B)U=T+R
we have that the eigenvalues pair up as described.
Note: We don’t necessarily need AandBto commute. We need only the
hypothesis that AandBare simultaneously diagonalizable.
IfAandBdo not c o m m u t el i t t l ec a nb es a i do f σ(A+B). In particular
σ(A+B)$σ(A)+σ(B). Indeed, by summing upper triangular and lower
triangular matrices we can exhibit a range of possibilities. Let A=[01
00],
B=[00
10],σ(A+B)={−1,1}butσ(A)=σ(B)={0}.
Corollary 4.2.1. Suppose A, B∈Mnare commuting matrices with eigen-
valuesα
1,... ,αnandβ1,... ,βnrespectively. If αi6=−βjfor all 1≤i, j≤n,
thenA+Bis nonsingular.
4.3. EXERCISES 177
Note: Giving conditions for the nonsingularity of A+Bin terms of various
conditions on AandBis a very di fferent problem. It has many possible
answers.
4.3 Exercises
1. Characterize all diagonal unitary matrices and all diagonal orthogonal
matrices in Mn.
2. Let T(θ)=£cosθ−sinθ
sinθcosθ¤
whereθis any real. Then T(θ)i srealorthog-
onal. Prove that if U∈M2(R) is real orthogonal, then Uhas the form
T(θ)f o rs o m e θor
U=·10
0−1¸
T(θ),
and conversely.
3. Let us de fine the n×nmatrix Pto be a w-permutation matrix if
for every vector x∈Rn, the vector Px has the same components
asxin value and number, though possibly permuted. Show that
w-permutation matrices are permutation matrices.
4. In R2identify all Householder transformations that are rotations.
5. Given any unit vector win the plane formed by eiandej.E x p r e s s
the Householder matrix for this vector.
6. Show that the set of unitary matrices on Cnforms a subgroup of the
subset of GL(n,C).
7. In R2,prove that the product of two (Householder) re flections is a
rotation. (Hint. If the re flection angles are θ1andθ2,then the rotation
angle is 2 ( θ1−θ2).)
8. Prove that the unitary matrices are closed the norm. That is, if the
sequence {Un}⊂Mn(C) are all unitary and lim n→∞Un=Uin the
k·k2norm, then Uis also unitary.
9. Let A∈Mnbe invertible. De fineG=Ak,(A−1)k,k=1,2,.... Show
thatGis a subgroup of Mn. Here the group multiplication is matrix
multiplication.
178 CHAPTER 4. UNITARY MATRICES
10. Prove that the unitary matrices are closed under pointwise conver-
gence. That is, if the sequence {Un}⊂Mn(C) are all unitary and
limn→∞Un=Ufor each ( ij)entry, then Uis also unitary.
11. Suppose that A∈Mnand that AB=BAfor all B∈Mn.Show that
Ais a multiple of the identity.
12. Prove that a unitary matrix Ucan be written as V−1W−1VW for
unitary V,W if and only if det U= 1. (Bellman, 1970)
13. Prove that the only triangular unitary matrices are diagonal matrices.
14. We know that given any bounded sequence of numbers, there is a
convergent subsequence. (Bolzano-Weierstrass Theorem). Show that
the same is true for matrices for any given matrix norm. In particular,show that if U
nis any sequence of unitary matrices, then there is a
convergent subsequence.
15. The Hadamard Gate (from quantum computing) is de fined by the
matrix H=1√
2·11
1−1¸
. Show that this transformation is a House-
holder matrix. What are its eigenvalues and eigenvectors?
16. Show that the unitary matrices do not form a subspace Mn(C).
17. Prove that if a matrix A∈Mn(C) preserves the orthonormality of one
orthonormal basis, then it must be unitary.
18. Call the matrix Aacheckerboard matrix if either
(I)aij=0 ifi+jis even or (II) aij=0ifi+jis odd
We call the matrices of type I even checkerboard and type II odd
checkerboard .D efineCH(n)t ob ea l l n×ninvertible checkerboard
matrices. The questions below all pertain to square matrices.
(a) Show that if nis odd there are no invertible even checkerboard
matrices.
(b) Prove that every unitary matrix Uhas determinant with modulus
one. (That is, |detU|=1.)
(c) Prove that nis odd the product of odd checkerboard matrices is
odd.
4.3. EXERCISES 179
(d) Prove that if nis even then the product of an even and an odd
checkerboard matrix is odd, while the product of two even (orodd) checkerboard matrices is even.
(e) Prove that if nis even then the inverse of any checkerboard matrix
is a checkerboard matrix of the same type. However, in light of(a), it is only true that if nis an invertible odd checkerboard
matrix, its inverse is odd.
(f) Suppose that n=2m. Characterize all the odd checkerboard
invertible matrices.
(g) Prove that CH(n) is a subgroup of GL(n).
(h) Prove that the invertible odd checkerboard matrices of any size
nforms a subgroup of CH(n).
In connection with chess, checkerboard matrices give the type of chess
board on which the maximum number of mutually non-attacking knightscan be placed on the even ( i+jis even) or odd ( i+jis odd) positions.
19. Prove that every unitary matrix can be written as the product of a
unitary diagonal matrix and another unitary matrix whose first column
has nonnegative entries.
180 CHAPTER 4. UNITARY MATRICES
Chapter 5
Hermitian Theory
Hermitian matrices form one of the most useful classes of square matri-
ces. They occur naturally in a variety of applications from the solution ofpartial di fferential equations to signal and image processing. Fortunately,
they possess the most desirable of matrix properties and present the user
with a relative ease of computation. There are several very powerful factsabout Hermitian matrices that have found universal application. First thespectrum of Hermitian matrices is real. Second, Hermitian matrices have acomplete set of orthogonal eigenvectors, which makes them diagonalizable.Third, these facts give a spectral repre sentation for Hermitian matrices and
a corresponding method to approximate them by matrices of less rank.
5.1 Diagonalizability of Hermitian Matrices
Let’s begin by recalling the basic de finition.
Definition 5.1.1. LetA∈Mn(C). We say that AisHermitian ifA=A∗,
where A∗=¯AT.A∗is called the adjoint of A. This, of course, is in con flict
with the other de finition of adjoint, which is given in terms of minors.
Recall the following facts and de finitions about subspaces of Cn:
•IfU, V are subspaces of Cn,w ed e fine the direct sum ofUandVby
U⊕V={u+v|u∈U, v∈V}.
•IfU, V are subspaces of Cn,w es a y UandVare orthogonal if hu, vi=0
for every u∈Uandv∈V.I nt h i sc a s ew ew r i t e U⊥V.
For example, a natural way to obtain orthogonal subspaces is from ortho-
normal bases. Suppose that {u1,...,u n}is an orthonormal basis of Cn.Let
181
182 CHAPTER 5. HERMITIAN THEORY
the integers {1,...,n }be divided into two (disjoint) subsets J1andJ2.Now
define
U1=S{ui|i∈J1}
U2=S{ui|i∈J2}
Then U1andU2are orthogonal, i.e. U1⊥U2,and
U1⊕U2=Cn
hAu, x i=hu, Ax i=λhu, xi=0
Our main result is that Hermitian matrices are diagonalizable. To prove it,
we reveal other interesting and importa nt properties of Hermitian matrices.
F o re x a m p l e ,c o n s i d e rt h ef o l l o w i n g .
Theorem 5.1.1. LetA∈Mn(C)be Hermitian. Then the spectrum of
A,σ(A),i sr e a l .
Proof. Letλ∈σ(A) with corresponding eigenvector x∈Cn.T h e n
hAx, x i=hx, Ax i=hx,λxi=¯λhx, xi
k
hλx, xi
k
λhx, xi.
Since we know kxk2=hx, xi6= 0, it follows that λ=¯λ, which is to say that
λis real.
Theorem 5.1.2. LetA∈Mn(C)be Hermitian and suppose that λandµ
are different eigenvalues with corresponding eigenvectors xandy.T h e n
x⊥y(i.e. hx, yi=0).
Proof. We know Ax=λxandAy=µy. Now compute
hAx, y i=hx, A∗yi=hx, Ay i=µhx, yi
k
λhx, yi.
Ifhx, yi6= 0, the equality above yields a contradiction and the result is
proved.
5.1. DIAGONALIZABILITY OF HERMITIAN MATRICES 183
Remark 5.1.1. This result also follows from the previously proved result
about the orthogonality of left and right eigenvectors pertaining to di fferent
eigenvalues.
Theorem 5.1.3. LetA∈Mn(C)be Hermitian, and let λbe an eigenvalue
ofA. Then the algebraic and geometric multiplicities of λare equal. In
symbols,
ma(λ)=mg(λ).
Proof. We prove this result by reconsideration of our main result on triangu-
larization of Aby a similarity transformation. Let x∈Cnbe an eigenvector
ofApertaining to λ.S o , Ax=λx.L e t u2,... ,u n⊂Cnbe a set of vectors
orthogonal to x,s ot h a t {x, u 2,... ,u n}is a basis of Cn. Indeed, it is an
orthogonal basis, and by normalizing the vectors it becomes an orthonor-
mal basis. We claim that U2=S(u2,... ,u n), the span of {u2,... ,u n}is
invariant under A. To see this, suppose
u=nX
j=2cjuj∈U2
and
Au=v+ax
where v∈U2anda6= 0. Then, on the one hand
hAu, x i=hu, Ax i=λhu, xi=0
On the other hand
hAu, x i=hv+ax, x i=ahx, xi=akxk26=0
This contraction establishes that the span of {u2,... ,u n}is invariant under
A.
LetU⊕V=Cnbe invariant subspaces with U⊥V. Suppose that UB
and VBare orthonormal bases of the orthogonal subspaces UandV.Define
the matrix P∈Mn(C) by taking for its columns first the basis vectors UB
and and then the basis vectors VB.L e t u s w r i t e P=[UB,VB]( w i t ho n l y
a small abuse of notation). Then, since Ais Hermitian,
B=P−1AP
=·Au0..........
0Av¸
184 CHAPTER 5. HERMITIAN THEORY
The whole process can be carried out exactly ma(λ)t i m e s ,e a c ht i m e
generating a new orthogonal eigenvector pertaining to λ.T h i s e s t a b l i s h e s
thatmg(λ)=ma(λ). A formal induction could have been given.
Remark 5.1.2. Note how we applied orthogonality and invariance to force
the triangular matrix of the previous result to become diagonal. This is
what permitted the successive extracti on of eigenvectors. Indeed, if for any
eigenvector xthe subspace of Cnorthogonal to xis invariant, we could have
carried out the same steps as above.
We are now in a position to state our main result, whose proof is implicit
in the three lemmas above.
Theorem 5.1.4. LetA∈Mn(C)be Hermitian. Then Ais diagonalizable.
The matrix Pfor which P−1AP is diagonal can be taken to be orthogonal.
Finally, if {λ1,...,λn}and {u1,...,u n}denote eigenvalues and pertaining
orthonormal eigenvectors for A,t h e n Aadmits the spectral representation
A=Pn
j=1λjujuT
j.
Corollary 5.1.1. LetA∈Mn(C)be Hermitian.
(i)Ahasnlinearly independent and orthogonal eigenvectors.
(ii)Ais unitarily equivalent to a diagonal matrix.
(iii) If A, B∈Mnare unitarily equivalent, then Ais Hermitian if and only
ifBis Hermitian.
Note that in part (iii) above, the condition of unitary equivalence cannot be
replaced by just similarity. (Why?)
Theorem 5.1.5. IfA, B∈Mn(C)andA∼Bwith Sas the similarity
transformation matrix, B=S−1AS.I f Ax=λxandy=S−1x,t h e n
By=λy.
If matrices are similar so also are their eigenstructures. It should establish
the very closeness that similarity implies. Later as we consider decomposi-
tion theorems, we will see even more remarkable consequences.
Though we have as yet no method of determining the eigenvalues of a
matrix beyond factoring the characteristic polynomial, it is instructive tosee how their existence impacts the fundamental problem of solving Ax=b.
Suppose that Ais Hermitian with eigenvalues λ
1,...,λn,c o u n t e da c c o r d i n g
to multiplicity and with o rthonormal eigenvectors {u1,...,u n}.C o n s i d e r
5.1. DIAGONALIZABILITY OF HERMITIAN MATRICES 185
the following solution method for the system Ax=b.Since the span of the
eigenvectors is Cnthen
b=nX
i=1biui
where as we know by the orthonormality of the vectors {u1,...,u n}that
bi=hb, uii. We can also write x=Pn
i=1xiui.Then, the system becomes
Ax =AÃnX
i=1xiui!
=nX
i=1xiλiui=nX
i=1biui
Therefore, the solution is
xi=bi
λi,i=1,...,n
Expanding the data vector bin the basis of eigenvectors yields a rapid
method to find the solution to the system. Nonetheless, this is not the
preferred method for solving linear systems when the coe fficient matrix is
Hermitian. Finding all the eigenvectors is usually costly, and other waysare available that are more e fficient. We will discuss a few of them in in
the section and in later chapters.
Approximating Hermitian matrices
With the spectral representation available, we have a tool to approximate thematrix, keeping the “important” part and discarding the less important part.Suppose the eigenvalues are arranged in decending order |λ
1|≥···≥|λn|.
Now approximate Aby
Ak=kX
j=1λjujuT
j (1)
This is an n×nmatrix. The di fference A−Ak=Pn
j=k+1λjujuT
j.We can
approximate the norm of the di fference by
(A−Ak)x=
nX
j=k+1λjujuT
j
x=nX
j=k+1λjxjuj
186 CHAPTER 5. HERMITIAN THEORY
where x=Pn
j=1xjuj.Assume kxk= 1. By the Cauchy-Schwartz inequality
k(A−Ak)xk2=°°°°°°nX
j=k+1λjxjuj°°°°°°2
≤nX
j=k+1|λj|2
Therefore, k(A−Ak)k≤³Pn
j=k+1|λj|2´1/2
. From this we can conclude
that if the smaller eigenvalues are su fficiently small, the matrix can be ac-
curately approximated by a matrix of lesser rank.
Example 5.1.1. The matrix
A=
0.5745−0.5005 0 .1005 0 .0000
−0.5005 1 .176−0.5756 0 .1005
0.1005−0.5756 1 .176−0.5005
0.0000 0 .1005−0.5005 0 .5745
has eigenvalues eigenvectors {2.004,0.9877,0.3219,0.1872 }with pertaining
eigenvectors
u1=
0.2740
−0.6519
0.6519
−0.2740
,u2=
0.4918
−0.5080
−0.5080
0.4918
,u3=
0.6519
0.2740
−0.2740
−0.6519
,u 4=
0.5080
0.4918
0.4918
0.5080
respectively. Neglecting the eigenvec tors pertaining to the two smaller
eigenvalues Ais approximated according as 1 the formula above by
A2=2X
j=1λjujuT
j=λ1u1uT
1+λ2u2uT
2
=2 .004
0.274
−0.6519
0.6519
−0.274
0.274
−0.6519
0.6519
−0.274
T
+0.9877
0.4918
−0.508
−0.508
0.4918
0.4918
−0.508
−0.508
0.4918
T
A2=
0.3893−0.6047 0 .1112 0 .0884
−0.6047 1 .107−0.5968 0 .1112
0.1112−0.5968 1 .107−0.6047
0.0884 0 .1112−0.6047 0 .3893
5.2. FINDING EIGENVECTORS 187
The difference
A−A2=
0.1852 0 .1042−0.0107−0.0884
0.1042 0 .069 0 .0212−0.0107
−0.0107 0 .0212 0 .069 0 .1042
−0.0884−0.0107 0 .1042 0 .1852
has 2-norm kA−A2k2=0.3218, while the 2-norm kAk2=2.004. The
relative error of approximation iskA−A2k2
kAk2=0.3218
2.004=0.1606.
To illustrate how this may be used, let us attempt to use A2to approx-
imate the solution of Ax=b,w h e r e b=[ 2.606,−4.087,1.113,0.346 4]T.
First of all the exact solution is x=[ 2.223,−2.688,−0.162 9,0.931 2]T.
Since the matrix A2has rank two, it is not solvable for every vector b.
We therefore project the vector binto the span of the range of A2,n a m e l y
u1andu2.T h u s
b2=hb, u1iu1+hb, u2iu2=[ 2.555,−4.118,1.109,0.358 4]T
Now solve A2x2=b2,t oo b t a i n x2=[ 2.023,−2.828,−0.220 2,0.927 4]T.
The 2-norm of the di fference is kx−x2k2=0.250 8. This error, though
not extremely small, can be accounted for by the fact that the data vector
bhas sizable u3andu4components. That is°°projS(u3,u4)b°°=0.250 8.
5.2 Finding eigenvectors
Recall that a zero of a polynomial is called simple if its multiplicity is one.
If the eigenvalues of A∈Mn(C) are distinct and the largest, λn, in modulus
is simple, then there is an iterative method to find it. Assume
(1) |λn|=ρ(A)
(2) ma(λn)=1 .
Moreover, without loss of generality we assume eigenvalues to be ordered
|λ1|≤···≤|λn−1|<|λn|=ρn. Select x(0)∈C(n).D efine
x(k+1)=1
kx(k)kAx(k).
By scaling we can assume that λn=1 ,a n dt h a t y(1),... ,y(n)are linearly
independent eigenvectors of A.S oAy(n)=y(n).W ec a nw r i t e
x(0)=c1y(1)+···+cny(n).
188 CHAPTER 5. HERMITIAN THEORY
Then, except for a scale factor (i.e. the factor kx(k)k−1)
x(k)=c1λk
1y(1)+···+cnλk
ny(n).
Since |λj|<1,j=1,2,... ,n−1, we have that |λk
j|→0i fj=1,2,... ,n−1.
Therefore, the limit of x(k)approaches a multiple of y(n).
This gives the following result.
Theorem 5.2.1. LetA∈Mn(C)have ndistinct eigenvalues and assume
the eigenvalue λnwith modulus ρ(A)is simple. If x(0)is not orthogonal to
the eigenvector y(n)pertaining to λn, then the sequence of vectors de fined by
x(k+1)=1
kx(k)kAx(k)
converges to a multiple of y(n).T h i si sc a l l e dt h e Power Method .
The rate of convergence is controlled by |λn−1|. The closer to 1 this
number is the slower the iterates converge. Also, if we know only that λn
(for which |λn|=ρ(A) is simple we can determine what it is by considering
theRayleigh quotient. Take
ρk=hAx(k),x(k)i
hx(k),x(k)i.
Then lim
k→∞ρk=λn. Thus the multiple of y(n)is indeed λn.
Tofind intermediate eigenvalues and eigenvectors we apply an adaptation
of the power method called the orthogonalization method . However, in order
to adapt the power method to determine λn−1,our underlying assumption
is that is also simple and morover |λn−2|<|λn−1|.
Assume y(n)andλna r ek n o w n . T h e nw er e s t a r tt h ei t e r a t i o n ,t a k i n g
the starting value
ˆx(0)=x(0)−hx(0),y(n)i
ky(n)k2y(n).
We know that the eigenvector y(n−1)pertaining to λn−1is orthogonal to
y(n). Thus, in theory all of the iterates
ˆx(k+1)=1
kˆx(k)kTˆx(k)
5.2. FINDING EIGENVECTORS 189
will remain orthogonal to y(n). Therefore,
lim
k→∞ˆx(k)=y(n−1).
–in theory. In practice, however, we must accept that y(n)has not been
determined exactly. This means ˆ x(0)has not been purged of all of y(n).B y
our previous reasoning, since λnis the dominant eigenvalue, the presence of
y(n)willcreep back into the iterates ˆ x(k). To reduce the contamination it is
best to purify the iterates ˆ x(k)periodically by the reduction
(?)ˆ x(k)−→ˆx(k)−hˆx(k),y(n)i
ky(n)ky(n)
before computing ˆ x(k+1). The previous argument can be applied to prove
that
(2) lim
k→∞x(k)=y(n−1)
(2) lim
k→∞hAx(k),x(k)i
kx(k)k2=λn−1.
Additionally, even if we know y(n)exactly, round-o fferror would reinstate
ay(n)component in our iterative computations. Thus the puri fication step
above, ( ?), should be applied in allcircumstances.
Finally, subsequent eigenvalues and eigenvectors may be determined by
successive orthogonalizations. Again th e eigenvalue simplicity and strict
inequality is needed for convergence. Speci fically, all eigenvectors can be
determined if we assume that eigenvalues to be strictly ordered |λ1|<···<
|λn−1|<|λn|=ρn. For example, we begin the iterations to determine y(n−j)
with
ˆx(0)=x(0)−j−1X
i=0hx(0),y(n−i)i
ky(n−i)k2y(n−i).
Don’t forget the re-orthogonalizations periodically throughout the iterative
process.
W h a tc a nb ed o n et o find intermediate eigenvalues and eigenvectors in
the case Ais not symmetric? The method above fails, but a variation of it
works.
What must be done is to generate the left and right eigenvectors, w(n)
andy(n),f o r A. Use the same process. To compute y(n−1)andw(n−1)we
190 CHAPTER 5. HERMITIAN THEORY
orthogonalize thusly:
ˆx(0)=x(0)−hx(0),w(n)i
kw(n)k2w(n)
ˆz(0)=z(0)−hz(0),y(n)i
ky(n)k2y(n)
where z(0)is the original starting value used to determine the left eigenvector
w(n).S i n c ew ek n o wt h a t
lim
k→∞x(k)=αy(n)
it is easy to see that
lim
k→∞Ax(k)=αλny(n).
Therefore,
lim
k→∞hAx(k),x(k)i
hx(k),x(k)i=λn.
Example 5.2.1. LetTbe the transformation of R2→R2that rotates a
vector by θradians. Then it is clear that no matter what nonzero vector
x(0)is selected the iterations x(k+1)=Tx(k)will never converge. (Assume
kx(0)k= 1.) Now the matrix representation of Tis
AT=·cosθ−sinθ
sinθcosθ¸
rotates counterclockwise
we have
pAT(λ)=d e t·λ−cosθ sinθ
−sinθλ−cosθ¸
=(λ−cosθ)2+s i n2θ.
The spectrum of ATis therefore
λ=c o sθ±isinθ.
Notice that the eigenvalues are discrete, but there are twoeigenvalues with
modulus ρ(AT) = 1. The above results therefore do not apply.
5.3. POSITIVE DEFINITE MATRICES 191
Example 5.2.2. Although the previous example is not based on a symmet-
ric matrix, it certainly illustrates non convergence of the power iterations.
The even simpler Householder matrix·10
0−1¸
furnishes us with a sym-
metric matrix for which the iterations also do not converge. In this case,
there are two eigenvalues with modulus equal to the spectral radius ( ±1).
With arbitrary starting vector x(0)=[a, b]T,i ti so b v i o u st h a tt h ee v e n
iterations are x(2i)=[a, b]Ta n dt h eo d di t e r a t i o n sa r e x(2i−1)=[a,−b]T.
Assuming again that the eigenvalues are distinct and even stronger, as-
suming that
|λ1|<|λ2|<···<|λn|
we can apply the process above to extract all the eigenvalues (Rayleigh
quotient) and the eigenvectors, one-by-one, when AAAis symmetric .
First of all, considering the matrix A−σIwe can shift the eigenvalues
to either the left or the right. Depending on the location of λnas ufficiently
large |σ|may be chosen so that |λ1−σ|=ρ(A−σI). The power method
can be applied to determine λ1−σand hence λ1.
5.3 Positive de finite matrices
Of the many important subclasses of Hermitian matrices, there is one class
that stands out.
Definition 5.3.1. We say that A∈Mn(C)i spositive de finiteifhAx, x i>
0 for every nonzero x∈Cn. Similarly, we say that A∈Mn(C)i spositive
semide finiteifhAx, x i≥0 for every nonzero x∈Cn.
It is easy to see that for positive de finite matrices all of the results are
true
Theorem 5.3.1. LetA, B∈Mn(C).T h e n
1. If Ais positive de finite, then σ(A)⊂R+
n
2. If Ais positive de finite, then Ais invertible.
3.B∗Bis positive semide finite.
4. If Bis invertible then B∗Bis positive de finite.
192 CHAPTER 5. HERMITIAN THEORY
5. If B∈Mn(C)is positive semide finite, then diag (B)is nonnegative,
and diag (B)is strictly positive when Bis postive de finite.
The proofs are all routine. Of course, every diagonal matrix with non-
negative entries is positive semide finite.
Square roots
Given a real matrix A∈Mn. It is sometimes desired to determine a square
root of A.B y t h i s w e m e a n a n y m a t r i x Bfor which B2=A.Moreover,
if possible, it is desired that the square root be real. Our experience withnumbers indicates that in order that a number have a positive square root,it must be positive. The analogue for matrices is the condition of beingpositive de finite.
Theorem 5.3.2. LetA∈M
nbe positive [semi-]de finite. Then Ahas a
real square root. Moreover, the square r oot can taken to be positive //[semi-
]definite.
Proof. We can write the diagonal matrix of the eigenvalues of Ain the equa-
tionA=P−1DP. E x t r a c tt h ep o s i t i v es q u a r er o o to f DasD1
2=diag(λ1/2
1,...,λ1/2
n).
Obviously D1
2D1
2=D.N o w d e fineA1
2byA1
2=P−1D1
2P.T h i s m a t r i x i s
real. It is simple to check that A1
2A1
2=A,and that this particular square
root is positive de finite.
Clearly any real diagonalizable matrix with nonnegative eigenvectors has a
real square root as well. However, bey ond that conditions for determin-
ing existence let alone determination of square roots take us into a very
specialized subject.
5.4 Singular Value Decomposition
Definition 5.4.1. For any A∈Mmn,t h e n×nHermitian matrix A∗Ais
positive semi-de finite. Denoting its eigenvalues by λjwe called the valuesp
λjthesingular values ofA.
Because r(A∗A)≤min ( r(A∗),r(A))≤min(m, n)t h e r ea r ea tm o s t
min ( m, n) nonzero singular values.
Lemma 5.4.1. LetA∈Mmn.There is an orthonormal basis {u1,...,u n}of
Cnsuch that {Au1,...,A u n}is orthogonal.
5.4. SINGULAR VALUE DECOMPOSITION 193
Proof. Proof. Consider the n×nHermitian matrix A∗A, and denote an
orthonormal basis of its eigenvectors by {u1,...,u n}.T h e ni ti se a s yt os e e
that {Au1,...,A u n}is an orthogonal set. For hAuj,A u ki=hA∗Auj,uki=
λjhuj,uki=0.
Lemma 5.4.2. Lemma 2 Let A∈Mmnand an orthonormal basis {u1,...,u n}of
Cn.D efine
vj=(
1
kAujkAujifkAujk6=0
0 ifkAujk=0
LetS=diag(kAu1k,..., kAunk),t h e n×nmatrix Uhaving rows given
by the basis {u1,...,u n}and ˆVthem×nmatrix given by the columns
{v1,...,v n}.T h e n A=ˆVS U .
Proof. Proof. Consider
ˆVS Uu j=
v1v2 vn
↓↓ ↓
kAu1k 0 ··· 0
0 kAu2k00
.........
00 ··· k Aunk
u1−→
u2−→
un−→
uj
=
v1v2 vn
↓↓ ↓
kAu
1k 0 ··· 0
0 kAu2k00
.........
00 ··· k Aunk
e
j
=
v1v2 vn
↓↓ ↓
kAujkej=kAujkvj=Auj
Thus both ˆVS U andAhave the same action on a basis. Therefore they are
equal.
It is easy to see that kAujk=p
λj, that is the singular values. While
A=ˆVS U could be called the singular value decomposition (SVD), what is
usually o ffered at the SVD is small modi fication of it. Rede fine the matrix
Uso that the firstrcolumns pertain to the nonzero singularvalues. De fine
Dto be the m×nmatrix consisting of the non zero singular values in the
194 CHAPTER 5. HERMITIAN THEORY
djjpositions, and filled in with zeros else where. De fine the matrix Vto be
thefirstrof the columns of ˆVand if r<m construct an additional m−r
orthonormal columns so that Vis an orthonormal basis of Cn.The resulting
product VD U ,c a l l e dt h e singular value decomposition ofA,i se q u a lt o
A,and moreover it follows that Vism×m, D ism×n,andUisn×n.
This gives the following theorem
Theorem 5.4.1. LetA∈Mmn.Then there is an m×morthogonal matrix
V,ann×northogonal matrix U,a n da n m×nmatrix Dwith only diagonal
entries such that A=VD U . The diagonal entries of Dare the singular
values of Aand the rows of Uare the eigenvectors of A∗A.
Example 5.4.1. The singular value decomposition can be used for image
compression. Here is the idea. Consider all the eigenvalues of A∗Aand
order them greatest to least. Zero the matrix Sfor all eigenvalues less than
some threshold. Then in the reconstruc tion and transmission of the matrix,
it is not necessary to include the vectors pertaining to these eigenvalues. Inthe example below, we have considered a 164 ×193 pixel image of C. F.
Gauss (1777-1855) on the postage stamp issued by Germany on Feb. 23,
1955, to commemorate the centenary of death. Therefore its spectrum has164 eigenvalues. The eigenvalues range from 26,603.0 to 1.895. A plotof the eigenvalues shown below. Now compress the image, retaining only afraction of the eigenvalues by e ffectively zeroing the smaller eigenvalues.
5.4. SINGULAR VALUE DECOMPOSITION 195
Note that the original image of the stamp has been enlarged and resampled
for more accurate comparisons. This image (stored at 221 dpi) is displayed
at effective 80 dpi with the enlargement. Below we show two plots where
we have retained respectively 30% and 10% of the eigenvalues. There is anapparent drastic decline in the image quality at roughly 10:1 compression.
In this image all eigenvalues smaller than λ= 860 have been zeroed.
Using 48 of 164 eigenvalues Using 16 of 164 eigenvalues
196 CHAPTER 5. HERMITIAN THEORY
5.5 Exercises
1. Prove Theorem 5.3.1 (i).
2. Prove Theorem 5.3.1 (ii).
3. Suppose that A, B∈Mn(C) are Hermitian. We will say A<0i fA
is non-negative de finite. Also, we say A<BifA−B<0. Is “ <”
an equivalence relation? If A<BandB<Cprove or disprove that
A<C.
4. Describe all Hermitian matrices of rank one.
5. Suppose that A, B∈Mn(C) are Hermitian and positive de finite. Find
necessary and su fficient conditions for ABto be Hermitian and also
positive de finite.
6. For any matrix A∈Mn(C) with eigenvalues λi,i=1,...,n .P r o v e
thatPn
i=1|λi|2=Pn
i,j=1|aij|2.
Chapter 6
Normal Matrices
Normal matrices are matrices that include Hermitian matrices and enjoy
several of the same properties as Hermitian matrices. Indeed, while we
proved that Hermitian matrices are uni tarily diagonalizable, we did not
establish any converse. That is, if a matrix is unitarily diagonalizable, thendoes it have any special property involving for example its spectrum or itsadjoint? As we shall see normal matrices are unitarily diagonalizable.
6.1 Introduction to Normal matrices
Definition 6.1.1. Am a t r i x A∈Mnis called normal ifA∗A=AA∗.
Proposition 6.1.1. A∈Mnis normal if and only if every matrix unitarily
equivalent to Ais normal.
Proof. Suppose Ais normal and B=U∗AU,w h e r e Uis unitary. Then
B∗B=U∗A∗AU=U∗AA∗U=U∗AUU∗A∗U=BB∗.I fU∗AUis normal
then it is easy to see that U∗AA∗U=U∗A∗AU. Multiply this equation on
the right by U∗a n do nt h el e f tb y Uto obtain AA∗=A∗A.
Examples.
(1) Unitary matrices are normal ( U∗U=I=UU∗).
(2) Hermitian matrices are normal ( AA∗=A2=A∗A).
(3) If A∗=−A,w eh a v e A∗A=AA∗=−A2. Hence matrices for which
A∗=−A,c a l l e d skew-Hermitian , are normal.
197
198 CHAPTER 6. NORMAL MATRICES
Example 6.1.1. Consider the arbitrary matrix N∈M2(R),written as
N=·ab
cd¸
. If we suppose that Nis normal then
N∗N=·ab
cd¸T·ab
cd¸
=·a2+c2ab+cd
ab+cd b2+d2¸
NN∗=·ab
cd¸·ab
cd¸T
=·a2+b2ac+bd
ac+bd c2+d2¸
From this we conclude that b2=c2,o rb=±c. Consider the cases in turn.
(i) If c=b,t h e n Nis Hermitian and thus normal.
(ii) If c=−b6=0,then ( N∗N)12=ab+cd=b(a−d). On the other
hand ( NN∗)12=ac+bd=(d−a)b.F o r b(a−d)=( d−a)b,w em u s t
have a=d.This gives that real 2 ×2 normal matrices are either symmetric
or have the form
N=·ab
−ba¸
Note this form includes both rotations and skew-symmetric matrices.
Recall the de finition of a unitarily diagonalizable matrix: A matrix A∈Mn
is called unitarily diagonalizable if there is a unitary matrix Ufor which
U∗AU is diagonal. A simple consequence of this is that if U∗AU =D
(where D= diagonal and U= unitary), then
AU=UD
and hence Ahasnorthonormal eigenvectors. This is just a part of the
spectral theorem for normal matrices .
Theorem 6.1.1 (Spectral theorem for normal matrices). IfA∈Mn
has eigenvalues λ1...λn, counted according to multiplicity, the following
statements are equivalent.
(a)Ais normal.
(b)Ais unitarily diagonalizable.
(c)Pn
i=1Pn
j=1|aij|2=Pn
j=1|λj|2.
(d) There is an orthonormal set of neigenvectors of A.
6.1. INTRODUCTION TO NORMAL MATRICES 199
Proof. (a)⇒(b). If Ais normal, then AA∗is Hermitian and therefore
unitarily diagonalizable. Thus U∗A∗AU=D=U∗AA∗U.A l s o , A,A∗,
A∗A=AA∗form a commuting family. This implies that eigenvectors of
A∗Aare also eigenvectors of A.S i n c e A∗Ahas a complete orthonormal set
we know that U∗AUis also diagonal. It is easy to see also that (b) ⇒(a)
We also note that (a) ⇒(d).
(b)⇒(c). Suppose U∗AU=D.T h e n U∗A∗U=D∗andU∗A∗AU=D∗D.
By Corollary 3.5.2 similarly preserves the trace. We know trace of A∗A
is tr ( A∗A)=Pn
j=1Pn
k=1a∗
jkakj=Pn
j=1Pn
k=1¯akjakj=Pn
j=1Pn
k=1|akj|2.
Since the trace of D∗DisΣ|λj|2, the result follows.
(c)⇒(b). We know that Ais unitarily equivalent to a upper triangular
matrix T. We also know that if A∼Bare unitarily equivalentP
ij|aij|2=
P
ij|bij|2. Application of this equality to the upper triangular matrix T yields
X
i,j|aij|2=X
|λj|2+X
j>i|tij|2=X
|λj|2.
Thus tij=0f o r j>i .T h u s Ais uniformly diagonalizable.
(d)⇒(b). Trivial.
Corollary 6.1.1. LetA∈MnandAis normal. If Uis unitary and if
U∗AU is upper triangular then U∗AU is diagonal.
Theorem 6.1.2. LetN∈Mn(R).T h e n Nis normal if and only if there
is a real orthogonal matrix Q∈Mn(R)such that
QTNQ=
A
1
A2 °
...
° An
(1)
where A
iis1×1(real) or Aiis2×2(real) of the form
Ai=·αiβj
−βjαi¸
.
Proof. First of all, any matrix Aof the form given by (1) is normal, and
therefore so also is any matrix unitarily similar (real orthogonally similar in
this case) to it.
200 CHAPTER 6. NORMAL MATRICES
To prove the converse we assume that N∈Mn(R)i sn o r m a l .W ek n o w
thatNis unitarily diagonalizable. That is, there is a unitary matrix Usuch
thatU∗NU=D, the diagonal matrix of its eigenvalues. Because Nis real,
all complex eigenvalues occur in compl ex conjugate pairs. Arrange them as
successive diagonal entries in D.I fλis a real eigenvalue, we can assume
without loss of generality that the corresponding eigenvector is real. For
complex eigenvalues, the corresponding eigenvectors also occur in conjugatepairs. Thus if α+iβis an eigenvector of Nwith corresponding eigenvector
written in real and complex parts u=u
r+ius.S i n c e Nis real we have
thatα−iβis also an eigenvector of Nwith corresponding eigenvector ¯ u=
ur−ius.B y t h e f a c t t h a t Nis unitarily diagonalible, these vectors are
orthogonal. This means hur,usi=0.
Replace the eigenvectors ur±iusby the real an imaginary parts in U.
This gives the matrix Q. Now compute QTNQ. It is easy to see that com-
puteNur=αur−βvsandNus=αus+βvr. When the first of these vectors
(αur−βvs) is multiplied by QTwe obtain the vector [0 ,...,α,−β,0,...0]T.
Multiplication by the second gives the vector [0 ,...,β,α,0,...0]T.I n t h i s
way the components·αβ
−βα¸
arise.
Corollary 6.1.2. (a)A∈Mnis symmetric if and only if (1) holds with all
blocks 1×1(and real).
(b)AAT=Iif and only if (1) has the form
λ1
...
λp ∗
A1
∗...
Ak
whereλ
j=±1andAj=hcosθj−sinθj
sinθjcosθji
θj∈R.
6.2 Exercises
1. If AandBcommute and if Ais normal, then A∗andBcommute.
Chapter 7
Factorization Theorems
This chapter highlights a few of the many factorization theorems for ma-
trices. While some factorization resul ts are relatively direct, others are it-
erative. While some factorization results serve to simplify the solution tolinear systems, others are concerned with revealing the matrix eigenvalues.We consider both types of results here.
7.1 The PLU Decomposition
The PLU decomposition (or factorization) To achieve LU factorization werequire a modi fied notion of the row reduced echelon form.
Definition 7.1.1. The modified row echelon form of a matrix is that form
which satis fies all the conditions of the modi fied row reduced echelon form
except that we do not require zeros to be above leading ones, and moreoverwe do not require leading ones, just nonzero entries.
For example the matrices below are in row echelon form.
A=
123
001000
B=
1 230
04−76
0 001
Most of the factorizations A∈M
n(C) studied so far require one essential
ingredient, namely the eigenvectors of A. While it was not emphasized when
we studied Gaussian elimination, there is a LU-type factorization there.
Assume for the moment that the only operations needed to carry Ato its
201
202 CHAPTER 7. FACTORIZATION THEOREMS
modi fied row echelon form are those that add a multiple of one row to
another. The modified row echelon form of a matrix is that form which
satisfies all the conditions of the modi fied row reduced echelon form except
that we do not require zeros to be above leading ones, and moreover wedo not require leading ones, just nonzero entries. Naturally it is easy to
make the leading nonzero entries into leading ones by the multiplication by
an appropriate identity matrix. That is not the point here. What wewant to observe is that in this case the reduction is accomplished by the leftmultiplication of Aby a sequence of lower triangular matrices of the form.
L=
1
01 0
...01
c...
0··· 1
Since we pivot at the (1 ,1)-entry first, we eliminate all the entries in the first
column below the first row. The product of all the matrices Lto accomplish
this has the form
L
1=
1
c
2110
c3101
......
cn10··· 1
where c
k1=−ak1
a11.Thus, with the notation that A=A1has entries a(1)
ijthis
first phase of the reduction renders the matrix A2with entries a(2)
ij
A2=L1A1=
a
(2)
11 ··· a(2)
1n
0a(2)
22 ···...
0a(2)
32a(2)33
.........
0a(2)
n2··· a(2)
nn
Since we have assumed that no row interchanges are necessary to carry out
the reduction we know that a
(2)
226=0.The next part of the reduction process
is the elimination of the elements in the second column below the second
7.1. THE PLU DECOMPOSITION 203
row, i.e. a(2)
32→0, ...a(2)
n2→0.Correspondingly, this can be achieved by
a matrix of the form
L2=
1
01 0
0c
22 1
.........
0cn2··· 1
(What are the values c
k2?) The result is the matrix A3given by
A3=L2A2=L2L1A1=
a
(3)
11 ··· a(3)
1n
0a(3)
22 ···...
00 a(3)
33...............
00 a
(3)
3n a(3)
nn
Proceeding in this way through all the rows (columns) there results
A
n=Ln−1An−1=Ln−1···L2L1A1=
a
(3)
11 ··· a(3)
1n
0a(3)
22 ···...
00 a(3)
33...............
000 a(3)
nn
The right side of the equation above is an upper triangular matrix. Denote
it by U.Since each of the matrices L
i,i=1,...n−1i s i n v e r t i b l e w e c a n
write
A=L−1
1···L−1
n−1U
The lemma below is useful in this.
Lemma 7.1.1. Suppose the lower triangular matrix L∈Mn(C)has the
204 CHAPTER 7. FACTORIZATION THEOREMS
form
L=
1
0... 0
1
01
......c
k+1,k...
......
0··· 0cnk 1
←−k
throw
Then Lis invertible with inverse given by
L−1=
1
0... 0
101
......−c
k+1,k...
......
0··· 0−cnk 1
←−k
throw
Proof. Trivial
Lemma 7.1.2. Suppose L1,L2,···,Ln−1are the matrices given above. Then
the matrix L=L−1
1···L−1
n−1has the form
L=
1
−c
21 10
−c31−c32 1
1
.........−ck+1,k...
......
−cn1−cn2···−cnk ··· 1
Proof. Trivial.
Applying these lemmas to the present situation we can say that when
no row interchanges are needed we can factor and matrix A∈M
n(C)a s
A=LU,where Lis lower triangular and Uis upper triangular. When row
7.1. THE PLU DECOMPOSITION 205
interchanges are needed and we let Pbe the permutation matrix that creates
these row interchanges then the LU-factorization above can be carried outfor the matrix PA. Thus PA=LU, where Lis lower triangular and Uis
upper triangular. We call this the PLU factorization. Let us summarize
this in the following theorem.
Theorem 7.1.1. LetA∈M
n(C). Then there is a permutation matrix
P∈Mn(C)and lower Land upper Utriangular matrices ( ∈Mn(C)), such
thatPA=LU. Moreover, Lcan be taken to have ones on its diagonal. That
is,`ii=1,i=1,...n .
By applying the result above to ATit is easy to see that the matrix U
can be taken to have the ones in its diagonal. The result is stated as a
corollary.
Corollary 7.1.1. LetA∈Mn(C). Then there is a permutation matrix
P∈Mn(C)and lower and upper triangular matrices ( ∈Mn(C)) respec-
tively, such that PA=LU. Moreover, Ucan be taken to have ones on its
diagonal ( uii=1,i=1,...n ).
The PLU decomposition can be put in service to solving the system
Ax=bas follows. Assume that A∈Mn(C) is invertible. Determine the
permutation matrix Pin order that PA=LU, where Lis lower triangular
andUis upper triangular. Thus, we have
Ax =b
PAx =Pb
LUx =Pb
Solve the systems
Ly =Pb
Ux =y
Then LUx =Ly=Pb.Hence xis a solution to the system. The advantages
of this formulation over the direct Gaussian elimination is that the systemsLy=PbandUx=yare triangular and hence are easy to solve. For example
for the first of the systems, Ly=Pb,let the vector Pb=h
ˆb
1,..., ˆbniT
.
Then it is easy to see that “back substitution” (aka “forward substitution”)
206 CHAPTER 7. FACTORIZATION THEOREMS
can be used to determine y. That is, we have the recursive relations
y1=ˆb1
l11
y2=ˆb2−l21y1
l22
...
yn=Ã
ˆbn−n−1X
m=1lnmym!
l−1
nn
A similar formula applies to solve Ux=y. I nt h i sc a s ew es o l v e first for
xn=yn/unn.The general formula is recursive with xkbeing determined
after xk+1,...,x n.are determined using the formula
xk=Ã
yk−nX
m=k+1ukmym!
u−1
kk
In practice the step of determining and then multiplying by the per-
mutation matrix is not actually carried out. Rather, an index array is
generated, while the elimination step is accomplished that e ffectively inter-
changes a “pointer” to the row interchanges. This saves considerable timein solving potentially very large systems.
More general and instructive methods are available for accomplishing
this LU factorization. Also, conditions are available for when no (nontrivial)permutation is required. We need the following lemma.
Lemma 7.1.3. LetA∈M
n(C)have the LU factorization A=LU,w h e r e
Lis lower triangular and Uis upper triangular. For any partition of the
matrix of the form
A=·A11A12
A21A22¸
there are corresponding decompositions of the matrices LandU
L=·L11 0
L21L22¸
and U=·U11U12
0U22¸
7.1. THE PLU DECOMPOSITION 207
where the Liiand the Uii.are lower and upper triangular respectively. More-
over, we have
A11=L11U11
A21=L21U11
A12=L12U22
A22=L21U12+L22U22
Thus L11U11is a LU factorization of A11.
With this lemma we can establish that almost every matrix can have a
LU factorization.
Definition 7.1.2. LetA∈Mn(C) and suppose that 1 ≤j≤n.T h e
expression det( A{1,...,j }) means the determininant of the upper left j×j
submatrix of A. These quaditities for j=1,...,n are called the principal
determinants of A.
Theorem 7.1.2. LetA∈Mn(C)and suppose that Ahas rank k.If
det(A{1,...,j })6=0 forj=1,...,k (1)
thenAhas a LU factorization A=LU,w h e r e Lis lower triangular and U
is upper triangular. Moreover, the factorization may be taken so that either
LorUis nonsingular. In the case k=nbothLandUwill be nonsingular.
Proof. We carry out this LU factorization as a direct calculation in compar-
ison to the Gaussian elimination method above. Let us propose to solve
the equation LU=Aexpressed as
l
11
l21l22 0
l31l32l33
............
...
ln1ln2··· ··· lnn
u
11u12u13··· u1n
u22u23··· u2n
u33
0......
...
unn
=
a
11a12a13 a1n
a21a22a23 a2n
a31a32a33
...............
...
an1an2··· ··· ann
208 CHAPTER 7. FACTORIZATION THEOREMS
It is easy to see that l11u11=a11.We can take, for example l11=1a n d
solve for u11.The detminant condition assures us that u116=0.Next solve
for the (2 ,1)-entry. We have l21u11=a21.Since u116=0,solve for l21.
For the (1 ,2)-entry we have l11u12=a12,w h i c hc a nb es o l v e df o r u12since
l116= 0. Finally, for the (2 ,2)-entry, l12u12+l22u22=a22is an equation
with two unknowns. Assign l22= 1 and solve for u22.What is important
to note is that the process carried out this way gives the factorization of theupper left 2 ×2 submatrix of A.Thus
·l
110
l21l22¸·u11u12
0u22¸
=·a11a12
a21a22¸
Since detµ·a11a12
a21a22¸¶
6=0,it follows that detµ·u11u12
0u22¸¶
6=0a n d
we know that·l110
l21l22¸
is nonsingular as the diagonal elements are ones.
Continue the factorization process through the k×kupper left submatrix
ofA.
Now consider the blocked matrix form form A
A=·A11A12
A21A22¸
where A11isk×kand has rank k. Thus we know that the rows of the lower
(n−k)×nmatrix above, that is£
A21A22¤
c a nb ew r i t t e na sau n i q u e
linear combination of the rows of the upper k×nmatrix£
A11A12¤
.Thus
£
A21A22¤
=C£
A11A12¤
for some ( n−k)×kmatrix C.Of course this means: A21=CA 11and
A22=CA 12. We consider the factorization
A=·A11A12
A21A22¸
=·L11 0
L21L22¸·U11U12
0U22¸
where the blocks L11andU11have just been determined. From the
equations in the lemma above we solve to get U12=L−1
11A12andL21=
7.2. LRLRLRFACTORIZATION 209
A12U−1
11.T h e n
A22=L21U12+L22U22
=A12U−1
11L−1
11A12+L22U22
=A12A−1
11A12+L22U22
=CA 11A−1
11A12+L22U22
=CA 12+L22U22
=A22+L22U22
Thus we solve L22U22=0.Obviously, we can take for L22any nonsingular
matrix we wish and solve for U22or conversely.
7.2LRLRLRfactorization
While the PLU factorization is useful for solving systems, the LR factoriza-
tion can be used to determine eigenvalues. .
LetA∈Mnbe given. Then
A=A1=L1R1.
Then
L−1
1A1L1=R1L1≡A2
A2=L2R2
L−1
2A2L2=R2L2≡A3.
Continue in this fashion to obtain
L−1
kAkLk=RkLk≡Ak+1 (?)
We de fine
Pk=L1L2...L k
Qk=Rk...R 2R1.
Then
PkAk+1=A1Pk
210 CHAPTER 7. FACTORIZATION THEOREMS
for
Ak+1=L−1
kAkLk
=L−1
kL−1
k−1Ak−1Lk−1Lk
...
=P−1
kA1Pk
or
PkAk+1=A1Pk.
Hence
PkQk=Pk−1AkQk−1
=A1Pk−1Qk−1
=A1Pk−2Ak−1Qk−2
=A2
1Pk−2Qk−2
...
=Ak
1.
Theorem 7.2.1 (Rutishauser). LetA∈Mnbe given. Assume the eigen-
values of Asatisfy
|λ1|>|λ2|>···>|λn|>0.
Then A∼Λ=diag(λ1...λn). Assume A=SΛS−1,a n d
Y≡S−1=LyRy X=S=LxRx
where LyandLxare lower unit triangular matrices and RyandRxare
upper triangular. Then Akdefined by (?)satisfy the result limAkis upper
triangular.
Proof. (Wilkinson) We have
Ak
1=XΛkY
=XΛkLyRy
=XΛkLyΛ−kΛkRy.
7.3. THE QRALGORITHM 211
By the strict inequalities between the eigenvalues we have
(ΛkLyΛ−k)ij=
1 i=j
µλi
λj¶k
`iji>j
0 i<j .
HenceΛkLyΛ−k→I(because|λi|
|λj|<1i fi>j ). Hence with
Ak
1=LxRx(ΛkLyΛ−k)ΛkRy
and
Ak
1=PkQk
we conclude that lim
k→∞Pk=Lx. Therefore
Lk=P−1
k−1Pk→I.
Finally we have that Akmust be upper triangular because
L−1
kAk=Rk
is upper triangular.
This exposes all the eigenvalues of A.T h e r e f o r e t h e e i g e n v e c t o r s c a n b e
determined.
7.3 The QRQRQRalgorithm
Certain numerical problems with the LUalgorithm have led to the QR
algorithm, which is based on the decomposition of the matrix Aas
A=QR
where Qis unitary and Ris upper triangular.
Theorem 7.3.1 (QR-factorization). (i) Suppose Ais inMn,mandn≥
m. Then there is a matrix Q∈Mn,mwith orthogonal columns and an
u p p e rt r i a n g u l a rm a t r i x
R∈Mmsuch that A=QR.
212 CHAPTER 7. FACTORIZATION THEOREMS
(ii) If n=m,t h e n Qis unitary. If Ais nonsingular the diagonal entries
ofRcan be chosen to be positive.
(iii) If Ais real; then QandRm a yb ec h o s e nt ob er e a l .
Proof. (i) We proceed inductively. Let a1,... , a ndenote the columns
ofAandq1,q2,... ,q mdenote the columns of Q. The basic idea of
the QR-factorization is to orthogonalize the columns of Afrom left
to right. Then the columns can be expressed by the formulas ak=Pk
i=1ckqk,k =1,...,n .T h e c o e fficients of the expansion become,
respectively, the entries of the kthcolumn of R,c o m p l e t e db y n−k
zeros. (Of course, if the rank of Ais less than m,w e fill in arbitrary
orthogonal vectors which we know exist as m≤n.) For the details,
first de fineq1=a1/ka1k. To compute q2we use the Gram—Schmidt
procedure.
ˆq2=a2−hq1,a1iq1
q2=ˆq2/kˆq2k.
Tracing backwards note that
a2=ˆq2+hq1,a1iq1
=kˆq2kq2+hq1,a1iq1.
So we have
·a1a2a3
↓↓↓ ...¸
=·q1q2q3
↓↓↓ ...¸
ka
1khq1,a1i...
0 kˆq2k
...0
00
.
Instead of the full inductive step we compute q
3andfinish at that
point
ˆq3=a3−hq1,a3iq1−hq2,a3iq2
q3=ˆq3/kˆq3k.
Hence
a3=kˆq3kq3+hq1,a3iq1+hq2,a3iq2.
7.3. THE QRALGORITHM 213
The third column of Ris thus given by
r3=[hq1,a3i,hq2,a3i,kˆq3k,0,0,... , 0]T.
In this way we see that the columns of Qare orthogonal and the matrix
Ris upper triangular, with an exception. That is the possibility that
ˆqk=0f o rs o m e k. In this degenerate case we take qkto be any
vector orthogonal to the span of a1,a2,... ,a m,a n dw et a k e rkj=0 ,
j=k,k+1...m . A l s ow en o t et h a ti fˆ qk=0 ,t h e n akis linearly
dependent on a1,a2,... ,a k−1, and hence on q1,q2,...q k−1. Select the
coefficients r1k,... ,r k−1kto reflect this dependence.
(ii) If m=n, the process above yields a unitary matrix. If Ais nonsingu-
lar, the process above yields a matrix Rwith a positive diagonal.
(iii) If Ais a real, all operators above can be carried out in real arithmetic.
Now what about the uniqueness of the decomposition? Essentially the
uniqueness is true up to a multiplication by a diagonal matrix, except inthe case when the matrix has rank is less than m, when there is no form of
uniqueness. Suppose that the rank of Aism.
Then application of the Gram-Schmidt procedure yields a matrix Rwith
positive diagonal. Suppose that Ahas two QR factorizations, QRandPS
with upper triangular factors having positive diagonals. Then
P
∗Q=SR−1
We have that SR−1is upper triangular and moreover has a positive diagonal.
Also, P∗Qis unitary. We know that the only upper triangular unitary
matrices are diagonal matrices, and finally the only unitary matrix with a
positive diagonal is the identity matrix. Therefore P∗Q=I,w h i c hi st o
say that P=Q.We summarize as
Corollary 7.3.1. Suppose Ais in Mn,mandn≥m.I f r a n k (A)=m
then the QR factorization of A=QRwith upper triangular matrix Rhaving
a positive diagonal is unique.
214 CHAPTER 7. FACTORIZATION THEOREMS
TheQRalgorithm
TheQRalgorithm parallels the LRalgorithm almost identically. Suppose
Ais inMnDefine
A1=Q1R1
A2≡R1Q1.
Also
Q∗
1A1Q=A2.
Then decompose A2into a QRdecomposition
A2=Q2R2
and
Q∗
2A2Q2=R2Q2≡A3.
Also
Q∗
2Q∗1A1Q1Q2=R2Q2=A3.
Proceed sequentially
Ak=QkRk
Ak+1=RkQk
Q∗
kAkQk=Ak+1.
Let
Pk=Q1Q2...Q k
Tk=RkRk−1...R 1.
Then
P∗
kA1Pk=Ak+1.
whence
PkAk+1=A1Pk.
7.3. THE QRALGORITHM 215
Also we have
PkTk=Pk−1QkRkTk−1
=Pk−1AkTk−1
=A1Pk−1Tk−1
=...
=Ak
1.
Theorem 7.3.2. LetA∈Mnbe given, and assume the eigenvalues of A
satisfy
|λ1|>|λ2|>···>|λn|>0.
Then the iterations Akconverge to a triangular matrix.
Proof. Our hypothesis gives that Ais diagonalizable, and we write A∼Λ=
diag(λ1...λn). That is,
A1=SΛS−1
whereΛ= diag(λ1...λn). Let
X=S=QxRx hereQR
Y=S−1=LyUyhereLU.
Then
Ak
1=QxRxΛkLyUy
=QxRxΛkLyΛ−kΛkUy
=Qx(I+RxEkR−1
x)RxΛkUY
where
Ek=ΛkLyΛ−k−I
(Ek)ij=
0 i=j
(λi/λj)k`iji>j
0 i<j .
It follows that I+RxEkR−1
x→I,a n d RxΛ−kUyis upper triangular. Thus
Qx(I+RxEkR−1
x)RxΛkUy=PkTk.
216 CHAPTER 7. FACTORIZATION THEOREMS
The matrix I+RxEkR−1
xcan be QR factored as ˜Uk˜Rk, and since I+
RxEkR−1
x→I, it follows that we can assume both ˜Uk→Iand ˜Rk→I.
Hence
Ak
1=Qx˜Uk[˜Rk(I+RxEkR−1
x)RxΛkUy]=PkTk.
with the first factor unitary and the second factor upper triangular. Since
we have assumed (by the eigenvalue condition) that Ais nonsingular, this
factorization is essentially unique, where possibly a multiplication by a di-
agonal matrix must be applied to give the upper triangular factor on theright a positive diagonal. Just what is the form of the diagonal matrix canbe seen from the following. Let Λ=|Λ|Λ
1,w h e r e |Λ|is the diagonal matrix
of moduli of the elements of Λand where Λ1is the unitary matrix of the
signs of each eigenvalue respectively. We also take Uy=Λ2(Λ∗
2Uy)w h e r e
Λ2is a unitary matrix chosen so that Λ∗
2Uhas a positive diagonal. Then
Ak
1=Qx˜UkΛ2Λk
1[³
Λ2Λk
1´−1˜Rk(I+RxEkR−1
x)Rx³
Λ2Λk
1´
|Λ|k(Λ∗
2Uy)] =PkTk.
From this we obtain Pkis essentially asymptotic to Qx˜UkΛ2Λk
1and from
this we obtain that
Qk=P−1
k−1Pk→Λ1
which is diagonal. Finally, it follows that Akis upper triangular since
Q−1
kAk=Rk
In the limit therefore Ais similar to an upper triangular matrix.
Example 7.3.1. Apply the QR method to the matrix
A:=
2.31 2
22 2 .1
320
The matrix Ahas eigvenvalues 5 .45,0.723,−1.87. The successive iterations
are
7.4. LEAST SQUARES 217
A2=
5.10−0.511 2 .13
0.631 0 .662 0 .136
1.42−0.0202−1.44
A3=
5.51−1.02−0.36
−0.0146 0 .666 0 .482
0.513 0 .240−1.84
A4=
5.46−1.41 0 .482
−0.0372 0 .495 0 .672
0.169 0 .815−1.62
A5=
5.47−0.366−1.26
−0.0404−0.462 1 .39
0.0430 1 .21−0.677
A6=
5.46−1.13−0.687
−0.0184−1.52 0 .813
0.00826 0 .983 0 .381
A7=
5.45 0 .529−1.18
−0.00682−1.78 0 .585
0.00115 0 .414 0 .638
A8=
5.43 0 .684−1.09
−0.000822 −1.87 0 .229
0.0000215 0 .0659 0 .729
Note the gradual appearance of the eigenvalues on the diagonal.
Remark. These iterations were carried out in precision 3 arithmetic, whichaffects the rate of convergence to triangular form.
7.4 Least Squares
As we know, if A∈Mn,mwith m<n it is generally not possible to solve
the overdetermined system
Ax=b.
For example, suppose we have the data {(xi,yi)}n
i=1,w i t ht h e x-coordinates
distinct. We may wish to “ fit” a straight to this data. This means we want
tofind coefficients mandbso that
b+mxi=yi,i =1,... ,n . (?)
Taking the matrix and data vector
A=
1x
1
1x2
...
1xn
b=
y
1
y2
...
yn
andz=[b, m]T, the system ( ?) becomes Az=b.U s u a l l y nÀ2. Hence
there is virtually no hope to determine a unique solution to system.
However, there are numerous ways to determine constants mandbso
that the resulting line represents the data. For example, owing to the dis-
tinctness of the x-coordinates, it is possible to solve any 2 ×2 subsystem of
218 CHAPTER 7. FACTORIZATION THEOREMS
Az=b. Other variations exist. A new 2 ×2 system could be created by
creating two averages of the data, say left and right, and solving. Assume
the sequence {xj}is ordered from least to greatest. De finex`=1
kkP
j=1xjand
xr=1
n−knP
j=k+1xj.L e t y`andyrdenote the corresponding averages for the
ordinates. Then de fine the intercept band slope mby solving the system
·1x`
1xr¸·b
m¸
=·y`
yr¸
While this will normally give a reasonable approximating line, its value has
little utility beyond its naive simplicity and visual appearance. What is
desired is to establish a criteria for choosing the line.
Define the residual of the approximation r=b−Az.I tm a k e sp e r f e c t
sense to consider finding z=[b, m]Tfor which the residual is minimized in
some norm. Any norm can be selected here, but on practical grounds thebest norm to use is the Euclidean norm k·k
2.T h e v e c t o r Azthat yields
the minimal norm residual is the one for which ( b−Az)⊥Aw, for we are
seeking the nearest value in the Awto the vector b. It can be found by
select the one for the solution, Az,f o rw h i c h
b−Ax⊥Aw allw.
This means
hb−Ax, Ay i=0 a l l y
or
hAT(b−Ay),yi=0 a l l y
or
AT(b−Ay)=0
ATAy=ATb.Normal
Equations
Theleast squares solution to Ax=bis given by the solution to the normal
equation
ATAy=ATb.
7.5. EXERCISES 219
Suppose we have the QRdecomposition for A.T h e ni f Ais real
ATA=RTQTQR=RTR
ATy=RTQy.
Hence the normal equations become
RTRx=RTQy.
Assuming that the rank of Aism,w em u s th a v et h a t Rand hence RT
is invertible. Therefore we have the least squares solution is given by the
triangular system
Rx=Qy.
7.5 Exercises
1. If A∈M(C)h a sr a n k k, show that there is a permutation matrix P
such that PAhas its firstkprincipal determinants nonzero.
2. For the least squares fit of a straight line determine RandQ.
3. In the case of data
ATA=·nΣxi
ΣxiΣx2
i¸
ATb=·Σyi
Σxiyi¸
.
4. In attempting to solve a quadratic fitw eh a v et h em o d e l
c+bxi+ax2
i=yi i=1,... ,n .
The system is
A=
1x1x2
1.........
1xnx2
n
b=
y1
y2
...
yn
.
The normal equations have the matrix and data given by
ATA=
nΣxiΣx2
i
ΣxiΣx2
iΣx3
i
Σx2
iΣx2
iΣx4
i
ATb=
Σyi
Σxiyi
Σx2
iyi
.
5. Find the normal equations for the least squares fito fd a t at oap o l y -
nomial of degree k.
Chapter 8
Jordan Normal Form
8.1 Minimal Polynomials
Recall pA(x)=d e t ( xI−A) is called the characteristic polynomial of the
matrix A.
Theorem 8.1.1. LetA∈Mn. Then there exists a unique monic polyno-
mial qA(x)of minimum degree for which qA(A)=0 .I fp(x)is any polyno-
mial such that p(A)=0 ,t h e n qA(x)divides p(x).
Proof. Since there is a polynomial pA(x)f o rw h i c h pA(A) = 0, there is one of
minimal degree, which we can assume is monic. by the Euclidean algorithm
pA(x)=qA(x)h(x)+r(x)
where deg r(x)<degqA(x). We know
pA(A)=qA(A)h(A)+r(A).
Hence r(A) = 0, and by the minimality assumption r(x)≡0. Thus qA
divided pA(x) and also any polynomial for which p(A) = 0. to establish
that qAis unique, suppose q(x) is another monic polynomial of the same
degree for which q(A)=0 . T h e n
r(x)=q(x)−qA(x)
is a polynomial of degree less than qA(x)f o rw h i c h r(a)=q(A)−qA(A)=0 .
This cannot be unless r(x)≡0=0 q=qA.
Definition 8.1.1. The polynomial qA(x) in the theorem above is called the
minimal polynomial.
221
222 CHAPTER 8. JORDAN NORMAL FORM
Corollary 8.1.1. IfA, B∈Mnare similar, then they have the same min-
imal polynomial.
Proof.
B=S−1AS
qA(B)=qA(S−1AS)=S−1qA(A)S=qA(A)=0.
If there is a minimal polynomial for Bof smaller degree, say qB(x), then
qB(A) = 0 by the same argument. This contradicts the minimality of qA(x).
Now that we have a minimum polynomial for any matrix, can we find a
matrix with a given polynomial as its minimum polynomial? Can the degreethe polynomial and the size of the matrix match? The answers to both
questions are a ffirmative and presented below in one theorem.
Theorem 8.1.2. For any n
thdegree polynomial
p(x)=xn+an−1xn−1+an−2xn−2+···+a1x+a0
there is a matrix A∈Mn(C)for which it is the minimal polynomial.
Proof. Consider the matrix given by
A=
00 ... ... −a
0
10 −a1
01 0...−a2
... 0...
0... 01−an−1
.
Observe that
Ie
1=e1=A0e1
Ae1=e2=Ae1
Ae2=e3=A2e1
...
Aen−1=en=An−1e1
8.1. MINIMAL POLYNOMIALS 223
and
Aen=−an−1en−an−2en−1−···−a1e2−a0e1
Since Aen=Ane1, it follows that
p(A)e1=Ane1+an−1An−1e1+an−2A−2e1+···+a1Ae1+a0Ie1=0
Also
p(A)ek=p(A)Ak−1e1=Ak−1p(A)e1=Ak−1(0) = 0 k=2,... ,n .
Hence p(A)ej=0 f o r j=1...n.T h u s p(A) = 0. We know also that
p(x) is monic. Suppose now that
q(x)=xm+bm−1xm−1···+b1x+b0
where m<n andq(A)=0 . T h e n
q(A)e1=Ame1+bm−1Am−1e1+···+b1Ae1+b0e1
=em+1+bm−1em+···+b1e2+b0e1=0.
But the vectors em+1...e 1are linear independent from which we conclude
thatq(A) = 0 is impossible. Thus p(x)i sm i n i m a l .
Definition 8.1.2. For a given monic polynomial p(x), the matrix Acon-
structed above is called the companion matrix to p.
The transpose of the companion matrix can also be used to generate a linear
differential system which has the same characteristic polynomial as a given
nthorder differential equation. Consider the linear di fferential equation
y(n)+an−1y(n−1)+···+a1y0+a0=0.
This nthorder ODE can be converted to a first order system as follows:
u1=y
u2=u0
1 =y0
u3=u0
2 =y00
......
un=u0
n−1=y(n−1)
224 CHAPTER 8. JORDAN NORMAL FORM
Then we have
u
1
u2
...
un
0
=
01
01 0
01
......
0... 1
a
0−a1 ... ... −an−1
u
1
u2
...
un
8.2 Invariant subspaces
There seems to be no truly simple way to the Jordan normal form. The
approach taken here is intended to reveal a number of features of a matrix,interesting in their own right. In particular, we will construct “generalized
eigenspaces” that envelop the entire connection of a matrix with its eigen-
values. We have in various ways considered subspaces VofC
nthat are
invariant under the matrix A∈Mn(C). Recall this means that AV⊂V.
For example, eigenvectors can be used to create invariant subspaces. Nullspaces, the eigenspace of the zero eigenvalue, are invariant as well. Triangu-lar matrices furnish an easily recognizable sequence of invariant subspaces.
Assuming T∈M
n(C) is upper triangular, it is easy to see that the sub-
spaces generated by the coordinate vectors {e1,...,e m}form=1,...,n are
invariant under T.
We now consider a speci fic type of invariant subspace that will lead the
so-called Jordan normal form of a ma trix, the closest matrix similar to A
that resembles a diagonal matrix.
Definition 8.2.1 (Generalized Eigenspace). LetA∈Mn(C)w i t hs p e c -
trumσ(A)={λ1,...,λk}.D e fine the generalized eigenspace pertaining to
λiby
Vλi={x∈Cn|(A−λiI)nx=0}
Observe that all the eigenvectors pertaining to λiare contained in Vλi.
If the span of the eigenvectors pertaining to λiis not equal to Vλithen there
must be a positive power pand a vector xsuch that ( A−λiI)px= 0 but that
y=(A−λiI)p−1x6=0 . T h u s yis an eigenvector pertaining to λi.F o rt h i s
reason we will call Vλithe space of generalized eigenvectors pertaining to
λi.O u r first result, that Vλiis invariant under A,i ss i m p l et op r o v e ,n o t i n g
that only closure under vector addition and scalar multiplication need be
established.
8.2. INVARIANT SUBSPACES 225
Theorem 8.2.1. LetA∈Mn(C)with spectrum σ(A)= {λ1,...,λk}.
Then for each i=1,...,k ,Vλiis an invariant subspace of A.
One might question as to whether Vλicould be enlarged by allowing
higher powers than nin the de finition. The negative answer is most simply
expressed by evoking the Hamilton-Cayley theorem. We write the charac-
teristic polynomial pA(λ)=Q(λ−λi)mA(λi),w h e r e mA(λi) is the algebraic
multiplicity of λi.S i n c e pA(A)=Q(A−λiI)mA(λi)= 0, it is an easy mat-
ter to see that we exhaust all of Cnwith the spaces Vλi.T h i s i s t o s a y t h a t
allowing higher powers in the de finition will not increase the subspaces Vλi.
Indeed, as we shall see, the power of ( λ−λi) can be decreased to the geomet-
ric multiplicity mg(λi)t h ep o w e rf o r λi. For now the general power nwill
suffice. One very important result, and an essential fir s ts t e pi nd e r i v i n g
the Jordan form, is to establish that any square matrix Ais similar to a
block diagonal matrix, with each block carrying a single eigenvalue.
Theorem 8.2.2. LetA∈Mn(C)with spectrum σ(A)= {λ1,...,λk}and
with invariant subspaces Vλi,i=1,2,...,k .T h e n ( i ) T h e s p a c e s Vλi,j=
1,...,k are mutually linearly independent. (ii)Lk
i=1Vλi=Cn(alternatively
Cn=S(Vλ1,...,V λk)) (iii) dimVλi=mA(λi).( i v ) Ais similar to a block
diagonal matrix with kblocks A1,...,A k.M o r e o v e r , σ(Ai)= {λi}and
dimAi=mA(λi).
Proof. (i) It should be clear that the subspaces Vλiare linearly independent
of each other. For if there is a vector xin both VλiandVλjthen there is
a vector for some integer q,it must be true that ( A−λjI)q−1x6= 0 but
(A−λjI)qx=0.This means that y=(A−λjI)q−1xis an eigenvector
pertaining to λj.S i n c e ( A−λiI)nx= 0 we must also have that
(A−λjI)q−1(A−λiI)nx=(A−λiI)n(A−λjI)q−1x
=(A−λiI)ny=0
=nX
k=0µn
k¶
(−λi)n−kAky
=nX
k=0µn
k¶
(−λi)n−kλjky
=(λj−λi)ny=0
This is impossible unless λj=λi. (ii) The key part of the proof is to block
diagonalize Awith respect to these invariant subspaces. To that end, let S
226 CHAPTER 8. JORDAN NORMAL FORM
be the matrix with columns generated from bases of the individual Vλitaken
in the order of the indices. Supposing there are more linearly independentvectors in C
nother than those already selected, fill out the matrix Swith
vectors linearly independent to the subspaces Vλi,i=1,...,k .N o w d e fine
˜A=S−1AS. We conclude by the invariance of the subspaces and their
mutual linear independence that ˜Ahas the following block structure.
˜A=S−1AS=
A
10 ··· 0∗
0 A2 0∗
.........
Ak∗
0 ··· 0 B
It follows that
p
A(λ)=p˜A(λ)=³Y
pAi(λ)´
pB(λ)
Any root rofpB(λ) must be an eigenvalue of A,sayλj,and there must be an
eigenvector xpertaining to λj. Moreover, due to the block structure we can
assume that x=[ 0,..., 0,x]T,where there are kblocked zeros of the sizes of
theAirespectively. Then it is easy to see that ASx =λjSx, and this implies
thatSx∈Vλj. Thus there is another vector in Vλj, which contradicts its
definition. Therefore Bis null, or what is the same thing, ⊕k
i=1Vλi=Cn.
(iii) Let di=d i m Vλi.F r o m ( i i ) k n o wPdi=n. Suppose that λi∈σ(Aj).
Then there is another eigenvector xpertaining to λiand for which Ajx=
λix.Moreover, this vector has the form x=[ 0,..., 0,x ,0,...0]T, analogous
to the argument above. By construction Sx /∈Vλi, but ASx =λiSx,and
this contradicts the de finition of Vλi. W et h u sh a v et h a t pAi(λ)=(λ−λi)di.
Since pA(λ)=QpAi(λ)=Q(λ−λi)mA(λi), it follows that di=mA(λi)
(iv) Putting (ii), and (iii) together gives the block diagonal structure as
required.
On account of the mutual linear independence of the invariant subspacesV
λiand the fact that they exhaust Cnthe following corollary is immediate.
Corollary 8.2.1. LetA∈Mn(C)with spectrum σ(A)=λ1,...,λkand
with generalized eigenspaces Vλi,i=1,2,...,k .T h e n e a c h x∈Cnhas a
unique representation x=Pk
i=1xiwhere xi∈Vλi.
Another interesting result which reveals how the matrix works as a linear
transformation is to decompose the it into components with respect to the
8.2. INVARIANT SUBSPACES 227
generalized eigenspaces. In particular, viewing the block diagonal form
˜A=S−1AS=
A10 ··· 0
0 A2 0
......
Ak
the space Cncan be split into a direct sum of subspaces E1,...,E kbased on
coordinate blocks. This is accomplished in such that any vector y∈Cncan
be written uniquely as y=Pk
i=1yiwhere the yi∈Ei. (Keep in mind that
eachyi∈Cn; its coordinates are zero outside the coordinate block pertaining
Ei.) Then ˜Ay=˜APk
i=1yi=Pk
i=1˜Ayi=Pk
i=1Aiyi.This provides a
computational tool – when this block diagonal form is known. Note that
the blocks correspond directly to the invariant subspaces by SEi=Vλi.W e
can use these invariant subspaces to get at the minimal polynomial. Foreach i=1,...,k define
m
i=m i n
j{(A−λiI)jx=0 |x∈Vλi}
Theorem 8.2.3. LetA∈Mn(C)with spectrum σ(A)=λ1,...,λkand
with invariant subspaces Vλi,i=1,2,...,k . Then the minimal polynomial
ofAis given by
q(λ)=kY
i=1(λ−λi)mi
Proof. Certainly we see that for any vector x∈Vλj
q(A)x=ÃkY
i=1(A−λiI)mi!
x=0
Hence, the minimal polynomial qA(λ)d i v i d e s q(A). To see that indeed
they are in fact equal, suppose that the minimal polynomial has the form
qA(λ)=kY
i=1(λ−λi)ˆmi
where ˆ mi≤mi,fori=1,...,k a n di np a r t i c u l a r ˆ mj<m j.B y c o n s t r u c t i o n
there must exist a vector x∈Vλjsuch that ( A−λjI)mjx= 0 but y=
228 CHAPTER 8. JORDAN NORMAL FORM
(A−λjI)mj−1x6=0.Then if
q(A)x=ÃkY
i=1(A−λiI)ˆmi!
x
=
kY
i=1
i6=j(A−λiI)ˆmi
y
=0
This cannot be because the contrary implies that there is another vector in
one of the invariant subspaces Vλk.
Just one more step is needed before the Jordan normal form can be derived.
For a given Vλiwe can interpret the spaces in a heirarchical viewpoint. We
know that Vλicontains all the eigenvectors pertaining to λi.C a l l t h e s e
eigenvectors the first order generalized eigenvectors . If the span of these
is not equal to Vλi, then there must be a vector x∈Vλifor which y=
(A−λiI)2x= 0 but ( A−λiI)x6=0 . T h a ti st os a y yis an eigenvector of
Apertaining to λi. Call such vectors second order generalized eigenvectors .
In general we call an x∈Vλia generalized eigenvector of order jify=
(A−λiI)jx= 0 but ( A−λiI)j−1x6= 0. In light of our previous discussion
Vλicontains generalized eigenvectors of order up to but not greater than
mλi.
Theorem 8.2.4. LetA∈Mn(C)with spectrum σ(A)= {λ1,...,λk}and
with invariant subspaces Vλi,i=1,2,...,k .
(i) Let x∈Vλibe a generalized eigenvector of order p. Then the vectors
x,(A−λiI)x,(A−λiI)2x ,..., (A−λiI)p−1x (1)
are linearly independent.
(ii) The subspace of Cngenerated by the vectors in (1) is an invariant
subspace of A.
Proof. (i) To prove linear independence of a set of vectors we suppose linear
dependence. That is there is a smallest integer kand constants bjsuch that
kX
j=0xj=kX
j=0bj(A−λiI)jx=0
8.2. INVARIANT SUBSPACES 229
where bk6=0.Solving we obtain bk(A−λiI)kx=−Pk−1
j=0bj(A−λiI)jx.
Now apply ( A−λiI)p−kto both sides and obtain a new linearly dependent
set as the following calculation shows.
0= bk(A−λiI)p−k+kx=−k−1X
j=0bj(A−λiI)j+p−kx
=−p−1X
j=p−kbj+p−k(A−λiI)jx
T h ek e yp o i n tt on o t eh e r ei st h a tt h el o w e rl i m i to ft h es u mi si n c r e a s e d .
This new linearly dependent set, which we denote with the notationPp−1
j=p−kcj(A−λiI)jx
can be split in the same way as before, where we assume with no loss in gen-erality that c
p−16=0 . T h e n
cp−1(A−λiI)p−1x=−p−2X
j=p−kcj(A−λiI)jx
Apply ( A−λiI) to both sides to get
0= cp−1(A−λiI)px=−p−2X
j=p−kcj(A−λiI)j+1x
=−p−1X
j=p−k+1cj−1(A−λiI)jx
Thus we have obtained another linearly independent set with the lower limit
of powers increased by one. Continue this process until the linear depen-dence of ( A−λ
iI)p−1xand ( A−λiI)p−2xis achieved. Thus we have
c(A−λiI)p−1x=d(A−λiI)p−2x
(A−λiI)y=d
cy
where y=(A−λiI)p−2x.T h i s i m p l i e s t h a t λi+d
cis a new eigenvalue
with eigenvector y∈Vλi, and of course this is a contradiction. (ii) The
invariance under Ais more straightforward. First note that while x1=x,
230 CHAPTER 8. JORDAN NORMAL FORM
x2=(A−λI)x=Ax−λxso that Ax=x2−λx1.Consider any vector y
defined by y=Pp−1
j=0bj(A−λiI)jxIt follows that
Ay =Ap−1X
j=0bj(A−λiI)jx
=p−1X
j=0bj(A−λiI)jAx
=p−1X
j=0bj(A−λiI)j(x2−λx1)
=p−1X
j=0bj(A−λiI)j[(A−λiI)x1−λx1]
=p−1X
j=0cj(A−λiI)jx
where cj=bj−1−λforj>0a n d c0=−λ, which proves the result.
8.3 The Jordan Normal Form
We need a lemma that points in the direction we are headed, that being the
use of invariant subspaces as a basis for the (Jordan) block diagonalization
of any matrix. These results were discussed in detail in the Section 8.2. Therestatement here illustrates the “invariant subspace” nature of the result,irrespective of generalized eigenspac es. Its proof is elementary and is left to
the reader.
Lemma 8.3.1. LetA∈M
n(C)with invariant subspace V⊂Cn.
(i) Suppose v1...v kis a basis for VandSis an invertible matrix with
thefirstkcolumns given by v1...v k.T h e n
1...k
S−1AS=·∗∗
0∗¸
.
8.3. THE JORDAN NORMAL FORM 231
(ii) Suppose that V1,V2⊂Cnare two invariant subspaces of AandCn=
V1⊕V2. Let the (invertible) matrix Sconsist respectively of bases from V1
andV2as its columns. Then
S−1AS=·∗0
0∗¸
.
Definition 8.3.1. Letλ∈C.A Jordan block Jk(λ)i sa k×kupper
triangular matrix of the form
Jk(λ)=
λ1
0
λ1
0...1
λ
.
AJordan matrix is any matrix of the form
J=
J
n1(λ1)0
...
0 Jnk(λk)
.
where the matrices Jn1are Jordan blocks. If J∈Mn(C), then n1+n2···+
nk=n.
Theorem 8.3.1 (Jordan normal form). LetA∈Mn(C). Then there is
a nonsingular matrix S∈Mnsuch that
A=S
J
n1(λ1)
0
...
0
Jnk(λk)
S
−1=SJS−1
where Jni(λi)is a Jordan block, where n1+n2+···+nk=n.Jis unique up
to permutations of the blocks. The eigenvalues λ1,... ,λkare not necessarily
distinct. If Ais real with real eigenvalues, then Scan be taken as real.
Proof. This result is proved in four steps.
(1) Block diagonalize (by similarity) into invariant subspaces pertaining to
σ(A). This is accomplished as follows. First block diagonalize the ma-
trix according to the generalized eigenspaces Vλi={x∈Cn|(A−λiI)nx=
232 CHAPTER 8. JORDAN NORMAL FORM
0}as discussed in the previous section. Beginning with the highest or-
der eigenvector in x∈Vλi, construct the invariant subspace as in (1)
of Section 8.2. Repeat this process until all generalized eigenvectorshave been included in an invariant subspace. This includes of coursefirst order eigenvectors that are not a ssociated with higher order eigen-
vectors. These invariant subspaces have dimension one. Each of these
invariant subspaces is linearly independent from the others. Continuethis process for all the generalized eigenspaces. This exhausts C
n.
Each of the blocks contains exactly one eigenvector. The dimensionsof these invariant subspaces can range from one to m
λi,t h e r eb e i n ga t
least one subspace of dimension mλi.
(2) Triangularize each block by Schur’s theorem, so that each block has
the form
K(λ)=
λ∗
...
0λ
You will note that
K(λ)=λI+N
where Nis nilpotent, or K(λ)i s1 ×1.
(3) “Jordanize” each triangular block. Assume that K1(λ)i sm×m,w h e r e
m> 1. By construction K1(λ) pertains to an invariant subspace for
which there is a unique vector xfor which
Nm−1x6=0 a n d Nmx=0.
Thus Nm−1xis an eigenvector of K1(λ), the unique eigenvector. De fine
yi=Ni−1xi =1,2,... ,m .
Expand the set {yi}m
i=1as a basis of Cm.D efine
S1="
ymym−1···y1
...... ···...#
.
Then
NS 1="
0ymym−1... y 2
......... ···...#
.
8.3. THE JORDAN NORMAL FORM 233
So
S−1
1NS 1=
01 0
01
0...
...1
00
.
We conclude that
S
−1
1K1(λ)S1=
λ10
λ1
λ1
......
...1
0 λ
(4) Assemble all of the blocks to form the Jordan form. For example,
the block K
1(λ) and the corresponding similarity transformation S1
studied above can be treated in the assembly process as follows: De fine
then×nmatrix
ˆS1=
I00
0S10
00 I
where S1i st h eb l o c kc o n s t r u c t e da b o v ea n dp l a c e di nt h e n×nmatrix
in the position that K1(λ) was extracted from the block triangular
form of A. Repeat this for each of the blocks pertaining to minimally
invariant subspaces. This gives a sequence of block diagonal matrices
ˆS1,ˆS2..., ˆSk.D efineT=ˆS1ˆS2...ˆSk. It has the form
T=
ˆS
1 0
ˆS2
...
0 ˆSk
Together with the original matrix Pthat transformed the matrix to
the minimal invariant subspace blocked form and the unitary matrix
234 CHAPTER 8. JORDAN NORMAL FORM
Vused to triangularize A, it follows that
A=PVT
J
n1(λ1)
0
...
0
Jnk(λk)
(PVT )
−1=SJS−1
with S=PVT .
Example 8.3.1. Let
J=
2
1
02
2
310
031
003
−1
In this example, there are four blocks, with two of the blocks pertaining to
the single eigenvalue 2. For the first block there is the single eigenvector e
1,
but the invariant subspace is S(e1,e2).For the second block, the eigenvector,
e3,generates the one dimensional invariant subspace. The block pertaining
to the eigenvector 3 has the single eigenvector e4while the minimal invariant
subspace is S(e4,e5,e6). Finally, the one dimensional subspace pertaining
to the eigenvector −1 is spanned by e7.The minimal invariant polynomial
isq(λ)=(λ−2)2(λ−3)3(λ+1 ) .
8.4 Convergent matrices
Using the Jordan normal form, the study of convergent matrices becomes
relatively straightforward and simpler.
Theorem 8.4.1. IfA∈Mnandρ(A)<1.T h e n
lim
k→∞A=0.
8.5. EXERCISES 235
Proof. We assume Ais a Jordan matrix. Each Jordan block Jk(λ)c a nb e
written as
Jk(λ)=λIk+Nk
where
Nk=
01 0
01
......
...1
00
is nilpotent.
Now
A=
J
n1(λ1)
...
Jmk(λk)
.
We compute, for m>n k
(Jnk(λk))M=(λI+N)m
=λmI+mkX
j=0λm−jNjµm
j¶
because Nj=0f o r j>n k.W eh a v e
λm−jµm
j¶
→0a sm→∞
since |λ|<1. The results follows.
8.5 Exercises
1. Prove Theorem 8.2.1.
2. Find a 3 ×3 matrix that has the same eigenvalues are the squares of
the roots of the equation λ3−3λ2+4λ−5=0 .
3. Suppose that Ais a square matrix with σ(A)= {3},ma(3) = 6, and
mg(3) = 3. Up to permutations of the blocks show all possible Jordan
normal forms for A.
236 CHAPTER 8. JORDAN NORMAL FORM
4. Let A∈Mn(C)a n dl e t x1∈Cn.Definexi+1=Axifori=1,...,n−1.
Show that V=S({x1,...,x n}) is an invariant subspace of A. Show
thatVcontains an eigenvector of A.
5. Referring to the previous problem, let A∈Mn(R)b eap e r m u t a t i o n
matrix. (i) Find starting vectors so that dim V=n.( i i ) F i n d s t a r t -
ing vectors so that dim V= 1. (iii) Show that if λ= 1 is a simple
eigenvalue of Athen dim V=1o rd i m V=n.
Chapter 9
Hermitian and Symmetric
Matrices
Example 9.0.1. Letf:D→R,D⊂Rn.T h e Hessian is defined by
H(x)=hij(x)≡∂f
∂xi∂xj∈Mn.
Since for functions f∈C2it is known that
∂2f
∂xi∂xj=∂2f
∂xj∂xi
it follows that H(x)i ss y m m e t r i c .
Definition 9.0.1. A function f:R→Risconvex if
f(λx+( 1−λ)y)≤λf(x)+( 1−λ)f(y)
forx, y∈D(domain) and 0 ≤λ≤1.
Proposition 9.0.1. Iff∈C2(D)andf00(x)≥0onDthenf(x)is convex.
Proof. Because f00≥0, this implies that f0(x) is increasing. Therefore if
x<x m<ywe must have
f(xm)≤f(x)+f0(xm)(xm−x)
and
f(xm)≤f(y)+f0(xm)(xm−y)
237
Chapter 9
Hermitian and Symmetric
Matrices
Example 9.0.1. Letf:D→R,D⊂Rn.T h e Hessian is defined by
H(x)=hij(x)≡∂f
∂xi∂xj∈Mn.
Since for functions f∈C2it is known that
∂2f
∂xi∂xj=∂2f
∂xj∂xi
it follows that H(x)i ss y m m e t r i c .
Definition 9.0.1. A function f:R→Risconvex if
f(λx+( 1−λ)y)≤λf(x)+( 1−λ)f(y)
forx, y∈D(domain) and 0 ≤λ≤1.
Proposition 9.0.1. Iff∈C2(D)andf00(x)≥0onDthenf(x)is convex.
Proof. Because f00≥0, this implies that f0(x) is increasing. Therefore if
x<x m<ywe must have
f(xm)≤f(x)+f0(xm)(xm−x)
and
f(xm)≤f(y)+f0(xm)(xm−y)
237
238 CHAPTER 9. HERMITIAN AND SYMMETRIC MATRICES
by the Mean Value Theorem. Therefore,
y−xm
y−xf(xm)+xm−x
y−xf(xm)≤y−xm
y−xf(x)+xm−x
y−xf(y)
or
f(xm)≤y−xm
y−xf(x)+(xm−x)
y−xf(y).
It is easy to see that
xm=y−xm
y−xx+xm−x
y−xy.
Definingλ=y−xm
y−x, it follows that 1 −λ=xm−x
y−x.
Definition 9.0.2. We say that f:D→R,w h e r e D⊂Rnis convex, is a
convex function if fis a convex function on every line in D.
Theorem 9.0.1. Suppose f∈C2(D)andH(x)is positive de finite. Then
fis convex on D.
Proof. Letx∈Dandηbe some direction. Then x+ληis a line in D.W e
computed2
dλ2f(x+λn)
d
dλf=∇f·n=
∂f
∂x1...
∂f
∂xn
·[n1...n n].
Now
d
dλµ∂f
∂x1¶
=∇∂f
∂x1·η=
∂2f
∂x1∂x1
∂2f
∂xn∂x1
·[η1,... ,ηn].
Hence we see that
d2
dλ2f(x+λη)=ηTH(x)η≥0
by assumption.
239
Example 9.0.2. LetA=[aij]∈Mn. Consider the quadratic form on Cn
orRndefined by
Q(x)=xTAx=Σaijxjxi
=1
2Σ(aij+aji)xjxi
=xT1
2(A+AT)x.
Since the matrix A+ATis symmetric the study of quadratic forms is reduced
to the symmetric case.
Example 9.0.3. LetLf=nP
i,j=1aij∂2f
∂xi∂xj.Lis called a partial di fferential
operator. By the combination of devices above (assuming f∈C2for exam-
ple) we can study the symmetric and equivalent partial di fferential operator
Lf=nX
i,j=11
2(aij+aji)∂2f
∂xi∂xj.
In particular, if A+ATis positive de finite the operator is called elliptic.
Other cases are
(1) hyperbolic,(2) degenerate/parabolic.
Characterizations of Her mitian matrices. Recall
(1)A∈M
nis Hermitian if A∗=A.
(2)A∈Mnis called skew-Hermitian if A=−A∗.
Here are some facts
(a) If Ais Hermitian the diagonal is real.
(b) If Ais skew-Hermitian the diagonal is imaginary.
(c)A+A∗,A A∗andA∗Aare all Hermitian if A∈Mn.
(d) If Ais Hermitian than Ak,k=0,1,... , are Hermitian. A−1is Her-
mitian if Ais invertible.
240 CHAPTER 9. HERMITIAN AND SYMMETRIC MATRICES
(e)A−A∗is skew-Hermitian.
(f)A∈Mnyields the decomposition
A=1
2(A+A∗)+1
2(A−A∗)
Hermitian Skew Hermitian
(g) If Ais Hermitian iAis skew-Hermitian. If Ais skew-Hermitian then
iAis Hermitian.
Theorem 9.0.2. LetA∈Mn.T h e n A=S+iTwhere SandTare
Hermitian. Moreover this is unique.
Proof.
A=1
2(A+A∗)+1
2(A−A∗)
=S+iT
where S=1
2(A+A∗)a n d
iT=1
2(A−A∗)
⇒T=−i
2(A−A∗).
Theorem 9.0.3. LetA∈Mnbe Hermitian. Then
(a)x∗Axis real for all x∈Cn;
(b) All the eigenvalues of Aare real;
(c)S∗ASis Hermitian for all S∈Mn.
Proof. For (a) we have
x∗Ax=Σaijxjxi.
The conjugate is
x∗Ax=Σ¯aij¯xjxi=Σaji¯xjxi
=Σaijxj¯xi+x∗Ax
Theorem 9.0.4. LetA∈Mn.T h e n Ais Hermitian if and only if at least
one of the following holds:
9.1. VARIATIONAL CHARACTERIZATIONS OF EIGENVALUES 241
(a)hAx, x i=x∗Axis real∀x∈Cn.
(b)Ais normal with real eigenvalues.
(c)S∗ASis Hermitian for all S∈Mn.
Proof. Hermitian ⇒(a), (b), or (c) are obvious. (a) ⇒Hermitian. Prove
using ejandek.W eh a v e
hAej+ek,ej+eki=ajj+akk+ajk+akj
⇒ajk+akjis real⇒Imajk=−Imakj.
Similarly
hAiej+ek,i ej+eki−ajj+akk+iakj−iajk
⇒i(akj−ajk)i sr e a l⇒Reakj=R e ajk.
All this gives A∗=A.
(c)⇒Hermitian S∗ASis Hermitian ⇒S∗ASis similar to a real diagonal
matrix. Therefore Ais similar to a real diagonal matrix. Just let S=Ito
getAis Hermitian.
Theorem 9.0.5 (Spectral Theorem). LetA∈Mnbe Hermitian. Then
Ais unitarily (similar) equivalent to a real diagonal matrix. If Ais real
Hermitian, then Ais orthogonally similar to a real diagonal matrix.
9.1 Variational Characterizations of Eigenvalues
LetA∈Mnbe Hermitian. Assume
λmin≤λ1≤λ2≤···≤λn−1≤λn=λmax.
Theorem 9.1.1 (Rayleigh—Ritz). LetA∈Mn, and let the eigenvalues
ofAbe ordered as above. Then
λmax=λn=m a x
x6=0hAx, x i
hx, xi
λmin=λ1=m i n
x6=0hAx, x i
hx, xi.
242 CHAPTER 9. HERMITIAN AND SYMMETRIC MATRICES
Proof. Letx1...x nbe the linearly independent and orthogonal eigenvectors
ofA. Any vector x∈Cnhas the representation
xΣαjxj.
Then hAx, x i=Σα2
jλjand hx, xi=Σα2
j.H e n c e
hAx, x i
hx, xi=X
jα2
jP
iα2iλj.
Since½
α2
j
Σα2
i¾
is nonnegative and sums to 1, it follows that
hAx, x i
hx, xi
is a convex combination of the eigenvalues, whence the theorem follows. In
particular take x=xn(x1) to achieve the maximum (minimum).
Remark 9.1.1. This result gives just the largest and smallest eigenvalues.
How can we achieve the intermediate eigenvalues? We have already consid-
ered this problem somewhat in conjunction with the power method. In thatconsideration we employed the bi-orthogonal eigenvectors. For a Hermitianmatrix, the families are the same. So we could characterize the eigenvaluesin a manner similar to that discussed previously. However, the following
characterization is simpler.
Theorem 9.1.2. LetA∈M
nbe Hermitian with eigenvalues as above and
corresponding eigenvectors x1...x n.T h e n
(∗) λn−k=m a x
x⊥{xn,... ,x n−k+1}
x 6=0hAx, x i
hx, xi.
The following result is even more general
(∗) λn−k=m i n
{wn,wn−1...wn−k+1}max
x⊥{wn,wn−1...wn−k+1}
x 6=0hAx, x i
hx, xi.
Proof. The Rayleigh—Ritz argument above gives ( ∗) directly.
9.2. MATRIX INEQUALITIES 243
9.2 Matrix inequalities
When the underlying matrix is symmetric or positive de finite, certain pow-
erful inequalities can be established. The first inequality is a consequence
of the cofactor result Proposition 2.5.1.
Theorem 9.2.1. LetA∈Mn(C)be positive de finite. Then detA≤a11a22···ann
Proof. Expanding in minors, the determinant of A
detA=a11¯¯¯¯¯¯¯a
22···a2n
.........
an2 ann¯¯¯¯¯¯¯+d e t¯¯¯¯¯¯¯¯¯0a
12···a1n
a21a22···a2n
............
an1an2 ann¯¯¯¯¯¯¯¯¯
Since Ais positive de finite, it follows that each of it principal submatrices
is also. Therefore, the first term in the right side of the equality above is
postive, while the second term is negative by the cofactor result Proposition2.5.1. To clarify , it is important to note that if Ais Hermitian and invert-
ible, so also is its inverse. In particular, it Ais positive de finite, we know
its determinant is positive and so when we write
det¯¯¯¯¯¯¯¯¯0a
12···a1n
a21a22···a2n
............
an1an2 ann¯¯¯¯¯¯¯¯¯=−X
a
1ja1kaik
it is clear this term is negative. Therefore,
detA≤a11¯¯¯¯¯¯¯a
22···a2n
.........
an2 ann¯¯¯¯¯¯¯
whence the result follows inductively.
Corollary 9.2.1. LetS1,S2,..., S rbe any partition of the integers {1,...,n },
and let A1,A2,...,A r., be the principal submatrices of Apertaining to the
indices of each subdivision. Then
detA=d e t A1detA2···detAr
An important consequence of Theorem 9.2.1 is one of the most famous
determinantal inequalities. It is due to Hadamard.
244 CHAPTER 9. HERMITIAN AND SYMMETRIC MATRICES
Theorem 9.2.1. Let B be an arbitrary nonsingular real square matrix. Then
(detB)2≤nY
i=1nX
j=1|bij|2
Proof. DefineA=BTB.Then
diagA=
nX
j=1|b1j|2,. . . ,nX
j=1|bnj|2
whence the result follows from Theorem 9.2.1 since det A=( d e t B)2.
9.3 Exercises
1. Prove Corollary 9.2.1.
2. Establish Theorem 9.2.1 in the case the matrix is complex.
3. Establish a result similar to Theorem 9.2.1 for rectangular matrices.
Chapter 10
Nonnegative Matrices
10.1 De finitions
Nonnegative matrices are simply those with all nonnegative entries. We
will eventually distinguish various types of such matrices substantially onhow they are nonnegative or conversely where they are zero. A particular
type of nonnegative matrix, the type that has no o ff-diagonal block to be
zero, will turn out to have the greatest value to us. If nonnegative matricesdistinguished themselves in only minor ways from general matrices, theywould not occupy the high position of importance they enjoy in the modernliterature. Indeed, many remarkable properties of nonnegative matriceshave been uncovered. Owing substantially to the many applications to
economics, probability, and engineering, the subject has now for more than
a century been one of the hottest areas of research within the subject.
Definition 10.1.1. A∈M
n(R)i sc a l l e d nonnegative ifaij≥0f o r1≤
i, j≤n. We use the notation A≥0 for nonnegative matrices. If aij>0
for all i, jwe say Aisfully (orstrictly )positive . (Notation. AÀ0.)
In this section we need some special notation. Let x∈Cn,x=(x1,... ,x n)T.
Then
|x|=(|x1|,... , |xn|)T.
We say that x∈Rnispositive ifxi≥0, 1≤i≤n.W es a yt h a t x∈Rnis
fully (orstrictly )positive ifxi>0, 1≤i≤n. We use the notation x≥0
andxÀ0 for positive and strictly positive vectors, respectively.
Note that if A≥0w eh a v e kAk∞=kAek∞where e=( 1,1,... , 1)T.
245
246 CHAPTER 10. NONNEGATIVE MATRICES
To give an idea of the direction we are heading, let us suppose that
A∈Mn(R) is a nonnegative matrix. Let λ6= 0 be an eigenvector of A
with pertaining eigenvector x≥0. Thus Ax=λx. (We will prove that
this comes to pass for every nonnegative square matrix.) Suppose furtherthat the vector Axis zero on the index set S.That is ( Ax)
i=0i f i∈S.
Sinceλ6=0i tf o l l o w st h a t xi=0i∈S.L e t SCbe the complement of
Sin the integers {1,...,n },a n dl e t PandP⊥be the projections to Sand
SC, respectively. So, Px=0a n d P⊥x=x. Also, it is easy to see that
In=P+P⊥,and hence
PAP⊥x=λPx=0
Therefore PAP⊥=0.When we write
A=³
P+P⊥´
A³
P+P⊥´
in block form
A=³
P+P⊥´
A³
P+P⊥´
=·PAP PAP⊥
P⊥AP P⊥AP⊥¸
we see that Ahas the special form
A=·PAP 0
P⊥AP P⊥AP⊥¸
This particular calculation reveals in a simple way the consequences of zero
components to eigenvectors of nonnegative matrices. Zero components ofeigenvectors implies zero blocks of A. I tw o u l db ec o r r e c tt oo b s e r v et h a t
this argument requires a nonnegative eig envector. Nonetheless, there are still
some very special properties of the remaining eigenvalues of A,p a r t i c u l a r l y
those with the modulus equal to the spectral radius. It will be convenient tohave a name for the index set where a nonnegative vector x∈R
nis strictly
positive.
Definition 10.1.2. Letx∈Rnbe nonnegative, x≥0. Let Sbe the index
set such that xi6=0i f i∈S.W ec a l l Sthesupport of the vector x.
10.2 General Theory
Definition 10.2.1. ForA∈Mn(R)a n dλ/∈σ(A)w ed e fine the resolvent
ofAby
R(λ)=(λI−A)−1.
10.2. GENERAL THEORY 247
This function behaves much like a rational function in complex variable
t h e o r y .I th a sp o l e sa te a c h λ∈σ(A). The order of the pole is the same as
t h eo r d e ro ft h ez e r oo f λin the minimal polynomial of A.
Theorem 10.2.1. Suppose A∈Mn(R)is nonnegative and λ>ρ(A),t h e n
R(λ)=(λI−A)−1≥0.
Proof. Apply the Neumann series
(?)( λI−A)−1=∞X
k=0λ−(k+1)Ak.
Sinceλ>0a n d A≥0, it is apparent that R(λ)≥0.
Remark 10.2.1. It is apparent that ( ?) holds upon noting that
(λI−A)−1Ã
1−µA
λ¶n+1!
=nX
0λ−(k+1)Ak
and that |λ|>ρ(A) yields convergence.
Theorem 10.2.2. Suppose A∈Mn(R)is nonnegative. Then the spectral
radiusρ(A)ofAis an eigenvalue of Awith at least one eigenvector, x≥0.
Proof. Suppose for each y≥0,R(λ)yremains bounded as λ↓ρ.I fx∈Cn
is arbitrary
|R(λ)x|=¯¯¯¯¯∞X
n=0λ−(n+1)Anx¯¯¯¯¯
≤∞X
n=0|λ|−(n+1)An|x|
=R(|λ|)|x|
for allλ,|λ|>ρ. It follows that R(λ)xis uniformly bounded in the region
|λ|>ρ, and this is impossible.
Now let y0≥0 be a vector for which R(λ)y0is unbounded as λ↓p,a n d
letkkdenote a vector norm. For λ>r,s e t
z(λ)=R(λ)y0/kR(λ)y0k,
where kR(λ)y0k↑∞.N o w z(λ)⊂{x|kxk=1}, and the latter is compact.
Therefore {z(λ)}λ>ρhas a cluster point x0with x0≥0a n d kx0k=1 . S i n c e
(ρI−A)z(λ)=(ρ−λ)z(λ)+y0/kR(λ)y0k
248 CHAPTER 10. NONNEGATIVE MATRICES
and since λ→ρand kR(λ)y0k↑∞ we have
lim
λ→ρ(ρI−A)z(λ)=l i m
λ→ρ(ρI−A)x0=0.
Corollary 10.2.1. Suppose A∈Mn(R)is nonnegative and Az=λzfor
some zÀ0,t h e nλ=ρ(A).
Proof. Sinceρ(A)⊂σ(AT) there is a vector y≥0f o rw h i c h ATy=ρy.I f
λ6=ρ,w em u s th a v e hz,yi=0 . B u t zÀ0 and this is impossible.
Corollary 10.2.2. Suppose A∈Mn(R)is strictly positive and z≥0such
thatAz=ρz,t h e n zÀ0.
Proof. Ifzi=0 ,t h e nnP
j=1aijzj= 0, and thus zj=0f o ra l l jsince aij>0,
for all j. Hence z= 0, a contradiction.
This result will be strengthened in the next section by another result that
has the same conclusion but a much weaker hypotheses.
Theorem 10.2.3. Suppose A∈Mn(R)is nonnegative with spectral radius
ρ(A)and de fine
rm=m i n
iX
jaij rM=m a x
iP
jaij
cm=m i n
jX
iaij cM=m a x
jP
iaij
then both inequalities below hold.
rm≤ρ(A)≤rM
cm≤ρ(A)≤cM
Moreover, if A, B∈Mn(R)thenρ(A+B)≥max(ρ(A),ρ(B)).
Proof. We prove the second set of inequalities. Let x∈Rnbe the eigenvec-
tor pertaining to ρ(A).Since x≥0 we can normalize xso thatPxi=1.
Then
nX
j=1aijxj=ρ(A)xi,i =1,...,n
10.2. GENERAL THEORY 249
Sum both side in ito obtain
nX
i=1nX
j=1aijxj=nX
i=1ρ(A)xi
nX
j=1ÃnX
i=1aij!
xj=ρ(A)
ReplacingPn
i=1aijbycMmakes the sum on the left larger, that is
ρ(A)≤nX
j=1ÃnX
i=1aij!
xj≤cMnX
i=1xi=cM
Similarly replacingPn
i=1aijbycmmakes the sum on the left smaller,
and therefore ρ(A)≥cm.Thus, cm≤ρ(A)≤cM.The other inequality,
rm≤ρ(A)≤rM,can be easily established by applying what has just been
proved to the matrix ATnoting of course that ρ(A)=ρ(AT).
The proof that ρ(A+B)≥max(ρ(A),ρ(B)) is an easy consequence of
a Neumann series argument.
We could just as well prove the first inequality and apply it to the transpose
to obtain the second. This proof is given below.
Alternative Proof. Letxbe the eigenvector pertaining to ρ(A)a n d xk=
max
ixi.T h e n
ρxk=nX
j=1aijxj≤
nX
j=1aij
maxxj.
Hence
ρ≤nX
j=1aij≤max
knX
j=1akj=t.
To obtain s≤ρ,l e ty≥0s a t i s f y
ATy=ρy
and suppose kyk1=1 .T h e nw eh a v eo nt h eo n eh a n d
he, ATyi=ρhe, yi=ρ
250 CHAPTER 10. NONNEGATIVE MATRICES
and on the other hand applying convexity
he, ATyi=X
iX
jaT
ijyj
=X
jyjÃX
iaji!
≥min
jX
iaji=s
Corollary 10.2.1. (i) Ifρ(A)is equal to one of the quantities cmorcM,
then all of the sumsP
iaijare equal for jin the support of the eigenvector
xpertaining to ρ.
(ii) Ifρ(A)is equal to one of the quantities rmorrM, then all of the sumsP
jaijare equal for iin the support of the eigenvector yofATpertaining
toρ. (iii) If all the row sums (resp. column sums) are equal, the spectral
radius is equal to this value.
Proof. Assume that ρ(A)=cM.L e t Sdenote the support (recall De finition
10.1.2) of the eigenvector x. Assume also that the eigenvector xis normal-
ized so thatPn
j=1xj=P
j∈Sxj= 1. Following the proof of Theorem 10.2.3
we have
nX
j=1ÃnX
i=1aij!
xj=X
j∈SÃnX
i=1aij!
xj=cM
Since1
cM(Pn
i=1aij)≤1f o re a c h j=1,...,n it follows that
X
j∈S1
cMÃnX
i=1aij!
xj≤X
j∈Sxj
Clearly if for some j∈Swe have1
cM(Pn
i=1aij)<1 the inequality above
will become a strict inequality, and the result is proved. The result is proved
similarly for cm
(ii) Apply the proof of (i) to AT.
10.3 Mean Ergodic Theorem
LetA∈Mn(R). We have seen that understanding the nature of powers Ak,
k=1,2,... leads to interesting conclusions. For example if lim
k→∞Ak=P,
10.3. MEAN ERGODIC THEOREM 251
it is easy to see that Pis idempotent ( P2=P)a n d AP=PA=P.I t i s
therefore a projection to the fixed space ofA. A less restrictive property is
mean convergence :
Mk=k−1(I+A+A2+···+Ak)
lim
k→∞Mk=P.
Note that if ρ(A)<1w eh a v e Ak→0. Hence Mk→0. On the other hand
ifρ(A)>1t h e Mkbecome unbounded. Hence mean convergence has value
only when ρ(A)=1 .
Theorem 10.3.1 (Mean Ergodic Theorem for matrices). LetA∈Mn(R)
withρ(A)=1 .I f{Ak}∞
k=1is bounded then lim
k→∞Mk=P,w h e r e Pis a pro-
jection, commuting with A,o n t ot h e fixed space of A.
Proof. We need a norm k·k, vector and matrix. We have kAkk<C for
k=1,2,... . It is easy to see that
(?) AM k−Mk=1
kÃkX
0Aj+1!
−1
kÃkX
0Aj!
=1
k(I+Ak+1)
≤1
k(1 +C).
Since the matrices Mkare bounded they have a cluster point P. This means
there is a subsequence Mkjthat converges to P. If there is another cluster
point Q(i.e. there is a subsequence M`j→Q), then we compare PandQ.
First we know
kMkj−Pk<ε
2(1 + C)
kM`j−Qk<ε
2(1 + C).
From ( ?)w eh a v e AP=PandAQ=Q, whence MkjP=PandMljQ=Q
for all k=1,2,... . Hence
P−Q=M`j(Mkj−P)−Mkj(M`j−Q)
or
kP−Qk≤ε.
252 CHAPTER 10. NONNEGATIVE MATRICES
Thus P=QandMkconverges to some matrix P.W eh a v e AP=PA=P
hence MkP=Pfor all kand therefore P2=P. Clearly the range of P
consists of all fixed vectors under A.
Corollary 10.3.1. LetA∈Mn(R)and suppose that lim
k→∞Ak=P,t h e n
lim
k→∞Mkexists and equals P.
Definition 10.3.1. LetA∈Mn(R)h a v eρ(A)=1 . T h e n Ais said to be
mean ergodic if
limk−1(I+A+···+Ak)
exists.
Clearly, if A∈Mn(R)s a t i s fiesρ(A)=1a n d Akis bounded, then Ais
mean ergodic. (Previous theorem.) Conversely, if Ais mean ergodic then
the sequence Akis bounded. This can be proved by resorting to the Jordan
canonical form and showing Akis bounded if and only if Ak
jis bounded,
where AJis the Jordan canonical form of A.
Lemma 10.3.1. LetA∈Mn(R)withρ(A)=1 .S u p p o s e λ=1 is a
simple pole of the resolvent, and the only eigenvalue of modulus 1. ThenA=P+Bwhere P=l i m
λ→1(λ−1)R(λ)is a projection onto the fixed space
ofA,PB=BP=0 andρ(B)<1.
Proof. We know that Pexists and using the Neumann series it is clear that
AP=PA=P
Alim
λ→1(λ−1)(λI−A)−1=A(λ−1)∞X
0λ−(k+1)Ak
=l i m
λ→1(λ−1)∞X
0λ−(k+1)Ak+1
=l i m
λ→1(λ−1)λ"X
0λ−(k+1)Ak−λ−1#
=l i m
λ→1(λ−1)R(λ)=P.
That is, AP=P. The other assertions follow similarly. Also, it follows that
P2=l i m
λ→1PR(λ)=P
10.4. IRREDUCIBLE MATRICES 253
which is to say that Pis a projection. We have, as well, that Ax=x⇒
Px=x, again using the Neumann series.
Now de fineB=A−P.T h e n BP=PB= 0. Suppose Bx=ax,w h e r e
|α|≥1. Then Px=0 ,f o r
αPx=PBx =0.
Therefore Ax=αx.H e r eα6=1i m p l i e s x= 0 by hypothesis and α=1
implies Px=x,o r x= 0, also. This contradicts the assumption that
|α|≥1.
Theorem 10.3.2. LetA∈Mn(R)satisfy A≥0andρ(A)=1 .T h e
following are equivalent:
(a)λ=1 is a simple pole of R(λ).
(b){Ak}∞
1is bounded.
(c)Ais mean ergodic.
Moreover
lim
λ→1(λ−1)R(λ)= l i m
k→∞1
k(I+A+···+Ak)
if either limit exists.
10.4 Irreducible Matrices
Irreducible matrices form one of the cornerstones of positive matrix theory.
Because the condition of irreducibility is both natural and often satis fied,
their applications are far and wide.
Definition 10.4.1. LetA∈Mn(C).Thezero pattern ofAis the set of
ordered pairs S={(i, j)|aij=0}.
Example 10.4.1. For example with
A=
012
100230
S={(1,1),(2,2),(2,3),(3,3)}.
254 CHAPTER 10. NONNEGATIVE MATRICES
Definition 10.4.2. A generalized permutation matrix Bis any matrix hav-
i n gt h es a m ez e r op a t t e ra sap e r m u t a t i o nm a t r i x .
This means of course that a generalized permutation matrix has ex-
actly one nonzero entry in each row and one nonzero entry in each column.Generalized permutation matrices ar e those nonnegative matrices with non-
negative inverses.
Theorem 10.4.1. LetA∈M
n(R)be invertible and A≥0.Then Ais
invertible with nonnegati ve inverse if and only if Ais a generalized permu-
tation matrix.
Proof. First, if Ais a generalized permutation matrix, de fine the matrix B
by
bij=
0i f aij=0
1
aijifaij6=0
Then A−1=BT≥0. On the other hand suppose A−1≥0a n d Ahas
on some row two non zero entries, say ai,j1andai,j2.S i n c e A−1≥0i s
invertible there is a nonzero entry in the ithcolumn, say a−1
ki>0. Now
compute multiply A−1A. It is easy to see that
¡
A−1A¢
k,j1>0a n d¡
A−1A¢
k,j2>0
which contradicts that A−1A=I.
Example 10.4.1. The analysis for 2 ×2 matrices can be carried out directly.
Suppose that
A=·ab
cd¸
and A−1=1
detA·d−b
−ca¸
There are two cases: (1) det A> 0. In this case we conclude that b=c=0.
Thus Ais a generalized permutation matrix. (2) det A< 0. In this case
we conclude that a=d=0.Again Ais a generalized permutation matrix.
As these are the only two cases, the result is veri fied.
Definition 10.4.3. LetA, B∈Mn(R). We say that Aiscogredient to
Bif there is a permutation matrix Psuch that
B=PTAP
10.4. IRREDUCIBLE MATRICES 255
Note that two cogredient matrices are similar, indeed they are unitarily
similar for the very special class of unitary matrices given by permutations.It is important to note that since APmerely interchanges the columns of
AandP
TAPthen interchanges the rows of AP.From this we observe that
both AandB=PTAPhave exactly the same elements, though permuted.
Definition 10.4.4. LetA∈Mn(R). We say that Aisirreducible if it is
notcogredient to a matrix of the form
·A10
A3A4¸
where the block matrices A1andA4are square matrices. Otherwise the
matrix is called reducible.
Example 10.4.2. All diagonal matrices are reducible.
There is another way to describe irreducible (and reducible) matrices in
terms of projections. Let S={j1,...,j k}⊂{1,2,...,n }and de finePSto
be the projection to the coordinate directions by
PSei=½1if i∈S
0if i /∈S
Furthermore de fine
P⊥
S=I−PS
Then PSandP⊥
Sare orthogonal projections. That is, the product PSP⊥
S=
0 and by de finition PS+P⊥
S=I.I f SCdenotes the complement of Sin
{1,2,...,n },it is easy to see that P⊥
S=PSC
Proposition 10.4.1. LetA∈Mn(R).T h e n Ais irreducible if and only
if there is no set of indices S={j1,...,j k}⊂{1,2,...,n }for which
PSAP⊥
S=0.
Proof. Suppose there is a set of indices S={j1,...,j k}⊂{1,2,...,n }for
which PSAP⊥
S=0.Define the permutation matrix as from the permutation
that takes the firstkintegers 1 ,..., k to the integers j1,...,j kand the
integers k+1,...,n to the complement SC.Then
PTAP=·A1A2
A3A4¸
256 CHAPTER 10. NONNEGATIVE MATRICES
T h ee n t r i e si nt h eb l o c k A2are those with rows given from the rows cor-
responding to the indices in Sand columns corresponding to the indices
inSC.T h u s Ais reducible. Conversely, suppose that Ais reducible and
that Pis a permutation matrix such that
PTAP=·A10
A3A4¸
For de finiteness, let us assume that A1isk×kandA4is (n−k)×(n−k).
DefineS={ji|pji,i=1,i=1,...,k }.a n d T={ji|pji,i=1,i=
k+1,...,n }Because Pis a permutation, it follows that T=SCand
PSAP⊥
S= 0, which proves the converse.
As we know for a nonnegative matrix A∈Mn(R) the spectral radius is
an eigenvalue of Awith pertaining eigenvector x≥0.When the additional
assumption of irreducibility of Ais added the conclusion can be strengthened
toxÀ0.To prove this result we establish a simple result from which the
xÀ0 follows almost directly.
Theorem 10.4.2. LetA∈Mn(R)be nonnegative and irreducible. Sup-
pose that y∈Rnis nonnegative with exactly 1≤k≤n−1nonzero entries.
Then (I+A)yhas strictly more nonzero entries.
Proof. Suppose the nonzero entries of yareS={j1,...,j k}.S i n c e ( I+A)y=
y+Ay, it follows immediately that (( I+A)y)ji6=0f o r ji∈S,a n dt h e r e -
fore there are at least as many nonzero entries in ( I+A)yas there are in
y.In order that the number of nonzero entries not increase, we must have
for each index i/∈Sthat ( Ay)i=0.With PSdefined as the projection
to the standard coordinates with indices from Swe conclude therefore that
P⊥
SAP=0,which means that Ais reducible, a contradiction.
Corollary 10.4.1. LetA∈Mn(R)be nonnegative. Then Ais irreducible
if and only if (I+A)n−1>0.
Proof. Suppose that Ais irreducible and y≥0i sa n yv e c t o ri n Rn.T h e n
(I+A)yhas strictly more nonzero coordinates and thus ( I+A)2yhas even
more nonzero coordinates. We see that the number of nonzero coordinatesof (I+A)
kymust increase by at least one until the maximum of nnonzero
coordinates is reached. This must occur by the ( n−1)thpower. Now
apply this to the vectors of the standard basis e1,...,e n. For example,
(I+A)n−1ekÀ0.This means that the kthcolumn of ( I+A)n−1is strictly
positive. The result follows.
10.4. IRREDUCIBLE MATRICES 257
Suppose conversely that ( I+A)n−1>0 but that Ais reducible. Then
there is a permutation matrix Psuch that PTAP has the form
PTAP=·A10
A3A4¸
It is easy to see that all the powers of Ahave the same form – with respect
to the same blocks. Also, ( I+A)n−1=I+Pn−1
i=1¡n−1
i¢
Ai.So
PT(I+A)n−1P=PTÃ
I+n−1X
i=1µn−1
i¶
Ai!
P
=I+n−1X
i=1µn−1
i¶
PTAiPT
has the same form
PT(I+A)n−1P=·˜A10
˜A3˜A4¸
a n dt h i sp r o v e st h er e s u l t .
Corollary 10.4.2. LetA∈Mn(R)be nonnegative. Then Ais irreducible
if and only if for each pair of indices (i, j)there a positive power 1≤k≤n
such that¡
Ak¢
ij>0.
Proof. Suppose that Ais irreducible. Since ( I+A)n−1=I+Pn−1
i=1¡n−1
i¢
Ai>
0, it follows that A(I+A)n−1=A+Pn−1
i=1¡n−1
i¢
Ai+1>0.Thus for each
pair of indices ( i, j),¡
Ak¢
ij>0f o r s o m e k.
On the other hand suppose that Ais reducible. Then there is a permu-
tation matrix Pfor which
PTAP=·A10
A3A4¸
Follow the same argument as above to establish that A(I+A)n−1must
h a v et h es a m ef o r ma s PTAPwith the same zero pattern. Thus there is
no positive power 1 ≤k≤nsuch that¡
Ak¢
ij>0 for any of the pairs of
indices corresponding to the (upper right) zero block above.
Another characterization of irreducibility arises when we have Ax≤axfor
some nonzero vector x≥0. This type of domination-type condition appears
to be quite weak. Nonetheless, it is equivalent to the others.
258 CHAPTER 10. NONNEGATIVE MATRICES
Corollary 10.4.3. LetA∈Mn(R)be nonnegative. Then Ais irreducible
if and only if Ax≤axfor some nonzero vector x≥0,t h e n xÀ0.
Proof. Ifxi=0,then aij= 0 whenever xj6=0.DefineS={1≤j≤
n|xj6=0}.Then with PSas the orthogonal projection to the standard
vectors ei,i∈S,it must follow that PSAP⊥
S=0 . S i n c e Sis strictly
contained in {1,2,...,n }, it follows that Ais reducible.
Now suppose that Ais reducible, which means we can assume it has
the form
A=·A10
A3A4¸
For de finiteness, let us assume that A1isk×kandA4is (n−k)×(n−k)
and of course k≥1. Let v∈Rn−kbe the nonzero nonnegative vector which
satisfiesA4v=ρ(A4)v.Letu=0∈Rk.D efine the vector x=u⊕v.T h e n
(Ax)i=½(A1u)i=0=ρ(A4)uiif 1≤i≤k
(A1v)i=ρ(A4)vi if k +1≤i≤n
The conditions of that the hypothesis are met without xÀ0.
Corollary 10.4.4 (Subinvariance). LetA∈Mn(R)be nonnegative and
irreducible with spectral radius ρ. Suppose that there is a positive number s
and nonnegative vector z≥0for which
(Az)i≤szi, for i=1,2,...,n
Then (i) zÀ0, and (ii) ρ≤s. If in addition ρ=s,t h e n Az=ρz.
Proof. Clearly, the same inequality holds for powers of A.F o r A(Az)≤
Az≤sz. Now apply induction. Next suppose that zi=0.By Corollary
10.4.2 we must have for each index pair i, jthat¡
Ak¢
ij>0 for some positive
integer kmaking¡
Akz¢
i≤sziimpossible for some k.T h u s zÀ0.To prove
the second assertion, let xbe the eigenvector of ATpertaining ρ.T h e n
hAz, x i=
z,ATx®
=ρhz,xi≤shz,xi (1)
Since xÀ0,the innerproduct hz,xi>0. Hence ρ≤s.
Finally, if ρ=sand ( Az)i<ρzifor some index ithen we conclude from
(1) that ρ<ρ, a contradiction.
10.4. IRREDUCIBLE MATRICES 259
Theorem 10.4.3. LetA∈Mn(R)be nonnegative and irreducible. If x≥
0is the eigenvector pertaining to ρ(A), i.e. Ax=ρ(A)x,t h e n xÀ0.
Proof. We have that ( I+A) is also nonnegative and irreducible, with spec-
tral radius 1 + ρ(A) and pertaining eigenvector x.S i n c e
(I+A)n−1x>(1 +ρ(A))n−1xÀ0
it follows that xÀ0.
Corollary 10.4.5. (i) Ifρ(A)is equal to one of the quantities cmorcM
from Theorem 10.2.3, then all of the sumsP
iaijare equal for j=1,...,n .
(ii) Ifρ(A)is equal to one of the quantities rmorrMfrom Theorem 10.2.3,
then all of the sumsP
jaijare equal for i=1,...,n . (iii) Conversely, if all
the row sums (resp. column sums) are equal, the spectral radius is equal to
this value.
Proof. Both are direct consequences of Corollary 10.2.1 and Theorem 10.4.3
which imply that the eigenvectors (of AandAT, resp.) pertaining to ρare
strictly positive.
10.4.1 Sharper estimates of the maximal eigenvalue
Sharper estimates for the maximal eigenvalue (spectral radius) are possible.The following results assembles much of what we have considered above. Webegin with a basic inequality that has an interesting geometric interpreta-
tion.
Lemma 10.4.1. Leta
i,bi,i=1,...n be two positive sequences. Then
min
jbj
aj≤Pn
j=1bjPn
j=1aj≤max
jbj
aj
Proof. We proceed by induction. The result is trivial for n=1 . F o r n=2
the result is a simple consequence of vector addition as follows. Considerthe ordered pairs ( a
i,bi),i=1,2. Their sum is ( a1+a2,b1+b2). It is a
simple matter to see that the slope of this vector must lie between the slopesof the summands, which is to say
min
jbj
aj≤P2
j=1bjP2
j=1aj≤max
jbj
aj
260 CHAPTER 10. NONNEGATIVE MATRICES
as is shown below.
Now assume by induction the result holds up to n−1.Then we must have
that
Pn
j=1bjPn
j=1aj=Pn−1
j=1bj+bnPn−1
j=1aj+an≤max"Pn−1
j=1bjPn−1
j=1aj,bn
an#
≤max·
max
1≤j≤n−1bj
aj,bn
an¸
=m a x
1≤j≤nbj
aj
A similar argument serves to establish the other inequality.
Theorem 10.4.4. LetA∈Mn(R)be nonnegative with row sums rkand
spectral radius ρ.T h e n
min
k
1
rknX
j=1akjrj
≤ρ≤max
k
1
rknX
j=1akjrj
(2)
Proof. First assume that Ais irreducible. Let x∈Rnbe the principal
nonnegative eigenvector of AT,s o ATx=ρxandxÀ0. Assume also
that xhas been normalized so thatPxi= 1. By summing both sides of
ATx=ρx,w es e et h a tPrkxk=ρ.S i n c eρ2is the spectral radius of A2,w e
also have that¡
AT¢2x=¡
A2¢Tx=ρ2x. Summing both sides, it follows
thatPn
k=1Pn
j=1rjaT
jkxk=ρ2.N o w
ρ=ρ2
ρ=Pn
k=1Pn
j=1rjaT
jkxk
rkxk=m a x
kPn
j=1akjrj
rk
by Lemma 10.4.1. The reverse inequality is proved similarly. In the
case that Ais not irreducible, we take the limit Ak↓Awhere the Akare
irreducible matrices. The result holds for each Ak, and the quantities in the
inequality are continuous in the limiting process. Some small care must be
t a k e ni nt h ec a s et h a t Ahas a zero row.
10.4. IRREDUCIBLE MATRICES 261
Of course, we have not yet established that the estimates (2) are sharper
that the basic inequality min krk≤ρ≤max krk. That is the content of
following result, whose proof is left as an exercise.
Corollary 10.4.6. Let A∈Mn(R)be nonnegative with row sums rk.
Then
min
krk≤min
k
1
rknX
j=1akjrj
≤ρ≤max
k
1
rknX
j=1akjrj
≤max
krk
We can continue this kind of argument even further. Let A∈Mn(R)
be nonnegative with row sums riand let xbe the nonnegative eigenvector
pertaining to ρ.T h e n¡
AT¢3x=ρ3xor expanded
X
m,k,jaT
ijaT
jkaTkmxm=ρ3xi
Summing both sides yields
X
m,k,jrjaT
jkaTkmxm=ρ3
Therefore,
ρ=ρ3
ρ2=P
mP
kP
jrjaT
jkaTkmxmP
mP
jrjaT
jmxm
≤max
mP
kP
jrjaT
jkaTkm
P
jrjaT
jm=m a x
mP
kamkP
jakjrjP
jamjrj
by Lemma 10.4.1, thus yielding another estimate for the spectral radius.
The estimate is two-sided as those above. This is summarized in the fol-lowing result.
Theorem 10.4.5. LetA∈M
n(R)be nonnegative with row sums ri.T h e n
min
mP
kamkP
jakjrjP
jamjrj≤ρ≤max
mP
kamkP
jakjrjP
jamjrj(3)
Remember the complete proof as developed above does require the assump-
tion irreducibility and the passage of the limit. Let’s consider an example
and see what these estimates provide.
262 CHAPTER 10. NONNEGATIVE MATRICES
Example 10.4.3. Let
A=
133
535
114
The eigenvalues of Aare−2,2, and 8. Thus ρ= 8. The row sums are
r1=7,r2=1 3 ,r3= 6, which yields the estimates 6 ≤ρ≤13. The
estimates from (2) are64
7,8,and22
3giving the estimates
7.333333333 ≤ρ≤9.142857143
Finally, from (3) the estimates for ρare
7.818181818 ≤ρ≤8.192307692
It remains to show that the estimates in (3) furnish an improvement to
the estimates in (2). This is furnished by the following corollary, which is
left as an exercise.
Corollary 10.4.7. LetA∈Mn(R)be nonnegative with row sums ri.T h e n
for each m=1,...,n
min
k
1
rknX
j=1akjrj
≤P
kamkP
jakjrjP
jamjrj≤max
k
1
rknX
j=1akjrj
(4)
These results are by no means the end of the story on the important
subject of eigenvalue estimation.
10.5 Stochastic Matrices
Definition 10.5.1. Am a t r i x A∈Mn(R),A≥0, is called (row) stochastic
if each row sum of Ais 1 or equivalently Ae=e.( R e c a l l , e=( 1,1,... , 1)T.)
Similarly, a matrix A∈Mn(R),A≥0, is called column stochastic if ATis
stochastic.
A direct consequence of Corollary 10.2. 1(iii) is that the spectral radius of
stochastic matrices must be one.
Theorem 10.5.1. LetA∈Mn(R)row or column stochastic . Then the
spectral radius of Aisρ(A)=1 .
10.5. STOCHASTIC MATRICES 263
Lemma 10.5.1. Let A, B∈Mn(R)be nonnegative row stochastic ma-
trices. Then for any number 0≤c≤1,t h em a t r i c e s cA+( 1−c)Band
AB are also row stochastic. The same result holds for column stochastic
matrices.
Theorem 10.5.2. Every stochastic matrix is mean ergodic, and the periph-
eral spectrum is fully cyclic.
Proof. IfA≥0 is stochastic, so also is Ak,k=1,2,... .T h e r e f o r e Akis
bounded. Also ρ(A)≤max
iP
kaij= 1. Hence the conditions of the previous
theorems are met. This proves Ais mean ergodic.
Stochastic matrices–as a set–have some interesting properties. Let Sn⊂
Mn(R) of stochastic matrices. Then Snis a convex set, it is also closed.
Whenever a closed convex set is determined, the characterization of its ex-
treme points is desired.
Recall, the (matrix) point A∈Snis called an extreme point if whenever
A=ΣλjAj
Σλj=1 ,λj≥0a n d Aj∈Snit follows that one of the Aj’s equals Aand
allλjbut one are 0.
Examples. We can also view Sn⊂Rn2. It is a convex subset of Rn2
and is moreover the intersection of the positive orthant of Rn2with the n
hyperplanes de fined byP
jaij=1 , i=1,... ,n . [Note that we are viewing
a vector x∈Rn2as
x=(α11,... ,α1n,α21,... ,α2n,... ,αn1,... ,αnn)T]
Snis therefore a convex polyhedron inRn2. The dimension of Snisn2−n.
Theorem 10.5.3. A∈Snis an extreme point of Snif and only if Ahas
exactly one 1 in each row, and all other entries are zero.
Proof. LetC=cijbe a (0,1)-matrix. (That is a matrix cij=½1
0or for
all 1≤i, j≤n.) Suppose that C∈Snthen there is a unique jfor which
cij=1 . I f C=λA+( 1−λ)Bit is easy to see that a1j=b1j=1a n d
moreover that a1k=b1k=0f o ra l lo t h e r k6=j. It follows that Cis an
264 CHAPTER 10. NONNEGATIVE MATRICES
extreme point. Conversely suppose C∈Snisnota (0,1) matrix. Fix one
row. Let
Aj=
e
j
c2→
...
cn→
where e
j=( 0,0,... , 1,
jthposition0...0) and c2...c nare the respective rows of C.
Then
C=nX
1C1jAj.
IfC1is not a (0,1) row we are finished, since Ccannot be an extreme point.
Otherwise select any non (0,1) row and repeat this argument with respectto that row.
Corollary 10.5.1. Permutation matrices are extreme points.
Definition 10.5.2. Alattice homomorphism ofCnis a linear map A:Cn→
Cnsuch that |Ax|=A|x|, for all x∈Cn.Ais a lattice isomorphism if both
AandA−1are lattice homomorphisms.
Theorem 10.5.4. A∈Snis a lattice homomorphism if and only if Ais an
extreme point of Sn.A∈Snis a lattice isomorphism if and only if Ais a
permutation matrix.
Theorem 10.5.5. A⊂Snis a permutation matrix if and only if σ(A)⊂Γ.
Γ={λ∈C||λ|=1}.Exercise.
1. Show that any row stochastic matrix (i.e. nonnegative entries with
row sums are equal 1) with at least one column having equal entries
is singular.
2. Show that the only nonnegative matrices with nonnegative inverses
must be diagonal matrices.
3. Adapt the argument from Section 10.1 to prove that the nonnegative
eigenvector xpertaining to the spectral radius of an irreducible matrix
must be strictly positive.
10.5. STOCHASTIC MATRICES 265
4. Suppose that A, B∈Mn(R) are nonnegative matrices. Prove or
disprove
(a) If Ais irreducible then ATis irreducible.
(b) If Ais irreducible then Apis irreducible, for every positive power
p≥1.
(c) If A2is irreducible then Ais irreducible.
(d) If A, B are irreducible, then ABis irreducible.
(e) If A, B are irreducible, then aA+bBis irreducible for all non-
negative constants a, b≥0.
5. Suppose that A, B∈Mn(R) are nonnegative matrices. Prove that
ifAis irreducible, then aA+bBis irreducible for all nonnegative
constants a, b > 0.
6. Suppose that A∈Mn(R) is nonnegative and irreducible and that the
trace of Ais positive, tr A> 0.Prove that Ak>0f o rs o m es u fficiently
large integer k.
7. Prove Corollary 10.2.1(ii) for ρ(A)=rm.
8. Let A∈Mn(R) be nonnegative with row sums ri(A)a n dc o l u m n
sums ci(A). Show thatPn
i=1ri¡
A2¢
=Pn
i=1ci(A)ri(A).
9. Prove Corollary 10.4.6.
10. Prove Corollary 10.4.7. (Hint. Use Lemma 10.4.1.)
11. Let pbe any positive integer and suppose that A∈Mn(R) is nonneg-
ative with row sums ri.D efiner=(r1,..., r n)T. Recall that ( Apr)m
refers to the mthcoordinate of the vector Apr.Show that
min
m(Apr)m
(Ap−1r)m≤ρ≤max
m(Apr)m
(Ap−1r)m
(Hint. Assume first that Ais irreducible.)
12. Use the estimates (2) and (3) to estimate the maximal eigenvalue of
A=
514
315
122
The maximal eigenvalue is approximately 7.640246936.
266 CHAPTER 10. NONNEGATIVE MATRICES
13. Use the estimates (2) and (3) to estimate the maximal eigenvalue of
A=
2144
573250605557
The maximal eigenvalue is approximately 14.34259731.
14. Suppose that A∈M
n(R) is nonnegative and irreducible with spectral
radiusρ, and suppose x∈Rnis any strictly positive vector, xÀ0.
Defineτ=m a x
i(Ax)i
xi. Show that ρ≤τ.
15.
A=·10
01¸
σ(A)={1} C=
001
100
010
B=·01
10¸
σ(B)={−1,1}ρC(λ)=λ3+1=0
λ=−1,eiπ/3,e−iπ/3