Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Math / Math Book Downloads / linear algebra

allen matrix notes

PDF · 253 pages · 2.0 MB
Open PDF file

Textbook-style notes on linear algebra and matrices, apparently by an author named Allen and stored among downloaded math books rather than as Phil's own work. The opening chapter defines vector spaces and gives examples (Rn, polynomial spaces, sequence spaces). It then covers subspaces, spans, linear dependence and independence, bases and the extension-to-a-basis theorem, with proofs. The text shown ends partway through the section on dimension.

AI-written summary; may contain errors. This description is approximate.

Extracted text (machine-read; may contain errors)
1 Chapter 1 Vectors and Vector Spaces 1.1 Vector Spaces Underlying every vector space (to be de fined shortly) is a scalar fieldF. Examples of scalar fields are the real and the complex numbers R:= real numbers C:= complex numbers. These are the only fields we use here. Definition 1.1.1. Avector space Vis a collection of objects with a (vector) addition and scalar multiplication de fined that closed under both operations and which in addition satis fies the following axioms: (i) (α+β)x=αx+βxfor all x∈Vandα,β∈F (ii)α(βx)=(αβ)x (iii)x+y=y+xfor all x, y∈V (iv)x+(y+z)=(x+y)+zfor all x, y, z∈V (v)α(x+y)=αx+αy (vi)∃O∈Vz0+x=x; 0 is usually called the origin (vii) 0 x=0 (viii) ex=xwhere eis the multiplicative unit in F. 7 8 CHAPTER 1. VECTORS AND VECTOR SPACES The “closed” property mentioned above means that for all α,β∈Fand x, y∈V αx+βy∈V (i.e. you can’t leave Vusing vector addition and scalar multiplication). Also, when we write for α,β∈Fandx∈V (α+β)x the ‘+’ is in the field, whereas when we write x+yforx, y∈V,t h e‘ + ’i s in the vector space. There is a multiple usage of this symbol. Examples. (1)R2={(a1,a2)|a1,a2∈R}two dimensional space. (2)Rn={(a1,a2,... ,a n)|a1,a2,... ,a n∈R},n dimensional space. (a1,a2,... ,a n)i sc a l l e da n n-tuple. (3)C2andCnrespectively to R2andRnwhere the underlying field isC, the complex numbers. (4)Pn=l n j=0ajxj|a0,a1,... ,a n∈RM is called the polynomial space of all polynomials of degree n. Note this includes not just the polynomials of exactly degree nbut also those of lesser degree. (5)fp={(ai,...)|ai∈R,Σ|ai|p<∞}. This space is comprised of vectors in the form of in finite-tuples of numbers. Properly we would write fp(R)o rfp(C) to designate the field. (6)TN=FN n=1ansinnπx|a1,... ,a n∈Rk , trigonometric polynomials. Standard vectors in Rn e1=( 1,0,... , 0) e2=( 0,1,0,... , 0) e3=( 0,0,1,0,... , 0) ... en=( 0,0,... , 0,1)These are the unit∗vec- tors whichpoint in thenorthogonal ∗ directions. 1.1. VECTOR SPACES 9 ∗Precise de finitions will be given later. ForR2, the standard vectors are 10 CHAPTER 1. VECTORS AND VECTOR SPACES e1=( 1,0) e2=( 0,1) (1,0)(0,1) (0,0)12 e Graphical representa- tion of e1ande2in the usual two dimensional plane. Recall the usual vector addition in the plane uses the parallelogram rule yx+y ForR3, the standard vectors are e1=( 1,0,0) e2=( 0,1,0) e3=( 0,0,1)(0,0,1) (1,0,0)ee e(0,1,0)23 1 Graphical representa- tion of e1,e2,a n d e3in the usual Linear algebra is the mathematics of v ector spaces and their subspaces. We will see that many questions about vector spaces can be reformulated asquestions about arrays of numbers. 1.1.1 Subspaces LetVbe a vector space and U⊂V.W e w i l l c a l l Uasubspace ofVifU is closed under vector addition, scalar multiplication and satis fies all of the vector space axioms. We also use the term linear subspace synonymously. 1.1. VECTOR SPACES 11 Examples. Proofs will be given later letV=R3={(a, b, c )|a, b, c∈R} (1.1) U={(a, b,0)|a, b∈R}. Clearly U⊂Vand also Uis a subspace of V. let v1,v2∈R3(1.2) W={av1+bv2|a, b∈R} Wis a subspace of R3. In this case we say Wis “spanned” by {v1,v2}. In general, let S⊂V,a vector space, have the form S={v1,v2,... ,v k}. Thespan ofSis the set U=  k3 j=1ajvj|a1,... ,a k∈R  . We will use the notion S(v1,v2,... ,v k) for the span of a set of vectors. Definition 1.1.2. We say that u=a1v1+···+akvk is alinear combination of the vectors v1,v2,... ,v k. Theorem 1.1.1. LetVbe a vector space and U⊂V.I fUis closed under vector addition and scalar multiplication, then Uis a subspace of V. Proof. We remark that this result provides a “short cut” to proving that a particular subset of a vector space is in fact a subspace. The actual proofof this result is simple. To show (i), note that if x∈Uthen x∈Vand so (ab)x=ax+bx. Nowax, bx, ax +bxand ( a+b)xareallinUby the closure hypothesis. The equality is due to vector space properties of V.T h u s( i )h o l d sf o r U.E a c h of the other axioms is proved similarly. 12 CHAPTER 1. VECTORS AND VECTOR SPACES A very important corollary follows about spans. Corollary 1.1.1. LetVbe a vector space and S={v1,v2,... ,v k}⊂V. ThenS(v1,... ,v k)is a linear subspace of V. Proof. We merely observe that S(v1,... ,v k)=lk3 1ajvj|a1,... ,a k∈RorCM . This means that the closure is built right into the de finition of span. Thus, if v=a1v1+···+akvk w=b1v1+···+bkvk then both v+w=(a1+b1)v1+···+(ak+bk)vk and cv=ca1v+ca2v+···+cakv are in U.T h u s Uis closed under both operations; therefore Uis a subspace ofV. Example 1.1.1. (Product spaces.) Let VandWbe vector spaces de fined over the same field. We de fine the new vector space Z=V×Wby Z={(v, w)|u∈V, w∈W} We de fine vector addition as ( v1,w1)+(v2,w2)=( v1+v2,w1+w2)a n d scalar multiplication by α(v, w)=(αv,αw). With these operations, Zis a vector space, sometimes called the product ofVandW. Example 1.1.2. Using set-builder notation, de fineV13={(a,0,b)|a, b,∈ R}.Then Uis a subspace of R3.It can also be realized as the subspace of the standard vectors e1=( 1,0,0) and e3=( 0,0,1), that is to say V13= S(e1,e3). 1.2. LINEAR INDEPENDENCE AND LINEAR DEPENDENCE 13 Example 1.1.3. More subspaces of R3.There are two other important methods to construct subspaces of R3. Besides the set builder notation used above, we have just considered the method of spanning sets. Forexample, let S={v 1,v2}⊂R3.ThenS(S) is a subspace of R3.Simi- larly, if T={v1}⊂R3.ThenS(T) is a subspace of R3.At h i r dw a y to construct subspaces is by using inner products. Let x, w∈R3.Ex- p r e s s e di nc o o r d i n a t e s x=(x1,x2,x3)a n d w=(w1,w2,w3).Define the inrner product of xandwbyx·w=x1w1+x2w2+x3w3.Then Uw={x∈R3|x·w=0}is a subpace of R3. To prove this it is neces- sary to prove closure under vector addition and scalar multiplication. Thelatter is easy to see because the inner product is homogeneous in α,that is, (αx)·w=αx 1w1+αx2w2+αx3w3=α(x·w).Therefore if x·w=0s o also is (αx)·w.The additivity is also straightforward. Let x, y∈U.T h e n the sum (x+y)·w=(x1+y1)w1+(x2+y2)w2+(x3+y3)w3 =(x1w1+x2w2+x3w3)+(y1w1+y2w2+y3w3) =0 + 0 = 0 However, by choosing two vectors v,w,∈R3we can de fineUv,w={x∈ R3|x·y=0a n d x·w=0}.E s t a b l i s h i n g Uv,wis a subspace of R3is proved similarly. In fact, what is that both these sets of subspaces, those formedby spanning sets and those formed from the inner products are the same set of subspaces. For example, referring to the previous example, it follows that V 13=S(e1,e3)=Ue2. Can you see how to correspond the others? 1.2 Linear independence and linear dependence One of the most important problems in vector spaces is to determine if a given subspace is the span of a collection of vectors and if so, to deter- mine a spanning set. Given the importance of spanning sets, we intend to examine the notion in more detail. In particular, we consider the conceptof uniqueness of representation. LetS={v 1,... ,v k}⊂V, a vector space, and let U=S(v1,... ,v k)( o r S(S) for simpler notation). Certainly we know that any vector v∈Uhas the representation v=a1v1+···+akvk for some set of scalars a1,... ,a k. Is this representation unique? Or,c a nw e 14 CHAPTER 1. VECTORS AND VECTOR SPACES find another set of scalars b1,... ,b kn o ta l lt h es a m ea s a1,... ,a krespec- tively for which v=b1v1+···+bkvk. We need more information about Sto answer this question either way. Definition 1.2.1. LetS={v1,... ,v k}⊂V, a vector space. We say that Sislinearly dependent (l.d.) if there are scalars a1,... ,a knot all zero for which a1v1+a2v2+···+akvk=0. (T) O t h e r w i s ew es a y Sislinearly independent (l.i.). Note. If we allow all the scalars to be zero we can always arrange for ( T) to hold, making the concept vacuous. Proposition 1.2.1. IfS={v1,... ,v k}⊂V, a vector space, is linearly dependent, then one member of this set can be expressed as a linear combi- nation of the others. Proof. We know that there are scalars a1,... ,a ksuch that a1v1+a2v2+···+akvk=0 Since not all of the coe fficients are zero, we can solve for one of the vectors as a linear combination of the other vectors. Remark 1.2.1. Actually we have shown that there is novector with a unique representation in S(S). Corollary 1.2.1. If0∈S={v1,... ,v k},t h e n Sis linearly dependent. Proof. Trivial. Corollary 1.2.2. IfS={v1,... ,v k}is linearly independent then every subset of Sis linearly independent. 1.3. BASES 15 1.3 Bases The idea of a basis is that of finding a minimal generating set for a vector space. Through basis, unicity of representation and a number of other usefulproperties, both theoretical and com putational, can be concluded. Thinking of the concept in operations research ideas, a basis will be a redundancy freeand complete generating set for a vector space Definition 1.3.1. LetVbe a vector space and S={v 1,... ,v k}⊂V.W e callSaspanning set for the subspace U=S(S). Suppose that Vis a vector space, and S={v1,... ,v k}is a linearly independent spanning set for V.T h e n Sis called a basis ofV.M o d i f yt h i s definition correspondingly for subspaces. Proposition 1.3.1. IfSis a basis of V, then every vector has a unique representation. Proof. LetS={v1,... ,v k}andv∈V.T h e n v=a1v1+···+akvk for some choice of scalars. If there is a second choice of scalars b1,... ,b k not all the same, respectively, as a1,... ,a k,w eh a v e v=b1v1+···+bkvk and 0=(a1−b1)v1+···+(ak−bk)vk. Since not all of the di fferences a1−b1,... ,a k−bkare zero we must have thatSis linearly dependent. This is a contradiction to our hypothesis, and the result is proved. Example. LetV=R3andS={e1,e2,e3}.T h e n Sis a basis for V. Proof. Clearly Vis spanned by S. Now suppose that 0=a1e1+a2e2+a3e3 or (0,0,0) =a1(1,0,0) +a2(0,1,0) +a3(0,0,1) =(a1,a2,a3). Hence a1=a2=a3=0 . T h u st h es e t {e1,e2,e3}is linearly independent. 16 CHAPTER 1. VECTORS AND VECTOR SPACES Remark 1.3.1. Note how we resolved the linearly dependent/linearly in- dependent issue by converting a vector problem to a numbers problem. Thisis at the heart of linear algebra. Exercise. LetS={v 1,v2}={(1,0,1),(1,−1,0)}⊂R3. Show that Sis linearly independent and therefore a basis of S(S). 1.4 Extension to a basis In this section, we show that given a linearly independent set of vectors from a vector space with a finite spanning set, it is possible add to this set more vectors until it becomes a basis. Thus any set of linearly independent vectors can be a part (subset) of a basis. Theorem 1.4.1 (Extension to a basis). Assume that the given vector space Vhas a finite spanning set S1, i.e. V=S(S1).L e t S0={x1,... ,x f}be a linearly independent subset of Vso that S(S0)V. Then, there is a subset SI 1ofS1,s u c ht h a t S0∪SIis a basis for V. Proof. Our intention is to add vectors to S0keeping it linearly independent and eventually becoming a basis. There are a couple of steps.Steps. 1. Since S(S 1)SS(S0), there is a vector y1∈S1such that S0,1= {S0,y1}is linearly independent and thus S(S0,1)SS(S0). 2. Continue this process generating sets S0,1={S0,y1} S0,2={S0,1,y2} ... S0,j={S0,j−1,yj−1} ... At each step S0,1,S0,2,... are linearly independent sets. Since S1is finite we must eventually have that S(S0,m)=S(S1)=V. 3. Since S0,mis linearly independent and spans V, it must be a basis. 1.5. DIMENSION 17 Remark 1.4.1. In the proof it was important to begin with anyspanning set for Vand to extract vectors from it as we did. Assuming merely that there exists a finite spanning set and extracting vectors directly from V leads to a problem of terminus. That is, when can we say that the new linearly independent set being generated in Step 2 above is a spanning set forV? What we would need is a theorem that says something to the e ffect that if Vhas a finite basis, then every linearly independent set having the same number of vectors is also a basis. This result is the content of the nextsection. However, to prove it we need the Extension theorem. Corollary 1.4.1. IfS={v 1,... ,v k}is linearly dependent then the repre- sentation of vectors in S(S)isnotunique. Proof. We know there are scalars a1,... ,a knot all zero, for which a1v1+···+akvk=0 letv∈S(S) have the representation v=b1v1+b2v2+···+bkvk. Then we also have the representation v=(a1+b1)v1+(a2+b2)v2+···+(ak+bk)vk establishing the result. Remark 1.4.2. The upshot of this construction is that we can always con- struct a basis from a spanning set. In actual practice this process may bequite difficult to carry out. In fact, we will spend some time achieving this goal. The main tool will be matrix theory. 1.5 Dimension One of the most remarkable features of vector spaces is the notion of dimension. We need one simple result that makes this happen, the basis theorem. Theorem 1.5.1 (Basis Theorem). LetS={v1,... ,v k}⊂Vbe a basis forV. Then every basis of Vhaskelements. 18 CHAPTER 1. VECTORS AND VECTOR SPACES Proof. We proceed by induction. Suppose S={v1}andT={w1,w2}are both bases of V.T h e ns i n c e Sis a basis w1=α1v1w2=α2v1 and therefore 1 α1w1−1 α2w2=0 which implies that Tis linearly dependent (we tacitly assumed that both α1andα2were nonzero. Why can we do this?) The next step is to assume the result holds for bases having up to k elements. Suppose that S={v1,... ,v k+1}andT={w1,... ,w k+2}are both bases of V. Now consider SI={v1,... ,v k}. We know that S(SI) S(S)=S(T)=V. By our extension of bases result, there is a vector wf1∈Tsuch that SI 1={v1,... ,v k,wf1} is linearly independent and S(SI 1)⊂S(S)=V.I fS(SI 1)V, our extension result applies again to give a vector vf1such that SI 11={v1,... ,v k,wf1,vf1} is linearly independent The only possible selection is vf1=vk+1. But in this casewfiwill depend on v1,... ,v k,vk+1, and that is a contradiction. Hence S(v1,... ,v k,wf1)=V. The next step is to remove the vector vkfrom SI 1and apply the extension to conclude that the span of the set SI 2={v1,... ,v k−1,wf1,wf2} isV. We continue in this way eventually concluding that SI k+1={wf1,wf2,... ,w fk+1} has span V.B u t SI k+1T, whence Tis linearly dependent. Proposition 1.5.1 (Reduced spanning sets). (a) Suppose that S= {v1,... ,v k}spans Vandvjdepends (linearly) on Sj={v1,... ,v j−1,vj+1...v k}. Then Sjalso spans V. 1.5. DIMENSION 19 ( b )I fa tl e a s to n ev e c t o ri n Sis nonzero (that is VW={0}, the smallest vector space), then there is a subset S0⊂Sthat is linearly independent and spans V. Proof. (Left to reader.) Definition 1.5.1. Thedimension of a vector space Vis the (unique) num- b e ro fv e c t o r si nab a s i so f V.W ew r i t ed i m ( V) for the dimension. Remark 1.5.1. This de finition make sense possible only because of our basis theorem from which we are assured all every linearly independentspanning sets of V, that is all bases, have the same number of elements. Examples. (1) dim( R n)=n, (2) dim( Pn)=n+1 . Exercise. LetM= all rectangular arrays of two rows and three columns with real entries. Find a basis for M,a n d find the dimension of M.N o t e M=F}abc def]eeeea, b, c, d, e, f ∈Rk Example 1.5.1. P n={anxn+an−1xn−1+···+a1x+a0=0 }is the vector space of polynomials of degree n. We claim that the powers, x0=1 , x, x2,... ,xnare linearly independent, and since Pn=S(1,x ,... ,xn) they form a basis of Pn. Proof. There are several ways we can prove this fact. Here is the most direct and it requires essentially no machinery. Suppose they are linearly dependent, which means that there are coe fficients a0,a1,... ,a nso that anxn+an−1xn−1+···+a1x+a0=0, (T) the function . (This functional view is critically important because every polynomial has roots.) There must be a coe fficient which is nonzero and which corresponds to the highest power. Let us assume that anW=0 ,f o r convenience, and with no loss in generality. 20 CHAPTER 1. VECTORS AND VECTOR SPACES Solve for xnto get xn=−an−1 anxn−1+···+−a1 anx−a0 an(TT) Now compute the ratio of this expression divided by xnon both sides, and letx→∞ . The left side of course will be 1. Again for convenience we take n= 2. So, condensing terms we will have b1x+b0 x2=b1w1 xW +b0w1 x2W =1 where bj=−aj/a2.B u t a s x→∞ the expression b1D1 xi +b0D1 x2i →0. This is a contradiction. It cannot be that the functions 1 ,x,a n d x2are linearly dependent. In the general case for nwe have bn−1w1 xW +bn−2w1 x2W +···+b0w1 xnW =1, where bj=−aj/an. Apply the same limiting argument to obtain the con- tradiction. Thus T={1,x ,... ,xn} is a basis of Pn. A calculus proof is available. It is also based on the fact that if the powers are linearly independent and ( T) holds, then we can assume that the same relation ( TT)i st r u e . N o wt a k et h e nthderivative of both sides. We obtain n!=0 a contraction, and the result if proved Finally, one more technique used to prove this result is by using the Fundamental Theorem of Algebra. Theorem 1.5.2. Every polynomial ( T)o f exactly nthdegree (i.e. with anW=0) has exactly nroots counted with mu ltiplicity (i.e. if q(x)=qnxn+ qn−1xn−1+···+q1x+q0∈Pn(C),qnW=0 t h e nt h en u m b e ro fs o l u t i o n so f q(x)=0 is exactly n). From (T) above we have an nthdegree polynomial that is zero for every x. Thus the polynomial is zero, and this means allthe coefficients are zero. This is a contradiction to the hypothesis, and therefore the theorem is proved. 1.5. DIMENSION 21 Remark 1.5.2. P0P1P2···Pn···. On the other hand this is not true for the Euclidean spaces R1,R2,... . However, we may say that there is a subspace of R3which is “like” R2in every possible way. Do you see this? We have R2={(a, b)|a, b∈R} R3={(a, b, c )|a, b, c∈R}. No element in R2, an ordered pair,c a nb ei n R3, a set of ordered triples. However U={(a, b,0)|a, b∈R} is “like” R2is just about every way. Later on we will give a precise mathe- matical meaning to this comparison. Example 1.5.2. Find a basis for the subspace V0ofR3of all solutions to x1+x2+x3=0 ( T) where x=(x1,x2,x3)∈R3. Solution. First show that the set V0={(x1,x2,x3)∈R3|x1+x2+x3=0} is in fact a subspace of R3. Clearly if x=(x1,x2,x3)∈V0andy= (y1,y2,y3)∈V0then x+y=(x1+y1,x2+y2,x3+y3)∈V0,p r o v i n g closure under vector addition. Similarly V0is closed under scalar multi- plication. Next, we seek a collection of vectors v1,v2,... ,v k∈V0so that S(v1,... ,v k)=V0.L e t x3=αandx2=βbe free parameters. Then x1=−(α+β). Hence all solutions of ( T)h a v et h ef o r m x=(−(α+β),β,α) x=α(−1,0,1) +β(−1,1,0). Obviously the vectors v1=(−1,0,1) and v2=(−1,1,0) are linearly inde- pendent, and xis expressed as being in the span of them. So, V0=S(v1,v2). V0has dimension 2. 22 CHAPTER 1. VECTORS AND VECTOR SPACES Theorem 1.5.3 (Uniqueness). LetS={v1,... ,v k}be a basis of V. Then each vector v∈Vhas a unique representation with respect to S. Proof. SinceS(S)=Vwe have that v=a1v1+a2v2+···+akvk for some coe fficients a1,a2,... ,a kin the given field. (This is the represen- tation of vwith respect to S.) If it is notunique there is another v=b1v1+b2v2+···+bkvk. So, subtracting we have (a1−b1)v1+(a2−b2)v2+···+(ak−bk)vk=0 where the di fferences aj−bjare not all zero. This implies that Sis a linearly dependent set. Theorem 1.5.4. Suppose that S={v1,... ,v k}is a basis of the vector space V. Suppose that T={w1,... ,w m}is a linearly independent subset ofV.T h e n m≤k. Proof. We know that Sis a linearly independent spanning set. This means that every linearly independent set of kvectors is also a spanning set. There- fore,m>k renders a contradiction as T0={w1,... ,w k}is a spanning set andwk+1∈S(T0). Definition 1.5.2. IfAis any set we de fine |A|:= cardinality of A, that is to say |A|is the number of elements of A. Example 1.5.3. LetT={1,x ,x2,x3}.T h e n |T|=4 . Theorem 1.5.5. Both RkandCkarek-dimensional and Sk={e1,e2,... ,e k} is a basis of both. Proof. It is easy to see that e1,... ,e kare linearly independent, and any vector xinRkhas the form x=a1e1+a2e2+···+akek fora1,... ,a k∈R.T h u s Skis a linearly independent spanning set and hence a basis of Rk. 1.6. NORMS 23 Question: What single change to the proof above gives the theorem for Ck? The following results follow easily from previous results. Theorem 1.5.6. LetVbe ak-dimensional vector space. (i) Every set Twith |T|>k is linearly dependent. (ii) If D={v1,... ,v j}is linearly independent and j<k , then there are vectors vf1,... ,v fk−j∈Vsuch that D∪{vf1,... ,v fk−j} is a basis of V. (iii) If D⊂V,|D|=k,a n d Dis either a spanning set for Vor linearly independent, then Dis a basis for V. 1.6 Norms Norms are a way of putting a measure of distance on vector spaces. The purpose is for the re fined analysis of vector spaces from the viewpoint of many applications. It is also to all the comparison of various vectors on the basis of their length. Ultimately , we wish to discuss vector spaces as representatives of points. Naturally, we are all accustomed to the “shortestdistance” distance from the Pythagorean theorem. This is an example of anorm, but we shall consider them as real valued functions with very special properties. Definition 1.6.1. Norms on vector spaces over C,orR.L e t Vbe a vec- tor space and suppose that ,·,:V→R +is a function from Vto the nonnegative reals for which (i),x,≥0 for all x∈Vand ,x,=0i fa n do n l yi f x=0 (ii),αx,=|α|,x,for allα∈C,Randx∈V (iii),x+y,≤, x,+,y,for all x, y∈V “The Triangle inequality”. Then,·,is called a norm onV. The second condition is often termed the (positive) homogeneity property. Remark 1.6.1. The notation is a substitute function notation. The ex- pression ,·,, without the vector, is just the way a norm is expressed. 24 CHAPTER 1. VECTORS AND VECTOR SPACES Examples. LetV=Rn(orCn). De fine for x=(x1,... ,x n) (i),x,2=(|x1|2+|x2|2+···+|xn|2)1/2Euclidean norm (ii),x,1=(|x1|+|x2|+···+|xn|)f1norm (iii),x,∞=m a x 1≤i≤n|xi|f∞norm Norm (ii) is read as: ell one norm. Norm (iii) is read as: ell in finity norm. Proof that (ii) is a norm. Clearly (i) holds. Next ,αx,1=(|αx1|+|αx2|+···+|αxn|) =(|α||x1|+|α||x2|+···+|α||xn|) =|α|(|x1|+|x2|+···+|xn|)=|α|,x,1 which is what we needed to prove. Also, ,x+y,1=(|x1+y1|+|x2+y2|+···+|xn+yn|) ≤(|x1|+|y1|+|x2|+|y2|+···+|xn|+|yn|) =(|x1|+|x2|+···+|xn|)+( |y1|+|y2|+···+|yn|) =,x,1+,y,1. Here we used the fact that |α+β|≤|α|+|β|for numbers. To prove that (i) is a norm we need a very famous inequality. Lemma 1.6.1 (Cauchy—Schwartz). Given that a1,... ,a nandb1,... ,b n are in C.T h e n n3 1|aibi|≤Xn3 1a2 i~1/2Xn3 1b2 i~1/2 . (T) Proof. We consider for the variable t Σ(ai+tbi)2=Σa2 i+2tΣaibi+t2Σb2 i. Note that ( T) is obvious ifn 1aibi=0 . I fn o tt a k e t=−n 1a2 i n 1aibi. 1.6. NORMS 25 Then Σ(ai+tbi)2=Σa2 i−2Σa2 i ΣaibiΣaibi+D Σa2 ii2 (Σaibi)2Σb2 i =−Σa2 i+D Σa2 ii2Σb2 i (Σaibi)2 =D Σa2 iiw −1+Σa2 iΣb2 i (Σaibi)2W . Since the left side is ≥0 and since Σa2 i≥0, we must have that w −1+Σa2 iΣb2 i (Σaibi)2W ≥0. Solving this inequality we have (Σaibi)2≤Σa2 iΣb2 i. Now that square roots to get the result. To prove that (i) is a norm, we note that conditions (i) and (ii) are straightforward. The truth of condition (iii) is a consequence of anotherfamous result. Theorem 1.6.1 (Minkowski). ,x+y, 2≤,x,2+,y,2. Proof. Σ(ai+bi)2=Σa2 i+2Σaibi+Σb2 i ≤Σa2 i+2D Σa2 ii1/2D Σb2 ii1/2+Σb2 i =pD Σa2 ii1/2+D Σb2 ii1/2Q2 . Taking square roots gives the result. Continuity and Equivalence of Norms Lemma 1.6.2. Every vector norm on Cnis continuous in the vector com- ponents. 26 CHAPTER 1. VECTORS AND VECTOR SPACES Proof. Letx∈Cnand,·,some norm on Cn. We need to show that if the vectorδ→0, in components, then ,x+δ,→, x,. First, by the triangle inequality ,x+δ,≤, x,+,δ,or ,x+δ,−,x,≤,δ, Similarly ,x,≤, x+δ−δ, ≤,x+δ,+,δ,or −,δ,≤, x+δ,−,x, Therefore |,x+δ,−,x,|≤,δ, Now expressing δin components and standard bases vectors, we write δ= δ1e1+···+δnenand ,δ,≤ |δ1|,e1,+···+|δn|,en, ≤max 1≤i≤n|δi|(,e1,+···+,en,) ≤Mmax 1≤i≤n|δi| where M=,e1,+···+,en,.We know that if δ→0i nc o m p o n e n t s ,t h e n max 1≤i≤n|δi|→0.Therefore |,x+δ,−,x,|→0, as well. Definition 1.6.2. Let,·,aand,·,bbe two vector norms on Cn.We say that these norms are equivalent if there are postive constants m, M such that for all x∈Cn m,x,a≤,x,b≤M,x,a The remarkable fact about vector norms on Cnis that they are allequiv- alent. The only tool we need to prove this is the following result: Every continuous function on a compact set of Cnassumes its maximum (and minimum) on that set. The term “compact” refers to a particular kind of setK, one which is both bounded and closed. Bounded means that for max x∈K,x,≤B<∞a n dc l o s e dm e a n st h a ti fl i m n→∞xn=x,t h e n x∈K. 1.6. NORMS 27 Theorem 1.6.2. All norms on Cnare equivalent. Proof. Since equivalence of norms is an equivalence condition, we can take one of the norms to be the in finity norm ,·,∞.Denote the other norm by ,·,.N o w d e fineK={x|,x,∞=1 }.This set, called the unit ball in the infinity norm, is compact. Now we de fine m=m i n x∈K,x,and M=m a x x∈K,x, Since,·,is a continuous function on K(from the lemma above) and since Kis compact, we have that both the minimum and maximum are attained by speci fic vectors in K. Since these vectors are nonzero (they’re in K)a n d because ,x,is positive for nonzero vectors, it must follow that 0 <m< M<∞. Hence, on K,t h er e l a t i o n m,x,≤, x,∞≤M,x, holds true. For any vector x∈Cnwe can write x=wx ,x,∞W ,x,∞and x ,x,∞∈K.Thus mEEEEx ,x,∞EEEE≤EEEEx ,x,∞EEEE ∞≤MEEEEx ,x,∞EEEE mEEEEx ,x,∞EEEE,x, ∞≤EEEEx ,x,∞EEEE ∞,x,∞≤MEEEEx ,x,∞EEEE,x, ∞ m,x,≤, x,∞≤M,x, and the theorem is proved. Example 1.6.1. Example. Find the estimates for the equivalence of ,·,2 and,·,∞ Solution. Letx∈Cn. Then, because we know for any finite sequences 28 CHAPTER 1. VECTORS AND VECTOR SPACES thatn i=1|aibi|≤max 1≤i≤n|ai|n i=1|bi| ,x,2=Xn3 i=1|xi|2~1 2 ≤Xn3 i=11·|xi|2~1 2 ≤w max 1≤i≤n|xi|2W1/2Xn3 i=11~1 2 =n1 2,x,∞ On the other hand, by the Cauchy-Schwartz inequality ,x,∞=m a x 1≤i≤n|xi| ≤n3 i=1|xi| ≤Xn3 i=11~1 2Xn3 i=1|xi|2~1/2 =n1 2,x,2 Putting these inequalities together we have n−1 2,x,2≤,x,∞≤n1 2,x,2 This makes m=n−1/2andM=n1 2. Remark 1.6.2. Note that the constants mandMdepend on the dimension of the vector space. Though not the rule in all cases, it is mostly thesituation. Norms on polynomial spaces Polynomial spaces, as we have considered earlier, can be given norms as well.Since they are function spaces, our norms usually need to consider all the values of the independent variable. In many, though not all, cases we need 1.6. NORMS 29 to restrict the domains of the polynomials. With that in mind we introduce the notation Pk(a, b)=Pkwith domain restricted to the interval [ a, b] We now de fine the function versions of the same three norms we have just studied. For functions p(x)i nPk(a, b)w ed e fine 1.,p(x),2=D$b a|p(x)|2dxi1 2 2.,p(x),1=$b a|p(x)|dx 3.,p(x),∞=m a x a≤x≤b|p(x)| The positivity and homogeneity properties are fairly easy to prove. The triangle property is a little more involved. However, it has essentially beenproved for the earlier norms. In the present case, one merely “integrates” over the inequality. Sometimes ,·, 2is called the energy norm. T h ei n t e g r a ln o r m sa r er e a l l yt h e norm for polynomial spaces. Alternate norms use pointwise evaluation or even derivatives depending on the appli-cation. Here is a common type of norm that features the first derivative. Forp∈P n(a, b)d efine ,p,=m a x a≤x≤b|p(x)|+m a x a≤x≤beepI(x)ee As is evident this norm becomes large not only when the polynomial is large but also when its derivative is large. If we remove the term max a≤x≤b|p(x)| from the norm above and de fine N(p)= m a x a≤x≤beepI(x)ee This function satis fies all the norm properties except one and thus is not a norm. (See the exercises.) Point evaluation-type norms take us too far a field of our goals partly because making poi nt evaluations into norms requires some knowledge of interpolation and related topics. Leave it said that the obvious point evaluation functions such as p(a)a n dt h el i k ew i l ln o tp r o v i d e us with norms. 30 CHAPTER 1. VECTORS AND VECTOR SPACES 1.7 Ordered Bases Given a vector space Vwith a basis S={v1,v2,... ,v k}we now know that every vector v∈Vhas a representation with respect to the basis v=a1v1+a2v2+···+akvk. But no order is implied. For example, for R2we have S={e1,e2}={e2,e1} shows us that there is no particular order convey through the de finition of a basis. When we place an order on a basis we will notice an underlyingalgebraic structure of all k-dimensional vector spaces. Definition 1.7.1. LetVbe a k-dimensional vector space with basis S= {v 1;v2;...;vk}is speci fied with a fixed and well de fined order as indicated by their relevant positions. Then Sis called an ordered basis. With or- dered bases we obtain coordinates .L e t Vbe a vector space of dimension kwith ordered basis S, and suppose v∈V.F o r 1 ≤i≤k,w ed e fine theithcoordinate ofvwith respect to Sto be the ithcoefficient aiin the representation v=a1v1+a2v2+···+aivi+···+akvk. In this way we can associate each v∈Vwith a k-tuple of numbers (a1,a2,... ,a k)∈Rkthat are the coe fficients of vwith respect to S.T h e k-tuple is unique, owing to the fixed ordering of S. Conversely, for each ordered k-tuple ( a1,a2,... ,a k)∈Rkthere is associated a unique vector v∈Vgiven by v=a1v1+a2v2+···+aivi+···+akvk. We will express this association as v∼(a1,a2,... ,a k) The following properties are each simple propositions: •Ifv∼(a1,a2,... ,a k)a n d w∼(b1,b2,... ,b k)t h e n v+w∼(a1+b1,a2+b2,... ,a k+bk) •Ifv∼(a1,a2,... ,a k)a n dα∈R(orC), then αv∼α(a1,a2,... ,a k)=(αa1,αa2,... ,αak) •Ifv∼(a1,a2,... ,a k)=( 0 ,0,...0), then v=0 . 1.7. ORDERED BASES 31 We now de fine a special type of linear function from one linear space to another. The special condition is linearity of the map. Definition 1.7.2. LetVandWbe two vector spaces. We say that Vand Warehomomorphic if there is a mapping Φbetween VandWfor which (1.) For vandwinV Φ(v+w)=Φ(v)+Φ(w) (2.) For vandα∈R(orC) Φ(αv)=αΦ(v) In this case we call Φahomomorphism from VtoW.F u r t h e r m o r e ,w es a y thatVandWare isomorphic if they are homomorphic and if (3.) For each w∈Wthere exists a unique v∈Vsuch that Φ(v)=w In this case we call Φaisomorphism from VtoW. We put all this together to show that finite dimensional vector spaces over the reals (resp. complex numers) and the standard Euclidean spaces Rk(resp. Ck) are very, very closely related. Indeed from the point of view of isometry, they are identical. Theorem 1.7.1. IfVis ak-dimensional vector space over R(respC), then Vis isomorphic to Rk(resp. Ck). This constitutes the beginning of the su fficiency of matrix theory as a tool to study finite dimentsional vector spaces. Definition 1.7.3. The mapping cs:V→Rkdefined by cs(v)=(a1,a2,... ,a k) where v∼(a1,a2,... ,a k)i st h es o - c a l l e d coordinate map. Example 1.7.1. We have shown that in R3the solutions to the equation x1+x2+x3=0f o r x=(x1,x2,x3)∈R3is a subspace V0with basis S={v1,v2}={(−1,0,1),(−1,1,0)}. With respect to this basis v0∈V0if there are constants α0,β0so that v0=α0(−1,0,1) +β0(−1,1,0) With respect to this basis the coordinate map has the form cs(v0)=(α0,β0) Therefore, we have established that V0is isomorphic to R2. 32 CHAPTER 1. VECTORS AND VECTOR SPACES 1.8 Exercises. 1. Show that {(a, b,0)|a, b∈R}is a subspace of R3by proving that it is spanned by vectors in R3. Find at least two sets of spanning sets. 2. Show that {(a, b,1)|a, b∈R}cannot be a subspace of R3. 3. Show that {(a−b,2b−a, a−b)|a, b∈R}is a subspace of R3by proving that it is spanned by vectors in R3. 4. For any w∈R3, show that Uw={x∈R3|x·w=dW=0}isnot subpace of R3. 5. Find a set of vectors in R3that spans the subspace Uw={x∈R3|x· w=0},w h e r e w=( 1,1,1). 6. Why can {(a−b, a2,a b)|a, b∈R}never be a subspace of R3? 7. Let Q={x1,...,x k}be a set of distinct points on the real line with k<n . Show that the subset PQof the polynomial space Pnof polynomials zero on the set Qis in fact a subspace of Pn. Characterize PQifk>n andk=n. 8. In the product space de fined above prove that de finitions given the result is a vector space. 9. What is the product space R2×R3? 10. Find a basis for Q={ax+bx3|a, b∈R}. 11. Let T⊂Pnbe those polynomials of exactly degree n. Show that Tis not a subspace of Pn. 12. What is the dimension of Q={ax+ax2+bx3|a, b∈Q}. What is a basis for Q? 13. Given that S={x1,x2, ..., x 2k}andT={y1,y2, ..., y 2k}are both bases of a vector space V.(Note, the space Vhas dimension 2 k.) Consider the set of any kintegers L={l1,..., l k}⊂{1,2,..., 2k}.(i) Show that associated with P={xl1,xl2, ..., x lk}there are exactly kvectors from T,s a yQ={ym1,ym2,. . . ,y mk}so that P∪Qis also ab a s i sf o r V.(ii) Is the set of vectors from Tunique? Why or why not? 1.8. EXERCISES. 33 14. Given Pn.D efineZn={pI(x)|p(x)∈Pn}.( T h e n o t a t i o n pI(x)i st h e standard notation for the derivative of the function p(x)w i t hr e s p e c t to the variable x.) What is another way to express Znin terms of previously de fined spaces? 15. Show that ,x,∞is a norm. (Hint. The condition (iii) should be the focal point of your e ffort.) 16. Let S={x1,x2,. . . , x n}⊂Rn.F o r e a c h j=1,2,...,n suppose xj∈Shas the property that its firstj−1 entries equal zero and the jthentry is nonzero. Show that Sis a basis of Rn. 17. Let w=(w1,w2,w3)∈R3,where all the components of ware strictly positive. De fine,·,wonR3by,x,w=p w1|x1|2+w2|x3|2+w2|x3|2Q1/2 . Show that ,·,wis a norm on R3. 18. Show that equivalence of norms is an equivalence relation.19. De fine for p∈P n(a, b) the function N(p)=m a x a≤x≤b|pI(x)|.Show that this is not a norm on Pn(a, b). 20. For p∈Pn(a, b)d efine,p,=m a x a≤x≤b|p(x)|+m a x a≤x≤b|pII(x)|. Show this is a norm on Pn(a, b). 21. Suppose that Vis a vector space with dimension k. Find two (linearly independent) spanning sets S={v1,v2,...,v k}andW={w1,w2,...,w k} ofVsuch that if any m<k vectors are chosen from Sand any k−m vectors are chosen from T,the resulting set will be a basis for V. 22. For p∈Pn(a, b)d e fineN(p)=eepDa+b 2iee.Show this is not a norm onPn(a, b). 23. Find the estimates for the equivalence of ,·,1and,·,∞. 24. Find the estimates for the equivalence of ,·,1and,·,2. 25. Show that Pndefined over the reals is isomorphic to Rn+1. 26. Show that Tn, the space of trigonometric polynomials, de fined over the reals is isomorphic to Rn. 27. Show that the product space Ck×Cmis isomorphic to Ck+m. 28. What is the relation between the product space Pn×PnandP2n? Find the polynomial space that is isomorphic to Pn×Pn. 34 CHAPTER 1. VECTORS AND VECTOR SPACES Terms. Field Vector space scalar multiplication Closed spaceOriginPolynomial spaceSubspace, linear subspaceSpan Spanning set RepresentationUniqueness of representationlinear dependencelinear independencelinear combination Basis Extension to a basisDimensionNormf 2norm;f1norm;f2∞norm Cauchy-Schwartz inequality Fundamental Theorem of AlgebraCardinalityTriangle inquality Chapter 2 Matrices and Linear Algebra 2.1 Basics Definition 2.1.1. Amatrix is an m×narray of scalars from a given field F. The individual values in the matrix are called entries . Examples. A=^ 21 3 −124„ B=^ 12 34„ Thesizeof the array is–written as m×n,w h e r e m×n cA number of rows number of columns Notation A= a 11a12... a 1n a21a22... a 2n an1an2... a mn A ←− rows t AAc columns A:= uppercase denotes a matrix a:= lower case denotes an entry of a matrix a∈F. Special matrices 33 34 CHAPTER 2. MATRICES AND LINEAR ALGEBRA (1) If m=n, the matrix is called square .I nt h i sc a s ew eh a v e (1a) A matrix Ais said to be diagonal if aij=0 iW=j. (1b) A diagonal matrix Amay be denoted by diag( d1,d2,... ,d n) where aii=diaij=0 jW=i. The diagonal matrix diag(1 ,1,... , 1) is called the identity matrix and is usually denoted by In= 10 ... 0 01 ...... 01  or simply I,w h e n nis assumed to be known. 0 = diag(0 ,... , 0) is called the zero matrix . (1c) A square matrix Lis said to be lower triangular if fij=0 i<j . (1d) A square matrix Uis said to be upper triangular if uij=0 i>j . (1e) A square matrix Ais called symmetric if aij=aji. (1f) A square matrix Ais called Hermitian if aij=¯aji(¯z:= complex conjugate of z). (1g) Eijhas a 1 in the ( i, j) position and zeros in all other positions. (2) A rectangular matrix Ais called nonnegative if aij≥0a l l i, j. It is called positive if aij>0a l l i, j. Each of these matrices has some speci al properties, which we will study during this course. 2.1. BASICS 35 Definition 2.1.2. The set of all m×nmatrices is denoted by Mm,n(F), where Fis the underlying field (usually RorC). In the case where m=n we write Mn(F) to denote the matrices of size n×n. Theorem 2.1.1. Mm,nis a vector space with basis given by Eij,1≤i≤ m,1≤j≤n. Equality, Additi on, Multiplication Definition 2.1.3. Two matrices AandBare equal if and only if they have t h es a m es i z ea n d aij=bijalli, j. Definition 2.1.4. IfAis any matrix and α∈Fthen the scalar multipli- cation B=αAis defined by bij=αaijalli, j. Definition 2.1.5. IfAandBare matrices of the same size then the sum AandBis defined by C=A+B,w h e r e cij=aij+bijalli, j We can also compute the difference D=A−Bby summing Aand (−1)B D=A−B=A+(−1)B. matrix subtraction. Matrix addition “inherits” many properties from the fieldF. Theorem 2.1.2. IfA, B, C∈Mm,n(F)andα,β∈F,t h e n (1)A+B=B+A commutivity (2)A+(B+C)=(A+B)+C associativity (3)α(A+B)=αA+αB distributivity of a scalar (4) If B=0 (a matrix of all zeros) then A+B=A+0= A (4)(α+β)A=αA+βA 36 CHAPTER 2. MATRICES AND LINEAR ALGEBRA (5)α(βA)=αβA (6)0A=0 (7)α0=0 . Definition 2.1.6. Ifxandy∈Rn, x=(x1...x n) y=(y1...y n). Then the scalar or dot product of xandyis given by x, yX=n3 i=1xiyi. Remark 2.1.1. (i) Alternate notation for the scalar product: x, yX=x·y. (ii) The dot product is de fined only for vectors of the same length. Example 2.1.1. Letx=( 1,0,3,−1) and y=( 0,2,−1,2) thenx, yX= 1(0) + 0(2) + 3( −1)−1(2) =−5. Definition 2.1.7. IfAism×nandBisn×p.L e t ri(A) denote the vector with entries given by the ithrow of A,a n dl e t cj(B) denote the vector with entries given by the jthrow of B. The product C=ABis the m×pmatrix defined by cij=ri(A),cj(B)X where ri(A) is the vector in Rnconsisting of the ithrow of Aand similarly cj(B) is the vector formed from the jthcolumn of B. Other notation for C=AB cij=n k=1aikbkj1≤i≤m 1≤j≤p. Example 2.1.2. Let A=}101 321] and B= 21 30 −11 . Then AB=}12 11 4] . 2.1. BASICS 37 Properties of matrix multiplication (1) If ABexists, does it happen that BAexists and AB=BA?T h e answer is usually no. First AB andBA exist if and only if A∈ Mm,n(F)a n d B∈Mn,m(F). Even if this is so the sizes of ABand BAare different ( ABism×mandBAisn×n) unless m=n. However even if m=nwe may have ABW=BA.S e e t h e e x a m p l e s below. They may be di fferent sizes and if they are the same size (i.e. AandBa r es q u a r e )t h ee n t r i e sm a yb ed i fferent A=[ 1,2]B=}−1 1] AB=[ 1 ] BA=}−1−2 12] A=}12 34] B=}−11 01] AB=}−13 −37] BA=}22 34] (2) If Ais square we de fine A1=A, A2=AA, A3=A2A=AAA An=An−1A=A···A(nfactors) . (3)I= diag(1 ,... , 1). If A∈Mm,n(F)t h e n AIn=Aand ImA=A. Theorem 2.1.3 (Matrix Multiplication Rules). Assume A, B ,a n d C are matrices for which all products below make sense. Then (1)A(BC)=(AB)C (2)A(B±C)=AB±ACand(A±B)C=AC±BC (3)AI=AandIA=A (4)c(AB)=(cA)B (5)A0=0 and0B=0 38 CHAPTER 2. MATRICES AND LINEAR ALGEBRA (6) For Asquare ArAs=AsArfor all integers r, s≥1. Fact: IfACandBCare equal, it does not follow that A=B. See Exercise 60. Remark 2.1.2. We use an alternate notation for matrix entries. For any matrix Bdenote the ( i, j)-entry by ( B)ij. Definition 2.1.8. LetA∈Mm,n(F). (i) De fine the transpose ofA, denoted by AT,t ob et h e n×mmatrix with entries (AT)ij=aji. (ii) De fine the adjoint ofA, denoted by A∗,t ob et h e n×mmatrix with entries (A∗)ij=¯ajicomplex conjugate Example 2.1.3. A=}123 541] AT= 15 24 31  In words ...“The rows of Abecome the columns of AT, taken in the same order.” The following results are easy to prove. Theorem 2.1.4 (Laws of transposes). (1)(AT)T=Aand(A∗)∗=A (2)(A±B)T=AT±BT(and for ∗) (3)(cA)T=cAT(cA)∗=¯cA∗ (4)(AB)T=BTAT (5) If Ais symmetric A=AT 2.1. BASICS 39 (6) If Ais Hermitian A=A∗. More facts about symmetry. Proof. (1) We know ( AT)ij=aji.S o( ( AT)T)ij=aij.T h u s( AT)T=A. (2) (A±B)T=aji±bji.S o( A±B)T=AT±BT. Proposition 2.1.1. (1)Ais symmetric if and only if ATis symmetric. (1)∗Ais Hermitian if and only if A∗is Hermitian. (2) If Ais symmetric, then A2is also symmetric. (3) If Ais symmetric, then Anis also symmetric for all n. Definition 2.1.9. A matrix is called skew-symmetric if AT=−A. Example 2.1.4. The matrix A= 012 −10−3 −23 0  is skew-symmetric. Theorem 2.1.5. (1) If Ais skew symmetric, then Ais a square matrix andaii=0,i=1,... ,n . (2) For any matrix A∈Mn(F) A−AT is skew-symmetric while A+ATis symmetric. (3) Every matrix A∈Mn(F)can be uniquely written as the sum of a skew-symmetric and symmetric matrix. Proof. (1) If A∈Mm,n(F), then AT∈Mn,m(F). So, if AT=−Awe must have m=n.A l s o aii=−aii fori=1,... ,n .S oaii=0f o ra l l i. 40 CHAPTER 2. MATRICES AND LINEAR ALGEBRA (2) Since ( A−AT)T=AT−A=−(A−AT), it follows that A−ATis skew-symmetric. (3) Let A=B+Cbe a second such decomposition. Subtraction gives 1 2(A+AT)−B=C−1 2(A−AT). The left matrix is symmetric while the right matrix is skew-symmetric. Hence both are the zero matrix. A=1 2(A+AT)+1 2(A−AT). Examples. A=J0−1 10o is skew-symmetric. Let B=}12 −14] BT=}1−1 24] B−BT=}03 −30] B+BT=}21 18] . Then B=1 2(B−BT)+1 2(B+BT). An important observation about matri x multiplication is related to ideas from vector spaces. Indeed, two very important vector spaces are associatedwith matrices. Definition 2.1.10. LetA∈M m,n(C). (i)Denote by cj(A): =jthcolumn of A cj(A)∈Cm. We call the subspace of Cmspanned by the columns of Athe column space ofA.W i t h c1(A),...,c n(A) denoting the columns of A 2.1. BASICS 41 the column space is S(c1(A),...,c n(A)). (ii) Similarly, we call the subspace of Cnspanned by the rows of Atherow space ofA.W i t h r1(A),...,r m(A) denoting the rows of Athe row space is therefore S(r1(A),...,r m(A)). Letx∈Cn,w h i c hw ev i e wa st h e n×1m a t r i x x=[x1...x n]T.T h e product Axis defined and Ax=n3 j=1xjcj(A). That is to say, Ax∈S(c1(A),... ,c n(A)) = column space of A. Definition 2.1.11. LetA∈Mn(F). The matrix Ais said to be invertible if there is a matrix B∈Mn(F) such that AB=BA=I. In this case Bis called the inverse ofA, and the notation for the inverse is A−1. Examples. (i) Let A=}13 −12] Then A−1=1 5}2−3 11] . (ii) For n=3w eh a v e A= 12−1 −13−1 −23−1  A−1= 01−1 −13−2 −37−5  A square matrix need not have an inverse, as will be discussed in the next section. As examples, the two matrices below do not have inverses A=}1−2 −12] B= 101 021122  42 CHAPTER 2. MATRICES AND LINEAR ALGEBRA 2.2 Linear Systems The solutions of linear systems is likely the single largest application of ma- trix theory. Indeed, most reasonable problems of the sciences and economicsthat have the need to solve problems of several variable almost without ex-ception are reduced to component parts where one of them is the solutionof a linear system. Of course the entire solution process may have the linear system solver as a relatively small component, but an essential one. Even the solution of nonlinear problems, esp ecially, employ linear systems to great and crucial advantage. To be precise, we suppose that the coe fficients a ij,1≤i≤mand 1≤ j≤nand the data bj,1≤j≤ma r ek n o w n . W ed e fine the linear system for the nunknowns x1,...,x nto be a11x1+a12x2+···+a1nxn=b1 a21x1+a22x2+···+a2nxn=b2 (∗) am1x1+am2x2+···+amnxn=bm The solution set is defined to be the subset of Rnof vectors ( x1,...,x n)t h a t satisfy each of the mequations of the system. The question of how to solve a linear system includes a vast literature of theoretical and computation methods. Certain systems form the model of what to do. In the systemsbelow we note that the first one has three highly coupled (interrelated) variables. 3x 1−2x2+4x3=7 x1−6x2−2x3=0 −x1+3x2+6x3=−2 The second system is more tractable because there appears even to the untrained eye a clear and direct method of solution. 3x1−2x2−x3=7 x2−2x3=1 2x3=−2 I n d e e d ,w ec a ns e er i g h to ffthatx3=−1.Substituting this value into the second equation we obtain x2=1−2=−1.Substituting both x2andx3 into the first equation, we obtain 2 x1−2(−1)−(−1) = 7 ,gives x1=2.The 2.2. LINEAR SYSTEMS 43 solution set is the vector (2 ,−1,−1).The virtue of the second system is that the unknowns can be determined on e-by-one, back substituting those already found into the next equation until all unknowns are determined. Soif we can convert the given system of the first kind to one of the second kind, we can determine the solution. This procedure for solving linear systems is therefore the applications of operations to e ffect the gradual elimination of unknowns from the equations until a new system results that can be solved by direct means. The oper-ations allowed in this process must have precisely one important property:They must not change the solution set by either adding to it or subtracting from it. There are exactly three such operations needed to reduce any set of linear equations so that it can be solved directly. (E1) Interchange two equations.(E2) Multiply any equation by a nonzero constant. (E3) Add a multiple of one equation to another. This can be summarized in the following theorem Theorem 2.2.1. Given the linear system (*). The set of equation opera- tions E1, E2, and E3 on the equations of (*) does not alter the solution setof the system (*). We leave this result to the exercises. Our main intent is to convert these operations into corresponding operations for matrices. Before we do this we clarify which linear systems can have a soltution. First, the system can be converted to matrix form by setting Aequal to the m×nmatrix of coefficients, bequal to the m×1 vector of data, and xequal to the n×1 vector of unknowns. Then the system (*) can be written as Ax=b In this way we see that with c i(A)d e n o t i n gt h e ithcolumn of A,the system is expressible as x1c1(A)+···+xncn(A)=b From this equation it is clear that the system has a solution if and only if the vector bis inS(c1(A),···,cn(A)). This is summarized in the following theorem. 44 CHAPTER 2. MATRICES AND LINEAR ALGEBRA Theorem 2.2.2. An e c e s s a r ya n ds u fficient condition that Ax=bhas a solution is that b∈S(c1(A)...c n(A)). In the general matrix product C=AB, we note that the column space of C⊂column space of A.I nt h ef o l l o w i n gd e finition we regard the matrix A as a function acting upon vectors in one vector space with range in anothervector space. This is entirely similar to the domain-range idea of function theory. Definition 2.2.1. Therange ofA={Ax|x∈R n(o rCn)}. It follows directly from our discussion above that the range of Aequals S(c1(A),... ,c n(A)). Row operations: To solve Ax=bwe use a process called Gaussian elimination , which is based on row operations. Type 1: Interchange two rows. (Notation: Ri←→Rj) Type 2: Multiply a row by a nonzero constant. (Notation: cRi→Ri) Type 3: Add a multiple of one row to another row. (Notation: cRi+Rj→ Rj) Gaussian elimination is the process of reducing a matrix to its RREF using these row operations. Each of these operations is the respective analogue of the equation operations described above, and each can be realized by leftmatrix multiplication. We have the following.Type 1 E 1= 1...... 1...... ... ... 0... ... ... 1... ... ... ...1... ......... ...1... ... ... 1... ... ... 0... ... ... ......1 ......... ......1 rowi row j column icolumn j 2.2. LINEAR SYSTEMS 45 Notation: Ri↔Rj Type 2 E2= 1... ...... 1... ... ... ... c ... ... ... ...1 ...... ...1 rowi column i Notation: cR i Type 3 E3= 1... 1... ... ...... ... ... c ... ... ... ... ...1 ...1 rowj column i Notation: cR i+Rj, the abbreviated form of cRi+Rj→Rj Example 2.2.1. The operations  21 0 02 1 −102 R1←→R2 → 02 1 21 0 −102 4R3 → 02 1 21 0 −408  46 CHAPTER 2. MATRICES AND LINEAR ALGEBRA can also be realized as R1←→ R2: 010 100001  21 0 02 1 −102 = 02 1 21 0 −102  4R 3 : 100 010004  02 1 21 0 −102 = 02 1 21 0 −408  The operations  21 0 02 1 −102 −3R 1+R2 → 2R1+R3 21 0 −6−11 32 2  can be realized by the left matrix multiplications  100 010 201  10 0 −310 00 1  21 0 02 1 −102 = 21 0 −6−11 32 2  Note there are two matrix multiplications them, one for each Type 3 ele- mentary operation. Row-reduced echelon form. To each A∈Mm,n(E) there is a canonical form also in Mm,n(E) which may be obtained by row operations. Called the RREF, it has the following properties. (a) Each nonzero row has a 1 as the first nonzero entry (:= leading one ). (b) All column entries above and below a leading one are zero. (c) All zero rows are at the bottom. (d) The leading one of one row is to the left of leading ones of all lower rows. Example 2.2.2. B= 1200 −1 0010 30001 00000 0 is in RREF. 2.2. LINEAR SYSTEMS 47 Theorem 2.2.3. LetA∈Mm,n(F). Then the RREF is necessarily unique. We defer the proof of this result. Let A∈Mm,n(F). Recall that the row space ofAis the subspace of Rn(orCn) spanned by the rows of A.I n symbols the row space is S(r1(A),... ,r m(A)). Proposition 2.2.1. ForA∈Mm,n(F)the rows of its RREF span the rows space of A. Proof. First, we know the nonzero rows of the RREF are linearly indepen- dent. And all row operations are linear combinations of the rows. Thereforet h er o ws p a c eg e n e r a t e df r o mt h eR R E Fi sc o n t a i n e di nt h er o ws p a c eo fA. If the containment is proper. That is there is a row of Athat is lin- early independent from the row space of the RREF, this is a contradiction because every row of Acan be obtained by the inverse row operations from the RREF. Proposition 2.2.2. IfA∈Mm,n(F)and a row operation is applied to A, then linearly dependent columns of Aremain linearly dependent and linearly independent columns of Aremain linearly independent. Proposition 2.2.3. The number of linearly independent columns of A∈ Mm,n(F)is the same as the number of leading ones in the RREF of A. Proof. LetS={i1...i k}be the columns of the RREF of Ahaving a lead- ing one. These columns of the RREF are linearly independent Thus these columns were originally linearly independent. If another column is linearlyindependent, this column of the RREF is linearly dependent on the columnswith a leading one. This is a contradiction to the above proposition. Proof of Theorem 2.2.3. By the way the RREF is constructed, left-to-right, and top-to-bottom, it should be apparent that if the right most row of the RREF is removed, there results the RREF of the m×(n−1) matrix formed from Aby deleting the nthcolumn. Similarly, if the bottom row of the RREF is removed there results a new matrix in RREF form, though notsimply related to the original matrix A. To prove that the RREF is unique, we proceed by a double induction, first on the number of columns. We take it as given that for an m×1m a t r i x the RREF is unique. It is either the zero m×1 matrix, which would be t h ec a s ei f Awas zero or the matrix with a 1 in the first row and zeros in 48 CHAPTER 2. MATRICES AND LINEAR ALGEBRA the other rows. Assume therefore that the RREF is unique if the number o fc o l u m n si sl e s st h a n n. Assume there are two RREF forms, B1and B2forA.N o w t h e R R E F o f Ais therefore unique through the ( n−1)st columns. The only di fference between the RREF’s B1andB2must occur in the nthcolumn. Now proceed by induction on the number of nonzero rows. Assume that AW=0 . I f Ahas just one row, the RREF of Ais simply thescalar multiple of Athat makes the first nonzero column entry a one. Thus it is unique. If A= 0, the RREF is also zero. Assume now that the RREF is unique for matrices with less than mrows. By the comments above that the only di fference between the RREF’s B1andB2can occur at the (m, n)-entry. That is ( B1)m,nW=(B2)m,n. They are therefore not leading ones. (Why?) There is a leading one in the mthrow, however, because it is a non zero row. Because the row spaces of B1andB2are identical, this results in a contradiction, and therefore the ( m, n)-entries must be equal. Finally, B1=B2.This completes the induction. (Alternatively, the two systems pertaining to the RREF’s must have the same solution set to thesystem Ax=0 . W i t h( B 1)m,nW=(B2)m,n, it is easy to see that the solution sets to B1x=0a n d B2x=0m u s td i ffer.) ¤ Definition 2.2.2. LetA∈Mm,nandb∈Rm(orCn). De fine [A|b]= a11... a 1nb1 a21... a 2nb2 am1... a mnbm  [A|b]i sc a l l e dt h e augmented matrix ofAbyb.[A|b]∈Mm,n+1(F). The augmented matrix is a useful notation for finding the solution of systems using row operations. Identical to other de finitions for solutions of equations, the equivalence of two systems is de fined via the idea of equality of the solution set. Definition 2.2.3. Two linear systems Ax=bandBx=care called equiv- alent if one can be converted to the other by elementary equation opera- tions. It is easy to see that this implies the followingTheorem 2.2.4. Two linear systems Ax=bandBx=care equivalent if and only if both [A|b]and[B|c]have the same row reduced echelon form. We leave the prove to the reader. (See Exercise 23.) Note that the solution set need not be a single vector; it can be null or in finite. 2.3. RANK 49 2.3 Rank Definition 2.3.1. Therank of any matrix A,d e n o t eb y r(A), is the di- mension of its column space. Proposition 2.3.1. (i) The rank of Aequals the number of nonzero rows of the RREF of A, i.e. the number of leading ones. (ii)r(A)=r(AT). Proof. (i) Follows from previous results. (ii) The number of linearly independent rows equals the number of lin- early independent columns. The number of linearly independent rows is the number of linearly independent columns of AT–by de finition. Hence r(A)=r(AT). Proposition 2.3.2. LetA∈Mm,n(C)andb∈Cm.T h e n Ax=bhas a solution if and only if r(A)=r([A|b]),w h e r e [A|b]is the augmented matrix. Remark 2.3.1. Solutions may exist and may not. However, even if a so- lution exists, it may not be unique. Indeed if it is not unique, there is an infinity of solutions. Definition 2.3.2. When Ax=bhas a solution we say the system is con- sistent . Naturally, in practical applications we want our systems to be consistent. When they are not, this can be an indicator that something is wrong withthe underlying physical model. In mathematics, we also want consistentsystems; they are usually far more interesting and o ffer richer environments for study. In addition to the column and row spaces, another space of great impor- tance is the so-called null space, the set of vectors x∈R nfor which Ax=0 . In contrast, when solving the simple single variable linear equation ax=b with aW= 0 we know there is always a unique solution x=b/a.I n s o l v i n g even the simplest higher dimensional systems, the picture is not as clear. Definition 2.3.3. LetA∈Mm,n(F). The null space ofAis defined to be Null( A)={x∈Rn|Ax=0}. It is a simple consequence of the linearity of matrix multiplication that Null( A) is a linear subspace of Rn. That is to say, Null( A) is closed under vector addition and scalar multiplication. In fact, A(x+y)=Ax+Ay= 0+0=0 , i f x, y∈Null( A). Also, A(αx)=αAx=0 ,i f x∈Null( A). We state this formally as 50 CHAPTER 2. MATRICES AND LINEAR ALGEBRA Theorem 2.3.1. LetA∈Mm,n(F).T h e n N u l l (A)is a subspace of Rn w h i l et h er a n g eo f Ais in Rm. Having such solutions gives valuable information about the solution set of the linear system Ax=b. For, if we have found asolution, x, and have any vector z∈Null( A), then x+zis a solution of the same linear system. Indeed, what is easy to see is that if uandvare both solutions to Ax=b, then A(u−v)=Au−Av= 0, or what is the same x−y∈Null( A). This means that to findallsolutions to Ax=b, we need only find a single solution and the null space. We summarize this as the following theorem. Theorem 2.3.2. LetA∈Mm,n(F)with null space Null (A).L e t xbe any nonzero solution to Ax=b. Then the set x+Null(A)is the entire solution set to Ax=b. Example 2.3.1. Find the null space of A=}13 −3−9] . Solution. Solve Ax=0.The RREF for Ais}13 00] .S o l v i n g x1+3x2=0, take x2=t, a “free” parameter and solve for x1to get x1=−3t.Thus every solution to Ax= 0 can be written in the form x=}−3t t] =t}−3 1] t∈R Expressed this way we see that Null( A)=F t}−3 1] |t∈Rk ,a subspace ofR2of dimension 1. Theorem 2.3.3 (Fundamental theorem on rank). A∈Mm,n(F).T h e following are equivalent (a)r(A)=k. (b) There exist exactly klinearly independent columns of A. (c) There exist exactly klinearly independent rows of A. (d) The dimension of the column space of Aisk(i.e. dim( Range A)=k). (e) There exists a set Sof exactly kvectors in Rmfor which Ax=bhas a solution for each b∈S(S). (f) The null space of Ahas dimension n−k. 2.3. RANK 51 Proof. The equivalence of (a), (b), (c) and (d) follow from previous con- siderations. To establish (e), let S={cf1,cf2,... ,c fk}denote the linearly independent column vectors of A.L e t T={ef1,ef2,... ,e fk}⊂Rnbe the standard vectors. Then Aefj=cfj.I fb∈S(S), then b=a1cf1+a2cf2+ ···+akcfk.As o l u t i o nt o Ax=bis given by x=a1ef1+a2ef2+···+akefk. Conversely, if (e) holds, then the set Smust be linearly independent for otherwise Scould be reduced to k−1 or fewer vectors. Similarly if Ahas k+ 1 linearly independent columns then set Scan be expanded. Therefore, the column space of Amust have exactly kvectors. To prove (f) we assume that S={v1,... ,v k}is a basis for the column space of A.L e t T={w1,... ,w k}⊂Rnfor which Awi=vi,i=1,... ,k . By our extension theorem, we select n−kvectors wk+1,... ,w nsuch that U={w1,... ,w k,wk+1,... ,w n}is a basis of Rn. We must have that Awk+1∈S(S). Hence there are scalars b1,... ,b ksuch that Awk+1=A(b1w1+···+bkwk) and thus wI k+1=wk+1−(b1w1+···+bkwk)i si nt h en u l ls p a c eo f A. Repeat this process for each wk+j,j=1,... ,n−k. We generate a total ofn−kvectors {wI k+1,... ,wI n}in this manner. This set must be linearly independent. (Why?) Therefore, the dimension of the null space must beat least n−k. Now we consider a new basis which consists of the original vectors and the n−kvectors {w I k+1,wI k+2,... ,wI n}for which Aw=0 . W e assert that the dimension of the null space is exactly n−k.F o ri f z∈Rnis av e c t o rf o rw h i c h Az=0 ,t h e n zcan be uniquely written as a component z1fromS(T) and a component z2fromS({wI k+1,... ,wI n}). But Az1W=0 andAz2= 0. Therefore Az= 0 is impossible unless the component z1=0 . Conversely, if (f) holds we take a basis for the null space T={u1,u2,... ,u n−k} a n de x t e n dt h eb a s i s TI=T∪{un−k+1,. . . ,u n} toRn.N e x ta r g u es i m i l a r l yt oa b o v et h a t Aun−k+1,A u n−k+2,... ,A u n must be linearly independent, for otherwise there is yet another linearly independent vector that can be added to its basis, a contradiction. Thereforethe column space must have dimension at least, and hence equal to k. The following corollary assembles many consequences of this theorem. 52 CHAPTER 2. MATRICES AND LINEAR ALGEBRA Corollary 2.3.1. (1)r(A)≤min(m, n). (2)r(AB)≤min(r(A),r(B)). (3)r(A+B)≤r(A)+r(B). (4)r(A)=r(AT)=r(A∗)=r(¯A). (5) If A∈Mm(F)andB∈Mm,n(F),a n di f Ais invertible, then r(AB)=r(B). Similarly, if C∈Mn(F)is invertible and B∈Mm,n(F) r(BC)=r(B). (6)r(A)=r(ATA)=r(A∗A). (7) Let A∈Mm,n(F),w i t h r(A)=k.T h e n A=XBY where X∈Mm,k, Y∈Mk,nandB∈Mkis invertible. (8) In particular, every rank 1 matrix has the form A=xyT,w h e r e x∈ Rmandy∈Rn.H e r e xyT= x1y1x1y2... x 1yn ......... xmy1xmy2... x myn . Proof. (1) The rank of any matrix is the number of linearly independent rows, which is the same as the number of linearly independent columns. The maximum this value can be is therefore the maximum of theminimum of the dimensions of the matrix, or r(A)≤min ( m, n). (2) The product ABc a nb ev i e w e di nt w ow a y s . T h e fir s ti sa sas e t of linear combinations of the rows of B,and the other is as a set of linear combinations of the columns of A.In either case the number of linear independent rows (or columns as the case may be) In otherwords, the rank of the product ABcannot be greater than the number of linearly independent columns of Anor greater than the number of linearly independent rows of B.Another way to express this is as r(AB)≤min(r(A),r(B)) 2.3. RANK 53 (3) Now let S={v1,...v r(A)}andT={w1,...,w r(B)}be basis of the column spaces of AandBrespectively. Then, the dimension of the union S∪T={v1,...v r(A),w1,...,w r(B)}cannot exceed r(A)+r(B). Also, every vector in the column space of A+Bis clearly in the span ofS∪T.The result follows. (4) The rank of Ais the number of linearly independent rows (and columns) ofA,which in turn is the number of linearly independent columns of AT,which in turn is the rank of AT.That is, r(A)=rD ATi .Similar proofs hold for A∗and ¯A. (5) Now suppose that A∈Mm(F) is invertible and B∈Mm,n.As we have emphasized many times the rows of the product ABcan be viewed as a set of linear combinations of the rows of B.Since Ahas rank m any set of linearly independent rows of Bremains linearly independent. To see why, let ri(AB)d e n o t et h e ithrow of the product AB. Then it is easy to see that ri(AB)=m3 j=1aijrj(B) Suppose we can determine constants c1,...,c mnot all zero so that 0=m3 j=1ciri(AB)=m3 i=1cim3 j=1aijrj(B) =m3 j=1rj(B)m3 i=1ciaij This linear combination of the rows of Bhas coefficient given by ATc, where c=[c1,...,c k]T.Because the rank of A(and AT)i sm,we can solve this system for any vector d∈Rm.Suppose that the row vectors rjl(B),f=1,...,r (B),are linearly independent. Arrange that the components of dto be zero for indices not included in the set jl,f=1,...,r (B) and not all zero otherwise. Then the conclusion 0=m j=1rj(B)m i=1ciaij=r(B) l=1rjl(B)djlis impossible. Indeed, the same basis of the row space of Bwill be a basis of the row space ofAB.T h i sp r o v e st h er e s u l t . (6) We postpone the proof of this result until we discuss orthogonality. 54 CHAPTER 2. MATRICES AND LINEAR ALGEBRA (7) Place Ain RREF, say ARREF .S i n c e r(A)=kwe know the top k rows of ARREF are linearly independent and the remaining rows are zero. De fineYto be the k×nmatrix consisting of these top krows. DefineB=Ik.Now the rows of Aare linear combinations of these rows. So, de fine the m×kmatrix Xto have rows as follows: The first row of consists of the coe fficients so thatx1jrj(Y)=r1(A). In general, the ithrow of Xis selected so that 3 xijrj(Y)=ri(A) (8) This is an application of (7) noting in this special case that Xis an m×1 matrix that can be interpretted as a vector x∈Rm. Similarly, Yis an 1 ×nmatrix that can be interpretted as a vector y∈Rn. Thus, with I=[ 1 ] ,w eh a v e A=xyT Example 2.3.2. Here is the decomposition of the form given in Lemma 2.3.1 (7). The 3 ×4m a t r i x Ahas rank 2. A= 12 −1 002 −1−23 240 = 1−1 02 −13 20 }10 01]}120 001] =XBY The matrix Yis the RREF of A. Example 2.3.3. Letx=[x1,x2,...,x m]T∈Rmandy=[y1,y2,...y n]T∈ Rn. Then the rank one m×nmatrix xyThas the form xyT= x 1y1x1y2··· x1yn x2y1x2y2 x2yn ......... xmy1xmy2···xmyn  In particular, with x=[ 1,3,5]T,a n d y=[−2,7]T,the rank one 3 ×2m a t r i x xyTis given by xyT= 1 35 [−2,7] = −27 −62 1 −10 35  2.3. RANK 55 Invertible Matrices A subclass matrices A∈Mn(F) that have only the zero kernel is very important in applications and theoretical developments. Definition 2.3.4. A∈Mnis called nonsingular ifAx= 0 implies that x=0 . In many texts such matrices are introduced though an equivalent alter- nate de finition involving rank. Definition 2.3.5. A∈Mnisnonsingular ifr(A)=n. We also say that nonsingular matrices have fullrank. That nonsingular matrices are invertible and conversely together with many other equivalencesis the content of the next theorem. Theorem 2.3.4. [Fundamental theorem on inverses] Let A∈M n(F).T h e n the following statements are equivalent. (a)Ais nonsingular. (b)Ais invertible. (c)r(A)=n. (d) The rows and columns of Aare linearly independent. (e)dim(Range( A)) =n. (f)dim(Null( A)) = 0 . (g)Ax=bis consistent for all b∈Rn(orCn). (h)Ax=bhas a unique solution for every x∈Rn(orCn). (i)Ax=0 has only the zero solution. (j)* 0 is not an eigenvalue of A. (k)* detAW=0. * The statements about eigenvalues and the determinant (det A)o fam a - trix will be clari fied later after they have been properly de fined. They are included now for completeness. 56 CHAPTER 2. MATRICES AND LINEAR ALGEBRA Definition 2.3.6. Two linear systems Ax=bandBx=care called equiv- alent if one can be converted to the other by elementary equation opera- tions. Equivalently, the systems are equivalent if [ A|b] can be converted to [B|c] by elementary row operations. Alternatively, the systems are equivalent if they have the same solution set which means of course that both can be reduced to the same RREF. Theorem 2.3.5. IfA∈Mn(F)andB∈Mn(F)with AB=I,t h e n Bis unique. Proof. IfAB=Ithen for every e1...e nthere is a solution to the system Abi=eifor all 1 = 1 ,2,... ,n .T h u st h es e t {bi}n i=1is linearly independent (because the set {ei}is) and moreover a basis. Similarly if AC=Ithen A(C−B) = 0, and there are ci∈Rn(orCn),i=1, ..., n such that Aci=ei. Suppose for example that c1−b1W=0 . S i n c et h e {bi}n i=1is a basis it follows that c1−b1=Σαjbj, where not all αjare zero. Therefore, A(c1−b1)=ΣαjAbj=ΣαjejW=0. and this is a contradiction. Theorem 2.3.6. LetA∈Mn(F).I fBis a right inverse, AB=I,t h e n B is a left inverse. Proof. DefineC=BA−I+B,a n da s s u m e CW=Bor what is the same thing that Bis not a left inverse. Then AC=ABA−A+AB=(AB)(A)−A+AB =A−A+AB=I This implies that Cis another right inverse of A, contradicting Theorem 2.3.5. 2.4 Orthogonality Let V be a vector space over C.W ed e fine an inner product ·,·XonV×V to be a function from VtoCthat satis fies the following properties: 1.av, wX=av,wXandv,awX=av,wX(ais the complex conjugate ofa) 2.v,wX=w,vX 2.4. ORTHOGONALITY 57 3.u+v,wX=u, wX+v,wX(linearity) 4.u, v+wX=u, vX+u, wX 5.v,vX≥ 0w i t hv,vX=0i fa n do n l yi f v=0. For inner products over real vector spaces, we neglect the complex con- jugate operation. In addition, we want our inner products to de fine anorm as follows: 6. For any v∈V,,v,2=v,vX We assume thoughout the text that all vector spaces with inner products have norms de fined exactly in this way. With the norm and vector vcan benormalized by dilating it to have length 1, say vn=v1 ,v,. The simplest type of inner product on Cnis given by v,wX=n3 i=1xi¯yi We call this the standard inner product. Using any inner product, we can de fine an angle between vectors. Definition 2.4.1. Theangleθxybetween vectors xandyinRnis defined by cosθxy=x, yX ,x,,y, =x, yX (x, xX)1/2(y,yX)1/2. This comes from the well known result in R2 x·y=,x,,y,cosθ which can be proved using the law of cosines. With angle comes the notion of orthogonality. Definition 2.4.2. Two vectors uandvare said to be orthogonal if the angle between them ifπ 2or what is the same thing u, vX=0 . I nt h i sc a s e we commonly write x⊥y. W ee x t e n dt h i sn o t a t i o nt os e t s Uwriting x⊥U to mean that x⊥ufor every u∈U. Similarly two sets UandVare called orthogonal if u⊥vfor every u∈Uandv∈V. 58 CHAPTER 2. MATRICES AND LINEAR ALGEBRA Remark 2.4.1. It is important to note that the notion of orthogonality depends completely on the inner product. For example, the weighted inner product defined byv,wX=n i=1wixi¯yiwhere the wi>0g i v e sv e r yd i fferent orthogonal vectors from the standard inner product. Example 2.4.1. InRnorCnthe standard unit vectors are orthogonal with respect to the standard inner product. Example 2.4.2. In the R3the vectors u=( 1,2,−1) and v=( 1,1,3) are orthogonal because x, yX=1( 1 )+2( 2 ) −1( 3 )=0 Note that in R3the complex conjugate is not written. The set of vectors (x1,x2,x3)∈R3orthogonal to u=( 1,2,−1) satis fies the equation x1+ 2x2−x3= 0 is recognizable as the plane with normal vector u. Definition 2.4.3. We de fine the projection Puvof one vector vin the di- rection of an other vector uto be Puv=u, vX ,u,2u A sy o uc a ns e e ,w eh a v em e r e l yw r i t t e na ne x p r e s s i o nf o rt h em o r ei n - tuitive version of the projection in question given by ,v,cosθuvu ,u,.I n t h e figure below, we show the fundamental diagram for the projection of one vector in the direction of another. /c113uv Pvu If the vectors uandvare orthogonal, it is easy to see that Puv=0. (Why?) Example 2.4.3. Find the projection of the vector v=( 1,2,1) on the vector u=(−2,1,3) 2.4. ORTHOGONALITY 59 Solution. We have Puv=u, vX ,u,2u=(1,2,1),(−2,1,3)X ,(−2,1,3),2(−2,1,3) =1(−2) + 2 (1) + 1 (3) 14(−2,1,3) =5 14(−2,1,3) We are now ready to findorthogonal sets of vectors and orthogonal bases. First we make an important de finition. Definition 2.4.4. LetVbe a vector space with an inner product. A set of vectors S={x1,... ,x n}inVis said to be orthogonal ifxi,xjX=0f o r iW=j.I ti sc a l l e d orthonormal if alsoxi,xiX= 1. If, in addition, Sis a basis it is called an orthogonal basis or orthonomal basis. Note: Sometimes the conditions for orthonormality are written as xi,xjX=δij whereδijis the “Dirac” delta: δij=0 ,iW=j,δii=1 . Theorem 2.4.1. Suppose Uis a subspace of the (inner product) vector space Vand that Uhas the basis S={x1...x k},t h e n Uhas an orthogonal basis. Proof. Definey1=x1 ,|x,|.T h u s y1is the “normalized” x1.N o w d e fine the new orthonormal basis recursively by yI j+1=xj+1−j3 i=1yi,xj+1Xyi yj+1=yI j+1 ,yI j+1, forj=1,2,...,k−1. Then (1)yj+1is orthogonal to y1,. . .,y j (2)yj+1W=0 . In the language above we have yi,yjX=δij. 60 CHAPTER 2. MATRICES AND LINEAR ALGEBRA Basically, what the proof accomplishes is to take the di fferences of the vector from the projections to the others. Referring to the figure above we compute v−Puva sn o t e di nt h e figure below. The process of orthogonal- ization described above is called the Gram—Schmidt process. /c113uv v PvPv uu- Representation of vectors One of the great advantages of orthonormal bases is that they make the representation of vectors particularl y easy. It is as simple as computing an inner product. Let Vbe a vector space with inner product ·,·Xand with subspace Uhaving basis S={u1,u2,...,u k}. Then for every u∈Uwe know there are constants a1,a2,...,a ksuch that x=a1u1+a2u2+···+akuk. Taking the inner product of both sides with ujand applying the orthogo- nality relations x, u jX=a1u1+a2u2+···+akuk.,ujX =k3 j=1aiui.,ujX=aj Thus aj=x, u jX,j=1,2, ..., k ,a n d x=k3 j=1u.,ujXuj Example 2.4.4. One basis of R2is given by the orthonormal vectors S= {u1,u2},w h e r e u1= 1√ 2,1√ 2=T and u2= 1√ 2,−1√ 2=T . The representa- tion of x=[ 3,2]Tis given by x=23 j=1u.,ujXuj=5 2√ 2}1√ 2,1√ 2]T +1 2√ 2}1√ 2,−1√ 2]T 2.4. ORTHOGONALITY 61 Orthogonal subspaces Definition 2.4.5. For any set of vectors Swe de fine S⊥={v∈V|v⊥S} That is, S⊥is the set of vectors orthogonal to S.O f t e n , S⊥is called the orthogonal complement ororthocomplement ofS. For example the orthocomplement of any vector v=[v1,v2,v3]T∈R3is the (unique) plane passing through the origin that is orthogonal to v.I ti se a s y to see that the equation of the plane is x1v1+x2v2+x3v3=0 . For any set of vectors Sthe orthocomplement S⊥has the remarkable property of being a subspace of V, and therefore it is must have an orthog- onal basis. Proposition 2.4.1. Suppose that Vis a vector space with an inner product, andS⊂V.T h e n S⊥is a subspace of V. Proof. Ify1,. . .,y m∈S⊥thenΣaiyi∈S⊥for every set of coe fficients a1,. . .,a minR(orC). Corollary 2.4.1. Suppose that Vi sav e c t o rs p a c ew i t ha ni n n e rp r o d u c t , andS⊂V. (i) If Sis a basis of V,S⊥={0}. (ii) If U=S(S),t h e n U⊥=S⊥. The proofs of these facts are elementary consequences of the proposition. An important decomposition result is based on orthogonality of subspaces.For example, suppose that Vis afinite dimensional inner product space and that Uis a subspace of V.L e t U ⊥be the orthocomplement of U,and let S={u1,u2,...,u k}be an orthonormal basis of U.L e t x∈V.D e fine x1=k j=1x, u jXuj,a n d x2=x−x1. Then it follows that x1∈Uand x2∈U⊥.M o r e o v e r , x=x1+x2. We summarize this in the following. Proposition 2.4.2. LetVis a vector space with inner product ·,·Xand with subspace U. Then every vector x∈Vc a nb ew r i t t e na sas u mo ft w o orthogonal vectors x=x1+x2,w h e r e x1∈Uandx2∈U⊥. 62 CHAPTER 2. MATRICES AND LINEAR ALGEBRA Geometrically what this results asserts is that for a given subspace of an inner product space, every vector has an orthogonal decomposition astwo unique sum of a vector from the subspace and its orthocomplement.We write the vector components as the respective projections of the givenvector to the orthogonal subspaces x 1=PUx x2=PU⊥x Such decompositions are important in the analysis of vector spaces and matrices. In the case of vector spaces, of course, the representation ofvectors is of great value. In the case of matrices, this type of decompositionserves to allow reductions of the matrices while preserving the informationthey carry. 2.4.1 An important equality for matrix multiplication and the inner product LetA∈Mmn(C). Then we know that both A∗AandAA∗(Alternatively, ATAandAATexist) exist, and we can surely inquire about the rank of these matrices. The main result of this section is on the rank of ATA,n a m e l y that r(A)=r(A∗A)=r(AA∗). The proof is quite simple but requires an important equality. Let A∈Mmn(C)a n d v∈Cnandw∈Cm.Then Av, wX=rm3 i=1(Av)i,wiS =m3 i=1n3 j=1aijvj¯wi =n3 j=1vjm3 i=1aij¯wi =n3 j=1vjm3 i=1¯aijwi =n3 j=1vj(A∗w)j =v,A∗wX 2.4. ORTHOGONALITY 63 As a consequence we have A∗Av, wX=Av, AwXif both v, w∈Cn.T h i s important equality allows the adjoint or transpose matrices to be used oneither side of the inner product, as needed. Indeed we shall use this below. Proposition 2.4.3. LetA∈M mn(C)have rank r(A).Then r(A)=r(A∗A)=r(AA∗) Proof. Assume that r(A)=k.Then there are kstandard vectors ej1,..., e jk such that for each l=1,2,...k, the vectors Aejlis one of the linearly independent columns of A.Moreover, it also follows that for every set of constants a1,...,a kthe vector ADalelji W=0.Now A∗ADalelji W=0 follows because ? A∗Ap3 aleljQ ,p3 aleljQ# =? Ap3 aleljQ ,Ap3 aleljQ# =EEEAp3 a leljQEEE2 W=0 This in turn establishes that A∗Acannot be zero on a linear space of di- mension kexcept for the zero element of course, and since the rank of A∗A cannot be larger than kthe result is proved. Remark 2.4.2. This establishes (6) of the Corollary 2.3.1 above. Also, it is easy to see that the result is also true for real matrices. 2.4.2 The Legendre Polynomials When a vector space has an inner product, it is possible to construct anorthogonal basis from any given basis. We do this now for the polynomial space P n(−1,1) and a particular basis. Consider the space the polynomials of degree ndefin e do nt h ei n t e r v a l [−1,1] over the reals .Recall that this is a vector space and has as a basis themonomials\ 1,x ,x2,...,xn‚ .We can de fine an assortment of inner products on this space, but the most common inner product is given by p, qX=81 −1p(x)q(x)dx Verifying the inner product properties is fairly straight forward and we leave it as an exercise. This inner product also de fines a norm ,p,2=81 −1|p(x)|2dx 64 CHAPTER 2. MATRICES AND LINEAR ALGEBRA This norm satis fies the triangle inequality requires an integral version of the Cauchy-Schwartz inequality. Now that we have an inner product and norm, we could proceed to find an othogonal basis of Pn(−1,1) by applying the Gram-Schmidt procedure to the basis\ 1,x ,x2,...,xn‚ . This procedure can be clumsy and tedious. It is easier to build an orthogonal basis from scratch. Following traditionwe will use capital letters P 0,P1,... to denote our orthogonal polynomials. Toward this end take P0= 1. Note we are numbering from 0 onwards so that the polynomial degree will agree with the index. Now let P1=ax+b. For orthogonality, we need 81 −1P0(x)P1(x)dx=81 −11·(ax+b)dx=2b=0 Thus b=0a n d acan be arbitrary. We take a=1.This gives y1=x.Now we assume the model for the next orthogonal function to be y2=ax2+bx+c. This time there are two orthogonality conditions to satisfy. 81 −1P0(x)P2(x)dx=81 −11·D ax2+bx+ci dx=2 3a+2c=0 81 −1P1(x)P2(x)dx=81 −1x·D ax2+bx+ci dx=2 3b=0 We conclude that b=0.From the equation2 3a+2c= 0, we can assign one of the variables and solve for the other one. Following tradition we takec=− 1 2and solve for ato get a=3 2. The next polynomial will be modeled as P3(x)=ax3+bx2+cx+d. Three orthogonality relations need to be satis fied. 81 −1P0(x)P3(x)dx=81 −11·D ax3+bx2+cx+di dx=2 3b+2d=0 81 −1P1(x)P3(x)dx=81 −1x·D ax3+bx2+cx+di dx=2 5a+2 3c=0 81 −1P2(x)P3(x)dx=81 −11 2(3x−1)D ax3+bx2+cx+di dx =3 5a−1 3b+c−d=0 It is easy to see that b=d=0( w h y ? ) a n df r o m2 5a+2 3c=0,we select 2.4. ORTHOGONALITY 65 c=−3 2anda=5 2.Our table of orthogonal polynomials so far is k Pk(x) 0 1 1 x 21 2(3x−1) 31 2D 5x3−3xi Continue in this fashion, generating polynomials of increasing order each orthogonal to all of the lower order ones. P0(x)=1 P1(x)= x P2(x)=3 /2x2−1/2 P3,x)=5 /2x3−3/2x P4(x)=35 8x4−15 4x2+3/8 P5(x)=63 8x5−35 4x3+15 8x P6(x)=231 16x6−315 16x4+105 16x2−5 16 P7(x)=429 16x7−693 16x5+315 16x3−35 16x P8(x)=6435 128x8−3003 32x6+3465 64x4−315 32x2+35 128 P9(x)=12155 128x9−6435 32x7+9009 64x5−1155 32x3+315 128x P10(x)=46189 256x10−109395 256x8+45045 128x6−15015 128x4+3465 256x2−63 256 2.4.3 Orthogonal matrices Besides sets of vectors being orthogonal, there is also a de finition of orthog- onal matrices. The two notions are closely linked. Definition 2.4.6. We say a matrix A∈Mn(C)i sorthogonal ifA∗A=I. The same de finition applies to matrices A∈Mn(R)w i t h A∗replaced by AT. 66 CHAPTER 2. MATRICES AND LINEAR ALGEBRA For example, the rotation matrices (Exercise ??)Bθ=}cosθ−sinθ sinθcosθ] are all orthogonal. A simple consequence of this de finition is that the rows and the columns ofAareorthonormal . We see for example that when Ais orthogonal then (A∗)2=(A−1)2=A−1A−1=(A2)−1. Such a de finition applies, as well to higher powers. For instance, if Ais orthogonal then Amis orthogonal for every positive integer m. One way to generate orthogonal matrices in Cn(orRn)i st ob e g i nw i t h an orthonormal basis and arrange it into an n×nmatrix either as its columns or rows. Theorem 2.4.2. (i) Let {xi},i=1,...,n be an orthonormal basis of Cn or(Rn). Then the matrices U= x1···xn ↓···↓ ··  and V= x1−→ · ...... xn−→ ·  formed by arranging the vectors xias its respective columns or rows are orthogonal. (ii) Conversely, Uis an orthogonal matrix, the sets of its rows and columns are each orthonormal, an d moreover each forms a basis of Cnor (Rn). The proofs are entirely trivial. We shall consider these types of results in more detail later in Chapter 4. In the meantime there are a few moreinteresting results that are direct consequences of the de finition and facts about the transpose (adjoint). Theorem 2.4.3. LetA, B∈M n(C)(orMn(R))be orthogonal matrices. Then(a)Ais invertible and A −1=A∗. (b) For each integer k=0,±1,±2,...,b o t h Akand−Akare orthogonal. (c)AB is orthogonal. 2.5 Determinants This section is about determinants that can be regarded as a measure of singularity of a matrix. More generally, in many applied situations that deal with complex objects, a single number is sought that will in some way 2.5. DETERMINANTS 67 classify an aspect of those objects. The determinant is such a measure for singularity of the matrix. The determinant is di fficult to calculate and of not much practical use. However, it has considerable theoretical value andcertainly has a place of historical interest. Definition 2.5.1. LetA∈M n(F). De fine the determinant ofAto be the value in F detA=3 σXn i=1aiσ(i)~ ·sgnσ whereσis a permutation of the integers {1,2,... ,n }and (1) σdenotes the sum over all permutations (2) sgnσ=s i g no f σ=±1 Atransposition is the exchange of two elements of an ordered list with all others staying the same. With respect to permutations, a transposition of one permutation is another permutation formed by the exchange of twovalues. For example a transposition of {1,4,3,2}is{1,3,4,2}.T h e sign of a given permutation σis (a) +1 ,if the number of transpositions required to bring σto{1,2,... ,n } is even. (b)−1,if the number of transpositions required to bring σto{1,2,... ,n } is odd. Alternatively, and what is the same thing, we may count the number mof transpositions required to bring σto{1,2,... ,n }and to compute the sign is (−1) m. Example 2.5.1. σ1={2,1,3}1↔2−−→ {1,2,3} odd σ2={2,3,1}3↔1−−→ {2,1,3}1↔2−−→ {1,2,3}even sgnσ1=−1s g n σ2=+ 1 Proposition 2.5.1. LetA∈Mn. 68 CHAPTER 2. MATRICES AND LINEAR ALGEBRA (i) If two rows of Aare interchanged to obtain B,t h e n detB=−detA. (ii) Given A∈Mn(F). If any row is multiplied by a scalar c,t h er e s u l t i n g matrix Bhas determinant detB=cdetA. (iii) If any two rows of A∈Mn(F)are equal, detA=0. Proof. (i) Suppose rows i1andi2are interchanged. Now for the given permutations σapply the transposition i1↔i2to getσ1.T h e n n i=1ai1σ(i)=n i=1bi2σ1(i) because ai1σ(i1)=bi2σ1(i2) as bi2j=ai1jandσ1(i2)=σ2(i1) and similarly ai2σ(i2)=bi1σ1(i1). All other terms are equal. In the computation of the full determinant with signs of the permutations, we see that the change is caused only by the fact sgn( σ1)=−sgn(σ). Thus, detB=−detA. (ii) Is trivial. (iii) If two rows are equal then by part (i) detA=−detA and this implies det A=0 . 2.5. DETERMINANTS 69 Corollary 2.5.1. LetA∈Mn.I fAhas two rows equal up to a multiplica- tive constant, it has has determinant zero. What happens to the determinant when two matrices are added. The result is too complicated to write down is not very important. However,when a single vector is added to a row or a column of a matrix, then theresult can be simply stated. Proposition 2.5.2. Suppose A∈M n(F).S u p p o s e Bis obtained from A by adding a vector vto a given row (resp. column) and Cis obtained from Aby replacing the given row (resp. column) by the vector v.T h e n detB=d e t A+d e t C. Proof. Assume the jthrow is altered. Using the de finition of the determi- nant, detA=3 σsgn(σ) ibiσ(i)=3 σsgn(σ)  iW=jbiσ(i) bjσ(j) =3 σsgn(σ)  iW=jaiσ(i) (a+v)jσ(j) =3 σsgn(σ)  iW=jaiσ(i) ajσ(j)+3 σsgn(σ)  iW=jaiσ(i) vjσ(j) detA+d e t C For column replacement the proof is similar, particularly using the alternate representation of the determinant given in Exercise 15. Corollary 2.5.2. Suppose A∈Mn(F)andBis obtained by multiplying a given row (resp. column) of Aby a scalar and adding it to another row (resp. column), then detB=d e t A. Proof. First note that in applying Proposition 2.5.2 Chas two rows equal up to a multiplicative constant. Thus det C=0 . Computing determinants is usually di fficult and many techniques have been devised to compute them out. As is evident from counting, computing 70 CHAPTER 2. MATRICES AND LINEAR ALGEBRA the determinant of an n×nmatrix using the de finition above would require the expression of all n! permutations of the integers {1,2,..., n }and the determination of their signs together with all the concommitant productsand summation. This method is prohibitively costly. Using elementaryrow operations and Gaussian elimination, the evaluation of the determinant becomes more manageable. First we need the result below. Theorem 2.5.1. For the elementary matrices the following results hold. (a) for Type 1 (row interchange) E 1 detE1=−1 (b) for Type 2 (multiply a row by a constant c)E2 detE2=c (c) for Type 3 (add a multiple of one row to another row) E3 detE3=1. Note that (c) is a consequence of Corollary 2.5.2. Proof of parts (a) and (b) are left as exercises. Thus for any matrix A∈Mn(F)w eh a v e det(E1A)=−detA=d e t E1detA det(E2A)=cdetA =d e t E2detA det(E3A)=d e t A =d e t E3detA. Suppose F1...F kis a sequence of row operations to reduce Ato its RREF. Then FkFk−1...F 1A=B. Now we see that detB=d e t ( FkFk−1...F 1A) =d e t ( Fk)d e t ( Fk−1...F 1A) =... =d e t ( Fk)d e t ( Fk−1)...det(F1)d e tA. ForBin RREF andB∈Mn(F), we have that Bis upper triangular. The next result establishes Theorem 2.3.4( k) about the determinant of singular and non singular matrices. Moreover, the determinant of triangular matrices is computed simply as the product of its diagonal elements. 2.5. DETERMINANTS 71 Proposition 2.5.3. LetA∈Mn.T h e n (i) If r(A)<n,t h e n detA=0. (ii) If Ais triangular then detA= aii (iii) If r(A)=n,t h e n detAW=0. Proof. (i) If r(A)<n, then its RREF has a row of zeros, and det A=0b y Theorem 2.5.1. (ii) If Ais triangular the only product without possible zero entries isaii.H e n c e d e t A=aii. (iii) If If r(A)=n, then its RREF has no nonzero rows. Since it is square and has a leading one in each column, it follows that the RREF is the identity matrix. Therefore det AW=0 . Now let A, B∈Mn.I fAis singular the RREF must have a zero row. It follows that det A=0 . I f Ais singular it follows that ABis singular. Therefore 0=d e t AB=d e t AdetB. The same reasoning applies if Bis singular. If AandBare not singular both AandBcan be row reduced to the identity. Let F1...F k1be the row operations that reduce AtoI,a n d G1...G kBbe the row operations that reduce BtoI.T h e n detA=[ d e t ( F1)...det(FkA)]−1 detB=[ d e t ( G1)...det(GkB)]−1. Also I=(GkB...G 1)(FkB...F 1)AB and we have detI=( d e t A)−1(detB)−1detAB. This proves the Theorem 2.5.2. IfA, B∈Mn(F),detAB=d e t AdetB. 72 CHAPTER 2. MATRICES AND LINEAR ALGEBRA 2.5.1 Minors and Determinants The method of row reduction is one of the simplest methods to compute the determinant of a matrix. Indeed, it is not necessary to use Type 2 el- ementary transformation. This results in the computing the determinantas the product of the diagonal elements of the resulting triangular matrixpossibly multiplied by a minus sign. An alternate approach to computingdeterminants using minors is both interesting and useful. However, unless the matrix has some special form, it does not provide a computational al- ternative to row reduction. Definition 2.5.2. LetA∈M n(C).For any row iand column jdefine the (ij)-minor ofAby Mij=d e t A ithrow removed jthcolumn removed The notation A ithrow removed jthcolumn removed denotes the ( n−1)×(n−1) matrix formed from Aby removing the ithrow andjthcolumn. With minors an alternative formulation of the determinant can be given. This method, while not of great value computationally, hassome theoretical importance. For example, the inverse of a matrix can beexpressed using minors. We begin by consideration the determinant. Theorem 2.5.3. LetA∈M n(C).(i) Fix any row, say row k.The de- terminant of Ais given by detA=n3 j=1akj(−1)k+jMkj (ii) Fix any column, say column m. The determinant of Ais given by detA=n3 j=1ajm(−1)m+jMjm Proof. (i) Suppose that k=1.Consider the quantity a11M11 2.5. DETERMINANTS 73 We observe that this is equivalent to all the products of the form sgn(σ)a11·a2σ(2)····· anσ(n) where only permutations that fix the integer (i.e. position) 1 are taken. Thusσ(1) = 1 .Since this position is fixed the signs taken in the determi- nant M11for permutations of n−1 integers are respectively the same as the signs for the new permutation of nintegers. Now consider all permutations that fix the integer 2 in the sense that σ(1) = 2. The quantity a12(−1)1+2M12consists of all the products of the form sgn(σ)a12·a2σ(1)a3σ(3)····· anσ(n) We need here the extra sign change because if the part of the permutation σof the integers {1,3,4,...,n }is of one sign, which is the sign used in the computation of det Mij, then the permutation of σof the integers {1,2,3,4,...,n }is of the other sign, and that sign is sgn(σ). When we proceed to the kthcomponent, we consider permutations that fixt h ei n t e g e r k.T h a t i s , σ(1) = k. In this case the quantity a1k(−1)1+kM1k consists of all products of the form sgn(σ)a1ka2σ(1)····ak−1σ(k−1)ak+1σ(k+1)····· anσ(n) Continuing in this way we exhaust all possible products a1σ(1)·a2σ(2)····· anσ(n)over all possible permutations of the integers {1,2,...,n }.T h i s proves the assertion. The proof for expanding from any row is similar, with only a possible change of sign needed, which is a prescribed. (ii) The proof is similar. Example 2.5.2. Find the determinant of A= 32−1 01 3 12−1  expanding across the first row and then expanding down the second column. Solution. Expanding across the first row gives detA=a11M11−a12M12+a13M12 =3 d e t}13 2−1] −2d e t}03 1−1] −1d e t}01 12] =3 (−7)−2(−3)−(−1) =−14 74 CHAPTER 2. MATRICES AND LINEAR ALGEBRA Expanding across the second column gives detA=−2d e t}03 1−1] +1d e t}3−1 1−1] −2d e t}3−1 03] =−2(−3) + (−2)−2( 9 )=−14 The inverse of the matrix can be formulated in terms of minors, which is formulated below. Definition 2.5.3. LetA∈Mn(C)( o r Mn(R)). De fine the adjugate (or adjoint )m a t r i x ˆAby ˆAij=(−1)i+jMji where Mjiis the jiminor. The adjugate has traditionally been call ed the “adjoint”, but that terminol- ogy is somewhat ambiguous in light of the previous de finition as complex conjugate transpose. Note that it is de fined for all square matrices; when restricted to invertible matrices the inverse appears. Theorem 2.5.4. LetA∈Mn(C)(orMn(R)) be invertible. Then A−1= 1 detAˆA Proof. A quick examination of the ij-entry of the product AˆAyields the following sum n3 j=1aijˆAjk=1 det (A)n3 j=1aij(−1)k+jMkj There are two possibilities. (1) If i=k,then the summation above is the summation to form the determinant as described in Theorem 2.5.3. (2) IfiW=k,the summation is the computation of the determinant of the matrix Awith the k throw replaced by the ithrow. Thus the determinant of a matrix with two identical rows is represented above and this must be zero. We conclude thatp AˆAQ ij=δij,the usual Kronecker ‘delta,’ and the result is proved. Example 2.5.1. Find the adjugate and inverse of A=}24 21] 2.5. DETERMINANTS 75 It is easy to see that ˆA=}1−4 −22] Also det A=−6. Therefore, the inverse A−1=−1 6}1−4 −22] Remark 2.5.1. The notation for cofactors of a square matrix Ais often used ˆaij=(−1)i+jMji Note the reversed order of the subscripts ijand then jiabove. Cramer’s Rule We know now that the solution to the system Ax=bis given by x=A−1b. Moreover, the inverse A−1is given by A−1=ˆA detA,w h e r e ˆAis the adjugate matrix. The the ithcomponent of the solution vector is therefore xi=1 detAn3 j=1ˆaijbj =1 detAn3 j=1(−1)i+jMjibj =detAi detA w h e r ew ed e fine the matrix Aito be the modi fication to Aby replacing its ithcolumn by the vector b. In this way we obtain a very compact formula for the solution of a linear system. Called Cramer’s rule we state thisconclusion as Theorem 2.5.1. (Cramer’s Rule.) Let A∈M n(C)be invertible and b∈ Cn.F o re a c h i=1,, n , define the matrix Aito be the modi fication of A by replacing its ithcolumn by the vector b. Then the solution to the linear system Ax=bis given by components xi=detAi detA,i=1,, n. 76 CHAPTER 2. MATRICES AND LINEAR ALGEBRA Example 2.5.2. Given the matrix A=}24 21] , and the vector b= }2 −1] .Solve the system Ax=bby Cramer’s rule. We have A1=}24 −11] and A2=}22 2−1] and det A1=6,detA2=−6,detA=−6. Therefore x1=−1a n d x2=1 A curious formula The useful formula using cofactors given below will have some consequence when we study positive de finite operators in Chapter ??. Proposition 2.5.1. Consider the matrix B= 0x 1x2··· xn x1a11a12···a1n x2a21a22···a2n ............... x nan1an2 ann  Then detB=−3 ˆa ijxixj where ˆaijis the ij-cofactor of A. Proof. Expand by minors along the top row to get detB=3 (−1)jxjM1j(B) Now expand the matrix of M1j(B)d o w nt h e first column. This gives M1j(B)=3 (−1)i−1xiMij(A) 2.6. PARTITIONED MATRICES 77 Combining we obtain detB=3 (−1)jxjM1j(B) =3 (−1)jxj3 (−1)i−1xiMij(A)= =33 (−1)i+j−1xjxiMij(A) =−33 ˆaijxjxi The reader may note that in the last line of the equation above, we should have used ˆ aij. However, the formulation given is correct, as well. (Why?) 2.6 Partitioned Matrices It is convenient to study partitioned or “blocked” matrices, or more graph- ically said, matrices whose entries are themselves matrices. For example,with I 2denoting the 2 ×2i d e n t i t ym a t r i xw ec a nc r e a t et h e4 ×4m a t r i x w r i t t e ni np a r t i t i o n e df o r ma n de x p a n d e df o r m . A=}aI2cI2 cI2dI2] = a0b0 0a0b c0d0 0c0d  Partitioning matrices allows our attention to focus on certain structural properties. In many applications part ititioned matrices appear in a natural way, with the particular blocks having some system context. Many similarsubclasses and processes apply to partitioned matrices. In speci fics i t u a t i o n s they can be added, multiplied, and inverted, just like regular matrices. It iseven possible to perform “blocked” version of Gaussian elimination. In the few results here, we touch on some of these possibilities. Definition 2.6.1. For each 1 ≤i≤mand 1≤j≤n,letA ijbe an mi×nj matrices where . Then the matrix A= A 11A12··· A1n A21A2n··· A2n ............ Am1Am2···Amn  78 CHAPTER 2. MATRICES AND LINEAR ALGEBRA is a partitioned matrix of order (m)i×(nj). The usual operations of addition and multiplication of partitioned ma- trices can be performed provided each of the operations makes sense. Foraddition of two partitioned matrices AandBit is necessary to have the same numbers of blocks of the respective same sizes. Then A+B= A 11A12··· A1n A21A2n··· A2n ............ Am1Am2···Amn + B11B12··· B1n B21B2n··· B2n ............ Bm1Bm2···Bmn  = A 11+B11 A12+B12··· A1n+B1n A21+B21 A2n+B2n··· A2n+B2n ............ Am1+Bm1Am2+Bm2···Amn+Bmn  For multiplication, the situation is a bit more complicated. For de finiteness, suppose that Bis a partitioned matrix with block sizes s i×tj,w h e r e1 ≤ i≤pand 1≤j≤qThe usual operations to construct C=AB, n3 j=1AijBjk then make sense provided p=nandnj=sj,1≤j≤n. A special category of partitioned matrices are the so-called quasi-triangular matrices, wherein Aij=0i f i>j for the “lower” triangular version. The special subclass of quasi-triangular matrices wherein Aij=0i f iW=jare called quasi-diagonal. In the case of the multiplication of partitioned matrices ( C=AB) with the left multiplicand Aa quasi-diagonal matrix, we have Cik=AiiBik. Thus the multiplication is similar in form to the usual multiplication of matrices where the left multiplicand is a diagonal matrix.In the case of the multiplication of partitioned matrices ( C=AB)w i t ht h e right multiplicand Ba quasi-diagonal matrix, we have C ik=AikBkk.F o r quasi-triangular matrices with square diagonal blocks, there is an interestingresult about the determinant. Theorem 2.6.1. LetAbe a quasi-triangular matrix, where the diagonal blocks A iiare square. Then detA= idetAii 2.7. LINEAR TRANSFORMATIONS 79 Proof. Apply row operations on each vertical block without row interchanges between blocks, without any Type 2 operations. The resulting matrix ineach diagonal block position ( i, i) is triangular. Be sure to multiply one of the diagonal entries by ±1,reflecting the number of row interchanges within a block. The resulting matrix c an still be regarded as partitioned, though the diagonal blocks are now actually upper triangular. Now apply Proposition 2.5.3, noting that the product of each of the diagonal entriespertaining to the i thblock is in fact det Aii. A simple consequence of this result, proved al´ a Gaussian elimination, is contained in the following corollary. Corollary 2.6.1. Consider the partitioned matrix A=}A11A12 A21A22] with square diagonal blocks and with A11invertible. Then the rank of Ais the same as the rank of A11if and only if A22=A21A−1 11A12. Proof. Multiplication of Aby the elementary partitioned matrix E=}I 0 −A21A−1 11I] yields EA =}I 0 −A21A−1 11I]}A11A12 A21A22] =}A11 A12 0A22−A21A−1 11A12] Since Ehas full rank, it follows that rank( EA)=r a n k A.S i n c e EAis quasi-triangular, it follows that the rank of Ai st h es a m ea st h er a n ko f A11 if and only if A22−A21A−1 11A12=0. 2.7 Linear Transformations Definition 2.7.1. A mapping Tfrom RntoRmis called a linear trans- formation if T(x+y)=Tx+Ty∀x, y∈Rn T(ax)=aTx ∀a∈F. 80 CHAPTER 2. MATRICES AND LINEAR ALGEBRA Note: We normally write Txinstead of T(x). Example 2.7.1. T:Rn→Rm.L e t a∈Rmandy∈Rn.T h e n f o r e a c h x∈Rn,Tx=x, yXais a linear transformation. Let S={v1...v n}be a basis of Rn, and de fine the m×nmatrix with columns given by the coordinates of Tv1,Tv 2,... ,Tv n. Then this matrix A=^ Tv1Tv2 Tvn ↓↓ ···↓„ is the matrix representation of Twith respect to the basis S.T h u s , i f x=Σaivi, whence [ x]S=(a1...a n), we have [Tx]S=A[x]S TThere is a duality between all linear transformations from RntoRm and the set Mm,n(F). Note that Mm,n(F) is itself a vector space over F. Hence L(Fn,Fm), the set of linear transformations from FntoFmis likewise. As such it has subspaces. Example 2.7.2. (1) Let ¯ x∈Fn.D efiniteJ={T∈L|T¯x=0}.T h e n Jis a subspace of L(Fn,Fm). (2) Let U={T∈L(Rn,Rn)|Tx≥0i fx≥0},w h e r e {x≥0}means the positive orthant of Rn.Uisnota linear subspace of L(Rn,Rn), though it is a convex set. (3) De fineT:Pn→PnbyTp=d dxp.Tis a linear transformation. Example 2.7.3. Express the linear transformation D:P3→P3given by Dp=d dxp(x) as a matrix with respect to the basis. S={1,x ,x2,x3}.W e have D1=0=0+0 x+0x2+0x3.A l s o [D1]S=[ 0,0,0,0]T similarly [Dx]S=[ 1,0,0,0]T [Dx2]S=[ 0,2,0,0]T [Dx3]S=[ 0,0,3,0]T. 2.7. LINEAR TRANSFORMATIONS 81 Hence [D]S= 0100 0020 00030000 . In this context the di fferentiation operator is rather simple. Example 2.7.4. Consider the linear transformation Tdefined by Tq= 3x d dxq+x2qforq∈P2.Find the matrix representation of T. Solution. First o ffwe notice that this transformation has range in P4.Let’s use the standard bases for this problem. We then determine the coordinates of Tfor vectors in the P2basis {1,x ,x2}in the P4basis {1,x ,x2,x3,x4}.Compute T(1) = x2 T(x)=3 x+x3 TD x2i =6 x2+x4 The coordinates of the input vectors we know are [1 ,0,0]T,[0,1,0]T,and [0,0,1]T.For the output vectors the coordinates are [0 ,0,1,0,0]T,[0,3,0,1,0]T, and [0 ,0,6,0,1]T.So, with respect to these two bases, the matrix of the transformation is A= 000 030 106 010001  Observe that the dimensionality corresponds with the dimentionality of the respective spaces. Example 2.7.5. LetV=R 2,w i t h S0={v1,v2}={[1 0],[1 1]},S1= {w1,w2}=\ [1 2],J−2 1o‚ ,a n d T=Ithe identity. The vectors above are expressed in the standard E={e1,e2},Tvj=Ivj=vj.T o find [vj]S1we solve vj=ajw1+βjw2 82 CHAPTER 2. MATRICES AND LINEAR ALGEBRA v1:}1−2 21]}α1 β1] =}1 0] −→}α1 β1] =}1 5 −2 5] v2:}1−2 21]}α2 β2] =}1 1] −→}α2 β2] =}3 5 −1 5]A tsolve linear systems Therefore S1[I]S0=}1 53 5 −2 5−1 5] ←change of basis matrix If [x]S0=}−1 2] [x]S1=S1[I]S0}−1 2] =1 5}13 −2−1]}−1 2] =1 5}5 0] =}1 0] . Note the necessity of using the standard basis to express the vectors in both bases S0andS1. 2.8 Change of Basis LetVbe a vector space with bases S0={v1...v n}andS1={w1...w n}, and suppose T:V→Vis a linear transformation. We want to find the representation of Tas a matrix that takes a vector xg i v e ni nt e r m so fi t s S0 coordinates and produces the vector Txg i v e ni nt e r m so fi t s S1coordinates. We know that x→[x]S0is well de fined. The action of Tis known if the nvector [ x]S0=}c1...cn] and the vectors Tv1,Tv 2,... ,Tv nare known, for if x=Σcjvj,t h e n Tx=ΣcjTvj,b yl i n e a r i t y . To determine [ Tx]S1we need to convert the Tvj,j=1,...,n to coor- dinates in the other S1basis, This is done as follows. Find [Tvj]S1= t 1j t2j ... tnj j=1,2,... ,n . 2.8. CHANGE OF BASIS 83 Then if x∈V [Tx]S1=[ΣcjTvj]S1=Σcj[Tvj]S1 = 3 jtijcj  = t11... t 1n tn1 tnn  c1 ... cn . This n×narray [ tij] depends on T,S 0andS1but not on x.W ed e fine the S0→S1basis representation of Tto be [ tij], and we write this as S1[T]S0= t11... t 1n ......... tn1... t nn . In the special case that Tis the identity operator the matrix S1[I]S0converts the coordinates of a vector in the basis S0to coordinates in the basis S1.I t is easy to see that S0[I]S1must be the inverse of S1[I]S0and thus S0[I]S1·S1[I]S0=I. We can also establish the equality S1[T]S1=S1[I]S0S0[T]S0S0[I]S1. In this way we see that the matrix representation of Tdepends on the bases involved. If Xi sa n yi n v e r t i b l em a t r i xi n Mn(F)w ec a nw r i t e B=X−1AX. The interpretation in this context is clear X: change of coordinate from one basis to another S0→S1 X−1: change of coordinate S1→S0 A: matrix of the linear transformation in the basis S0 B: matrix of the same linear transformation in the basis S1. With this in mind it seems prudent to study linear transformations in the basis that makes their matrix representation as simple as possible. 84 CHAPTER 2. MATRICES AND LINEAR ALGEBRA Example 2.8.1. LetA=}13 −11] be the matrix representation of a lin- ear transformation given with respect to the standard basis S0={e1,e2}= {(1,0),(0,1)}Find the matrix representation of this transformation with resepect to the basis S1={v1,v2}={(2,1),(1,1)}. Solution. According to the analysis above we need to determine S0[I]S1and S1[I]S0.Of course S1[I]S0=S0[I]−1 S1.Since the coordinates of the vectors inS1are expressed in terms of the basis vectors S0we obtain directly S0[I]S1=}21 11] Its inverse is given by S0[I]−1 S1=}1−1 −12] Assembling these matrices we have the final matrix converted to the new basis. S1[A]S1= S1[I]S0AS0[I]S1 =}1−1 −12]}13 −11]}21 11] =}64 −7−4] Example 2.8.2. Consider the same problem as above except that the ma- trixAi sg i v e ni nt h eb a s i s S1. Find matrix representation of this transfor- mation with resepect to the basis S0. Solution. To solve this problem we need to determine S0[A]S0=S0[I]S1AS1[I]S0. As we already have these matrices, we determine that S0[A]S0= S0[I]S1AS1[I]S0. =}21 11]}13 −11]}1−1 −12] =}−61 3 −48] 2.9. APPENDIX A – SOLVING LINEAR SYSTEMS 85 2.9 Appendix A – Solving linear systems The key to solving linear systems is to reduce the augmented system to RREF and solve the resulting equations. While this may be so, there is anintermediate step that occurs about half way through the computation ofthe RREF where the reduced matrix achieves an upper triangular form. Atthis point the solution can be determined directly by back substitution. To clarify the rules on back substitution, suppose that we have the triangular form  a 11a12···a1n 0a22···a2n ......... 0··· 0anneeeeeeeeeb 1 b2 ... bn  Assuming that the diagonal part consists of all nonzero terms, we can solve this system by back substitution. First solve for x n=bn ann. Now inductively solve for the remaining solution coordinates using the formula xn−j=1 bn−j,n−j^j−13 k=0an−j,n−kxn−k„ ,j =1,2,..., n−1 This inconvenient looking formula can be replaced by xj=1 bjj n3 k=j+1ajkxk ,j =n−1,n−2,..., 1 where the index runs from j=n−1u pt o j=1.The upshot is that the row reduction process can be halted when a triangular-like form has beenattained. The applies as well to nonsingular and non square systems, wherethe the process is stopped when all the leading ones have been identi fied, entries below them have been zeroed out, and all the zero rows are present. The principle reason for using back substitution is to reduce the number of computations required, an important consideration in numerical linearalgebra. In the example below we solve a 3 ×3 nonsingular system. Example 2.9.1. Solve Ax=bwhere A= 120 22−1 −13 2 b= 3 6 −2  86 CHAPTER 2. MATRICES AND LINEAR ALGEBRA Solution. Find the RREF of [ A|b]. Then solve Ax=b.  120 22−1 −13 2eeeeee3 6 −2 −2R 1+R2 → R1+R3 12 0 0−2−1 05 2eeeeee3 0 1  − 1 2R2 → 120 011 2 052eeeeee3 0 1  −5R 2+R3 → 12 0 011 2 00−1 2eeeeee3 01  (∗) − 1 2R3+R2 → 120 010001eeeeee3 1 −2  −2R 3 → 120 011 2 001eeeeee3 0 −2  −2R 2+R1 → 100 010001eeeeee1 1 −2  Hence solving we obtain x 3=−2,x 2=1,andx1=1 . T h i si s fine, but there is a faster way to solve this system. Stop the reduction when the system attains a triangular form at ( ∗).From this point solve to obtain x3=−2.Now back substitute x3= 2 into the second row (equation) to solve for x2.T h u s x2=−1 2(−2) = 1 .Finally, back substitute x3=2 a n d x2= 1 into the first row (equation) to solve for x1.T h u s x1=3−2( 1 )=1 . Sometimes the form ( ∗) is called the row reduced form. Example 2.9.2. Given the augmented system for Ax=bis in RREF.  12000 0 00120 −1 00001 300000 0eeeeeeee4 110  2.9. APPENDIX A – SOLVING LINEAR SYSTEMS 87 Find the solution. Solution. The leading ones occur in columns 1, 3, and 5. The values in columns 2, 4, and 6 can be taken as free parameters. So, take x2=r, x 4=s, andx6=t. Now solving for the other varables we have x1=4−2r x3=1−2s+t x5=1−3t The solution set is comprised of the vectorx=[ 4−2r, r,1−2s+t, s, 1−t, t] T =[ 4 ,0,1,0,1,0]T+r[−2,1,0,0,0,0]T+s[0,0,−2,1,0,0]T+t[0,0,1,0−3,1]T for all r, s, andt.We can rewrite this as the set S=   4 0 1010 +r −2 1 0000 +s 0 0 −2 100 +t 0 0 10 −3 1 eeeeeeeeeeeer, s, t∈RorC   This representation shows better the connection between the free constants and the component vectors that make up the solution. Note this expressionalso reveals the solution of the homogeneous solution Ax=0a st h es e t + r[−2,1,0,0,0,0] T+s[0,0,−2,1,0,0]T+t[0,0,1,0−3,1]Teeer, s, t∈RorC Indeed, this is a full subspace. Example 2.9.3. The RREF can be used to determine the inverse, as well. Given the matrix A∈M n,the inverse is given by the matrix Xfor which AX =I. I nt u r nw i t h x1, ..., x nrepresenting the columns of Xande1, ..., e nrepresenting the standard vectors we see that Axj= ej,j=1,2,...,n . To solve for these vectors, form the augmented matrix [A|ej],j=1,2,..., n and row reduce as above. A massive short cut to this process is to augment all the standard vectors at one and row reduce the 88 CHAPTER 2. MATRICES AND LINEAR ALGEBRA resulting n×2nmatrix [ A|I]. If Ais invertible, its RREF is the identity. Therefore, [A|I]row → operations[I|X] and, of course, A−1=X.T h u s , f o r A= −21 0 1123−2−1  we row reduce [ A|I]a sf o l l o w s [A|I]= −21 0 1123−2−1eeeeee100 010001  row → operations 100 010001eeeeee312 724 −5−1−3  2.10 Exercises 1. Consider the di fferential operator T=2xd dx(·)−4 acting on the vector space of cubic polynomials, P3.Show that Tis a linear transformation andfind a matrix representation of it. Assume the basis is given by {1,x ,x2,x3}. 2. (i) Find matrices AandB, each with positive rank, for which r(A+ B)=r(A)+r(B). (ii) Find matrices AandB,e a c hw i t hp o s i t i v e rank, for which r(A+B) = 0. (iii) Give a method to find two nonzero matrices AandBfor which the sum has any preassigned rank. Of course, the matrix sizes may depend on this value. 3. Find square matrices AandBfor which r(A)=r(B)=2a n df o r which r(AB)=0 . 4. Suppose that A is an m×nmatrix and that xis a solution of Ax=b over the prescribed field. Show that every solution of Ax=bhave the form x+x0,w h e r e x0is a solution of Ax0=0 . 2.10. EXERCISES 89 5. Find a matrix A∈Mnof rank n−1f o rw h i c h r(Ak)=n−kfork≤n. Is it possible to begin this process with a matrix A∈Mnof rank n and for which r(Ak)=n−k+1 f o r k≤n? 6. Show that Ax=bhas a solution if and only if yTb=0i fa n do n l yi f yTA= 0 for some column vector. 7. In R2the linear transformation that rotates any vector by θradians counter clockwise Thas matrix representation with respect to the standard basis given by A=}cosθ−sinθ sinθcosθ] What is the matrix representation with respect to the standard basis of the transformation that rotates any vector by θradians clockwise? What is the relation between the matrices? 8. Show that if B,C∈Mn(F), where Bis symmetric and Cis skew- symmetric, then B=Cimplies that B=C=0 . 9. Prove the general formula for the inverse of the 2 ×2m a t r i x A=}ab cd] isA−1=1 detA}d−b −ca] . 10. Prove Theorem 2.5.1(a). 11. Prove Theorem 2.5.1(b). 12. Prove that every permutation σmust have an inverse σ−1(i . e .σ−1(σ(j)) = j), and the signs of σ−1andσare the same. 13. Show that the sign of every transposition is −1. 14. Prove that det A= σ(−1)sgn(σ)w iaσ(i)iW 15. Prove Proposition 2.5.2 using minors. 16. Suppose that the n×nmatrix Ais singular. Show that each column of the adjugate matrix ˆAis a solution of Ax=0 . ( M c D u ffee, Chapter 3, Theorem 29.) 90 CHAPTER 2. MATRICES AND LINEAR ALGEBRA 17. Suppose that Ais an ( n−1)×nmatrix, and consider the homogeneous system Ax=0f o r x∈Rn.D e finehito be the determinant of the (n−1)×(n−1) matrix formed by removing the ithcolumn of A. Show that the vector h=(h1, ..., h n)Tis a solution to Ax=0 . ( M c D u ffee, Chapter 3, Corollary 29.) 18. Show that A=J1−1 −11o has no inverse by trying to solve AB=IThat is, assume the form B=}ab cd] multiply the matrices ( AandB) together, and then solve for the un- knowns a, b, c, andd. (This is not a very e fficient way to determine inverses of matrices. Try the same thing for any 3 ×3m a t r i x . ) 19. Prove that the elementary equation operations do not change the so- lution set of a linear system. 20. Find the inverses of E1,E2,a n d E3. 21. Find the matrix representation of linear transformation Tthat rotates any vector by θradians counter clockwise (ccw) with respect to the basis S={(2,1),(1,1)}. 22. Consider R3.Suppose that we have angles {θi}3 i=1and pairs of co- ordinate vectors {(e1,e2),(e1,e3),(e2,e3)}.LetTbe the linear trans- formation that successively rotates a vector in the respective planes {(ei1,ei2)}k i=1through the respective angles {θi}k i=1.Find the matrix representation of Twith respect to the standard basis .Prove that it is invertible. 23. Prove Theorem 2.2.4.24. Prove or disprove the equivalence of the linear systems. 2x−3y=−1 x+4y=5−x+4y=3 x+2y=3 25. Find basis for the orthoc omplement of the subspace of R 3spanned by the vectors {[2,1,1]T,[1,1,2]T}. 26. Find basis for the orthoc omplement of the subspace of R3spanned by the vector [1 ,1,1]T. 2.10. EXERCISES 91 27. Consider planar rotations in Rnwith respect to the standard bases elements .Prove that there must ben(n−1) 2of them – discounting the particular angle. Display the general representation of any of them. Prove or disprove that any two of them are commutative. Thatis for two angles {θ i}2 i=1and pairs of coordinate vectors {(ei1,ei2)}k i=1 the respective counter clockwise rotations are commutative. 28. Suppose that A∈Mmk,B∈Mknand both have rank k.Show that the rank of ABisk. 29. Suppose that A∈Mmkhas rank k.P r o v e t h a t ARREF =}Ik 0] where Ikis the identity matrix of size kand 0 is the m−k×kzero matrix. 30. Suppose that B∈Mknhas rank k.P r o v e t h a t BRREF =J Ik0o where Ikis the identity matrix of size kand 0 is the k×n−kzero matrix. 31. Determine and prove a version of Corollary 2.6.1 for 3 ×3b l o c k e d matrices, where we assume the diagonal blocks A11is invertible and wish to conclude the result that the rank of Ais the equal to the rank ofA11. 32. Suppose that we have angles {θi}k i=1and pairs of coordinate vectors {(ei1,ei2)}k i=1.LetTbe the linear transformation that successively rotates a vector in the respective planes {(ei1,ei2)}k i=1through the respective angles {θi}k i=1.Prove that the matrix representation of the linear transformation with respect to any basis must be invertible. 33. The super-diagonal of a matrix is the set of elements ai,i+1.The subdi- agonal of a matrix is the set of elements ai−1,i.A tri-banded matrix is one for which the entries are zero above the super-diagonal and belowthe subdiagonal. Suppose that for an n×ntri-banded matrix T,w e have a i−1,i=a, a ii=0,andai,i+1=c.Prove the following facts: (a) If nis odd det A=0. (b) If n=2mis even det T=(−1)mamcm. 34. For the banded matrix of the previous example, prove the following for the powers TpofT. 92 CHAPTER 2. MATRICES AND LINEAR ALGEBRA (a) If pis odd, prove that ( Tp)ij=0i f i+jis even. (b) If pis even, prove that ( Tp)ij=0i f i+jis odd. 35. Consider the vector space P2(1,2) with inner product de fined byp, qX=$2 1p(x)q(x)dx.Find an orthogonal basis of P2(1,2).(Hint. Begin with the standard basis {1,x ,x2}.Apply the Gram-Schmidt procedure.) 36. For what values of aandbis the matrix below singular A= a21 21 b 1a−2  37. The Vandermonde matrix, de fined for a sequence of numbers {x1,...x n}, is given by the n×nmatrix Vn= 1x 1x2 1···xn−1 1 1x2x2 2···xn−1 2 ............... 1xnx2 n···xn−1 n  Prove that the determinant is given by detV n=n i>j=1(xi−xj) 38. In the case the x-values are the integers {1,...,n },p r o v et h a td e t Vn is divisible byn i=1(i−1)!.(These numbers are called superfactorials.) 39. Prove that for the weighted functional de fined in Remark 2.4.1, it is necessary and su fficient that the weights be strictly positive for it to be an inner product. 40. For what values of aandbis the matrix below singular A= b0a0 00 ba 0a0b ba 00  2.10. EXERCISES 93 41. Find an orthogonal basis for R2from the vectors {(1,2),(2,1)}. 42. Find an orthogonal basis of the subspace of R3spanned by {(1,0,1),(0,1,−1)} 43. Suppose that Vis a vector space with an inner product, and S⊂V. Show that if Sis a basis of V,S⊥={0}. 44. Suppose that Vis a vector space with an inner product, and S⊂V. Show that if U=S(S), then U⊥=S⊥. 45. Let A∈Mmn(F). Show it may not be true that r(A)=r(ATA)= r(AAT) unless F=R, in which case it is true. 46. If A∈Mn(C) is orthogonal, show that the rows and columns of Aare orthogonal. 47. If Ais orthogonal then Amis orthogonal for every positive integer m. (This is a part of Theorem 2.4.3(b).) 48. Consider the polynomial space Pn[−1,1] with the inner product p, qX=$1 −1p(t)q(t)dt.Show that every polynomial p∈Pnfor which p(1) = p(−1) = 0 is orthogonal to its derivative. 49. Consider the polynomial space Pn[−1,1] with the inner product p, qX=$1 −1p(t)q(t)dt.Show that the subspace of polynomials in even pow- ers (e.g. p(t)=t2−5t6) is orthogonal to the subspace of polynomials in odd powers. 50. Let A=}13 −11] be the matrix representation of a linear trans- formation given with respect to the standard basis S0={e1,e2}= {(1,0),(0,1)}Find the matrix representation of this transformation with resepect to the basis S1={v1,v2}={(2,−3),(1,−2)}. 51. Show that the sign of every transposition from the set {1,2, ..., n }is −1. 52. What are the signs of the permutations {7, 6, 5, 4, 3, 2, 1 }and {7,1, 6, 4, 3, 5, 2 }of the integers {1, 2, 3, 4, 5, 6, 7 }? 53. Prove that the sign of the permutation {m, m−1,..., 2,1}is (−1)m. 54. Suppose A, B∈Mn(F). IfAis singular, use a row space argument to show that det AB=0 . 94 CHAPTER 2. MATRICES AND LINEAR ALGEBRA 55. Show that if A∈Mm,nandB∈Mm,nand both r(A)=r(B)=m.I f r(AB)=m−k, what can be said about n? 56. Prove that deteeeeeeeex 1x2x3x4 −x2x1−x4x3 −x3x4x1−x2 −x4−x3−x2x1eeeeeeee=D x 2 1+x2 2+x2 3+x2 4i2 57. Prove that there is no invertible 3 ×3 matrix that has all the same cofactors. What similar statement can be made for n×nmatrices? 58. Show by example that there are matrices AandBfor which lim n→∞An and lim n→∞Bnboth exist, but for which lim n→∞(AB)ndoes not ex- ist. 59. Let A∈M2(C). Show that there is no matrix solution B∈M2(C) toAB−BA=I. What can you say about the same problem with A, B∈Mn(C)? 60. Show by example that if AC=BCthen it does not follow that A=B. However, show that if Cis inveritble the conclusion A=Bis valid. Chapter 3 Eigenvalues and Eigenvectors In this chapter we begin our study of the most important, and certainly the most dominant aspect, of matrix theory. Called spectral theory, it allows usto give fundamental structure theorems for matrices and to develop power tools for comparing and computing w i t hm a t r i c e s . W eb e g i nw i t has t u d y of norms on matrices. 3.1 Matrix Norms We know Mnis a vector space. It is most useful to apply a metric on this vector space. The reasons are manifold, ranging from general information ofa metrized system to perturbation theory where the “smallness” of a matrixmust be measured. For that reason we de fine metrics called matrix norms that are regular norms with one additional property pertaining to the matrix product. Definition 3.1.1. LetA∈M n.R e c a l l t h a t a norm ,,·,,o na n yv e c t o r space sati fies the properties: (i),A,≥0a n d |A,=0i fa n do n l yi f A=0 (ii),cA,=|c|,A,forc∈R (iii),A+B,≤, A,+,B,. There is a true vector product on Mndefined by matrix multiplication. In this connection we say that the norm is submultiplicative if (iv),AB,≤, A,,B, 95 96 CHAPTER 3. EIGENVALUES AND EIGENVECTORS In the case that the norm ,·,satifies all four properties (i) - (iv) we call it amatrix norm . Here are a few simple consequences for matrix norms. The proofs are straightforward. Proposition 3.1.1. Let,·,be a matrix norm on Mn, and suppose that A∈Mn.T h e n (a),A,2≤,A,2,,Ap,≤, A,p,p=2,3,. . . (b) If A2=Athen,A,≥1 (c) If Ais invertible, then ,A−1,≥,I, ,A, (d),I,≥1. Proof. The proof of (a) is a consequence of induction. Supposing that A2=A,we have by the submultiplicativity property that ,A,=EEA2EE≤ ,A,2. Hence ,A,≥ 1, and therefore (b) follows. If Ais invertible, we apply the submultiplicativity again to obtain ,I,=EEAA−1EE≤,A,EEA−1EE, whence (c) follows. Finally, (d) follows because I2=Iand (b) applies. Matrices for which A2=Aare called idempotent . Idempotent matrices turn up in most unlikely places and are useful for applications. Examples. We can easily apply standard vector space type norms, i.e. f1, f2,a n df∞to matrices. Indeed, an n×nmatrix can clearly be viewed as an element of Cn2w i t ht h ec o o r d i n a t es t a c k e di nr o w so f nnumbers each. The trick is usually to verify the submultiplicativity condition (iv).1.f 1.D efine ,A,1=3 i,j|aij| The usual norm conditions (i)—(iii) hold. To show submultiplicativity we write ,AB,1=3 ijeeeee3 kaikbkjeeeee ≤3 ij3 k|aik|bkj| ≤3 ijkm|aik||bmj|=3 i,k|aik|3 mj|bmj| =,A,1,B,1. 3.1. MATRIX NORMS 97 Thus,A,1is a matrix norm. 2.f2.D efine ,A,2= 3 i,ja2 ij 1/2 conditions (i)—(iii) clearly hold. ,A,2is also a matrix norm as we see by application of the Cauchy—Schwartz inequality. We have ,AB,2 2=3 ijX3 kaikbkj~2 ≤3 i,jX3 ka2 ik~X3 mb2 jm~ = 3 i,k|aik|2  3 j,m|bjm|2  =,A,2 2,B,2 2. This norm has three common names: The (a) Frobenius norm, (b) Schur norm, and (c) Hilbert—Schmidt norm. It has considerable importance inmatrix theory. 3.f ∞.D efine for A∈Mn(R) ,A,∞=s u p i,j|aij|=m a x i,j|aij|. Note that if J=[11 11],,J,∞=1 . A l s o J2=2J.T h u s,J2,=2,J,=1W≤ ,J,2.S o,A,∞is not a matrix norm, though it is a vector space norm. We can make it into a matrix norm by ,A,=n,A,∞. Note |||AB|||=nmax i,jeeeee3 kaikbkjeeeee ≤nmax ijnmax k|aik||bkj| ≤n2max i,k|aik|max k,j|bkj| =|||A||| ||| B|||. 98 CHAPTER 3. EIGENVALUES AND EIGENVECTORS In the inequalities above we use the fundamental inequality 3 k|ckdk|≤max k|dk|3 k|ck| (See Exercise 4.) While these norms have some use in general matrix theory, most of the widely applicable norms are those that are subordinate to vectornorms in the manner de fined below. Definition 3.1.2. Let,·,be a vector norm on R n(orCn). For A∈Mn(R) (orMn(C)) we de fine the norm ,A,onMnby ,A,=m a x ,x,=1,Ax,. (T) and call ,A,the norm subordinate to the vector norm. Note the use of the same notation for both the vector and subordinate norms. Theorem 3.1.1. The subordinate norm is a matrix norm and ,Ax,≤ ,A,,x,. Proof. We need to verify conditions (i)—(iv). Conditions (i) and (ii) are obvious and are left to the reader . To show (iii), we have ,A+B,=m a x ,x,=1,(A+B)x,≤max ,x,=1(,Ax,+,Bx,) ≤max ,x,=1,Ax,+m a x ,x,=1,Bx, =,A,+,B,. Note that ,Ax,=,x,Awx ,x,W ≤,A,,x, sinceEEEx ,x,EEE= 1. Finally, it follows that for any x∈R n ,ABx,≤, A,,Bx,≤, A,,B,,x, and therefore ,AB,≤, A,,B,. Corollary 3.1.1. (i),I,=1. 3.2. CONVERGENCE AND PERTURBATION THEORY 99 (ii) If Ais invertible, then ,A−1,≥(,A,)−1. Proof. For (i) we have ,I,=m a x ,x,=1,Ix,=m a x ,x,=1,x,=1. To prove (ii) begin with A−1A=I. Then by the submultiplicativity and (i) 1=,I,≤, A−1,,A, and so,A−1,≥1/,A,. There are many results connected with matrix norms and eigenvectors that we shall explore before long. The relation between the norm of the matrixand its inverse is important in computational linear algebra. The quantity ,A −1,,A,thecondition number of the matrix A. When it is very large, the solution of the linear system Ax=bby general methods such as Gaussian elimination may produce results with considerable error. The conditionn u m b e r ,t h e r e f o r e ,t i p so ffinvestigators to this possibility. Naturally enough t h ec o n d i t i o nn u m b e rm a yb ed i fficult to compute accurately in exactly these circumstances. Alternative and very approximate methods are often used as reliable substitutes for the condition number. A special type of matrix, one for which ,Ax,=,x,for every x∈C, is called an isometry . Such matrices which do not “stretch” any vectors have remarkable spectral properties and play an important roll in spectraltheory. 3.2 Convergence and perturbation theory It will often be necessary to compare o ne matrix with another matrix that isnearby in some sense. When a matrix norm at is hand it is possible to measure the proximity of two matrices by computing the norm of their difference. This is just as we do for numbers. We begin with this study by showing that if the norm of a matrix is less than one, then its di fference with the identity is invertible. Again, this is just as with numbers; that is,if|r|<1, 1 1−ris defined. Let us assume R∈Mnand,, is some norm on Mn. We want to show that if ,R,<1t h e n( I−R)−1exists. Toward this end we prove the following lemma. 100 CHAPTER 3. EIGENVALUES AND EIGENVECTORS Lemma 3.2.1. For every R∈Mn (I−R)(I+R+R2+···+Rn)=I−Rn+1. Proof. This result for matrices is the direct analog of the result for numbers (1−r)(1+r+r2+···+rn)=1−rn+1, also often written as 1+ r+r2+···+rn= 1−rn+1 1−r. We prove the result inductively. If n=1t h er e s u l tf o l l o w sf r o m direct computation, ( I−R)(I+R)=I−R2. Assume the result holds up ton−1. Then (I−R)(I+R+R2+···+Rn−1+Rn)=(I−R)(I+R+···+Rn−1) +(I−R)Rn =(I−Rn)+Rn−Rn+1=I−Rn+1 by our inductive hypothesis. This calculation completes the induction, and hence the proof. Remark 3.2.1. Sometimes the proof is presented in a “quasi-inductive” manner. That is, you will see (I−R)(I+R+R2+···+Rn)=(I+R+R2+···+Rn) −(R+R2+···+Rn+1)(∗) =I−Rn+1 This is usually considered acceptable b ecause the correct induction is trans- parent in the calculation. Below we will show that if ,R,=λ<1, then ( I+R+R2+···)= (I−R)−1.I t w o u l d b e incorrect to apply the obvious fact that ,Rn+1,< λn+1→∞ to draw the conclusion from the equality ( ∗)a b o v ew i t h o u t first establishing convergence of the series∞ 0Rk. A crucial step in showing that an in finite series is convergent is showing that its partial sums satisfy the Cauchy criterion:∞ k=1akconverges if and only if for each ε>0,there exists an integer Nsuch that if m, n > N, theneen k=m+1akee<ε.(See Appendix A.) There is just one more aspect of this problem. While it is easyto establishe the Cauchy criterion for our present situation, we still need toresolve the situation between norm convergence andpointwise convergence . We need to conclude that if ,R,<1t h e nl i m n→∞Rn=0,and by this expression we mean that ( Rn)ij→0 for all 1 ≤i, j≤n. 3.2. CONVERGENCE AND PERTURBATION THEORY 101 Lemma 3.2.2. Suppose that the norm ,·,is a subordinate norm on Mn andR∈Mn. (i) If,R,<ε, then there is a constant Msuch that |pij|<Mε. (ii) If limn→∞,Rn,=0,t h e n limn→∞Rn=0. Proof. (i) If,R,<6, if follows that ,Rx,<6for each vector x,a n db y selecting the standard vectors ejin turn, it follows that from which it follows that,r∗j,<6,w h e r e r∗jdenotes the jthcolumn of R.By Theorem 1.6.2 all norms are equivalent. It follows that there is a fixed constant Mindependent of6andRsuch that |rij|<M6. (ii) Suppose for some increasing subsequence of powers nk→∞ it happens thateee(Rnk)ijeee≥r.Select the standard unit vector e j.A little computation shows that ,Rnkej,≥r,whence,Rnk,≥r, contradicting the known limit limn→∞,Rn,= 0. The conclusion lim n→∞Rn=0f o l l o w s . Lemma 3.2.3. Suppose that the norm ,·,is a subordinate norm on Mn. If,R,=λ<1,t h e n I+R+R2+···+Rk+···converges. Proof. LetPn=I+R+···+Rn. To show convergence we establish that {Pn}is a Cauchy sequence. For n>m we have Pn−Pm=n3 k=m+1Rk Hence ,Pn−Pm,=EEEEEn3 k=m+1RkEEEEE ≤n3 k=m+1,Rk, ≤n3 k=m+1,R,k =n3 m+1λk=λm+1n−m−13 j=0λj ≤λm+1(1−λ)−1→0 102 CHAPTER 3. EIGENVALUES AND EIGENVECTORS where in the second last step we used the inequality, which is valid for 0≤λ≤1.n−m−1 0λj≤∞ 0λj<(1−λ)−1. We conclude by Lemma 3.2.2 that the individual matrix entries of the parital sums converge and thus the series itself converges. Note that this result is independent of the particular norm. In practice it isoften necessary to select a convenient norm to actually carry out or verifyparticular computations are valid. In the theorem below we complete theanalysis of the matrix version of the geometric series, stating that when thenorm of a matrix is less than one, the geometric series based on that matrix converges and the inverse of the di fference with the identity exists. Theorem 3.2.1. IfR∈M n(F)and,R,<1for some norm, then (I−R)−1 exists and (I−R)−1=I+R+R2+···=∞3 k=0Rk. Proof. Apply the two previous lemmas. T h ep e r t u r b a t i o nr e s u l ta l l u d e dt oa b o v ec a nn o wb es t a t e da n de a s i l y proved. In words this result states that i fw eb e g i nw i t ha ni n v e r t i b l em a t r i x and additively perturb it by a su fficiently small amount the result remains invertible. Overall, this is the first of a series of results where what is proved is that some property of a matrix is preserved under additive perturbations. Corollary 3.2.1. IfA, B∈MnandAis invertible, then A+λBis invert- ible for su fficiently small |λ|(inRorC). Proof. A sa b o v ew ea s s u m et h a t ,·,is a norm on Mn(F). It is any easy computation to see that A+λB=A(I+λA−1B). Selectλsufficiently small so that ,λA−1B,=|λ|,A−1B,<1. Then by the theorem above, I+λA−1Bis invertible. Therefore (A+λB)−1=(I+λA−1B)−1A−1 and the result follows. 3.3. EIGENVECTORS AND EIGENVALUES 103 Another way of stating this is to say that if A, B∈MnandAhas a nonzero determinant, then for su fficiently small λthe matrix A+λBalso has a nonzero determinant. This corollary can be applied directly to the identitymatrix itself being perturbed by a rank one matrix. In this case the λcan be speci fied in terms of the two vectors comprising the matrix. (Recall Theorem 2.3.1(8).) Corollary 3.2.2. Letx, y∈R nsatisfy |x, yX|=|λ|<1.T h e n I+xyTis invertible and (I+xyT)−1=I−xyT(1 +λ)−1. Proof. We have that ( I+xyT)−1exists by selecting a norm ,·,consistent with the inner product ·,·X.( F o re x a m p l e ,t a k e ,A,=s u p ,x,2=1,Ax,2,w h e r e ,·,2is the Euclidean norm.) It is easy to see that ( xyT)k=λk−1xyT. Therefore (I+xyT)−1=I−xyT+(xyT)2−(xyT)3+··· =I−xyT+λxyT−λ2xyT+··· =I−xyTXn3 k=0(−λ)k~ . Thus (I+xyT)−1=I−xyT(1 +λ)−1 and the result is proved. In words we conclude that the perturbation of the identity by a small rank 1m a t r i xh a sa computable inverse. 3.3 Eigenvectors and Eigenvalues Throughout this section we will consider only matrices A∈Mn(C)o rMn(R). Furthermore, we suppress the field designation unless it is relevant. Definition 3.3.1. IfA∈Mnandx∈CnorRn.I f t h e r e i s a c o n s t a n t λ∈Cand a vector xW=0f o rw h i c h Ax=λx 104 CHAPTER 3. EIGENVALUES AND EIGENVECTORS we callλaneigenvalue ofAandxits corresponding eigenvector .A l - ternatively, we call xthe eigenvector pertaining to the eigenvalue λ,a n d vice-versa. Definition 3.3.2. ForA∈Mn,d efine (1)σ(A)={λ|Ax=λxhas a solution for a nonzero vector x}.σ(A)i s called the spectrum ofA. (2)ρ(A)= s u p λ∈σ(A)|λ|p or equivalently max λ∈σ(A)|λ|Q .ρ(A)i sc a l l e dt h e spectral radius . Example 3.3.1. LetA=[21 12]. Then λ= 1 is an eigenvalue of Awith eigenvector x=[−1,1]T.A l s oλ= 3 is an eigenvalue of Awith eigenvector x=( 1,1)T. The spectrum of Aisσ(A)={1,3}and the spectral radius of Aisρ(A)=3 . Example 3.3.2. The 3 ×3m a t r i x B= −30 6 −12 9 26 4−4−9 has eigenvalues: −1,−3,1. Pertaining to the eigenvalues are the eigenvectors    3 11   ↔1,   1 10   ↔− 3   3 −2 2   ↔− 1 The characteristic polynomial To say that Ax=λxhas a nontrivial solution ( xW=0 )f o rs o m e λ∈Cis the same as the assertion that ( A−λI)x= 0 has a nontrivial solution. This means that det(A−λI)=0 or what is more commonly written det(λI−A)=0 . From the original de finition (De finition 2.5.1) the determinant is sum of products of individual matrix entries. Therefore, det( λI−A)m u s tb ea polynomial in λ.T h i sm a k e st h ed e finition: 3.3. EIGENVECTORS AND EIGENVALUES 105 Definition 3.3.3. LetA∈Mn. The determinant pA(λ)=d e t (λI−A) is called the characteristic polynomial ofA. Its zeros1are the called theeigenvalues ofA.T h e s e t σ(A) of all eigenvalues of Ais called the spectrum ofA. A simple consequence of the nature of the determinant of det( λI−A)i s the following. Proposition 3.3.1. IfA∈Mn,t h e n pA(λ)has degree exactly n. See Appendix A for basic information on solving polynomials equations p(λ) = 0. We may note that even though A∈Mn(C)h a s n2entries in its de finition, its spectrum is completely determined by the ncoefficients of pA(λ). Procedure. The basic procedure of determining eigenvalues and eigenvec- tors is this: (1) Solve det( λI−A) = 0 for the eigenvalues λand (2) for any given eigenvalue λsolve the system ( A−λI)x= 0 for the pertaining eigenvector(s). Though this procedure is not practical in general it can beeffective for small sized matrices, and for matrices with special structures. Theorem 3.3.1. LetA∈M n.The set of eigenvectors pertaining to any particular eigenvalue is a subspace of the given vector space Cn. Proof. Letλ∈σ(A).The set of eigenvectors pertaining to λis the null space of ( A−λI).The proof is complete by application of Theorem 2.3.1 that states the null space of any matrix is a subspace of the underlying vector space. In light of this theorem the following de finition makes sense and is a most important concept in the study of eigenvalues and eigenvectors. Definition 3.3.4. LetA∈Mnand letλ∈σ(A). The null space of (A−λI)i sc a l l e dt h e eigenspace ofApertaining to λ. Theorem 3.3.2. LetA∈Mn.T h e n (i) Eigenvectors pertaining to di fferent eigenvalues are linearly indepen- dent. 1The zeros of a polynomial (or more generally a function) p(λ) are the solutions to the equation p(λ)=0 . As o l u t i o nt o p(λ)=0i sa l s oc a l l e da root of the equation. 106 CHAPTER 3. EIGENVALUES AND EIGENVECTORS (ii) Suppose λ∈σ(A)with eigenvector xis different from the set of eigenvalues {µ1,...,µ k}⊂σ(A)andVµis the span of the pertaining eigenspaces. Then x/∈Vµ. Proof. (i) Let µ,λ∈σ(A)w i t h µW=λpertaining eigenvectors xandy resepectively. Suppose these vectors are linearly dependent; that is, y=cx. Then µy=µcx=cµx =cAx=A(cx) =Ay =λy This is a contradiction, and (i) is proved. (ii) Suppose the contrary holds, namely that x∈Vµ.T h e n x=a1y1+···+ akymwhere {y1,···,ym}are linearly independent vectors of Vµ.E a c h o f the vectors yiis an eigenvector, we know. Assume the pertaining eigenvalues denoted by µji.T h a ti st os a y , Ayi=µjiyi,for each i=1,...,m . Then λx=Ax=A(a1y1+···+amym) =a1µj1y1+···+akµjmym IfλW=0w eh a v e x=a1y1+···+amym=a1µj1 λy1+···+akµjm λym We know by the previous part of this result that at least two of the co- efficients aimust be nonzero. (Why?) Thus we have two di fferent rep- resentations of the same vector by linearly independent vectors, which is impossible. On the other hand, if λ=0t h e n a1µ1y1+···+akµkyk=0 , which is also impossible. Thus, (ii) is proved. The examples below will illustrate the spectra of various matrices. Example 3.3.1. LetA=Jab cdo .T h e n pA(λ)i sg i v e nb y pA(λ)=d e t (λI−A)=d e t}λ−a−b −cλ−d] =(λ−a)(λ−d)−bc =λ2−(a+d)λ+ad−bc. 3.3. EIGENVECTORS AND EIGENVALUES 107 The eigenvalues are the roots of pA(λ)=0 λ=a+d±0 (a+d)2−4(ad−bc) 2 =a+d±0 (a−d)2+4bc 2. For this quadratic there are three possibilities: (a) Two real rootsc 9different values equal values (b) Two complex roots Here are three 2 ×2 examples that illustrate each possibility. The reader should compute the characteristic polynomials to verify these computation. B1=}1−1 1−1] .Then pB1(λ)=λ2λ=0,0 B2=}0−1 10] .Then pB2(λ)=λ2+1λ=±i B3=}01 10] .Then pB3(λ)=λ2−1λ=±1. Example 3.3.2. Consider the rotation in the x, z-plane through an angle θ B= cosθ0−sinθ 01 0 sinθ0c o sθ  The characteristic polynomial is given by pB(λ)=d e t λ 100 010001 − cosθ0−sinθ 01 0 sinθ0c o sθ   =−1+( 1+2c o s θ)λ−(1 + 2 cos θ)λ 2+λ3 The eigenvalues are 1 ,cosθ+√ cos2θ−1,cosθ−√ cos2θ−1.Whenθ is not equal to an even multiple of π, exactly two of the roots are complex numbers. In fact, they are complex conjugate pairs, which can also be written as cos θ+isinθ,cosθ−isinθ. The magnitude of each eigenvalue 108 CHAPTER 3. EIGENVALUES AND EIGENVECTORS is 1, which means all three eigenvalues lie on the unit circle in the complex plane. An interesting observation is that the characteristic polynomial andhence the eigenvalues are the same regardless of which pair of axes ( x-z, x-y,o ry-z) is selected for the rotation. Matrices of the form Bare actually called rotations. In two dimensions the counter-clockwise rotations through the angle θare given by B θ=}cosθ−sinθ sinθcosθ] The eigenvalues for all θis not equal to an even multiple of πare±i.( S e e Exercise 2.) Example 3.3.3. IfTis upper triangular with diag T=[t11,t22,... ,t nn]. ThenλI−Tis upper triangular with diag[ λ−t11,λ−t22,... ,λ−tnn]. Thus the determinant of λI−Tgives the characteristic polynomial of Tto be pT(λ)=n i=1(λ−tii) The eigenvalues of Tare the diagonal elements of T. By expanding this product we see that pT(λ)=λn−(Σtii)λn−1+ lower order terms. The constant term of pT(λ)i s(−1)nn i=1tii=(−1)ndetT.W ed e fine trT=n3 i=1tii and call it the trace ofT. The same statements apply to lower triangular matrices. Moreover, the trace de finition applies to all matrices, not just to triangular ones, and the result will be the same. Example 3.3.4. Suppose that Ais rank 1. Then there are two vectors w,z∈Cnfor which A=wzT. Tofind the spectrum of Awe consider the equation Ax=λx 3.3. EIGENVECTORS AND EIGENVALUES 109 or z,xXw=λx. From this we see that x=wis an eigenvector with eigenvalue z,wX.I f z⊥x, x is an eigenvector pertaining to the eigenvalue 0. Therefore, σA={z,wX,0}. T h ec h a r a c t e r i s t i cp o l y n o m i a li s pλ(λ)=(λ−z,wX)λn−1. Ifwandzare orthogonal then pA(λ)=λn. Ifwandzare not orthogonal though there are just two eigenvalues, we say that 0 is an eigenvalue of multiplicity n−1, the order of the factor ( λ−0). Alsoz,wXhas multiplicity 1. This is the subject of the next section. For instance, suppose w=( 1,−1,2) and z=( 0,1,−3). Then spectrum of the matrix A=wzTis given by σ(A)={−7,0}. The eigenvalue pertain- ing toλ=−7i swand we may take x=( 0,3,1) and ( c,3,1), for any c∈C, to be eigenvectors pertaining to λ=0 . N o t et h a t {w,(0,3,1),(1,3,1)}form ab a s i sf o r R3. To complete our discussion of characteristic polynomials we prove a re- sult that every nthdegree polynomial with lead coe fficient one is the char- acteristic polynomial of some matrix. You will note the similarity of thisresult and the analogous result for di fferential systems. Theorem 3.3.3. Every polynomial of n thdegree with lead coe fficient 1, that is q(λ)=λn+b1λn−1+···+bn−1λ+bn is the characteristic polynomial of some matrix. Proof. We consider the n×nmatrix B= 01 0 ··· 0 00 1 0 ...... −b n−bn−1 −b1  110 CHAPTER 3. EIGENVALUES AND EIGENVECTORS ThenλI−Bhas the form λI−B= λ−10 ··· 0 0λ−10 ... λ... b nbn−1 λ+b1  Now expand in minors across the bottom row to get det (λI−B)= b n(−1)n+1det −10 ··· 0 λ−10 λ...  +bn−1(−1)n+2det λ0··· 0 0−10 ...λ... +··· +b1(−1)n+ndet λ−10 ··· 0λ−1 ... λ...  =bn(−1)n+1(−1)n−1+bn−1(−1)n+2λ(−1)n−2+··· +(λ+b1)(−1)n+nλn−1 =bn+bn−1λ+···+b1λn−1+λn which is what we set out to prove. (The reader should check carefully the term with bn−2to fully understand the nature of this proof.) Multiplicity LetA∈Mn(C). Since pA(λ) is a polynomial of degree exactly n,i tm u s t have exactly neigenvectors (i.e. roots) λ1,λ2,... ,λncounted according to multiplicity. Recall that the multiplicity of an eigenvalue is the number oftimes the monomial ( λ−λ i) is repeated in the factorization of pA(λ). For example the multiplicity of the root 2 in the polynomial ( λ−2)3(λ−5) is 3. Suppose µ1,... ,µ kare the distinct eigenvalues with multiplicities m1,m2,... ,m krespectively. Then the characteristic polynomial can be 3.3. EIGENVECTORS AND EIGENVALUES 111 rewritten as pA(λ)=d e t (λI−A)=n 1(λ−λi) =k 1(λ−µi)mi. More precisely, the multiplicities m1,m2,... ,m kare called the algebraic multiplicities of the respective eigenvalues. This factorization will be very useful later. We know that for each eigenvalue λof any multiplicity m,t h e r em u s t be at least oneeigenvector pertaining to λ. What is desired, but not always possible, is to findµlinearly independent eigenvectors corresponding to the eigenvalue λof multiplicity m. This state of a ffairs makes matrix theory at once much more challenging but also much more interesting. Example 3.3.3. For the matrix A= 20 0 07 41 4√ 3 01 4√ 35 4  the characteristic polynomial is det(λI−A)=pA(λ)=λ3−5λ2+8λ− 4,which can be factored as pA(λ)=(λ−1) (λ−2)2We see that the eigenvalues are 1, 2, and 2. So, th e multiplicity of the eigenvalue λ=1 is 1, and the multiplicity of the eigenvalue λ=2i s2 . T h ee i g e n s p a c e pertaining to the eigenvalue λ= 1 is generated by the vector 0 −1 3√ 3 1 , and the dimension of this eigenspace is one. The eigenspace pertaining to λ=2i sg e n e r a t e db y 1 00 and 0 1 1 3√ 3 .(That is, these two vectors form a basis of the eigenspace.) To summarize, for the given matrix there is one eigenvector for the eigenvalue λ= 1, and there are two linearly independent eigenvectors for the eigenvalue λ= 2. The dimension of the eigenspace pertaining to λ= 1 is one, and the dimension of the eigenspace pertaining toλ=2i st w o . Now contrast the above example where the eigenvectors span the space C3 and the next example where we have an eigenvalue of multiplicity three but the eigenspace is of dimension one. 112 CHAPTER 3. EIGENVALUES AND EIGENVECTORS Example 3.3.4. Consider the matrix A= 2−10 020 102  The characteristic polynomial is given by ( λ−2)3.Hence the eigenvalue λ= 2 has multiplicity three. The eigenspace pertaining to λ=2i sg e n e r - ated by the single vector [0 ,0,1]TTo see this we solve (A−2I)x=  2−10 020102 −2 100 010001  x = 0−10 000100 x=0 The row reduced echelon form for 0−10 000 100 is 100 010 000 .F r o m this it is apparent that we may take x 3=t,but that x1=x2=0.Now assign t= 1 to obtain the generating vector [0 ,0,1]T.This type of example and its consequences seriously complexi fies the study of matrices. Symmetric Functions Definition 3.3.5. Letnbe a positive integer and Λ={λ1,λ2,... ,λn}be given numbers. Suppose that kis a positive integer with 1 ≤k≤n.T h e kthelementary symmetric function on theΛis defined by Sk(λ1,... ,λn)=3 1≤i1<···<ik≤nk j=1λij. It is easy to see that S1(λ1,... ,λn)=n3 1λi Sn(λ1,... ,λn)=n 1λi. 3.3. EIGENVECTORS AND EIGENVALUES 113 For a given matrix A∈Mnthere are nsymmetric functions de fined with respect to its eigenvalues Λ={λ1,λ2,... ,λn}. The symmetric functions are sometimes called the invariants of matrices as they are invariant undersimilarity transformations that will be in Section 3.5. They also furnishdirectly the coe fficients of the characteristic polynomial. Thus specifying thensymmetric functions of an n×nmatrix is su fficient to determine its eigenvalues. Theorem 3.3.4. LetA∈M nhave symmetric functions Sk,k=1,2, ..., n . Then det(λI−A)=n i=1(λ−λi)=λn+n3 k=1(−1)kSkλn−k. Proof. The proof is a consequence of actually expanding the productn i=1(λ− λi). Each term in the expansion has exactly nterms multiplied together that are combinations of the factor λand the−λI is. For example, for the power λn−kthe coefficient is obtained by computing the total number of products ofk“different”−1λI is.( T h e t e r m d i fferent is in quotes because it refers to different indices not actual values.) Co l l e c t i n ga l lt h e s et e r m si sa c c o m - plished by addition. Now the number of ways we can obtain products ofthese kdifferent (−1)λ I isis easily seen to be the number of sequences in the set{1≤i1<···<ik≤n}. The sum of these is clearly ( −1)kSk,w i t ht h e (−1)kfactor being the collected product of k−1’s. Two of the symmetric functions are familiar. In the following we restate this using familiar terms. Theorem 3.3.5. LetA∈Mn(C).T h e n pA(λ)=d e t (λI−A)=n3 k=0pkλn−k where p0=1 and (i)p1=−tr(A)=−n3 1aii (ii) pn=−detA. (iii) pk=(−1)kSk,for1<k<n . 114 CHAPTER 3. EIGENVALUES AND EIGENVECTORS Proof. Note that pA(0) =−det(A)=pn. This gives (ii). To establish (i), we consider det λ·a 11−a12 ...−a1n −a21λ−a22 −a2n ...... −an1 ... λ−ann . Clearly the productn 1(λ−aii) is one of the selections of products in the calculation process. In every other product there must be no more than n−2 diagonal terms. Hence pA(λ)=d e t (λI−A)=n 1(λ−aii)+pn−2(λ), where pn−2(λ) is a polynomial of degree n−2. The coe fficient ofλn−1is −n 1aiiby the Theorem 3.3.4, and this is (i). As a final note, observe that the characteristic polynomial is de fined by knowing the nsymmetric functions. However, the matrix itself has n2en- tries. Therefore, one may expect that knowing the only characteristic poly- nomial of a matrix is insu fficient to characterize it. This is correct. Many matrices having rather di fferent properties can have the same characteristic polynomial. 3.4 The Hamilton-Cayley Theorem The Hamilton-Cayley Theorem opens the doors to a finer analysis of a ma- trix through the use of polynomials, which in turn is an important tool of spectral analysis. The results states that any square matrix satis fies its own charactertic polynomial, that is pA(A)=0 . T h ep r o o fi sn o td i fficult, but we need some preliminary results about fac toring matrix-valued polynomials. Preceding that we need to consider matrix polynomials in some detail. Matrix polynomials One of the very important results of matrix theory is the Hamilton-Cayleytheorem which states that a matrix satis fies its only characteristic equation. This implies we need the notion of a matrix polynomial. It is an easy idea 3.4. THE HAMILTON-CAYLEY THEOREM 115 – just replace the coe fficients of any polynomial by matrices – but it bears some important consequences. Definition 3.4.1. LetA0,A1,... ,A mbe square n×nmatrices. We can define the polynomial with matrix coe fficients A(λ)=A0λm+A1λm−1+···+Am−1λ+Am. Thedegree ofA(λ)i sm,p r o v i d e d A0W=0 . A(λ) is called regular if det A0W= 0. In this case we can construct an equivalent monic2polynomial. ˜A(λ)=A−1 0A(λ)=A−1 0A0λm+A−1 0A1λm−1+···+A−1 0Am =Iλm+˜A1λm−1+···+˜Am. The algebra of matrix polynomials mim ics the normal polynomial algebra. Let A(λ)=m3 0Aiλm−kB(λ)=m3 0Biλm−k. (1)Addition: A(λ)±B(λ)=m3 0(Ak±Bk)λm−k. (2)Multiplication: A(λ)B(λ)=m3 i=0λmwm3 k=0AiBm−iW The termm k=0AiBm−iis called the Cauchy product of the sequences. Note that the matrices Ajalways multiply on the left of the Bk. (3)Division: LetA(λ)a n d B(λ)b et w om a t r i xp o l y n o m i a l so fd e g r e e m(as above) and suppose B(λ) is regular, i.e. det B0W=0 . W es a y thatQr(λ)a n d Rr(λ)a r eright quotient andremainder ofA(λ)u p o n division by B(λ)i f A(λ)=Qr(λ)B(λ)+Rr(λ)( 1 ) 2Recall that a monic polynomial is a polynomial where coe ffic i e n to ft h eh i g h e s tp o w e r is one. For matrix polynomials the corresponding coe fficient is I, the identity matrix. 116 CHAPTER 3. EIGENVALUES AND EIGENVECTORS i ft h ed e g r e eo f Rr(λ)i slessthan that of B(λ). Similarly Qf(λ)a n d Rf(λ) are respectively the leftquotient andremainder ofA(λ)u p o n division by B(λ)i f A(λ)=B(λ)Qf(λ)+Rf(λ)( 2 ) i ft h ed e g r e eo f Rf(λ)i slessthan that of B(λ). I nt h ec a s e( 1 )w es e e A(λ)B−1(λ)=Qr(λ)+Rr(λ)B−1(λ), (3) which looks much likea b=q+r b, a way to write the quotient and remainder of a divided by bwhen aandbare numbers. Also, the form (3) may not properly exist for all λ. Lemma 3.4.1. LetBi∈Mn(C),i =0,... , n with B0nonsingular. Then the polynomial B(λ)=B0λn+B1λn−1+···+Bn is invertible for su fficiently large |λ|. Proof. We factor B(λ)a s B(λ)= B0λnD I+B−1 0B1λ−1+···+B−1 0Bnλ−ni =B0λnJ I+λ−1D B−1 0B1+···+B−1 0Bnλ1−nio Forλ>1, the norm of the term B−1 0B1+···+B−1 0Bnλ1−nis bounded by EEB−1 0B1+···+B−1 0Bnλ1−nEE≤EEB−1 0B1EE+EEB−1 0B2EE|λ|−1+···EEB−1 0BnEE|λ|1−n ≤p 1+|λ|−1+···+|λ|1−nQ max 1≤i≤nEEB−1 0BiEE ≤1 1−|λ|−1max 1≤i≤nEEB−1 0BiEE Thus the conditions of our perturbation theorem hold and for su fficiently large |λ|,it followsEEλ−1D B−1 0B1+···+B−1 0Bnλ1−niEE<1.Hence I+λ−1D B−1 0B1+···+B−1 0Bnλ1−ni is invertible and therefore B(λ) is also invertible. 3.4. THE HAMILTON-CAYLEY THEOREM 117 Theorem 3.4.1. LetA(λ)andB(λ)be matrix polynomials in Mn(C)or (Mn(R)). Then both left and right division of A(λ)byB(λ)is possible and the respective quotients an d remainders are unique. Proof. We proceed by induction on deg B, and clearly if deg B=0t h er e s u l t holds. If deg B=p> degA(λ)=mthen the result follows simply. For, takeQr(λ)=0a n d Rr(λ)=A(λ). The conditions of right division are met. Now suppose that p≤m. It is easy to see that A(λ)=A0λm+A1λm−1+···+Am−1λ+Am =A0B−1 0λm−pB(λ)−p3 j=1A0B−1 0Bjλm−j+m3 j=1Ajλm−j =Q1(λ)B(λ)+A1(λ) where deg A1(λ)<degA(λ). Our inductive hypothesis assumed the division was possible for matrix polynomials A(λ)o fd e g r e e <p.T h e r e f o r e , A1(λ)= Q2(λ)B() + R(λ), where the degree of B(λ)<p . Finally, with Q(λ)= Q(λ)+Q2(λ), there results A(λ)=Q(λ)B() +R(λ). To establish uniqueness we assume two right divisors and quotients have been determined. Thus A=Qr1(λ)B(λ)+Rr1(λ) A=Qr2(λ)B(λ)+Rr2(λ) Subtract to get 0=( Qr1(λ)−Qr2(λ))B(λ)+Rr1(λ)−Rr2(λ). IfQr1(λ)−Qr2(λ)W= 0, we know the degree of ( Qr1(λ)−Qr2(λ))B(λ)i s greater than the degree of R1(λ)−R2(λ). This contradiction implies that Qr1(λ)−Qr2(λ)=0 ,w h i c hi nt u r ni m p l i e st h a t Rr1(λ)−Rr2(λ)=0 . H e n c e the decomposition is unique. Hamilton-Cayley Theorem Let B(λ)=B0λm+B1λm−1+···+Bm−1λ+Bm with B0W= 0. We can also, write B(λ)=m i=0λm−iBi. Both versions are the same. However, when A∈Mn(F), there are two possible evaluations of B(A). 118 CHAPTER 3. EIGENVALUES AND EIGENVECTORS Definition 3.4.2. LetB(λ),A∈Mn(C)( o r Mn(R)) and B(λ)=B0λm+ B1λm−1+···+Bm−1λ+Bm.D efine B(A)=B0Am+B1Am−1+···+Bm “right value” B(A)=AmB0+Am−1B1+···+Bm “left value” The generalized B´ ezout theorem gives the remainder of B(λ)d i v i d e db y λI−A.I nf a c t ,w eh a v e Theorem 3.4.2 (Generalized B´ ezout Theorem). The right division of B(λ)byλI−Ahas remainder Rr(λ)=B(A) Similarly, the left division of B(λ)by(λI−A)has remainder Rf(λ)=B(A). Proof. In the case deg B(λ)=1 ,w eh a v e B0λ+B1=B0(λI−A)+B0A+B1. The remainder Rr(λ)=B0A+B1=B(A). Assume the result holds for all polynomials up to degree p−1. We have B(λ)=B0λp+B1λp−1+···+Bp =B0λp−1(λI−A)+B0Aλp−1+B1λp−1+··· =B0λp−1(λI−A)+B1(λ) where deg B1(λ)≤p−1. By induction B(λ)=B0λp−1(λI−A)+Qr(λ)(λI−A)+B1(A) B1(A)=( B0A+B1)Ap−1+B2Ap−2+···+Bp−1A+Bp=B(A). This proves the result. Corollary 3.4.1. (λI−A)divides B(λ)if and only if B(A)=0 (resp B(A)=0 ). Combining the B´ ezout result and the adjoint formulation of the matrix inverse, we can establish the important Hamilton-Cayley theorem. Theorem 3.4.3 (Hamilton-Cayley). LetA∈Mn(C)(orMn(R))w i t h characteristic polynomial pA(λ).T h e n pA(A)=0 . 3.4. THE HAMILTON-CAYLEY THEOREM 119 Proof. Recall the adjoint formulation of the inverse as ˆC=1 det·CCT ij=C−1. Now let B=adj(A−λI). Then B(λI−A)=d e t (λI−A)I (λI−A)B=d e t (λI−A)I. These equations show that p(λ)I=d e t (λI−A)Iis divisible on the right andthe left by ( λI−A) without remainder. It follows from the generalized B´ezout theorem that this is possible only if pA(A)=0 . LetA∈Mn(C).Now that we know any Asatisfies its characteristic polynomial, we might also ask if there are polynomials of lower degree that it also satis fies. In particular, we will study the so-called minimal polynomial that a matrix satis fies. The nature of this polynomial will shed considerable light on the fundamental structure of A. For example, both matrices below have the same characteristic polynomial P(λ)=(λ−2)3. A= 200 020002 andB= 210 021002  Henceλ= 2 is an eigenvalue of multiplicity three. However, Asatifies the much simpler first degree polynomial ( λ−2) while there is no polynomial of degree less that three that Bsatisfies. By this time you recognize that Ahas three linearly independent eigenvectors, while Bhas only one eigenvector. We will take this subject up in a later chapter. Biographies Arthur Cayley (1821-1895), one of the most proli fic mathematicians of his era and of all time, born in Richmond, Surrey, and studied mathematics atCambridge. For four years he taught at Cambridge having won a Fellowshipand, during this period, he published 28 papers in the Cambridge Mathe-matical Journal. A Cambridge fellowship had a limited tenure so Cayley had to find a profession. He chose law and was admitted to the bar in 1849. He spent 14 years as a lawyer, but Cayley always considered it as a meansto make money so that he could pursue mathematics. During these 14 yearsas a lawyer Cayley published about 250 mathematical papers! Part of thattime he worked in collaboration with James Joseph Sylvester 3(1814- 3In 1841 he went to the United States to become professor at the University of Virginia, but just four years later resigned and returned to England. He took to teaching private 120 CHAPTER 3. EIGENVALUES AND EIGENVECTORS 1897), another lawyer. Together, but not in collaboration, they founded the algebraic theory of invariants 1843. In 1863 Cayley was appointed Sadleirian professor of Pure Mathemat- ics at Cambridge. This involved a very large decrease in income. However Cayley was very happy to devote himself entirely to mathematics. He pub-lished over 900 papers and notes covering nearly every aspect of modernmathematics. The most important of his work is in developing the algebra of matrices, work in non-Euclidean geometry and n-dimensional geometry. Importantly, he also clari fied many of the theorems of algebraic geometry that had previ- ously been only hinted at, and he was among the first to realize how many different areas of mathematics were linked together by group theory. As early as 1849 Cayley wrote a paper l inking his ideas on permutations with Cauchy’s. In 1854 Cayley wrote two papers which are remarkable forthe insight they have of abstract groups. At that time the only knowngroups were groups of permutations and even this was a radically new area, yet Cayley de fines an abstract group and gives a table to display the group multiplication. Cayley developed the theory of algebraic invariance, and his develop- ment of n-dimensional geometry has been applied in physics to the study of the space-time continuum. His work on matrices served as a founda- tion for quantum mechanics, which was developed by Werner Heisenberg in1925. Cayley also suggested that Euclidean and non-Euclidean geometryare special types of geometry. He united projective geometry and metricalgeometry which is dependent on sizes of angles and lengths of lines. Heinrich Weber (1842-1913) was born and educated in Heidelberg, where he became professor 1869. He then taught at a number of institutions inGermany and Switzerland. His main work was in algebra and number theory. He is best known for his outstanding text Lehrbuch der Algebra published in 1895. Weber worked hard to connect the various theories even fundamental concepts such as a fie l da n dag r o u p ,w h i c hw e r es e e na st o o l sa n dn o t properly developed as theories in his Die partiellen Di fferentialgleichungen der mathematischen Physik 1900-01, which was essentially a reworking of a book of the same title based on lectures given by Bernhard Riemann and pupils and had among them Florence Nightingale. By 1850 he became a barrister, and by 1855 returned to an academic life at the Royal Military Academy in Woolwich, London. Hereturned to the US again in 1877 to become prof essor at the new Johns Hopkins University, but returned to England once again in 1877. Sylvester coined the term ‘matrix’ in 1850. 3.4. THE HAMILTON-CAYLEY THEOREM 121 written by Karl Hattendor ff. Etienne B´ ezout (1730-1783) was a mathematician who represents a char- acteristic aspect of the subject at that time. One of the many successful textbook projects produced in the 18thcentury was B´ ezout’s Cours de math- ematique ,as i xv o l u m ew o r kt h a t first appeared in 1764-1769, which was almost immediately issued in a new e dition of 1770-1772, and which boasted many versions in French and other languages. (The first American textbook in analytic geometry, incidentally, was derived in 1826 from B´ ezout’s Cours .) It was through such compilations, rather than through the original works of the authors themselves, that the mathematical advances of Euler and d’Alembert became widely known. B´ ezout’s name is familiar today in con- nection with the use of determinants in algebraic elimination. In a memoirof the Paris Academy for 1764, and more extensively in a treatise of 1779entitled Theorie generale des equations algebriques ,B ´ezout gave arti ficial rules, similar to Cramer’s, for solving nsimultaneous linear equations in n unknowns. He is best known for an extension of these to a system of equa- tions in one or more unknowns in which it is required to find the condition on the coe fficients necessary for the equations to have a common solution. To take a very simple case, one might ask for the condition that the equa-tions a 1x+b1y+c1=0 , a2x+b2y+c2=0 , a3x+bay+c3=0h a v ea common solution. The necessary condition is that the eliminant a special case of the “Bezoutiant,” should be 0. Somewhat more complicated eliminants arise when conditions are sought for two polynomial equations of unequal degree to have a common solution.B´ezout also was the first one to give a satisfactory proof of the theorem, known to Maclaurin and Cramer, that two algebraic curves of degrees mand nrespectively intersect in general in m·npoints; hence, this is often called B´ezout’s theorem. Euler also had contributed to the theory of elimination, but less extensively than did B´ ezout. Taken from A History of Mathematics by Carl Boyer 122 CHAPTER 3. EIGENVALUES AND EIGENVECTORS William Rowen Hamilton Born Aug. 3/4, 1805, Dublin, Ire. and died Sept. 2, 1865, Dublin Irish mathematician and astronomer who developed the theory of quater- nions, a landmark in the development of algebra, and discovered the phe-nomenon of conical refraction. His uni fication of dynamics and optics, more- over, has had lasting in fluence on mathematical physics, even though the full signi ficance of his work was not fully appreciated until after the rise of quantum mechanics. Like his English contemporaries Thomas Babington Macaulay and John Stuart Mill, Hamilton showed unusual intellect as a child. Before the age of three his parents sent him to live with his father’s brother, James, a learned clergyman and schoolmaster at an Anglican school at Trim, a small town near Dublin, where he remained until 1823, when he entered Trinity College, Dublin. Within a few months of his arrival at his uncle’s he couldread English easily and was advanced in arithmetic; at five he could translate Latin, Greek, and Hebrew and recite Homer, Milton and Dryden. Beforehis 12th birthday he had compiled a grammar of Syriac, and by the age of 14 he had su fficient mastery of the Persian language to compose a welcome to the Persian ambassador on his visit to Dublin. Hamilton became interested in mathematics after a meeting in 1820 with Zerah Colburn, an American who cou ld calculate mentally with aston- ishing speed. Having read the El´ements d’alg` ebre of Alexis—Claude Clairaut and Isaac Newton’s Principia ,Hamilton h a di m m e r s e dh i m s e l fi nt h e five volumes of Pierre—Simon Laplace’s Trait´ ed em ´ ecanique c´ eleste (1798-1827; Celestial Mechanics ) by the time he was 16. His detection of a flaw in Laplace’s reasoning brought him to the attention of John Brinkley, pro-fessor of astronomy at Trinity College. When Hamilton was 17, he sent Brinkley, then president of the Royal Irish Academy, an original memoir about geometrical optics. Brinkley, in forwarding the memoir to the Acad- emy, is said to have remarked: “This young man, I do not say will be , but is,t h e first mathematician of his age.” In 1823 Hamilton entered Trinity College, from which he obtained the highest honours in both classics and m athematics. Meanwhile, he continued his research in optics and in April 1827 submitted this “theory of Systems of Rays” to the Academy. The paper tr ansformed geometrical optics into a new mathematical science by establishing one uniform method for thesolution of all problems in that field.Hamilton started from the principle, originated by the 17th-century French mathematician Pierre de Fermat, thatlight takes the shortest possible time in going from one point to another, whether the path is straight or is bent by refraction. Hamilton ’s key idea 3.5. SIMILARITY 123 was to consider the time (or a related quantity called the “action”) as a function of the end points between which the light passes and to show thatthis quantity varied when the coordinates of the end points varied, accordingto a law that he called the law of varying action. He showed that the entiretheory of systems of rays is reducible to the study of this characteristic function. Shortly after - Hamilton submitted his paper and while still an under- graduate, Trinity College elected h im to the post of Andrews professor of astronomy and royal astronomer of Ireland, to succeed Brinkley, who hadbeen made bishop. Thus an undergraduate (not quite 22 years old) becameex officio an examiner of graduates who were candidates for the Bishop Law Prize in mathematics. The electors’ object was to provide Hamilton with a research post free from heavy teachi ng duties. Accordingly, in October 1827Hamilton took up residence next to Dunsink Observatory, 5 miles (8 km) from Dublin, where he lived for the rest of his life. He proved to be anunsuccessful observer, but large audiences were attracted by the distinctlyliterary flavour of his lectures on astronomy. Throughout his life Hamilton was attracted to literature and considered the poet William Wordsworth among his fiends, although Wordsworth advised him to write mathematics rather than poetry. With eigenvalues we are able to begin spectral analysis. That part is the derivation of the various normal forms for matrices. We begin with a relatively weak form of the Jordan form, which is coming up. First of all, as you have seen diagona l matrices furnish the easiest form for matrix analysis. Also, linear systems are very simple to solve for diagonalmatrices. The next simplest class of matrices are the triangular matrices.We begin with the following result based on the idea of similarity. 3.5 Similarity Definition 3.5.1. Am a t r i x B∈Mnis said to be similar toA∈Mnif there exists a nonsingular matrix S∈Mnsuch that B=S−1AS. The transformation A→S1ASis called a similarity transformation .S o m e - times we write A∼B. Note that similarity is an equivalence relation: (i) A∼A reflexivity (ii) B∼A⇒A∼B symmetry (iii) B∼AandA∼C⇒B∼Ctransitivity 124 CHAPTER 3. EIGENVALUES AND EIGENVECTORS Theorem 3.5.1. Similar matrices have the same characteristic polynomial. Proof. We suppose A, B,a n d S∈Mnwith Sinvertible and B=S−1AS. Then λI−B=λI−S−1AS =λS−1IS−S−1AS =S−1(λI−A)S. Hence det(λI−B)=d e t ( S−1(λI−A)S) =d e t ( S−1)d e t (λI−A)d e tS =d e t (λI−A) because 1 = det I=d e t ( SS−1)=d e t SdetS−1. A simple consequence of Theorem 3.5.1 and Corollary 2.3.1 follows. Corollary 3.5.1. IfAandBare in Mnand if AandBare similar, then they have the same eigenvalues counted ac cording to multiplicity, and there- f o r et h es a m er a n k . We remark that even though [00 00]a n d[0100] have the same eigenvalues, 0 and 0, they are not similar. Hence the converse is not true. Another immediate corollary of Theorem 3.5.1 can be expressed in terms of the invariance of the trace and determinant of similarity transformations , that is functions on Mn(F)d efined for any invertible matrix SbyTS(A)= S−1AS. Note that such transformations are linear mappings from Mn(F)→ Mn(F). We shall see how important are those properties of matrices that are invariant (i.e. do not change) under similarity transformations. Corollary 3.5.2. IfA, B∈Mnare similar, then they have the same trace and same determinant. That is, tr(A) = tr(B) anddetA=d e t B. Theorem 3.5.2. IfA∈Mn,t h e n Ais similar to a triangular matrix. Proof. The following sketch shows the first two steps of the proof. A formal induction can be applied to achieve the full result. Letλ1be an eigenvalue of Awith eigenvector u1. Select a basis , say S1,o fCnand arrange these vectors into columns of the matrix P1,w i t h u1 3.5. SIMILARITY 125 in the first column. De fineB1=P−1 1AP1.T h e n B1is the representation of Ain the new basis and so B1= λα 1...αn−1 0 ... A2 0 where A 2is (n−1)×(n−1) because Au1=λ1u1. Remembering that B1is the representation of Ain the basis S1,t h e n[ u1]S1=e1and hence B1e1=λe1=λ1[1,0,... , 0]T.B y similarity, the characteristic polynomial of B1i st h es a m ea st h a to f A, but more importantly (using expansion by minors down the first column) det(λI−B1)=(λ−λ1)d e t (λIn−1−A2). Now select an eigenvalue λ2ofA2and pertaining eigenvector v2∈Cn;s o A2v2=λ2v2.N o t et h a t λ2is also an eigenvalue of A. With this vector we define ˆu1= 1 0 ... 0 ˆu2= 0 v2 . Select a basis of Cn−1,w i t h v2selected first and create the matrix P2with this basis as columns P2= 10 ... 0 0 ...P 2 0 . It is an easy matter to see that P 2is invertible and B2=P−1 2B1P2 = λ 1∗... ... ∗ 0λ2∗...∗ 00 A3 ...... 00 . Of course, B 2∼B1, and by the transitivity of similarity B2∼A.C o n - tinue this process, deriving ultimately the triangular matrix Bn∼Bn−1and hence Bn∼A. This completes the proof. 126 CHAPTER 3. EIGENVALUES AND EIGENVECTORS Definition 3.5.2. We say that A∈Mnisdiagonalizable ifAis similar to a diagonal matrix. Suppose P∈Mnis nonsingular and D∈Mnis a diagonal matrix. Let A=PDP−1. Suppose the columns of Pare the vectors v1,v2,... ,v n.T h e n Avj=PDP−1vj =PDe j=λjPej=λjvj whereλjis the jthdiagonal element of D,a n d ejis the jthstandard vector. Similarly, if u1,... ,u narenlinearly independent vectors of Apertain- ing to eigenvalues λ1,λ2,... ,λn,t h e nw i t h Q, the matrix with columns u1...u n,w eh a v e AQ=QD where D= λ 1 λ2s ... s λn . Therefore Q−1AQ=D. We have thus established the Theorem 3.5.3. LetA∈Mn.T h e n Ais diagonalizable if and only if A hasnlinearly independent eigenvectors. As a practical measure, the conditions of this theorem are remarkably difficult to verify. Example 3.5.1. The matrix A=[01 00]i snotdiagonalizable. Solution. F i r s tw en o t et h a tt h es p e c t r u m σ(A)={0}.S o l v i n g Ax=0 we see that x=c(1,0)Tis the only solution. That is to say, there are not twolinearly independent eigenvectors. hence the result. 3.5. SIMILARITY 127 Corollary 3.5.3. IfA∈Mnis diagonalizable and Bis similar to A,t h e n Bis diagonalizable. Proof. The proof follows directly from transitivity of similarity. However, more directly, suppose that B∼AandSis the invertible matrix such that B=SAS−1Then BS=SA.I fuis an eigenvector of Awith eigenvalue λ, then SAu =λSuand therefore BSu =λSu. This is valid for each eigenvec- tor. We see that if u1,...,u nare the eigenvalues of A,t h e n Su1,...,Su n are the eigenvalues of B. The similarity matrix converts the eigenvectors of one matrix to the eigenvectors of the transformed matrix. In light of these remarks, we see that similarity transforms preserve com-pletely the dimension of eigenspaces. It is just as signi ficant to note that if a similarity transformation diagonalizes a given matrix A,t h es i m i l a r i t y matrix must consist of the eigenvectors of A. Corollary 3.5.4. LetA∈M nbe nonzero and nilpotent. Then Ais not diagonalizable. Proof. Suppose A∼T,w h e r e Tis triangular. Since Ais nilpotent (i.e.( Am= 0),Tis nilpotent, as well. Therefore the diagonal entries of Tare zero. The spectrum of Tand hence Ais zero, it’s null space has dimension n.T h e r e - fore, by Theorem 2.3.3(f), its rank is zero. Therefore T= 0, and hence A= 0. The result is proved. Alternatively, if the nilpotent matrix Ais similar to a diagonal matrix with zero diagonal entries, then Ai ss i m i l a rt ot h ez e r om a t r i x . T h u s Ais itself the zero matrix. From the obvious fact that the power of a similarity trans- formations is the similarity transformation of the power of a matrix, that is(SAS −1)m=SAmS−1(see Exercise 26), we have Corollary 3.5.5. Suppose A∈Mn(C)is diagonalizable. (i) Then Amis diagonalizable for every positive integer m. (ii) If p(·)is any polynomial, then p(A)is diagonalizable. Corollary 3.5.6. If all the eigenvalues of A∈Mn(C)are distinct, then A is diagonalizable. Proposition 3.5.1. LetA∈Mn(C)andε>0. Then for any matrix norm ,·,there is a matrix Bwith norm ,B,<εfor which A+Bis diagonalizable and for each λ∈σ(A)there is a µ∈σ(B)for which |λ−µ|<ε. 128 CHAPTER 3. EIGENVALUES AND EIGENVECTORS Proof. First triangularize AtoTby a similarity transformation, T=SAS−1. Now add to Tany diagonal matrix Dso that the resulting triangular ma- trixT+Dhas all distinct values. Moreover this can be accomplished by a diagonal matrix of arbitrarily small norm for any matrix norm. Then wehave S −1(T+D)S=S−1TS+S−1DS =A+B where B:=S−1DS. Now by the submultiplicative property of matrix norms ,B,≤EES−1DSEE≤EES−1EE,D,,S,. Thus to obtain the estimate ,B,<ε, it is sufficient to take ,D,≤6 ,S−1,,S,, which is possible as established above. Since the spectrum σ(A+B)o fA+Bhasndistinct values, it follows that A+Bis diagonalizable. There are many, many results on diagonalizable and non-diagonalizable ma- trices. Here is an interesting class o f nondiagonalizable matrices we will encounter later. Proposition 3.5.2. pro Every matrix of the form A=λI+N where Nis nilpotent and is not zero is not diagonalizable. An important subclass has the form: A= λ1 λ1s λ1 s...1 λ “Jordan block”. Eigenvectors Once an eigenvalue is determined, it is a relatively simple matter to find the pertaining eigenvectors. Just solve the homogeneous system ( λI−A)x=0 . Eigenvectors have a more complex structure and their study merits ourattention. Facts (1)σ(A)=σ(A T) including mutliplicities. 3.5. SIMILARITY 129 (2)σ(A)=σ(A∗), including multiplicities. Proof. det(λI−A)=d e t ( ( λI−A)T)=d e t (λI−AT). Similarly for A∗. Definition 3.5.3. The linear space spanned by all the eigenvectors per- taining to an eigenvalue λis called the eigenspace corresponding to the eigenvalue λ. For any A∈Mnany subspace V∈Cnfor which AV⊂V is called an invariant subspace ofA. The determination of invariant sub- spaces for linear transformations has been an important question for decades. Example 3.5.2. For any upper triangular matrix Tthe spaces Vj=S(e1,e2,... ,e j), j=1,... ,n are invariant. Example 3.5.3. Given A∈Mn, with eigenvalue λ. The eigenspace cor- responding to λand all of its subspaces are invariant subspaces. Corollary to this, the null space N(A)= {x|Ax=0}is an invariant subspace of A corresponding to the eigenvalue λ=0 . Definition 3.5.4. LetA∈Mnwith eigenvalue λ. The dimension of the eigenspace corresponding to λis called the geometric multiplicity ofλ.T h e multiplicity of λas a zero of the characteristic polynomial pA(λ)i sc a l l e d thealgebraic multiplicity . Theorem 3.5.4. IfA∈Mnandλ0is an eigenvalue of Awith geometric and algebraic multiplicities mgandma, respectively. Then mg≤ma. Proof. Letu1...u mgbe linearly independent eigenvectors pertaining to λ0. LetSbe a matrix consisting of a basis of Cnwith u1...u mgselected among them and placed in the firstmgcolumns. Then B=S−1AS= λ0I...∗............ 0...B  where I=Img. It is easy to see that pB(λ)=pA(λ)=(λ−λ0)mgp0B(λ), whence the algebraic multiplicity ma≥mg. 130 CHAPTER 3. EIGENVALUES AND EIGENVECTORS Example 3.5.4. ForA=[01 00], the algebraic multiplicity of ais 2, while the geometric multiplicity of 0 is 1. Hence, the equality mg=mais not possible, in general. Theorem 3.5.5. IfA∈Mnand for each eigenvalue µ∈σ(A),mg(µ)= ma(µ),t h e n Ais diagonalizable. The converse is also true. Proof. Extend the argument given in the theorem just above. Alternatively, we can see that Amust have nlinearly independent eigenvalues, which follows from the Lemma 3.5.1. IfA∈MnandµW=λare eigenvalues with eigenspaces Eµ andEλrespectively. Then EµandEλare linearly independent. Proof. Suppose u∈Eµc a nb ee x p r e s s e da s u=k3 cjvj where, of course uW=0a n d vjW=0,j=1,... ,k where v1...v k∈Eλ.T h e n Au=AΣcjvj ⇒ µu=λΣcjvj ⇒ u=λ µΣcjvj,ifµW=0. Sinceλ/µW= 1, we have a contradiction. If µ=0 ,t h e n Σcjvj=0a n dt h i s implies u= 0. In either case we have a contradiction, the result is therefore proved. While both AandAThave the same eigenvalues, counted even with multiplicities, the eigenspaces can be very much di fferent. Consider the example where A=}11 02] and AT=}10 12] . The eigenvalue λ= 1 has eigenvector u=[1 0]f o rAandJ1 −1o forAT. A new concept of left eigenvector yield some interesting results. Definition 3.5.5. We sayλis aleft eigenvector ofA∈Mnif there is a vector y∈Cnfor which y∗A=λy∗ 3.6. EQUIVALENT NORMS AND CONVERGENT MATRICES 131 or inRnifyTA=λyT. Taking adjoints the two sides of this equality become (y∗A)∗=A∗y (λy∗)∗=¯λy. Putting these lines together A∗y=¯λyand hence ¯λis an eigenvalue of A∗. Here’s the big result. Theorem 3.5.6. LetA∈Mn(C)with eigenvalues λand eigenvector u. Letu, v be left eigenvectors pertaining to µ,λ, respectively. If µW=λ,v∗u= u, vX=0. Proof. We have v∗u=1 µv∗Aµ=λ µv∗u. Assuming µW= 0 we have a contradiction. If µ=0 ,a r g u ea s v∗u=1 λv∗Au= µ λc∗u= 0, which was to be proved. The result is proved. Note that left eigenvectors of Aare (right) eigenvectors of A∗(ATin the real case). 3.6 Equivalent norms and convergent matrices Equivalent norms S of a rw eh a v ed e fined a number of di fferent norms. Just what “di fferent” m e a n si ss u b j e c tt od i fferent interpretations. For example, we might agree that different means that the two norms have a di fferent value for some matrix A. On the other hand if we have two matrix norms ,·,aand,·,b, we might be prepared to say that if for two positive constants mM > 0 0<m≤,A,a ,A,b≤M for all A∈Mn(F) then these norms are not really so di fferent but rather are equivalent be- cause “small” in one norm implies “small” in the other and the same for“large.” Indeed, when two norms satisfy the condition above we will callthem equivalent. The remarkable fact is that all subordinate norms on afinite dimensional space are equivalent. This is a direct consequence of the similar result for vector norms. We state it below but leave the details of the proof to the reader. 132 CHAPTER 3. EIGENVALUES AND EIGENVECTORS Theorem 3.6.1. Any two matrix norms, ,·,aand,·,b,o n Mn(C)are equivalent in the sense that ther e exist two positive constants mM > 0 0<m≤,A,a ,A,b≤M for all A∈Mn(F) We have already proved one convergence result about the invertibility of I−Awhen,A,<1. A deeper version of this result can be proved based on a new matrix norm. This result has important consequences in general matrix theory and particularly in com putational matrix theory. It is most important in applications, where having ρ(A)<1 can yield the same results as having ,A,<1. Lemma 3.6.1. Let,·,be a matrix norm that is subordinate to the vector norm (on Cn),·,. Then for each A∈Mn ρ(A)≤,A,. Note: We use the same notation for both vector and matrix norms. Proof. Letλ∈σ(A). Then with corresponding eigenvector xλwe have Axλ=λxλ.N o r m a l i z i n g xλso that,xλ,=1 ,w eh a v e ,A,=m a x ,x,=1,Ax,≥ ,Axλ,=|λ|,xλ,=|λ|.H e n c e ,A,≥ max λ∈σ(A)|λ|=ρ(A). Theorem 3.6.2. For each A∈Mnandε>0there is a vector norm ,·, onCnfor which the subordinate matrix norm ,·,satisfies ρ(A)≤,A,≤ρ(A)+ε. Proof. The proof follows in a series of simple steps, the first of which is interesting in its own right. First we know that Ais similar to a triangular matrix B–from a previous theorem B=SAS−1=Λ+U whereΛis the diagonal part of BandUis strictly upper triangular part of B,w i t hz e r o s filled in elsewhere. 3.6. EQUIVALENT NORMS AND CONVERGENT MATRICES 133 Note that the diagonal of Bisthe spectrum of Band hence that of A. This is important. Now select a δ>0 and form the diagonal matrix D= 1 δ−1s ... s δ1−n . A brief computation reveals that C=DBD−1=DSAS−1D−1 =D(Λ+U)D−1 =Λ+DUD−1=Λ+V where V=DUD−1and more speci fically vij=δj−iuijforj>i and of courseΛis a diagonal matrix with the spectrum of Afor the diagonal ele- ments. In this way we see that for δsmall enough we have arranged that Ais similar to the diagonal matrix of its spectral elements plus a small triangular perturbation. We now de fine the new vector norm on Cnby ,x,A=(DS)∗(DS)x, xX1/2. We recognise this to be a norm from a previous result. Now compute the matrix norm ,A,A=m a x ,x,A=1,Ax,A. Thus ,Ax,2 A=DSAx,DSAx X =CDSx,CDSx X ≤,C,2 2,DSx,2 =,C∗C,2,x,2 A. From C=Λ+V, it follows that C∗C=Λ∗Λ+Λ∗V+V∗Λ+V∗V. Because the last three terms have a δpmultiplier for various positive powers p,w ec a nm a k et h et e r m s Λ∗V+V∗Λ+V∗Vsmall in anynorm by taking δsufficiently small. Also the diagonal elements of Λ∗Λhave the form |λi|2λi∈S(A). 134 CHAPTER 3. EIGENVALUES AND EIGENVECTORS Talkingδsufficiently small so that each of ,Λ∗V,2,,V∗Λ,2and,V∗V,2is less than 6/3. With ,x,2 A=1a sr e q u i r e d ,w eh a v e ,Ax,A≤(ρ(A)+ε),x,A and we know already that ,A,A≥ρ(A). An important corollary places a lower bound on vector norms is given below. Corollary 3.6.1. For any A∈Mn ρ(A)= i n f ,,(m a x ,x,=1,Ax,) where the in fimum is taken over all vector norms. Convergent matrices Definition 3.6.1. We say that a matrix A∈Mn(F)f o r F=RorCis convergent if lim m→∞Am=0 ←the zero matrix. That is, for each 1 ≤i, j≤nthelimm→∞(Am)ij=0 . T h i si ss o m e t i m e s called pointwise convergence. Theorem 3.6.3. The following three statements are equivalent: (a)Ais convergent. (b) lim n→∞,Am,=0,f o rs o m em a t r i xn o r m . (c)ρ(A)<1. Proof. These results follow substantially from previously proven results. However, for completeness, assume (a) holds. Then lim n→∞max ij|(An)ij|=0 or what is the same we have lim n→∞,An,∞=0 3.6. EQUIVALENT NORMS AND CONVERGENT MATRICES 135 which is (b). Now suppose that (b) holds. If ρ(A)≥1t h e r ei sav e c t o r xfor which ,Anx,=,ρ(A),n,x,.T h e r e f o r e ,An,≥1, which contradicts (b). Thus (c) holds. Next if (c) holds we can apply the above theorem toestablish that there is a norm for which ,A,<1. Hence (b) follows. Finally, we know that ,A m,∞≤M,Am, (T) whence lim n→∞Am=0 . ( T h a ti s( a ) ≡(b)). In sum we have shown that (a) ≡(b) and (b) ≡(c). To establish ( T)w en e e dt h ef o l l o w i n g . Theorem 3.6.4. If,, and,·,Iare two vector norms on Cnthen there are constants mandMso that m,x,I≤,x,<M,x,I for all x∈Cn. This result carries over to matrix norms subordinate to vector norms by simple inheritance. To prove this result we use compactness ideas. Supposethere is a sequence of vectors x jfor which 1≤,xj,<1 j,xj,I and for which the components of xjare bounded in modulus. By compact- ness there is a convergent subsequence of the xj,f o rw h i c hl i m xj=xW=0 . But,x,I<∞. Hence,x,= 0, a contradiction. Here is the new and improved version of our previous result. Theorem 3.6.5. Ifρ(A)<1then (I−A)−1exists and (I−A)−1=I+A+A2+···. Proof. Select a matrix norm ,·,for which ,A,<1. Apply previous calcu- lations. We know that every matrix can be triangularized and “almost” di- agonalized we may ask if we can eliminate altogether the o ff-diagonal terms. The answer is unfortunately no. But we can resolve the diagonal question completely. 136 CHAPTER 3. EIGENVALUES AND EIGENVECTORS 3.7 Exercises 1. Suppose that U∈Mnis orthogonal and let ,·,be a matrix norm. Show that ,U,≥1. 2. In two dimensions the counter-clockwise rotations through the angle θare given by Bθ=}cosθ−sinθ sinθcosθ] Find the eigenvalues and eigenvectors for all θ. (Note the two special cases,θis not equal to an even multiple of πandθ=0 . 3. Given two matrices A, B∈Mn(C). De fine the commutant ofAand Bby [A, B]−AB−BA. Prove that tr[ A, B]=0 . 4. Given two finite sequences {ck}n k=1and{dk}n k=1.Prove that k|ckdk|≤ max k|dk| k|ck|. 5. Verify that a matrix norm which is subordinate to a vector norm sat- isfies norm conditions (i) and (ii). 6. Let A∈Mn(C).Show that the matrix norm subordinate to the vector norm ,·,∞is given by ,A,∞=m a x i,ri(A),1 where as usual ri(A)d e n o t e st h e ithrow of the matrix Aand,·,1is thef1norm. 7. The Hilbert matrix, Hnof order nis defined by hij=1 i+j−11≤i, j≤n 8. Show that ,Hn,1<lnn. 9. Show that ,Hn,∞=1 . 10. Show that ,Hn,2∼n1 2. 11. Show that the spectral radius of Hnis bounded by 1. 12. Show that for each ε>0 there exists an integer Nsuch that if n>N there is a vector x∈Rnwith,x,2=1 such that ,Hnx,2<ε. 3.7. EXERCISES 137 13. Same as the previous question except that you need to show that N=! 1 ε1/2 +1w i l lw o r k . 14. Show that the matrix A= 110 031 1−12 is not diagonalizable.Let A∈Mn(C). 15. We know that the characteristic polynomial of a matrix A∈M12(C) is equal to pA(λ)=(λ−1)12−1.Show that Ais not similar to the identity matrix. 16. We know that the spectrum of A∈M3(C)i sσ(A)={1,1,−2}and the corresponding eigenvectors are {[1,2,1]T,[2,1,−1]T,[1,1,2]T}. (i) Is it possible to determine A? Why or why not? If so, prove it. If not show two di fferent matrices with the given spectrum and eigenvectors. (ii) Is Adiagonalizable? 17. Prove that if A∈Mn(C) is diagonalizable then for each λ∈σ(A), the algebraic and geometric multiplicities are equal. That is, ma(λ)= mg(λ). 18. Show that B−1(λ)e x i s t sf o r |λ|sufficiently large. 19. We say that a matrix A∈Mn(R) is row stochastic if all its entries are non negative and the sum of the entries of each row is one. (i) Prove that 1 ∈σ(A). (ii) Prove that ρ(A)=1 . 20. We say that A, B∈Mn(C)a r e quasi -commutative if the spectrum of AB−BAis just zero, i.e. σ(AB−BA)={0}. 21. Prove that if ABis nilpotent then so is BA. 22. A matrix A∈Mn(C) has a square root if there is a matrix B∈Mn(C) such that B2=A.Show that if Ais diagonalizable then it has a square root. 23. Prove that every matrix that commutes with every diagonal matrix is itself diagonal. 138 CHAPTER 3. EIGENVALUES AND EIGENVECTORS 24. Suppose that if A∈Mn(C) is diagonalizable and for each λ∈σ(A), |λ|<1.Prove directly from similarity ideas that lim n→∞An= 0.(That is, do not apply the more general theorem from the lecture notes.) 25. We say that Aisright quasi- invertible if there exists a matrix B∈ Mn(C) such that AB∼D,w h e r e Dis a diagonal matrix with diagonal entries nonzero. Similarly, we say that Aisleft quasi- invertible if there exists a matrix B∈Mn(C) such that BA∼D,w h e r e Dis a diagonal matrix with diagonal entries nonzero. (i) Show that if Aisright quasi- invertible then it is invertible. (ii) Show that if Aisright quasi- invertible then it is left quasi- invertible. (iii) Prove that quasi- invertibility is not an equivalence relation. (Hint. How do we usually show that an assertion is false?) 26. Suppose that A, S∈Mn,w i t h Sinvertible, and mis a positive integer. Show that ( SAS−1)m=SAmS−1. 27. Consider the rotation  10 0 0c o sθ−sinθ 0s i nθcosθ  Show that the eigenvalues are the same as for the rotation  cosθ0−sinθ 01 0 sinθ0c o sθ  See Example 2 of section 3.2. 28. Let A∈Mn(C). De fineeA=∞ n=0An n!. (i) Prove that exists. (ii) Suppose that A, B∈Mn(C). Is eA+B=eAeB?I f n o t , w h e n i s i t true? 29. Suppose that A, B∈Mn(C). We say A`Bif [A, B]=0 ,w h e r e [·,·] is the commutant. Prove or disprove that “ `” is an equivalence relation. Answer the same question in the case of quasi-commutivity.(See Exercise 20 30. Let u, v∈C n.Find ( I+uv∗)m. 3.8. APPENDIX A 139 31. De fine the function fonMm×n(C)b y f(A)=r a n k A, and suppose that,·,is any matrix norm. Prove the following. (1) If for any matrix Afor which f(A)=nshow that fis continuous in ,·,.T h a t is, for every ε>0t h e n f(B)=nfor every matrix Bwith,BA,<ε. is continuous in an ε-neighborhood of A.T h a t i s , f o r a n y Bwith ,A−B,<ε,t h e n f(B)=f(A). (This deviates slightly from the usual de finition of continuity because the function fis integer valued.) (2) If f(A)<n,t h e n fis not continuous in every ε-neighborhood of A. 3.8 Appendix A It is desired to solve the equation p(λ)=0f o rc o e fficients in p0,...,p n∈C orR. The basic theorem on this subject is called the Fundamental Theo- rem of Algebra (FTA), whose importance is manifest by its hundreds ofapplications. Concommitant with the FTA is the notion of reducible andirreducible factors. Theorem 3.8.1 (Theorem Fundamental Theorem of Algebra). Given any polynomial p(λ)=p 0λn+p1λn−1+···+pnwith coefficients in p0,...,p n∈ C. There is at least one solution λ∈Cto the equation p(λ)=0 . Though proving this result would take us too far a field of our subject, we remarks that the simplest proof of this result no doubt comes as a di-rect application of Liouville’s Theorem, a result in complex analysis. As a corollary to the FTA, we have that there are exactly nsolutions to p(λ)=0 when counted with multiplicitiy. Proved originally by Gauss, the proof ofthis theorem eluded mathematicians for many years. Let us assume thatp 0= 1 to make the factorization simpler to write. Thus p(λ)=k i=1(λ−λi)mi(4) . As we know, in the case that the coe fficients p0,...,p nare real, there may be complex solutions. As is easy to prove, complex solutions must come in complex conjugate pairs. For if λ=r+isis a solution p(r+is)=(r+is)n+p1(r+is)n−1+···pn−1(r+is)+pn=0 Because the coe fficients are real, the real and imaginary parts of the powers (r+is)jremain respectively real or imaginary upon multiplication by pj. 140 CHAPTER 3. EIGENVALUES AND EIGENVECTORS Thus p(r+is)=R e ( r+is)n+p1Re (r+is)n−1+···pn−1Re (r+is)+pn +ip Im (r+is)n+p1Im (r+is)n−1+···pn−1Im (r+is)+pnQ =0 Hence the real and imaginary parts are each zero. Since Re ( r−is)j= Re (r+is)jand Im ( r−is)j=−Im (r+is)jit follows that p(r−is)=0 . In the case that the coe fficients are real it may be of interest to note what statement of factorization analogous to (4) above. To this end we need the de finition Definition 3.8.1. The real polynomial x2+bx+cis called irreducible if it has no real zeros. In general, any polynomial that cannot be factored over the underlying field is called irreducible. Irreducible polynomials have played a very impor- tant role in fields such as abstract algebra and number theory. Indeed, they were central in early attempts to prove Fermat’s last theorem and also so but more indirectly to the acutal proof. Our attention here is restricted to thefieldsCandR. Theorem 3.8.2. Letp(λ)=λn+p1λn−1+···+pnwith coefficients in p0,...,p n∈R.T h e n p(λ)can be factored as a product of linear factors pertaining to real zeros of p(λ)=0 and irreducible quadratic factors. Proof. The proof is an easy application of the FTA and the observation above. If λk=r+isis a zero of p(λ), then so also is ¯λk=r−is.T h e r e f o r e the product ( λ−λk)D λ−¯λki =λ2−2λr+r2+s2is an irreducible quadratic. Combine such terms with the linear factors generated from the real zeros,and use (4). Remark 3.8.1. It is worth noting that there are nohigher order irreducible factors with real coe fficients. Even though there are certainly higher order polynomials with complex roots, they can always factored as products of either real linear or real quadratic f actors. Proving this without the FTA may prove challenging. 3.9. APPENDIX B 141 3.9 Appendix B 3.9.1 In finite Series Definition 3.9.1. Aninfiniteseries, denoted by a0+a1+a2+··· is a sequence {un},w h e r e unis defined by un=a0+a1+···+an If the sequence {un}converges to some limit A,w es a yt h a tt h ei n finite series converges to Aand use the notation a0+a1+a2+···=A We also say that the sum of the in finite series is A.I f {un}diverges, the infinite series a0+a1+a2+···is said to be divergent. The sequence {un}is called the sequence ofpartial sums, and the se- quence {an}is called the sequence ofterms of the in finite series a0+a1+ a2+···. Let us now return to the formula 1+r+r2+···+rn=1−rn+1 1−r where rW= 1. Since the sequence {rn+1}converges (and, in fact, to 0) if and only if−1<r< 1( o r |r|<1),rbeing different from 1, the in finite series 1+r+r2+··· converges to (1 −r)−1if and only if |r|<1. This in finite series is called the geometric series with ratio r. Definition 3.9.2. Geometric Series with Ratio r 1+r+r2+···=1 1−r,if|r|<1 and diverges if |r|≥1. Multiplying both sides by ryields the following. 142 CHAPTER 3. EIGENVALUES AND EIGENVECTORS Definition 3.9.3. If|r|<1, then r+r2+r3+···=r 1−r Example 3.9.1. Investigate the convergence of each of the following in fi- nite series. If it converges, determine its limit (sum).a. 1− 1 2+1 4−1 8+··· b. 1 +2 3+4 9+8 27+··· c.3 4+9 16+27 64+··· Solution A careful study reveals that each in finite series is a geometric series. In fact, the ratios are, respectively, (a) −1 2,( b )2 3,a n d( c )3 4.S i n c e they are all of absolute value less than 1, they all converge. In fact, we have:a. 1− 1 2+1 4−1 8+···=1 1−(−1 2)=2 3 b. 1 +2 3+4 9+8 27+···=1 1−2 3=3 1=3 c.3 4+9 16+27 64+···=3 4 1−3 4=3 4·4 1=3 Since an in finite series is de fined as a sequence of partial sums, the prop- erties of sequences carry over to in finite series. The following two properties are especially important. 1. Uniqueness of Limit Ifa0+a1+a2+···converges, then it converges to a unique limit. This property explains the notation a0+a1+a2+···=A where Ais the limit of the in finite series. Because of this notation, we often say that the in finite series sums to A,o rt h e sum oftheinfinitely many terms a0,a1,a2,... isA. We remark, however, that the order of the terms cannot be changed arbitrarily in general. Definition 3.9.4. 2. Sum of In finite Series If a0+a1+a2+···=Aandb0+b1+b2+···=B,t h e n (a0+b0)+(a1+b1)+(a2+b2)+··· =(a0+a1+a2+···) +(b0+b1+b2+···) =A+B This property follows from the observation that (a0+b0)+(a1+b1)+···+(an+bn)=( a0+a1+···+an) +(b0+b1+···+bn) converges to A+B. Another property is 3.9. APPENDIX B 143 Definition 3.9.5. 3. Constant Multiple of In finite Series Ifa0+a1+ a2+···=Aandcis a constant, then ca0+ca1+ca2+···=cA Example 3.9.2. Determine the sums of the following convergent in finite series.a. 5 3·2+13 9·4+35 27·8+··· b.1 3·2+1 9·2+1 27·2+· Solution a. By the Sum Property, the sum of the first in finite series is w1 3+1 2W +w1 9+1 4W +w1 27+1 8W +··· =w1 3+1 9+1 27+···W +w1 2+1 4+1 8+···W =1 3 1−1 3+1 2 1−1 2=1 2+1=3 2 b. By the Constant-Multiple Property, the sum of the second in finite series is 1 2w1 3+1 9+1 27+···W =1 2·1 3 1−1 3=1 2·1 2=1 4 Another type of in finite series can be illustrated by using Taylor poly- nomial extrapolation as follows. Example 3.9.3. In extrapolating the value of ln 2, the nth-degree Taylor polynomial Pn(x)o fl n ( 1+ x)a tx= 0 was evaluated at x=1i nE x a m p l e 4 of Section 12.3. Interpret ln 2 as the sum of an in finite series whose nth partial sum is Pn(1). Solution We have seen in Example 3 of Section 12.3 that Pn(x)=x−1 2x2+1 3x3−···+(−1)n−11 nxn so that Pn(1) = 1−1 2+1 3−···+(−1)n−11 n,a n di ti st h e nth partial sum of the in finite series 1−1 2+1 3−1 4+··· Since {Pn(1)}converges to ln 2 (see Example 4in Section 12.3), we have ln 2 = 1−1 2+1 3−1 4−··· 144 CHAPTER 3. EIGENVALUES AND EIGENVECTORS It should be noted that the terms of the in finite series in the preceding example alternate in sign. In general, an in finite series of this type always converges. Alternating Series Test Leta0≥a1≥···≥0. If anap- proaches 0, then the in finite series a0−a1+a2−a3+··· converges to some limit A,w h e r e0 <A<a 0. The condition that the term antends to zero is essential in the above Alternating Series Test. In fact, if the terms do not tend to zero, the corre-sponding in finite series, alternating or not, must diverge. Divergence Test Ifa ndoes not approach zero as n→+∞, then the in finite series a0+a1+a2+··· diverges. Example 3.9.4. Determine the convergence or divergence of the following infinite series. a. 1−1 3+1 5−1 7+··· b.−2 3+4 5−6 7+··· Solution a. This series is an alternating series with 1>1 3>1 5>···>0 and with the terms approaching 0. In fact, the general ( nth) term is (−1)n1 2n+1 Hence, by the Alternating Series Test, the in finite series 1 −1 3+1 5−1 7+··· converges. b. Let a1=−2 3,a2=4 5,a3=−6 7,... . It is clear that andoes not approach 0, and it follows from the Divergence Test that the in finite series is divergent. [The nth term is ( −1)n2n/(2n+ 1).] Note, however, that the terms do alternate in sign. 3.9. APPENDIX B 145 Another useful tool for testing convergence or divergence of an in finite series is the Ratio Test, to be discussed below. Recall that the geometric series 1+r+r2+··· where the nth term is an=rn,c o n v e r g e si f |r|<1 and diverges otherwise. Note also that the ratio of two consecutive terms is an+1 an=rn+1 rn=r Hence, the geometric series converges if and only if this ratio is of absolute value less than 1. In general, it is possible to draw a similar conclusion if thesequence of the absolute values of the ratios of consecutive terms converges. Ratio Test Suppose that the sequence {|a n+1|/|an|} converges to some limit R. Then, the in finite series a0+a1+ a2+···converges if R< 1 and diverges if R> 1. Example 3.9.5. In each of the following, determine all values of rfor which the in finite series converges. a. 1 + r+r2 2!+r3 3!+··· b. 1 + 2 r+3r3+4r3+··· Solution a. The nth term is an=rn/n!, so that an+1 an=rn+1 (n+1 ) !·n! rn=r n+1 For each value of r,w eh a v e |an+1| |an|=1 n+1|r|→0<1 Hence, by the Ratio Test, the in finite series converges for all values of r. b. Let an=(n+1 )rn. Then, |an+1| |an|=(n+2 )|r|n+1 (n+1 )|r|n=n+2 n+1|r|→|r| Hence, the Ratio Test says that the in finite series converges if |r|<1a n d diverges if |r|>1. For |r|=1 ,w eh a v e |an|=n+1, which does not converge to 0, so that the corresponding in finite series must be divergent as a result of applying the Divergence Test. Sometimes it is possible to compare the partial sums of an in finite series with certain integrals, as illustrated in the following example. 146 CHAPTER 3. EIGENVALUES AND EIGENVECTORS Example 3.9.6. Show that 1+1 2+···+1/nis larger than the de fine integral$n+1 1dx/x , and conclude that the in finite series 1 +1 2+1 3+···diverges. Solution Consider the function f(x)=1 /x. Then, the de finite integral 8n+1 1dx x=l n ( n+1 )−ln 1 = ln( n+1 ) is the area under the graph of the function y=f(x) between x=1a n d x=n+ 1. On the other hand, the sum 1+1 2+···+1 n=[f(1) + f(2) + ···+f(n)]∆x where∆x= 1, is the sum of the areas of nrectangles with base ∆x=1a n d heights f(1),... ,f (n), consecutively, as shown in Figure 12.5. Since the union of these rectangles covers the region bounded by the curve y=f(x), thex-axis, and the vertical lines x=1a n d x=n+1 ,w eh a v e : 1+1 2+···+1 n>8n+1 1dx x=l n ( n+1 ) Now recall that ln( n+ 1) approaches ∞asnapproaches ∞. Hence, the sequence of partial sums of the in finite series 1+1 2+1 3+··· must be divergent. The above in finite series is called the harmonic series. It diverges “to infinity” in the sense that its sequence of partial sums becomes arbitrarily large for all large values of n.I f a n i n finite series diverges to in finity, we also say that it sums toinfinity and use the notation “= ∞” accordingly. We have TheHarmonic Series is defined by 1+1 2+1 3+···=∞ Chapter 4 Unitary Matrices 4.1 Basics This chapter considers a very important class of matrices that are quite use- ful in proving a number of structure theorems about all matrices. Calledunitary matrices, they comprise a class of matrices that have the remarkable properties that as transformations they preserve length, and preserve the an- gle between vectors. This is of course true for the identity transformation.Therefore it is helpful to regard unitary matrices as “generalized identities,”though we will see that they form quite a large class. An important exam-ple of these matrices, the rotations, have already been considered. In this chapter, the underlying field is usually C, the underlying vector space is C n, and almost without exception the underlying norm is k·k 2.W e b e g i n b y recalling a few important facts. Recall that a set of vectors x1,...,x k∈Cnis called orthogonal if x∗ jxm=hxm,xji=0f o r1 ≤j6=m≤k.T h es e ti s orthonormal if x∗ jxm=δmj=½1j=m 0j6=m. An orthogonal set of vectors can be made orthonormal by scaling: xj−→1 (x∗ jxj)1/2xj. Theorem 4.1.1. Every set of orthonormal vecto rs is linearly independent. Proof. The proof is routine, using a common technique. Suppose S= {xj}k j=1is orthonormal and linearly dependent. Then, without loss of gen- 157 158 CHAPTER 4. UNITARY MATRICES erality (by relabeling if needed), we can assume uk=k−1X j=1cjuj. Compute u∗ kukas 1=u∗ kuk=u∗ k k−1X j=1cjuj  =k−1X j=1cju∗ kuj=0. This contradiction proves the result. Corollary 4.1.1. IfS={u1...u k}⊂Cnis orthonormal then k≤n. Corollary 4.1.2. Every k-dimensional subspace of Cnhas an orthonormal basis. Proof. Apply the Gram—Schmidt process to any basis to orthonormalize it. Definition 4.1.1. Am a t r i x U∈Mnis said to be unitary ifU∗U=I.[ I f U∈Mn(R)a n d UTU=I,t h e n Uis called real orthogonal .] Note: A linear transformation T:Cn→Cnis called an isometry if kTxk=kxkfor all x∈Cn. Proposition 4.1.1. Suppose that U∈Mnis unitary. (i) Then the columns ofUform an orthonormal basis of Cn,o rRn,i fUis real. (ii) The spectrum σ(u)⊂{z||z|=1}. (iii) |detU|=1. Proof. The proof of (i) is a consequence of the de finition. To prove (ii), first denote the columns of Ubyui,i=1,...,n .I fλis an eigenvalue of Uwith pertaining eigenvector x,t h e n kUxk=kPxiuik=(P|xi|2)1/2=kxk= |λ|kxk.H e n c e |λ|= 1. Finally, (iii) follows directly because det U=Qλi. Thus |detU|=Q|λi|=1 . This important result is just one of many equivalent results about unitary matrices. In the result below, a number of equivalences are established. Theorem 4.1.2. LetU∈Mn. The following are equivalent. 4.1. BASICS 159 (a)Uis unitary. (b)Uis nonsingular and U∗=U−1. (c)UU∗=I. (d)U∗is unitary. (e) The columns of Uform an orthonormal set. (f) The rows of Uform an orthonormal set. (g)Uis an isometry. (h)Ucarries every set of orthonormal vectors to a set of orthonormal vectors. Proof. (a)⇒(b) follows from the de finition of unitary and the fact that the inverse is unique. (b)⇒(c) follows from the fact that a left inverse is also a right inverse. (a)⇒(d)UU∗=(U∗)∗U∗=I. (d)⇒(e) (e) ≡(b)⇒ u∗ jukδjk,w h e r e u1...u nare the columns of U. Similarly (b) ⇒(e). (d)≡(f) same reasoning. (e)⇒(g) We know the columns of Uare orthonormal. Denoting the columns byu1...u n,w eh a v e Ux=nX 1xiui where x=(x1,... ,x n)T. It is an easy matter to see that kUxk2=nX 1|xi|2=kxk2. (g)⇒(e). Consider x=ej.T h e n Ux=uj.H e n c e 1 = kejk=kUejk= kujk.T h e c o l u m n s o f Uhave norm one. Now let x=αei+βejbe chosen 160 CHAPTER 4. UNITARY MATRICES such that kxk=kαei+βejk=q |α|2+|β|2=1 . T h e n 1= kUxk2 =kU(αei+βej)k2 =hU(αei+βej),U(αei+βej)i =|α|2hUei,Ue ii+|β|2hUej,Ue ji+α¯βhUei,Ue ji+¯αβhUej,Ue ii =|α|2hui,uii+|β|2huj,uji+α¯βhui,uji+¯αβhuj,uii =|α|2+|β|2+2<¡ α¯βhui,uji¢ =1 + 2 <¡ α¯βhui,uji¢ Thus <¡ α¯βhui,uji¢ =0.Now suppose hui,uji=s+it. Selecting α= β=1√ 2we obtain that <hui,uji= 0, and selecting α=iβ=1√ 2we obtain that =hui,uji=0 . T h u s hui,uji= 0. Since the coordinates iandjare arbitrary, it follows that the columns of Uare orthogonal. (g)⇒(h) Suppose {v1,...,v n}is orthogonal. For any two of them kU(vj+ vk)k2=kvj+vkk2. Hence hUvj,U v ki=0 . (h)⇒(e) The orthormal set of the standard unit vectors ej,j=1,...,n is carried to the columns of U.T h a t i s Uej=uj,t h e jthcolumn of U. Therefore the columns of Uare orthonormal. Corollary 4.1.3. IfU∈Mn(C)is unitary, then the transformation de fined byUpreserves angles. Proof. We have for any vectors x, y∈Cnthat the angle θis completely determined from the inner product via cos θ=hx,yi kxkkyk.S i n c e Uis unitary (and thus an isometry) it follows that hUx,Uy i=hU∗Ux,y i=hx, yi This proves the result. Example 4.1.1. LetT(θ)=£cosθ−sinθ sinθcosθ¤ whereθis any real. Then T(θ)i s realorthogonal. Proposition 4.1.1. IfU∈M2(R)is real orthogonal, then Uhas the form T(θ)for some θor the form U=·10 0−1¸ T(θ)=·cosθsinθ sinθ−cosθ¸ Finally, we can easily establish the di agonalizability of unitary matrices. 4.1. BASICS 161 Theorem 4.1.3. IfU∈Mnis unitary, then it is diagonalizable. Proof. To prove this we need to revisit the proof of Theorem 3.5.2. As before, select the first vector to be a normalized eigenvector u1pertaining toλ1.Now choose the remaining vectors to be orthonormal to u1.T h i s makes the matrix P1with all these vectors as columns a unitary matrix. Therefore B1=P−1UPis also unitary. However it has the form B1= λα 1...αn−1 0 ... A2 0 where A 2is (n−1)×(n−1) Since it is unitary, it must have orthogonal columns by Theorem 4.1.2. It follows then that α1=α2=···=αn=0 a n d B1= λ0... 0 0 ... A2 0  At this point one may apply and inductive hypothesis to conclude that A2 is similar to a diagonal matrix. Thus by the manner in which the full similarity was constructed, we see that Amust also be similar to a diagonal matrix. Corollary 4.1.1. LetU∈Mnbe unitary. Then (i) Then Uhas a set of northogonal eigenvectors. (ii) Let {λ1,...,λn}and{v1,...,v n}denote respectively the eigenvalues and their pertaining orthonormal eigenvectors of U.Then Uhas the representation as the sum of rank one matrices given by U=nX j=1λjvjvT j This representation is often called the spectral respresentation or spectral decomposition of U. 162 CHAPTER 4. UNITARY MATRICES 4.1.1 Groups of matrices Invertible and unitary matrices have a fundamental structure that makes possible a great many general statements about their nature and the waythey act upon vectors other vectors matrices. A group is a set with a math-ematical operation, product, that obeys some minimal set of properties so as to resemble the nonzero numbers under multiplication. Definition 4.1.2. Agroup Gi sas e tw i t hab i n a r yo p e r a t i o n G×G→G which assigns to every pair a, bof elements of Ga unique element abinG. The operation, called the product ,s a t i s fies four properties: 1. Closure. If a, b∈G,t h e n ab∈G. 2. Associativity. If a, b, c∈G,t h e n a(bc)=(ab)c. 3. Identity. There exists an element e∈Gsuch that ae=ea=afor every a∈G.eis called the identity ofG. 4. Inverse. For each a∈G, there exists an element ˆ a∈Gsuch that aˆa=ˆaa=e.ˆais called the inverse of aand is often denoted by a −1. A subset of Gthat is itself a group under the same product is called a subgroup ofG. It may be interesting to note that removal of any of the properties 2-4 leads to other categories of sets that have interest, and in fact applications, in their own right. Moreover, many groups have additional properties such ascommutativity, i.e. ab=bafor all a, b∈G.B e l o w a r e a f e w e x a m p l e s o f matrix groups. Note matrix addition is not involved in these de finitions. Example 4.1.2. As usual M nis the vector space of n×nmatrices. The product in these examples is the usual matrix product. •The group GL(n, F) is the group of invertible n×nmatrices. This is the so-called general linear group. The subset of Mnof invertible lower (resp. upper) triangular matrices is a subgroup of GL(n, F). •Theunitary group Unof unitary matrices in Mn(C). •Theorthogonal group Onorthogonal matrices in Mn(R). The sub- group of Ondenoted by SOnconsists of orthogonal matrices with determinant 1. 4.1. BASICS 163 Because element inverses are required, it is obvious that the only subsets of invertible matrices in Mnwill be groups. Clearly, GL(n, F) is a group because the properties follow from those matrix of multiplication. We havealready established that invertible lo wer triangular matrices have lower tri- angular inverses. Therefore, they form a subgroup of GL(n, F). We consider the unitary and orthogonal groups below. Proposition 4.1.2. For any integer n=1,2,... the set of unitary matrices U n(resp. real orthogonal) forms a group. Similarly Onis a group, with subgroup SOn. Proof. The result follows if we can show that unitary matrices are closed under multiplication. Let UandVbe unitary. Then (UV)∗(UV)=V∗U∗UV =V∗V=I For orthogonal matrices the proof is essentially identical. That SOnis a group follows from the determinant equality det( AB)=d e t AdetB.T h e r e - fore it is a subgroup of On. 4.1.2 Permutation matrices Another example of matrix groups comes from the idea of permutations of integers. Definition 4.1.3. The matrix P∈Mn(C)i sc a l l e da permutation matrix if each row and each column has exactly one 1, the rest of the entries being zero. Example 4.1.3. Let P= 100 001010  Q= 0001 01001000 0010  PandQare permutation matrices. Another way to view a permutation matrix is with the game of chess. On an n×nchess board place nrooks in positions where none of them attack one another. Viewing the board as an n×nmatrix with ones where the rooks are and zeros elsewhere, this matrix will be a permutation matrix. 164 CHAPTER 4. UNITARY MATRICES Of course there are n! such placements, exactly the number of permutations of the integers {1,2,...,n }. Permutation matrices are closely linked with permutations as discussed in Chapter 2.5. Let σbe a permutation of the integers {1,2,...,n }.Define the matrix Aby aij=½1i f j=σ(i) 0i fo t h e r w i s e Then Ais a permutation matrix. We could also use the Dirac notation to express the same matrix, that is to say aij=δiσ(j).The product of permutation matrices is again a permutation matrix. This is apparent bystraight multiplication. Let PandQbe two n×npermutation matrices with pertaining permutations σ PandσQof the integers {1,2,...,n }.Then theithrow of PiseσP(i)and the ithrow of QiseσQ(i). (Recall the eiare the usual standard vectors.) Now the ithrow of the product PQcan be computed by nX j=1pijeσQ(j)=nX j=1δiσ(j)eσQ(j)=eσQ(σP(i)) Thus the multiplication ithrow of PQis a standard vector. Since the σP(i) ranges over the integers {1,2,...,n },i ti st r u ea l s ot h a t σQ(σP(i)) does likewise. Therefore the “product” σQ(σP(i)) is also a permutation. We conclude that the standard vectors constitute the rows of PQ.T h u s p e r m u - tation matrices are orthogonal under multiplication. Moreover the inverse ofevery permutation is permutation matrix, with the inverse describe throughthe inverse of the pertaining permutation of {1,2,...,n }. Therefore, we have the following result. Proposition 4.1.3. Permutation matrices are orthogonal. Permutation matrices form a subgroup of O n. 4.1.3 Unitary equivalence Definition 4.1.4. Am a t r i x B∈Mnis said to be unitarily equivalent toA if there is a unitary matrix U∈Mnsuch that B=U∗AU. ( I nt h er e a lc a s ew es a y Bis orthogonally equivalent to A.) 4.1. BASICS 165 Theorem 4.1.4. IfBandAare unitarily equivalent. Then nX i,j=1|bij|2=nX i,j=1|aij|2. Proof. We havenP i,j=1|aij|2=t rA∗A,a n d Σ|bij|2=t rB∗B=t r( U∗AU)∗U∗AU=t rU∗A∗AU =t rA∗A, since the trace is invariant under similarity transformations. Alternatively, we have BU∗=U∗A.S i n c e U(and U∗) are isometries we have each column of U∗Ahas the same norm as the norm of the same column of A. The same holds for the rows of BU∗, whence the result. Example 4.1.4. B=£31 −20¤ andA=[11 02] are similar but not unitarily equivalent. AandBare similar because (1) they have the same spectrum, σ(A)=σ(B)={1,2}and (2) they have two linearly independent eigenvec- tors. They are not unitarily equivalent because the conditions of the above theorem are not met. Remark 4.1.1. Unitary equivalence is a finer classi fication than similarity. Indeed, consider the two sets S(A)={B|Bis similar to A}andU(A)= {B|Bis unitarily equivalent to A},t h e n U(A)⊂S(A). (Can you show that U(A)$S(A) for some large class of A∈Mn?) 4.1.4 Householder transformations An important class of unitary transformations are elementary re flections. These can be realized as transformations that re flect one vector to its neg- a t i v ea n dl e a v ei n v a r i a n tt h eo r t h o c o m p l e m e n to fv e c t o r s . Definition 4.1.5. Simple Householder transformation.) Suppose w∈Cn, kwk=1 . D e fine the Householder transformation Hwby Hw=I−2ww∗ 166 CHAPTER 4. UNITARY MATRICES Computing HwH∗ w=(I−2ww∗)(I−2ww∗)∗ =I−2ww∗−2ww∗+4ww∗ww∗ =I−4ww∗+4hw,wiww∗=I it follows that Hwis unitary. Example 4.1.5. Consider the vector w=·cosθ sinθ¸ and the Householder transformation Hw=I−2wwT=·1−2c o s2θ−2c o sθsinθ −2c o sθsinθ1−2s i n2θ¸ =·−cos 2θ−sin 2θ −sin 2θcos 2θ¸ The transformation properties for the standard vectors are Hwe1=·−cos 2θ −sin 2θ¸ and Hwe2=·cos 2θ −sin 2θ¸ This is shown below. It is evident that this unitary transformation is not a rotation. Though, it can be imagined as a “rotation with a one dimensionalreflection.” ww wH e H eee2 1 θ2θ2θ Householder transformation Example 4.1.6. Letθbe real. For any n≥2a n d1≤i, j≤nwith i6=j 4.1. BASICS 167 define Un(θ,i ,j)i j 10 ......0 ......... 01...... ........... c o s θ............ −sinθ........... ...1 ... 1... ........... s i n θ............ c o s θ.................1... ......0...0 0.........1  ij Then U n(θ;i, j) is a rotation and is unitary. Proposition 4.1.4 (Limit theorems). (i) Show that the unitary matri- ces are closed with respect to any norm. That is, if the sequence {Un}⊂ Mn(C)are all unitary and the limn→∞Un=U in the k·k2norm, then U is also unitary. (ii) The unitary matrices are closed under pointwise convergence. That is, if the sequence {Un}⊂Mn(C)are all unitary and limn→∞Un=Ufor each ( ij)entry, then Uis also unitary. Householder transformations can als ob eu s e dt ot r i a n g u l a r i z eam a t r i x . The procedure successively removes the lower triangular portion of a matrixcolumn by column in a way similar to Gaussian elimination. The elemen- tary row operations are replaced by elementary re flectors. The result is a triangular matrix T=H vn···Hv2Hv1A.Since these re flectors are unitary, the factorization yields an e ffective method for solving linear systems. The actual process is rather straightforward. Construct the vector vsuch that ¡ I−2vvT¢ A=¡ I−2vvT¢ a 11a12···a1n a21a22···a2n ............ an1an2···ann = ˆa 11ˆa12··· ˆa1n 0ˆa22··· ˆa2n ............ 0ˆan2··· ˆann  168 CHAPTER 4. UNITARY MATRICES This is accomplished as follows. This means we wish to find a vector v such that ¡ I−2vvT¢ [a11,a21,···,an1]T=[ ˆa11,0,···,0]T For notational convenience and to emphasize the construction is vector based, relabel the column vector [ a11,a21,···,an1]Tas [x1,x2,...,x n]T vj=xj 2hv,xi forj=2,3,...,n .D e fine v1=x1+α 2hv,xi Then hv,xi=hx, xi 2hv,xi+αx1 2hv,xi=kxk2 2hv,xi+αx1 2hv,xi 4hv,xi2=2 kxk2+2αx1 where k·kdenotes the Euclidean norm. Also, for Hvto be unitary we need 1= hv,vi=1 4hv,xi2³ kxk2+2αx1+α2´ 4hv,xi2=kxk2+2αx1+α2 Equating the two expressions for 4 hv,xi2gives 2kxk2+2αx1=kxk2+2αx1+α2 α2=kxk2 α=±kxk We now have that 4hv,xi2=2 kxk2±2kxkx1 hv,xi=µ1 2³ kxk2±kxkx1´¶1 2 This makes v1=x1±kxk (1 2(kxk2±kxkx1))1 2.With the construction of Hvto “elim- inate” the first column of A, we relabel the vector vasv1(with the small 4.1. BASICS 169 possibility of notational confusion) and move to describe Hv2using a trans- formation of the same kind with a vector of the type v=( 0,x2,x3,...,x n). Such a selection will not a ffect the structure of the first column. Continue this until the matrix is triangularized. The upshot is that every matrix canbe factored as A=UT, where Uis unitary and Tis upper triangular. The main result for this section is the factorization theorem. As it turns out, every Unitary matrix can be written as a product of elementary re flectors. The proof requires a slightly more general notion of re flector. Definition 4.1.6. The general form of the Householder matrix, also called anelementary re flector , has the form H v=I−τvv∗ where the vector v∈Cn. In order for Hvto be unitary, it must be true that HvH∗ v=(I−τvv∗)(I−τvv∗)∗ =(I−τvv∗)(I−¯τvv∗) =I−τvv∗−¯τvv∗+|τ|2(v∗v)vv∗ =I−2Re (τ)vv∗+|τ|2|v|2vv∗ =I Therefore, for v6=0w em u s th a v e −2Re (τ)vv∗+|τ|2|v|2vv∗=³ −2Re (τ)+|τ|2|v|2´ vv∗ =0 or −2Re (τ)+|τ|2|v|2=0 Now suppose that Q∈Mn(R) is orthogonal and that the spectrum σ(Q)⊂ {−1,1}.Suppose Qhas a complete set of normalized orthogonal eigen- vectors1, it can be expressed as Q=nP i=1λivivT iwhere the set v1,...,v nare the eigenvectors and λi⊂σ(Q). Now assume the eigenvectors have been arranged so that this simpli fies to Q=−kX i=1vivT i+nX i=k+1vivT i 1This is in fact a theorem that will be established in Chapter 4.2. It follows as a consequence of Schur’s theorem 170 CHAPTER 4. UNITARY MATRICES Here we have just arrange the eigenvectors with eigenvalue −1t oc o m e first. Define Hj=I−2vjvT j,j =1,...,k It follows that Q=kY j=1Hj=kY j=1¡ I−2vjvT j¢ for it is easy to check that Uvm=kY j=1Hj=½−vmifm≤k vm ifm>k Since we have agreement with Qon a basis, the equality follows. In words we may say that an orthogonal matrix Uwith spectrum σ(Q)⊂{−1,1}with can be written as a product of re flectors. With this simple case out of t h ew a y ,w ec o n s i d e rm o r eg e n e r a lc a s ew ew r i t e Q=nP i=1λivivT i.where the setv1,...,v nare the eigenvectors and λi⊂σ(Q). We wish to represent Q similar to the above formula as a product of elementary re flectors. Q=kY j=1Hj=kY j=1¡ I−τjwjw∗ j¢ where wi=αivifor some scalars αi. On the one hand it must be true that−2Re (τi)+|τi|2|wi|2=0f o re a c h i=1,...n , and on the other hand it must follow that (I−τiwiw∗ i)vm=½λiviifm=i vmifm6=i The second of these relations is automatically satis fied by the orthogonality of the eigenvectors. The second relation can needs to be solved. This simpli fies to ( I−τiwiw∗ i)vi=³ 1−τi|αi|2´ vi=λivi.T h e r e f o r e , i t i s necessary to solve the system −2Re (τi)+|τi|2|vi|2=0 1−τi|αi|2=λi Having done so there results the factorization. 4.2. SCHUR’S THEOREM 171 Theorem 4.1.5. LetQ∈Mnbe real orthogonal or unitary. Then Qcan be factored as the product of elementary re flectors Q=nQ j=1³ I−τjwjw∗ j´ , where the wjare the eigenvectors of Q. Note that the product written here is up to nwhereas the earlier product was just up to k. The difference here is slight, for by taking the scalar ( α) equal zero when necessary, it is possible to equate the second form to thefirst form when the spectrum is contained in the set {−1,1}. 4.2 Schur’s theorem It has already been established in Theorem 3.5.2 that every matrix is similar to a triangular matrix. A far stronger result is possible. Called Schur’stheorem, this result proves that the similarity is actually unitary similarity. Theorem 4.2.1 (Schur’s Theorem). Every matrix A∈M n(C)is uni- tarily equivalent to a triangular matrix. Proof. We proceed by induction on the size of the matrix n. First suppose theAis 2×2.Then for a given eigenvalue λand normalized eigenvector vform the matrix Pwith its first column vand second column any vector orthogonal to vwith norm one. Then Pis an unitary matrix and P∗AP=·λ∗ 0∗¸ . This is the desired triangular form. Now assume the Schur factorization is possible for matrices up to size ( n−1)×(n−1). For the given n×nmatrix Aselect any eigenvalue λ. With its pertaining normalized eigenvector vconstruct the matrix Pwith vin the first column and an orthonormal complementary basis in the remaining n−1 columns. Then P∗AP= λˆa 12··· ˆa1n 0ˆa22··· ˆa2n ............ 0ˆan2··· ˆann = λˆa 12··· ˆa1n 0 ... A2 0  The eigenvalues of A 2together with λconstitute the eigenvalues of A.B y the inductive hypothesis there is an ( n−1)×(n−1) unitary matrix ˆQsuch that ˆQ∗A2ˆQ=T2,where T2is triangular. Now embed ˆQin an n×nmatrix 172 CHAPTER 4. UNITARY MATRICES Qas shown below Q= λ0··· 0 0 ... ˆQ 0  It follows that Q ∗P∗APQ = λ0··· 0 0 ... T 2 0 =T is triangular, as indicated. Therefore the factorization is complete upon defining the similarity transformation U=PQ. Of course, the eigenvalues ofA 2together with λconstitute the eigenvalues of A, and therefore by similarity the diagonal of Tcontains only eigenvalues of A. The triangular matrix above is not unique as is easy to see by mixing or permuting the eigenvalues. An import ant consequence of Schur’s theorem pertains to unitary matrices. Suppose that Bis unitary and we apply the Schur factorization to write B=U∗TU.T h e n T=UBU∗. It follows that the upper triangular matrix is itself unitary, which is to say T∗T=I.It is a simple fact to prove this implies that Tis in fact a diagonal matrix. Thus, the following result is proved. Corollary 4.2.1. Every unitary matrix is diagon alizable. Moreover, every unitary matrix has northogonal eigenvectors. Proposition 4.2.1. IfB,A∈M2are similar and tr (B∗B)=tr(A∗A),t h e n BandAare unitarily equivalent. This result higher dimensions. Theorem 4.2.2. Suppose that A∈Mn(C)has distinct eigenvalues λ1...λk with multiplicities n1,n2,... ,n krespectively. Then Ais similar to a matrix of the form  T 1 0 T2 0... Tk  4.2. SCHUR’S THEOREM 173 where Ti∈Mniis upper triangular with diagonal entries λi. Proof. LetE1be eigenspace of λ1. Assume dim E1=m1.L e t u1...u m1be an orthogonal basis of E1, and suppose v1...v n−m1is an orthonormal basis ofCn−E1.T h e n w i t h P=[u1...u m1,v1...v n−m1]w eh a v e P−1APhas the block structure ·T10 0A2¸ . Continue this way, as we have done before. However, if dim Ei<m iwe must proceed a di fferent way. First apply Schur’s Theorem to triangularize. Arrange the firstm1eigenvalues to the firstm1diagonal entries of T,t h e similar triangular matrix. Let Er,sb et h em a t r i xw i t ha1i nt h e r, sposition and 0’s elsewhere. We assume throughout that r6=s. Then, it is easy to see that I+αEr,sis invertible and that ( I+αErs)−1=I−αErs.N o w i f we de fine (I+αErs)−1T(I+αErs) we see that trs→trs+α(trr−tss). We know trr−tss6=0i f randspertain to values with di fferent eigenvalues. Now this value is changed, but so also are the values above and to the right of trs.This is illustrated below. columns → s row r ↑ ∗↑ ∗→ →  To see how to use these similarity transformation to zero out the upper blocks, consider the special case with just two blocks A= T 1 T12 0 T2  We will give a procedure to use elementary matrices αE ijto zero-out each of the entries in T12.The matrix T12hasm1·m2entries and we will need 174 CHAPTER 4. UNITARY MATRICES to use exactly that many of the matrices αEij.In the diagram below we illustrate the order of removal (zeroing-out) of upper entries  T 1m12m1···m1·m2 ......... 2... 1m1+1 ··· ··· 0 T2  W h e np e r f o r m e di nt h i sw a yw ed e v e l o p m 1·m2constants αij.We proceed in the upper right block ( T12) in the left-most column and bottom entry, that is position ( m1,m1+1 ).For the ( m1,m1+ 1) position we use the elementary matrixαEm1,m1+1to zero out this position. This is possible because the diagonal entries of T1andT2are different. Now use the elementary matrix αEm1−1,m1+1to zero-out the position ( m1−1,m1+ 1). Proceed up this column to the first row each time using a new α.W e a r e finished when all the entries in column m1+1 f r o m r o w m1to row 1 are zero. Now focus on the next column to the right, column m1+ 2. Proceed in the same way from the ( m1,m1+ 2) position to the (1 ,m1+ 1) position, using elementary matrices αEk,m 1+2zeroing out the entries ( k,m 1+ 2) positions, k=m1,..., 1. Ultimately, we can zero out every entry in the upper right block with this marching left-to-right, bottom-to-top procedure. In the general scheme with k×kblocks, next use the matrices T2and T3to zero-out the block matrix T23.(See below.) T= T 10T13 T2T23 T3... ... 0... Tk−1Tk−1,k Tk  After that it is possible using blocks T 1andT3to zero-out the block T13. This is the general scheme for the entire triangular structure, moving downthe super-diagonal ( j.j+ 1) blocks and then up the (block) columns. In this way we can zero-out any value not in the square diagonal blocks pertaining to the eigenvalues λ 1...λk. 4.2. SCHUR’S THEOREM 175 Remark 4.2.1. IfA∈Mn(R)a n dσ(A)⊂R, then all operations can be carried out with real numbers. Lemma 4.2.1. LetJ⊂Mnbe a commuting family. Then there is a vector x∈Cnfor which xis an eigenvector for every A∈J. Proof. LetW⊂Cnbe a subspace of minimal dimension that is invariant under J.S i n c e Cnis invariant, we know Wexists. Since Cnis invariant, we know Wexists. Suppose there is an A∈Jfor which there is a vector inWwhich is not an eigenvector of A.D e fineW0={y∈W|Ay= λyfor some λ}.T h a t i s , W0is a set of eigenvectors of A.S i n c e Wis invariant under A, it follows that W06=φ. Also, by assumption. W0$W. For any x∈W0 ABx =(AB)x=B(Ax)=λBx and so Bx∈W0. It follows that W0isJinvariant, and W0has lower positive dimension than W. As a consequence we have the following result. Theorem 4.2.3. LetJbe a commuting family in Mn.I fA∈Jis diago- nalizable, then Jis simultaneously diagonalizable. Proof. Since Ais diagonalizable there are nlinearly independent eigenvec- tors of Aand if Sis in Mnand consists of those eigenvectors the matrix S−1ASis diagonal. Since eigenvectors of Aare the same as eigenvectors of B∈J, it follows that S−1BSis also diagonal. Thus S−1JS=D:= all diagonal matrices. Theorem 4.2.4. LetJ⊂Mnbe a commuting family. There is a unitary matrix U∈Mnsuch that U∗AU is upper triangular for each A∈J. Proof. From the proof of Schur’s Theorem we have that the eigenvectors chosen for Uare the same for all A∈J. This follows because after the first step we have reduced AandBto ·A11A12 0A22¸ and·B11B12 0B22¸ respectively. Commutativity is preserved under simultaneous similarity and therefore A22andB22commutes. Therefore at the second step the same eigenvector can be selected for allB22⊂J2, a commuting family. 176 CHAPTER 4. UNITARY MATRICES Theorem 4.2.5. IfA∈Mn(R), there is a real orthogonal matrix Q∈ Mn(R)such that (?) QTAQ= A1 A2? 0... Ak 1≤k≤n where each Aiis a real 1×1matrix or a real 2×2matrix with a non real pair of complex conjugate eigenvalues. Theorem 4.2.6. LetJ⊂Mn(R)be a commuting family. There is a real orthogonal matrix Q∈Mn(R)for which QTAQ has the form (?)for every A∈J. Theorem 4.2.7. Suppose A, B∈Mn(C)have eigenvalues α1,... ,αnand β1,... ,βnrespectively. If AandBcommute then there is a permutation i1...i nof the integers 1,... ,n for which the eigenvalues of A+Bareαj+ βij,j=1,... ,n .T h u sσ(A+B)⊂σ(A)+σ(B). Proof. Since J={A, B}f o r m sac o m m u t i n gf a m i l yw eh a v et h a te v e r y eigenvector of Ais an eigenvector of B, and conversely. Thus if Ax=αjx we must have that Bx=Bijx. But how do we get the permutation? The answer is to simultaneously triangularize with U∈Mn.W eh a v e U∗AU=Tand U∗BU=R. Since U∗(A+B)U=T+R we have that the eigenvalues pair up as described. Note: We don’t necessarily need AandBto commute. We need only the hypothesis that AandBare simultaneously diagonalizable. IfAandBdo not c o m m u t el i t t l ec a nb es a i do f σ(A+B). In particular σ(A+B)$σ(A)+σ(B). Indeed, by summing upper triangular and lower triangular matrices we can exhibit a range of possibilities. Let A=[01 00], B=[00 10],σ(A+B)={−1,1}butσ(A)=σ(B)={0}. Corollary 4.2.1. Suppose A, B∈Mnare commuting matrices with eigen- valuesα 1,... ,αnandβ1,... ,βnrespectively. If αi6=−βjfor all 1≤i, j≤n, thenA+Bis nonsingular. 4.3. EXERCISES 177 Note: Giving conditions for the nonsingularity of A+Bin terms of various conditions on AandBis a very di fferent problem. It has many possible answers. 4.3 Exercises 1. Characterize all diagonal unitary matrices and all diagonal orthogonal matrices in Mn. 2. Let T(θ)=£cosθ−sinθ sinθcosθ¤ whereθis any real. Then T(θ)i srealorthog- onal. Prove that if U∈M2(R) is real orthogonal, then Uhas the form T(θ)f o rs o m e θor U=·10 0−1¸ T(θ), and conversely. 3. Let us de fine the n×nmatrix Pto be a w-permutation matrix if for every vector x∈Rn, the vector Px has the same components asxin value and number, though possibly permuted. Show that w-permutation matrices are permutation matrices. 4. In R2identify all Householder transformations that are rotations. 5. Given any unit vector win the plane formed by eiandej.E x p r e s s the Householder matrix for this vector. 6. Show that the set of unitary matrices on Cnforms a subgroup of the subset of GL(n,C). 7. In R2,prove that the product of two (Householder) re flections is a rotation. (Hint. If the re flection angles are θ1andθ2,then the rotation angle is 2 ( θ1−θ2).) 8. Prove that the unitary matrices are closed the norm. That is, if the sequence {Un}⊂Mn(C) are all unitary and lim n→∞Un=Uin the k·k2norm, then Uis also unitary. 9. Let A∈Mnbe invertible. De fineG=Ak,(A−1)k,k=1,2,.... Show thatGis a subgroup of Mn. Here the group multiplication is matrix multiplication. 178 CHAPTER 4. UNITARY MATRICES 10. Prove that the unitary matrices are closed under pointwise conver- gence. That is, if the sequence {Un}⊂Mn(C) are all unitary and limn→∞Un=Ufor each ( ij)entry, then Uis also unitary. 11. Suppose that A∈Mnand that AB=BAfor all B∈Mn.Show that Ais a multiple of the identity. 12. Prove that a unitary matrix Ucan be written as V−1W−1VW for unitary V,W if and only if det U= 1. (Bellman, 1970) 13. Prove that the only triangular unitary matrices are diagonal matrices. 14. We know that given any bounded sequence of numbers, there is a convergent subsequence. (Bolzano-Weierstrass Theorem). Show that the same is true for matrices for any given matrix norm. In particular,show that if U nis any sequence of unitary matrices, then there is a convergent subsequence. 15. The Hadamard Gate (from quantum computing) is de fined by the matrix H=1√ 2·11 1−1¸ . Show that this transformation is a House- holder matrix. What are its eigenvalues and eigenvectors? 16. Show that the unitary matrices do not form a subspace Mn(C). 17. Prove that if a matrix A∈Mn(C) preserves the orthonormality of one orthonormal basis, then it must be unitary. 18. Call the matrix Aacheckerboard matrix if either (I)aij=0 ifi+jis even or (II) aij=0ifi+jis odd We call the matrices of type I even checkerboard and type II odd checkerboard .D efineCH(n)t ob ea l l n×ninvertible checkerboard matrices. The questions below all pertain to square matrices. (a) Show that if nis odd there are no invertible even checkerboard matrices. (b) Prove that every unitary matrix Uhas determinant with modulus one. (That is, |detU|=1.) (c) Prove that nis odd the product of odd checkerboard matrices is odd. 4.3. EXERCISES 179 (d) Prove that if nis even then the product of an even and an odd checkerboard matrix is odd, while the product of two even (orodd) checkerboard matrices is even. (e) Prove that if nis even then the inverse of any checkerboard matrix is a checkerboard matrix of the same type. However, in light of(a), it is only true that if nis an invertible odd checkerboard matrix, its inverse is odd. (f) Suppose that n=2m. Characterize all the odd checkerboard invertible matrices. (g) Prove that CH(n) is a subgroup of GL(n). (h) Prove that the invertible odd checkerboard matrices of any size nforms a subgroup of CH(n). In connection with chess, checkerboard matrices give the type of chess board on which the maximum number of mutually non-attacking knightscan be placed on the even ( i+jis even) or odd ( i+jis odd) positions. 19. Prove that every unitary matrix can be written as the product of a unitary diagonal matrix and another unitary matrix whose first column has nonnegative entries. 180 CHAPTER 4. UNITARY MATRICES Chapter 5 Hermitian Theory Hermitian matrices form one of the most useful classes of square matri- ces. They occur naturally in a variety of applications from the solution ofpartial di fferential equations to signal and image processing. Fortunately, they possess the most desirable of matrix properties and present the user with a relative ease of computation. There are several very powerful factsabout Hermitian matrices that have found universal application. First thespectrum of Hermitian matrices is real. Second, Hermitian matrices have acomplete set of orthogonal eigenvectors, which makes them diagonalizable.Third, these facts give a spectral repre sentation for Hermitian matrices and a corresponding method to approximate them by matrices of less rank. 5.1 Diagonalizability of Hermitian Matrices Let’s begin by recalling the basic de finition. Definition 5.1.1. LetA∈Mn(C). We say that AisHermitian ifA=A∗, where A∗=¯AT.A∗is called the adjoint of A. This, of course, is in con flict with the other de finition of adjoint, which is given in terms of minors. Recall the following facts and de finitions about subspaces of Cn: •IfU, V are subspaces of Cn,w ed e fine the direct sum ofUandVby U⊕V={u+v|u∈U, v∈V}. •IfU, V are subspaces of Cn,w es a y UandVare orthogonal if hu, vi=0 for every u∈Uandv∈V.I nt h i sc a s ew ew r i t e U⊥V. For example, a natural way to obtain orthogonal subspaces is from ortho- normal bases. Suppose that {u1,...,u n}is an orthonormal basis of Cn.Let 181 182 CHAPTER 5. HERMITIAN THEORY the integers {1,...,n }be divided into two (disjoint) subsets J1andJ2.Now define U1=S{ui|i∈J1} U2=S{ui|i∈J2} Then U1andU2are orthogonal, i.e. U1⊥U2,and U1⊕U2=Cn hAu, x i=hu, Ax i=λhu, xi=0 Our main result is that Hermitian matrices are diagonalizable. To prove it, we reveal other interesting and importa nt properties of Hermitian matrices. F o re x a m p l e ,c o n s i d e rt h ef o l l o w i n g . Theorem 5.1.1. LetA∈Mn(C)be Hermitian. Then the spectrum of A,σ(A),i sr e a l . Proof. Letλ∈σ(A) with corresponding eigenvector x∈Cn.T h e n hAx, x i=hx, Ax i=hx,λxi=¯λhx, xi k hλx, xi k λhx, xi. Since we know kxk2=hx, xi6= 0, it follows that λ=¯λ, which is to say that λis real. Theorem 5.1.2. LetA∈Mn(C)be Hermitian and suppose that λandµ are different eigenvalues with corresponding eigenvectors xandy.T h e n x⊥y(i.e. hx, yi=0). Proof. We know Ax=λxandAy=µy. Now compute hAx, y i=hx, A∗yi=hx, Ay i=µhx, yi k λhx, yi. Ifhx, yi6= 0, the equality above yields a contradiction and the result is proved. 5.1. DIAGONALIZABILITY OF HERMITIAN MATRICES 183 Remark 5.1.1. This result also follows from the previously proved result about the orthogonality of left and right eigenvectors pertaining to di fferent eigenvalues. Theorem 5.1.3. LetA∈Mn(C)be Hermitian, and let λbe an eigenvalue ofA. Then the algebraic and geometric multiplicities of λare equal. In symbols, ma(λ)=mg(λ). Proof. We prove this result by reconsideration of our main result on triangu- larization of Aby a similarity transformation. Let x∈Cnbe an eigenvector ofApertaining to λ.S o , Ax=λx.L e t u2,... ,u n⊂Cnbe a set of vectors orthogonal to x,s ot h a t {x, u 2,... ,u n}is a basis of Cn. Indeed, it is an orthogonal basis, and by normalizing the vectors it becomes an orthonor- mal basis. We claim that U2=S(u2,... ,u n), the span of {u2,... ,u n}is invariant under A. To see this, suppose u=nX j=2cjuj∈U2 and Au=v+ax where v∈U2anda6= 0. Then, on the one hand hAu, x i=hu, Ax i=λhu, xi=0 On the other hand hAu, x i=hv+ax, x i=ahx, xi=akxk26=0 This contraction establishes that the span of {u2,... ,u n}is invariant under A. LetU⊕V=Cnbe invariant subspaces with U⊥V. Suppose that UB and VBare orthonormal bases of the orthogonal subspaces UandV.Define the matrix P∈Mn(C) by taking for its columns first the basis vectors UB and and then the basis vectors VB.L e t u s w r i t e P=[UB,VB]( w i t ho n l y a small abuse of notation). Then, since Ais Hermitian, B=P−1AP =·Au0.......... 0Av¸ 184 CHAPTER 5. HERMITIAN THEORY The whole process can be carried out exactly ma(λ)t i m e s ,e a c ht i m e generating a new orthogonal eigenvector pertaining to λ.T h i s e s t a b l i s h e s thatmg(λ)=ma(λ). A formal induction could have been given. Remark 5.1.2. Note how we applied orthogonality and invariance to force the triangular matrix of the previous result to become diagonal. This is what permitted the successive extracti on of eigenvectors. Indeed, if for any eigenvector xthe subspace of Cnorthogonal to xis invariant, we could have carried out the same steps as above. We are now in a position to state our main result, whose proof is implicit in the three lemmas above. Theorem 5.1.4. LetA∈Mn(C)be Hermitian. Then Ais diagonalizable. The matrix Pfor which P−1AP is diagonal can be taken to be orthogonal. Finally, if {λ1,...,λn}and {u1,...,u n}denote eigenvalues and pertaining orthonormal eigenvectors for A,t h e n Aadmits the spectral representation A=Pn j=1λjujuT j. Corollary 5.1.1. LetA∈Mn(C)be Hermitian. (i)Ahasnlinearly independent and orthogonal eigenvectors. (ii)Ais unitarily equivalent to a diagonal matrix. (iii) If A, B∈Mnare unitarily equivalent, then Ais Hermitian if and only ifBis Hermitian. Note that in part (iii) above, the condition of unitary equivalence cannot be replaced by just similarity. (Why?) Theorem 5.1.5. IfA, B∈Mn(C)andA∼Bwith Sas the similarity transformation matrix, B=S−1AS.I f Ax=λxandy=S−1x,t h e n By=λy. If matrices are similar so also are their eigenstructures. It should establish the very closeness that similarity implies. Later as we consider decomposi- tion theorems, we will see even more remarkable consequences. Though we have as yet no method of determining the eigenvalues of a matrix beyond factoring the characteristic polynomial, it is instructive tosee how their existence impacts the fundamental problem of solving Ax=b. Suppose that Ais Hermitian with eigenvalues λ 1,...,λn,c o u n t e da c c o r d i n g to multiplicity and with o rthonormal eigenvectors {u1,...,u n}.C o n s i d e r 5.1. DIAGONALIZABILITY OF HERMITIAN MATRICES 185 the following solution method for the system Ax=b.Since the span of the eigenvectors is Cnthen b=nX i=1biui where as we know by the orthonormality of the vectors {u1,...,u n}that bi=hb, uii. We can also write x=Pn i=1xiui.Then, the system becomes Ax =AÃnX i=1xiui! =nX i=1xiλiui=nX i=1biui Therefore, the solution is xi=bi λi,i=1,...,n Expanding the data vector bin the basis of eigenvectors yields a rapid method to find the solution to the system. Nonetheless, this is not the preferred method for solving linear systems when the coe fficient matrix is Hermitian. Finding all the eigenvectors is usually costly, and other waysare available that are more e fficient. We will discuss a few of them in in the section and in later chapters. Approximating Hermitian matrices With the spectral representation available, we have a tool to approximate thematrix, keeping the “important” part and discarding the less important part.Suppose the eigenvalues are arranged in decending order |λ 1|≥···≥|λn|. Now approximate Aby Ak=kX j=1λjujuT j (1) This is an n×nmatrix. The di fference A−Ak=Pn j=k+1λjujuT j.We can approximate the norm of the di fference by (A−Ak)x= nX j=k+1λjujuT j x=nX j=k+1λjxjuj 186 CHAPTER 5. HERMITIAN THEORY where x=Pn j=1xjuj.Assume kxk= 1. By the Cauchy-Schwartz inequality k(A−Ak)xk2=°°°°°°nX j=k+1λjxjuj°°°°°°2 ≤nX j=k+1|λj|2 Therefore, k(A−Ak)k≤³Pn j=k+1|λj|2´1/2 . From this we can conclude that if the smaller eigenvalues are su fficiently small, the matrix can be ac- curately approximated by a matrix of lesser rank. Example 5.1.1. The matrix A= 0.5745−0.5005 0 .1005 0 .0000 −0.5005 1 .176−0.5756 0 .1005 0.1005−0.5756 1 .176−0.5005 0.0000 0 .1005−0.5005 0 .5745  has eigenvalues eigenvectors {2.004,0.9877,0.3219,0.1872 }with pertaining eigenvectors u1= 0.2740 −0.6519 0.6519 −0.2740 ,u2= 0.4918 −0.5080 −0.5080 0.4918 ,u3= 0.6519 0.2740 −0.2740 −0.6519 ,u 4= 0.5080 0.4918 0.4918 0.5080  respectively. Neglecting the eigenvec tors pertaining to the two smaller eigenvalues Ais approximated according as 1 the formula above by A2=2X j=1λjujuT j=λ1u1uT 1+λ2u2uT 2 =2 .004 0.274 −0.6519 0.6519 −0.274  0.274 −0.6519 0.6519 −0.274 T +0.9877 0.4918 −0.508 −0.508 0.4918  0.4918 −0.508 −0.508 0.4918 T A2= 0.3893−0.6047 0 .1112 0 .0884 −0.6047 1 .107−0.5968 0 .1112 0.1112−0.5968 1 .107−0.6047 0.0884 0 .1112−0.6047 0 .3893  5.2. FINDING EIGENVECTORS 187 The difference A−A2= 0.1852 0 .1042−0.0107−0.0884 0.1042 0 .069 0 .0212−0.0107 −0.0107 0 .0212 0 .069 0 .1042 −0.0884−0.0107 0 .1042 0 .1852  has 2-norm kA−A2k2=0.3218, while the 2-norm kAk2=2.004. The relative error of approximation iskA−A2k2 kAk2=0.3218 2.004=0.1606. To illustrate how this may be used, let us attempt to use A2to approx- imate the solution of Ax=b,w h e r e b=[ 2.606,−4.087,1.113,0.346 4]T. First of all the exact solution is x=[ 2.223,−2.688,−0.162 9,0.931 2]T. Since the matrix A2has rank two, it is not solvable for every vector b. We therefore project the vector binto the span of the range of A2,n a m e l y u1andu2.T h u s b2=hb, u1iu1+hb, u2iu2=[ 2.555,−4.118,1.109,0.358 4]T Now solve A2x2=b2,t oo b t a i n x2=[ 2.023,−2.828,−0.220 2,0.927 4]T. The 2-norm of the di fference is kx−x2k2=0.250 8. This error, though not extremely small, can be accounted for by the fact that the data vector bhas sizable u3andu4components. That is°°projS(u3,u4)b°°=0.250 8. 5.2 Finding eigenvectors Recall that a zero of a polynomial is called simple if its multiplicity is one. If the eigenvalues of A∈Mn(C) are distinct and the largest, λn, in modulus is simple, then there is an iterative method to find it. Assume (1) |λn|=ρ(A) (2) ma(λn)=1 . Moreover, without loss of generality we assume eigenvalues to be ordered |λ1|≤···≤|λn−1|<|λn|=ρn. Select x(0)∈C(n).D efine x(k+1)=1 kx(k)kAx(k). By scaling we can assume that λn=1 ,a n dt h a t y(1),... ,y(n)are linearly independent eigenvectors of A.S oAy(n)=y(n).W ec a nw r i t e x(0)=c1y(1)+···+cny(n). 188 CHAPTER 5. HERMITIAN THEORY Then, except for a scale factor (i.e. the factor kx(k)k−1) x(k)=c1λk 1y(1)+···+cnλk ny(n). Since |λj|<1,j=1,2,... ,n−1, we have that |λk j|→0i fj=1,2,... ,n−1. Therefore, the limit of x(k)approaches a multiple of y(n). This gives the following result. Theorem 5.2.1. LetA∈Mn(C)have ndistinct eigenvalues and assume the eigenvalue λnwith modulus ρ(A)is simple. If x(0)is not orthogonal to the eigenvector y(n)pertaining to λn, then the sequence of vectors de fined by x(k+1)=1 kx(k)kAx(k) converges to a multiple of y(n).T h i si sc a l l e dt h e Power Method . The rate of convergence is controlled by |λn−1|. The closer to 1 this number is the slower the iterates converge. Also, if we know only that λn (for which |λn|=ρ(A) is simple we can determine what it is by considering theRayleigh quotient. Take ρk=hAx(k),x(k)i hx(k),x(k)i. Then lim k→∞ρk=λn. Thus the multiple of y(n)is indeed λn. Tofind intermediate eigenvalues and eigenvectors we apply an adaptation of the power method called the orthogonalization method . However, in order to adapt the power method to determine λn−1,our underlying assumption is that is also simple and morover |λn−2|<|λn−1|. Assume y(n)andλna r ek n o w n . T h e nw er e s t a r tt h ei t e r a t i o n ,t a k i n g the starting value ˆx(0)=x(0)−hx(0),y(n)i ky(n)k2y(n). We know that the eigenvector y(n−1)pertaining to λn−1is orthogonal to y(n). Thus, in theory all of the iterates ˆx(k+1)=1 kˆx(k)kTˆx(k) 5.2. FINDING EIGENVECTORS 189 will remain orthogonal to y(n). Therefore, lim k→∞ˆx(k)=y(n−1). –in theory. In practice, however, we must accept that y(n)has not been determined exactly. This means ˆ x(0)has not been purged of all of y(n).B y our previous reasoning, since λnis the dominant eigenvalue, the presence of y(n)willcreep back into the iterates ˆ x(k). To reduce the contamination it is best to purify the iterates ˆ x(k)periodically by the reduction (?)ˆ x(k)−→ˆx(k)−hˆx(k),y(n)i ky(n)ky(n) before computing ˆ x(k+1). The previous argument can be applied to prove that (2) lim k→∞x(k)=y(n−1) (2) lim k→∞hAx(k),x(k)i kx(k)k2=λn−1. Additionally, even if we know y(n)exactly, round-o fferror would reinstate ay(n)component in our iterative computations. Thus the puri fication step above, ( ?), should be applied in allcircumstances. Finally, subsequent eigenvalues and eigenvectors may be determined by successive orthogonalizations. Again th e eigenvalue simplicity and strict inequality is needed for convergence. Speci fically, all eigenvectors can be determined if we assume that eigenvalues to be strictly ordered |λ1|<···< |λn−1|<|λn|=ρn. For example, we begin the iterations to determine y(n−j) with ˆx(0)=x(0)−j−1X i=0hx(0),y(n−i)i ky(n−i)k2y(n−i). Don’t forget the re-orthogonalizations periodically throughout the iterative process. W h a tc a nb ed o n et o find intermediate eigenvalues and eigenvectors in the case Ais not symmetric? The method above fails, but a variation of it works. What must be done is to generate the left and right eigenvectors, w(n) andy(n),f o r A. Use the same process. To compute y(n−1)andw(n−1)we 190 CHAPTER 5. HERMITIAN THEORY orthogonalize thusly: ˆx(0)=x(0)−hx(0),w(n)i kw(n)k2w(n) ˆz(0)=z(0)−hz(0),y(n)i ky(n)k2y(n) where z(0)is the original starting value used to determine the left eigenvector w(n).S i n c ew ek n o wt h a t lim k→∞x(k)=αy(n) it is easy to see that lim k→∞Ax(k)=αλny(n). Therefore, lim k→∞hAx(k),x(k)i hx(k),x(k)i=λn. Example 5.2.1. LetTbe the transformation of R2→R2that rotates a vector by θradians. Then it is clear that no matter what nonzero vector x(0)is selected the iterations x(k+1)=Tx(k)will never converge. (Assume kx(0)k= 1.) Now the matrix representation of Tis AT=·cosθ−sinθ sinθcosθ¸ rotates counterclockwise we have pAT(λ)=d e t·λ−cosθ sinθ −sinθλ−cosθ¸ =(λ−cosθ)2+s i n2θ. The spectrum of ATis therefore λ=c o sθ±isinθ. Notice that the eigenvalues are discrete, but there are twoeigenvalues with modulus ρ(AT) = 1. The above results therefore do not apply. 5.3. POSITIVE DEFINITE MATRICES 191 Example 5.2.2. Although the previous example is not based on a symmet- ric matrix, it certainly illustrates non convergence of the power iterations. The even simpler Householder matrix·10 0−1¸ furnishes us with a sym- metric matrix for which the iterations also do not converge. In this case, there are two eigenvalues with modulus equal to the spectral radius ( ±1). With arbitrary starting vector x(0)=[a, b]T,i ti so b v i o u st h a tt h ee v e n iterations are x(2i)=[a, b]Ta n dt h eo d di t e r a t i o n sa r e x(2i−1)=[a,−b]T. Assuming again that the eigenvalues are distinct and even stronger, as- suming that |λ1|<|λ2|<···<|λn| we can apply the process above to extract all the eigenvalues (Rayleigh quotient) and the eigenvectors, one-by-one, when AAAis symmetric . First of all, considering the matrix A−σIwe can shift the eigenvalues to either the left or the right. Depending on the location of λnas ufficiently large |σ|may be chosen so that |λ1−σ|=ρ(A−σI). The power method can be applied to determine λ1−σand hence λ1. 5.3 Positive de finite matrices Of the many important subclasses of Hermitian matrices, there is one class that stands out. Definition 5.3.1. We say that A∈Mn(C)i spositive de finiteifhAx, x i> 0 for every nonzero x∈Cn. Similarly, we say that A∈Mn(C)i spositive semide finiteifhAx, x i≥0 for every nonzero x∈Cn. It is easy to see that for positive de finite matrices all of the results are true Theorem 5.3.1. LetA, B∈Mn(C).T h e n 1. If Ais positive de finite, then σ(A)⊂R+ n 2. If Ais positive de finite, then Ais invertible. 3.B∗Bis positive semide finite. 4. If Bis invertible then B∗Bis positive de finite. 192 CHAPTER 5. HERMITIAN THEORY 5. If B∈Mn(C)is positive semide finite, then diag (B)is nonnegative, and diag (B)is strictly positive when Bis postive de finite. The proofs are all routine. Of course, every diagonal matrix with non- negative entries is positive semide finite. Square roots Given a real matrix A∈Mn. It is sometimes desired to determine a square root of A.B y t h i s w e m e a n a n y m a t r i x Bfor which B2=A.Moreover, if possible, it is desired that the square root be real. Our experience withnumbers indicates that in order that a number have a positive square root,it must be positive. The analogue for matrices is the condition of beingpositive de finite. Theorem 5.3.2. LetA∈M nbe positive [semi-]de finite. Then Ahas a real square root. Moreover, the square r oot can taken to be positive //[semi- ]definite. Proof. We can write the diagonal matrix of the eigenvalues of Ain the equa- tionA=P−1DP. E x t r a c tt h ep o s i t i v es q u a r er o o to f DasD1 2=diag(λ1/2 1,...,λ1/2 n). Obviously D1 2D1 2=D.N o w d e fineA1 2byA1 2=P−1D1 2P.T h i s m a t r i x i s real. It is simple to check that A1 2A1 2=A,and that this particular square root is positive de finite. Clearly any real diagonalizable matrix with nonnegative eigenvectors has a real square root as well. However, bey ond that conditions for determin- ing existence let alone determination of square roots take us into a very specialized subject. 5.4 Singular Value Decomposition Definition 5.4.1. For any A∈Mmn,t h e n×nHermitian matrix A∗Ais positive semi-de finite. Denoting its eigenvalues by λjwe called the valuesp λjthesingular values ofA. Because r(A∗A)≤min ( r(A∗),r(A))≤min(m, n)t h e r ea r ea tm o s t min ( m, n) nonzero singular values. Lemma 5.4.1. LetA∈Mmn.There is an orthonormal basis {u1,...,u n}of Cnsuch that {Au1,...,A u n}is orthogonal. 5.4. SINGULAR VALUE DECOMPOSITION 193 Proof. Proof. Consider the n×nHermitian matrix A∗A, and denote an orthonormal basis of its eigenvectors by {u1,...,u n}.T h e ni ti se a s yt os e e that {Au1,...,A u n}is an orthogonal set. For hAuj,A u ki=hA∗Auj,uki= λjhuj,uki=0. Lemma 5.4.2. Lemma 2 Let A∈Mmnand an orthonormal basis {u1,...,u n}of Cn.D efine vj=( 1 kAujkAujifkAujk6=0 0 ifkAujk=0 LetS=diag(kAu1k,..., kAunk),t h e n×nmatrix Uhaving rows given by the basis {u1,...,u n}and ˆVthem×nmatrix given by the columns {v1,...,v n}.T h e n A=ˆVS U . Proof. Proof. Consider ˆVS Uu j= v1v2 vn ↓↓ ↓  kAu1k 0 ··· 0 0 kAu2k00 ......... 00 ··· k Aunk  u1−→ u2−→ un−→ uj = v1v2 vn ↓↓ ↓  kAu 1k 0 ··· 0 0 kAu2k00 ......... 00 ··· k Aunk e j = v1v2 vn ↓↓ ↓ kAujkej=kAujkvj=Auj Thus both ˆVS U andAhave the same action on a basis. Therefore they are equal. It is easy to see that kAujk=p λj, that is the singular values. While A=ˆVS U could be called the singular value decomposition (SVD), what is usually o ffered at the SVD is small modi fication of it. Rede fine the matrix Uso that the firstrcolumns pertain to the nonzero singularvalues. De fine Dto be the m×nmatrix consisting of the non zero singular values in the 194 CHAPTER 5. HERMITIAN THEORY djjpositions, and filled in with zeros else where. De fine the matrix Vto be thefirstrof the columns of ˆVand if r<m construct an additional m−r orthonormal columns so that Vis an orthonormal basis of Cn.The resulting product VD U ,c a l l e dt h e singular value decomposition ofA,i se q u a lt o A,and moreover it follows that Vism×m, D ism×n,andUisn×n. This gives the following theorem Theorem 5.4.1. LetA∈Mmn.Then there is an m×morthogonal matrix V,ann×northogonal matrix U,a n da n m×nmatrix Dwith only diagonal entries such that A=VD U . The diagonal entries of Dare the singular values of Aand the rows of Uare the eigenvectors of A∗A. Example 5.4.1. The singular value decomposition can be used for image compression. Here is the idea. Consider all the eigenvalues of A∗Aand order them greatest to least. Zero the matrix Sfor all eigenvalues less than some threshold. Then in the reconstruc tion and transmission of the matrix, it is not necessary to include the vectors pertaining to these eigenvalues. Inthe example below, we have considered a 164 ×193 pixel image of C. F. Gauss (1777-1855) on the postage stamp issued by Germany on Feb. 23, 1955, to commemorate the centenary of death. Therefore its spectrum has164 eigenvalues. The eigenvalues range from 26,603.0 to 1.895. A plotof the eigenvalues shown below. Now compress the image, retaining only afraction of the eigenvalues by e ffectively zeroing the smaller eigenvalues. 5.4. SINGULAR VALUE DECOMPOSITION 195 Note that the original image of the stamp has been enlarged and resampled for more accurate comparisons. This image (stored at 221 dpi) is displayed at effective 80 dpi with the enlargement. Below we show two plots where we have retained respectively 30% and 10% of the eigenvalues. There is anapparent drastic decline in the image quality at roughly 10:1 compression. In this image all eigenvalues smaller than λ= 860 have been zeroed. Using 48 of 164 eigenvalues Using 16 of 164 eigenvalues 196 CHAPTER 5. HERMITIAN THEORY 5.5 Exercises 1. Prove Theorem 5.3.1 (i). 2. Prove Theorem 5.3.1 (ii). 3. Suppose that A, B∈Mn(C) are Hermitian. We will say A<0i fA is non-negative de finite. Also, we say A<BifA−B<0. Is “ <” an equivalence relation? If A<BandB<Cprove or disprove that A<C. 4. Describe all Hermitian matrices of rank one. 5. Suppose that A, B∈Mn(C) are Hermitian and positive de finite. Find necessary and su fficient conditions for ABto be Hermitian and also positive de finite. 6. For any matrix A∈Mn(C) with eigenvalues λi,i=1,...,n .P r o v e thatPn i=1|λi|2=Pn i,j=1|aij|2. Chapter 6 Normal Matrices Normal matrices are matrices that include Hermitian matrices and enjoy several of the same properties as Hermitian matrices. Indeed, while we proved that Hermitian matrices are uni tarily diagonalizable, we did not establish any converse. That is, if a matrix is unitarily diagonalizable, thendoes it have any special property involving for example its spectrum or itsadjoint? As we shall see normal matrices are unitarily diagonalizable. 6.1 Introduction to Normal matrices Definition 6.1.1. Am a t r i x A∈Mnis called normal ifA∗A=AA∗. Proposition 6.1.1. A∈Mnis normal if and only if every matrix unitarily equivalent to Ais normal. Proof. Suppose Ais normal and B=U∗AU,w h e r e Uis unitary. Then B∗B=U∗A∗AU=U∗AA∗U=U∗AUU∗A∗U=BB∗.I fU∗AUis normal then it is easy to see that U∗AA∗U=U∗A∗AU. Multiply this equation on the right by U∗a n do nt h el e f tb y Uto obtain AA∗=A∗A. Examples. (1) Unitary matrices are normal ( U∗U=I=UU∗). (2) Hermitian matrices are normal ( AA∗=A2=A∗A). (3) If A∗=−A,w eh a v e A∗A=AA∗=−A2. Hence matrices for which A∗=−A,c a l l e d skew-Hermitian , are normal. 197 198 CHAPTER 6. NORMAL MATRICES Example 6.1.1. Consider the arbitrary matrix N∈M2(R),written as N=·ab cd¸ . If we suppose that Nis normal then N∗N=·ab cd¸T·ab cd¸ =·a2+c2ab+cd ab+cd b2+d2¸ NN∗=·ab cd¸·ab cd¸T =·a2+b2ac+bd ac+bd c2+d2¸ From this we conclude that b2=c2,o rb=±c. Consider the cases in turn. (i) If c=b,t h e n Nis Hermitian and thus normal. (ii) If c=−b6=0,then ( N∗N)12=ab+cd=b(a−d). On the other hand ( NN∗)12=ac+bd=(d−a)b.F o r b(a−d)=( d−a)b,w em u s t have a=d.This gives that real 2 ×2 normal matrices are either symmetric or have the form N=·ab −ba¸ Note this form includes both rotations and skew-symmetric matrices. Recall the de finition of a unitarily diagonalizable matrix: A matrix A∈Mn is called unitarily diagonalizable if there is a unitary matrix Ufor which U∗AU is diagonal. A simple consequence of this is that if U∗AU =D (where D= diagonal and U= unitary), then AU=UD and hence Ahasnorthonormal eigenvectors. This is just a part of the spectral theorem for normal matrices . Theorem 6.1.1 (Spectral theorem for normal matrices). IfA∈Mn has eigenvalues λ1...λn, counted according to multiplicity, the following statements are equivalent. (a)Ais normal. (b)Ais unitarily diagonalizable. (c)Pn i=1Pn j=1|aij|2=Pn j=1|λj|2. (d) There is an orthonormal set of neigenvectors of A. 6.1. INTRODUCTION TO NORMAL MATRICES 199 Proof. (a)⇒(b). If Ais normal, then AA∗is Hermitian and therefore unitarily diagonalizable. Thus U∗A∗AU=D=U∗AA∗U.A l s o , A,A∗, A∗A=AA∗form a commuting family. This implies that eigenvectors of A∗Aare also eigenvectors of A.S i n c e A∗Ahas a complete orthonormal set we know that U∗AUis also diagonal. It is easy to see also that (b) ⇒(a) We also note that (a) ⇒(d). (b)⇒(c). Suppose U∗AU=D.T h e n U∗A∗U=D∗andU∗A∗AU=D∗D. By Corollary 3.5.2 similarly preserves the trace. We know trace of A∗A is tr ( A∗A)=Pn j=1Pn k=1a∗ jkakj=Pn j=1Pn k=1¯akjakj=Pn j=1Pn k=1|akj|2. Since the trace of D∗DisΣ|λj|2, the result follows. (c)⇒(b). We know that Ais unitarily equivalent to a upper triangular matrix T. We also know that if A∼Bare unitarily equivalentP ij|aij|2= P ij|bij|2. Application of this equality to the upper triangular matrix T yields X i,j|aij|2=X |λj|2+X j>i|tij|2=X |λj|2. Thus tij=0f o r j>i .T h u s Ais uniformly diagonalizable. (d)⇒(b). Trivial. Corollary 6.1.1. LetA∈MnandAis normal. If Uis unitary and if U∗AU is upper triangular then U∗AU is diagonal. Theorem 6.1.2. LetN∈Mn(R).T h e n Nis normal if and only if there is a real orthogonal matrix Q∈Mn(R)such that QTNQ= A 1 A2 ° ... ° An (1) where A iis1×1(real) or Aiis2×2(real) of the form Ai=·αiβj −βjαi¸ . Proof. First of all, any matrix Aof the form given by (1) is normal, and therefore so also is any matrix unitarily similar (real orthogonally similar in this case) to it. 200 CHAPTER 6. NORMAL MATRICES To prove the converse we assume that N∈Mn(R)i sn o r m a l .W ek n o w thatNis unitarily diagonalizable. That is, there is a unitary matrix Usuch thatU∗NU=D, the diagonal matrix of its eigenvalues. Because Nis real, all complex eigenvalues occur in compl ex conjugate pairs. Arrange them as successive diagonal entries in D.I fλis a real eigenvalue, we can assume without loss of generality that the corresponding eigenvector is real. For complex eigenvalues, the corresponding eigenvectors also occur in conjugatepairs. Thus if α+iβis an eigenvector of Nwith corresponding eigenvector written in real and complex parts u=u r+ius.S i n c e Nis real we have thatα−iβis also an eigenvector of Nwith corresponding eigenvector ¯ u= ur−ius.B y t h e f a c t t h a t Nis unitarily diagonalible, these vectors are orthogonal. This means hur,usi=0. Replace the eigenvectors ur±iusby the real an imaginary parts in U. This gives the matrix Q. Now compute QTNQ. It is easy to see that com- puteNur=αur−βvsandNus=αus+βvr. When the first of these vectors (αur−βvs) is multiplied by QTwe obtain the vector [0 ,...,α,−β,0,...0]T. Multiplication by the second gives the vector [0 ,...,β,α,0,...0]T.I n t h i s way the components·αβ −βα¸ arise. Corollary 6.1.2. (a)A∈Mnis symmetric if and only if (1) holds with all blocks 1×1(and real). (b)AAT=Iif and only if (1) has the form  λ1 ... λp ∗ A1 ∗... Ak  whereλ j=±1andAj=hcosθj−sinθj sinθjcosθji θj∈R. 6.2 Exercises 1. If AandBcommute and if Ais normal, then A∗andBcommute. Chapter 7 Factorization Theorems This chapter highlights a few of the many factorization theorems for ma- trices. While some factorization resul ts are relatively direct, others are it- erative. While some factorization results serve to simplify the solution tolinear systems, others are concerned with revealing the matrix eigenvalues.We consider both types of results here. 7.1 The PLU Decomposition The PLU decomposition (or factorization) To achieve LU factorization werequire a modi fied notion of the row reduced echelon form. Definition 7.1.1. The modified row echelon form of a matrix is that form which satis fies all the conditions of the modi fied row reduced echelon form except that we do not require zeros to be above leading ones, and moreoverwe do not require leading ones, just nonzero entries. For example the matrices below are in row echelon form. A= 123 001000 B= 1 230 04−76 0 001  Most of the factorizations A∈M n(C) studied so far require one essential ingredient, namely the eigenvectors of A. While it was not emphasized when we studied Gaussian elimination, there is a LU-type factorization there. Assume for the moment that the only operations needed to carry Ato its 201 202 CHAPTER 7. FACTORIZATION THEOREMS modi fied row echelon form are those that add a multiple of one row to another. The modified row echelon form of a matrix is that form which satisfies all the conditions of the modi fied row reduced echelon form except that we do not require zeros to be above leading ones, and moreover wedo not require leading ones, just nonzero entries. Naturally it is easy to make the leading nonzero entries into leading ones by the multiplication by an appropriate identity matrix. That is not the point here. What wewant to observe is that in this case the reduction is accomplished by the leftmultiplication of Aby a sequence of lower triangular matrices of the form. L= 1 01 0 ...01 c... 0··· 1  Since we pivot at the (1 ,1)-entry first, we eliminate all the entries in the first column below the first row. The product of all the matrices Lto accomplish this has the form L 1= 1 c 2110 c3101 ...... cn10··· 1  where c k1=−ak1 a11.Thus, with the notation that A=A1has entries a(1) ijthis first phase of the reduction renders the matrix A2with entries a(2) ij A2=L1A1= a (2) 11 ··· a(2) 1n 0a(2) 22 ···... 0a(2) 32a(2)33 ......... 0a(2) n2··· a(2) nn  Since we have assumed that no row interchanges are necessary to carry out the reduction we know that a (2) 226=0.The next part of the reduction process is the elimination of the elements in the second column below the second 7.1. THE PLU DECOMPOSITION 203 row, i.e. a(2) 32→0, ...a(2) n2→0.Correspondingly, this can be achieved by a matrix of the form L2= 1 01 0 0c 22 1 ......... 0cn2··· 1  (What are the values c k2?) The result is the matrix A3given by A3=L2A2=L2L1A1= a (3) 11 ··· a(3) 1n 0a(3) 22 ···... 00 a(3) 33............... 00 a (3) 3n a(3) nn  Proceeding in this way through all the rows (columns) there results A n=Ln−1An−1=Ln−1···L2L1A1= a (3) 11 ··· a(3) 1n 0a(3) 22 ···... 00 a(3) 33............... 000 a(3) nn  The right side of the equation above is an upper triangular matrix. Denote it by U.Since each of the matrices L i,i=1,...n−1i s i n v e r t i b l e w e c a n write A=L−1 1···L−1 n−1U The lemma below is useful in this. Lemma 7.1.1. Suppose the lower triangular matrix L∈Mn(C)has the 204 CHAPTER 7. FACTORIZATION THEOREMS form L= 1 0... 0 1 01 ......c k+1,k... ...... 0··· 0cnk 1 ←−k throw Then Lis invertible with inverse given by L−1= 1 0... 0 101 ......−c k+1,k... ...... 0··· 0−cnk 1 ←−k throw Proof. Trivial Lemma 7.1.2. Suppose L1,L2,···,Ln−1are the matrices given above. Then the matrix L=L−1 1···L−1 n−1has the form L= 1 −c 21 10 −c31−c32 1 1 .........−ck+1,k... ...... −cn1−cn2···−cnk ··· 1  Proof. Trivial. Applying these lemmas to the present situation we can say that when no row interchanges are needed we can factor and matrix A∈M n(C)a s A=LU,where Lis lower triangular and Uis upper triangular. When row 7.1. THE PLU DECOMPOSITION 205 interchanges are needed and we let Pbe the permutation matrix that creates these row interchanges then the LU-factorization above can be carried outfor the matrix PA. Thus PA=LU, where Lis lower triangular and Uis upper triangular. We call this the PLU factorization. Let us summarize this in the following theorem. Theorem 7.1.1. LetA∈M n(C). Then there is a permutation matrix P∈Mn(C)and lower Land upper Utriangular matrices ( ∈Mn(C)), such thatPA=LU. Moreover, Lcan be taken to have ones on its diagonal. That is,`ii=1,i=1,...n . By applying the result above to ATit is easy to see that the matrix U can be taken to have the ones in its diagonal. The result is stated as a corollary. Corollary 7.1.1. LetA∈Mn(C). Then there is a permutation matrix P∈Mn(C)and lower and upper triangular matrices ( ∈Mn(C)) respec- tively, such that PA=LU. Moreover, Ucan be taken to have ones on its diagonal ( uii=1,i=1,...n ). The PLU decomposition can be put in service to solving the system Ax=bas follows. Assume that A∈Mn(C) is invertible. Determine the permutation matrix Pin order that PA=LU, where Lis lower triangular andUis upper triangular. Thus, we have Ax =b PAx =Pb LUx =Pb Solve the systems Ly =Pb Ux =y Then LUx =Ly=Pb.Hence xis a solution to the system. The advantages of this formulation over the direct Gaussian elimination is that the systemsLy=PbandUx=yare triangular and hence are easy to solve. For example for the first of the systems, Ly=Pb,let the vector Pb=h ˆb 1,..., ˆbniT . Then it is easy to see that “back substitution” (aka “forward substitution”) 206 CHAPTER 7. FACTORIZATION THEOREMS can be used to determine y. That is, we have the recursive relations y1=ˆb1 l11 y2=ˆb2−l21y1 l22 ... yn=à ˆbn−n−1X m=1lnmym! l−1 nn A similar formula applies to solve Ux=y. I nt h i sc a s ew es o l v e first for xn=yn/unn.The general formula is recursive with xkbeing determined after xk+1,...,x n.are determined using the formula xk=à yk−nX m=k+1ukmym! u−1 kk In practice the step of determining and then multiplying by the per- mutation matrix is not actually carried out. Rather, an index array is generated, while the elimination step is accomplished that e ffectively inter- changes a “pointer” to the row interchanges. This saves considerable timein solving potentially very large systems. More general and instructive methods are available for accomplishing this LU factorization. Also, conditions are available for when no (nontrivial)permutation is required. We need the following lemma. Lemma 7.1.3. LetA∈M n(C)have the LU factorization A=LU,w h e r e Lis lower triangular and Uis upper triangular. For any partition of the matrix of the form A=·A11A12 A21A22¸ there are corresponding decompositions of the matrices LandU L=·L11 0 L21L22¸ and U=·U11U12 0U22¸ 7.1. THE PLU DECOMPOSITION 207 where the Liiand the Uii.are lower and upper triangular respectively. More- over, we have A11=L11U11 A21=L21U11 A12=L12U22 A22=L21U12+L22U22 Thus L11U11is a LU factorization of A11. With this lemma we can establish that almost every matrix can have a LU factorization. Definition 7.1.2. LetA∈Mn(C) and suppose that 1 ≤j≤n.T h e expression det( A{1,...,j }) means the determininant of the upper left j×j submatrix of A. These quaditities for j=1,...,n are called the principal determinants of A. Theorem 7.1.2. LetA∈Mn(C)and suppose that Ahas rank k.If det(A{1,...,j })6=0 forj=1,...,k (1) thenAhas a LU factorization A=LU,w h e r e Lis lower triangular and U is upper triangular. Moreover, the factorization may be taken so that either LorUis nonsingular. In the case k=nbothLandUwill be nonsingular. Proof. We carry out this LU factorization as a direct calculation in compar- ison to the Gaussian elimination method above. Let us propose to solve the equation LU=Aexpressed as  l 11 l21l22 0 l31l32l33 ............ ... ln1ln2··· ··· lnn  u 11u12u13··· u1n u22u23··· u2n u33 0...... ... unn  = a 11a12a13 a1n a21a22a23 a2n a31a32a33 ............... ... an1an2··· ··· ann  208 CHAPTER 7. FACTORIZATION THEOREMS It is easy to see that l11u11=a11.We can take, for example l11=1a n d solve for u11.The detminant condition assures us that u116=0.Next solve for the (2 ,1)-entry. We have l21u11=a21.Since u116=0,solve for l21. For the (1 ,2)-entry we have l11u12=a12,w h i c hc a nb es o l v e df o r u12since l116= 0. Finally, for the (2 ,2)-entry, l12u12+l22u22=a22is an equation with two unknowns. Assign l22= 1 and solve for u22.What is important to note is that the process carried out this way gives the factorization of theupper left 2 ×2 submatrix of A.Thus ·l 110 l21l22¸·u11u12 0u22¸ =·a11a12 a21a22¸ Since detµ·a11a12 a21a22¸¶ 6=0,it follows that detµ·u11u12 0u22¸¶ 6=0a n d we know that·l110 l21l22¸ is nonsingular as the diagonal elements are ones. Continue the factorization process through the k×kupper left submatrix ofA. Now consider the blocked matrix form form A A=·A11A12 A21A22¸ where A11isk×kand has rank k. Thus we know that the rows of the lower (n−k)×nmatrix above, that is£ A21A22¤ c a nb ew r i t t e na sau n i q u e linear combination of the rows of the upper k×nmatrix£ A11A12¤ .Thus £ A21A22¤ =C£ A11A12¤ for some ( n−k)×kmatrix C.Of course this means: A21=CA 11and A22=CA 12. We consider the factorization A=·A11A12 A21A22¸ =·L11 0 L21L22¸·U11U12 0U22¸ where the blocks L11andU11have just been determined. From the equations in the lemma above we solve to get U12=L−1 11A12andL21= 7.2. LRLRLRFACTORIZATION 209 A12U−1 11.T h e n A22=L21U12+L22U22 =A12U−1 11L−1 11A12+L22U22 =A12A−1 11A12+L22U22 =CA 11A−1 11A12+L22U22 =CA 12+L22U22 =A22+L22U22 Thus we solve L22U22=0.Obviously, we can take for L22any nonsingular matrix we wish and solve for U22or conversely. 7.2LRLRLRfactorization While the PLU factorization is useful for solving systems, the LR factoriza- tion can be used to determine eigenvalues. . LetA∈Mnbe given. Then A=A1=L1R1. Then L−1 1A1L1=R1L1≡A2 A2=L2R2 L−1 2A2L2=R2L2≡A3. Continue in this fashion to obtain L−1 kAkLk=RkLk≡Ak+1 (?) We de fine Pk=L1L2...L k Qk=Rk...R 2R1. Then PkAk+1=A1Pk 210 CHAPTER 7. FACTORIZATION THEOREMS for Ak+1=L−1 kAkLk =L−1 kL−1 k−1Ak−1Lk−1Lk ... =P−1 kA1Pk or PkAk+1=A1Pk. Hence PkQk=Pk−1AkQk−1 =A1Pk−1Qk−1 =A1Pk−2Ak−1Qk−2 =A2 1Pk−2Qk−2 ... =Ak 1. Theorem 7.2.1 (Rutishauser). LetA∈Mnbe given. Assume the eigen- values of Asatisfy |λ1|>|λ2|>···>|λn|>0. Then A∼Λ=diag(λ1...λn). Assume A=SΛS−1,a n d Y≡S−1=LyRy X=S=LxRx where LyandLxare lower unit triangular matrices and RyandRxare upper triangular. Then Akdefined by (?)satisfy the result limAkis upper triangular. Proof. (Wilkinson) We have Ak 1=XΛkY =XΛkLyRy =XΛkLyΛ−kΛkRy. 7.3. THE QRALGORITHM 211 By the strict inequalities between the eigenvalues we have (ΛkLyΛ−k)ij=  1 i=j µλi λj¶k `iji>j 0 i<j . HenceΛkLyΛ−k→I(because|λi| |λj|<1i fi>j ). Hence with Ak 1=LxRx(ΛkLyΛ−k)ΛkRy and Ak 1=PkQk we conclude that lim k→∞Pk=Lx. Therefore Lk=P−1 k−1Pk→I. Finally we have that Akmust be upper triangular because L−1 kAk=Rk is upper triangular. This exposes all the eigenvalues of A.T h e r e f o r e t h e e i g e n v e c t o r s c a n b e determined. 7.3 The QRQRQRalgorithm Certain numerical problems with the LUalgorithm have led to the QR algorithm, which is based on the decomposition of the matrix Aas A=QR where Qis unitary and Ris upper triangular. Theorem 7.3.1 (QR-factorization). (i) Suppose Ais inMn,mandn≥ m. Then there is a matrix Q∈Mn,mwith orthogonal columns and an u p p e rt r i a n g u l a rm a t r i x R∈Mmsuch that A=QR. 212 CHAPTER 7. FACTORIZATION THEOREMS (ii) If n=m,t h e n Qis unitary. If Ais nonsingular the diagonal entries ofRcan be chosen to be positive. (iii) If Ais real; then QandRm a yb ec h o s e nt ob er e a l . Proof. (i) We proceed inductively. Let a1,... , a ndenote the columns ofAandq1,q2,... ,q mdenote the columns of Q. The basic idea of the QR-factorization is to orthogonalize the columns of Afrom left to right. Then the columns can be expressed by the formulas ak=Pk i=1ckqk,k =1,...,n .T h e c o e fficients of the expansion become, respectively, the entries of the kthcolumn of R,c o m p l e t e db y n−k zeros. (Of course, if the rank of Ais less than m,w e fill in arbitrary orthogonal vectors which we know exist as m≤n.) For the details, first de fineq1=a1/ka1k. To compute q2we use the Gram—Schmidt procedure. ˆq2=a2−hq1,a1iq1 q2=ˆq2/kˆq2k. Tracing backwards note that a2=ˆq2+hq1,a1iq1 =kˆq2kq2+hq1,a1iq1. So we have ·a1a2a3 ↓↓↓ ...¸ =·q1q2q3 ↓↓↓ ...¸ ka 1khq1,a1i... 0 kˆq2k ...0 00 . Instead of the full inductive step we compute q 3andfinish at that point ˆq3=a3−hq1,a3iq1−hq2,a3iq2 q3=ˆq3/kˆq3k. Hence a3=kˆq3kq3+hq1,a3iq1+hq2,a3iq2. 7.3. THE QRALGORITHM 213 The third column of Ris thus given by r3=[hq1,a3i,hq2,a3i,kˆq3k,0,0,... , 0]T. In this way we see that the columns of Qare orthogonal and the matrix Ris upper triangular, with an exception. That is the possibility that ˆqk=0f o rs o m e k. In this degenerate case we take qkto be any vector orthogonal to the span of a1,a2,... ,a m,a n dw et a k e rkj=0 , j=k,k+1...m . A l s ow en o t et h a ti fˆ qk=0 ,t h e n akis linearly dependent on a1,a2,... ,a k−1, and hence on q1,q2,...q k−1. Select the coefficients r1k,... ,r k−1kto reflect this dependence. (ii) If m=n, the process above yields a unitary matrix. If Ais nonsingu- lar, the process above yields a matrix Rwith a positive diagonal. (iii) If Ais a real, all operators above can be carried out in real arithmetic. Now what about the uniqueness of the decomposition? Essentially the uniqueness is true up to a multiplication by a diagonal matrix, except inthe case when the matrix has rank is less than m, when there is no form of uniqueness. Suppose that the rank of Aism. Then application of the Gram-Schmidt procedure yields a matrix Rwith positive diagonal. Suppose that Ahas two QR factorizations, QRandPS with upper triangular factors having positive diagonals. Then P ∗Q=SR−1 We have that SR−1is upper triangular and moreover has a positive diagonal. Also, P∗Qis unitary. We know that the only upper triangular unitary matrices are diagonal matrices, and finally the only unitary matrix with a positive diagonal is the identity matrix. Therefore P∗Q=I,w h i c hi st o say that P=Q.We summarize as Corollary 7.3.1. Suppose Ais in Mn,mandn≥m.I f r a n k (A)=m then the QR factorization of A=QRwith upper triangular matrix Rhaving a positive diagonal is unique. 214 CHAPTER 7. FACTORIZATION THEOREMS TheQRalgorithm TheQRalgorithm parallels the LRalgorithm almost identically. Suppose Ais inMnDefine A1=Q1R1 A2≡R1Q1. Also Q∗ 1A1Q=A2. Then decompose A2into a QRdecomposition A2=Q2R2 and Q∗ 2A2Q2=R2Q2≡A3. Also Q∗ 2Q∗1A1Q1Q2=R2Q2=A3. Proceed sequentially Ak=QkRk Ak+1=RkQk Q∗ kAkQk=Ak+1. Let Pk=Q1Q2...Q k Tk=RkRk−1...R 1. Then P∗ kA1Pk=Ak+1. whence PkAk+1=A1Pk. 7.3. THE QRALGORITHM 215 Also we have PkTk=Pk−1QkRkTk−1 =Pk−1AkTk−1 =A1Pk−1Tk−1 =... =Ak 1. Theorem 7.3.2. LetA∈Mnbe given, and assume the eigenvalues of A satisfy |λ1|>|λ2|>···>|λn|>0. Then the iterations Akconverge to a triangular matrix. Proof. Our hypothesis gives that Ais diagonalizable, and we write A∼Λ= diag(λ1...λn). That is, A1=SΛS−1 whereΛ= diag(λ1...λn). Let X=S=QxRx hereQR Y=S−1=LyUyhereLU. Then Ak 1=QxRxΛkLyUy =QxRxΛkLyΛ−kΛkUy =Qx(I+RxEkR−1 x)RxΛkUY where Ek=ΛkLyΛ−k−I (Ek)ij=  0 i=j (λi/λj)k`iji>j 0 i<j . It follows that I+RxEkR−1 x→I,a n d RxΛ−kUyis upper triangular. Thus Qx(I+RxEkR−1 x)RxΛkUy=PkTk. 216 CHAPTER 7. FACTORIZATION THEOREMS The matrix I+RxEkR−1 xcan be QR factored as ˜Uk˜Rk, and since I+ RxEkR−1 x→I, it follows that we can assume both ˜Uk→Iand ˜Rk→I. Hence Ak 1=Qx˜Uk[˜Rk(I+RxEkR−1 x)RxΛkUy]=PkTk. with the first factor unitary and the second factor upper triangular. Since we have assumed (by the eigenvalue condition) that Ais nonsingular, this factorization is essentially unique, where possibly a multiplication by a di- agonal matrix must be applied to give the upper triangular factor on theright a positive diagonal. Just what is the form of the diagonal matrix canbe seen from the following. Let Λ=|Λ|Λ 1,w h e r e |Λ|is the diagonal matrix of moduli of the elements of Λand where Λ1is the unitary matrix of the signs of each eigenvalue respectively. We also take Uy=Λ2(Λ∗ 2Uy)w h e r e Λ2is a unitary matrix chosen so that Λ∗ 2Uhas a positive diagonal. Then Ak 1=Qx˜UkΛ2Λk 1[³ Λ2Λk 1´−1˜Rk(I+RxEkR−1 x)Rx³ Λ2Λk 1´ |Λ|k(Λ∗ 2Uy)] =PkTk. From this we obtain Pkis essentially asymptotic to Qx˜UkΛ2Λk 1and from this we obtain that Qk=P−1 k−1Pk→Λ1 which is diagonal. Finally, it follows that Akis upper triangular since Q−1 kAk=Rk In the limit therefore Ais similar to an upper triangular matrix. Example 7.3.1. Apply the QR method to the matrix A:= 2.31 2 22 2 .1 320  The matrix Ahas eigvenvalues 5 .45,0.723,−1.87. The successive iterations are 7.4. LEAST SQUARES 217 A2= 5.10−0.511 2 .13 0.631 0 .662 0 .136 1.42−0.0202−1.44  A3= 5.51−1.02−0.36 −0.0146 0 .666 0 .482 0.513 0 .240−1.84  A4= 5.46−1.41 0 .482 −0.0372 0 .495 0 .672 0.169 0 .815−1.62  A5= 5.47−0.366−1.26 −0.0404−0.462 1 .39 0.0430 1 .21−0.677  A6= 5.46−1.13−0.687 −0.0184−1.52 0 .813 0.00826 0 .983 0 .381  A7= 5.45 0 .529−1.18 −0.00682−1.78 0 .585 0.00115 0 .414 0 .638  A8= 5.43 0 .684−1.09 −0.000822 −1.87 0 .229 0.0000215 0 .0659 0 .729  Note the gradual appearance of the eigenvalues on the diagonal. Remark. These iterations were carried out in precision 3 arithmetic, whichaffects the rate of convergence to triangular form. 7.4 Least Squares As we know, if A∈Mn,mwith m<n it is generally not possible to solve the overdetermined system Ax=b. For example, suppose we have the data {(xi,yi)}n i=1,w i t ht h e x-coordinates distinct. We may wish to “ fit” a straight to this data. This means we want tofind coefficients mandbso that b+mxi=yi,i =1,... ,n . (?) Taking the matrix and data vector A= 1x 1 1x2 ... 1xn b= y 1 y2 ... yn  andz=[b, m]T, the system ( ?) becomes Az=b.U s u a l l y nÀ2. Hence there is virtually no hope to determine a unique solution to system. However, there are numerous ways to determine constants mandbso that the resulting line represents the data. For example, owing to the dis- tinctness of the x-coordinates, it is possible to solve any 2 ×2 subsystem of 218 CHAPTER 7. FACTORIZATION THEOREMS Az=b. Other variations exist. A new 2 ×2 system could be created by creating two averages of the data, say left and right, and solving. Assume the sequence {xj}is ordered from least to greatest. De finex`=1 kkP j=1xjand xr=1 n−knP j=k+1xj.L e t y`andyrdenote the corresponding averages for the ordinates. Then de fine the intercept band slope mby solving the system ·1x` 1xr¸·b m¸ =·y` yr¸ While this will normally give a reasonable approximating line, its value has little utility beyond its naive simplicity and visual appearance. What is desired is to establish a criteria for choosing the line. Define the residual of the approximation r=b−Az.I tm a k e sp e r f e c t sense to consider finding z=[b, m]Tfor which the residual is minimized in some norm. Any norm can be selected here, but on practical grounds thebest norm to use is the Euclidean norm k·k 2.T h e v e c t o r Azthat yields the minimal norm residual is the one for which ( b−Az)⊥Aw, for we are seeking the nearest value in the Awto the vector b. It can be found by select the one for the solution, Az,f o rw h i c h b−Ax⊥Aw allw. This means hb−Ax, Ay i=0 a l l y or hAT(b−Ay),yi=0 a l l y or AT(b−Ay)=0 ATAy=ATb.Normal Equations Theleast squares solution to Ax=bis given by the solution to the normal equation ATAy=ATb. 7.5. EXERCISES 219 Suppose we have the QRdecomposition for A.T h e ni f Ais real ATA=RTQTQR=RTR ATy=RTQy. Hence the normal equations become RTRx=RTQy. Assuming that the rank of Aism,w em u s th a v et h a t Rand hence RT is invertible. Therefore we have the least squares solution is given by the triangular system Rx=Qy. 7.5 Exercises 1. If A∈M(C)h a sr a n k k, show that there is a permutation matrix P such that PAhas its firstkprincipal determinants nonzero. 2. For the least squares fit of a straight line determine RandQ. 3. In the case of data ATA=·nΣxi ΣxiΣx2 i¸ ATb=·Σyi Σxiyi¸ . 4. In attempting to solve a quadratic fitw eh a v et h em o d e l c+bxi+ax2 i=yi i=1,... ,n . The system is A= 1x1x2 1......... 1xnx2 n  b= y1 y2 ... yn . The normal equations have the matrix and data given by ATA= nΣxiΣx2 i ΣxiΣx2 iΣx3 i Σx2 iΣx2 iΣx4 i  ATb= Σyi Σxiyi Σx2 iyi . 5. Find the normal equations for the least squares fito fd a t at oap o l y - nomial of degree k. Chapter 8 Jordan Normal Form 8.1 Minimal Polynomials Recall pA(x)=d e t ( xI−A) is called the characteristic polynomial of the matrix A. Theorem 8.1.1. LetA∈Mn. Then there exists a unique monic polyno- mial qA(x)of minimum degree for which qA(A)=0 .I fp(x)is any polyno- mial such that p(A)=0 ,t h e n qA(x)divides p(x). Proof. Since there is a polynomial pA(x)f o rw h i c h pA(A) = 0, there is one of minimal degree, which we can assume is monic. by the Euclidean algorithm pA(x)=qA(x)h(x)+r(x) where deg r(x)<degqA(x). We know pA(A)=qA(A)h(A)+r(A). Hence r(A) = 0, and by the minimality assumption r(x)≡0. Thus qA divided pA(x) and also any polynomial for which p(A) = 0. to establish that qAis unique, suppose q(x) is another monic polynomial of the same degree for which q(A)=0 . T h e n r(x)=q(x)−qA(x) is a polynomial of degree less than qA(x)f o rw h i c h r(a)=q(A)−qA(A)=0 . This cannot be unless r(x)≡0=0 q=qA. Definition 8.1.1. The polynomial qA(x) in the theorem above is called the minimal polynomial. 221 222 CHAPTER 8. JORDAN NORMAL FORM Corollary 8.1.1. IfA, B∈Mnare similar, then they have the same min- imal polynomial. Proof. B=S−1AS qA(B)=qA(S−1AS)=S−1qA(A)S=qA(A)=0. If there is a minimal polynomial for Bof smaller degree, say qB(x), then qB(A) = 0 by the same argument. This contradicts the minimality of qA(x). Now that we have a minimum polynomial for any matrix, can we find a matrix with a given polynomial as its minimum polynomial? Can the degreethe polynomial and the size of the matrix match? The answers to both questions are a ffirmative and presented below in one theorem. Theorem 8.1.2. For any n thdegree polynomial p(x)=xn+an−1xn−1+an−2xn−2+···+a1x+a0 there is a matrix A∈Mn(C)for which it is the minimal polynomial. Proof. Consider the matrix given by A= 00 ... ... −a 0 10 −a1 01 0...−a2 ... 0... 0... 01−an−1 . Observe that Ie 1=e1=A0e1 Ae1=e2=Ae1 Ae2=e3=A2e1 ... Aen−1=en=An−1e1 8.1. MINIMAL POLYNOMIALS 223 and Aen=−an−1en−an−2en−1−···−a1e2−a0e1 Since Aen=Ane1, it follows that p(A)e1=Ane1+an−1An−1e1+an−2A−2e1+···+a1Ae1+a0Ie1=0 Also p(A)ek=p(A)Ak−1e1=Ak−1p(A)e1=Ak−1(0) = 0 k=2,... ,n . Hence p(A)ej=0 f o r j=1...n.T h u s p(A) = 0. We know also that p(x) is monic. Suppose now that q(x)=xm+bm−1xm−1···+b1x+b0 where m<n andq(A)=0 . T h e n q(A)e1=Ame1+bm−1Am−1e1+···+b1Ae1+b0e1 =em+1+bm−1em+···+b1e2+b0e1=0. But the vectors em+1...e 1are linear independent from which we conclude thatq(A) = 0 is impossible. Thus p(x)i sm i n i m a l . Definition 8.1.2. For a given monic polynomial p(x), the matrix Acon- structed above is called the companion matrix to p. The transpose of the companion matrix can also be used to generate a linear differential system which has the same characteristic polynomial as a given nthorder differential equation. Consider the linear di fferential equation y(n)+an−1y(n−1)+···+a1y0+a0=0. This nthorder ODE can be converted to a first order system as follows: u1=y u2=u0 1 =y0 u3=u0 2 =y00 ...... un=u0 n−1=y(n−1) 224 CHAPTER 8. JORDAN NORMAL FORM Then we have  u 1 u2 ... un 0 = 01 01 0 01 ...... 0... 1 a 0−a1 ... ... −an−1  u 1 u2 ... un  8.2 Invariant subspaces There seems to be no truly simple way to the Jordan normal form. The approach taken here is intended to reveal a number of features of a matrix,interesting in their own right. In particular, we will construct “generalized eigenspaces” that envelop the entire connection of a matrix with its eigen- values. We have in various ways considered subspaces VofC nthat are invariant under the matrix A∈Mn(C). Recall this means that AV⊂V. For example, eigenvectors can be used to create invariant subspaces. Nullspaces, the eigenspace of the zero eigenvalue, are invariant as well. Triangu-lar matrices furnish an easily recognizable sequence of invariant subspaces. Assuming T∈M n(C) is upper triangular, it is easy to see that the sub- spaces generated by the coordinate vectors {e1,...,e m}form=1,...,n are invariant under T. We now consider a speci fic type of invariant subspace that will lead the so-called Jordan normal form of a ma trix, the closest matrix similar to A that resembles a diagonal matrix. Definition 8.2.1 (Generalized Eigenspace). LetA∈Mn(C)w i t hs p e c - trumσ(A)={λ1,...,λk}.D e fine the generalized eigenspace pertaining to λiby Vλi={x∈Cn|(A−λiI)nx=0} Observe that all the eigenvectors pertaining to λiare contained in Vλi. If the span of the eigenvectors pertaining to λiis not equal to Vλithen there must be a positive power pand a vector xsuch that ( A−λiI)px= 0 but that y=(A−λiI)p−1x6=0 . T h u s yis an eigenvector pertaining to λi.F o rt h i s reason we will call Vλithe space of generalized eigenvectors pertaining to λi.O u r first result, that Vλiis invariant under A,i ss i m p l et op r o v e ,n o t i n g that only closure under vector addition and scalar multiplication need be established. 8.2. INVARIANT SUBSPACES 225 Theorem 8.2.1. LetA∈Mn(C)with spectrum σ(A)= {λ1,...,λk}. Then for each i=1,...,k ,Vλiis an invariant subspace of A. One might question as to whether Vλicould be enlarged by allowing higher powers than nin the de finition. The negative answer is most simply expressed by evoking the Hamilton-Cayley theorem. We write the charac- teristic polynomial pA(λ)=Q(λ−λi)mA(λi),w h e r e mA(λi) is the algebraic multiplicity of λi.S i n c e pA(A)=Q(A−λiI)mA(λi)= 0, it is an easy mat- ter to see that we exhaust all of Cnwith the spaces Vλi.T h i s i s t o s a y t h a t allowing higher powers in the de finition will not increase the subspaces Vλi. Indeed, as we shall see, the power of ( λ−λi) can be decreased to the geomet- ric multiplicity mg(λi)t h ep o w e rf o r λi. For now the general power nwill suffice. One very important result, and an essential fir s ts t e pi nd e r i v i n g the Jordan form, is to establish that any square matrix Ais similar to a block diagonal matrix, with each block carrying a single eigenvalue. Theorem 8.2.2. LetA∈Mn(C)with spectrum σ(A)= {λ1,...,λk}and with invariant subspaces Vλi,i=1,2,...,k .T h e n ( i ) T h e s p a c e s Vλi,j= 1,...,k are mutually linearly independent. (ii)Lk i=1Vλi=Cn(alternatively Cn=S(Vλ1,...,V λk)) (iii) dimVλi=mA(λi).( i v ) Ais similar to a block diagonal matrix with kblocks A1,...,A k.M o r e o v e r , σ(Ai)= {λi}and dimAi=mA(λi). Proof. (i) It should be clear that the subspaces Vλiare linearly independent of each other. For if there is a vector xin both VλiandVλjthen there is a vector for some integer q,it must be true that ( A−λjI)q−1x6= 0 but (A−λjI)qx=0.This means that y=(A−λjI)q−1xis an eigenvector pertaining to λj.S i n c e ( A−λiI)nx= 0 we must also have that (A−λjI)q−1(A−λiI)nx=(A−λiI)n(A−λjI)q−1x =(A−λiI)ny=0 =nX k=0µn k¶ (−λi)n−kAky =nX k=0µn k¶ (−λi)n−kλjky =(λj−λi)ny=0 This is impossible unless λj=λi. (ii) The key part of the proof is to block diagonalize Awith respect to these invariant subspaces. To that end, let S 226 CHAPTER 8. JORDAN NORMAL FORM be the matrix with columns generated from bases of the individual Vλitaken in the order of the indices. Supposing there are more linearly independentvectors in C nother than those already selected, fill out the matrix Swith vectors linearly independent to the subspaces Vλi,i=1,...,k .N o w d e fine ˜A=S−1AS. We conclude by the invariance of the subspaces and their mutual linear independence that ˜Ahas the following block structure. ˜A=S−1AS= A 10 ··· 0∗ 0 A2 0∗ ......... Ak∗ 0 ··· 0 B  It follows that p A(λ)=p˜A(λ)=³Y pAi(λ)´ pB(λ) Any root rofpB(λ) must be an eigenvalue of A,sayλj,and there must be an eigenvector xpertaining to λj. Moreover, due to the block structure we can assume that x=[ 0,..., 0,x]T,where there are kblocked zeros of the sizes of theAirespectively. Then it is easy to see that ASx =λjSx, and this implies thatSx∈Vλj. Thus there is another vector in Vλj, which contradicts its definition. Therefore Bis null, or what is the same thing, ⊕k i=1Vλi=Cn. (iii) Let di=d i m Vλi.F r o m ( i i ) k n o wPdi=n. Suppose that λi∈σ(Aj). Then there is another eigenvector xpertaining to λiand for which Ajx= λix.Moreover, this vector has the form x=[ 0,..., 0,x ,0,...0]T, analogous to the argument above. By construction Sx /∈Vλi, but ASx =λiSx,and this contradicts the de finition of Vλi. W et h u sh a v et h a t pAi(λ)=(λ−λi)di. Since pA(λ)=QpAi(λ)=Q(λ−λi)mA(λi), it follows that di=mA(λi) (iv) Putting (ii), and (iii) together gives the block diagonal structure as required. On account of the mutual linear independence of the invariant subspacesV λiand the fact that they exhaust Cnthe following corollary is immediate. Corollary 8.2.1. LetA∈Mn(C)with spectrum σ(A)=λ1,...,λkand with generalized eigenspaces Vλi,i=1,2,...,k .T h e n e a c h x∈Cnhas a unique representation x=Pk i=1xiwhere xi∈Vλi. Another interesting result which reveals how the matrix works as a linear transformation is to decompose the it into components with respect to the 8.2. INVARIANT SUBSPACES 227 generalized eigenspaces. In particular, viewing the block diagonal form ˜A=S−1AS= A10 ··· 0 0 A2 0 ...... Ak  the space Cncan be split into a direct sum of subspaces E1,...,E kbased on coordinate blocks. This is accomplished in such that any vector y∈Cncan be written uniquely as y=Pk i=1yiwhere the yi∈Ei. (Keep in mind that eachyi∈Cn; its coordinates are zero outside the coordinate block pertaining Ei.) Then ˜Ay=˜APk i=1yi=Pk i=1˜Ayi=Pk i=1Aiyi.This provides a computational tool – when this block diagonal form is known. Note that the blocks correspond directly to the invariant subspaces by SEi=Vλi.W e can use these invariant subspaces to get at the minimal polynomial. Foreach i=1,...,k define m i=m i n j{(A−λiI)jx=0 |x∈Vλi} Theorem 8.2.3. LetA∈Mn(C)with spectrum σ(A)=λ1,...,λkand with invariant subspaces Vλi,i=1,2,...,k . Then the minimal polynomial ofAis given by q(λ)=kY i=1(λ−λi)mi Proof. Certainly we see that for any vector x∈Vλj q(A)x=ÃkY i=1(A−λiI)mi! x=0 Hence, the minimal polynomial qA(λ)d i v i d e s q(A). To see that indeed they are in fact equal, suppose that the minimal polynomial has the form qA(λ)=kY i=1(λ−λi)ˆmi where ˆ mi≤mi,fori=1,...,k a n di np a r t i c u l a r ˆ mj<m j.B y c o n s t r u c t i o n there must exist a vector x∈Vλjsuch that ( A−λjI)mjx= 0 but y= 228 CHAPTER 8. JORDAN NORMAL FORM (A−λjI)mj−1x6=0.Then if q(A)x=ÃkY i=1(A−λiI)ˆmi! x = kY i=1 i6=j(A−λiI)ˆmi y =0 This cannot be because the contrary implies that there is another vector in one of the invariant subspaces Vλk. Just one more step is needed before the Jordan normal form can be derived. For a given Vλiwe can interpret the spaces in a heirarchical viewpoint. We know that Vλicontains all the eigenvectors pertaining to λi.C a l l t h e s e eigenvectors the first order generalized eigenvectors . If the span of these is not equal to Vλi, then there must be a vector x∈Vλifor which y= (A−λiI)2x= 0 but ( A−λiI)x6=0 . T h a ti st os a y yis an eigenvector of Apertaining to λi. Call such vectors second order generalized eigenvectors . In general we call an x∈Vλia generalized eigenvector of order jify= (A−λiI)jx= 0 but ( A−λiI)j−1x6= 0. In light of our previous discussion Vλicontains generalized eigenvectors of order up to but not greater than mλi. Theorem 8.2.4. LetA∈Mn(C)with spectrum σ(A)= {λ1,...,λk}and with invariant subspaces Vλi,i=1,2,...,k . (i) Let x∈Vλibe a generalized eigenvector of order p. Then the vectors x,(A−λiI)x,(A−λiI)2x ,..., (A−λiI)p−1x (1) are linearly independent. (ii) The subspace of Cngenerated by the vectors in (1) is an invariant subspace of A. Proof. (i) To prove linear independence of a set of vectors we suppose linear dependence. That is there is a smallest integer kand constants bjsuch that kX j=0xj=kX j=0bj(A−λiI)jx=0 8.2. INVARIANT SUBSPACES 229 where bk6=0.Solving we obtain bk(A−λiI)kx=−Pk−1 j=0bj(A−λiI)jx. Now apply ( A−λiI)p−kto both sides and obtain a new linearly dependent set as the following calculation shows. 0= bk(A−λiI)p−k+kx=−k−1X j=0bj(A−λiI)j+p−kx =−p−1X j=p−kbj+p−k(A−λiI)jx T h ek e yp o i n tt on o t eh e r ei st h a tt h el o w e rl i m i to ft h es u mi si n c r e a s e d . This new linearly dependent set, which we denote with the notationPp−1 j=p−kcj(A−λiI)jx can be split in the same way as before, where we assume with no loss in gen-erality that c p−16=0 . T h e n cp−1(A−λiI)p−1x=−p−2X j=p−kcj(A−λiI)jx Apply ( A−λiI) to both sides to get 0= cp−1(A−λiI)px=−p−2X j=p−kcj(A−λiI)j+1x =−p−1X j=p−k+1cj−1(A−λiI)jx Thus we have obtained another linearly independent set with the lower limit of powers increased by one. Continue this process until the linear depen-dence of ( A−λ iI)p−1xand ( A−λiI)p−2xis achieved. Thus we have c(A−λiI)p−1x=d(A−λiI)p−2x (A−λiI)y=d cy where y=(A−λiI)p−2x.T h i s i m p l i e s t h a t λi+d cis a new eigenvalue with eigenvector y∈Vλi, and of course this is a contradiction. (ii) The invariance under Ais more straightforward. First note that while x1=x, 230 CHAPTER 8. JORDAN NORMAL FORM x2=(A−λI)x=Ax−λxso that Ax=x2−λx1.Consider any vector y defined by y=Pp−1 j=0bj(A−λiI)jxIt follows that Ay =Ap−1X j=0bj(A−λiI)jx =p−1X j=0bj(A−λiI)jAx =p−1X j=0bj(A−λiI)j(x2−λx1) =p−1X j=0bj(A−λiI)j[(A−λiI)x1−λx1] =p−1X j=0cj(A−λiI)jx where cj=bj−1−λforj>0a n d c0=−λ, which proves the result. 8.3 The Jordan Normal Form We need a lemma that points in the direction we are headed, that being the use of invariant subspaces as a basis for the (Jordan) block diagonalization of any matrix. These results were discussed in detail in the Section 8.2. Therestatement here illustrates the “invariant subspace” nature of the result,irrespective of generalized eigenspac es. Its proof is elementary and is left to the reader. Lemma 8.3.1. LetA∈M n(C)with invariant subspace V⊂Cn. (i) Suppose v1...v kis a basis for VandSis an invertible matrix with thefirstkcolumns given by v1...v k.T h e n 1...k S−1AS=·∗∗ 0∗¸ . 8.3. THE JORDAN NORMAL FORM 231 (ii) Suppose that V1,V2⊂Cnare two invariant subspaces of AandCn= V1⊕V2. Let the (invertible) matrix Sconsist respectively of bases from V1 andV2as its columns. Then S−1AS=·∗0 0∗¸ . Definition 8.3.1. Letλ∈C.A Jordan block Jk(λ)i sa k×kupper triangular matrix of the form Jk(λ)= λ1 0 λ1 0...1 λ . AJordan matrix is any matrix of the form J= J n1(λ1)0 ... 0 Jnk(λk) . where the matrices Jn1are Jordan blocks. If J∈Mn(C), then n1+n2···+ nk=n. Theorem 8.3.1 (Jordan normal form). LetA∈Mn(C). Then there is a nonsingular matrix S∈Mnsuch that A=S J n1(λ1) 0 ... 0 Jnk(λk) S −1=SJS−1 where Jni(λi)is a Jordan block, where n1+n2+···+nk=n.Jis unique up to permutations of the blocks. The eigenvalues λ1,... ,λkare not necessarily distinct. If Ais real with real eigenvalues, then Scan be taken as real. Proof. This result is proved in four steps. (1) Block diagonalize (by similarity) into invariant subspaces pertaining to σ(A). This is accomplished as follows. First block diagonalize the ma- trix according to the generalized eigenspaces Vλi={x∈Cn|(A−λiI)nx= 232 CHAPTER 8. JORDAN NORMAL FORM 0}as discussed in the previous section. Beginning with the highest or- der eigenvector in x∈Vλi, construct the invariant subspace as in (1) of Section 8.2. Repeat this process until all generalized eigenvectorshave been included in an invariant subspace. This includes of coursefirst order eigenvectors that are not a ssociated with higher order eigen- vectors. These invariant subspaces have dimension one. Each of these invariant subspaces is linearly independent from the others. Continuethis process for all the generalized eigenspaces. This exhausts C n. Each of the blocks contains exactly one eigenvector. The dimensionsof these invariant subspaces can range from one to m λi,t h e r eb e i n ga t least one subspace of dimension mλi. (2) Triangularize each block by Schur’s theorem, so that each block has the form K(λ)= λ∗ ... 0λ  You will note that K(λ)=λI+N where Nis nilpotent, or K(λ)i s1 ×1. (3) “Jordanize” each triangular block. Assume that K1(λ)i sm×m,w h e r e m> 1. By construction K1(λ) pertains to an invariant subspace for which there is a unique vector xfor which Nm−1x6=0 a n d Nmx=0. Thus Nm−1xis an eigenvector of K1(λ), the unique eigenvector. De fine yi=Ni−1xi =1,2,... ,m . Expand the set {yi}m i=1as a basis of Cm.D efine S1=" ymym−1···y1 ...... ···...# . Then NS 1=" 0ymym−1... y 2 ......... ···...# . 8.3. THE JORDAN NORMAL FORM 233 So S−1 1NS 1= 01 0 01 0... ...1 00 . We conclude that S −1 1K1(λ)S1= λ10 λ1 λ1 ...... ...1 0 λ  (4) Assemble all of the blocks to form the Jordan form. For example, the block K 1(λ) and the corresponding similarity transformation S1 studied above can be treated in the assembly process as follows: De fine then×nmatrix ˆS1= I00 0S10 00 I  where S1i st h eb l o c kc o n s t r u c t e da b o v ea n dp l a c e di nt h e n×nmatrix in the position that K1(λ) was extracted from the block triangular form of A. Repeat this for each of the blocks pertaining to minimally invariant subspaces. This gives a sequence of block diagonal matrices ˆS1,ˆS2..., ˆSk.D efineT=ˆS1ˆS2...ˆSk. It has the form T= ˆS 1 0 ˆS2 ... 0 ˆSk  Together with the original matrix Pthat transformed the matrix to the minimal invariant subspace blocked form and the unitary matrix 234 CHAPTER 8. JORDAN NORMAL FORM Vused to triangularize A, it follows that A=PVT J n1(λ1) 0 ... 0 Jnk(λk) (PVT ) −1=SJS−1 with S=PVT . Example 8.3.1. Let J= 2 1 02 2 310 031 003 −1  In this example, there are four blocks, with two of the blocks pertaining to the single eigenvalue 2. For the first block there is the single eigenvector e 1, but the invariant subspace is S(e1,e2).For the second block, the eigenvector, e3,generates the one dimensional invariant subspace. The block pertaining to the eigenvector 3 has the single eigenvector e4while the minimal invariant subspace is S(e4,e5,e6). Finally, the one dimensional subspace pertaining to the eigenvector −1 is spanned by e7.The minimal invariant polynomial isq(λ)=(λ−2)2(λ−3)3(λ+1 ) . 8.4 Convergent matrices Using the Jordan normal form, the study of convergent matrices becomes relatively straightforward and simpler. Theorem 8.4.1. IfA∈Mnandρ(A)<1.T h e n lim k→∞A=0. 8.5. EXERCISES 235 Proof. We assume Ais a Jordan matrix. Each Jordan block Jk(λ)c a nb e written as Jk(λ)=λIk+Nk where Nk= 01 0 01 ...... ...1 00 is nilpotent. Now A= J n1(λ1) ... Jmk(λk) . We compute, for m>n k (Jnk(λk))M=(λI+N)m =λmI+mkX j=0λm−jNjµm j¶ because Nj=0f o r j>n k.W eh a v e λm−jµm j¶ →0a sm→∞ since |λ|<1. The results follows. 8.5 Exercises 1. Prove Theorem 8.2.1. 2. Find a 3 ×3 matrix that has the same eigenvalues are the squares of the roots of the equation λ3−3λ2+4λ−5=0 . 3. Suppose that Ais a square matrix with σ(A)= {3},ma(3) = 6, and mg(3) = 3. Up to permutations of the blocks show all possible Jordan normal forms for A. 236 CHAPTER 8. JORDAN NORMAL FORM 4. Let A∈Mn(C)a n dl e t x1∈Cn.Definexi+1=Axifori=1,...,n−1. Show that V=S({x1,...,x n}) is an invariant subspace of A. Show thatVcontains an eigenvector of A. 5. Referring to the previous problem, let A∈Mn(R)b eap e r m u t a t i o n matrix. (i) Find starting vectors so that dim V=n.( i i ) F i n d s t a r t - ing vectors so that dim V= 1. (iii) Show that if λ= 1 is a simple eigenvalue of Athen dim V=1o rd i m V=n. Chapter 9 Hermitian and Symmetric Matrices Example 9.0.1. Letf:D→R,D⊂Rn.T h e Hessian is defined by H(x)=hij(x)≡∂f ∂xi∂xj∈Mn. Since for functions f∈C2it is known that ∂2f ∂xi∂xj=∂2f ∂xj∂xi it follows that H(x)i ss y m m e t r i c . Definition 9.0.1. A function f:R→Risconvex if f(λx+( 1−λ)y)≤λf(x)+( 1−λ)f(y) forx, y∈D(domain) and 0 ≤λ≤1. Proposition 9.0.1. Iff∈C2(D)andf00(x)≥0onDthenf(x)is convex. Proof. Because f00≥0, this implies that f0(x) is increasing. Therefore if x<x m<ywe must have f(xm)≤f(x)+f0(xm)(xm−x) and f(xm)≤f(y)+f0(xm)(xm−y) 237 Chapter 9 Hermitian and Symmetric Matrices Example 9.0.1. Letf:D→R,D⊂Rn.T h e Hessian is defined by H(x)=hij(x)≡∂f ∂xi∂xj∈Mn. Since for functions f∈C2it is known that ∂2f ∂xi∂xj=∂2f ∂xj∂xi it follows that H(x)i ss y m m e t r i c . Definition 9.0.1. A function f:R→Risconvex if f(λx+( 1−λ)y)≤λf(x)+( 1−λ)f(y) forx, y∈D(domain) and 0 ≤λ≤1. Proposition 9.0.1. Iff∈C2(D)andf00(x)≥0onDthenf(x)is convex. Proof. Because f00≥0, this implies that f0(x) is increasing. Therefore if x<x m<ywe must have f(xm)≤f(x)+f0(xm)(xm−x) and f(xm)≤f(y)+f0(xm)(xm−y) 237 238 CHAPTER 9. HERMITIAN AND SYMMETRIC MATRICES by the Mean Value Theorem. Therefore, y−xm y−xf(xm)+xm−x y−xf(xm)≤y−xm y−xf(x)+xm−x y−xf(y) or f(xm)≤y−xm y−xf(x)+(xm−x) y−xf(y). It is easy to see that xm=y−xm y−xx+xm−x y−xy. Definingλ=y−xm y−x, it follows that 1 −λ=xm−x y−x. Definition 9.0.2. We say that f:D→R,w h e r e D⊂Rnis convex, is a convex function if fis a convex function on every line in D. Theorem 9.0.1. Suppose f∈C2(D)andH(x)is positive de finite. Then fis convex on D. Proof. Letx∈Dandηbe some direction. Then x+ληis a line in D.W e computed2 dλ2f(x+λn) d dλf=∇f·n= ∂f ∂x1... ∂f ∂xn ·[n1...n n]. Now d dλµ∂f ∂x1¶ =∇∂f ∂x1·η= ∂2f ∂x1∂x1 ∂2f ∂xn∂x1 ·[η1,... ,ηn]. Hence we see that d2 dλ2f(x+λη)=ηTH(x)η≥0 by assumption. 239 Example 9.0.2. LetA=[aij]∈Mn. Consider the quadratic form on Cn orRndefined by Q(x)=xTAx=Σaijxjxi =1 2Σ(aij+aji)xjxi =xT1 2(A+AT)x. Since the matrix A+ATis symmetric the study of quadratic forms is reduced to the symmetric case. Example 9.0.3. LetLf=nP i,j=1aij∂2f ∂xi∂xj.Lis called a partial di fferential operator. By the combination of devices above (assuming f∈C2for exam- ple) we can study the symmetric and equivalent partial di fferential operator Lf=nX i,j=11 2(aij+aji)∂2f ∂xi∂xj. In particular, if A+ATis positive de finite the operator is called elliptic. Other cases are (1) hyperbolic,(2) degenerate/parabolic. Characterizations of Her mitian matrices. Recall (1)A∈M nis Hermitian if A∗=A. (2)A∈Mnis called skew-Hermitian if A=−A∗. Here are some facts (a) If Ais Hermitian the diagonal is real. (b) If Ais skew-Hermitian the diagonal is imaginary. (c)A+A∗,A A∗andA∗Aare all Hermitian if A∈Mn. (d) If Ais Hermitian than Ak,k=0,1,... , are Hermitian. A−1is Her- mitian if Ais invertible. 240 CHAPTER 9. HERMITIAN AND SYMMETRIC MATRICES (e)A−A∗is skew-Hermitian. (f)A∈Mnyields the decomposition A=1 2(A+A∗)+1 2(A−A∗) Hermitian Skew Hermitian (g) If Ais Hermitian iAis skew-Hermitian. If Ais skew-Hermitian then iAis Hermitian. Theorem 9.0.2. LetA∈Mn.T h e n A=S+iTwhere SandTare Hermitian. Moreover this is unique. Proof. A=1 2(A+A∗)+1 2(A−A∗) =S+iT where S=1 2(A+A∗)a n d iT=1 2(A−A∗) ⇒T=−i 2(A−A∗). Theorem 9.0.3. LetA∈Mnbe Hermitian. Then (a)x∗Axis real for all x∈Cn; (b) All the eigenvalues of Aare real; (c)S∗ASis Hermitian for all S∈Mn. Proof. For (a) we have x∗Ax=Σaijxjxi. The conjugate is x∗Ax=Σ¯aij¯xjxi=Σaji¯xjxi =Σaijxj¯xi+x∗Ax Theorem 9.0.4. LetA∈Mn.T h e n Ais Hermitian if and only if at least one of the following holds: 9.1. VARIATIONAL CHARACTERIZATIONS OF EIGENVALUES 241 (a)hAx, x i=x∗Axis real∀x∈Cn. (b)Ais normal with real eigenvalues. (c)S∗ASis Hermitian for all S∈Mn. Proof. Hermitian ⇒(a), (b), or (c) are obvious. (a) ⇒Hermitian. Prove using ejandek.W eh a v e hAej+ek,ej+eki=ajj+akk+ajk+akj ⇒ajk+akjis real⇒Imajk=−Imakj. Similarly hAiej+ek,i ej+eki−ajj+akk+iakj−iajk ⇒i(akj−ajk)i sr e a l⇒Reakj=R e ajk. All this gives A∗=A. (c)⇒Hermitian S∗ASis Hermitian ⇒S∗ASis similar to a real diagonal matrix. Therefore Ais similar to a real diagonal matrix. Just let S=Ito getAis Hermitian. Theorem 9.0.5 (Spectral Theorem). LetA∈Mnbe Hermitian. Then Ais unitarily (similar) equivalent to a real diagonal matrix. If Ais real Hermitian, then Ais orthogonally similar to a real diagonal matrix. 9.1 Variational Characterizations of Eigenvalues LetA∈Mnbe Hermitian. Assume λmin≤λ1≤λ2≤···≤λn−1≤λn=λmax. Theorem 9.1.1 (Rayleigh—Ritz). LetA∈Mn, and let the eigenvalues ofAbe ordered as above. Then λmax=λn=m a x x6=0hAx, x i hx, xi λmin=λ1=m i n x6=0hAx, x i hx, xi. 242 CHAPTER 9. HERMITIAN AND SYMMETRIC MATRICES Proof. Letx1...x nbe the linearly independent and orthogonal eigenvectors ofA. Any vector x∈Cnhas the representation xΣαjxj. Then hAx, x i=Σα2 jλjand hx, xi=Σα2 j.H e n c e hAx, x i hx, xi=X jα2 jP iα2iλj. Since½ α2 j Σα2 i¾ is nonnegative and sums to 1, it follows that hAx, x i hx, xi is a convex combination of the eigenvalues, whence the theorem follows. In particular take x=xn(x1) to achieve the maximum (minimum). Remark 9.1.1. This result gives just the largest and smallest eigenvalues. How can we achieve the intermediate eigenvalues? We have already consid- ered this problem somewhat in conjunction with the power method. In thatconsideration we employed the bi-orthogonal eigenvectors. For a Hermitianmatrix, the families are the same. So we could characterize the eigenvaluesin a manner similar to that discussed previously. However, the following characterization is simpler. Theorem 9.1.2. LetA∈M nbe Hermitian with eigenvalues as above and corresponding eigenvectors x1...x n.T h e n (∗) λn−k=m a x x⊥{xn,... ,x n−k+1} x 6=0hAx, x i hx, xi. The following result is even more general (∗) λn−k=m i n {wn,wn−1...wn−k+1}max x⊥{wn,wn−1...wn−k+1} x 6=0hAx, x i hx, xi. Proof. The Rayleigh—Ritz argument above gives ( ∗) directly. 9.2. MATRIX INEQUALITIES 243 9.2 Matrix inequalities When the underlying matrix is symmetric or positive de finite, certain pow- erful inequalities can be established. The first inequality is a consequence of the cofactor result Proposition 2.5.1. Theorem 9.2.1. LetA∈Mn(C)be positive de finite. Then detA≤a11a22···ann Proof. Expanding in minors, the determinant of A detA=a11¯¯¯¯¯¯¯a 22···a2n ......... an2 ann¯¯¯¯¯¯¯+d e t¯¯¯¯¯¯¯¯¯0a 12···a1n a21a22···a2n ............ an1an2 ann¯¯¯¯¯¯¯¯¯ Since Ais positive de finite, it follows that each of it principal submatrices is also. Therefore, the first term in the right side of the equality above is postive, while the second term is negative by the cofactor result Proposition2.5.1. To clarify , it is important to note that if Ais Hermitian and invert- ible, so also is its inverse. In particular, it Ais positive de finite, we know its determinant is positive and so when we write det¯¯¯¯¯¯¯¯¯0a 12···a1n a21a22···a2n ............ an1an2 ann¯¯¯¯¯¯¯¯¯=−X a 1ja1kaik it is clear this term is negative. Therefore, detA≤a11¯¯¯¯¯¯¯a 22···a2n ......... an2 ann¯¯¯¯¯¯¯ whence the result follows inductively. Corollary 9.2.1. LetS1,S2,..., S rbe any partition of the integers {1,...,n }, and let A1,A2,...,A r., be the principal submatrices of Apertaining to the indices of each subdivision. Then detA=d e t A1detA2···detAr An important consequence of Theorem 9.2.1 is one of the most famous determinantal inequalities. It is due to Hadamard. 244 CHAPTER 9. HERMITIAN AND SYMMETRIC MATRICES Theorem 9.2.1. Let B be an arbitrary nonsingular real square matrix. Then (detB)2≤nY i=1nX j=1|bij|2 Proof. DefineA=BTB.Then diagA= nX j=1|b1j|2,. . . ,nX j=1|bnj|2  whence the result follows from Theorem 9.2.1 since det A=( d e t B)2. 9.3 Exercises 1. Prove Corollary 9.2.1. 2. Establish Theorem 9.2.1 in the case the matrix is complex. 3. Establish a result similar to Theorem 9.2.1 for rectangular matrices. Chapter 10 Nonnegative Matrices 10.1 De finitions Nonnegative matrices are simply those with all nonnegative entries. We will eventually distinguish various types of such matrices substantially onhow they are nonnegative or conversely where they are zero. A particular type of nonnegative matrix, the type that has no o ff-diagonal block to be zero, will turn out to have the greatest value to us. If nonnegative matricesdistinguished themselves in only minor ways from general matrices, theywould not occupy the high position of importance they enjoy in the modernliterature. Indeed, many remarkable properties of nonnegative matriceshave been uncovered. Owing substantially to the many applications to economics, probability, and engineering, the subject has now for more than a century been one of the hottest areas of research within the subject. Definition 10.1.1. A∈M n(R)i sc a l l e d nonnegative ifaij≥0f o r1≤ i, j≤n. We use the notation A≥0 for nonnegative matrices. If aij>0 for all i, jwe say Aisfully (orstrictly )positive . (Notation. AÀ0.) In this section we need some special notation. Let x∈Cn,x=(x1,... ,x n)T. Then |x|=(|x1|,... , |xn|)T. We say that x∈Rnispositive ifxi≥0, 1≤i≤n.W es a yt h a t x∈Rnis fully (orstrictly )positive ifxi>0, 1≤i≤n. We use the notation x≥0 andxÀ0 for positive and strictly positive vectors, respectively. Note that if A≥0w eh a v e kAk∞=kAek∞where e=( 1,1,... , 1)T. 245 246 CHAPTER 10. NONNEGATIVE MATRICES To give an idea of the direction we are heading, let us suppose that A∈Mn(R) is a nonnegative matrix. Let λ6= 0 be an eigenvector of A with pertaining eigenvector x≥0. Thus Ax=λx. (We will prove that this comes to pass for every nonnegative square matrix.) Suppose furtherthat the vector Axis zero on the index set S.That is ( Ax) i=0i f i∈S. Sinceλ6=0i tf o l l o w st h a t xi=0i∈S.L e t SCbe the complement of Sin the integers {1,...,n },a n dl e t PandP⊥be the projections to Sand SC, respectively. So, Px=0a n d P⊥x=x. Also, it is easy to see that In=P+P⊥,and hence PAP⊥x=λPx=0 Therefore PAP⊥=0.When we write A=³ P+P⊥´ A³ P+P⊥´ in block form A=³ P+P⊥´ A³ P+P⊥´ =·PAP PAP⊥ P⊥AP P⊥AP⊥¸ we see that Ahas the special form A=·PAP 0 P⊥AP P⊥AP⊥¸ This particular calculation reveals in a simple way the consequences of zero components to eigenvectors of nonnegative matrices. Zero components ofeigenvectors implies zero blocks of A. I tw o u l db ec o r r e c tt oo b s e r v et h a t this argument requires a nonnegative eig envector. Nonetheless, there are still some very special properties of the remaining eigenvalues of A,p a r t i c u l a r l y those with the modulus equal to the spectral radius. It will be convenient tohave a name for the index set where a nonnegative vector x∈R nis strictly positive. Definition 10.1.2. Letx∈Rnbe nonnegative, x≥0. Let Sbe the index set such that xi6=0i f i∈S.W ec a l l Sthesupport of the vector x. 10.2 General Theory Definition 10.2.1. ForA∈Mn(R)a n dλ/∈σ(A)w ed e fine the resolvent ofAby R(λ)=(λI−A)−1. 10.2. GENERAL THEORY 247 This function behaves much like a rational function in complex variable t h e o r y .I th a sp o l e sa te a c h λ∈σ(A). The order of the pole is the same as t h eo r d e ro ft h ez e r oo f λin the minimal polynomial of A. Theorem 10.2.1. Suppose A∈Mn(R)is nonnegative and λ>ρ(A),t h e n R(λ)=(λI−A)−1≥0. Proof. Apply the Neumann series (?)( λI−A)−1=∞X k=0λ−(k+1)Ak. Sinceλ>0a n d A≥0, it is apparent that R(λ)≥0. Remark 10.2.1. It is apparent that ( ?) holds upon noting that (λI−A)−1à 1−µA λ¶n+1! =nX 0λ−(k+1)Ak and that |λ|>ρ(A) yields convergence. Theorem 10.2.2. Suppose A∈Mn(R)is nonnegative. Then the spectral radiusρ(A)ofAis an eigenvalue of Awith at least one eigenvector, x≥0. Proof. Suppose for each y≥0,R(λ)yremains bounded as λ↓ρ.I fx∈Cn is arbitrary |R(λ)x|=¯¯¯¯¯∞X n=0λ−(n+1)Anx¯¯¯¯¯ ≤∞X n=0|λ|−(n+1)An|x| =R(|λ|)|x| for allλ,|λ|>ρ. It follows that R(λ)xis uniformly bounded in the region |λ|>ρ, and this is impossible. Now let y0≥0 be a vector for which R(λ)y0is unbounded as λ↓p,a n d letkkdenote a vector norm. For λ>r,s e t z(λ)=R(λ)y0/kR(λ)y0k, where kR(λ)y0k↑∞.N o w z(λ)⊂{x|kxk=1}, and the latter is compact. Therefore {z(λ)}λ>ρhas a cluster point x0with x0≥0a n d kx0k=1 . S i n c e (ρI−A)z(λ)=(ρ−λ)z(λ)+y0/kR(λ)y0k 248 CHAPTER 10. NONNEGATIVE MATRICES and since λ→ρand kR(λ)y0k↑∞ we have lim λ→ρ(ρI−A)z(λ)=l i m λ→ρ(ρI−A)x0=0. Corollary 10.2.1. Suppose A∈Mn(R)is nonnegative and Az=λzfor some zÀ0,t h e nλ=ρ(A). Proof. Sinceρ(A)⊂σ(AT) there is a vector y≥0f o rw h i c h ATy=ρy.I f λ6=ρ,w em u s th a v e hz,yi=0 . B u t zÀ0 and this is impossible. Corollary 10.2.2. Suppose A∈Mn(R)is strictly positive and z≥0such thatAz=ρz,t h e n zÀ0. Proof. Ifzi=0 ,t h e nnP j=1aijzj= 0, and thus zj=0f o ra l l jsince aij>0, for all j. Hence z= 0, a contradiction. This result will be strengthened in the next section by another result that has the same conclusion but a much weaker hypotheses. Theorem 10.2.3. Suppose A∈Mn(R)is nonnegative with spectral radius ρ(A)and de fine rm=m i n iX jaij rM=m a x iP jaij cm=m i n jX iaij cM=m a x jP iaij then both inequalities below hold. rm≤ρ(A)≤rM cm≤ρ(A)≤cM Moreover, if A, B∈Mn(R)thenρ(A+B)≥max(ρ(A),ρ(B)). Proof. We prove the second set of inequalities. Let x∈Rnbe the eigenvec- tor pertaining to ρ(A).Since x≥0 we can normalize xso thatPxi=1. Then nX j=1aijxj=ρ(A)xi,i =1,...,n 10.2. GENERAL THEORY 249 Sum both side in ito obtain nX i=1nX j=1aijxj=nX i=1ρ(A)xi nX j=1ÃnX i=1aij! xj=ρ(A) ReplacingPn i=1aijbycMmakes the sum on the left larger, that is ρ(A)≤nX j=1ÃnX i=1aij! xj≤cMnX i=1xi=cM Similarly replacingPn i=1aijbycmmakes the sum on the left smaller, and therefore ρ(A)≥cm.Thus, cm≤ρ(A)≤cM.The other inequality, rm≤ρ(A)≤rM,can be easily established by applying what has just been proved to the matrix ATnoting of course that ρ(A)=ρ(AT). The proof that ρ(A+B)≥max(ρ(A),ρ(B)) is an easy consequence of a Neumann series argument. We could just as well prove the first inequality and apply it to the transpose to obtain the second. This proof is given below. Alternative Proof. Letxbe the eigenvector pertaining to ρ(A)a n d xk= max ixi.T h e n ρxk=nX j=1aijxj≤ nX j=1aij maxxj. Hence ρ≤nX j=1aij≤max knX j=1akj=t. To obtain s≤ρ,l e ty≥0s a t i s f y ATy=ρy and suppose kyk1=1 .T h e nw eh a v eo nt h eo n eh a n d he, ATyi=ρhe, yi=ρ 250 CHAPTER 10. NONNEGATIVE MATRICES and on the other hand applying convexity he, ATyi=X iX jaT ijyj =X jyjÃX iaji! ≥min jX iaji=s Corollary 10.2.1. (i) Ifρ(A)is equal to one of the quantities cmorcM, then all of the sumsP iaijare equal for jin the support of the eigenvector xpertaining to ρ. (ii) Ifρ(A)is equal to one of the quantities rmorrM, then all of the sumsP jaijare equal for iin the support of the eigenvector yofATpertaining toρ. (iii) If all the row sums (resp. column sums) are equal, the spectral radius is equal to this value. Proof. Assume that ρ(A)=cM.L e t Sdenote the support (recall De finition 10.1.2) of the eigenvector x. Assume also that the eigenvector xis normal- ized so thatPn j=1xj=P j∈Sxj= 1. Following the proof of Theorem 10.2.3 we have nX j=1ÃnX i=1aij! xj=X j∈SÃnX i=1aij! xj=cM Since1 cM(Pn i=1aij)≤1f o re a c h j=1,...,n it follows that X j∈S1 cMÃnX i=1aij! xj≤X j∈Sxj Clearly if for some j∈Swe have1 cM(Pn i=1aij)<1 the inequality above will become a strict inequality, and the result is proved. The result is proved similarly for cm (ii) Apply the proof of (i) to AT. 10.3 Mean Ergodic Theorem LetA∈Mn(R). We have seen that understanding the nature of powers Ak, k=1,2,... leads to interesting conclusions. For example if lim k→∞Ak=P, 10.3. MEAN ERGODIC THEOREM 251 it is easy to see that Pis idempotent ( P2=P)a n d AP=PA=P.I t i s therefore a projection to the fixed space ofA. A less restrictive property is mean convergence : Mk=k−1(I+A+A2+···+Ak) lim k→∞Mk=P. Note that if ρ(A)<1w eh a v e Ak→0. Hence Mk→0. On the other hand ifρ(A)>1t h e Mkbecome unbounded. Hence mean convergence has value only when ρ(A)=1 . Theorem 10.3.1 (Mean Ergodic Theorem for matrices). LetA∈Mn(R) withρ(A)=1 .I f{Ak}∞ k=1is bounded then lim k→∞Mk=P,w h e r e Pis a pro- jection, commuting with A,o n t ot h e fixed space of A. Proof. We need a norm k·k, vector and matrix. We have kAkk<C for k=1,2,... . It is easy to see that (?) AM k−Mk=1 kÃkX 0Aj+1! −1 kÃkX 0Aj! =1 k(I+Ak+1) ≤1 k(1 +C). Since the matrices Mkare bounded they have a cluster point P. This means there is a subsequence Mkjthat converges to P. If there is another cluster point Q(i.e. there is a subsequence M`j→Q), then we compare PandQ. First we know kMkj−Pk<ε 2(1 + C) kM`j−Qk<ε 2(1 + C). From ( ?)w eh a v e AP=PandAQ=Q, whence MkjP=PandMljQ=Q for all k=1,2,... . Hence P−Q=M`j(Mkj−P)−Mkj(M`j−Q) or kP−Qk≤ε. 252 CHAPTER 10. NONNEGATIVE MATRICES Thus P=QandMkconverges to some matrix P.W eh a v e AP=PA=P hence MkP=Pfor all kand therefore P2=P. Clearly the range of P consists of all fixed vectors under A. Corollary 10.3.1. LetA∈Mn(R)and suppose that lim k→∞Ak=P,t h e n lim k→∞Mkexists and equals P. Definition 10.3.1. LetA∈Mn(R)h a v eρ(A)=1 . T h e n Ais said to be mean ergodic if limk−1(I+A+···+Ak) exists. Clearly, if A∈Mn(R)s a t i s fiesρ(A)=1a n d Akis bounded, then Ais mean ergodic. (Previous theorem.) Conversely, if Ais mean ergodic then the sequence Akis bounded. This can be proved by resorting to the Jordan canonical form and showing Akis bounded if and only if Ak jis bounded, where AJis the Jordan canonical form of A. Lemma 10.3.1. LetA∈Mn(R)withρ(A)=1 .S u p p o s e λ=1 is a simple pole of the resolvent, and the only eigenvalue of modulus 1. ThenA=P+Bwhere P=l i m λ→1(λ−1)R(λ)is a projection onto the fixed space ofA,PB=BP=0 andρ(B)<1. Proof. We know that Pexists and using the Neumann series it is clear that AP=PA=P Alim λ→1(λ−1)(λI−A)−1=A(λ−1)∞X 0λ−(k+1)Ak =l i m λ→1(λ−1)∞X 0λ−(k+1)Ak+1 =l i m λ→1(λ−1)λ"X 0λ−(k+1)Ak−λ−1# =l i m λ→1(λ−1)R(λ)=P. That is, AP=P. The other assertions follow similarly. Also, it follows that P2=l i m λ→1PR(λ)=P 10.4. IRREDUCIBLE MATRICES 253 which is to say that Pis a projection. We have, as well, that Ax=x⇒ Px=x, again using the Neumann series. Now de fineB=A−P.T h e n BP=PB= 0. Suppose Bx=ax,w h e r e |α|≥1. Then Px=0 ,f o r αPx=PBx =0. Therefore Ax=αx.H e r eα6=1i m p l i e s x= 0 by hypothesis and α=1 implies Px=x,o r x= 0, also. This contradicts the assumption that |α|≥1. Theorem 10.3.2. LetA∈Mn(R)satisfy A≥0andρ(A)=1 .T h e following are equivalent: (a)λ=1 is a simple pole of R(λ). (b){Ak}∞ 1is bounded. (c)Ais mean ergodic. Moreover lim λ→1(λ−1)R(λ)= l i m k→∞1 k(I+A+···+Ak) if either limit exists. 10.4 Irreducible Matrices Irreducible matrices form one of the cornerstones of positive matrix theory. Because the condition of irreducibility is both natural and often satis fied, their applications are far and wide. Definition 10.4.1. LetA∈Mn(C).Thezero pattern ofAis the set of ordered pairs S={(i, j)|aij=0}. Example 10.4.1. For example with A= 012 100230  S={(1,1),(2,2),(2,3),(3,3)}. 254 CHAPTER 10. NONNEGATIVE MATRICES Definition 10.4.2. A generalized permutation matrix Bis any matrix hav- i n gt h es a m ez e r op a t t e ra sap e r m u t a t i o nm a t r i x . This means of course that a generalized permutation matrix has ex- actly one nonzero entry in each row and one nonzero entry in each column.Generalized permutation matrices ar e those nonnegative matrices with non- negative inverses. Theorem 10.4.1. LetA∈M n(R)be invertible and A≥0.Then Ais invertible with nonnegati ve inverse if and only if Ais a generalized permu- tation matrix. Proof. First, if Ais a generalized permutation matrix, de fine the matrix B by bij=  0i f aij=0 1 aijifaij6=0 Then A−1=BT≥0. On the other hand suppose A−1≥0a n d Ahas on some row two non zero entries, say ai,j1andai,j2.S i n c e A−1≥0i s invertible there is a nonzero entry in the ithcolumn, say a−1 ki>0. Now compute multiply A−1A. It is easy to see that ¡ A−1A¢ k,j1>0a n d¡ A−1A¢ k,j2>0 which contradicts that A−1A=I. Example 10.4.1. The analysis for 2 ×2 matrices can be carried out directly. Suppose that A=·ab cd¸ and A−1=1 detA·d−b −ca¸ There are two cases: (1) det A> 0. In this case we conclude that b=c=0. Thus Ais a generalized permutation matrix. (2) det A< 0. In this case we conclude that a=d=0.Again Ais a generalized permutation matrix. As these are the only two cases, the result is veri fied. Definition 10.4.3. LetA, B∈Mn(R). We say that Aiscogredient to Bif there is a permutation matrix Psuch that B=PTAP 10.4. IRREDUCIBLE MATRICES 255 Note that two cogredient matrices are similar, indeed they are unitarily similar for the very special class of unitary matrices given by permutations.It is important to note that since APmerely interchanges the columns of AandP TAPthen interchanges the rows of AP.From this we observe that both AandB=PTAPhave exactly the same elements, though permuted. Definition 10.4.4. LetA∈Mn(R). We say that Aisirreducible if it is notcogredient to a matrix of the form ·A10 A3A4¸ where the block matrices A1andA4are square matrices. Otherwise the matrix is called reducible. Example 10.4.2. All diagonal matrices are reducible. There is another way to describe irreducible (and reducible) matrices in terms of projections. Let S={j1,...,j k}⊂{1,2,...,n }and de finePSto be the projection to the coordinate directions by PSei=½1if i∈S 0if i /∈S Furthermore de fine P⊥ S=I−PS Then PSandP⊥ Sare orthogonal projections. That is, the product PSP⊥ S= 0 and by de finition PS+P⊥ S=I.I f SCdenotes the complement of Sin {1,2,...,n },it is easy to see that P⊥ S=PSC Proposition 10.4.1. LetA∈Mn(R).T h e n Ais irreducible if and only if there is no set of indices S={j1,...,j k}⊂{1,2,...,n }for which PSAP⊥ S=0. Proof. Suppose there is a set of indices S={j1,...,j k}⊂{1,2,...,n }for which PSAP⊥ S=0.Define the permutation matrix as from the permutation that takes the firstkintegers 1 ,..., k to the integers j1,...,j kand the integers k+1,...,n to the complement SC.Then PTAP=·A1A2 A3A4¸ 256 CHAPTER 10. NONNEGATIVE MATRICES T h ee n t r i e si nt h eb l o c k A2are those with rows given from the rows cor- responding to the indices in Sand columns corresponding to the indices inSC.T h u s Ais reducible. Conversely, suppose that Ais reducible and that Pis a permutation matrix such that PTAP=·A10 A3A4¸ For de finiteness, let us assume that A1isk×kandA4is (n−k)×(n−k). DefineS={ji|pji,i=1,i=1,...,k }.a n d T={ji|pji,i=1,i= k+1,...,n }Because Pis a permutation, it follows that T=SCand PSAP⊥ S= 0, which proves the converse. As we know for a nonnegative matrix A∈Mn(R) the spectral radius is an eigenvalue of Awith pertaining eigenvector x≥0.When the additional assumption of irreducibility of Ais added the conclusion can be strengthened toxÀ0.To prove this result we establish a simple result from which the xÀ0 follows almost directly. Theorem 10.4.2. LetA∈Mn(R)be nonnegative and irreducible. Sup- pose that y∈Rnis nonnegative with exactly 1≤k≤n−1nonzero entries. Then (I+A)yhas strictly more nonzero entries. Proof. Suppose the nonzero entries of yareS={j1,...,j k}.S i n c e ( I+A)y= y+Ay, it follows immediately that (( I+A)y)ji6=0f o r ji∈S,a n dt h e r e - fore there are at least as many nonzero entries in ( I+A)yas there are in y.In order that the number of nonzero entries not increase, we must have for each index i/∈Sthat ( Ay)i=0.With PSdefined as the projection to the standard coordinates with indices from Swe conclude therefore that P⊥ SAP=0,which means that Ais reducible, a contradiction. Corollary 10.4.1. LetA∈Mn(R)be nonnegative. Then Ais irreducible if and only if (I+A)n−1>0. Proof. Suppose that Ais irreducible and y≥0i sa n yv e c t o ri n Rn.T h e n (I+A)yhas strictly more nonzero coordinates and thus ( I+A)2yhas even more nonzero coordinates. We see that the number of nonzero coordinatesof (I+A) kymust increase by at least one until the maximum of nnonzero coordinates is reached. This must occur by the ( n−1)thpower. Now apply this to the vectors of the standard basis e1,...,e n. For example, (I+A)n−1ekÀ0.This means that the kthcolumn of ( I+A)n−1is strictly positive. The result follows. 10.4. IRREDUCIBLE MATRICES 257 Suppose conversely that ( I+A)n−1>0 but that Ais reducible. Then there is a permutation matrix Psuch that PTAP has the form PTAP=·A10 A3A4¸ It is easy to see that all the powers of Ahave the same form – with respect to the same blocks. Also, ( I+A)n−1=I+Pn−1 i=1¡n−1 i¢ Ai.So PT(I+A)n−1P=PTà I+n−1X i=1µn−1 i¶ Ai! P =I+n−1X i=1µn−1 i¶ PTAiPT has the same form PT(I+A)n−1P=·˜A10 ˜A3˜A4¸ a n dt h i sp r o v e st h er e s u l t . Corollary 10.4.2. LetA∈Mn(R)be nonnegative. Then Ais irreducible if and only if for each pair of indices (i, j)there a positive power 1≤k≤n such that¡ Ak¢ ij>0. Proof. Suppose that Ais irreducible. Since ( I+A)n−1=I+Pn−1 i=1¡n−1 i¢ Ai> 0, it follows that A(I+A)n−1=A+Pn−1 i=1¡n−1 i¢ Ai+1>0.Thus for each pair of indices ( i, j),¡ Ak¢ ij>0f o r s o m e k. On the other hand suppose that Ais reducible. Then there is a permu- tation matrix Pfor which PTAP=·A10 A3A4¸ Follow the same argument as above to establish that A(I+A)n−1must h a v et h es a m ef o r ma s PTAPwith the same zero pattern. Thus there is no positive power 1 ≤k≤nsuch that¡ Ak¢ ij>0 for any of the pairs of indices corresponding to the (upper right) zero block above. Another characterization of irreducibility arises when we have Ax≤axfor some nonzero vector x≥0. This type of domination-type condition appears to be quite weak. Nonetheless, it is equivalent to the others. 258 CHAPTER 10. NONNEGATIVE MATRICES Corollary 10.4.3. LetA∈Mn(R)be nonnegative. Then Ais irreducible if and only if Ax≤axfor some nonzero vector x≥0,t h e n xÀ0. Proof. Ifxi=0,then aij= 0 whenever xj6=0.DefineS={1≤j≤ n|xj6=0}.Then with PSas the orthogonal projection to the standard vectors ei,i∈S,it must follow that PSAP⊥ S=0 . S i n c e Sis strictly contained in {1,2,...,n }, it follows that Ais reducible. Now suppose that Ais reducible, which means we can assume it has the form A=·A10 A3A4¸ For de finiteness, let us assume that A1isk×kandA4is (n−k)×(n−k) and of course k≥1. Let v∈Rn−kbe the nonzero nonnegative vector which satisfiesA4v=ρ(A4)v.Letu=0∈Rk.D efine the vector x=u⊕v.T h e n (Ax)i=½(A1u)i=0=ρ(A4)uiif 1≤i≤k (A1v)i=ρ(A4)vi if k +1≤i≤n The conditions of that the hypothesis are met without xÀ0. Corollary 10.4.4 (Subinvariance). LetA∈Mn(R)be nonnegative and irreducible with spectral radius ρ. Suppose that there is a positive number s and nonnegative vector z≥0for which (Az)i≤szi, for i=1,2,...,n Then (i) zÀ0, and (ii) ρ≤s. If in addition ρ=s,t h e n Az=ρz. Proof. Clearly, the same inequality holds for powers of A.F o r A(Az)≤ Az≤sz. Now apply induction. Next suppose that zi=0.By Corollary 10.4.2 we must have for each index pair i, jthat¡ Ak¢ ij>0 for some positive integer kmaking¡ Akz¢ i≤sziimpossible for some k.T h u s zÀ0.To prove the second assertion, let xbe the eigenvector of ATpertaining ρ.T h e n hAz, x i=­ z,ATx® =ρhz,xi≤shz,xi (1) Since xÀ0,the innerproduct hz,xi>0. Hence ρ≤s. Finally, if ρ=sand ( Az)i<ρzifor some index ithen we conclude from (1) that ρ<ρ, a contradiction. 10.4. IRREDUCIBLE MATRICES 259 Theorem 10.4.3. LetA∈Mn(R)be nonnegative and irreducible. If x≥ 0is the eigenvector pertaining to ρ(A), i.e. Ax=ρ(A)x,t h e n xÀ0. Proof. We have that ( I+A) is also nonnegative and irreducible, with spec- tral radius 1 + ρ(A) and pertaining eigenvector x.S i n c e (I+A)n−1x>(1 +ρ(A))n−1xÀ0 it follows that xÀ0. Corollary 10.4.5. (i) Ifρ(A)is equal to one of the quantities cmorcM from Theorem 10.2.3, then all of the sumsP iaijare equal for j=1,...,n . (ii) Ifρ(A)is equal to one of the quantities rmorrMfrom Theorem 10.2.3, then all of the sumsP jaijare equal for i=1,...,n . (iii) Conversely, if all the row sums (resp. column sums) are equal, the spectral radius is equal to this value. Proof. Both are direct consequences of Corollary 10.2.1 and Theorem 10.4.3 which imply that the eigenvectors (of AandAT, resp.) pertaining to ρare strictly positive. 10.4.1 Sharper estimates of the maximal eigenvalue Sharper estimates for the maximal eigenvalue (spectral radius) are possible.The following results assembles much of what we have considered above. Webegin with a basic inequality that has an interesting geometric interpreta- tion. Lemma 10.4.1. Leta i,bi,i=1,...n be two positive sequences. Then min jbj aj≤Pn j=1bjPn j=1aj≤max jbj aj Proof. We proceed by induction. The result is trivial for n=1 . F o r n=2 the result is a simple consequence of vector addition as follows. Considerthe ordered pairs ( a i,bi),i=1,2. Their sum is ( a1+a2,b1+b2). It is a simple matter to see that the slope of this vector must lie between the slopesof the summands, which is to say min jbj aj≤P2 j=1bjP2 j=1aj≤max jbj aj 260 CHAPTER 10. NONNEGATIVE MATRICES as is shown below. Now assume by induction the result holds up to n−1.Then we must have that Pn j=1bjPn j=1aj=Pn−1 j=1bj+bnPn−1 j=1aj+an≤max"Pn−1 j=1bjPn−1 j=1aj,bn an# ≤max· max 1≤j≤n−1bj aj,bn an¸ =m a x 1≤j≤nbj aj A similar argument serves to establish the other inequality. Theorem 10.4.4. LetA∈Mn(R)be nonnegative with row sums rkand spectral radius ρ.T h e n min k 1 rknX j=1akjrj ≤ρ≤max k 1 rknX j=1akjrj  (2) Proof. First assume that Ais irreducible. Let x∈Rnbe the principal nonnegative eigenvector of AT,s o ATx=ρxandxÀ0. Assume also that xhas been normalized so thatPxi= 1. By summing both sides of ATx=ρx,w es e et h a tPrkxk=ρ.S i n c eρ2is the spectral radius of A2,w e also have that¡ AT¢2x=¡ A2¢Tx=ρ2x. Summing both sides, it follows thatPn k=1Pn j=1rjaT jkxk=ρ2.N o w ρ=ρ2 ρ=Pn k=1Pn j=1rjaT jkxk rkxk=m a x kPn j=1akjrj rk by Lemma 10.4.1. The reverse inequality is proved similarly. In the case that Ais not irreducible, we take the limit Ak↓Awhere the Akare irreducible matrices. The result holds for each Ak, and the quantities in the inequality are continuous in the limiting process. Some small care must be t a k e ni nt h ec a s et h a t Ahas a zero row. 10.4. IRREDUCIBLE MATRICES 261 Of course, we have not yet established that the estimates (2) are sharper that the basic inequality min krk≤ρ≤max krk. That is the content of following result, whose proof is left as an exercise. Corollary 10.4.6. Let A∈Mn(R)be nonnegative with row sums rk. Then min krk≤min k 1 rknX j=1akjrj ≤ρ≤max k 1 rknX j=1akjrj ≤max krk We can continue this kind of argument even further. Let A∈Mn(R) be nonnegative with row sums riand let xbe the nonnegative eigenvector pertaining to ρ.T h e n¡ AT¢3x=ρ3xor expanded X m,k,jaT ijaT jkaTkmxm=ρ3xi Summing both sides yields X m,k,jrjaT jkaTkmxm=ρ3 Therefore, ρ=ρ3 ρ2=P mP kP jrjaT jkaTkmxmP mP jrjaT jmxm ≤max mP kP jrjaT jkaTkm P jrjaT jm=m a x mP kamkP jakjrjP jamjrj by Lemma 10.4.1, thus yielding another estimate for the spectral radius. The estimate is two-sided as those above. This is summarized in the fol-lowing result. Theorem 10.4.5. LetA∈M n(R)be nonnegative with row sums ri.T h e n min mP kamkP jakjrjP jamjrj≤ρ≤max mP kamkP jakjrjP jamjrj(3) Remember the complete proof as developed above does require the assump- tion irreducibility and the passage of the limit. Let’s consider an example and see what these estimates provide. 262 CHAPTER 10. NONNEGATIVE MATRICES Example 10.4.3. Let A= 133 535 114  The eigenvalues of Aare−2,2, and 8. Thus ρ= 8. The row sums are r1=7,r2=1 3 ,r3= 6, which yields the estimates 6 ≤ρ≤13. The estimates from (2) are64 7,8,and22 3giving the estimates 7.333333333 ≤ρ≤9.142857143 Finally, from (3) the estimates for ρare 7.818181818 ≤ρ≤8.192307692 It remains to show that the estimates in (3) furnish an improvement to the estimates in (2). This is furnished by the following corollary, which is left as an exercise. Corollary 10.4.7. LetA∈Mn(R)be nonnegative with row sums ri.T h e n for each m=1,...,n min k 1 rknX j=1akjrj ≤P kamkP jakjrjP jamjrj≤max k 1 rknX j=1akjrj  (4) These results are by no means the end of the story on the important subject of eigenvalue estimation. 10.5 Stochastic Matrices Definition 10.5.1. Am a t r i x A∈Mn(R),A≥0, is called (row) stochastic if each row sum of Ais 1 or equivalently Ae=e.( R e c a l l , e=( 1,1,... , 1)T.) Similarly, a matrix A∈Mn(R),A≥0, is called column stochastic if ATis stochastic. A direct consequence of Corollary 10.2. 1(iii) is that the spectral radius of stochastic matrices must be one. Theorem 10.5.1. LetA∈Mn(R)row or column stochastic . Then the spectral radius of Aisρ(A)=1 . 10.5. STOCHASTIC MATRICES 263 Lemma 10.5.1. Let A, B∈Mn(R)be nonnegative row stochastic ma- trices. Then for any number 0≤c≤1,t h em a t r i c e s cA+( 1−c)Band AB are also row stochastic. The same result holds for column stochastic matrices. Theorem 10.5.2. Every stochastic matrix is mean ergodic, and the periph- eral spectrum is fully cyclic. Proof. IfA≥0 is stochastic, so also is Ak,k=1,2,... .T h e r e f o r e Akis bounded. Also ρ(A)≤max iP kaij= 1. Hence the conditions of the previous theorems are met. This proves Ais mean ergodic. Stochastic matrices–as a set–have some interesting properties. Let Sn⊂ Mn(R) of stochastic matrices. Then Snis a convex set, it is also closed. Whenever a closed convex set is determined, the characterization of its ex- treme points is desired. Recall, the (matrix) point A∈Snis called an extreme point if whenever A=ΣλjAj Σλj=1 ,λj≥0a n d Aj∈Snit follows that one of the Aj’s equals Aand allλjbut one are 0. Examples. We can also view Sn⊂Rn2. It is a convex subset of Rn2 and is moreover the intersection of the positive orthant of Rn2with the n hyperplanes de fined byP jaij=1 , i=1,... ,n . [Note that we are viewing a vector x∈Rn2as x=(α11,... ,α1n,α21,... ,α2n,... ,αn1,... ,αnn)T] Snis therefore a convex polyhedron inRn2. The dimension of Snisn2−n. Theorem 10.5.3. A∈Snis an extreme point of Snif and only if Ahas exactly one 1 in each row, and all other entries are zero. Proof. LetC=cijbe a (0,1)-matrix. (That is a matrix cij=½1 0or for all 1≤i, j≤n.) Suppose that C∈Snthen there is a unique jfor which cij=1 . I f C=λA+( 1−λ)Bit is easy to see that a1j=b1j=1a n d moreover that a1k=b1k=0f o ra l lo t h e r k6=j. It follows that Cis an 264 CHAPTER 10. NONNEGATIVE MATRICES extreme point. Conversely suppose C∈Snisnota (0,1) matrix. Fix one row. Let Aj= e j c2→ ... cn→  where e j=( 0,0,... , 1, jthposition0...0) and c2...c nare the respective rows of C. Then C=nX 1C1jAj. IfC1is not a (0,1) row we are finished, since Ccannot be an extreme point. Otherwise select any non (0,1) row and repeat this argument with respectto that row. Corollary 10.5.1. Permutation matrices are extreme points. Definition 10.5.2. Alattice homomorphism ofCnis a linear map A:Cn→ Cnsuch that |Ax|=A|x|, for all x∈Cn.Ais a lattice isomorphism if both AandA−1are lattice homomorphisms. Theorem 10.5.4. A∈Snis a lattice homomorphism if and only if Ais an extreme point of Sn.A∈Snis a lattice isomorphism if and only if Ais a permutation matrix. Theorem 10.5.5. A⊂Snis a permutation matrix if and only if σ(A)⊂Γ. Γ={λ∈C||λ|=1}.Exercise. 1. Show that any row stochastic matrix (i.e. nonnegative entries with row sums are equal 1) with at least one column having equal entries is singular. 2. Show that the only nonnegative matrices with nonnegative inverses must be diagonal matrices. 3. Adapt the argument from Section 10.1 to prove that the nonnegative eigenvector xpertaining to the spectral radius of an irreducible matrix must be strictly positive. 10.5. STOCHASTIC MATRICES 265 4. Suppose that A, B∈Mn(R) are nonnegative matrices. Prove or disprove (a) If Ais irreducible then ATis irreducible. (b) If Ais irreducible then Apis irreducible, for every positive power p≥1. (c) If A2is irreducible then Ais irreducible. (d) If A, B are irreducible, then ABis irreducible. (e) If A, B are irreducible, then aA+bBis irreducible for all non- negative constants a, b≥0. 5. Suppose that A, B∈Mn(R) are nonnegative matrices. Prove that ifAis irreducible, then aA+bBis irreducible for all nonnegative constants a, b > 0. 6. Suppose that A∈Mn(R) is nonnegative and irreducible and that the trace of Ais positive, tr A> 0.Prove that Ak>0f o rs o m es u fficiently large integer k. 7. Prove Corollary 10.2.1(ii) for ρ(A)=rm. 8. Let A∈Mn(R) be nonnegative with row sums ri(A)a n dc o l u m n sums ci(A). Show thatPn i=1ri¡ A2¢ =Pn i=1ci(A)ri(A). 9. Prove Corollary 10.4.6. 10. Prove Corollary 10.4.7. (Hint. Use Lemma 10.4.1.) 11. Let pbe any positive integer and suppose that A∈Mn(R) is nonneg- ative with row sums ri.D efiner=(r1,..., r n)T. Recall that ( Apr)m refers to the mthcoordinate of the vector Apr.Show that min m(Apr)m (Ap−1r)m≤ρ≤max m(Apr)m (Ap−1r)m (Hint. Assume first that Ais irreducible.) 12. Use the estimates (2) and (3) to estimate the maximal eigenvalue of A= 514 315 122  The maximal eigenvalue is approximately 7.640246936. 266 CHAPTER 10. NONNEGATIVE MATRICES 13. Use the estimates (2) and (3) to estimate the maximal eigenvalue of A= 2144 573250605557  The maximal eigenvalue is approximately 14.34259731. 14. Suppose that A∈M n(R) is nonnegative and irreducible with spectral radiusρ, and suppose x∈Rnis any strictly positive vector, xÀ0. Defineτ=m a x i(Ax)i xi. Show that ρ≤τ. 15. A=·10 01¸ σ(A)={1} C= 001 100 010  B=·01 10¸ σ(B)={−1,1}ρC(λ)=λ3+1=0 λ=−1,eiπ/3,e−iπ/3