[Roman.S][GTM135]Advanced Linear Algebra 3ed
PDF · 528 pages · 2.5 MB
Open PDF file
Published textbook by Steven Roman (Springer, 2008), not Phil's own work. The visible text covers the prefaces and contents: vector spaces, linear transformations, modules over a PID, canonical forms, inner product and metric spaces, Hilbert spaces, tensor products, affine geometry, associative algebras and umbral calculus. It likely sits in the folder for its tensor product chapter.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
Graduate Texts
inMathematics
Graduate Texts in Mathematics 135
Editorial Board
S. Axler
K.A. Ribet
Graduate Texts in Mathematics
1T AKEUTI /ZARING . Introduction to Axiomatic
Set Theory. 2nd ed.
2O XTOBY . Measure and Category. 2nd ed.
3S CHAEFER . Topological Vector Spaces.
2nd ed.
4H ILTON /STAMMBACH . A Course in
Homological Algebra. 2nd ed.
5M ACLANE. Categories for the Working
Mathematician. 2nd ed.
6H UGHES /PIPER. Projective Planes.
7J . - P . S ERRE. A Course in Arithmetic.
8T AKEUTI /ZARING . Axiomatic Set Theory.
9H UMPHREYS . Introduction to Lie Algebras and
Representation Theory.
10 C OHEN . A Course in Simple Homotopy
Theory.
11 C ONW AY . Functions of One Complex
Variable I. 2nd ed.
12 B EALS. Advanced Mathematical Analysis.
13 A NDERSON /FULLER . Rings and Categories of
Modules. 2nd ed.
14 G OLUBITSKY /GUILLEMIN . Stable Mappings and
Their Singularities.
15 B ERBERIAN . Lectures in Functional Analysis
and Operator Theory.
16 W INTER . The Structure of Fields.
17 R OSENBLATT . Random Processes. 2nd ed.
18 H ALMOS . Measure Theory.
19 H ALMOS . A Hilbert Space Problem Book.
2nd ed.
20 H USEMOLLER . Fibre Bundles. 3rd ed.
21 H UMPHREYS . Linear Algebraic Groups.
22 B ARNES /MACK. An Algebraic Introduction to
Mathematical Logic.
23 G REUB. Linear Algebra. 4th ed.
24 H OLMES . Geometric Functional Analysis and
Its Applications.
25 H EWITT /STROMBERG . Real and Abstract
Analysis.
26 M ANES. Algebraic Theories.
27 K ELLEY . General Topology.
28 Z ARISKI /SAMUEL . Commutative Algebra.
Vo l . I .
29 Z ARISKI /SAMUEL . Commutative Algebra.
V ol. II.
30 J ACOBSON . Lectures in Abstract Algebra I.
Basic Concepts.
31 J ACOBSON . Lectures in Abstract Algebra II.
Linear Algebra.
32 J ACOBSON . Lectures in Abstract Algebra III.
Theory of Fields and Galois Theory.
33 H IRSCH . Differential Topology.
34 S PITZER . Principles of Random Walk. 2nd ed.
35 A LEXANDER /WERMER . Several Complex
Variables and Banach Algebras. 3rd ed.
36 K ELLEY /NAMIOKA et al. Linear Topological
Spaces.
37 M ONK. Mathematical Logic.38 G RAUERT /FRITZSCHE . Several Complex
Variables.
39 A RVESON . An Invitation to C∗-Algebras.
40 K EMENY /SNELL/KNAPP. Denumerable Markov
Chains. 2nd ed.
41 A POSTOL . Modular Functions and Dirichlet
Series in Number Theory. 2nd ed.
42 J.-P. S ERRE. Linear Representations of Finite
Groups.
43 G ILLMAN /JERISON . Rings of Continuous
Functions.
44 K ENDIG . Elementary Algebraic Geometry.
45 L OÈVE. Probability Theory I. 4th ed.
46 L OÈVE. Probability Theory II. 4th ed.
47 M OISE. Geometric Topology in Dimensions 2
and 3.
48 S ACHS/WU. General Relativity for
Mathematicians.
49 G RUENBERG /WEIR. Linear Geometry. 2nd ed.
50 E DW ARDS . Fermat’s Last Theorem.
51 K LINGENBERG . A Course in Differential
Geometry.
52 H ARTSHORNE . Algebraic Geometry.
53 M ANIN. A Course in Mathematical Logic.
54 G RAVER /WATKINS . Combinatorics with
Emphasis on the Theory of Graphs.
55 B ROWN /PEARCY . Introduction to Operator
Theory I: Elements of Functional Analysis.
56 M ASSEY . Algebraic Topology: An
Introduction.
57 C ROWELL /FOX. Introduction to Knot Theory.
58 K OBLITZ .p-adic Numbers, p-adic Analysis,
and Zeta-Functions. 2nd ed.
59 L ANG. Cyclotomic Fields.
60 A RNOLD . Mathematical Methods in Classical
Mechanics. 2nd ed.
61 W HITEHEAD . Elements of Homotopy Theory.
62 K ARGAPOLOV /MERIZJAKOV . Fundamentals of
the Theory of Groups.
63 B OLLOBAS . Graph Theory.
64 E DW ARDS . Fourier Series. V ol. I. 2nd ed.
65 W ELLS. Differential Analysis on Complex
Manifolds. 3rd ed.
66 W ATERHOUSE . Introduction to Affine Group
Schemes.
67 S ERRE. Local Fields.
68 W EIDMANN . Linear Operators in Hilbert
Spaces.
69 L ANG. Cyclotomic Fields II.
70 M ASSEY . Singular Homology Theory.
71 F ARKAS /KRA. Riemann Surfaces. 2nd ed.
72 S TILLWELL . Classical Topology and
Combinatorial Group Theory. 2nd ed.
73 H UNGERFORD .A l g e b r a .
74 D AV EN PO RT . Multiplicative Number Theory.
3rd ed.
75 H OCHSCHILD . Basic Theory of Algebraic
Groups and Lie Algebras.
(continued after index)
Steven Roman
Advanced Linear Algebra
Third Edition
Steven Roman
8 Night Star
Irvine, CA 92603
USA
[email protected]
Editorial Board
S. Axler K.A. Ribet
Mathematics Department Mathematics Department
San Francisco State University University of California at Berkeley
San Francisco, CA 94132 Berkeley, CA 94720-3840
USA USA
[email protected] [email protected]
ISBN-13: 978-0-387-72828-5 e-ISBN-13: 978-0-387-72831-5
Library of Congress Control Number: 2007934001
Mathematics Subject Classification (2000): 15-01
c/circlecopyrt2008 Springer Science+Business Media, LLC
All rights reserved. This work may not be translated or copied in whole or in part without the written
permission of the publisher (Springer Science+Bu siness Media, LLC, 233 Spring Street, New York,
NY 10013, USA), except for brief excerpts in connec tion with reviews or scholarly analysis. Use in
connection with any form of information storage and retri eval, electronic adaptation, computer software,
or by similar or dissimilar methodology now known or hereafter developed is forbidden.
The use in this publication of trade names, trademarks, service marks, and similar terms, even if they are
not identified as such, is not to be taken as an expre ssion of opinion as to whether or not they are subject
to proprietary rights.
Printed on acid-free paper.
987654321
springer.com
To Donna
and to
Rashelle, Carol and Dan
Preface to the Third Edition
Let me begin by thanking the readers of the second edition for their many
helpful comments and suggestions, with special thanks to Joe Kidd and Nam
Trang. For the third edition, I have corrected all known errors, polished and
refined some arguments (such as the discussion of reflexivity, the rational
canonical form, best approximations and the definitions of tensor products) and
upgraded some proofs that were originally done only for finite-dimensional/rank
cases. I have also moved some of the material on projection operators to an
earlier position in the text.
A few new theorems have been added in this edition, including the spectral
mapping theorem and a theore m to the effect that , with dim dim²= ³ ²= ³i
equality if and only if is finite-dimensional. =
I have also added a new chapter on associative algebras that includes the well-
known characterizations of the finite-dime nsional division algebras over the real
field (a theorem of Frobenius) and over a finite field (Wedderburn's theorem).
The reference section has been enlarged considerably, with over a hundred
references to books on linear algebra.
Steven Roman Irvine, California, May 2007
Preface to the Second Edition
Let me begin by thanking the readers of the first edition for their many helpful
comments and suggestions. The second edition represents a major change from
the first edition. Indeed, one might say th at it is a totally new book, with the
exception of the general range of topics covered.
The text has been completely rewritten. I hope that an additional 12 years and
roughly 20 books worth of experience has enabled me to improve the quality of
my exposition. Also, the exercise sets have been completely rewritten.
The second edition contains two new chapters: a chapter on convexity,
separation and positive solutions to linear systems Chapter 15) and a chapter on (
the QR decomposition, singular values and pseudoinverses Chapter 17). The (
treatments of tensor products and the umbral calculus have been greatly
expanded and I have included discussions of determinants in the chapter on (
tensor products), the complexification of a real vector space, Schur's theorem
and Geršgorin disks.
Steven Roman Irvine, California February 2005
Preface to the First Edition
This book is a thorough introduction to linear algebra, for the graduate or
advanced undergraduate student. Prerequisites are limited to a knowledge of the
basic properties of matrices and determinants. However, since we cover the
basics of vector spaces and linear transformations rather rapidly, a prior course
in linear algebra even at the sophomore level), along with a certain measure of (
“mathematical maturity,” is highly desirable.
Chapter 0 contains a summary of certain topics in modern algebra that are
required for the sequel. This chapter should be skimmed quickly and then used
primarily as a reference . Chapters 1–3 contain a discussion of the basic
properties of vector spaces and linear transformations.
Chapter 4 is devoted to a discussi on of modules, emphasizing a comparison
between the properties of modules and those of vector spaces. Chapter 5
provides more on modules. The main goals of this chapter are to prove that any
two bases of a free module have the same cardinality and to introduce
Noetherian modules. However, the instructor may simply skim over this
chapter, omitting all proofs. Chapter 6 is devoted to the theory of modules over
a principal ideal domain, establishing the cyclic decomposition theorem for
finitely generated modules. This theorem is the key to the structure theorems for
finite-dimensional linear operators, discussed in Chapters 7 and 8.
Chapter 9 is devoted to real and complex inner product spaces. The emphasis
here is on the finite-dimensional case, in order to arrive as quickly as possible at
the finite-dimensional spectral theore m for normal operators, in Chapter 10.
However, we have endeavored to state as many results as is convenient for
vector spaces of arbitrary dimension.
The second part of the book consists of a collection of independent topics, with
the one exception that Chapter 13 requires Chapter 12. Chapter 11 is on metric
vector spaces, where we describe the structure of symplectic and orthogonal
geometries over various base fields. Chapter 12 contains enough material on
metric spaces to allow a unified treatment of topological issues for the basic
xii Preface
Hilbert space theory of Chapter 13. The rather lengthy proof that every metric
space can be embedded in its completion may be omitted.
Chapter 14 contains a brief introducti on to tensor products. In order to motivate
the universal property of tensor products, without getting too involved in
categorical terminology, we first treat both free vector spaces and the familiar
direct sum, in a universal way. Chapter 15 (Chapter 16 in the second edition) is
on affine geometry, emphasizing algebraic , rather than geometric, concepts.
The final chapter provides an introduction to a relatively new subject, called the
umbral calculus. This is an algebraic theory used to study certain types of
polynomial functions that play an impor tant role in applied mathematics. We
give only a brief introduction to the subject emphasizing the algebraic c
aspects, rather than the a pplications. This is the first time that this subject has
appeared in a true textbook.
One final comment. Unless otherwise me ntioned, omission of a proof in the text
is a tacit suggestion that the reader attempt to supply one.
Steven Roman Irvine, California
Contents
Preface to the Third Edition, vii
Preface to the Second Edition, ix
Preface to the First Edition, xi
Preliminaries, 1
Part 1: Preliminaries, 1
Part 2: Algebraic Structures, 17
Part I—Basic Linear Algebra, 33
1 Vector Spaces, 35
Vector Spaces, 35
Subspaces, 37
Direct Sums, 40
Spanning Sets and Linear Independence, 44
The Dimension of a Vector Space, 48
Ordered Bases and Coordinate Matrices, 51
The Row and Column Spaces of a Matrix, 52
The Complexification of a Real Vector Space, 53
Exercises, 55
2 Linear Transformations, 59
Linear Transformations, 59
The Kernel and Image of a Linear Transformation, 61
Isomorphisms, 62
The Rank Plus Nullity Theorem, 63
Linear Transformations from to , 64 --
Change of Basis Matrices, 65
The Matrix of a Linear Transformation, 66
Change of Bases for Linear Transformations, 68
Equivalence of Matrices, 68
Similarity of Matrices, 70
Similarity of Operators, 71
Invariant Subspaces and Reducing Pairs, 72
Projection Operators, 73
xiv Contents
Topological Vector Spaces, 79
Linear Operators on , 82 =d
Exercises, 83
3 The Isomorphism Theorems, 87
Quotient Spaces, 87
The Universal Property of Quotients and
the First Isomorphism Theorem, 90
Quotient Spaces, Complements and Codimension, 92
Additional Isomorphism Theorems, 93
Linear Functionals, 94
Dual Bases, 96
Reflexivity, 100
Annihilators, 101
Operator Adjoints, 104
Exercises, 106
4 Modules I: Basic Properties, 109
Motivation, 109
Modules, 109
Submodules, 111
Spanning Sets, 112
Linear Independence, 114
Torsion Elements, 115
Annihilators, 115
Free Modules, 116
Homomorphisms, 117
Quotient Modules, 117
The Correspondence and Isomorphism Theorems, 118
Direct Sums and Direct Summands, 119
Modules Are Not as Nice as Vector Spaces, 124
Exercises, 125
5 Modules II: Free and Noetherian Modules, 127
The Rank of a Free Module, 127
Free Modules and Epimorphisms, 132
Noetherian Modules, 132
The Hilbert Basis Theorem, 136
Exercises, 137
6 Modules over a Principal Ideal Domain, 139
Annihilators and Orders, 139
Cyclic Modules, 140
Free Modules over a Principal Ideal Domain, 142
Torsion-Free and Free Modules, 145
The Primary Cyclic Decomposition Theorem, 146
The Invariant Factor Decomposition, 156
Characterizing Cyclic Modules, 158
Contents xv
Indecomposable Modules, 158
Exercises, 159
7 The Structure of a Linear Operator, 163
The Module Associated with a Linear Operator, 164
The Primary Cyclic Decomposition of , 167 =
The Characteristic Polynomial, 170
Cyclic and Indecomposable Modules, 171
The Big Picture, 174
The Rational Canonical Form, 176
Exercises, 182
8 Eigenvalues and Eigenvectors, 185
Eigenvalues and Eigenvectors, 185
Geometric and Algebraic Multiplicities, 189
The Jordan Canonical Form, 190
Triangularizability and Schur's Theorem, 192
Diagonalizable Operators, 196
Exercises, 198
9 Real and Complex Inner Product Spaces, 205
Norm and Distance, 208
Isometries, 210
Orthogonality, 211
Orthogonal and Orthonormal Sets, 212
The Projection Theorem and Best Approximations, 219
The Riesz Representation Theorem, 221
Exercises, 223
10 Structure Theory for Normal Operators, 227
The Adjoint of a Linear Operator, 227
UnitaryDiagonalizability,233
Normal Operators, 234
Special Types of Normal Operators, 238
Self-Adjoint Operators, 239
Unitary Operators and Isometries, 240
The Structure of Normal Operators, 245
Functional Calculus, 247
Positive Operators, 250
The Polar Decomposition of an Operator, 252
Exercises, 254
Part II—Topics, 257
11 Metric Vector Spaces: The Theory of Bilinear Forms, 259
Symmetric, Skew-Symmetric and Alternate Forms, 259
The Matrix of a Bilinear Form, 261Orthogonal Projections, 231
xvi Contents
Quadratic Forms, 264
Orthogonality, 265
Linear Functionals, 268
Orthogonal Complements and Orthogonal Direct Sums, 269
Isometries, 271
Hyperbolic Spaces, 272
Nonsingular Completions of a Subspace, 273
The Witt Theorems: A Preview, 275
The Classification Problem for Metric Vector Spaces, 276
Symplectic Geometry, 277
The Structure of Orthogonal Geometries: Orthogonal Bases, 282
The Classification of Orthogonal Geometries:
Canonical Forms, 285
The Orthogonal Group, 291
The Witt Theorems for Orthogonal Geometries, 294
Maximal Hyperbolic Subspaces of an Orthogonal Geometry, 295
Exercises, 297
12 Metric Spaces, 301
The Definition, 301
Open and Closed Sets, 304
Convergence in a Metric Space, 305
The Closure of a Set, 306
Dense Subsets, 308
Continuity, 310
Completeness, 311
Isometries, 315
The Completion of a Metric Space, 316
Exercises, 321
13 Hilbert Spaces, 325
A Brief Review, 325
Hilbert Spaces, 326
Infinite Series, 330
An Approximation Problem, 331
Hilbert Bases, 335
Fourier Expansions, 336
A Characterization of Hilbert Bases, 346
Hilbert Dimension, 346
A Characterization of Hilbert Spaces, 347
The Riesz Representation Theorem, 349
Exercises, 352
14 Tensor Products, 355
Universality, 355
Bilinear Maps, 359
Tensor Products, 361
Contents xvii
When Is a Tensor Product Zero?, 367
Coordinate Matrices and Rank, 368
Characterizing Vectors in a Tensor Product, 371
Defining Linear Transformations on a Tensor Product, 374
The Tensor Product of Linear Transformations, 375
Change of Base Field, 379
Multilinear Maps and Iterated Tensor Products, 382
Tensor Spaces, 385
Special Multilinear Maps, 390
Graded Algebras, 392
The Symmetric and Antisymmetric
The Determinant, 403
Exercises, 406
15 Positive Solutions to Linear Systems:
Convexity and Separation, 411
Convex, Closed and Compact Sets, 413
Convex Hulls, 414
Linear and Affine Hyperplanes, 416
Separation, 418
Exercises, 423
16 Affine Geometry, 427
Affine Geometry, 427
Affine Combinations, 428
Affine Hulls, 430
The Lattice of Flats, 431
Affine Independence, 433
Affine Transformations, 435
Projective Geometry, 437
Exercises, 440
17 Singular Values and the Moore–Penrose Inverse, 443
Singular Values, 443
The Moore–Penrose Generalized Inverse, 446
Least Squares Approximation, 448
Exercises, 449
18 An Introduction to Algebras, 451
Motivation, 451
Associative Algebras, 451
Division Algebras, 462
Exercises, 469
19 The Umbral Calculus, 471
Formal Power Series, 471
The Umbral Algebra, 473 Tensor Algebras, 392
xviii Contents
Formal Power Series as Linear Operators, 477
Sheffer Sequences, 480
Examples of Sheffer Sequences, 488
Umbral Operators and Umbral Shifts, 490
Continuous Operators on the Umbral Algebra, 492
Operator Adjoints, 493
Umbral Operators and Automorphisms
of the Umbral Algebra, 494
Umbral Shifts and Derivations of the Umbral Algebra, 499
The Transfer Formulas, 504
A Final Remark, 505
Exercises, 506
References, 507
Index of Symbols, 513
Index, 515
Preliminaries
In this chapter, we briefly discuss some topics that are needed for the sequel.
This chapter should be skimmed quickly and used primarily as a reference.
Part 1 Preliminaries
Multisets
The following simple concept is much more useful than its infrequent
appearance would indicate.
Definition Let be a nonempty set. A with is a:4 : multiset underlying set
set of ordered pairs
4~¸ ² Á³ : Á Á £ £ ¹ b{ for
where . The number is referred to as the of the{b ~ ¸ÁÁù multiplicity
elements in . If the underlying set of a multiset is finite, we say that the 4
multiset is . The of a finite multis et is the sum of the multiplicities finite size 4
of all of its elements.
For example, is a multiset with underlying set 4 ~ ¸²Á³Á²Á³Á²Á³¹
:~¸ÁÁ¹ . The element has multiplicity . One often writes out the
elements of a multiset according to multiplicities, as in . 4~¸ÁÁÁÁÁ¹
Of course, two mutlisets are equal if their underlying sets are equal and if the
multiplicity of each element in the common underlying set is the same in both
multisets.
Matrices
The set of matrices with entries in a field is denoted by or d - ²-³ CÁ
by when the field does not requi re mention. The set is denoted CC <Á Á ²³
by or If , the th entry of will be denoted by .CC C Á ²-³ À ( ²Á³ ( (
The identity matrix of size is denot ed by . The elements of the base d 0
2 Advanced Linear Algebra
field are called . We expect that the reader is familiar with the basic - scalars
properties of matrices, including matrix addition and multiplication.
The of an matrix is the sequence of entriesmain diagonal d (
(Á (Á Ã Á (Á Á Á
where .~ ¸ Á ¹ min
Definition The of is the matrix defined by transpose ( (CÁ!
²( ³ ~ (!
Á Á
A matrix is if and if . (( ~ ( ( ~ c ( symmetric skew-symmetric!!
Theorem 0.1 Properties of the transpose Let , . Then() ()CÁ
1)²( ³ ~ (!!
2)²(b)³ ~ ( b)!!!
3 for all )²(³ ~ ( -!!
4 provided that the product is defined)²()³ ~ ) ( ()!! !
5 .)d e t d e t²( ³ ~ ²(³!
Partitioning and Matrix Multiplication
Let be a matrix of size . If and , then4 d ) ¸ÁÃÁ¹ * ¸ÁÃÁ¹
the is the matrix obtained from by keeping only thesubmatrix 4´)Á*µ 4
rows with index in and the columns w ith index in . Thus, all other rows and )*
columns are discarded and has size . 4´)Á*µ ) d * (( ((
Suppose that and . Let 4 5CCÁ Á
1) be a partition of F~¸) ÁÃÁ) ¹ ¸ÁÃÁ¹
2) be a partition of G~¸*ÁÃÁ*¹ ¸ÁÃÁ¹
3) be a partition of H~¸+ÁÃÁ+¹ ¸ÁÃÁ¹
(Partitions are defined formally later in this chapter.) Then it is a very useful fact
that matrix multiplication can be performed at the block level as well as at the
entry level. In particular, we have
´45µ´) Á+ µ ~ 4´) Á* µ5´* Á+ µ
*
G
When the partitions in question cont ain only single-element blocks, this is
precisely the usual formul a for matrix multiplication
´45µ ~ 4 5 Á Á Á
~
Preliminaries 3
Block Matrices
It will be convenient to introduce the notational device of a block matrix. If )Á
are matrices of the appropriate sizes, then by the block matrix
4~))Ä )
ÅÅ Å
))Ä )vy
wzÁ Á Á
Á Á Áblock
we mean the matrix whose upper left is , and so on. Thus, the submatrix )Á
)4Á's are of and not entries. A square matrix of the form submatrices
4~) Ä
Æ ÆÅ
ÅÆ Æ
Ä)vy
x{x{
wz
block
where each is square and is a zero submatrix, is said to be a ) block
diagonal matrix .
Elementary Row Operations
Recall that there are three types of elementary row operations. Type 1
operations consist of multiplying a row of by a nonzero scalar. Type 2 (
operations consist of interchanging two rows of . Type 3 operations consist of (
adding a scalar multiple of one row of to another row of . ((
If we perform an elementary operation of type to an identity matrix , the0
result is called an of type . It is easy to see that all elementary matrix
elementary matrices are invertible.
In order to perform an elementary row operation on we can perform (CÁ
that operation on the identity , to obtain an elementary matrix and then take 0,
the product . Note that multiplying on the right by has the effect of ,( ,
performing column operations.
Definition A matrix is said to be in if 9 reduced row echelon form
1 All rows consisting only of 's appear at the bottom of the matrix.)
2 In any nonzero row, the first nonzero entry is a . This entry is called a)
leading entry .
3 For any two consecutive rows, the leading entry of the lower row is to the )
right of the leading entry of the upper row.
4 Any column that contains a leading entry has 's in all other positions. )
Here are the basic facts concerning reduced row echelon form.
4 Advanced Linear Algebra
Theorem 0.2 Matrices are , denoted by , (Á) ( )CÁ row equivalent
if either one can be obtained from the other by a series of elementary row
operations.
1 Row equivalence is an equivalence relation. That is,)
a )((
b )()¬)(
c , . )() )*¬(*
2 A matrix is row equivalent to one and only one matrix that is in) (9
reduced row echelon form. The matrix is called the 9 reduced row
echelon form of . Furthermore,(
9~,Ä ,(
where are the elementary matrices required to reduce to reduced row,(
echelon form.
3 is invertible if and only if its reduced row echelon form is an identity)(
matrix. Hence, a matrix is invertible if and only if it is the product of
elementary matrices.
The following definition is probably well known to the reader.
Definition A square matrix is if all of its entries below the upper triangular
main diagonal are . Similarly, a square matrix is if all of its lower triangular
entries above the main diagonal are . A square matrix is if all of its diagonal
entries off the main diagonal are .
Determinants
We assume that the reader is familiar with the following basic properties of
determinants.
Theorem 0.3 Let . Then is an element of . Furthermore,( ² -³ ² ( ³ -CÁ det
1 For any ,) ) ² -³C
det det det²()³ ~ ²(³ ²)³
2 is nonsingular invertible if and only if .)( )(² ( ³ £ det
3 The determinant of an upper triangular or lower triangular matrix is the )
product of the entries on its main diagonal.
4 If a square matrix has the block diagonal form) 4
4~) Ä
Æ ÆÅ
ÅÆ Æ
Ä)vy
x{x{
wz
block
then . det det²4³ ~ ²) ³
Preliminaries 5
Polynomials
The set of all polynomials in the variab le with coefficients from a field is%-
denoted by . If , we say that is a polynomial . If -´%µ ²%³ -´%µ ²%³ - over
²%³ ~ b %bÄb %
is a polynomial with , then is called the of £ ² % ³ leading coefficient
and the of is , written . For convenience, the degree degree ²%³ ²%³ ~ deg
of the zero polynomial is . A polynomial is if its leading coefficient cB monic
is .
Theorem 0.4 Let where .()Division algorithm ²%³Á²%³ -´%µ ²%³ deg
Then there exist unique polynomials for which ²%³Á²%³ -´%µ
²%³ ~ ²%³²%³b²%³
where or .²%³ ~ ²%³ ²%³ deg deg
If , that is, if there exists a polynomial for which²%³ ²%³ ²%³ divides
²%³ ~ ²%³²%³
then we write . A nonzero polynomial is said to ²%³ ²%³ ²%³ -´%µ split
over if can be written as a product of linear factors- ² % ³
²%³~²%c³Ä²%c ³
where . -
Theorem 0.5 Let . The of and²%³Á²%³ -´%µ ²%³ greatest common divisor
²%³ ²²%³Á²%³³ ²%³ -, denoted by , is the unique monic polynomial over gcd
for which
1 and )²%³ ²%³ ²%³ ²%³
2 if and then .)²%³ ²%³ ²%³ ²%³ ²%³ ²%³
Furthermore, there exist polynomials and over for which ²%³ ²%³ -
gcd²²%³Á²%³³ ~ ²%³²%³b²%³²%³
Definition The polynomials are if ²%³Á²%³ -´%µ relatively prime
gcd²²%³Á²%³³ ~ ²%³ ²%³ . In particular, and are relatively prime if and
only if there exist polynomials and over for which ²%³ ²%³ -
²%³²%³b²%³²%³ ~
Definition A nonconstant polynomial is if whenever ²%³-´%µ irreducible
²%³ ~ ²%³²%³ ²%³ ²%³ , then one of and must be constant.
The following two theorems support the view that irreducible polynomials
behave like prime numbers.
6 Advanced Linear Algebra
Theorem 0.6 A nonconstant polynomial is irreducible if and only if it has ²%³
the property that whenever , then either or ²%³ ²%³²%³ ²%³ ²%³
²%³²%³ .
Theorem 0.7 Every nonconstant polynomial in can be written as a product -´%µ
of irreducible polynomials. Moreover, this expression is unique up to order of
the factors and multiplication by a scalar.
Functions
To set our notation, we should make a few comments about functions.
Definition Let be a function from a set to a set .¢:¦; : ;
1 The of is the set and the of is .) domain range : ;
2 The of is the set .)i m image ²³~¸² ³ :¹
3 is , or an , if .)( )% £ & ¬ ² % ³ £ ² & ³injective one-to-one injection
4 is , or a , if .)( ) i m; ² ³ ~ ;surjective onto surjection
5 is , or a , if it is both injective and surjective.)bijective bijection
6 Assuming that , the of is) ; support
supp² ³~¸ :² ³£ ¹
If is injective, then its inverse exists and is well-¢:¦; ¢ ²³¦:cim
defined as a function on . im²³
It will be convenient to apply to subsets of and . In particular, if : ; ? :
and if , we set@;
²?³~¸²%³%?¹
and
² @ ³ ~ ¸ : ² ³ @ ¹c
Note that the latter is defined even if is not injective.
Let . If , the of to is the function ¢:¦; (: ( O ¢(¦; restriction (
defined by
O ²³~²³(
for all . Clearly, the restriction of an injective map is injective. (
In the other direction, if and if , then an of to is ¢:¦; :< < extension
a function for which . ¢< ¦; O ~ :
Preliminaries 7
Equivalence Relations
The concept of an equivalence relati on plays a major role in the study of
matrices and linear transformations.
Definition Let be a nonempty set. A binary relation on is called an: :
equivalence relation on if it satisfies the following conditions::
1)( )Reflexivity
for all .:
2)( )Symmetry
¬
for all .Á :
3)( )Transitivity
Á¬
for all .ÁÁ :
Definition Let be an equivalence relation on . For , the set of all: :
elements equivalent to is denoted by
´ µ~¸ : ¹
and called the of . equivalence class
Theorem 0.8 Let be an equivalence relation on . Then:
1) ´µ ¯ ´µ ¯ ´µ ~ ´µ
2 For any , we have either or .) Á : ´µ ~ ´µ ´µq´µ ~ J
Definition A of a nonempty set is a collection ofpartition :¸ ( Á Ã Á ( ¹
nonempty subsets of , called the of the partition, for which : blocks
1 for all )(q (~J £
2 .):~( rÄr(
The following theorem sheds considerable light on the concept of an
equivalence relation.
Theorem 0.9
1 Let be an equivalence relation on . Then the set of equivalence): distinct
classes with respect to are the blocks of a partition of . :
2 Conversely, if is a partition of , the binary relation defined by) F :
if and lie in the same block of F
8 Advanced Linear Algebra
is an equivalence relation on , whose equivalence classes are the blocks :
of .F
This establishes a one-to-one correspondence between equivalence relations on
:: and partitions of .
The most important problem related to equi valence relations is that of finding an
efficient way to determine when two elements are equivalent. Unfortunately, in
most cases, the definition does not provide an efficient test for equivalence and
so we are led to the following concepts.
Definition Let be an equivalence relation on . A function , where: ¢ : ¦ ;
; is any set, is called an of if it is constant on the equivalence invariant
classes of , that is,
¬² ³~² ³
and a if it is constant and distinct on the equivalence complete invariant
classes of , that is,
¯² ³~² ³
A collection of invariants is called a ¸ ÁÃÁ ¹ complete system of
invariants if
¯² ³~² ³ ~ ÁÃÁ for all
Definition Let be an equivalence relation on . A subset is said to be: * :
a set of or just a for if for every , canonical forms canonical form () :
there is such that . Put another way, each equivalence exactly one *
class under contains member of . * exactly one
Example 0.1 Define a binary relation on by letting if and - ´ % µ ² % ³ ² % ³
only if for some nonzero constant . This is easily seen to be²%³ ~ ²%³ -
an equivalence relation. The function that assigns to each polynomial its degree
is an invariant, since
²%³ ²%³ ¬ ²²%³³ ~ ²²%³³ deg deg
However, it is not a complete invariant, since there are inequivalent polynomials
with the same degree. The set of all monic polynomials is a set of canonical
forms for this equivalence relation.
Example 0.2 We have remarked that row equivalence is an equivalence relation
on . Moreover, the subset of reduced row echelon form matrices is aCÁ²-³
set of canonical forms for row equivalence, since every matrix is row equivalent
to a unique matrix in reduced row echelon form.
Preliminaries 9
Example 0.3 Two matrices , are row equivalent if and only if () ² - ³C
there is an invertible matrix such that . Similarly, and are7( ~ 7 ) ( )
column equivalent , that is, can be reduced to using elementary column ()
operations, if and only if there exists an invertible matrix such that . 8( ~ ) 8
Two matrices and are said to be if there exist invertible () equivalent
matrices and for which 78
(~7) 8
Put another way, and are equivalent if can be reduced to by () ( )
performing a series of elementary row a nd/or column operations. The use of the (
term equivalent is unfortunate, since it a pplies to all equivalence relations, not
just this one. However, the terminology is standard, so we use it here.)
It is not hard to see that an matrix that is in both reduced row echelon d 9
form and reduced column echelon form must have the block form
1~0
Á c
cÁ cÁc>?
block
We leave it to the reader to show that every matrix in is equivalent to (C
exactly one matrix of the form and so the set of these matrices is a set of 1
canonical forms for equivalence. Moreover, the function defined by
²(³~ (1 , where , is a complete invariant for equivalence.
Since the rank of is and since neither row nor column operations affect the 1
rank, we deduce that the rank of is . Hence, rank is a complete invariant for (
equivalence. In other words, two matrices are equivalent if and only if they have
the same rank.
Example 0.4 Two matrices , are said to be if there exists () ² - ³C similar
an invertible matrix such that 7
(~7) 7c
Similarity is easily seen to be an equivalence relation on . As we will learn, C
two matrices are similar if and only if they represent the same linear operators
on a given -dimensional vector space . Hence, similarity is extremely =
important for studying the structure of linea r operators. One of the main goals of
this book is to develop canonical forms for similarity.
We leave it to the reader to show th at the determinant function and the trace
function are invariants for similarity. Ho wever, these two invariants do not, in
general, form a complete system of invariants.
Example 0.5 Two matrices , are said to be if there () ² - ³C congruent
exists an invertible matrix for which 7
10 Advanced Linear Algebra
(~7) 7!
where is the transpose of . This relati on is easily seen to be an equivalence 77!
relation and we will devote some effort to finding canonical forms for
congruence. For some base fields such as , or a finite field , this is -()sd
relatively easy to do, but for other base fields such as , it is extremely ()r
difficult.
Zorn's Lemma
In order to show that any vector space has a basis, we require a result known as
Zorn's lemma . To state this lemma, we need some preliminary definitions.
Definition A is a pair where is a nonempty setpartially ordered set ²7Á ³ 7
and is a binary relation called a , read “less than or equal to,” partial order
with the following properties:
1 For all ,)( )Reflexivity 7
2 For all ,)( )Antisymmetry Á 7
~ and implies
3 For all ,)( )Transitivity ÁÁ 7
and implies
Partially ordered sets are also called . posets
It is customary to use a phrase such as “Let be a partially ordered set” when 7
the partial order is understood. Here are some key terms related to partially
ordered sets.
Definition Let be a partially ordered set.7
1 The , element of , should it exist, is an element)( ) maximum largest top 7
47 7 with the property that all elemen ts of are less than or equal to
4, that is,
7¬4
Similarly, the , , element of , should it mimimum least smallest bottom () 7
exist, is an element with the property that all elements of are 57 7
greater than or equal to , that is, 5
7¬5
2 A is an element with the property that there is no) maximal element 7
larger element in , that is, 7
7Á¬~
Preliminaries 11
Similarly, a is an element with the property that minimal element 7
there is no smaller element in , that is, 7
7Á¬~
3 Let . Then is an for and if)Á 7 " 7 upper bound
" " and
The unique smallest upper bound for and , if it exists, is called the least
upper bound of and and is denoted by . ¸ Á ¹ lub
4 Let . Then is a for and if)Á 7 M 7 lower bound
M M and
The unique largest lower bound for and , if it exists, is called the
greatest lower bound of and and is denoted by . ¸ Á ¹ glb
Let be a subset of a partially ordered set . We say that an element is:7 " 7
an for if for all . Lower bounds are definedupper bound : " :
similarly.
Note that in a partially ordered set, it is possible that not all elements are
comparable. In other words, it is possible to have with the property %Á& 7
that and .%& & %
Definition A partially ordered set in which every pair of elements is
comparable is called a , or a . Any totally ordered set linearly ordered set
totally ordered subset of a partially ordered set is called a in . 77 chain
Example 0.6
1 The set of real numbers, with the usual binary relation , is a partially)s
ordered set. It is also a totally ordered set. It has no maximal elements.
2 The set of natural numbers, together with the binary) o~ ¸ÁÁù
relation of divides, is a partially ordered set. It is customary to write
to indicate that divides . The s ubset of consisting of all powers of : o
is a totally ordered subset of , th at is, it is a chain in . The setoo
7 ~ ¸ÁÁÁÁ
Á¹ is a partially ordered set under . It has two maximal
elements, namely and . The subset is a partially 8 ~ ¸ÁÁ ÁÁ¹
ordered set in which every element is both maximal and minimal!
3 Let be any set and let be the power set of , that is, the set of all):² : ³: F
subsets of . Then , together with the subset relation , is a partially :² : ³ F
ordered set.
Now we can state Zorn's lemma, which gives a condition under which a
partially ordered set has a maximal element.
12 Advanced Linear Algebra
Theorem 0.10 If is a partially ordered set in which every()Zorn's lemma 7
chain has an upper bound, then has a maximal element. 7
We will use Zorn's lemma to prove that every vector space has a basis. Zorn's
lemma is equivalent to the famous axiom of choice. As such, it is not subject to
proof from the other axioms of ordinary (ZF) set theory. Zorn's lemma has many
important equivalancies, one of which is the . A well-ordering principle well
ordering on a nonempty set is a total order on with the property that every ??
nonempty subset of has a least element. ?
Theorem 0.11 Every nonempty set has a well()Well-ordering principle
ordering.
Cardinality
Two sets and have the same , written :; cardinality
(( ((:~;
if there is a bijective function a one-to-one correspondence between the sets. ()
The reader is probably aware of the fact that
(( (( (( (({o ro~~ and
where denotes the natural numbers, the integers and the rationalo{ r
numbers.
If is in one-to-one correspondence with a of , we write . If:; : ; subset (( ((
:; ; is in one-to-one correspondence with a subset of but not all of , proper
then we write . The second condition is necessary, since, for instance, (( ((:;
o{ o is in one-to-one correspondence with a proper subset of and yet is also in
one-to-one correspondence with itself. Hence, . {o { (( ((~
This is not the place to enter into a detailed discussion of cardinal numbers. The
intention here is that the cardinality of a set, whatever that is, represents the
“size” of the set. It is actually easier to talk about two sets having the same, or
different, size cardinality than it is to explicitly define the size cardinality of () ()
a given set.
Be that as it may, we associate to each set a cardinal number, denoted by :: ((
or , that is intended to measure th e size of the set. Actually, cardinal card²:³
numbers are just very special types of sets. However, we can simply think of
them as vague amorphous objects that measure the size of sets.
Definition
1 A set is if it can be put in one-t o-one correspondence with a set of the ) finite
form , for some nonnegative integer . A set that is{~ ¸ÁÁà Ác¹
Preliminaries 13
not finite is . The or of a finite set is infinite cardinal number cardinality ()
just the number of elements in the set.
2 The of the set of natural numbers is read “aleph)( cardinal number o L
nought” , where is the first letter of the Hebrew alphabet. Hence, ) L
(( (( ((o{r~~~ L
3 Any set with cardinality is called a set and any finite) L countably infinite
or countably infinite set is called a set. An infinite set that is not countable
countable is said to be . uncountable
Since it can be shown that , the real numbers are uncountable. (( ((so
If and are sets, then it is well known that:; finite
(( (( (( (( (( ((:; ;:¬:~; and
The first part of the next theorem tells us that this is also true for infinite sets.
The reader will no doubt recall that the of a set is the set of power setF²:³ :
all subsets of . For finite sets, the power set of is always bigger than the set ::
itself. In fact,
(( ( (:~ ¬ ² : ³~ F
The second part of the next theorem sa ys that the power set of any set is :
bigger has larger cardinality than itself. On the other hand, the third part of () :
this theorem says that, fo r infinite sets , the set of all subsets of is the :: finite
same size as . :
Theorem 0.12
1 – For any sets and ,)( )Schroder Bernstein theorem ¨ :;
(( (( (( (( (( ((:; ;: :~; and ¬
2 If denotes the power set of , then)( )Cantor's theorem F²:³ :
(( ( (: ² : ³F
3 If denotes the set of all finite subsets of and if is an infinite set, )F²:³ : :
then
(( ( (:~ ² : ³F
Proof. We prove only parts 1 and 2 . Let be an injective function )) ¢:¦;
from into and let be an in jective function from into . We :; ¢ ; ¦ : ;:
want to use these functions to create a bijective function from to . For this :;
purpose, we make the following definitions. The of an element descendants
: are the elements obtained by repeated alternate applications of the
functions and , namely
14 Advanced Linear Algebra
² ³Á²² ³³Á²²² ³³³ÁÃ
If is a descendant of , then is an of . Descendants and ancestors! ! ancestor
of elements of are defined similarly. ;
Now, by tracing an element's ancestry to its beginning, we find that there are
three possibilities: the elemen t may originate in , or in , or it may have no :;
point of origin. Accordingly, we can write as the union of three disjoint sets:
I
I
I:
;
B~¸ : :¹
~¸ : ;¹
~¸ : ¹ originates in
originates in
has no originator
Similarly, is the di sjoint union of , and . ; JJ J:; B
Now, the restriction
O ¢ ¦I:IJ::
is a bijection. To see this, note th at if , then originated in and! ! :J:
therefore must have the form for some . But and its ancestor have ² ³ : !
the same point of origin and so implies . Thus, is surjective ! OJI:: I:
and hence bijective. We leave it to the reader to show that the functions
²O ³ ¢ ¦ O ¢ ¦J I ; Bc
;; BBIJ IJ and
are also bijections. Putting these three bijections together gives a bijection
between and . Hence, , as desired. :; : ~ ; (( ((
We now prove Cantor's theorem. The map defined by F ¢:¦ ²:³ ² ³~¸ ¹
is an injection from to a nd so . To complete the proof we :² : ³ : ² : ³FF (( ( (
must show that no injective map can be surjective. To this end, let ¢:¦ ²:³F
?~¸ : ¤² ³ ¹ ² :³ F
We claim that is not in . For suppose that for some . ? ²³ ?~²%³ %: im
Then if , we have by the definition of that . On the other hand, if%? ? %¤?
%¤? ? %? , we have again by the definiti on of that . This contradiction
implies that and so is not surjective. ?¤ ² ³ im
Cardinal Arithmetic
Now let us define addition, multiplication and exponentiation of cardinal
numbers. If and are sets, the is the set of all :; : d ; cartesian product
ordered pairs
:d;~¸ ² Á! ³ :Á!;¹
The set of all functions from to is denoted by . ;: :;
Preliminaries 15
Definition Let and denote cardinal numbe rs. Let and be disjoint sets :;
for which and . (( ((:~ ;~
1 The is the cardinal number of .) sumb: r ;
2 The is the cardinal number of .) product :d;
3 The is the cardinal number of .) power:;
We will not go into the details of why these definitions make sense. For (
instance, they seem to depend on the sets and , but in fact they do not. It :; )
can be shown, using these definiti ons, that cardinal addition and multiplication
are associative and commutative and that multiplication distributes over
addition.
Theorem 0.13 Let , and be cardinal numbers. Then the following
properties hold:
1)( )Associativity
b ²b³ ~ ²b³ b ² ³ ~ ² ³ and
2)( )Commutativity
b~b ~ and
3)( )Distributivity
²b³ ~ b
4 Properties of Exponents)( )
a ) b~
b )²³ ~
c )²³ ~
On the other hand, the arithmetic of cardinal numbers can seem a bit strange, as
the next theorem shows.
Theorem 0.14 Let and be cardinal numbers, at least one of which is
infinite. Then
b~ ~ ¸ Á ¹ max
It is not hard to see that there is a one-to-one correspondence between the power
set of a set and the set of all functions from to . This leads toF²:³ : : ¸Á¹
the following theorem.
Theorem 0.15 For any cardinal
1 I f , t h e n ) (( ( (:~ ² : ³~ F
2)
16 Advanced Linear Algebra
We have already observed that . It can be shown that is the smallest ((o~L L
infinite cardinal, that is,
L ¬ 0 is a natural number
It can also be shown that the set of real numbers is in one-to-one s
correspondence with the power set of the natural numbers. Therefore, Fo²³
((s~L
The set of all points on the real lin e is sometimes called the and so continuum
L is sometimes called the and denoted by . power of the continuum
Theorem 0.14 shows that cardinal a ddition and multiplication have a kind of
“absorption” quality, which makes it hard to produce larger cardinals from
smaller ones. The next theorem demonstrates this more dramatically.
Theorem 0.16
1 Addition applied a countable number of times or multiplication applied a)
finite number of times to the cardi nal number , does not yield anything L
more than . Specifically, for any nonzero , we have L o
Lh L~ L L~ L and
2 Addition and multiplication applied a countable number of times to the)
cardinal number does not yield more than . Specifically, we have LL
Lh ~ ² ³ ~ LL L LL and
Using this theorem, we can establish other relationships, such as
²L ³ ² ³ ~ LL L L L
which, by the Schro ¨der–Bernstein theorem, implies that
²L ³ ~ LL
We mention that the problem of evaluati ng in general is a very difficult one
and would take us far beyond the scope of this book.
We will have use for the following reasonable-sounding result, whose proof is
omitted.
Theorem 0.17 Let be a collection of sets, indexed by the set ,¸( 2¹ 2
with . If for all , then(( ( (2~ ( 2
ee
2(
Let us conclude by describing the cardinality of some famous sets.
Preliminaries 17
Theorem 0.18
1 The following sets have cardinality .) L
a The rational numbers . ) r
b The set of all finite subsets of . ) o
c The union of a countable number of countable sets. )
d The set of all ordered -tuples of integers. ){
2 The following sets have cardinality .) L
a The set of all points in . ) s
b The set of all infinite sequences of natural numbers. )
c The set of all infinite sequences of real numbers. )
d The set of all finite subsets of . ) s
e The set of all irrational numbers. )
Part 2 Algebraic Structures
We now turn to a discussion of some of the many algebraic structures that play a
role in the study of linear algebra.
Groups
Definition A is a nonempty set , together with a binary operationgroup .
denoted by *, that satisfies the following properties:
1 For all ,)( )Associativity ÁÁ .
²i³i ~ i²i³
2 There exists an element for which)( )Identity .
i ~ i ~
for all ..
3 For each , there is an element for which)( )Inverses . .c
i ~ i ~ c c
Definition A group is , or , if . abelian commutative
i ~ i
for all . When a group is abelian, it is customary to denote theÁ .
operation by +, thus writing as . It is also customary to refer to the i i b
identity as the and to denote the inverse by , referred to as zero element c c
the of .negative
Example 0.7 The set of all bijective functions from a set to is a group< ::
under composition of functions. However, in general, it is not abelian.
Example 0.8 The set is an abelian group under addition of matrices.CÁ²-³
The identity is the zero matrix 0 of size . The set is not aÁ d ²-³ C
group under multiplication of matrices, since not all matrices have multiplicative
18 Advanced Linear Algebra
inverses. However, the set of invertible matrices of size is a nonabelian d ()
group under multiplication.
A group is if it contains only a finite number of elements. The . finite
cardinality of a finite group is called its and is denoted by or . ² . ³ order
simply . Thus, for example, is a finite group under ((. ~ ¸ÁÁÃÁc¹ {
addition modulo , but is not finite. ² ³CsÁ
Definition A of a group is a nonempty subset of that is asubgroup .: .
group in its own right, using the same operations as defined on . .
Cyclic Groups
If is a formal symbol, we can define a group to be the set of all integral.
powers of :
.~¸ ¹{
where the product is defined by the formal rules of exponents:
~ b
This group is denoted by and called the . The º» cyclic group generated by
identity of is . In general, a group is if it has the form º» ~ .cyclic
.~º » . for some .
We can also create a finite group of arbitrary positive order by *² ³
declaring that . Thus, ~
* ²³ ~ ¸ ~ ÁÁ ÁÃÁ ¹ c
where the product is defined by the formal rules of exponents, followed by
reduction modulo :
~ ² b ³ mod
This defines a group of order , called a . The inverse cyclic group of order
of is .² c ³ mod
Rings
Definition A is a nonempty set , together with two binary operations,ring 9
called denoted by and denoted by juxtaposition , addition multiplication () ( ) b
for which the following hold:
1 is an abelian group under addition)9
2 For all ,)( )Associativity ÁÁ 9
²³ ~ ²³
Preliminaries 19
3 For all ,)( )Distributivity ÁÁ 9
²b³ ~ b ²b³ ~ b and
A ring is said to be if for all . If a ring 9 ~ Á 9 9 commutative
contains an element with the property that
~ ~
for all , we say that is a . The identity is usually9 9 ring with identity
denoted by .
A is a commutative ring with identity in which each nonzero elementfield-
has a multiplicative inverse, that is, if is nonzero, then there is a - -
for which . ~
Example 0.9 The set is a commutative ring under{~ ¸ÁÁà Ác¹
addition and multiplication modulo
l~²b³ Á p~ mod mod
The element is the identity. {
Example 0.10 The set of even integers is a commutative ring under the usual,
operations on , but it has no identity. {
Example 0.11 The set is a noncommutative ring under matrix additionC²-³
and multiplication. The identity matr ix is the identity for .0² - ³ C
Example 0.12 Let be a field. The set of all polynomials in a single-- ´ % µ
variable , with coefficients in , is a commutative ring under the usual %-
operations of polynomial addition and multiplication. What is the identity for
-´%µ -´% ÁÃÁ% µ ? Similarly, the set of polynomials in variables is a
commutative ring under the usual addition and multiplication of polynomials.
Definition If and are rings, then a function is a 9: ¢ 9 ¦ : ring
homomorphism if
²b³ ~ b
²³ ~ ²³ ²³
~
for all .Á 9
Definition A of a ring is a subset of that is a ring in its own subring 9: 9
right, using the same operations as defined on and having the same 9
multiplicative identity as . 9
20 Advanced Linear Algebra
The condition that a subring have the same multiplicative identity as is:9
required. For example, the set of all matrices of the form : d
(~
>?
for is a ring under addition and multiplication of matrices isomorphic to- (
-: ( 0). The multiplicative identity in is the matrix , wh ich is not the identity
of . Hence, is a ring under the same operations as but it isCCÁ Á ²-³ : ²-³
not a subring of . CÁ²-³
Applying the definition is not generally th e easiest way to show that a subset of
a ring is a subring. The following character ization is usually easier to apply.
Theorem 0.19 A nonempty subset of a ring is a subring if and only if :9
1 The multiplicative identity of is in ) 9 :9
2 is closed under subtraction, that is,):
Á:¬c:
3 is closed under multiplication, that is,):
Á : ¬ :
Ideals
Rings have another important substructure besides subrings.
Definition Let be a ring. A nonempty subset of is called an if99 ? ideal
1 is a subgroup of the abelian gr oup , that is, is closed under )?? 9
subtraction:
Á ¬ c ??
2 is closed under multiplication by ring element, that is,)? any
Á9¬ ?? ? and
Note that if an ideal contains the unit element , then . ?? ~ 9
Example 0.13 Let be a polynomial in . The set of all multiples of²%³ -´%µ
²%³,
º²%³» ~ ¸²%³²%³ ²%³ -´%µ¹
is an ideal in , called the . -´%µ ²%³ ideal generated by
Definition Let be a subset of a ring with identity. The set:9
º:»~¸ bÄb 9Á :Á¹
Preliminaries 21
of all finite linear combinations of elem ents of , with coefficients in , is an:9
ideal in , called the . It is the smallest in the sense of set 9: ideal generated by (
inclusion ideal of containing . If is a finite set, we write ) 9 : :~¸ ÁÃÁ ¹
º ÁÃÁ »~¸ bÄb 9Á :¹
Note that in the previous definition, we require that have an identity. This is 9
to ensure that . :º : »
Theorem 0.20 Let be a ring.9
1 The intersection of any collection of ideals is an ideal.) ¸ 2 ¹?
2 If is an ascending sequence of ideals, each one contained in)?? Ä
the next, then the union is also an ideal. ?
3 More generally, if)
9?~¸ 0¹
is a chain of ideals in , then the union is also an ideal in . 9~ 9 @? 0
Proof. To prove 1 , let . Then if , we have for all )@? @ ?~ Á Á
2 c 2 c . Hence, for all and so . Hence, is closed ?@ @
under subtraction. Also, if , then for all and so . Of 9 2 ?@
course, part 2 is a special case of part 3 . To prove 3 , if , then )) ) Á @?
and for some . Since one of and is contained in the other, we Á0?? ?
may assume that . It follows that and so and if ?? ? ?@ Á c
9 , then . Thus is an ideal. ?@ @
Note that in general, the union of ideals is not an ideal. However, as we have
just proved, the union of any of ideals is an ideal. chain
Quotient Rings and Maximal Ideals
Let be a subset of a commutative ring with identity. Let be the binary:9
relation on defined by 9
¯ c:
It is easy to see that is an equivalence relation. When , we say that
and are . The term “mod” is used as a colloquialism for: congruent modulo
modulo and is often written
: mod
As shorthand, we write .
22 Advanced Linear Algebra
To see what the equivalence classes look like, observe that
´ µ~¸ 9 ¹
~¸ 9c:¹
~¸ 9~b :¹
~¸ b :¹
~b: for some
The set
b:~¸ b :¹
is called a of in . The element is called a for coset coset representative :9
b: .
Thus, the equivalence classes for congruence mod are the cosets of : b : :
in . The set of all cosets is denoted by9
9°: ~ ¸b: 9¹
This is read “ mod .” We would like to place a ring structure on . 9: 9 ° :
Indeed, if is a subgroup of the abelian group , then is easily seen to be :9 9 ° :
an abelian group as well under coset addition defined by
²b:³b²b:³~²b³b:
In order for the product
²b:³²b:³~b:
to be well-defined, we must have
b:~b:¬ b:~ b:ZZ
or, equivalently,
c :¬ ² c³:ZZ
But may be any element of and may be any element of and so thisc : 9Z
condition implies that must be an ideal . Conversely, if is an ideal, then ::
coset multiplication is well defined.
Theorem 0.21 Let be a commutative ring with identity. Then the quotient9
9°?? is a ring under coset addition and multiplication if and only if is an
ideal of . In this case, is called the of , where 99 ° 9 ?? quotient ring modulo
addition and multiplication are defined by
²b:³b²b:³~²b³b:
²b:³²b:³~b:
Preliminaries 23
Definition An ideal in a ring is a if and if whenever?? 9£ 9 maximal ideal
@? @ @ ? @ is an ideal satisfying , then either or . 9 ~ ~ 9
Here is one reason why maximal ideals are important.
Theorem 0.22 Let be a commutative ring with identity. Then the quotient9
ring is a field if and only if is a maximal ideal.9°??
Proof. First, note that for any ideal of , the ideals of are precisely the ??99 °
quotients where is an ideal for which . It is clear that @? @ ? @ @?° 9 °
is an ideal of . Conversely, if is an ideal of , then let 9° 9°?A ?Z
A? A~¸ 9b ¹Z
It is easy to see that is an ideal of for which . A? A 9 9
Next, observe that a commutative ring with identity is a field if and only if ::
has no nonzero proper ideals. For if is a field and is an ideal of :: ?
containing a nonzero element , then and so . Conversely, ~ ~:c??
if has no nonzero proper ideals and , then the ideal must be : £ : º » :
and so there is an for which . Hence, is a field. : ~ :
Putting these two facts together proves the theorem.
The following result says that maximal ideals always exist.
Theorem 0.23 Any nonzero commutative ring with identity contains a 9
maximal ideal.
Proof. Since is not the zero ring, the ideal is a proper ideal of . Hence,9 ¸¹ 9
the set of all proper ideals of is nonempty. IfI 9
9?~¸ 0¹
is a chain of proper ideals in , then the union is also an ideal. 9~ @? 0
Furthermore, if is not proper, then and so , for some , @@ ?~9 0
which implies that is not proper. Hence, . Thus, any chain in ?@ I I~9
has an upper bound in and so Zorn's lemma implies that has a maximal II
element. This shows that has a maximal ideal. 9
Integral Domains
Definition Let be a ring. A nonzero element r is called a if9 9 zero divisor
there exists a nonzero for which . A commutative ring with 9 ~ 9
identity is called an if it contains no zero divisors. integral domain
Example 0.14 If is not a prime number, then the ring has zero divisors and {
so is not an integral domain. To see this, observe that if is not prime, then
~ Á in , where . But in , we have{{
24 Advanced Linear Algebra
p~ ~ mod
and so and are both zero divisors. As we will see later, if is a prime, then
{ is a field which is an integral domain, of course . ()
Example 0.15 The ring is an integral domain, since implies -´%µ ²%³²%³ ~
that or .²%³ ~ ²%³ ~
If is a ring and where , then we cannot in general cancel9 % ~ & Á%Á& 9
the 's and conclude that . For instance, in , we have , but %~& h~h {
canceling the 's gives . However, it is precisely the integral domains in ~
which we can cancel. The simple proof is left to the reader.
Theorem 0.24 Let be a commutative ring with identity. Then is an integral 99
domain if and only if the cancellation law
% ~ &Á £ ¬ % ~ &
holds.
The Field of Quotients of an Integral Domain
Any integral domain can be embedded in a field. The or 9 quotient field field (
of quotients ) of is a field that is constructed from just as the field of99
rational numbers is constructed from the ring of integers. In particular, we set
9 ~¸²Á³Á9Á£¹b
where if and only if . Addition and multiplication of²Á³ ~ ² Á ³ ~ ZZ Z Z
fractions is defined by
²Á³b²Á ³ ~ ² bÁ ³
and
²Á³h²Á ³~²Á ³
It is customary to write in the form . Note that if has zero divisors, ²Á³ ° 9
then these definitions do not make sense, because may be even if and
are not. This is why we require that be an integral domain. 9
Principal Ideal Domains
Definition Let be a ring with identity and let . The 9 9 principal ideal
generated by is the ideal
º» ~ ¸ 9¹
An in which every ideal is a principal ideal is called aintegral domain 9
principal ideal domain .
Preliminaries 25
Theorem 0.25 The integers form a principal ideal domain. In fact, any ideal ?
in is generated by the smallest positi ve integer a that is contained in . {?
Theorem 0.26 The ring is a principal ideal domain. In fact, any ideal is -´%µ ?
generated by the unique monic polynomia l of smallest degree contained in . ?
Moreover, for polynomials , ²%³ÁÃÁ ²%³
º ²%³Áà Á ²%³» ~ º ¸ ²%³Áà Á ²%³¹» gcd
Proof. Let be an ideal in and let be a monic polynomial of? -´%µ ²%³
smallest degree in . First, we observe that there is only one such polynomial in ?
??. For if is monic and , then²%³ ²²%³³ ~ ²²%³³ deg deg
²%³ ~ ²%³c²%³ ?
and since , we must have and so deg deg²²%³³ ²²%³³ ²%³ ~
²%³ ~ ²%³ .
We show that . Since , we have . To establish ?? ?~ º²%³» ²%³ º²%³»
the reverse inclusion, if , then dividing by gives ²%³ ²%³ ²%³?
²%³ ~ ²%³²%³b²%³
where or deg deg . But since is an ideal,²%³ ~ ²%³ ²%³ ?
²%³ ~ ²%³c²%³²%³ ?
and so is impossible. Hence, and ² % ³ ² % ³ ² % ³~deg deg
²%³ ~ ²%³²%³ º²%³»
This shows that and so . ?? º²%³» ~ º²%³»
To prove the second statement, let . Then, by what we ?~º ²%³ÁÃÁ ²%³»
have just shown,
?~º ²%³ÁÃÁ ²%³»~º²%³»
where is the unique monic polynomial in of smallest degree. In²%³ ²%³ ?
particular, since , we have for each . ²%³ º²%³» ²%³ ²%³ ~ Áà Á
In other words, is a common divisor of the 's. ²%³ ²%³
Moreover, if for all , then for all , which implies ²%³ ²%³ ²%³ º²%³»
that
²%³ º²%³» ~ º ²%³Áà Á ²%³» º²%³»
and so . This shows that is the common divisor of the²%³ ²%³ ²%³ greatest
² % ³'s and completes the proof.
26 Advanced Linear Algebra
Example 0.16 The ring of polynomials in two variables and is 9~-´ % Á& µ % &
not a principal ideal domain. To see this, observe that the set of all ?
polynomials with zero constant term is an ideal in . Now, suppose that is the9 ?
principal ideal . Since , there exist polynomials ??~º²%Á&³» %Á& ²%Á&³
and for which²%Á&³
% ~ ²%Á&³²%Á&³ & ~ ²%Á&³²%Á&³ and 0.1 ()
But cannot be a constant, for then we would have . Hence,²%Á&³ ~ 9 ?
deg²²%Á&³³ ²%Á&³ ²%Á&³ and so and must both be constants, which
implies that 0.1 cannot hold. ()
Theorem 0.27 Any principal ideal domain satisfies the 9 ascending chain
condition , that is, cannot have a stric tly increasing sequence of ideals 9
?? Ä
where each ideal is properly contained in the next one.
Proof. Suppose to the contrary that there is such an increasing sequence of
ideals. Consider the ideal
<~ ?
which must have the form for some . Since for some , <~º » < ?
we have for all , contradicting the fact that the inclusions are??~
proper.
Prime and Irreducible Elements
We can define the notion of a prime element in any integral domain. For
Á 9 % 9 , we say that written if there exists an for divides ()
which . ~%
Definition Let be an integral domain.9
1 An invertible element of is called a . Thus, is a unit if ) 9" 9 " # ~ unit
for some . #9
2 Two elements are said to be if there exists a unit for ) Á 9 " associates
which . We denote this by writing .~"
3 A nonzero nonunit is said to be if) 9 prime
¬ or
4 A nonzero nonunit is said to be if) 9 irreducible
~ ¬ or is a unit
Note that if is prime or irreducib le, then so is for any unit . " "
The property of being associate is clearly an equivalence relation.
Preliminaries 27
Definition We will refer to the equivalence classes under the relation of being
associate as the of . associate classes 9
Theorem 0.28 Let be a ring.9
1 An element is a unit if and only if .) "9 º " »~9
2 if and only if .) º »~º »
3 divides if and only if .) º » º »
4 , that is, where is not a unit, if and only if) ~ % %properly divides
º » º» .
In the case of the integers, an integer is prime if and only if it is irreducible. In
any integral domain, prime elements are irreducible, but the converse need not
hold. In the ring the irreducible element ( {{´c µ ~ ¸ b c Á ¹ jj
divides the product but does not divide either ²b c ³²c c ³ ~
jj
factor.)
However, in principal ideal domains, the two concepts are equivalent.
Theorem 0.29 Let be a principal ideal domain.9
1 An is irreducible if and only if the ideal is maximal.)9 º »
2 An element in is prime if and only if it is irreducible.) 9
3 The elements are , that is, have no common) Á 9 relatively prime
nonunit factors, if and only if there exist for which Á 9
b ~
This is denoted by writing . ²Á³ ~
Proof. To prove 1 , suppose that is irreducible and that . Then ) º »º »9
º » ~% %9 and so for some . The irreducibility of implies that or
% º »~9 % º »~º % »~º » is a unit. If is a unit, then and if is a unit, then .
This shows that is maximal. We have , since is not a unit. º» º» £ 9 ()
Conversely, suppose that is not irreducible, that is, where neither nor ~
º »º »9 º »~º » is a unit. Then . But if , then , which implies that
º » £ º » º » ~ 9 is a unit. Hence . Also, if , then must be a unit. So we
conclude that is not maximal, as desired. º»
To prove 2 , assume first that is prime and . Then or . We ) ~
may assume that . Therefore, . Canceling 's gives ~% ~% ~%
and so is a unit. Hence, is irre ducible. Note that this argument applies in (
any integral domain.)
Conversely, suppose that is irreducible and let . We wish to prove that
º» ºÁ» ~ º» ºÁ» ~ 9 or . The ideal is maximal and so or . In the
former case, and we are done. In the latter case, we have
~% b&
28 Advanced Linear Algebra
for some . Thus, %Á& 9
~% b&
and since divides both term s on the right, we have .
To prove 3 , it is clear that if , th en and are relatively prime. For ) b ~
the converse, consider the ideal , which must be principal, say ºÁ»
ºÁ» ~ º%» % % % . Then and and so must be a unit, which implies that
ºÁ»~9 Á 9 b ~ . Hence, there exist for which .
Unique Factorization Domains
Definition An integral domain is said to be a 9 unique factorization domain
if it has the following factorization properties:
1 Every nonzero nonunit element can be written as a product of a finite) 9
number of irreducible elements . ~Ä
2 The factorization into i rreducible elements is unique in the sense that if )
~Ä ~Ä ~ and are two such factorizations, then and
after a suitable reindexing of the factors, .
Unique factorization is clearly a desira ble property. Fortunately, principal ideal
domains have this property.
Theorem 0.30 Every principal ideal domain is a unique factorization 9
domain.
Proof. Let be a nonzero nonunit. If is irreducible, then we are done. If9
not, then , where neither factor is a unit. If and are irreducible, we ~
are done. If not, suppose that is not irreducible. Then , where ~
neither nor is a unit. Continuing in this way, we obtain a factorization of
the form after renumbering if necessary ()
~ ~² ³~² ³ ² ³~² ³ ² ³~Ä
Each step is a factorization of into a product of nonunits. However, this
process must stop after a finite number of steps, for otherwise it will produce an
infinite sequence of nonunits of for which properly divides . Á Á Ã 9 b
But this gives the ascending chain of ideals
º »º »º »º »Ä
where the inclusions are proper. But this contradicts the fact that a principal
ideal domain satisfies the ascending ch ain condition. Thus, we conclude that
every nonzero nonunit has a factorization into irreducible elements.
As to uniqueness, if and are two such factorizations, ~Ä ~Ä
then because is an integral domain, we may equate them and cancel like 9
factors, so let us assume this has been done. Thus, for all . If there are £ Á
no factors on either side, we are done. If exactly one side has no factors left,
Preliminaries 29
then we have expressed as a product of irreducible elements, which is not
possible since irreducible elements are nonunits.
Suppose that both sides have factors left, that is,
Ä ~Ä
where . Then , which implies that for some . We can£ Ä
assume by reindexing if necessary that . Since is irreducible ~
must be a unit. Replacing by and canceling gives
Ä ~ Ä c c
This process can be repeated until we r un out of 's or 's. If we run out of 's
first, then we have an equation of the form where is a unit, " Ä ~ "
which is not possible since the 's ar e not units. By the same reasoning, we
cannot run out of 's first and so and the 's and 's can be paired off as ~
associates.
Fields
For the record, let us give the definition of a field a concept that we have been (
using . )
Definition A is a set , contai ning at least two elemen ts, together with two field -
binary operations, called denoted by and addition multiplication () b
()denoted by juxtaposition , for which the following hold:
1 is an abelian group under addition.)-
2 The set of all elements in is an abelian group under) --inonzero
multiplication.
3 For all ,)( )Distributivity ÁÁ -
²b³ ~ b ²b³ ~ b and
We require that have at least two elements to avoid the pathological case in -
which .~
Example 0.17 The sets , and , of all rational, real and complex numbers,rs d
respectively, are fields, under the usual operations of addition and multiplication
of numbers.
Example 0.18 The ring is a field if and only if is a prime number. We{
have already seen that is not a field if is not prime, since a field is also an {
integral domain. Now suppose that is a prime. ~
We have seen that is an integral domain and so it remains to show that every {
nonzero element in has a multiplicative inverse. Let . Since {{ £
, we know that and are relativel y prime. It follows that there exist
integers and for which"#
30 Advanced Linear Algebra
"b# ~
Hence,
" ²c#³ mod
and so in , that is, is the multiplicative inverse of ."p~ " {
The previous example shows that not all fields are infinite sets. In fact, finite
fields play an extremely important role in many areas of abstract and applied
mathematics.
A field is said to be if every nonconstant polynomial- algebraically closed
over has a root in . This is equivalent to saying that every nonconstant--
polynomial splits over . For example, the complex field is algebraically - d
closed but the real field is not. We mention w ithout proof that every field is s -
contained in an algebraically closed field , called the of . -- algebraic closure
For example, the algebraic closure of the real field is the complex field.
The Characteristic of a Ring
Let be a ring with identity. If is a positive integer, then by , we simply 9 h
mean
h~ bÄb
terms
Now, it may happen that there is a positive integer for which
h~
For instance, in , we have . On the other hand, in , the {{ h~~
equation implies and so no such positive integer exists. h~ ~
Notice that in any ring, there must ex ist such a positive integer , since the finite
members of the infinite sequence of numbers
hÁhÁhÁÃ
cannot be distinct and so for some , whence . h~h ² c ³h~
Definition Let be a ring with identity. The smallest positive integer for 9
which is called the of . If no such number exists, weh~ 9 characteristic
say that has characteristic . The characteristic of is denoted by 9 9
char²9³.
If , then for any , we havechar²9³ ~ 9
h~ bÄb ~²bÄb³~h~
terms terms
Preliminaries 31
Theorem 0.31 Any finite ring has nonzero characteristic. Any finite integral
domain has prime characteristic.
Proof. We have already seen that a finite ring has nonzero characteristic. Let -
be a finite integral domain and suppose that . If , where char²-³ ~ ~
Á h~ ²h³²h³~ h~ , then . Hence, , implying that or
h~ . In either case, we have a contradiction to the fact that is the smallest
positive integer such that . Hence, must be prime. h~
Notice that in any field of characteristic , we have for all . - ~ -
Thus, in , -
~c - for all
This property takes a bit of getting used to and makes fields of characteristic
quite exceptional. As it happens, there are many important uses for fields of (
characteristic . It can be shown that all finite fields have size equal to a )
positive integral power of a prime and for each prime power , there is a
finite field of size . In fact, up to isomorphi sm, there is exactly one finite field
of size .
Algebras
The final algebraic structure of which we will have use is a combination of a
vector space and a ring. We have not yet officially defined vector spaces, but (
we will do so before needing the following definition, which is placed here for
easy reference.)
Definition An over a field is a nonempty set , together with algebra77 -
three operations, called denoted by , denoted by addition multiplication () ( b
juxtaposition and also denoted by juxtaposition , for )( ) scalar multiplication
which the following properties hold:
1 is a vector space over under addition and scalar multiplication.)7 -
2 is a ring under addition and multiplication.)7
3 If and , then)- Á 7
²³ ~ ²³ ~ ²³
Thus, an algebra is a vector space in which we can take the product of vectors,
or a ring in which we can multiply each element by a scalar subject, of course, (
to additional requirements as given in the definition . )
Part I—Basic Linear Algebra
Chapter 1
Vector Spaces
Vector Spaces
Let us begin with the definition of one of our principal objects of study.
Definition Let be a field, whose elements are referred to as . A - scalars vector
space vectors over is a nonempty set , whose elements are referred to as ,-=
together with two operations. Th e first operation, called and denoted addition
by , assigns to each pair of vectors in a vector in . Theb² " Á # ³ = " b # =
second operation, called and denoted by juxtaposition, scalar multiplication
assigns to each pair a vector in . Furthermore, the ²Á"³ - d= " =
following properties must be satisfied:
1 For all vectors ,)( )Associativity of addition "Á#Á$ =
"b²#b$³ ~ ²"b#³b$
2 For all vectors ,)( )Commutativity of addition "Á# =
"b#~#b"
3 There is a vector with the property that)( )Existence of a zero =
b"~"b~"
for all vectors . "=
4 For each vector , there is a vector)( )Existence of additive inverses "=
in , denoted by , with the property that=c "
"b² c " ³~² c " ³b"~
36 Advanced Linear Algebra
5 For all scalars F and for all)( )Properties of scalar multiplication Á
vectors ,"Á# =
²"b#³~"b#
²b³" ~ "b"
²³" ~ ²"³
" ~ "
Note that the first four properties in the definition of vector space can be
summarized by saying that is an abelian group under addition. =
A vector space over a field is sometimes called an . A vector space - --space
over the real field is called a and a vector space over the real vector space
complex field is called a . complex vector space
Definition Let be a nonempty subset of a vector space . A := linear
combination of vectors in is an expression of the form :
# bÄb #
where and . The scalars are called the# ÁÃÁ# : ÁÃÁ -
coefficients trivial of the linear combination. A linear combination is if every
coefficient is zero. Otherwise, it is . nontrivial
Examples of Vector Spaces
Here are a few examples of vector spaces.
Example 1.1
1 Let be a field. The set of all functions from to is a vector space)-- - --
over , under the operations of ordinary addition and scalar multiplication-
of functions:
² b³²%³ ~ ²%³b²%³
and
²³²%³ ~ ²²%³³
2 The set of all matrices with entries in a field is a vector)CÁ²-³ d -
space over , under the operations of matrix addition and scalar -
multiplication.
3 The set of all ordered -tuples whose components lie in a field , is a) - -
vector space over , with addition and scalar multiplication defined -
componentwise:
² ÁÃÁ ³b² ÁÃÁ ³~² b ÁÃÁ b ³
and
Vector Spaces 37
² ÁÃ Á ³ ~ ² ÁÃ Á ³
When convenient, we will also write the elements of in column form. -
When is a finite field with elements, we write for .-- = ² Á ³ -
4 Many sequence spaces are vector spaces. The set Seq of all infinite) ²-³
sequences with members from a field is a vector space under the -
componentwise operations
² ³b²! ³ ~ ² b! ³
and
² ³ ~ ² ³
In a similar way, the set of all sequences of complex numbers that
converge to is a vector space, as is the set of all bounded complex MB
sequences. Also, if is a positive in teger, then the set of all complex M
sequences for which ² ³
((
~B
B
is a vector space under componentwise operations. To see that addition is a
binary operation on , one verifies MMinkowski's inequality
89 8 9 8 9 (( ( ( ( (
~ ~ ~BB B
° ° °
b ! b !
which we will not do here.
Subspaces
Most algebraic structures contain substructures, and vector spaces are no
exception.
Definition A of a vector space is a subset of that is a vectorsubspace =: =
space in its own right under the opera tions obtained by restricting the
operations of to . We use the not ation to indicate that is a =: : = :
subspace of and to indicate that is a of , that is, =: = : = proper subspace
: = : £ = = ¸¹ but . The of is . zero subspace
Since many of the properties of add ition and scalar multiplication hold a fortiori
in a nonempty subset , we can establish that is a subspace merely by ::
checking that is closed under the operations of . :=
Theorem 1.1 A nonempty subset of a vector space is a subspace of if := =
and only if is closed under addition and scalar multiplication or, equivalently, :
38 Advanced Linear Algebra
: is closed under linear combinations, that is,
Á -Á"Á# : ¬ "b# :
Example 1.2 Consider the vector space of all binary -tuples, that is, =² Á³
² # ³ # = ² Á ³-tuples of 's and 's. The of a vector is the number weightM
of nonzero coordinates in . For instance, . Let be the set of # ²³ ~ , M
all vectors in of even weight. Then is a subspace of . =, = ² Á ³
To see this, note that
MM M M²"b#³ ~ ²"³b ²#³c ²"q#³
where is the vector in whose th component is the product of the"q# =²Á³
" #th components of and , that is,
²"q#³ ~ " h#
Hence, if and are both even, so is . Finally, scalarMM M²"³ ²#³ ²"b#³
multiplication over is trivial and so is a subspace of , known as -, = ² Á ³
the of .even weight subspace =² Á³
Example 1.3 Any subspace of the vector space is called a . =² Á³ linear code
Linear codes are among the most important and most studied types of codes,
because their structure allows for efficient encoding and decoding of
information.
The Lattice of Subspaces
The set of all subspaces of a vector space is partially ordered by setI²= ³ =
inclusion. The zero subspace is the smallest element in and the entire ¸¹ ²= ³ I
space is the largest element.=
If , then is the largest subspace of that is contained in:Á; ²=³ :q; =I
both and . In terms of set inclusion, is the of :; : q ; : greatest lower bound
and :;
:q;~ ¸ :Á;¹ glb
Similarly, if is any collection of subspaces of , then their ¸: 2¹ =
intersection is the greatest lower bound of the subspaces:
2:~ ¸ : 2 ¹ glb
On the other hand, if and is infinite , then if : Á ; ² =³ - :r; ² =³II ()
and only if or . Thus, the union of two subspaces is never a :; ;:
subspace in any “interesting” case. We also have the following.
Vector Spaces 39
Theorem 1.2 A nontrivial vector space over an infinite field is not the =-
union of a finite number of proper subspaces.
Proof. Suppose that , where we may assume that =~ :r Ä r :
:\: rÄr:
Let and let . Consider the infinite set$: ±²: rÄr: ³ #¤:
(~¸ $b#-¹
which is the “line” through , parallel to . We want to show that each #$ :
contains at most one vector from the infin ite set , which is contrary to the fact(
that . This will prove the theorem.=~ :r Ä r :
If for , then implies , contrary to assumption. $b#: £ $: #:
Next, suppose that and , for , where . $b#: $b#: £
Then
: ²$b#³c²$b#³~² c³$
and so , which is also contrary to assumption.$:
To determine the smallest subspace of containing the subspaces and , we =: ;
make the following definition.
Definition Let and be subspaces of . The is defined by:; = : b ; sum
:b;~¸ "b#":Á#;¹
More generally, the of any collection of subspaces is the set sum ¸: 2¹
of all finite sums of vectors from the union : :
HI c
2 2 :~ b Ä b :
It is not hard to show that the sum of any collection of subspaces of is a =
subspace of and that the sum is the least upper bound under set inclusion: =
:b;~ ¸ :Á;¹ lub
More generally,
2:~ ¸ : 2 ¹ lub
If a partially ordered set has the property that every pair of elements has a 7
least upper bound and greatest lower bound, then is called a . If has 77 lattice
a smallest element and a largest element and has the property that every
collection of elements has a least upper bound and greatest lower bound, then 7
40 Advanced Linear Algebra
is called a . The least upper bound of a collection is also called complete lattice
the of the collection and the greatest lower bound is called the .join meet
Theorem 1.3 The set of all subspaces of a vector space is a completeI²= ³ =
lattice under set inclusion, with smallest element , largest element , meet ¸¹ =
glb¸: 2¹ ~ :
2
and join
lub¸: 2¹ ~ :
2
Direct Sums
As we will see, there are many ways to construct new vector spaces from old
ones.
External Direct Sums
Definition Let be vector spaces over a field . The =ÁÃÁ= - external direct
sum of , denoted by=ÁÃÁ=
=~ = Ä = ^^
is the vector space whose elements are ordered -tuples: =
= ~¸²# ÁÃÁ# ³# =Á~ÁÃÁ¹
with componentwise operations
²" ÁÃÁ" ³b²# ÁÃÁ# ³~²" b# ÁÃÁ" b# ³
and
²# ÁÃ Á# ³ ~ ²# ÁÃ Á# ³
for all .-
Example 1.4 The vector space is the external direct sum of copies of , - -
that is,
-~ - Ä -^^
where there are summands on the right-hand side.
This construction can be generalized to any collection of vector spaces by
generalizing the idea that an ordered -tuple is just a function ² # Á Ã Á # ³
¢¸Áà Á¹ ¦ = ¸ÁÃÁ¹ from the to the union of the spaces index set
with the property that . ²³=
Vector Spaces 41
Definition Let be any family of vector spaces over . The<~¸ = 2¹ -
direct product of is the vector space<
HI d
2 2 =~ ¢ 2¦ = ² ³ =
thought of as a subspace of the vector space of all functions from to . 2=
It will prove more useful to restrict the set of functions to those with finite
support.
Definition Let be a family of vector spaces over . The<~¸ = 2¹ -
support of a function is the set ¢2¦ =
supp² ³~¸ 2² ³£ ¹
Thus, a function has if for all but a finite number of ² ³ ~ finite support
2 . The of the family is the vector space external direct sum <
HI d
2 2 ext=~ ¢ 2¦ = ² ³ = , has finite support
thought of as a subspace of the vector space of all functions from to . 2=
An important special case occurs when for all . If we let =~ = 2 =2
denote the set of all functions from to and denote the set of all 2= ² = ³2
functions in that have finite support, then =2
2 222
=~ = =~ ² = ³ andext
Note that the direct product and the external direct sum are the same for a finite
family of vector spaces.
Internal Direct Sums
An internal version of the direct sum construction is often more relevant.
Definition A vector space is the of a family = ()internal direct sum
<~¸ : 0¹ = of subspaces of , written
=~ =~ : <or
0
if the following hold:
42 Advanced Linear Algebra
1 is the sum join of the family :)( ) ( )Join of the family = <
=~ :
0
2 For each ,)( )Independence of the family 0
: q : ~ ¸¹
£ps
qt
In this case, each is called a of . If is a := ~ ¸ : Á Ã Á : ¹ direct summand <
finite family, the direct sum is often written
=~ :l Ä l :
Finally, if , then is called a of in . =~ :l ; ; : = complement
Note that the condition in part 2) of the previous definition is than stronger
saying simply that the member s of are pairwise disjoint:<
:q :~J
for all .£0
A word of caution is in order here: If and are subspaces of , then we may :; =
always say that the sum exists. However, to say that the direct sum of :b; :
and exists or to write is to imply that . Thus, while the; : l; : q; ~ ¸¹
sum of two subspaces always exists, the sum of two subspaces does not direct
always exist. Similar statements apply to families of subspaces of . =
The reader will be asked in a later chapter to show that the concepts of internal
and external direct sum are essentially equivalent isomorphic . For this reason, ()
the term “direct sum” is often used without qualification.
Once we have discussed the concept of a basis, the following theorem can be
easily proved.
Theorem 1.4 Any subspace of a vector space has a complement, that is, if is a :
subspace of , then there exists a subspace for which . =; = ~ : l ;
It should be emphasized that a subspace generally has many complements
()although they are isomorphic . The reader can easily find examples of this in
s.
We can characterize the uniqueness part of the definition of direct sum in other
useful ways. First a remark. If and are distinct subspaces of and if :; =
%Á& : q; %b& , then the sum can be thought of as a sum of vectors from the
Vector Spaces 43
same subspace (say ) or from different subspaces—one from and one from ::
;#. When we say that a vector cannot be written as a sum of vectors from the
distinct subspaces and , we mean that cannot be written as a sum :; # % b &
where and as coming from different subspaces, even if%& can be interpreted
they can also be interpreted as coming from the same subspace. Thus, if
%Á& : q; # ~ %b& # , then express as a sum of vectors from distinct does
subspaces.
Theorem 1.5 Let be a family of distinct subspaces of . The<~¸ : 0¹ =
following are equivalent:
1 For each ,)( )Independence of the family 0
: q : ~ ¸¹
£ps
qt
2 The zero vector cannot be written as a)( )Uniqueness of expression for
sum of nonzero vectors from distinct subspaces of . <
3 Every nonzero has a unique, except for)( )Uniqueness of expression #=
order of terms, expression as a sum
#~ bÄb
of nonzero vectors from distinct subspaces in . <
Hence, a sum
=~ :
0
is direct if and only if any one of 1 3 holds. )– )
Proof. Suppose that 2) fails, that is,
~ bÄb
where the nonzero 's are from distinct subspaces . Then and so :
c ~ bÄb
which violates 1). Hence, 1) implies 2). If 2) holds and
#~ bÄb #~! bÄb! and
where the terms are nonzero and the 's belong to distinct subspaces in and <
similarily for the 's, then !
~ bÄb c! cÄc!
By collecting terms from the same subspaces, we may write
~² c! ³bÄb² c! ³b bÄb c! cÄc! b b
44 Advanced Linear Algebra
Then 2) implies that and for all . Hence, 2) ~~ ~! "~ ÁÃÁ ""
implies 3).
Finally, suppose that 3) holds. If
£#:q :
£ps
qt
then and#~ :
~ b Ä b
where are nonzero. But this violates 3). :
Example 1.5 Any matrix can be written in the form (C
(~ ²(b(³b ²(c(³~)b*
!!()1.1
where is the transpose of . It is easy to verify that is symmetric and is(( ) *!
skew-symmetric and so 1.1 is a decomposition of as the sum of a symmetric () (
matrix and a skew-symmetric matrix.
Since the sets Sym and SkewSym of all symmetric and skew-symmetric
matrices in are subspaces of , we haveCC
C~bSym SkewSym
Furthermore, if , where and are symmetric and and :b;~: b; : : ; ;ZZ Z Z
are skew-symmetric, then the matrix
<~:c:~;c;ZZ
is both symmetric and skew-symme tric. Hence, provided that , we char²-³ £
must have and so and . Thus, <~ :~: ;~;ZZ
C~lSym SkewSym
Spanning Sets and Linear Independence
A set of vectors a vector space if every vector can be written as a linear spans
combination of some of the vectors in th at set. Here is the formal definition.
Definition The or by a nonempty set subspace spanned subspace generated ()
:= : of vectors in is the set of all linear combinations of vectors from :
º:» ~ ²:³ ~ ¸ # bÄb # -Á# :¹ span
Vector Spaces 45
When is a finite set, we use the notation or:~¸# ÁÃÁ# ¹ º# ÁÃÁ# »
span²# ÁÃÁ# ³ : = = = . A set of vectors in is said to , or , if span generate
=~ ² : ³ span .
It is clear that any superset of a spanning set is also a spanning set. Note also
that all vector spaces have spanning sets, since spans itself. =
Linear Independence
Linear independence is a fundamental concept.
Definition Let be a vector space. A nonempty set of vectors in is=: =
linearly independent if for any distinct vectors in , Á ÃÁ :
bÄb ~ ¬ ~ for all
In words, is linearly independent if the only linear combination of vectors :
from that is equal to is the tr ivial linear combination, all of whose :
coefficients are . If is not linear ly independent, it is said to be : linearly
dependent .
It is immediate that a linearly independe nt set of vectors cannot contain the zero
vector, since then violates th e condition of linear independence. h~
Another way to phrase the definition of lin ear independence is to say that is :
linearly independent if the zero vector ha s an “as unique as possible” expression
as a linear combination of vectors from . We can never prevent the zero vector :
from being written in the form , but we can prevent from ~ bÄb
being written in any other way as a lin ear combination of the vectors in . :
For the introspective reader, the expression has two ~ b²c ³
interpretations. One is where and , but this does ~ b ~ ~c
not involve distinct vectors so is not relevant to the question of linear
independence. The other interpretation is where ~ b! ! ~c £
(assuming that ). Thus, if is linearly independent, then cannot £ : :
contain both and . c
Definition Let be a nonempty set of vectors in . To say that a nonzero:=
vector is an linear combination of the vectors in is#= : essentially unique
to say that, up to order of terms, there is one and only one way to express as a #
linear combination
#~ bÄb
where the 's are distinct vectors in and the coefficients are nonzero. More :
explicitly, is an essentially unique linear combination of the vectors in #£ :
if and if whenever#º :»
46 Advanced Linear Algebra
#~ bÄb #~! bÄb ! and
where the 's are distinct, the 's are distinct and all coefficients are nonzero, !
then and after a reindexing of the 's if necessary, we have and~ ! ~
~! ~ÁÃÁ for all . Note that this is stronger than saying that (
~! .)
We may characterize linear independence as follows.
Theorem 1.6 Let be a nonempty set of vectors in . The following are: £ ¸¹ =
equivalent:
1 is linearly independent.):
2 Every nonzero vector is an essentially unique linear)s p a n # ² :³
combination of the vectors in . :
3 No vector in is a linear combination of other vectors in .) ::
Proof. Suppose that 1 holds and that )
£#~ bÄb ~! bÄb !
where the 's are distinct, the 's are distinct and the coefficients are nonzero. !
By subtracting and grouping 's and 's that are equal, we can write !
~² c ³ bÄb² c ³
b bÄb
c ! cÄc !
b b
b b
and so 1 implies that and and for all . ) ~~ ~ ~! ~ ÁÃÁ "" ""
Thus, 1 implies 2 . ))
If 2) holds and can be written as :
~ bÄb
where are different from , then we may collect like terms on the right :
and then remove all terms with coefficient. The resulting expression violates
2). Hence, 2) implies 3). If 3) holds and
bÄb ~
where the 's are distinct and , then and we may write £
~c ² bÄb ³
which violates 3 . )
The following key theorem relates the notions of spanning set and linear
independence.
Vector Spaces 47
Theorem 1.7 Let be a set of vectors in . The following are equivalent::=
1 is linearly independent and spans .):=
2 Every nonzero vector is an essentially unique linear combination of) #=
vectors in . :
3 is a minimal spanning set, that is , spans but any proper subset of ):: = :
does not span . =
4 is a maximal linearly independent se t, that is, is linearly independent, )::
but any proper superset of is not linearly independent. :
A set of vectors in that satisfies any and hence all of these conditions is = ()
called a for . basis =
Proof. We have seen that 1 and 2 are equivalent. Now suppose 1 holds. Then )) )
:: : = is a spanning set. If some proper subset of also spanned , then anyZ
vector in would be a linear combination of the vectors in , :c: :Z Z
contradicting the fact that the vectors in are linearly independent. Hence 1 : )
implies 3 . )
Conversely, if is a minimal spanning se t, then it must be linearly independent. :
For if not, some vector would be a linear combination of the other vectors :
in and so would be a proper spanning subset of , which is not:: c ¸ ¹ :
possible. Hence 3 implies 1 . ))
Suppose again that 1 holds. If were not maximal, there would be a vector ) :
#=c: :r¸ # ¹ # for which the set is linearly independent. But then is not
in the span of , contradicting the fact that is a spanning set. Hence, is a :: :
maximal linearly independent set and so 1 implies 4 . ))
Conversely, if is a maximal linearly i ndependent set, then must span , for :: =
if not, we could find a vector that is not a linear combination of the #=c:
vectors in . Hence, would be a linearly independent proper superset of :: r ¸ # ¹
:, which is a contradiction. Thus, 4 implies 1 . ))
Theorem 1.8 A finite set of vectors in is a basis for if :~¸# ÁÃÁ# ¹ = =
and only if
= ~ º# »lÄlº# »
Example 1.6 The th in is the vector that has 's in all - standard vector
coordinate positions except the th, where it has a . Thus,
~²ÁÁÃÁ³Á ~²ÁÁÃÁ³ ÁÃÁ ~²ÁÃÁÁ³
The set is called the for .¸ ÁÃÁ ¹ -standard basis
The proof that every nontrivial vector space has a basis is a classic example of
the use of Zorn's lemma.
48 Advanced Linear Algebra
Theorem 1.9 Let be a nonzero vector space. Let be a linearly independent=0
set in and let be a spanning set in containing . Then there is a basis =: = 0 8
for for which . In particular,=0 : 8
1 Any vector space, except the zero space , has a basis.) ¸¹
2 Any linearly independent set in is contained in a basis.) =
3 Any spanning set in contains a basis.) =
Proof. Consider the collection of all linearly independent subsets of 7 =
containing and contained in . This collection is not empty, since . 0: 0 7
Now, if
9~¸ 0 2¹
is a chain in , then the union 7
<~ 0
2
is linearly independent and satisfies , that is, . Hence, every 0<: < 7
chain in has an upper bound in and according to Zorn's lemma, must77 7
contain a maximal element , which is linearly independent. 8
Now, is a basis for the vector space , for if any is not a linear8 º:» ~ = :
combination of the elements of , then is linearly independent,88 r¸ ¹:
contradicting the maximality of . Hence and so . 88 8 : º» =~ º : » º»
The reader can now show, using Theorem 1.9, that any subspace of a vector
space has a complement.
The Dimension of a Vector Space
The next result, with its classical elegant proof, says that if a vector space has =
a spanning set , then the size of any linearly independent set cannotfinite :
exceed the size of . :
Theorem 1.10 Let be a vector space and assume that the vectors = # ÁÃÁ#
are linearly independent and the vectors span . Then . Á ÃÁ =
Proof. First, we list the two sets of v ectors: the spanning set followed by the
linearly independent set:
ÁÃÁ Â# ÁÃÁ#
Then we move the first vector to the front of the first list: #
#Á Á Ã Á Â #Á Ã Á #
Since span , is a linear combin ation of the 's. This implies that Á ÃÁ = #
we may remove one of the 's, which by reindexing if necessary can be ,
from the first list and still have a spanning set
#Á Á Ã Á Â #Á Ã Á #
Vector Spaces 49
Note that the first set of vectors still spans and the second set is still linearly =
independent.
Now we repeat the process, moving from the second list to the first list #
#Á #Á Á Ã Á Â #Á Ã Á #
As before, the vectors in the first list ar e linearly dependent, since they spanned
=# # before the inclusion of . However, since the 's are linearly independent,
any nontrivial linear combination of the v ectors in the first list that equals
must involve at least one of the 's. Hence, we may remove that vector, which
again by reindexing if necessary may be taken to be and still have a spanning
set
#Á #Á Á Ã Á Â #Á Ã Á #
Once again, the first set of vectors spans and the second set is still linearly =
independent.
Now, if , then this process will ev entually exhaust the 's and lead to the
list
#Á #Á Ã Á # Â # Á Ã Á # b
where span , which is clearly not possible since is not in the#Á #Á Ã Á # = #
span of . Hence, .#Á #Á Ã Á #
Corollary 1.11 If has a spanning set, then any two bases of have the== finite
same size.
Now let us prove the analogue of Corollary 1.11 for arbitrary vector spaces.
Theorem 1.12 If is a vector space, then any two bases for have the same==
cardinality.
Proof. We may assume that all bases for ar e infinite sets, for if any basis is=
finite, then has a finite spanning set and so Corollary 1.11 applies. =
Let be a basis for and let be another basis for . Then any89~¸ 0¹ = =
vector can be written as a finite lin ear combination of the vectors in , 98
where all of the coefficients are nonzero, say
~
<
But because is a basis, we must have 9
9<~ 0
50 Advanced Linear Algebra
for if the vectors in can be expressed as finite linear combinations of the 9
vectors in a subset of , then spans , which is not the case. proper 88 8ZZ=
Since for all , Theorem 0.17 implies that ((< L 9
( ( (( (( ((89 9~0 L ~
But we may also reverse the roles of and , to conclude that and so 89 98 (( ( (
the Schro ¨der–Bernstein theorem implies that . (( ( (89~
Theorem 1.12 allows us to make the following definition.
Definition A vector space is if it is the zero space , or = ¸¹ finite-dimensional
if it has a finite basis. All other vector spaces are . The infinite-dimensional
dimension of the zero space is and the of any nonzero vector dimension
space is the cardinality of any basis for . If a vector space has a basis of== =
cardinality , we say that is and write . =² = ³ ~-dimensional dim
It is easy to see that if is a subspace of , then . If in : = ²:³ ²= ³ dim dim
addition, , then . dim dim²:³ ~ ²= ³ B : ~ =
Theorem 1.13 Let be a vector space. =
1 If is a basis for and if and , then)8 888 88 =~ r q ~ J
=~ º » l º »88
2 Let . If is a basis for and is a basis for , then)=~ :l ; : ; 88
88 888 q~ J ~r = and is a basis for .
Theorem 1.14 Let and be subspaces of a vector space . Then :; =
dim dim dim dim²:³b ²;³ ~ ²: b;³b ²: q;³
In particular, if is any complement of in , then ;: =
dim dim dim²:³b ²;³ ~ ²= ³
that is,
dim dim dim²: l;³ ~ ²:³b ²;³
Proof. Suppose that is a basis for . Extend this to a basis 8~¸ 0¹ :q;
78 7 8 8r: ~ ¸ 1 ¹ for where is disjoint from . Also, extend to a
basis for where is disjoint from . We claim that89 9 8r; ~ ¸ 2 ¹
789 789rr : b ; ºrr » ~ : b ; is a basis for . It is clear that .
To see that is linearly independent, suppose to the contrary that789rr
Vector Spaces 51
#b Ä b #~
where and for all . There must be vectors in this# r r £ # 789
expression from both and , since and are linearly independent. 79 7 88 9 rr
Isolating the terms involving the vectors from on one side of the equality 7
shows that there is a nonzero vector in . But then %º »qº r » %:q;78 9
and so , which implies that , a contradiction. Hence, %º »qº » %~78
789rr : b ; is linearly independent and a basis for .
Now,
dim dim
dim
dim dim²:³b ²;³ ~ r b r
~bbb
~bb b² : q ; ³
~ ² :b;³b ² :q;³(( ( (
(( (( (( ( (
(( (( ( (78 89
7889
789
as desired.
It is worth emphasizing that while the equation
dim dim dim dim²:³b ²;³ ~ ²: b;³b ²: q;³
holds for all vector spaces, we cannot write
dim dim dim dim²: b;³ ~ ²:³b ²;³c ²: q;³
unless is finite-dimensional.:b;
Ordered Bases and Coordinate Matrices
It will be convenient to consider bases that have an order imposed on their
members.
Definition Let be a vector space of dimension . An for is= = ordered basis
an ordered -tuple of vectors for which the set is a ²# ÁÃÁ# ³ ¸# ÁÃÁ# ¹
basis for . =
If is an ordered basis for , then for each there is a8~² #ÁÃÁ#³ = #=
unique ordered -tuple of scalars for which ² Á Ã Á ³
#~# bÄb #
Accordingly, we can define the by coordinate map 8¢= ¦-
88²#³ ~ ´#µ ~
Å
vy
wz
()1.3
52 Advanced Linear Algebra
where the column matrix is known as the of with ´#µ #8 coordinate matrix
respect to the ordered basis . Clearly , knowing is equivalent to knowing 8 ´#µ #8
()assuming knowledge of . 8
Furthermore, it is easy to see that the coordinate map is bijective and 8
preserves the vector space operations, that is,
88 8²# bÄb # ³~ ²#³bÄb ²# ³
or equivalently
´# bÄb # µ ~´#µ bÄb ´# µ 88 8
Functions from one vector space to another that preserve the vector space
operations are called and form the objects of study in the linear transformations
next chapter.
The Row and Column Spaces of a Matrix
Let be an matrix over . The rows of span a subspace of known( d - ( -
as the of and the columns of span a subspace of known as row space ((-
the of . The dimensions of these spaces are called the column space row rank (
and , respectively. We denote the row space and row rank by column rank
rs rrk cs crk²(³ ²(³ ²(³ ²(³ and and the column space and column rank by and .
It is a remarkable and useful fact that the row rank of a matrix is always equal to
its column rank, despite the fact that if , the row space and column space £
are not even in the same vector space!
Our proof of this fact hinges on the following simple observation about
matrices.
Lemma 1.15 Let be an matrix. Then elementary column operations do( d
not affect the row rank of . Similarly, elementary row operations do not affect (
the column rank of . (
Proof. The second statement follows from the first by taking transposes. As to
the first, the row space of is (
rs²(³ ~ º (ÁÃÁ (»
where are the standard basis vectors in . Performing an elementary-
column operation on is equivalent to multiplying on the right by an ((
elementary matrix . Hence the row space of is ,( ,
rs²(,³~º (,ÁÃÁ (,»
and since is invertible, ,
Vector Spaces 53
rrk rs rs rrk²(³ ~ ² ²(³³ ~ ² ²(,³³ ~ ²(,³ dim dim
as desired.
Theorem 1.16 If , then . This number is called the ( ² ( ³~ ² ( ³CÁ rrk crk
rank of and is denoted by .(² ( ³ rk
Proof. According to the previous lemma, we may reduce to reduced column (
echelon form without affecting the row rank. But this reduction does not affect
the column rank either. Then we may further reduce to reduced row echelon (
form without affecting either rank. The resulting matrix has the same row 4
and column ranks as . But is a ma trix with 's followed by 's on the main (4
diagonal entries and 's elsewhere. Hence, ()4Á 4Á Ã Á Á
rrk rrk crk crk²(³ ~ ²4³ ~ ²4³ ~ ²(³
as desired.
The Complexification of a Real Vector Space
If is a complex vector space that is, a vector space over , then we can> () d
think of as a real vector space simply by restricting all scalars to the field .> s
Let us denote this real vector space by and call it the of . >>s real version
On the other hand, to each real vector space , we can associate a complex =
vector space . This “complexification” process will play a useful role when =d
we discuss the structure of linear operators on a real vector space. Throughout (
our discussion will denote a real vector space. = )
Definition If is a real vector space, then the set of ordered== ~ = d =d
pairs, with componentwise addition
²"Á#³b²%Á&³~²"b%Á#b&³
and scalar multiplication over defined by d
²b³²"Á#³ ~ ²"c#Á#b"³
for is a complex vector space, called the of .Á =s complexification
It is convenient to introduce a notati on for vectors in that resembles the =d
notation for complex numbers. In particular, we denote by ²"Á#³ = "b#d
and so
=~ ¸ " b # " Á # = ¹d
Addition now looks like ordinary addition of complex numbers,
²"b#³b²%b&³~²"b%³b²#b&³
and scalar multiplication looks like ordinary multiplication of complex numbers,
54 Advanced Linear Algebra
²b³²"b#³~²"c#³b²#b"³
Thus, for example, we immediately have for , Á s
²"b#³ ~ "b#
²"b#³ ~ c#b"
²b³"~"b"
²b³#~c#b#
The of is and the of is . real part imaginary part '~"b# "= ' #=
The essence of the fact that is really an ordered pair is that is '~"b# = 'd
if and only if its real and imaginary parts are both .
We can define the by complexification map cpx¢= ¦=d
cpx²#³ ~ #b
Let us refer to as the , or of . #b #= complexification complex version
Note that this map is a group homomorphism, that is,
cpx cpx cpx cpx²³ ~ b ²"f#³ ~ ²"³f ²#³ and
and it is injective:
cpx cpx²"³ ~ ²#³ ¯ " ~ #
Also, it preserves multiplication by scalars: real
cpx cpx² " ³~ "b ~ ² "b ³~ ² " ³
for . However, the complexification map is not surjective, since it givess
only “real” vectors in . =d
The complexification map is an injective linear transformation defined in the (
next chapter from the real vector space to the real version of the ) =² = ³ds
complexification , that is, to the complex vector space provided that ==dd
scalars are restricted to real numbers. In this way, we see that contains an =d
embedded copy of . =
The Dimension of =d
The vector-space dimensions of and are the same. This should not ==d
necessarily come as a surprise because although may seem “bigger” than , ==d
the field of scalars is also “bigger.”
Theorem 1.17 If is a basis for over , then the8s~¸ # 0¹ =
complexification of ,8
cpx²³ ~ ¸ #b # ¹88
Vector Spaces 55
is a basis for the vector space over . Hence, =dd
dim dim²= ³ ~ ²= ³d
Proof. To see that spans over , let . Then cpx²³ = % b & = % Á & =8ddd
and so there exist real numbers and some of which may be for which ()
%b&~ # b #
~² # b # ³
~ ² b ³²# b³@A
~ ~11
~1
~1
To see that is linearly independent, if cpx²³8
~1
² b³²# b³~b
then the previous computations show that
~ ~11
# ~ # ~ and
The independence of then implies that and for all . 8 ~ ~
If and is a basis for , then we may write#= ~¸ # 0¹ =8
#~ #
~
for . Since the coefficients are real, we haves
#b~ ²# b³
~
and so the coordinate matrices are equal:
´#bµ ~ ´#µ cpx²³88
Exercises
1. Let be a vector space over . Prove that and for all = - #~ ~ #=
and . Describe the different 's in these equations. Prove that if-
#~ ~ #~ #~# #~ ~ , then or . Prove that implies that or .
56 Advanced Linear Algebra
2. Prove Theorem 1.3.
3. a Find an abelian group and a field for which is a vector space ) =- =
over in at least two different ways, that is, there are two different-
definitions of scalar multiplication making a vector space over . =-
b Find a vector space over and a subset of that is 1 a )( ) =- : =
subspace of and 2 a vector space using operations that differ from = ()
those of .=
4. Suppose that is a vector space with basis and is a =~ ¸ 0 ¹ : 8
subspace of . Let be a partition of . Then is it true that =¸ ) Á Ã Á ) ¹ 8
:~ ² :qº )» ³
~
What if for all ?: qº) » £ ¸¹
5. Prove Theorem 1.8.
6. Let . Show that if , then:Á;Á< ²=³ < :I
: q²; b<³ ~ ²: q;³b<
This is called the for the lattice . modular law I²= ³
7. For what vector spaces does the distributive law of subspaces
: q²; b<³ ~ ²: q;³b²: q<³
hold?
8. A vector is called if for all #~² ÁÃÁ ³ s strongly positive
~ ÁÃÁ .
a Suppose that is strongly positive. Show that any vector that is “close ) #
enough” to is also strongly positive. Formulate carefully what “close # (
enough” should mean.)
b Prove that if a subspace of contains a strongly positive vector, ) :s
then has a basis of strongly positive vectors.:
9. Let be an matrix whose rows are linearly independent. Suppose4 d
that the columns of span the column space of . Let be ÁÃÁ 4 4 *
the matrix obtained from by deleting all columns except .4 ÁÃÁ
Show that the rows of are also linearly independent. *
10. Prove that the first two statements in Theorem 1.7 are equivalent.
11. Show that if is a subspace of a vector space , then . : = ²:³ ²= ³ dim dim
Furthermore, if then . Give an example to dim dim²:³ ~ ²= ³ B : ~ =
show that the finiteness is required in the second statement.
12. Let and suppose that . What can you dim² =³B =~<l: ~<l:
say about the relationship between and ? What can you say if ::
: : ?
13. What is the relationship between and ? Is the direct sum :l; ;l:
operation commutative? Formulate and prove a similar statement
concerning associativity. Is there an “identity” for direct sum? What about
“negatives”?
Vector Spaces 57
14. Let be a finite-dimensional vector space over an infinite field . Prove=-
that if are subspaces of of equal dimension, then there is a:ÁÃÁ: =
subspace of for which for all . In other words, ; = = ~: l; ~ÁÃÁ
;: is a common complement of the subspaces .
15. Prove that the vector space of all continuous functions from to is 9s s
infinite-dimensional.
16. Show that Theorem 1.2 need not hol d if the base field is finite. -
17. Let be a subspace of . The set is called an := # b : ~ ¸ # b : ¹
affine subspace of .=
a Under what conditions is an affine subspace of a subspace of ? ) ==
b Show that any two affine subspaces of the form and are ) #b: $b:
either equal or disjoint.
18. If and are vector spaces over for which , then does it=> - = ~ > (( ( (
follow that ? dim dim²= ³ ~ ²>³
19. Let be an -dimensional real vector space and suppose that is a = :
subspace of with . Define an equivalence relation on =² : ³ ~ c dim
the set by if the “line segment”=±: #$
3²#Á$³ ~ ¸#b²c³$ ¹
has the property that . Prove that is an equivalence 3²#Á$³q: ~ J
relation and that it has exactly two equivalence classes.
20. Let be a field. A of is a subset of that is a field in its-- 2 - subfield
own right using the same operations as defined on . -
a Show that is a vector space over any subfield of . ) -2 -
b Suppose that is an -dimensional vector space over a subfield of ) - 2
-= - =. If is an -dimensional vector space over , show that is also a
vector space over . What is the dimension of as a vector space 2=
over ?2
21. Let be a finite field of size a nd let be an -dimensional vector space - =
over . The purpose of this exercise is to show that the number of-
subspaces of of dimension is =
45 ² c³Ä² c³
² c³Ä² c³² c³Ä² c³~
c
The expressions are called and have properties ²³
Gaussian coefficients
similar to those of the binomial coefficients. Let be the number of :²Á³
=-dimensional subspaces of .
a Let be the number of -tuples of linearly independent vectors )5²Á³
²# ÁÃÁ# ³ = in . Show that
5²Á³ ~ ² c³² c³Ä² c ³ c
b Now, each of the -tuples in a can be obtained by first choosing a ))
subspace of of dimension and then selecting the vectors from this =
subspace. Show that for any -dimensional subspace of , the number =
58 Advanced Linear Algebra
of -tuples of independent vectors in this subspace is
² c³² c³Ä² c ³ c
c Show that )
5²Á³ ~ :²Á³² c³² c³Ä² c ³ c
How does this complete the proof?
22. Prove that any subspace of is a closed set or, equivalently, that its set :s
complement is open, that is, for any there is an open :~ ± : % : s
ball centered at with radius for which .)²%Á ³ % )²%Á ³ :
23. Let and be bases for a vector space .89~¸ ÁÃÁ¹ ~¸ ÁÃÁ¹ =
Let . Show that there is a permutation of such c ¸ÁÃÁ¹
that
ÁÃÁ Á ÁÃÁ ²b³ ²³
and
ÁÃÁ Á ÁÃÁ²³ ²³ b
are both bases for . : You may use the fact that if is an invertible =4Hint
d matrix and if , then it is possible to reorder the rows so
that the upper left submatrix and the lower right d ²c³d²c³
submatrix are both invertible. This follows, for example, from the general (
Laplace expansion theorem for determinants.)
24. Let be an -dimensional vector space over an infinite field and = -
suppose that are subspaces of with . Prove :ÁÃÁ: = ² :³ dim
that there is a subspace of of dimension for which ;= c
; q: ~ ¸¹ for all .
25. What is the dimension of the co mplexification thought of as a real =d
vector space?
26. When is a subspace of a complex vector space a complexification? Let () =
be a real vector space with complexification and let be a subspace of =<d
=: =d. Prove that there is a subspace of for which
<~: ~¸ b! Á !: ¹d
if and only if is closed under complex conjugation defined <¢ = ¦ = dd
by .²"b#³~"c#
Chapter 2
Linear Transformations
Linear Transformations
Loosely speaking, a linear transformation is a function from one vector space to
another that the vector space operations. Let us be more precise. preserves
Definition Let and be vector spaces over a field . A function => - ¢ = ¦ >
is a iflinear transformation
²"b #³ ~ ²"³b ²#³
for all scalars and vectors , . The set of all linear Á - " # =
transformations from to is denoted by . => ² = Á > ³ B
1 A linear transformation from to is called a on . The) == = linear operator
set of all linear operators on is denoted by . A linear operator on a =² = ³ B
real vector space is called a and a linear operator on a real operator
complex vector space is called a . complex operator
2 A linear transformation from to the base field thought of as a vector)( =-
space over itself is called a on . The set of all linear ) linear functional =
functionals on is denoted by and called the of . == =idual space
We should mention that some authors use the term linear operator for any linear
transformation from to . Also, the application of a linear transformation =>
on a vector is denoted by or by , parentheses being used when #² # ³ #
necessary, as in , or to improve readability, as in rather than ²"b#³ ² "³
²² " ³ ³ .
Definition The following terms are also employed:
1 for linear transformation)homomorphism
2 for linear operator)endomorphism
3 or for injective lin ear transformation )( )monomorphism embedding
4 for surjective lin ear transformation )epimorphism
5 for bijective linear transformation.)isomorphism
60 Advanced Linear Algebra
6 for bijective linear operator.)automorphism
Example 2.1
1 The derivative is a linear operator on the vector space of all) +¢= ¦ = =
infinitely differentiable functions on . s
2 The integral operator defined by) ¢-´%µ¦-´%µ
~ ² ! ³ !
%
is a linear operator on . -´%µ
3 Let be an matrix over . The function defined by)( d - ¢ - ¦ - (
(#~( # , where all vectors are written as column vectors, is a linear
transformation from to . This function is just multiplication by . -- (
4 The coordinate map of an -dimensional vector space is a) ¢= ¦-
linear transformation from to . =-
The set is a vector space in its own right and has the structure ofBB²= Á>³ ²= ³
an algebra, as defined in Chapter 0.
Theorem 2.1
1 The set is a vector space under ordinary addition of functions) B²= Á>³
and scalar multiplication of functions by elements of . -
2 If and , then the composition is in .)B B B² < Á = ³ ² = Á > ³ ² < Á > ³
3 If is bijective then .)B B² = Á > ³ ² > Á = ³c
4 The vector space is an algebra, where multiplication is composition) B²= ³
of functions. The identity map is the multiplicative identity andB² = ³
the zero map is the additive identity. ² =³B
Proof. We prove only part 3 . Let be a bijective linear )¢= ¦>
transformation. Then is a well-defined function and since any two c¢> ¦=
vectors and in have the form and , we have$$ > $ ~ #$ ~ #
c c
c
c c
²$ b$³~ ² # b #³
~ ² ²# b# ³³
~ # b #
~ ² $³b ² $³
which shows that is linear. c
One of the easiest ways to define a linear transformation is to give its values on
a basis. The following theorem says that we may assign these values arbitrarily
and obtain a unique linear transformati on by linear extension to the entire
domain.
Theorem 2.2 Let and be vector spaces and let be a => ~ ¸ # 0 ¹ 8
basis for . Then we can define a linear transformation by = ² = Á > ³ B
Linear Transformations 61
specifying the values of for all and extending to by 8 ## =arbitrarily
linearity, that is,
²# bÄb # ³~ # bÄb #
This process defines a unique linear transformation, that is, if BÁ² = Á > ³
satisfy for all then . 8 #~ # # ~
Proof. The crucial point is that the extension by linearity is well-defined, since
each vector in has an essentially unique representation as a linear =
combination of a finite number of vecto rs in . We leave the details to the8
reader.
Note that if and if is a subspace of , then the restriction of B ² = Á > ³ : = O :
to is a linear transformation from to .:: >
The Kernel and Image of a Linear Transformation
There are two very important vector spaces associated with a linear
transformation from to . =>
Definition Let . The subspaceB² = Á > ³
ker² ³~¸ #= #~ ¹
is called the of and the subspace kernel
im² ³~¸ ##=¹
is called the of . The dimension of is called the of and is image nullity ker²³
denoted by . The dimension of is called the of and is null²³ ²³ im rank
denoted by . rk²³
It is routine to show that is a subspace of and is a subspace of ker²³ = ²³ im
>. Moreover, we have the following.
Theorem 2.3 Let . ThenB² = Á > ³
1 is surjective if and only if )i m ²³ ~ >
2 is injective if and only if ) ker² ³ ~ ¸¹
Proof. The first statement is merely a restatement of the definition of
surjectivity. To see the validity of the second statement, observe that
"~ #¯ ² "c# ³~¯"c# ² ³ ker
Hence, if , then , which shows that is injective. ker² ³~¸ ¹ "~ #¯"~#
Conversely, if is injective and , then and so . This " ² ³ "~ "~ker
shows that . ker² ³ ~ ¸¹
62 Advanced Linear Algebra
Isomorphisms
Definition A bijective linear transformation is called an ¢= ¦>
isomorphism from to . When an isomorphism from to exists, we say=> =>
that and are and write .=> = > isomorphic
Example 2.2 Let . For any ordered basis of , the coordinate dim²= ³ ~ = 8
map that sends each vector to its coordinate matrix8¢= ¦- #=
´#µ - -8 is an isomorphism. Hence, any -dimensional vector space over is
isomorphic to . -
Isomorphic vector spaces share many properties, as the next theorem shows. If
B² = Á > ³ : = and we write
:~¸ : ¹
Theorem 2.4 Let be an isomorphism. Let . ThenB² = Á > ³ : =
1 spans if and only if spans .):= :>
2 is linearly independent in if and only if is linearly independent in):= :
>.
3 is a basis for if and only if is a basis for .):= :>
An isomorphism can be characterized as a linear transformation that ¢= ¦>
maps a basis for to a basis for . =>
Theorem 2.5 A linear transformation is an isomorphism if and B² = Á > ³
only if there is a basis for for which is a basis for . In this case, 8 8 =>
maps any basis of to a basis of . =>
The following theorem says that, up to isomorphism, there is only one vector
space of any given dimension over a given field.
Theorem 2.6 Let and be vector spaces over . Then if and only=> - = >
if .dim dim²= ³ ~ ²>³
In Example 2.2, we saw that any -dimensional vector space is isomorphic to
-) ² - ³ ) . Now suppose that is a set of cardinality and let be the vector
space of all functions from to with finite support. We leave it to the reader )-
to show that the functions defined for all by )² - ³ )
²%³ ~% ~
% £ Fif
if
form a basis for , called the . Hence, . ²- ³ ²²- ³ ³ ~ ))) standard basis dim ((
It follows that for any cardinal number , there is a vector space of dimension .
Also, any vector space of dimension is isomorphic to . ²- ³)
Linear Transformations 63
Theorem 2.7 If is a natural number, then any -dimensional vector space
over is isomorphic to . If is any cardinal number and if is a set of-- )
cardinality , then any -dimensional vector space over is isomorphic to the -
vector space of all functions from to with finite support. ²- ³ ) -)
The Rank Plus Nullity Theorem
Let . Since any subspace of has a complement, we can writeB² = Á > ³ =
=~ ²³ l ²³ ker ker
where is a complement of in . It follows that ker ker²³ ²³ =
dim dim ker dim ker²= ³ ~ ² ² ³³b ² ² ³ ³
Now, the restriction of to , ker²³
¢² ³ ¦ >ker
is injective, since
ker ker ker² ³ ~ ² ³q ² ³ ~ ¸¹
Also, . For the reverse inclusion, if , then since im im im²³ ² ³ # ² ³
# ~ " b $" ² ³ $ ² ³ for and , we have ker ker
#~ "b $~ $~ $ ² ³im
Thus . It follows that im im²³ ~ ² ³
ker²³ ²³im
From this, we deduce the following theorem.
Theorem 2.8 Let .B² = Á > ³
1 Any complement of is isomorphic to )i m ker²³ ²³
2)( )The rank plus nullity theorem
dim ker dim dim²² ³ ³ b ² ² ³ ³ ~ ² = ³ im
or, in other notation,
rk null²³ b ²³ ~ ² = ³ dim
Theorem 2.8 has an important corollary.
Corollary 2.9 Let , where . Then isB ² =Á>³ ² =³~ ² >³B dim dim
injective if and only if it is surjective.
Note that this result fails if the vector spaces are not finite-dimensional. The
reader is encouraged to find an example to support this statement.
64 Advanced Linear Algebra
Linear Transformations from to --
Recall that for any matrix over the multiplication map d ( -
(²#³ ~ (#
is a linear transformation. In fact, any linear transformation has B² - Á -³
this form, that is, is just multiplication by a matrix, for we have
23 23 Ä ~ Ä ~ ²³
and so , where~(
(~ Ä 23
Theorem 2.10
1 If is an matrix over then .)( d - ² - Á - ³ B(
2 If then , where)B ² - Á -³ ~(
(~² Ä ³
The matrix is called the of . ( matrix
Example 2.3 Consider the linear transformation defined by ¢- ¦-
²%Á&Á'³~²%c&Á'Á%b&b'³
Then we have, in column form,
vy v y v y vy
wz w z w z wz% %c& c %
& ' &
' %b&b' '~~
and so the standard matrix of is
(~c
vy
wz
If , then since the image of is the column space of , we have( (CÁ (
dim ker dim²² ³ ³ b² ( ³ ~ ² - ³(rk
This gives the following useful result.
Theorem 2.11 Let be an matrix over . ( d -
1 is injective if and only if n.)r k(¢- ¦- ²(³~
2 is surjective if and only if m. )r k(¢- ¦- ²(³~
Linear Transformations 65
Change of Basis Matrices
Suppose that and are ordered bases for a 89~² ÁÃÁ³ ~² ÁÃÁ³
vector space . It is natural to ask how the coordinate matrices and are =´ # µ ´ # µ 89
related. Referring to Figure 2.1,
VFn
FnIB
ICIC(IB)-1
Figure 2.1
the map that takes to is and is called the ´#µ ´#µ ~89 8 9 9 8 Ácchange of basis
operator change of coordinates operator or . Since is an operator on() 89Á
-( , it has the form , where
(~² ² ³Ä ² ³ ³
~ ² ²´ µ ³ Ä ²´ µ ³³
~ ²´ µ Ä ´ µ ³³
89 89
98 9888
99Á Á
c c
We denote by and call it the from to . (489, change of basis matrix 89
Theorem 2.12 Let and be ordered bases for a vector space89~² ÁÃÁ³
=~ -. Then the change of basis operator is an automorphism of , 89 9 8 Ác
whose standard matrix is
4 ~² ´ µ Ä´ µ³ ³89 9 9,
Hence
´#µ ~ 4 ´#µ98 9 8 Á
and .4~ 498 89 Ác
,
Consider the equation
(~489Á
or equivalently,
(~² ´ µ Ä´ µ³ ³ 99
Then given any two of an invertible matrix an ordered basis for ( d Á() ( 8
--)( ) and an ordered basis for , the third component is uniquely9
determined by this equation. This is clear if and are given or if and are 89 9 (
66 Advanced Linear Algebra
given. If and are given, then there is a unique for which and (( ~ 489cÁ98
so there is a unique for which . 9 (~489Á
Theorem 2.13 If we are given any two of the following:
1 an invertible matrix ) d (
2 an ordered basis for ) 8-
3 an ordered basis for .) 9-
then the third is uniquely determined by the equation
(~489Á
The Matrix of a Linear Transformation
Let be a linear transformation, where and¢= ¦> ²=³~ dim
dim²>³ ~ ~ ² ÁÃÁ ³ = and let be an ordered basis for and an89
ordered basis for . Then the map >
¢´#µ ¦ ´ #µ89
is a of as a linear transformation from to , in the senserepresentation --
that knowing along with and , of c ourse is equivalent to knowing . Of 8 9 ()
course, this representation depends on the choice of ordered bases and . 89
Since is a linear transformation fro m to , it is just multiplication by an --
d ( matrix , that is,
´# µ~ ( ´ # µ98
Indeed, since , we get the columns of as follows: ´ µ ~ (8
(~ ( ~ ( ´ # µ ~ ´ µ²³
89
Theorem 2.14 Let and let and be orderedB 8 9² = Á > ³ ~ ² Á Ã Á ³
bases for and , respectively. Then can be represented with respect to => 8
and as matrix multiplication, that is,9
´# µ~ ´µ ´ # µ98 9 8 ,
where
´µ ~ ²89,´µ Ä ´µ³99
is called the and . When and matrix of with respect to the bases8 9 =~ >
89 ~´ µ ´ µ, we denote by and so 88 8,
´# µ ~ ´µ´ # µ88 8
Example 2.4 Let be the derivative operator, defined on the vector+¢ ¦FF
space of all polynomials of degree at most . Let . Then ~ ~ ² Á % Á % ³89
Linear Transformations 67
´ + ² ³ µ~ ´ µ~ ´ + ² % ³ µ~ ´ µ~ Á ´ + ² % ³ µ~ ´ % µ~
99 99 9 9vy vy vy
wz wz wz,
and so
´+µ ~
8vy
wz
Hence, for example, if , then ²%³~ b%b%
´+²%³µ ~ ´+µ ´²%³µ ~ ~
988v y vy vy
w z wz wz
and so .+²%³ ~ b%
The following result shows that we may work equally well with linear
transformations or with the matrices th at represent them with respect to fixed (
ordered bases and . This applies not only to addition and scalar 89 )
multiplication, but also to matrix multiplication.
Theorem 2.15 Let and be finite-dimensional vector spaces over , with => -
ordered bases and , respectively. 89~² ÁÃÁ³ ~² ÁÃÁ ³
1 The map defined by) B C¢² = Á > ³ ¦ ² - ³ Á
²³ ~ ´µ 89,
is an isomorphism and so . Hence, BC²= Á>³ ²-³ Á
dim dim² ² =Á>³ ³~ ² ² -³ ³~dBC Á
2 If and and if , and are ordered bases for)B B 8 9 :² < Á = ³ ² = Á > ³
<= >, and , respectively, then
´µ ~ ´ µ´ µ 8: 9: 89,, ,
Thus, the matrix of the product composition is the product of the ()
matrices of and . In fact, this is the prima ry motivation for the definition
of matrix multiplication.
Proof. To see that is linear, observe that for all ,
´ b! µ ´ µ ~ ´² b! ³² ³µ
~´ ² ³b! ² ³ µ
~ ´ ² ³µ b!´ ² ³µ
~ ´ µ ´ µ b! ´ µ ´ µ
~² ´ µ b! ´ µ ³ ´ µ
89 8 9
9
99
89 8 89 8
89 89 8Á
Á Á
ÁÁ
68 Advanced Linear Algebra
and since is a standard basis vector, we conclude that ´ µ ~ 8
´ b! µ ~ ´ µ b!´ µ 89 89 89ÁÁ Á
and so is linear. If , we define by the condition ,C ( ´ µ ~( Á ²³9
whence and is surjective. Also, since ² ³~( ² ³~¸ ¹ ´ µ ~ ker 8
implies that . Hence, the map is an isomorphism. To prove part 2 , we ~ )
have
´ µ ´ # µ ~ ´²# ³ µ ~ ´µ ´# µ~ ´µ ´µ ´ # µ 8: 8 : 9: 9 9: 89 8ÁÁ Á ,
Change of Bases for Linear Transformations
Since the matrix that represents depends on the ordered bases and , it ´µ 8 989,
is natural to wonder how to choose these bases in order to make this matrix as
simple as possible. For instance, can we always choose the bases so that is
represented by a diagonal matrix?
As we will see in Chapter 7, the answer to this question is no. In that chapter,
we will take up the general question of how best to represent a linear operator
by a matrix. For now, let us take the fi rst step and describe the relationship
between the matrices and of with respect to two different pairs ´µ ´µ 89 89ÁÁZZ
²Á³ ² Á ³ ´µ ´ # µ89 89 and of ordered bases. Multiplication by sends toZZÁ89 8ZZ Z
´# µ8 89Z. This can be reproduced by first switching from to , then applyingZ
´µ9 989ÁZ and finally switching from to , that is,
´µ ~ 4 ´µ 4 ~ 4 ´µ 4 89 99 89 88 99 89 88ZZ Z Z Z Z ,, , ÁÁ Á Ác
Theorem 2.16 Let , and let and be pairs of orderedB 8 9 8 9² = > ³ ² Á ³ ²Á³ZZ
bases of and , respectively. Then=>
´µ ~ 4 ´µ 489 99 89 88ZZ Z ZÁÁ Á Á (2.1)
When is a linear operator on , it is generally more convenient toB² = ³ =
represent by matrices of the form , where the ordered bases used to ´µ8
represent vectors in the domain and image are the same. When , Theorem 89~
2.16 takes the following important form.
Corollary 2.17 Let and let and be ordered bases for . Then theB 8 9² = ³ =
matrix of with respect to can be exp ressed in terms of the matrix of with 9
respect to as follows:8
´µ~ 4 ´µ498 9 8 89 Á Ác(2.2)
Equivalence of Matrices
Since the change of basis matrices are precisely the invertible matrices, 2.1 has ()
the form
Linear Transformations 69
´µ ~ 7 ´µ 889 89ZZÁÁc
where and are invertible matrices. This motivates the following definition. 78
Definition Two matrices and are if there exist invertible () equivalent
matrices and for which 78
)~7( 8c
We have remarked that is equivalent to if and only if can be obtained )( )
from by a series of elementary row and column operations. Performing the(
row operations is equivalent to multiply ing the matrix on the left by and (7
performing the column operations is e quivalent to multiplying on the right by (
8c.
In terms of 2.1 , we see that performing row operations premultiplying by () ( ) 7
is equivalent to changing the basis used to represent vectors in the image and
performing column operations postm ultiplying by is equivalent to () 8c
changing the basis used to represent vectors in the domain.
According to Theorem 2.16, if and are matrices that represent with ()
respect to possibly different ordered bases, then and are equivalent. The ()
converse of this also holds.
Theorem 2.18 Let and be vector spaces with and => ² = ³ ~ dim
dim²>³ ~ d ( ) . Then two matrices and are equivalent if and only if
they represent the same linear transformation , but possibly with B² = Á > ³
respect to different ordered bases. In this case, and represent exactly the ()
same set of linear transformations in . B²= Á>³
Proof. If and represent , that is, if()
( ~ ´µ )~ ´µ89 89,,and ZZ
for ordered bases and , then Theorem 2.16 shows that and are 898 9ÁÁ ( )ZZ
equivalent. Now suppose that and are equivalent, say ()
)~7( 8c
where and are invertible. Suppose also that represents a linear78 (
transformation for some ordered bases and , that is, B 8 9² = Á > ³
(~´ µ89Á
Theorem 2.9 implies that there is a unique ordered basis for for which 8Z=
8~4 > 7~488 99 Á ÁZZ Z and a unique ordered basis for for which . Hence 9
)~4 ´µ 4 ~´µ99 89 88 89ÁÁÁ ÁZZZ Z
70 Advanced Linear Algebra
Hence, also represents . By symmetry, we see that and represent the)( )
same set of linear transformations. This completes the proof.
We remarked in Example 0.3 that every matrix is equivalent to exactly one
matrix of the block form
1~0
Á c
cÁ cÁc>?
block
Hence, the set of these matrices is a set of canonical forms for equivalence.
Moreover, the rank is a complete invariant for equivalence. In other words, two
matrices are equivalent if and only if they have the same rank.
Similarity of Matrices
When a linear operator is represen ted by a matrix of the form , B ² = ³ ´ µ 8
equation 2.2 has the form ()
´µ ~ 7 ´µ788Zc
where is an invertible matrix. Th is motivates the following definition. 7
Definition Two matrices and are , denoted by , if there () ( ) similar
exists an invertible matrix for which 7
)~7( 7c
The equivalence classes associated with similarity are called similarity
classes .
The analog of Theorem 2.18 for square matrices is the following.
Theorem 2.19 Let be a vector space of dimension . Then two = d
matrices and are similar if and only if they represent the same linear ()
operator , but possibly with respect to different ordered bases. In this B² = ³
case, and represent exactly the same set of linear operators in .() ² = ³ B
Proof. If and represent , that is, if() ² = ³ B
( ~ ´µ )~ ´µ89and
for ordered bases and , then Corollary 2.17 shows that and are similar. 89 ()
Now suppose that and are similar, say ()
)~7( 7c
Suppose also that represents a linear operator for some ordered ( ² = ³ B
basis , that is,8
(~´ µ8
Theorem 2.9 implies that there is a unique ordered basis for for which 9=
Linear Transformations 71
7~489Á. Hence
)~4 ´µ4 ~´µ89 8 9 89 Á Ác
Hence, also represents . By symmetry, we see that and represent the)( )
same set of linear operators. This completes the proof.
We will devote much effort in Chapter 7 to finding a canonical form for
similarity.
Similarity of Operators
We can also define similarity of operators.
Definition Two linear operators are , denoted by , if B Á² = ³ similar
there exists an automorphism for which B² = ³
~c
The equivalence classes associated with similarity are called similarity
classes .
Note that if and are ordered bases for , then 89~² ÁÃÁ³ ~² ÁÃÁ³ =
4 ~² ´ µ Ä´ µ ³98 8 8Á
Now, the map defined by is an automorphism of and ² ³ ~ =
4 ~ ²´ ² ³µ Ä ´ ² ³µ ³ ~ ´ µ98 8 8 8Á
Conversely, if is an automorphism and is an ordered 8¢= ¦= ~² ÁÃÁ ³
basis for , then is also a basis: = ~² ~ ² ³ÁÃÁ ~ ² ³³9
´µ~ ² ´² ³ µ Ä ´² ³ µ³ ~ 4 88 8 9 8 Á
The analog of Theorem 2.19 for linear operators is the following.
Theorem 2.20 Let be a vector space of dimension . Then two linear =
operators and on are similar if and only if there is a matrix that C =(
represents both operators, but with re spect to possibly different ordered bases.
In this case, and are represented by exactly the same set of matrices in . C
Proof. If and are represented by , that is, if C (
´ µ ~(~´ µ89
for ordered bases and , then 89
´µ~ ´µ~ 4 ´µ4 989 8 9 8 9 ÁÁ
As remarked above, if is defined by , then ¢= ¦= ²³~
72 Advanced Linear Algebra
´µ~ 498 9 Á
and so
´µ~ ´µ ´µ´µ~ ´ µ 99 999c c
from which it follows that and are similar. Conversely, suppose that and
are similar, say
~c
where is an automorphism of . Suppose also that is represented by the =
matrix , that is,(C
(~´ µ8
for some ordered basis . Then and so 8´µ~ 489 8 Á
´µ~ ´ µ~ ´µ´µ´µ ~ 4 ´µ4 88 8 89 8 8 89 8c c c
Á Á
It follows that
( ~ ´µ~ 4 ´µ4 ~ ´µ88 9 8 9 89 Á Ác
and so also represents . By symmetry, we see that and are represented(
by the same set of matrices. This completes the proof.
We can summarize the sitiation with resp ect to similarity in Figure 2.2. Each
similarity class in corresponds to a similarity class in : is IB JC J²= ³ ²-³
the set of all matrices that represent an y and is the set of all operatorsI I
in that are represented by any .BJ²= ³ 4
W similarity classes
of L(V)
[W]B
[W]CSimilarity classes
of matricesWV
V
[V]B
[V]CI
J
Figure 2.2
Invariant Subspaces and Reducing Pairs
The restriction of a linear operator to a subspace of is not B² = ³ : =
necessarily a linear operator on . This prompts the following definition. :
Linear Transformations 73
Definition Let . A subspace of is said to be orB ² = ³ : = invariant under
- if , that is, if for all . Put another way, isinvariant :: : : :
invariant under if the restriction is a linear operator on . O::
If
=~ :l ;
then the fact that is -invariant does not imply that the complement is also :;
s-invariant. The reader may wish to supply a simple example with . () =~
Definition Let . If and if both and are -invariant,B ² = ³ = ~ : l ; : ;
we say that the pair . ²:Á;³ reduces
A reducing pair can be used to decompose a linear operator into a direct sum as
follows.
Definition Let . If reduces we writeB ² = ³ ² : Á ; ³
~OlO:;
and call the of and . Thus, the expression direct sum OO:;
~l
means that there exist subspaces and of for which reduces and :; = ² : Á ; ³
and ~O ~O:;
The concept of the direct sum of linear operators will play a key role in the
study of the structure of a linear operator.
Projection Operators
We will have several uses for a special type of linear operator that is related to
direct sums.
Definition Let . The linear operator defined by=~ :l ; ¢ =¦ = :Á;
:Á;² b!³ ~
where and is called onto . : !; : ; projection along
Whenever we say that the operator is a projection, it is with the:Á;
understanding that . The following theorem describes a few basic =~ :l ;
properties of projection operators. We leave proof as an exercise.
Theorem 2.21 Let be a vector space and let .= ² = ³ B
74 Advanced Linear Algebra
1 I f t h e n)=~ :l ;
:Á; ;Á:b~
2 I f t h e n)~:Á;
im²³ ~ : ²³ ~ ; and ker
and so
=~ ²³ l ²³ im ker
In other words, is projection onto its image along its kernel. Moreover,
# ² ³ ¯ #~#im
3 If has the property that)B² = ³
=~ ²³ l ²³ O ~ im ker and im²³
then is projection onto along . im²³ ²³ ker
Projection operators are easy to characterize.
Definition A linear operator is if . B ² = ³ ~ idempotent
Theorem 2.22 A linear operator is a projection if and only if it is B² = ³
idempotent.
Proof. If , then for any and ,~ : ! ;:Á;
² b! ³~ ~ ~ ² b! ³
and so . Conversely, suppose that is idempotent. If , ~# ² ³ q ² ³ im ker
then and so#~ %
~ #~ %~ %~#
Hence . Also, if , then im² ³q ² ³ ~ ¸¹ # = ker
#~² #c # ³b # ² ³l ² ³ ker im
and so . Finally, and so .=~ ²³ l ²³ ²% ³ ~ %~ % O ~ ker im²³im
Hence, is projection onto along . im²³ ²³ ker
Projections and Invariance
Projections can be used to characterize invariant subspaces. Let and B² = ³
let be a subspace of . Let for any complement of . The key is:= ~ ; : :Á;
that the elements of can be characterized as those vectors fixed by , that is, :
Linear Transformations 75
: ~ if and only if . Hence, the following are equivalent:
::
: :
² ³ ~ :
² ³ ~ : for all
for all
for all
Thus, is -invariant if and only if for all vectors . But this is:~ :
also true for all vectors in , since both sides are equal to on . This proves ; ;
the following theorem.
Theorem 2.23 Let . Then a subspace of is -invariant if and onlyB ² = ³ : =
if there is a projection for which ~:Á;
~
in which case this holds fo r all projections of the form . ~:Á;
We also have the following relationshi p between projections and reducing pairs.
Theorem 2.24 Let . Then reduces if and only if =~ :l ; ² : Á ; ³ ² = ³ B
commutes with . :Á;
Proof. Theorem 2.23 implies that and are -invariant if and only if :;
:Á; :Á; :Á; :Á; :Á; :Á; ~ ²c ³²c ³ ~ ²c ³ and
and a little algebra shows that this is equivalent to
:Á; :Á; :Á; :Á; :Á; ~~ and
which is equivalent to . :Á; :Á;~
Orthogonal Projections and Resolutions of the Identity
Observe that if is a projection, then
²c³ ~ ²c³~
Definition Two projections are , written , if B Á² = ³ orthogonal
~~
Note that if and only if
im im²³ ²³ ²³ ²³ ker ker and
The following example shows that it is not enough to have in the ~
definition of orthogonality. In fact, it is possible for and yet is not ~
even a projection.
76 Advanced Linear Algebra
Example 2.5 Let and consider the - and -axes and the diagonal:=~ - ? @
? ~ ¸²%Á³ % -¹
@ ~ ¸²Á&³ & -¹
+ ~ ¸²%Á%³ % -¹
Then
+Á? +Á@ +Á@ +Á? +Á@ +Á? ~£~
From this we deduce that if and are projections, it may happen that both
products and are projections, but that they are not equal. We leave it to
the reader to show that (which is a projection), but that @Á? ?Á+ ?Á+ @Á? ~
is not a projection.
Since a projection is idempotent, we can write the identity operator as s sum
of two orthogonal projections:
b ²c³ ~Á ²c³
Let us generalize this to more than two projections.
Definition A on is a sum of the formresolution of the identity =
bÄb ~
where the 's are pairwise orthogonal projections, that is, for . £
There is a connection between the resolu tions of the identity on and direct =
sum decompositions of . In general terms, if =
bÄb ~
for any linear operators , then for all , B² = ³ # =
#~ #bÄb # ² ³bÄb ² ³ im im
and so
=~ ² ³ b Ä b ² ³ im im
However, the sum need not be direct.
Theorem 2.25 Let be a vector space. Resolutions of the identity on ==
correspond to direct sum decompositions of as follows: =
1 If is a resolution of the identity, then) bÄb ~
=~ ² ³ l Ä l ² ³ im im
Linear Transformations 77
and is projection onto along im²³
ker²³ ~ ²³
£im
2 Conversely, if)
=~ :l Ä l :
and if is projection onto along the direct sum ,, then £ ::
bÄb ~ is a resolution of the identity.
Proof. To prove 1), if is a resolution of the identity, then bÄb ~
=~ ² ³ b Ä b ² ³ im im
Moreover, if
%b Ä b %~
then applying gives and so the sum is direct. As to the kernel of , %~
we have
im im im²³ l ²³ ~ = ~ ²³ l ²³
£kerps
qt
and since , it follows that~
£ im²³ ²³ ker
and so equality must hold. For part 2), suppose that
=~ :l Ä l :
and is projection onto along . If , then £ :: £
im²³ ~ : ²³ ker
and so . Also, if for , then #~ bÄb :
#~ bÄb ~ #bÄb #~² bÄb ³#
and so is a resolution of the identity. ~b Ä b
The Algebra of Projections
If and are projections, it does not necessarily follow that , or bc
is a projection. For example, the sum is a projection if and only if b
²b³~ b
78 Advanced Linear Algebra
which is equivalent to
~c
Of course, this holds if , that is, if . But the converse is also ~~
true, provided that . To see this, we simply evaluate in two char²-³ £
ways:
²³ ~ c ²³ ~ c
and
²³ ~ c ²³ ~ c
Hence, and so . It follows that and so ~~ c ~ ~ c ~
² - ³ £ b. Thus, for , we have is a projection if and only if char
.
Now suppose that is a projection. For the kernel of , note that bb
² b ³ #~ ¬ ² b ³ #~ ¬ #~
and similarly, . Hence, . But the reverse #~ ² b ³ ² ³q ² ³ ker ker ker
inclusion is obvious and so
ker ker ker²b³ ~ ²³ q ²³
As to the image of , we have b
# ² b ³ ¬ #~² b ³ #~ #b # ² ³b ² ³im im im
and so . For the reverse inclusion, if , im im im²b³ ²³ b ²³ # ~% b&
then
²b³ # ~ ²b³ ² % b& ³ ~% b& ~ #
and so . Thus, . Finally, # ² b ³ ² b ³~ ² ³b ² ³ ~im im im im
implies that and so the sum is direct and im²³ ²³ ker
im im im²b³ ~ ²³ l ²³
The following theorem also describes the situation for the difference and
product. Proof in these cases is left for the exercises.
Theorem 2.26 Let be a vector space over a field of characteristic and=- £
let and be projections.
1 The sum is a projection if and only if , in which case) b
im im im²b³ ~ ²³ l ²³ ²b³ ~ ²³ q ²³ and ker ker ker
2 The difference is a projection if and only if) c
~~
Linear Transformations 79
in which case
im im im²c³ ~ ²³ q ²³ ²c³ ~ ²³ l ²³ ker ker ker and
3 If and commute, then is a projection, in which case)
im im im²³ ~² ³ q² ³ ²³ ~ ² ³ b ² ³ and ker ker ker
()Example 2.5 shows that the converse may be false.
Topological Vector Spaces
This section is for readers with some familiarity with point-set topology.
The Definition
A pair where is a real vector space and is a topology on the set²= Á ³ = =JJ
= is called a if the operations of addition topological vector space
77¢= d= ¦=Á ²#Á$³~#b$
and scalar multiplication
Cs C¢d = ¦ = Á ² Á # ³ ~ #
are continuous functions.
The Standard Topology on s
The vector space is a topological vector space under the , sstandard topology
which is the topology for which the set of open rectangles
8s~¸ 0 dÄd0 0 ¹ 's are open intervals in
is a base, that is, a subset of is open if and only if it is a union of open s
rectangles. The standard t opology is also the topology induced by the Euclidean
metric on , since an open rectangle is the union of Euclidean open balls ands
an open ball is the union of open rectangles.
The standard topology on has the property that the addition function s
7s s s¢d¦¢ ² # Á $ ³ ¦ # b $
and the scalar multiplication function
Cs s s¢d ¦ ¢ ² Á # ³ ¦ #
are continuous and so is a topological vector space under this topology. s
Also, the linear functionals are continuous maps. ¢ ¦ss
For example, to see that addition is continuous, if
²" ÁÃÁ" ³b²# ÁÃÁ# ³² Á ³dÄd² Á ³ 8
80 Advanced Linear Algebra
then and so there is an for which"b # ² Á ³
²" c Á" b ³b²# c Á# b ³ ² Á ³
for all . It follows that if
²" ÁÃÁ" ³0²" c Á" b ³dÄd²" c Á" b ³ 8
and
²# ÁÃÁ# ³ 1 ²# c Á# b ³dÄd²# c Á# b ³ 8
then
²" ÁÃÁ" ³b²# ÁÃÁ# ³ ²0Á1³² Á ³dÄd² Á ³ 7
The Natural Topology on =
Now let be a real vector space of dimension and fix an ordered basis =
8J~² #ÁÃÁ#³ = for . We wish to show that there is precisely one topology
on for which is a topological vector space and all linear functionals=² = Á ³ J
are continuous. This topology is called the on . natural topology =
Our plan is to show that if is a topological vector space and if all linear ²= Á ³J
functionals on are continuous, then the coordinate map is a =¢ = s8
homeomorphism. This implies that if does exist, it must be unique. Then we J
use to move the standard topol ogy from to , thus giving a s~= =8c
topology for which is a homeomorphi sm. Finally, we show that is J J 8 ²= Á ³
a topological vector space and that all linear functionals on are continuous. =
The first step is to show that if is a topological vector space, then is ²= Á ³J
continuous. Since where is defined by s~¢ ¦ =
² ÁÃÁ ³~#
it is sufficient to show that these maps are continuous. The sum of continuous (
maps is continuous. ) Let be an open set in . Then6 J
Csc²6³ ~ ¸²Á%³ d= % 6¹
is open in . This implies that if , then there is an open interval sd= %6
0 s containing for which
0%~¸ % 0¹6
We need to show that the set is open. But c²6³
s
sss ssc
²6³ ~ ¸² ÁÃÁ ³ # 6¹
~ dÄd d¸ # 6¹d dÄd
In words, an -tuple is in if the th coordinate times is ² Á Ã Á ³ ² 6 ³ # c
Linear Transformations 81
in . But if , then there is an open interval for which and 6 # 6 0 0 s
0# 6 . Hence, the entire open set
<~ dÄd d0d dÄdss ss
where the factor is in the th position is in , that is, 0 ² 6 ³ c
² ÁÃÁ ³< ²6³ c
Thus, is open and , and therefore also , is continuous. c ²6³
Next we show that if every linear functional on is continuous under a =
topology on , then the coordinate map is continuous. If denote by J=# =
´#µ ´#µ ¢= ¦ # ~ ´#µ88 8Á Á the th coordinate of . The map defined by is a s
linear functional and so is continuous by assumption. Hence, for any open
interval the set0s
( ~¸ #= ´ # µ 0¹Á 8
is open. Now, if are open intervals in , then 0 s
c
²0 dÄd0 ³~¸#= ´#µ 0 dÄd0 ¹~ ( 8
is open. Thus, is continuous.
We have shown that if a topology has the property that is a JJ ²= Á ³
topological vector space under which every linear functional is continuous, then
J and are homeomorphisms. This means that if exists, its open sets~c
must be the images under of the open sets in the standard topology of . It s
remains to prove that the topology on that makes a homeomorphism J=
makes a topological vector space for which any linear functional on ²= Á ³ =J
is continuous.
The addition map on is a composition =
7 7~ k k ²d³c Z
where is addition in and since each of the maps on the7s s s sZ ¢d¦
right is continuous, so is . 7
Similarly, scalar multiplication in is =
C C~kk ² d ³c Z
where is scalar multiplication in . Hence, isCs s s s CZ ¢d ¦
continuous.
Now let be a linear functional. Since is continuous if and only if is k c
continuous, we can confine attention to . In this case, if is the =~ Á Ã Á s
standard basis for for any s and for all , then ((²³ 4
82 Advanced Linear Algebra
%~² ÁÃÁ ³ s, we have
((²%³ ~ ²³ ²³ 4 cc (( ( ( ((
Now, if , then (( ( (% ² % ³ ° 4 ° 4 (( and so , which implies that
is continuous at . %~
According to the Riesz representation theorem (Theorem 9.18) and the Cauchy–
Schwarz inequality, we have
)) ) ) ) )²%³ %H
where . Hence, implies and so by linearity, 9 %¦ ² % ³ ¦ %¦ % s
implies and so is continuous at all .²% ³¦% %
Theorem 2.27 Let be a real vector space of dimension . There is a unique =
topology on , called the , for which is a topological vector == natural topology
space and for which all linear functional s on are continuous. This topology is=
determined by the fact that the coordinate map is a s¢= ¦
homeomorphism, where has the standard topology induced by the Euclidean s
metric.
Linear Operators on =d
A linear operator on a real vector space can be extended to a linear operator =
dd on the complexifica tion by defining=
d²"b#³ ~ ²"³b ²#³
Here are the basic properties of this of . complexification
Theorem 2.28 If , then BÁ² = ³
1 , )² ³ ~ sdd
2)²b³~ b ddd
3)²³ ~ dd d
4 .)´# µ ~ ² #³dd d
Let us recall that for any ordered basis for and any vector we have 8=# =
´#bµ ~ ´#µ cpx²³88
Now, if is an ordered basis for , then the th column of is8 = ´ µ 8
´ µ ~ ´b µ ~ ´ ² b ³ µ ²³ ²³ 8 88d
cpx cpx
which is the th column of the coordina te matrix of with respect to the basis d
cpx²³8. Thus we have the following theorem.
Linear Transformations 83
Theorem 2.29 Let where is a real vector space. The matrix of B ² = ³ =d
with respect to the ordered basis is equal to the matrix of with respect cpx²³8
to the ordered basis : 8
´µ ~ ´ µd
88 cpx²³
Hence, if a real matrix represents a linear operator on , then also (= (
represents the complexification of on . dd=
Exercises
1. Let have rank . Prove that there are matrices and( ?CCÁ Á
@ (~? @ ( CÁ, both of rank , for which . Prove that has rank if
and only if it has the form where and are row matrices. (~%& % &!
2. Prove Corollary 2.9 and find an example to show that the corollary does not
hold without the finiteness condition.
3. Let . Prove that is an isomorphism if and only if it carries aB ² = Á > ³
basis for to a basis for . =>
4. If and we define the external direct sumB B² = Á > ³ ² = Á > ³
^ B ^ ^² = = Á > > ³ by
² ³²²# Á# ³³ ~ ² # Á # ³^
Show that is a linear transformation.^
5. Let . Prove that . Thus, internal and external =~ :l ; :l ; : ; ^
direct sums are equivalent up to isomorphism.
6. Let and consider the external direct sum . Define a=~ ( b ) ,~ ( ) ^
map by . Show that is linear. What is the^ ¢( )¦= ²#Á$³~#b$
kernel of ? When is an isomorphism?
7. Let where . Let . Suppose thatB C² = ³ ² = ³ ~ B ( ² - ³- dim
there is an isomorphism with the property that . ¢= - ² #³~(² #³
Prove that there is an ordered basis for which . 8 (~´ µ8
8. Let be a subset of . A subspace of is if is -JB J ²= ³ : = : -invariant
invariant for every . Also, is if the only -invariant J J J= -irreducible
subspaces of are and . Prove the following form of = ¸¹ = Schur's lemma.
Suppose that and and is -irreducible and JB JB J=> =² = ³ ² > ³ = >
is -irreducible. Let satisfy , that is, for anyJ B J J >= > ² = Á > ³ ~
J J J ~ => > there is a such that and for any there is a
J ~ ~ = such that . Prove that or is an isomorphism.
9. Let where . If show thatB ² =³ ² =³B ² ³~ ² ³ dim rk rk
im² ³q ² ³ ~ ¸¹ ker .
10. Let , and . Show thatB B² < = ³ ² = Á > ³
minrk rk rk²³ ¸ ² ³ Á² ³ ¹
11. Let and . Show thatB B² < Á = ³ ² = Á > ³
null null null²³ ² ³ b ² ³
84 Advanced Linear Algebra
12. Let where is invertible. Show that B Á² = ³
rk rk rk²³ ~²³ ~² ³
13. Let . Show that BÁ² = Á > ³
rk rk rk²b³ ²³ b ²³
14. Let be a subspace of . Show that there is a for which := ² = ³ B
ker² ³~: ² =³ ² ³~: B . Show also that there exists a for which . im
15. Suppose that . BÁ² = ³
a Show that for some . )i m i m ~² ³ ² ³ ² = ³B if and only if
b Show that for some . ) ~² ³ ² ³ ² = ³B if and only if ker ker
16. Let and suppose that satisfies . Show that dim²= ³ B ²= ³ ~ B
² ³ ² = ³rk dim .
17. Let be an matrix over . What is the relationship between the ( d -
linear transformation and the system of equations ? (¢- ¦- (?~)
Use your knowledge of linear transformations to state and prove various
results concerning the system , especially when . (? ~ ) ) ~
18. Let have basis and assume that the base field for = ~¸# ÁÃÁ# ¹ - = 8
has characteristic . Suppose that for each we define Á
BÁ² = ³ by
Á
²# ³ ~# £
#b # ~Fif
if
Prove that the are invertible and form a basis for . BÁ ²= ³
19. Let . If is a -invariant subspace of must there be a subspaceB ² = ³ : =
;= ² : Á ; ³ of for which reduces ?
20. Find an example of a vector space and a proper subspace of for =: =
which .= :
21. Let . If , prove that implies that and dim²= ³ B ²= ³ ~ B
are invertible and that for some polynomial . ~ ² ³ ² % ³-´ % µ
22. Let . If for all show that , for someB B ² = ³ ~ ² = ³ ~
- , where is the identity map.
23. Let be a vector space over a field of characteristic and let and =- £
be projections. Prove the following:
a The difference is a projection if and only if ) c
~~
in which case
im im im²c³ ~ ²³ q ²³ ²c³ ~ ²³ l ²³ ker ker ker and
Hint: is a projection if and only if is a projection and so cc
is a projection if and only if
Linear Transformations 85
~c ²c³ ~ ² c³ b
is a projection.
b If and commute, then is a projection, in which case )
im im im²³ ~² ³ q² ³ ²³ ~ ² ³ b ² ³ and ker ker ker
24 be a continuous function with the property that. Let ¢ss¦
²%b&³~²%³b²&³
Prove that is a linear functional on . s
25. Prove that any linear functional is a continuous map. ¢ ¦ss
26. Prove that any subspace of is a closed set or, equivalently, that :s
:~ ± : % : ) ² % Á³ s is open, that is, for any there is an open ball
centered at with radius for which . % ) ² % Á ³ :
27. Prove that any linear transformation is continuous under the ¢= ¦>
natural topologies of and . =>
28. Prove that any surjective linear transformation from to both finite- => (
dimensional topological vector spaces under the natural topology is an )
open map, that is, maps open sets to open sets.
29. Prove that any subspace of a finite-dimensional vector space is a :=
closed set or, equivalently, that is open, that is, for any there is:% :
an open ball centered at with radius for which )²%Á ³ %
)²%Á ³ :.
30. Let be a subspace of with .:= ² = ³ B dim
a Show that the subspace topology on inherited from is the natural ) :=
topology.
b Show that the natural topology on is the topology for which the ) =°:
natural projection map continuous and open. ¢= ¦=°:
31. If is a real vector space, then is a complex vector space. Thinking of==d
=² = ³² = ³dd dss as a vector space over , show that is isomorphic to the s
external direct product . ==^
32. When is a complex linear map a complexification? Let be a real vector() =
space with complexification and let . Prove that is a = ² = ³ddB
complexification, that is, has the form for some if and only Bd² = ³
if commutes with the conjugate map defined by ¢= ¦=dd
²"b#³~"c# .
33. Let be a complex vector space.>
a Consider replacing the scalar multiplication on by the operation ) >
²'Á$³ ¦ '$
where and . Show that the re sulting set with the addition ' $>d
defined for the vector space and with this scalar multiplication is a >
complex vector space, which we denote by . >
b Show, without using di mension arguments, that . ) ²> ³ > >sd^
Chapter 3
The Isomorphism Theorems
Quotient Spaces
Let be a subspace of a vector space . It is easy to see that the binary relation:=
on defined by=
"# ¯ "c#:
is an equivalence relation. When , we say that and are "# " # congruent
modulo . The term is used as a colloquialism for modulo and is:" # mod
often written
"# : mod
When the subspace in question is clear, we will simply write . "#
To see what the equivalence classes look like, observe that
´ # µ~¸ "= "# ¹
~¸ "= "c#:¹
~¸ "= "~#b :¹
~¸ #b :¹
~#b: for some
The set
´ # µ~#b:~¸ #b :¹
is called a of in and is called a for . coset coset representative := # # b :
()Thus, any member of a coset is a coset representative.
The set of all cosets of in is denoted by :=
=°:~¸#b:#=¹
This is read “ mod ” and is called the . Of =: = : quotient space of modulo
88 Advanced Linear Algebra
course, the term space is a hint that we intend to define vector space operations
on .=°:
The natural choice for these vector space operations is
²"b:³b²#b:³~²"b#³b:
and
²"b:³ ~ ²"³b:
but we must check that these operations are well-defined, that is,
1)" b:~" b: Á# b:~# b:¬² " b#³b:~² " b#³b:
2)"b : ~ "b : ¬ "b : ~ "b :
Equivalently, the equivalence relation must be with the vector consistent
space operations on , that is, =
3)" " Á # #¬ ² "b # ³ ² "b # ³
4)" "¬ " "
This senario is a recurring one in algebra. An equivalence relation on an
algebraic structure, such as a group, ring, module or vector space is called a
congruence relation if it preserves the algebraic operations. In the case of a
vector space, these are conditions 3) and 4) above.
These conditions follow easily from the fact that is a subspace, for if :" "
and , then# #
"c " : Á #c # : ¬ ² "c "³ b ² #c #³ :
¬ ²" b #³c²" b #³:
¬ "b # "b #
which verifies both conditions at once. We leave it to the reader to verify that
=°: - is indeed a vector space over under these well-defined operations.
Actually, we are lucky here: For subspace of , the quotient is a any := = ° :
vector space under the natural operations. In the case of groups, not all
subgroups have this property. Indeed, it is precisely the subgroups of normal 5
.. ° 5 that have the property that the quotient is a group. Also, for rings, it is
precisely the (not the subrings) that have the property that the quotient is ideals
a ring.
Let us summarize.
The Isomorphism Theorems 89
Theorem 3.1 Let be a subspace of . The binary relation:=
"# ¯ "c#:
is an equivalence relation on , whose equivalence classes are the = cosets
#b:~¸ #b :¹
of in . The set of all cos ets of in , called the of := = ° : := = quotient space
modulo , is a vector space under the well-defined operations:
²"b:³~"b:
²"b:³b²#b:³~²"b#³b:
The zero vector in is the coset . =°: b:~:
The Natural Projection and the Correspondence Theorem
If is a subspace of , then we can define a map by sending:= ¢ = ¦ = ° : :
each vector to the coset containing it:
:²#³ ~ #b:
This map is called the or of onto canonical projection natural projection =
=°: : , or simply . (Not to be confused with the projection projection modulo
operators .) It is easily seen to be linear, for we have writing for :Á; : ()
²"b #³ ~ ²"b #³b: ~ ²"b:³b ²#b:³ ~ ²"³b ²#³
The canonical projection is clearly surjective. To determine the kernel of , note
that
# ² ³¯ ² # ³~¯#b:~:¯#:ker
and so
ker²³ ~ :
Theorem 3.2 The canonical projection defined by :¢= ¦=°:
:²#³ ~ #b:
is a surjective linear transformation with . ker²³ ~ ::
If is a subspace of , then the subspaces of the quotient space have the:= = ° :
form for some intermediate subspace satisfying . In fact, as;°: ; :; =
shown in Figure 3.1, the projection map provides a one-to-one :
correspondence between intermediate subspaces and subspaces of :;=
the quotient space . The proof of the following theorem is left as an =°:
exercise.
90 Advanced Linear Algebra
V
V/S
{0}S T/ST
{0}
Figure 3.1: The correspondence theorem
Theorem 3.3 The correspondence theorem() Let be a subspace of . Then:=
the function that assigns to ea ch intermediate subspace the :;=
subspace of is an order-preserving with respect to set inclusion ;°: =°: ()
one-to-one correspondence between the set of all subspaces of containing =:
and the set of all subspaces of . =°:
Proof. We prove only that the correspondence is surjective. Let
?~¸ "b:"<¹
be a subspace of and let be the union of all cosets in : =°: ; ?
;~ ² "b: ³
"<
We show that and that . If , then and :;= ;° :~? % Á&; %b:
&b: ? ?=°: are in and since , we have
%b:Á²%b&³b:?
which implies that . Hence, is a subspace of containing . %Á%b& ; ; = :
Moreover, if , then and so . Conversely, if !b:;°: !; !b:?
"b:? "; "b:;°: ?~;°: , then and therefore . Thus, .
The Universal Property of Quotients and the First
Isomorphism Theorem
Let be a subspace of . The pair has a very special property,:= ² = ° : Á ³ :
known as the —a term that comes from the world of category universal property
theory.
Figure 3.2 shows a linear transformation , along with the B² = Á > ³
canonical projection from to the quotient space . :== ° :
The Isomorphism Theorems 91
V W
V/SSsW
W'
Figure 3.2: The universal property
The universal property states that if , then there is a unique ker²³ :
Z¢=°:¦> for which
Z
:k~
Another way to say this is that any such can be B² = Á > ³ factored through
the canonical projection . :
Theorem 3.4 Let be a subspace of and let satisfy := ² = Á > ³ B
: ²³ ker. Then, as pictured in Figure 3.2, there is a unique linear
transformation with the property that Z¢=°:¦>
Z
:k~
Moreover, and . ker ker²³ ~ ² ³ ° : ²³ ~ ² ³ ZZim im
Proof. We have no other choice but to define by the condition , ZZ:k~
that is,
Z²#b:³ ~ #
This function is well-defined if and only if
#b:~"b:¬ ² #b: ³~ ² "b: ³ ZZ
which is equivalent to each of the following statements:
#b:~"b:¬ #~ "
#c":¬ ²#c"³~
%:¬ %~
: ²³
ker
Thus, is well-defined. Also,Z¢=°:¦>
im im² ³~¸ ² #b:³#=¹~¸ ##=¹~ ² ³ ZZ
and
92 Advanced Linear Algebra
ker
ker
ker²³ ~ ¸ # b : ² # b : ³ ~ ¹
~¸ #b: #~ ¹
~¸ #b:# ² ³ ¹
~² ³ ° :
ZZ
The uniqueness of is evident. Z
Theorem 3.4 has a very important corollary, which is often called the first
isomorphism theorem and is obtained by taking . :~ ²³ ker
Theorem 3.5 The Let be a linear () first isomorphism theorem ¢= ¦>
transformation. Then the linea r transformation defined by Z¢=° ² ³¦>ker
Z²#b ² ³³ ~ # ker
is injective and
=
²³² ³kerim
According to Theorem 3.5, the image of any linear transformation on is =
isomorphic to a quotient space of . Conversely, any quotient space of == ° : =
is the image of a linear transformation on : the canonical projection . Thus, = :
up to isomorphism, quotient spaces are equivalent to homomorphic images.
Quotient Spaces, Complements and Codimension
The first isomorphism theorem gives some insight into the relationship between
complements and quotient spaces. Let be a subspace of and let be a := ;
complement of , that is, :
=~ :l ;
Applying the first isomorphism theorem to the projection operator ;Á:¢= ¦;
gives
;=° :
Theorem 3.6 Let be a subspace of . All complements of in are := : =
isomorphic to and hence to each other. =°:
The previous theorem can be rephrased by writing
(l)~(l*¬)*
On the other hand, quotients and complements do not behave as nicely with
respect to isomorphisms as one might casua lly think. We leave it to the reader to
show the following:
The Isomorphism Theorems 93
1 It is possible that)
(l)~*l+
with but . Hence, does imply that a complement(* )+ (* ° not
of is isomorphic to a complement of .(*
2 It is possible that and) = >
=~ :l ) >~ :l + and
but . Hence, does imply that . However,)+ => =° :>° :° not (
according to the previous theorem, if then . => ) +equals )
Corollary 3.7 Let be a subspace of a vector space . Then:=
dim dim dim²= ³ ~ ²:³b ²= °:³
Definition If is a subspace of , then is called the of:= ² = ° : ³ dim codimension
:= ² : ³ ² : ³ in and is denoted by or . codim codim =
Thus, the codimension of in is the dimension of any complement of in := :=
and when is , we have = finite-dimensional
codim =²:³ ~ ²= ³c ²:³ dim dim
(This makes no sense, in general, if is not finite-dimensional, since infinite=
cardinal numbers cannot be subtracted.)
Additional Isomorphism Theorems
There are other isomorphism theorems that are direct consequences of the first
isomorphism theorem. As we have seen, if then . This can =~ :l ; = ° ; :
be written
:l; :
;: q ;
This applies to nondirect sums as well.
Theorem 3.7 The Let be a vector space () second isomorphism theorem =
and let and be subspaces of . Then:; =
:b; :
;: q ;
Proof. Let be defined by¢² :b;³¦:° ² :q;³
² b!³ ~ b²: q;³
We leave it to the reader to show th at is a well-defined surjective linear
transformation, with kernel . An application of the first isomorphism theorem ;
then completes the proof.
94 Advanced Linear Algebra
The following theorem demonstrates one way in which the expression =°:
behaves like a fraction.
Theorem 3.8 The Let be a vector space and () third isomorphism theorem =
suppose that are subspaces of . Then :;= =
=°: =
;°: ;
Proof. Let be defined by . We leave it to the¢= °: ¦ = °; ²#b:³ ~ #b;
reader to show that is a well-defined surjective linear transformation whose
kernel is . The rest follows from the first isomorphism theorem. ;°:
The following theorem demonstrates one way in which the expression =°:
does not behave like a fraction.
Theorem 3.9 Let be a vector space and let be a subspace of . Suppose =: =
that and with . Then=~ =l = :~ :l : : =
== l ===
:: l : ::~
^
Proof. Let be defined by^¢=¦² =° :³ =° :³ ²
²# b#³~²# b:Á# b:³
This map is well-defined, since the sum is direct. We leave it to =~ =l =
the reader to show that is a surjective linear transformation, whose kernel is
:l : . The rest follows from the first isomorphism theorem.
Linear Functionals
Linear transformations from to the base field thought of as a vector space =- (
over itself are extremely important. )
Definition Let be a vector space over . A linear transformation=-
² =Á- ³ -B whose values lie in the base field is called a linear functional
() ( )or simply on . Some authors use the term . The functional = linear function
vector space of all linear functionals on is denoted by and is called the ==*
algebraic dual space of .=
The adjective is needed here, since there is another type of dual space algebraic
that is defined on general normed vector spaces, where continuity of linear
transformations makes sense. We will discuss the so-called continuous dual
space briefly in Chapter 13. However, until then, the term “dual space” will
refer to the algebraic dual space.
The Isomorphism Theorems 95
To help distinguish linear functionals from other types of linear transformations,
we will usually denote linear functionals by lowercase italic letters, such as ,
and .
Example 3.1 The map defined by is a linear ¢-´%µ¦- ²²%³³~²³
functional, known as . evaluation at
Example 3.2 Let denote the vector space of all continuous functions on9´Áµ
´Áµ ¢ ´Áµ ¦s9 s. Let be defined by
² ²%³³ ~ ²%³%
Then . ´ Á µ9i
For any , the rank plus nullity theorem is=*
dim ker dim dim² ²³³b ² ²³³~ ²=³ im
But since , we have either , in which case is the zero im im²³ - ²³ ~ ¸¹
linear functional, or , in which case is surjective. In other words, a im²³ ~ -
nonzero linear functional is surjective. Moreover, if , then £
codim²² ³ ³ ~ ~ =
²³ker dimker67
and if , then dim²= ³ B
dim ker dim²² ³ ³ ~ ² = ³ c
Thus, in dimensional terms, the kernel of a linear functional is a very “large”
subspace of the domain . =
The following theorem will prove very useful.
Theorem 3.10
1 For any nonzero vector , there exists a linear functional for) #= =*
which .²#³£
2 A vector is zero if and only if for all .) #= ² # ³~ =*
3 Let . If , then)= ² % ³£i
=~ º % » l ² ³ ker
4 Two nonzero linear functionals have the same kernel if and only) Á=i
if there is a nonzero scalar such that . ~
Proof. For part 3 , if , then and for ) £ # º%»q ²³ ²#³ ~ # ~ % ker
£ - ²%³ ~ º%»q ²³ ~ ¸¹ , whence , which is false. Hence, and ker
the direct sum exists. Also, for any we have :~º % »l ² ³ #= ker
96 Advanced Linear Algebra
# ~ %b #c % º%»b ²³²#³ ²#³
²%³ ²%³67 ker
and so .=~ º % » l ² ³ ker
For part 4 , if for , then . Conversely, if )~ £ ² ³~ ² ³ ker ker
2 ~ ²³ ~ ²³ % ¤ 2 ker ker , then for we have by part 3 , )
=~ º % » l 2
Of course, for any . Therefore, if , it follows that O ~ O ~²%³°²%³22
²%³ ~ ²%³ ~ and hence .
Dual Bases
Let be a vector space with basis . For each , we can= ~¸ # 0¹ 0 8
define a linear functional by the orthogonality condition # =i
*
#² #³~i
Á
where is the , defined byÁ Kronecker delta function
Á~ ~
£ Fif
if
Then the set is linearly independent, since applying the 8ii
~¸ # 0¹
equation
~ # bÄb #ii
to the basis vector gives #
~ #² #³ ~ ~
~ ~
Á i
for all .
Theorem 3.11 Let be a vector space with basis . =~ ¸ # 0 ¹ 8
1 The set is linearly independent.)8ii
~¸ # 0¹
2 If is finite-dimensional, then is a basis for , called the of)== 8iidual basis
8.
Proof. For part 2 , for any , we have ) =i
Á i
²#³# ²#³~ ²#³ ~²#³
and so is in the span of . Hence, is a basis for .~ ² #³ # =ii i
88*
The Isomorphism Theorems 97
It follows from the previous theorem that if , then dim²= ³ B
dim dim²= ³ ~ ²= ³i
since the dual vectors also form a basis for . Our goal now is to show that the =i
converse of this also holds. But first, let us consider an example.
Example 3.3 Let be an infinite-dimensional vector space over the field=
- ~ ~ ¸Á¹ - {8 , with basis . Since the only coefficients in are and , a
finite linear combination over is just a finite sum. Hence, is the set of all -=
finite sums of vectors in and so according to Theorem 0.12, 8
(( ( ( ((= ²³ ~F8 8
On the other hand, each linear functional is uniquely defined by =i
specifying its values on the basis . Since these values must be either or , 8
specifying a linear functional is equivale nt to specifying the subset of on 8
which takes the value . In other words, there is a one-to-one correspondence
between linear functionals on and all subsets of . Hence, = 8
((( (( (( (=~² ³ =iF8 8
This shows that cannot be isomorphic to , nor to any proper subset of . == =i
Hence, . dim dim²= ³ ²= ³i
We wish to show that the behavior in the previous example is typical, in
particular, that
dim dim²= ³ ²= ³i
with equality if and only if is finite- dimensional. The proof uses the concept =
of the of a field , which is defi ned as the smallest subfield of prime subfield 2
the field . Since , it follows that contains a copy of the integers 2 Á 2 2
ÁÁ ~ bÁ ~ bbÁÃ
If has prime characteristic , then and so contains the elements2 ~ 2
{~ ¸ÁÁÁÁÃÁ c¹
which form a subfield of . Since any subfi eld of contains and , we see 2- 2
that and so is the prime subfield of . On the other hand, if has{{- 2 2
characteristic , then contains a “copy” of the integers and therefore also 2 {
the rational numbers , which is the pr ime subfield of . Our main interest in r 2
the prime subfield is that in eith er case, the prime subfield is . countable
Theorem 3.12 Let be a vector space. Then=
dim dim²= ³ ²= ³i
with equality if and only if is finite-dimensional. =
98 Advanced Linear Algebra
Proof. For any vector space , we have =
dim dim²= ³ ²= ³i
since the dual vectors to a basis for are linearly independent in . We 8==i
have already seen that if is finite-dimensional, then . We =² = ³ ~ ² = ³ dim dimi
wish to show that if is in finite-dimensional, then . The=² = ³ ² = ³ dim dimi(
author is indebted to Professor Rich ard Foote for suggesting this line of proof.)
If is a basis for and if is the base field for , then Theorem 2.7 implies8 =2 =
that
= ² 2³8
where is the set of all functions with finite support from to and²2 ³ 28 8
= 2i8
where is the set of all functions from to . Thus, we can work with the 2288
vector spaces and . ²2 ³ 288
The plan is to show that if is a c ountable subfield of and if is infinite,-2 8
then
dim dim dim dim2 - - 223 23 2 3²2 ³ ~ ²- ³ ²- ³ 288 8 8
Since we may take to be the prime s ubfield of , this will prove the theorem. -2
The first equality follows from the fact that the -space and the -space 2² 2 ³ -8
²- ³8 each have a basis consisting of the “standard” linear functionals
¸ ¹8 defined by
# ~ Á
for all , where is the Kronecker delta function.# Á 8
For the final inequality, suppose that is linearly independent over ¸ ¹ - -8
and that
~
where . If is a basis for over , then for Á Á 2 ¸ ¹ 2 - ~ -
and so
~ ~
Á
Evaluating at any gives #8
The Isomorphism Theorems 99
~ ² # ³~ ² # ³ 89
Á Á
and since the inner sums are in and is -independent, the inner sums -¸ ¹ -
must be zero:
Á ² # ³ ~
Since this holds for all , we have #8
Á ~
which implies that for all . Hence, is linearly independent over ~ Á ¸ ¹Á
2² - ³ 2. This proves that . dim dim-288 23
For the center inequality, it is clear that
dim dim- -23²- ³ ²- ³88
We will show that the inequality must be strict by showing that the cardinality
of is whereas the cardinality of is greater than . To this end, the ²- ³ -88(( ((88
set can be partitioned into blocks based on the support of the function. In²- ³8
particular, for each finite subset of , if we let :8
(~ ¸ ² -³ ² ³ ~ : ¹:8supp
then
²- ³ ~ (8
:
:
:8
finite
where the union is disjoint. Moreover, if , then ((:~
((( (( - L:
and so
bb ( ( (( (( (( ²- ³ ~ ( hL ~ ² ÁL ³ ~8
:
:
:8
finite88 8 max
But since the reverse inequality is easy to establish, we have
bb (( ²- ³ ~8
8
As to the cardinality of , for each subset of , there is a function -; -888 ;
that sends every element of to and every element of to . Clearly, ; ± ; 8
each distinct subset gives rise to a distinct function and so Cantor's ; ;
100 Advanced Linear Algebra
theorem implies that
bbb b b b (( - ~ ² - ³88 88
This shows that
dim dim- -23²- ³ ²- ³88
and completes the proof.
Reflexivity
If is a vector space, then so is the dual space and so we may form the==i
double algebraic dual space ( ) , which consists of all linear functionals =ii
¢= ¦- =i. In other words, an element of is a linear functional that**
assigns a scalar to each linear functional on . =
With this firmly in mind, there is one rather obvious way to obtain an element of
=# = # ¢ = ¦ -ii i. Namely, if , consider the map defined by
#²³ ~ ²#³
which sends the linear functional to the scalar . The map is called ² # ³ #
evaluation at # #= Á= Á-. To see that , if and , thenii i
#² b³ ~ ² b³²#³ ~ ²#³b²#³ ~ #²³b#²³
and so is indeed linear.#
We can now define a map by ¢= ¦=ii
#~#
This is called the or the from to . This canonical map natural map () ==ii
map is injective and hence in the fin ite-dimensional case, it is also surjective.
Theorem 3.13 The canonical map defined by , where is ¢= ¦= #~# #ii
evaluation at , is a monomorphism. If is finite-dimensional, then is an #=
isomorphism.
Proof. The map is linear since
"b#²³ ~ ²"b#³ ~ ²"³b²#³ ~ ²"b#³²³
for all . To determine the kernel of , observe that=i
#~¬#~
¬# ² ³~ =
¬² # ³~ =
¬# ~ for all
for all i
i
by Theorem 3.10 and so . In the finite-dimensional case, since ker² ³ ~ ¸¹
The Isomorphism Theorems 101
dim dim dim² = ³~ ² = ³~ ² =³ii i
it follows that is also surjective, hence an isomorphism.
Note that if , then since the dime nsions of and are the same, dim²= ³ B = =ii
we deduce immediately that . This is not the point of Theorem 3.13. = =ii
The point is that the is an isomorphism. Because of this, natural map #¦# =
is said to be . Theorem 3.13 and Theorem 3.12 together algebraically reflexive
imply that a vector space is algebraically reflexive if and only if it is finite-
dimensional.
If is finite-dimensional, it is customary to identify the double dual space ==ii
with and to think of the elements of simply as vectors in . Let us == =ii
consider a specific example to show how algebraic reflexiv ity fails in the
infinite-dimensional case.
Example 3.4 Let be the vector space over with basis= {
~²ÁÃÁÁÁÁó
where the is in the th position. Thus , is the set of all infinite binary =
sequences with a finite number of 's. Define the of any to be ² # ³ # = order
the largest coordinate of with value . Then for all . # ² # ³ B # =
Consider the dual vectors , defined as usual by i()
² ³~i
Á
For any , the evaluation functional has the property that#= #
# ² ³~² # ³~ ² # ³ii if
However, since the dual vectors are linearly independent, there is a linear i
functional for which =ii
² ³~i
for all . Hence, does not have the form for any . This shows that # #=
the canonical map is not surjective and so is not algebraically reflexive.=
Annihilators
The functions are defined on vectors in , but we may also define on = = i
subsets of by letting4=
²4³~¸²#³#4¹
102 Advanced Linear Algebra
Definition Let be a nonempty subset of a vector space . The 4= annihilator
44 of is
4 ~ ¸ = ²4³ ~ ¸¹¹i
The term annihilator is quite descriptiv e, since consists of all linear 4
functionals that send to every vector in . It is not hard to see annihilate () 4
that is a subspace of , even when is not a subspace of .4= 4 =i
The basic properties of annihilators ar e contained in the following theorem.
Theorem 3.14
1 I f )( )Order-reversing 45 = and are nonempty subsets of , then
45 5 4 ¬
2 If , then for any nonempty subset of the natural map)d i m ²= ³ B 4 =
¢span²4³ 4
is an isomorphism from . In particular, if is a span²4³ 4 onto :
subspace of , then . =: :
3 If and are subspaces of , then):; =
²:q;³ ~: b; ²:b;³ ~: q; and
Proof. We leave proof of part 1 for the reader. For part 2 , since ))
4~ ² ² 4 ³ ³ span
it is sufficient to prove that ¢:: : is an isomorphism, where is a
subspace of . Now, we know that is a monomorphism, so it remains to prove =
that . If , then has the property that for all ,:~: : ~ :
²³ ~ ~
and so , which implies that . Moreover, if , then ~ : : #: :
for all we have:
²#³~#²³~
and so every linear functional that a nnihilates also annihilates . But if , :# # ¤ :
then there is a linear functional for which and . = ²:³ ~ ¸¹ ²#³ £ i
()We leave proof of this as an exercise. Hence, and so and #: #~ #:
so .::
For part 3 , it is clear that annihilates if and only if annihilates both ) : b ;
:; and . Hence, ²:b;³ ~: q; ~b: b; . Also, if where
: ; Á² :q;³ ² :q;³ and , then and so . Thus,
The Isomorphism Theorems 103
:b ; ² : q ; ³
For the reverse inclusion, suppose that . Write ² :q; ³
=~ :l ² :q ; ³ l ;l <ZZ
where and . Define by:~:l² :q;³ ;~² :q;³l; =ZZ i
O ~Á O ~O ~ Á O ~ Á O ~:: q ; : q ; ;<ZZ
and define by =i
O ~ Á O ~O ~ Á O ~Á O ~:: q ; : q ;;<ZZ
It follows that , and . ; : b~
Annihilators and Direct Sums
Consider a direct sum decomposition
=~ :l ;
Then any linear functional can be extended to a linear functional on ; =i
by setting . Let us call this . Clearly, and it is ²:³~ : extension by
easy to see that the extension by map is an isomorphism from to ¦ ;i
:;, whose inverse is the restriction to .
Theorem 3.15 Let . =~ :l ;
a The extension by map is an isomorphism from to and so) ; :i
; :i
b If is finite-dimensional, then)=
dim dim dim²: ³ ~ ²:³ ~ ²= ³c ²:³codim =
Example 3.5 Part b of Theorem 3.15 may fail in the infinite-dimensional case, )
since it may easily happen that . As an example, let be the vector : = =i
space over with a countably infinite ordered basis . Let {8 ~² Á Á Ã ³
:~º » ;~º Á Á Ã » : ; = ii and . It is easy to see that and that
dim dim²= ³ ²= ³i.
The annihilator provides a way to describe the dual space of a direct sum.
Theorem 3.16 A linear functional on the direct sum can be written =~ :l ;
as a sum of a linear functional that annihilates and a linear functional that :
annihilates , that is, ;
²: l;³ ~ : l;i
104 Advanced Linear Algebra
Proof. Clearly , since any functional that annihilates both and : q; ~ ¸¹ :
;: l ; ~ = : b ; must annihilate . Hence, the sum is direct. The rest
follows from Theorem 3.14, since
= ~ ¸¹ ~i² : q ; ³~ :b ;~ :l ;
Alternatively, since is the id entity map, if , then we can ;:ib~ =
write
~k² b ³~² k ³b² k ³: l; ;: ; :
and so .=~ :l ;i
Operator Adjoints
If , then we may define a map byB ² = Á > ³ ¢ >¦ =di *
d²³ ~ k ~
for . We will write composition as juxtaposition. Thus, for any ,> #=*()
´² ³ µ ² # ³ ~ ² # ³d
The map is called the of and can be described by thedoperator adjoint
phrase “apply first.”
Theorem 3.17 Properties of the Operator Adjoint ()
1 For and ,) BÁ ²=Á>³ Á-
² b ³ ~ b ddd
2 For and ,)B B² = Á > ³ ² > Á < ³
²³ ~ dd d
3 For any invertible ,) B² = ³
²³ ~ ² ³c cdd
Proof. Proof of part 1 is left for the reader. For part 2 , we have for all , )) <i
² ³ ²³ ~ ² ³ ~ ² ³ ~ ² ²³³ ~ ² ³²³ dd d d d d
Part 3 follows from part 2 and ))
dc d c d d²³ ~ ² ³ ~~
and in the same way, . Hence . ²³ ~ ²³ ~ ² ³ c d d c d d c
If , then and so . Of course,B B B² = Á > ³ > Á = ² =Á >³di i d d i i i i()
dd is not equal to . However, in the finite-dimensional case, if we use the
natural maps to identify with a nd with , then we can think of == >>ii ii dd
The Isomorphism Theorems 105
as being in . Using these identifications, we do have equality in the B²= Á>³
finite-dimensional case.
Theorem 3.18 Let and be finite-dimensional and let . If we => ² = Á > ³ B
identify with and with using the natural maps, then is== >>ii ii dd
identified with .
Proof. For any let the corresponding element of be denoted by and %= = %ii
similarly for . Then before making any identifications, we have for , ># =
dd d²#³²³ ~ #´ ²³µ ~ #² ³ ~ ² #³ ~ #²³
for all and so>*
dd ii²#³ ~ # >
Therefore, using the canonical identifications for both and we have =>ii ii
dd²#³ ~ #
for all .#=
The next result describes the kerne l and image of the operator adjoint.
Theorem 3.19 Let . ThenB² = Á > ³
1)i mker²³ ~² ³d
2)i m²³ ~ ² ³dker
Proof. For part 1 , )
ker²³ ~ ¸ > ² ³ ~ ¹
~ ¸ > ² =³ ~ ¸¹¹
~ ¸ > ² ² ³³ ~ ¸¹¹
~² ³
di d
i
i
im
im
For part 2 , if , then and so )i m~ ~ ² ³ ²³ ² ³ ddker ker
²³ker.
For the reverse inclusion, let . We wish to show that ²³ =keri
~ ~ > 2~ ²³ di for some . On , there is no problem since ker
and agree on for any . Let be a complement of . di~ 2 > : ² ³ ker
Then maps a basis for to a linearly independent set8 ~¸ 0¹ :
8 ~¸ 0¹
in and so we can define on by setting> >i8
² ³ ~
and extending to all of . Then on and therefore on . Thus, > ~ ~ : 8d
~ ² ³ddim .
106 Advanced Linear Algebra
Corollary 3.20 Let , where and are finite-dimensional.B² = Á > ³ = >
Then . rk rk²³ ~ ² ³d
In the finite-dimensional case, and can both be represented by matrices. d
Let
89~² ÁÃÁ³ ~² ÁÃÁ ³ and
be ordered bases for and , respectively, and let =>
89ii i ii i
~² ÁÃÁ³ ~² ÁÃÁ ³ and
be the corresponding dual bases. Then
² ´ µ ³ ~² ´ µ³ ~´ µ89 9Á Á i
and
²´ µ ³ ~ ²´ ² ³µ ³ ~ ´ ² ³µ ~ ² ³² ³ ~ ² ³ dd i i i d i d i i
Á Á 98 8ii i
Comparing the last two expressions we see that they are the same except that the
roles of and are reversed. Hence, the matrices in question are transposes.
Theorem 3.21 Let , where and are finite-dimensional. If B 8² = Á > ³ = >
and are ordered bases for and , respectively, and and are the98 9 =>ii
corresponding dual bases, then
´µ ~ ² ´ µ³d!
ÁÁ98 8 9ii
In words, the matrices of and its ope rator adjoint are transposes of one d
another.
Exercises
1. If is infinite-dimensional and is an infinite-dimensional subspace, must=:
the dimension of be finite? Explain. =°:
2. Prove the correspondence theorem.
3. Prove the first isomorphism theorem.
4. Complete the proof of Theorem 3.9.
5. Let be a subspace of . Starting with a basis for how: = ¸ ÁÃÁ ¹ :Á
would you find a basis for ? =°:
6. Use the first isomorphism theorem to prove the rank-plus-nullity theorem
rk null²³ b ²³ ~ ² = ³ dim
for and .B² = Á > ³ ² = ³ B dim
7. Let and suppose that is a subspace of . Define a mapB² = ³ : =
Z¢=°: ¦ =°: by
The Isomorphism Theorems 107
Z²#b:³ ~ #b:
When is well-defined? If is well-defined, is it a linear transformation?ZZ
What are and ? im²³ ²³ZZker
8. Show that for any nonzero vector , there exists a linear functional #=
= ² # ³£i for which .
9. Show that a vector is zero if and only if for all . #= ² # ³~ =i
10. Let be a proper subspace of a finite-dimensional vector space and let:=
#=±: = . Show that there is a linear functional for whichi
²#³~ ² ³~ : and for all .
11. Find a vector space and decompositions =
=~ ( l )~ *l +
with but . Hence, does imply that .(* )+ (* ( * ° not
12. Find isomorphic vectors spaces and with =>
=~ :l ) >~ :l + and
but . Hence, does imply that .)+ => =° :>° :° not
13. Let be a vector space with=
=~ :l ;~ :l ;
Prove that if and have finite codimension in , then so does :: = : q :
and
codim²: q: ³ ²; ³b ²; ³ dim dim
14. Let be a vector space with=
=~ :l ;~ :l ;
Suppose that and have finite codimension. Hence, by the previous ::
exercise, so does . Find a direct sum decomposition :q : =~ > l ?
for which 1 has finite codimension, 2 and 3 () () ()>> : q :
?; b; .
15. Let be a basis for an infinite-dimensional vector space and define, for8 =
all , the map by if and otherwise, for all = ² ³~ ~ 8Zi Z
¸ ¹ =88. Does form a basis for ? What do you conclude aboutZi
the concept of a dual basis?
16. Prove that if and are subspaces of , then . :; = ² : l ; ³ : ;iii^
17. Prove that and where is the zero linear operator and is ~ ~ dd
the identity.
18. Let be a subspace of . Prove that .:= ² = ° : ³ :i
19. Verify that
a for . )²b³~ b Á ² = Á > ³ Bddd
b for any and )² ³ ~ - ²= Á>³ Bdd
20. Let , where and are finite-dimensional. Prove thatB² = Á > ³ = >
rk rk²³ ~ ² ³d.
Chapter 4
Modules I: Basic Properties
Motivation
Let be a vector space over a field and let . Then for any=- ² = ³ B
polynomial , the operator is well-defined. For instance, if ²%³ -´%µ ² ³
²%³ ~ b%b%, then
² ³ ~ b b
where is the identity operator and is the threefold composition . kk
Thus, using the operator we can define the product of a polynomial
² % ³-´ % µ #= and a vector by
²%³# ~ ² ³²#³ ()4.1
This product satisfies the usual properties of scalar multiplication, namely, for
all and ,²%³Á ²%³ -´%µ "Á# =
²%³²"b#³ ~ ²%³"b²%³#
²²%³b ²%³³" ~ ²%³"b ²%³"
´²%³ ²%³µ" ~ ²%³´ ²%³"µ
" ~ "
Thus, for a fixed , we can think of as being endowed with the B² = ³ =
operations of addition and multiplication of an element of by a in = polynomial
-´%µ -´%µ = . However, since is not a field, these two operations do not make
into a vector space. Nevertheless, the situation in which the scalars form a ring
but not a field is extremely important, not only in this context but in many
others.
Modules
Definition Let be a commutative ring with identity, whose elements are9
called . An or a is a nonempty set , scalars -module module over 99 4 ()
110 Advanced Linear Algebra
together with two operations. Th e first operation, called and denoted addition
by , assigns to each pair , an element . Theb² " Á # ³ 4 d 4 " b # 4
second operation, denoted by juxtaposition, assigns to each pair
²Á#³ 9 d4 # 4 , an element . Furthermore, the following properties
must hold:
1 is an abelian group under addition.)4
2 For all and ) Á 9 "Á# 4
²"b#³~"b#
²b ³" ~ "b "
² ³" ~ ² "³
" ~ "
The ring is called the of . 94 base ring
Note that vector spaces are just special types of modules: a vector space is a
module over a field.
When we turn in a later chapter to the study of the structure of a linear
transformation , we will think of as having the structure of a vector B² = ³ =
space over as well as a module over and we will use the notation . Put -- ´ % µ =
another way, is an abelian group under addition, with two scalar =
multiplications—one whose scalars are elements of and one whose scalars are -
polynomials over . This viewpoint will be of tremendous benefit for the study -
of . For now, we concentrate only on modules.
Example 4.1
1 If is a ring, the set of all ordered -tuples whose components lie in )99 9
is an -module, with addition and scalar multiplication defined9
componentwise just as in , () -
² ÁÃÁ ³b² ÁÃÁ ³~² b ÁÃÁ b ³
and
² ÁÃÁ ³ ~ ² ÁÃÁ ³
for , . For example, is the -module of all ordered -tuples Á 9 {{
of integers.
2 If is a ring, the set of all matrices of size is an -)9² 9 ³ d 9 CÁ
module, under the usual operations of matrix addition and scalar
multiplication over . Since is a ring, we can also take the product of 99
matrices in . One important example is , whence CÁ²9³ 9 ~ -´%µ
CÁ²-´%µ³ -´%µ d is the -module of all matrices whose entries are
polynomials.
3 Any commutative ring with identity is a module over itself, that is, is) 99
an -module. In this case, scalar multiplication is just multiplication by9
Modules I: Basic Properties 111
elements of , that is, scalar multip lication is the ring multiplication. The 9
defining properties of a ring imply th at the defining properties of the - 9
module are satisfied. We shall use this example many times in the9
sequel.
Importance of the Base Ring
Our definition of a module requires that the ring of scalars be commutative. 9
Modules over noncommutative rings can exhibit quite a bit more unusual
behavior than modules over commutative rings. Indeed, as one would expect,
the general behavior of -modules improves as we impose more structure on 9
the base ring . If we impose the very strict structure of a field, the result is the 9
very well behaved vector space.
To illustrate, we will give an example of a module over a ring noncommutative
that has a basis of size for every integer ! As another example, if the
base ring is an integral domain, then whenever are linearly #Á Ã Á #
independent over so are for any nonzero . This can fail 9 # Á Ã Á # 9
when is not an integral domain.9
We will also consider the property on the base ring that all of its ideals are 9
finitely generated. In this case, any finitely generated -module has the 94
property that all of its submodules are al so finitely generated. This property of
99-modules fails if does not have the stated property.
When is a principal ideal domain such as or , each of its ideals is9- ´ % µ (){
generated by a single element. In this case, the -modules are “reasonably” well 9
behaved. For instance, in general, a module may have a basis and yet possess a
submodule that has no basis. However, if is a principal ideal domain, this9
cannot happen.
Nevertheless, even when is a principal ideal domain, -modules are less well 99
behaved than vector spaces. For example, there are modules over a principal
ideal domain that do not have any linear ly independent elements. Of course,
such modules cannot have a basis.
Submodules
Many of the basic concepts that we defined for vector spaces can also be
defined for modules, although their proper ties are often quite different. We
begin with submodules.
Definition A of an -module is a nonempty subset of thatsubmodule 94 : 4
is an -module in its own right, under the operations obtained by restricting the 9
operations of to . We write to denote the fact that is a submodule 4: : 4 :
of .4
112 Advanced Linear Algebra
Theorem 4.1 A nonempty subset of an -module is a submodule if and :9 4
only if it is closed under the taking of linear combinations, that is,
Á 9Á"Á# : ¬ "b # :
Theorem 4.2 If and are submodules of , then and are also:; 4 : q ;: b ;
submodules of . 4
We have remarked that a commutative ring with identity is a module over 9
itself. As we will see, this type of module provides some good examples of non-
vector-space-like behavior.
When we think of a ring as an -module rather than as a ring, multiplication 99
is treated as multiplication. This has some important implications. In scalar
particular, if is a submodule of , then it is closed under scalar multiplication, :9
which means that it is closed under multiplication by elements of the ring . all 9
In other words, is an ideal of the ring . Conversely, if is an ideal of the :9 ?
ring , then is also a submodule of the module . Hence, 99? the submodules of
the -module are precisely the ideals of the ring .99 9
Spanning Sets
The concept of spanning set carries over to modules as well.
Definition The or by a subset of a module submodule spanned generated () :
4: is the set of all of elements of : linear combinations
º º:» »~¸# bÄb # 9Á# :Á¹
A subset is said to or if . :4 4 4 4~º º : » » span generate
We use a double angle bracket notation for the submodule generated by a set
because when we study the -vector space/ -module , we will need to -- ´ % µ =
make a distinction between the subspace generated by and the º#» ~ -# # =
submodule generated by . ºº#»» ~ - ´%µ# #
One very important point to note is that if a nontrivial linear combination of the
elements in an -module is , #Á Ã Á # 9 4
# bÄb # ~
where not all of the coefficients are , then we conclude, as we could in cannot
a vector space, that one of the elements is a linear combination of the others. #
After all, this involves dividing by one of the coefficients, which may not be
possible in a ring. For instance, for the -module we have {{ { d
²Á
³c²Á³ ~ ²Á³
but neither nor is an integer multiple of the other. ²Á
³ ²Á³
Modules I: Basic Properties 113
The following simple submodules play a special role in the theory.
Definition Let be an -module. A submodule of the form49
ºº#»» ~ 9# ~ ¸# 9¹
for is called the generated by .#4 # cyclic submodule
Of course, any finite-dimensional vector space is the direct sum of cyclic
submodules, that is, one-dimensional subspaces. One of our main goals is to
show that a finitely generated modul e over a principal ideal domain has this
property as well.
Definition An -module is said to be if it contains a94 finitely generated
finite set that generates . More specifically, is if it has a 44 -generated
generating set of size although it may have a smaller generating set as (
well .)
Of course, a vector space is finitely generated if and only if it has a finite basis,
that is, if and only if it is finite-dimensional. For modules, life is more
complicated. The following is an exampl e of a finitely generated module that
has a submodule that is not finitely generated.
Example 4.2 Let be the ring of all polynomials in infinitely9 -´% Á% Áõ
many variables over a field . It will be convenient to use to denote -?
%Á %Á Ã 9 ² ? ³ and write a polynomial in in the form . Each polynomial in (
99, being a finite sum, involves only fin itely many variables, however. Then )
is an -module and as such, is finitely generated by the identity element9
²?³ ~ .
Now consider the submodule of all polynomials with zero constant term. This :
module is generated by the variables themselves,
:~º º %Á %Á Ã » »
However, is not finitely generated. To see this, suppose that :. ~ ¸ Á Ã Á ¹
is a finite generating set for . Choose a variable that does not appear in any :%
of the polynomials in . Then no linear combination of the polynomials in ..
can be equal to . For if %
% ~ ²?³ ²?³
~
114 Advanced Linear Algebra
then let where does not involve . This gives ²?³~% ²?³b ²?³ ²?³ %
% ~ ´% ²?³b ²?³µ ²?³
~ % ²?³ ²?³b ²?³ ²?³
~
~ ~
The last sum does not involve and so it must equal . Hence, the first sum %
must equal , which is not possibl e since has no constant term. ² ? ³
Linear Independence
The concept of linear independence also carries over to modules.
Definition A subset of an -module is if for any :9 4 linearly independent
distinct and , we have# ÁÃÁ# : ÁÃÁ 9
# bÄb # ~¬ ~ for all
A set that is not linearly independent is .: linearly dependent
It is clear from the definition that an y subset of a linearly independent set is
linearly independent.
Recall that in a vector space, a set of vectors is linearly dependent if and only :
if some vector in is a linear comb ination of the other vectors in . For ::
arbitrary modules, this is not true.
Example 4.3 Consider as a -module. The elements are linearly{{ { Á
dependent, since
²³c²³ ~
but neither one is a linear combination i.e., integer multiple of the other. ()
The problem in the previous example as noted earlier is that ()
# bÄb # ~
implies that
# ~c#cÄc#
but in general, we cannot divide bot h sides by , since it may not have a
multiplicative inverse in the ring . 9
Modules I: Basic Properties 115
Torsion Elements
In a vector space over a field , singleton sets where are linearly =- ¸ # ¹ # £
independent. Put another way, and imply . However, in a £ #£ #£
module, this need not be the case.
Example 4.4 The abelian group is a -module, with {{~ ¸ÁÁÃÁc¹
scalar multiplication defined by , for all and . ' ~²'h³ ' mod {{
However, since for all , no singleton set is linearly ~ ¸¹ {
independent. Indeed, has no linearly independent sets. {
This example motivates the following definition.
Definition Let be an -module. A nonzero element for which 4 9 #4 #~
for some nonzero is called a of . A module that has no 9 4 torsion element
nonzero torsion elements is said to be . If all elements of are torsion-free 4
torsion elements, then is a . Th e set of all torsion elements of 4 torsion module
44, together with the zero element, is denoted by . tor
If is a module over an , it is not hard to see that is a44 integral domain tor
submodule of and that is torsion-free. We will define quotient 44 ° 4 tor (
modules shortly: they are defined in the same way as for vector spaces.)
Annihilators
Closely associated with the notion of a torsion element is that of an annihilator.
Definition Let be an -module. The of an element is49 # 4 annihilator
ann² # ³~¸ 9 #~ ¹
and the of a submodule of is annihilator 54
ann²5³ ~ ¸ 9 5 ~ ¸¹¹
where . Annihilators are also called .5 ~ ¸# # 5¹ order ideals
It is easy to see that and are ideals of . Clearly, is a ann ann²#³ ²5³ 9 # 4
torsion element if and only if . Also, if and are submodules of ann²#³ £ ¸¹ ( )
4, then
( ) ¬ ²)³ ²(³ ann ann
(note the reversal of order).
Let be a finitely generated module over an integral domain4~º º "ÁÃÁ"» »
9" and assume that each of the generators is torsion, that is, for each , there is
a nonzero . Then, the nonzero product annihilates each ² " ³ ~ Ä ann
generator of and therefore every element of , that is, . This 44 ² 4 ³ ann
116 Advanced Linear Algebra
shows that . On the other hand, this may fail if is not an ann²4³ £ ¸¹ 9
integral domain. Also, there are torsion m odules whose annihilators are trivial.
()We leave verification of these statements as an exercise.
Free Modules
The definition of a basis for a module parallels that of a basis for a vector space.
Definition Let be an -module. A subset of is a if is linearly49 4 88 basis
independent and spans . An -module is said to be if or if 4 9 4 4 ~ ¸¹ free
44 4 has a basis. If is a basis for , we say that is . 88 free on
We have the following analog of part of Theorem 1.7.
Theorem 4.3 A subset of a module is a basis if and only if every nonzero 8 4
#4 is an essentially unique linear combination of the vectors in . 8
In a vector space, a set of vectors is a basis if and only if it is a minimal
spanning set, or equivalently, a maximal linearly independent set. For modules,
the following is the best we can do in general. We leave proof to the reader.
Theorem 4.4 Let be a basis for an -module . Then8 94
1 is a minimal spanning set.)8
2 is a maximal linearly independent set.)8
The -module has no basis since it has no linearly independent sets. But{{
since the entire module is a spanning set, we deduce that a minimal spanning set
need not be a basis. In the exercises, the reader is asked to give an example of a
module that has a finite basis, but with the property that not every spanning 4
set in contains a basis and not ev ery linearly independent set in is 44
contained in a basis. It follows in th is case that a maximal linearly independent
set need not be a basis.
The next example shows that even free modules are not very much like vector
spaces. It is an example of a free module that has a submodule that is not free.
Example 4.5 The set is a free module over itself, using componentwise{{d
scalar multiplication
²Á³²Á³ ~ ²Á³
with basis . But the submodule is not free since it has no ¸²Á³¹ d¸¹ {
linearly independent elements and hence no basis.
Theorem 2.2 says that a linear transformation can be defined by specifying its
values arbitrarily on a basis. The same is true for modules. free
Modules I: Basic Properties 117
Theorem 4.5 Let and be -modules where is free with basis45 9 4
8~¸ 0¹ 9 ¢4¦5 . Then we can define a unique -map by specifying
the values of arbitrarily for all and then extending to by 8 4
linearity, that is,
²# bÄb # ³~ # bÄb #
Homomorphisms
The term is special to vector spaces. However, the linear transformation
concept applies to most algebraic structures.
Definition Let and be -modules. A function is an -45 9 ¢ 4 ¦ 5 9
homomorphism -map or if it preserves the module operations, that is,9
²"b #³ ~ ²"³b ²#³
for all and . The set of all -homomorphisms from to isÁ 9 "Á# 4 9 4 5
denoted by . The following terms are also employed: hom9²4Á5³
1 An - is an -homomorphism from to itself.)994endomorphism
2 An - or - is an injective -homomorphism.)99 9monomorphism embedding
3 An - is a surjective -homomorphism.)99epimorphism
4 An - is a bijective -homomorphism.)99isomorphism
It is easy to see that is itself an -module under addition of hom9²4Á5³ 9
functions and scalar multiplication defined by
² ³²#³ ~ ² #³ ~ ²#³
Theorem 4.6 Let . The kernel and image of , defined as for² 4 Á 5 ³hom9
linear transformations by
ker² ³~¸ #4 #~ ¹
and
im² ³~¸ ##4¹
are submodules of and , respectively. Moreover, is a monomorphism if 45
and only if . ker² ³ ~ ¸¹
If is a submodule of the -module , then the map defined by59 4 ¢ 5 ¦ 4
²#³ ~ # 9 5 4 is evidently an -monomorphism, called of into . injection
Quotient Modules
The procedure for defining quotient modules is the same as that for defining
quotient vector spaces. We summarize in the following theorem.
118 Advanced Linear Algebra
Theorem 4.7 Let be a submodule of an -module . The binary relation:9 4
"#¯"c#:
is an equivalence relation on , whose equivalence classes are the 4 cosets
#b:~¸ #b :¹
of in . The set of all co sets of in , called the of :4 4 ° : :4 quotient module
4: 9 , is an -module under the well-defined operationsmodulo
²"b:³b²#b:³~²"b#³b:
²"b:³~"b:
The zero element in is the coset . 4°: b: ~:
One question that immediately comes to mind is whether a quotient module of a
free module must be free. As the next example shows, the answer is no.
Example 4.6 As a module over itself, is free on the set . For any , { ¸¹
the set is a free cyclic submodule of , but the quotient -{{ { {~¸ '' ¹
module is isomorphic to via the map{{ {°
{²"b ³ ~ " mod
and since is not free as a -module, neither is .{{ { { °
The Correspondence and Isomorphism Theorems
The correspondence and isomorphism theorems for vector spaces have analogs
for modules.
Theorem 4.8 The correspondence theorem() Let be a submodule of .:4
Then the function that assigns to each intermediate submodule the :;4
quotient submodule of is an order-preserving with respect to set ;°: 4°: (
inclusion one-to-one correspondence be tween submodules of containing ) 4:
and all submodules of . 4°:
Theorem 4.9 The Let be an -() first isomorphism theorem ¢4¦5 9
homomorphism. Then the map defined by Z¢4° ² ³¦5 ker
Z²#b ² ³³ ~ # ker
is an -embedding and so9
4
²³² ³kerim
Modules I: Basic Properties 119
Theorem 4.10 The Let be an -module() second isomorphism theorem 49
and let and be submodules of . Then:; 4
:b; :
;: q ;
Theorem 4.11 The Let be an -module and() third isomorphism theorem 49
suppose that are submodules of . Then :; 4
4°: 4
;°: ;
Direct Sums and Direct Summands
The definition of direct sum of a family of submodules is a direct analog of the
definition for vector spaces.
Definition The of -modules , denoted by external direct sum 9 4 ÁÃÁ4
4~4 Ä 4 ^^
is the -module whose elements are ordered -tuples
4~¸²# ÁÃÁ# ³# 4Á~ÁÃÁ¹
with componentwise operations
²" ÁÃÁ" ³b²# ÁÃÁ# ³~²" b# ÁÃÁ" b# ³
and
²# ÁÃ Á# ³ ~ ²# ÁÃ Á# ³
for .9
We leave it to the reader to formulate the definition of external direct sums and
products for arbitrary families of modules, in direct analogy with the case of
vector spaces.
Definition An -module is the of a family94 ()internal direct sum
<~¸ : 0¹ 4 of submodules of , written
4~ 4~ : <or
0
if the following hold:
1 is the sum join of the family :)( ) ( )Join of the family 4 <
=~ :
0
120 Advanced Linear Algebra
2 For each ,)( )Independence of the family 0
: q : ~ ¸¹
£ps
qt
In this case, each is called a of . If is :4 ~ ¸ : Á Ã Á : ¹ direct summand <
a finite family, the direct sum is often written
4~:lÄl:
Finally, if , then is said to be and is called a 4~:l; : ; complemented
complement of in .:4
As with vector spaces, we have the following useful characterization of direct
sums.
Theorem 4.12 Let be a family of distinct submodules of an -<~¸ : 0¹ 9
module . The following are equivalent:4
1 For each ,)( )Independence of the family 0
: q : ~ ¸¹
£ps
qt
2 The zero element cannot be written as)( )Uniqueness of expression for
a sum of nonzero elements from distinct submodules in . <
3 Every nonzero has a unique, except for)( )Uniqueness of expression #4
order of terms, expression as a sum
#~ bÄb
of nonzero elements from distinct submodules in . <
Hence, a sum
4~ :
0
is direct if and only if any one of 1 3 holds. )– )
In the case of vector spaces, every subspace is a direct summand, that is, every
subspace has a complement. However, as the next example shows, this is not
true for modules.
Example 4.7 The set of integers is a -module. Since the submodules of {{ {
are precisely the ideals of the ring and since is a principal ideal domain, the {{
submodules of are the sets {
º º » »~ ~¸ '' ¹{{
Modules I: Basic Properties 121
Hence, any two nonzero proper submodules of have nonzero intersection, for {
if , then£
{{ {q ~
where . It follows that the only complemented submodules of ~ ¸ Á ¹ lcm {
are and .{¸¹
In the case of vector spaces, there is an intimate connection between subspaces
and quotient spaces, as we saw in Theorem 3.6. The problem we face in
generalizing this to modules is that not all submodules are complemented.
However, this is the only problem.
Theorem 4.13 Let be a complemented submodule of . All complements of:4
:4 ° : are isomorphic to and hence to each other.
Proof. For any complement of , the first isomorphism theorem applied to ;:
the projection gives . ;Á:¢4¦; ; 4°:
Direct Summands and Extensions of Isomorphisms
Direct summands play a role in questi ons relating to whether certain module
homomorphisms can be extended from a submodule to the ¢5¦4 54
full module . The discussion will be a bit simpler if we restrict attention to 4
epimorphisms.
If , then a module epimorphism can be extended to an4~5l/ ¢5¦4
epimorphism simply by sending the elements of to zero, that is, ¢4¦4 /
by setting
²b³ ~
This is easily seen to be an -map with 9
ker ker²³ ~ ²³ l /
Moreover, if is another extension of with the same kernel as , then and
agree on as well as on , whence . Thus, there is a extension of /5 ~ unique
with kernel . ker²³ l /
Now suppose that is an isomorphism. If is complemented, that is, ¢54 5
if
.~5l/
then we have seen that there is a extension of for which . unique ker²³ ~ /
Thus, the correspondence
/ª ² ³~/, where ker
from complements of to extensions of is an injection. To see that this 5
correspondence is a bijection, if is an extension of , then ¢4¦4
122 Advanced Linear Algebra
4~5l ²³ ker
To see this, we have
5 q ² ³ ~ ² ³ ~ ¸¹ ker ker
and if , then there is a for which and so4 5 ~
²c³ ~ c ~
Thus,
~b² c ³5b ² ³ ker
which shows that is a complement of . ker²³ 5
Theorem 4.14 Let and be -modules and let .44 9 5 4
1 If , then any -epimorphism has a unique)4~5l/ 9 ¢5¦4
extension to an epimorphism with¢4¦4
ker ker²³ ~ ²³ l /
2 Let be an -isomorphism. Then the correspondence)¢54 9
/ª ² ³~/, where ker
is a bijection from complements of onto the extensions of . Thus, an5
isomorphism has an extension to if and only if is ¢54 4 5
complemented.
Definition Let . When the identity map has an extension to54 ¢55
¢4¦5 5 4 , the submodule is called a of and is called the retract
retraction map .
Corollary 4.15 A submodule is a retract of if and only if has a 54 4 5
complement in . 4
Direct Summands and One-Sided Invertibility
Direct summands are also related to one-sided invertibility of -maps. 9
Definition Let be a module homomorphism.¢(¦)
1 A of is a module homomorphism for which) left inverse 3¢)¦(
3k~ .
2 A of is a module homomorphism for which) right inverse 9¢)¦(
k~9 .
Left and right inverses are called . An ordinary inverse is one-sided inverses
called a . two-sided inverse
Unlike a two-sided inverse, one-sided inverses need not be unique.
Modules I: Basic Properties 123
A left-invertible homomorphism must be injective, since
~ ¬ k ~ k ¬~ 33
Also, a right-invertible homomorphism must be surjective, since if¢(¦)
) , then
~ ´ ² ³ µ ² ³ 9 im
For functions, the converses of these stat ements hold: is left-invertible if set
and only if it is injective and is right -invertible if and only if it is surjective.
However, this is not the case for -maps. 9
Let be an injective -map. Referring to Figure 4.1,¢4¦4 9
M M1im(V)H
V_im(V))-1V_im(V)
Figure 4.1
the map obtained from by restricting its range to is O ¢ 4 ²³ ²³im²³im im
an isomorphism and the left inverses of are precisely the extensions of3
²O ³ ¢ ²³ 4 4im²³c im to . Hence, Theorem 4.14 says that the
correspondence
/ª ² O ³ / extension of with kernel im²³c
is a bijection from the complements of onto the left inverses of ./² ³ im
Now let be a surjective -map. Referring to Figure 4.2,¢4¦4 9
M M1ker(V)
HV|H
VR=(V|H)-1
Figure 4.2
if is complemented, that is, ifker²³
4~ ²³l/ ker
124 Advanced Linear Algebra
then is an isomorphism. Thus, a map is a rightO¢ / 4 ¢ 4¦ 4/
inverse of if and only if is a of , the only range-extension ²O³ ¢ 4 //c
difference being in the ranges of the two functions. Hence, ²O³ ¢ 4¦ 4/c
is the only right inverse of with im age . It follows that the correspondence /
/ª² O ³ ¢4 ¦4/c
is an injection from the complements of to the right inverses of ./² ³ ker
Moreover, this map is a bijection, since if is a right inverse of , 9¢4 ¦4
then and is an extension of , which 9 9 9 9c¢4 ² ³ ¢ ² ³4 im im
implies that
4~ ² ³l ²³ im9 ker
Theorem 4.16 Let and be -modules and let be an -map.44 9 ¢ 4 ¦ 4 9
1 Let be injective. The map)Æ¢4 4
/ª ² O ³ / extension of with kernel im²³c
is a bijection from the complements of onto the left inverses of ./² ³ im
Thus, there is exactly one left inverse of for each complement of im²³
and that complement is the kernel of the left inverse.
2 Let be surjective. The map)¢4¦4
/ª² O ³ ¢4 ¦4/c
is a bijection from the complements of to the right inverses of ./² ³ ker
Thus, there is exactly one right inverse of for each complement of /
ker²³ and that complement is the im age of the right inverse. Thus,
4~ ²³l/ ²³ ²³ ker ker ^ im
The last part of the previous theorem is worth further comment. Recall that if
¢= ¦> is a linear transformation on vector spaces, then
= ²³ ²³ ker^ im
This holds for modules as well . provided that is a direct summand ker²³
Modules Are Not as Nice as Vector Spaces
Here is a list of some of the properties of modules over commutative rings with (
identity that emphasize the differences between modules and vector spaces. )
1 A submodule of a module need not have a complement.)
2 A submodule of a finitely generated m odule need not be finitely generated. )
3 There exist modules with no linearly independent elements and hence with)
no basis.
4 A minimal spanning set or maximal linearly independent set is not)
necessarily a basis.
Modules I: Basic Properties 125
5 There exist free modules with submodules that are not free.)
6 There exist free modules with linearly independent sets that are not)
contained in a basis and spanning sets that do not contain a basis.
Recall also that a module over a ring may have bases of noncommutative
different sizes. However, all bases for a free module over a commutative ring
with identity have the same size, as we will prove in the next chapter.
Exercises
1. Give the details to show that an y commutative ring with identity is a
module over itself.
2. Let be a subset of a module . Prove that is:~¸# ÁÃÁ# ¹ 4 5~º º:» »
the submodule of containing . First you will need to formulatesmallest 4:
precisely what it means to be the smallest submodule of containing . 4:
3. Let be an -module and let be an ideal in . Let be the set of all49 0 9 0 4
finite sums of the form
# bÄb #
where and . Is a submodule of ? 0 # 4 0 4 4
4. Show that if and are submodules of , then with respect to set :; 4 (
inclusion)
:q;~ ¸ :Á;¹ :b;~ ¸ :Á;¹ glb lub and
5. Let be an ascending sequence of submodules of an - : : Ä 9
module . Prove that the union is a submodule of .4: 4
6. Give an example of a module that ha s a finite basis but with the property 4
that not every spanning set in cont ains a basis and not every linearly4
independent set in is contained in a basis. 4
7. Show that, just as in the case of vector spaces, an -homomorphism can be 9
defined by assigning arbitrary values on the elements of a basis and
extending by linearity.
8. Let be an -isomorphism. If is a basis for , prove8² 4 Á 5 ³ 9 4hom9
that is a basis for .8 8~¸ ¹ 5
9. Let be an -module and let be an -endomorphism.4 9 ²4Á4³ 9 hom9
If is , that is, if , show that idempotent~
4~ ²³l ²³ ker im
Does the converse hold?
10. Consider the ring of polynomials in two variables. Show that 9~-´ % Á& µ
the set consisting of all polynomials in that have zero constant term is 49
an -module. Show that is not a free -module.94 9
11. Prove that if is an integral domain, then all -modules have the 99 4
following property: If is linearly independent over , then so is #Á Ã Á # 9
# ÁÃÁ# 9 for any nonzero .
126 Advanced Linear Algebra
12. Prove that if a nonzero commutative ri ng with identity has the property9
that every finitely generated -module is free then is a field. 99
13. Let and be -modules. If is a submodule of and is a 45 9 : 4;
submodule of show that 5
4l5 4 5
:l; : ;^
14. If is a commutative ring with identity and is an ideal of , then is an 99 ??
9-module. What is the maximum size of a linearly independent set in ? ?
Under what conditions is free? ?
15. a the set of all ) Show that for any module over an integral domain 44 tor
torsion elements in a module is a submodule of . 44
b Find an example of a ring with the property that for some -module ) 99
44 the set is not a submodule. tor
c ) Show that for any module over an integral domain, the quotient 4
module is torsion-free.4°4 tor
16. a Find a module that is finitely generated by torsion elements but for ) 4
which . ann²4³ ~ ¸¹
b Find a torsion module for which . )a n n 4 ²4³ ~ ¸¹
17. Let be an abelian group together with a scalar multiplication over a ring5
99 # that satisfies all of the properties of an -module except that does not
necessarily equal for all . Show that can be written as a direct ## 5 5
sum of an -module and another “pseudo -module” . 95 9 5
18. Prove that is an -module under addition of functions and hom9²4Á5³ 9
scalar multiplication defined by
² ³²#³ ~ ² #³ ~ ²#³
19. Prove that any -module is isomorphic to the -module . 94 9 ² 9 Á 4 ³ hom9
20. Let and be commutative rings with identity and let be a ring9: ¢ 9 ¦ :
homomorphism. Show that any -module is also an -module under the :9
scalar multiplication
# ~ ²³#
21. Prove that where . hom gcd{²Á ³ ~ ² Á ³{{ {
22. Suppose that is a commutative ring with identity. If and are ideals of 9 ?@
99 ° 9 ° 9 ~ for which as -modules, then prove that . Is the ?@ ? @
result true if as rings? 9° 9°?@
Chapter 5
Modules II: Free and Noetherian Modules
The Rank of a Free Module
Since all bases for a vector space have the same cardinality, the concept of =
vector space dimension is well-defined. A similar statement holds for free - 9
modules when the base ring is commutative but not otherwise . ()
Theorem 5.1 Let be a free module over a commutative ring with identity.49
1 Then any two bases of have the same cardinality.) 4
2 The cardinality of a spanning set is greater than or equal to that of a basis.)
Proof. The plan is to find a vector space with the property that, for any basis =
for , there is a basis of the same cardi nality for . Then we can appeal to the 4=
corresponding result for vector spaces.
Let be a maximal ideal of , which exists by Theorem 0.23. Then is a?? 99 °
field. Our first thought might be that is a vector space over , but that is 49 ° ?
not the case. In fact, scalar multiplication using the field , 9°?
²b ³# ~ #?
is not even well-defined, since this would require that . On the other ?4 ~ ¸¹
hand, we can fix precisely this probl em by factoring out the submodule
??4~¸ #bÄb# Á#4 ¹
Indeed, is a vector space over , with scalar multiplication defined4° 4 9°??
by
² b³ ² " b4 ³ ~ " b4?? ?
To see that this is well-defined, we must show that the conditions
b ~ b
"b 4~" b 4??
??Z
Z
imply
128 Advanced Linear Algebra
"b 4 ~ " b 4??ZZ
But this follows from the fact that
"c " ~ ²"c" ³b²c ³" 4ZZ Z Z Z?
Hence, scalar multiplication is well-defined. We leave it to the reader to show
that is a vector space over .4° 4 9°??
Consider now a set and the corresponding set 8~¸ 0¹4
8? ??b4 ~ ¸ b4 0 ¹ 4
4
If spans over , then spans over . To see this, note88 ? ? ?49 b 4 4 ° 49 °
that any has the form for and so#4 #~ 9 '
#b 4~ b 4
~ ² b 4 ³
~ ² b ³² b 4³??
?
??89
which shows that spans . 8? ?b4 4 °4
Now suppose that is a basis for over . We show that 8~¸ 0¹ 4 9
8? ? ? 8?b4 4 °4 9 ° b4 is a basis for over . We have seen that spans
4° 4?. Also, if
² b ³² b 4³ ~ 4???
then and so 4?
~
where . From the linear independence of we deduce that for all ?8 ?
b ~ b 4 and so . Hence is linearly independent and therefore a ?? 8?
basis, as desired.
To see that , note that if , then (( ( (88 ? ? ?~ b4 b4 ~ b4
c ~
where . If , then the coefficient of on the right must be equal to £ ?
Modules II: Free and Noetherian Modules 129
and so , which is not possible since is a maximal ideal. Hence, ??
~ .
Thus, if is a basis for over , then8 49
(( ( (88 ? ?~b 4 ~ ² 4 ° 4 ³ dim9°?
and so all bases for over have the same cardinality, which proves part 1 . 49 )
Finally, if spans over , then spans and so88 ? ?49 b 4 4 ° 4
dim9°?²4° 4³ b 4 ?8 ?8 (( ( (
Thus, has cardinality at least as great as that of any basis for over .8 49
The previous theorem allows us to define the of a free module. The term rank (
dimension is not used for modules in general.)
Definition Let be a commutative ring with identity. The of a94 ³ rank rk²
nonzero free -module is the cardinality of any basis for . The rank of the 94 4
trivial module is . ¸¹
Theorem 5.1 fails if the underlying ring of scalars is not commutative. The next
example describes a module over a noncommutative ring that has the
remarkable property of possessing a basi s of size for any positive integer .
Example 5.1 Let be a vector space over with a countably infinite basis=-
8B~¸ Á Á Ã ¹ ² =³ = . Let be the ring of linear operators on . Observe that
B²= ³ is not commutative, since composition of functions is not commutative.
The ring is an -module and as such, the identity map forms a basisBB ²= ³ ²= ³
for . However, we can also construct a basis for of any desired finiteBB²= ³ ²= ³
size . To understand the idea, consider the case and define the operators ~
and by
b ² ³ ~ Á ² ³ ~
and
b ² ³ ~ Á ² ³ ~
These operators are linearly independent essentially because they are surjective
and their supports are disjoint. In particular, if
b ~
then
~ ² b ³² ³ ~ ² ³
130 Advanced Linear Algebra
and
~ ² b ³² ³ ~ ² ³ b
which shows that and . Moreover, if , then we define ~ ~ ² =³ B
and by
² ³~² ³
² ³ ~ ² ³
b
from which it follows easily that
~ b
which shows that is a basis for . ¸Á¹ ² = ³ B
More generally, we begin by partitioning into blocks. For each 8
~ ÁÃÁc , let
8 ~¸ ¹ mod
Now we define elements by B ² = ³
b ! ! Á ² ³ ~
where and where is the Kronecker delta function. These functions! !Á
are surjective and have disjoint support. It follows that is 9 c ~¸ ÁÃÁ ¹ 0
linearly independent. For if
~ bÄb c c
where , then, applying this to givesB b !² = ³
~ ² ³~ ² ³ !! b ! !
for all . Hence, .~ !
Also, spans , for if , we define by9B B B ²= ³ ²= ³ ²= ³
² ³ ~ ² ³ +
to get
² bÄb ³² ³ ~ ² ³ ~ ² ³ ~ ² ³ c c ! !! ! ! ! ++ +
and so
~b Ä b c c
Thus, is a basis for of size .9 B c ~¸ ÁÃÁ ¹ ² =³ 0
Recall that if is a basis for a vector space over , then is isomorphic to )= - =
the vector space of all functions from to that have finite support. A ²- ³ ) -)
Modules II: Free and Noetherian Modules 131
similar result holds for free -modules. We begin with the fact that is a 9² 9 ³)
free -module. The simple proof is left to the reader.9
Theorem 5.2 Let be any set and let be a commutative ring with identity. )9
The set of all functions from to that have finite support is a free -²9 ³ ) 9 9)
module of rank with basis where (() 8~¸ ¹
²%³ ~% ~
% £ Fif
if
This basis is referred to as the for . standard basis ²9 ³)
Theorem 5.3 Let be an -module. If is a basis for , then is 49 ) 4 4
isomorphic to . ²9 ³)
Proof. Consider the map defined by setting ¢4¦²9 ³)
~
where is defined in Theorem 5.2 and extending to by linearity. Since 4
maps a basis for to a basis for , it follows that is an 4~ ¸ ¹ ² 9 ³ 8 )
isomorphism from to . 4² 9 ³)
Theorem 5.4 Two free -modules over a commutative ring are isomorphic if 9 ()
and only if they have the same rank.
Proof. If , then any isomorphism from to maps a basis for to45 4 5 4
a basis for . Since is a bijection, we have . Conversely, 5² 4 ³ ~ ² 5 ³ rk rk
suppose that . Let be a basis for and let be a basis for . rk rk²4³ ~ ²5³ 4 5 89
Since , there is a bijective map . This map can be extended by (( ( (89 8 9~¢ ¦
linearity to an isomorphism of onto and so . 45 4 5
We have seen that the cardinality of a minimal spanning set for a free module ()
4² 4 ³ is at least equal to . Let us now speak about the cardinality of maximal rk
linearly independent sets.
Theorem 5.5 Let be an integral domain and let be a free -module. Then94 9
all linearly independent sets have cardinality at most . rk²4³
Proof. Since we need only prove the result for . Let be the4² 9³ ² 9³ 8
field of quotients of . Then is a vector space. Now, if 9² 8 ³
8~¸ # 0¹² 9³ ² 8³
is linearly independent over as a subset of , then is clearly linearly 8² 8 ³8
independent over as a subset of . Conversely, suppose that is linearly 9² 9 ³ 8
independent over and 9
#b Ä b #~
132 Advanced Linear Algebra
where for all and for some . Multiplying by £ £ ~ Ä £
produces a nontrivial linear dependency over , 9
# bÄb # ~
which implies that for all . Thus is linearly dependent over if and ~ 9 8
only if it is linearly dependent over . But in the vector space , all sets of 8² 8 ³
cardinality greater than are linearly de pendent over and hence all subsets of 8
²9 ³ 9 of cardinality greater than are linearly dependent over .
Free Modules and Epimorphisms
If is a module epimorphism where is free on , then it is easy to8¢4¦- -
define a right inverse for , since we can define an -map by 9¢ - ¦ 4 9
specifying its values arbitrarily on and extending by linearity. Thus, we take8
9c²³ ²³ ² ³ to be any member of . Then Theorem 4.16 implies that is a ker
direct summand of and 4
4 ²³ - ker^
This discussion applies to th e canonical projection provided that ¢4¦4°:
the quotient is free. 4°:
Theorem 5.6 Let be a commutative ring with identity.9
1 If is an -epimorphism and is free, then is)¢4¦- 9 - ² ³ ker
complemented and
4~ ²³l5 ²³ - ker ker ^
where .5-
2 If is a submodule of and if is free, then is complemented and ):4 4 ° : :
4:4
:^
If and are free, then4Á: 4°:
rk rk rk²4³ ~ ²:³b4
:67
and if the ranks are all finite, then
rk rk rk674
:~² 4 ³ c² : ³
Noetherian Modules
One of the most desirable properties of a finitely generated -module is that 94
all of its submodules be finitely generated:
Modules II: Free and Noetherian Modules 133
4: 4 ¬ : finitely generated, finitely generated
Example 4.2 shows that this is not always the case and leads us to search for
conditions on the ring that will guarantee this property for -modules. 99
Definition An -module is said to satisfy the 94 ascending chain condition
()abbreviated ACC on submodules if every ascending sequence of submodules
: : : Ä
of is eventually constant, that is, there exists an index for which4
:~ : ~ : ~ Ä b b k
Modules with the ascending chain condition on submodules are also called
Noetherian modules after Emmy Noether, one of the pioneers of module (
theory . )
Since a ring is a module over itself and since the submodules of the module 99
are precisely the ideals of the ring , the preceding definition can be formulated 9
for rings as follows.
Definition A ring is said to satisfy the 9 ascending chain condition
()abbreviated ACC on ideals if any ascending sequence
??? Ä
of ideals of is eventually constant, that is, there exists an index for which 9
?? ? b b~~~ Ä 2
A ring that satisfies the ascending c hain condition on ideals is called a
Noetherian ring .
The following theorem describes the releva nce of this to the present discussion.
Theorem 5.7
1 An -module is Noetherian if and only if every submodule of is)94 4
finitely generated.
2 In particular, a ring is Noetherian if and only if every ideal of is) 99
finitely generated.
Proof. Suppose that all submodules of are finitely generated and that 44
contains an infinite ascending sequence
: : : Ä 3 ()5.1
of submodules. Then the union
134 Advanced Linear Algebra
:~ :
is easily seen to be a submodule of . Hence, is finitely generated, say 4:
:~º º "ÁÃÁ"» » " : " : . Since , there exists an index such that .
Therefore, if , we have ~ ¸ ÁÃÁ ¹ max
¸" ÁÃÁ" ¹:
and so
: ~ º º " Á Ã Á " » » : : : Ä : b b
which shows that the chain 5.1 is eventually constant.()
For the converse, suppose that satisfies the ACC on submodules and let be 4:
a submodule of . Pick and consider the submodule 4" : : ~ º º " » » :
generated by . If , then is finitely generated. If , then there is " : ~: : : £:
a . Now let , . If , then is finitely generated." :c: : ~º º " "» » : ~: :
If , then pick and consider the submodule:£ : " : c : 3
: ~ ºº" " " »»33 ,, .
Continuing in this way, we ge t an ascending chain of submodules
º º "» »º º "Á "» »º º "Á "Á "» »Ä:
If none of these submodules were equal to , we would have an infinite :
ascending chain of submodules, each properly contained in the next, which
contradicts the fact that satisfies the ACC on submodules. Hence, 4
:~º º "ÁÃÁ"» » : for some and so is finitely generated.
Our goal is to find conditions under whic h all finitely generated -modules are 9
Noetherian. The very pleasing answer is that all finitely generated -modules 9
are Noetherian if and only if is Noet herian as an -module, or equivalently, 99
as a ring.
Theorem 5.8 Let be a commutative ring with identity.9
1 is Noetherian if and only if every finitely generated -module is)99
Noetherian.
2 Let be a principal ideal domain. If an -module is -generated, then)99 4
any submodule of is also -generated. 4
Proof. For part 1 , one direction is evident. Assume that is Noetherian and ) 9
let be a finitely generated -module. Consider the4~º º "ÁÃÁ"» » 9
epimorphism defined by ¢9 ¦4
² ÁÃÁ ³~ " bÄb "
Let be a submodule of . Then:4
Modules II: Free and Noetherian Modules 135
c ² :³~¸ "9 ":¹
is a submodule of and . If every submodule of is finitely 9² : ³ ~ : 9c
generated, then is finitely generated and so . c c ²:³ ²:³~º º# ÁÃÁ# » »
Then is finitely generated by . Thus, it is sufficient to prove the :¸ # Á Ã Á # ¹
theorem for , which we do by induction on . 9
If , any submodule of is an ideal of , which is finitely generated by~ 9 9
assumption. Assume that every submodule of is finitely generated for all9
: 9 and let be a submodule of .
If , we can extract from something that is isomorphic to an ideal of : 9
and so will be finitely generated. In particular, let be the “last coordinates” in :
:, specifically, let
: ~ ¸²ÁÃÁÁ ³ ² ÁÃÁ Á ³ : ÁÃÁ 9¹ c c for some
The set is isomorphic to an ideal of and is therefore finitely generated, say :9
: ~ ºº »» ~ ¸ Á à Á ¹ : ==, where is a finite subset of .
Also, let
: ~ ¸# : # ~ ² ÁÃÁ Á³ ÁÃÁ 9¹ c c for some
be the set of all elements of that ha ve last coordinate equal to . Note that : :
is a submodule of and is isomorphic to a submodule of . Hence, the 99 c
inductive hypothesis implies that is finitely generated, say , where:: ~ º º » » =
= is a finite subset of . :
By definition of , each has the form : =
~²ÁÃÁÁ ³ Á
for where there is a of the form 9 :Á
~² ÁÃÁ Á ³ Á Ác Á
Let . We claim that is generated by the finite set .== = ~¸ ÁÃÁ¹ : r
To see this, let . Then and so #~² ÁÃÁ ³: ²ÁÃÁÁ ³:
²ÁÃÁÁ ³~
~
for . Consider now the sum 9
136 Advanced Linear Algebra
$~ º º » »
~
=
The last coordinate of this sum is
~
Á ~
and so the difference has last coordinate and is thus in . # c $ : ~ ºº »» =
Hence
# ~ ²# c $³ b $ ºº »» b ºº »» ~ ºº r »» === =
as desired.
For part 2 , we leave it to the reader to review the proof and make the necessary )
changes. The key fact is that is isomorphic to an ideal of , which is :9
principal. Hence, is generated by a single element of . :4
The Hilbert Basis Theorem
Theorem 5.8 naturally leads us to ask which familiar rings are Noetherian. The
following famous theorem describes one very important case.
Theorem 5.9 Hilbert basis theorem () If a ring is Noetherian, then so is the 9
polynomial ring . 9´%µ
Proof. We wish to show that any ideal in is finitely generated. Let ?9´%µ 3
denote the set of all leading coefficients of polynomials in , together with the ?
element of . Then is an ideal of . 93 9
To see this, observe that if is the leading coefficient of and if ?3 ² % ³
9 ~ ² % ³ , then either or else is the leading coefficient of . In ?
either case, . Similarly, suppose that is the leading coefficient of 3 3
²%³ ²%³ ~ ²%³ ~ ?. We may assume that and , with . Then deg deg
²%³ ~ % ²%³c is in , has leading coefficient and has the same degree as?
²%³ c c. Hence, either is or is the leading coefficient of
²%³c²%³ c 3 ? . In either case .
Since is an ideal of the Noetherian ri ng , it must be finitely generated, say 39
3~º ÁÃÁ » 3 ²%³ . Since , there exist polynomials with leading ?
coefficient . By multiplying each by a suitable power of , we may ² % ³ %
assume that
deg max deg² % ³~~ ¸ ² % ³ ¹
for all .~ ÁÃÁ
Modules II: Free and Noetherian Modules 137
Now for let be the set of all leading coefficients of ~ ÁÃÁc 3
polynomials in of degree , together with the element of . A similar ? 9
argument shows that is an ideal of and so is also finitely generated. 39 3
Hence, we can find polynomials in whose 7 ~ ¸ ²%³ÁÃÁ ²%³¹ Á Á ?
leading coefficients constitute a generating set for . 3
Consider now the finite set
7 ~ 7 r¸ ²%³ÁÃÁ ²%³¹89
~c
If is the ideal generated by , then . An induction argument can be@@ ? 7
used to show that . If ha s degree , then it is a linear @? ?~ ² % ³
combination of the elements of which are constants and is thus in . 7() @
Assume that any polynomial in of de gree less than is in and let ?@ ? ² % ³
have degree .
If , then some linear combination over of the polynomials in ² % ³ 9 7
has the same leading coefficient as and if , then some linear ²%³
combination of the polynomials ²%³
BC% ²%³ÁÃÁ% ²%³ c c
@
has the same leading coefficient as . In either case, there is a polynomial ²%³
²%³ ²%³ ²%³c²%³ @? that has the same leading coefficient as . Since
has degree strictly smaller than that of the induction hypothesis implies that²%³
²%³c²%³ @
and so
²%³ ~ ´²%³c²%³µb²%³ @
This completes the induction and shows that is finitely generated. ?@~
Exercises
1. If is a free -module and is an epimorphism, then must 49 ¢ 4 ¦ 5 5
also be free?
2. Let be an ideal of . Prove that if is a free -module, then is the?? ? 99 ° 9
zero ideal.
3. Prove that the union of an ascendi ng chain of submodules is a submodule.
4. Let be a submodule of an -module . Show that if is finitely :9 4 4
generated, so is the quotient module . 4°:
5. Let be a submodule of an -module. Show that if both and are:9 : 4 ° :
finitely generated, then so is . 4
6. Show that an -module satisfies the ACC for submodules if and only if 94
the following condition holds. Every nonempty collection of submodules I
138 Advanced Linear Algebra
of has a maximal element. That is, for every nonempty collection of4 I
submodules of there is an with the property that 4: I
; ¬;:I .
7. Let be an -homomorphism.¢4¦5 9
a Show that if is finitely generated, then so is . )i m 4² ³
b Show that if and are finitely generated, then )i m ker²³ ²³
4~ ²³b: : 4 ker where is a finitely generated submodule of .
Hence, is finitely generated.4
8. If is Noetherian and is an ideal of show that is also Noetherian.99 9 ° ??
9. Prove that if is Noetherian, then so is . 9 9´% ÁÃÁ% µ
10. Find an example of a commutative ri ng with identity that does not satisfy
the ascending chain condition.
11. a Prove that an -module is cyclic if and only if it is isomorphic to ) 94
9° 9?? where is an ideal of .
b Prove that an -module is and has no proper )( 9 4 4 £ ¸¹ 4 simple
nonzero submodules if and only if it is isomorphic to where is ) 9°??
a maximal ideal of . 9
c Prove that for any nonzero commu tative ring with identity, a simple ) 9
9-module exists.
12. Prove that the condition that be a principal ideal domain in part 2 of9 )
Theorem 5.8 is required.
13. Prove Theorem 5.8 in the following way.
a Show that if are submodules of and if and are ) ;: 4 ; : ° ;
finitely generated, then so is . :
b The proof is again by induction. Assuming it is true for any module )
generated by elements, let and let 4~º º# ÁÃÁ# » » b
4~ º º # Á Ã Á #» » ;~ : q 4ZZ . Then let in part a . )
14. Prove that any -module is isomorphic to the quotient of a free module 94
-4 -. If is finitely generated, then can also be taken to be finitely
generated.
15 Prove that if and are isomorphic submodules of a module it does. :; 4
not necessarily follow that the quotient modules and are 4°: 4°;
isomorphic. Prove also that if as modules it does not :l; :l;
necessarily follow that . Prove that these statements do hold if all ; ;
modules are free and have finite rank.
Chapter 6
Modules over a Principal Ideal Domain
We remind the reader of a few of the basic properties of principal ideal
domains.
Theorem 6.1 Let be a principal ideal domain.9
1 An element is irreducible if and only if the ideal is maximal.) 9 º »
2 An element in is prime if and only if it is irreducible.) 9
3 is a unique factorization domain.)9
4 satisfies the ascending chain condition on ideals. Hence, so does any)9
finitely generated -module . Moreover, if is -generated, then any 94 4
submodule of is -generated. 4
Annihilators and Orders
When is a principal ideal domain, all annihilators are generated by a single 9
element. This permits the following definition.
Definition Let be a principal ideal domain and let be an -module.94 9
1 If is a submodule of , then any generator of is called an )a n n54 ² 5 ³ order
of .5
2 An of an element is an order of the submodule .) order #4 º º # » »
For readers acquainted with group theory, we mention that the order of a
module corresponds to the smallest exponent of a group, to the order of the not
group.
Theorem 6.2 Let be a principal ideal domain and let be an -module.94 9
1 If is an order of , then the orders of are precisely the) 54 5
associates of . We denote any order of by and, as is customary, 5 ² 5 ³
refer to as “the” order of .²5³ 5
2 I f , t h e n)4~(l)
²4³ ~ ²²(³Á²)³³ lcm
140 Advanced Linear Algebra
that is, the orders of are precisely the least common multiples of the 4
orders of and . ()
Proof. We leave proof of part 1) for the reader. For part 2), suppose that
²4³ ~ Á ²(³ ~ Á ²)³ ~ Á ~ ² Á ³ lcm
Then and imply that and and so . On the ( ~ ¸¹ ) ~ ¸¹
other hand, annihilates both and and therefore also . Hence, () 4 ~ ( l )
4 and so is an order of .
Cyclic Modules
The simplest type of nonzero module is cl early a cyclic module. Despite their
simplicity, cyclic modules will play a very important role in our study of linear
operators on a finite-dimensional vector space and so we want to explore some
of their basic properties, including their composition and decomposition.
Theorem 6.3 Let be a principal ideal domain.9
1 If is a cyclic -module with annihilator , then the multiplication)ºº#»» 9 º »
map defined by is an -epimorphism with kernel . ¢9¦º º#» » ~# 9 º »
Hence the induced map
¢¦ º º # » »9
º»
defined by
²bº »³ ~ #
is an isomorphism. In other words, cyclic -modules are isomorphic to 9
quotient modules of the base ring . 9
2 Any submodule of a cyclic -module is cyclic.) 9
3 If is a cyclic submodule of of order , then for ,)ºº#»» 4 9
²ºº #»»³ ~²Á³
gcd
Also,
ºº #»» ~ ºº#»» ¯ ²²#³Á ³ ~ ¯ ² #³ ~ ²#³
Proof. We leave proof of part 1 as an exercise. For part 2 , let . Then )) :º º # » »
0~¸ 9 #: ¹ 9 0~º » 9 is an ideal of and so for some . Thus,
:~0#~9 #~º º # » »
For part 3 , we have if and only if , that is, if and only if ) ² #³ ~ ² ³# ~
, which is equivalent to
Modules Over a Principal Ideal Domain 141
²Á³gcdd
Thus, if and only if and so . For the second ² # ³ º » ² # ³~º »ann ann
statement, if then there exist for which and so ²Á³ ~ Á 9 b ~
# ~ ² b ³# ~ # ºº #»» ºº#»»
and so . Of course, if then . Finally, ifºº #»» ~ ºº#»» ºº #»» ~ ºº#»» ² #³ ~
² #³ ~ , then
~ ² # ³~²Á³gcd
and so .²Á³ ~
The Decomposition of Cyclic Modules
The following theorem shows how cyclic modules can be composed and
decomposed.
Theorem 6.4 Let be an -module.49
1 If have relatively prime)( )Composing cyclic modules "Á Ã Á " 4
orders, then
²" bÄb" ³ ~ ²" ³Ä²" ³
and
ºº" »» l Ä l ºº" »» ~ ºº" b Ä b " »»
Consequently, if
4~(bÄb(
where the submodules have relatively prime orders, then the sum is (
direct.
2 If where the 's are)( )Decomposing cyclic modules ²#³ ~ Ä
pairwise relatively prime, then has the form #
#~" bÄb"
where and so²" ³ ~
ºº#»» ~ ºº" b Ä b " »» ~ ºº" »» l Ä l ºº" »»
Proof. For part 1), let , and . Then ~²"³ Ä #" bÄb"
since annihilates , the order of di vides . If is a proper divisor of , ## ² # ³
then for some index , there is a prime for which annihilates . But ° #
° " £ annihilates each for . Thus,
142 Advanced Linear Algebra
~ #~ " ~ "
67
Since and are relatively prime, the order of is equal to²" ³ ° ² ° ³"
²" ³ ~ ²#³ ~, which contradicts the equation above. Hence, .
It is clear that . For the reverse ºº" b Ä b " »» ºº" »» l Ä l ºº" »»
inclusion, since and are relatively prime, there exist for which ° Á 9
b ~
Hence
"~ b "~ "~ ² "b Ä b "³ º º "b Ä b "» »
67
Similarly, for all and so we get the reverse inclusion. " º º "b Ä b " » »
Finally, to see that the sum above is direct, note that if
#b Ä b #~
where , then each must be , for otherwise the order of the sum on the# ( #
left would be different from .
For part 2 , the scalars are relatively prime and so there exist ) ~° 9
for which
b Ä b ~
Hence,
#~² bÄb ³#~ #bÄb #
Since and since and are relatively prime,² #³ ~ ° ² Á ³ ~ gcd
we have . The second statement follows from part 1 . ² #³ ~ )
Free Modules over a Principal Ideal Domain
We have seen that a submodule of a free module need not be free: The
submodule of the module over itself is not free. However, if {{ {d¸¹ d 9
is a principal ideal domain this cannot happen.
Theorem 6.5 Let be a free module over a principal ideal domain . Then49
any submodule of is also free and . : 4 ²:³ ²4³ rk rk
Proof. We will give the proof first for modules of finite rank and then
generalize to modules of arbitrary rank. Since where is 49 ~ ² 4 ³rk
finite, we may in fact assume that . For each , let 4~9
Modules Over a Principal Ideal Domain 143
0 ~ ¸ 9 ² ÁÃÁ ÁÁÁÃÁ³ : ÁÃÁ 9¹ c c for some
Then it is easy to see that is an ideal of and so for some . 09 0 ~ º » 9
Let
" ~² ÁÃÁ Á ÁÁÃÁ³: c
We claim that
8~¸ " ~ ÁÃÁ £ ¹ and
is a basis for . As to linear independence, suppose that :
8~¸ " ÁÃÁ" ¹
and that
" bÄb " ~
Then comparing the th coordinates gives and since , it ~ £
follows that . In a similar way, all coefficients are and so is linearly ~ 8
independent.
To see that spans , we partition the elements according to the largest 8 :% :
coordinate index with nonzero entry and induct on . If , then ²%³ ²%³ ²%³ ~
%~ %: ² % ³ , which is in the span of . Suppose that all with are in 8
the span of and let , that is,8 ²%³ ~
%~² ÁÃÁ ÁÁÃÁ³
where . Then and so and for some . £ 0 £ ~ 9
Hence, and so and therefore .²% c " ³ & ~ % c " ºº »» % ºº »» 88
Thus, is a basis for .8 :
The previous proof can be generalized in a more or less direct way to modules
of arbitrary rank. In this case, we may assume that is the -module 4~² 9³ 9
of functions with finite support from to , where is a cardinal number. We 9
use the fact that is a well-ordered set, that is, is a totally ordered set in which
any nonempty subset has a smallest element. If , the ´ Á µ closed interval
is
´ Á µ~¸ % % ¹
Let . For each , let:4
4 ~ ¸ : ²³ ´Á µ¹ supp
Then the set
0~ ¸ ²³ 4 ¹
144 Advanced Linear Algebra
is an ideal of and so for some . We show that 90 ~ º ² ³ » :
8 ~¸ Á² ³£ ¹
is a basis for . First, suppose that :
bÄb ~
where for . Applying this to gives
² ³~
and since is an integral domain, . Similarly, for all and so is 9 ~ ~ 8
linearly independent.
To show that spans , since any has finite support, there is a largest 8 : :
index for which . Now, if , then since is well- 8 ~ ² ³ ² ³£ º º » »:
ordered, we may choose a for which is as small as : ± ºº »» ~ ~ ²³8
possible. Then . Moreover, since , it follows that 4 £ ² ³0
²³£ ²³~ ²³ 9 and for some . Then
supp²c ³ ´Á µ
and
² c ³ ²³ ~ ²³ c ²³ ~
and so , which implies that . But then²c ³ c º º » » 8
~²c ³b º º » » 8
a contradiction. Thus, is a basis for . 8 :
In a vector space of dimension , any set of linearly independent vectors is a
basis. This fails for modules. For example, is a -module of rank but the {{
independent set is not a basis. On th e other hand, the fact that a spanning set ¸¹
of size is a basis does hold for modules over a principal ideal domain, as we
now show.
Theorem 6.6 Let be a free -module of finite rank , where is a principal49 9
ideal domain. Let be a spanning set for . Then is a basis :~¸ ÁÃÁ ¹ 4 :
for .4
Proof. Let be a basis for and define the map by8~¸ ÁÃÁ¹ 4 ¢4¦4
~ 9 4 and extending to a surjective -homomorphism. Since is free,
Theorem 5.6 implies that
4 ²³ ²³ ~ ²³ 4 ker ker^ ^ im
Since is a submodule of the free module and since is a principal ideal ker²³ 9
domain, we know that is free of rank at most . It follows that ker²³
Modules Over a Principal Ideal Domain 145
rk rk rk²4³ ~ ² ² ³³b ²4³ ker
and so , that is, , which implies that is an - rk² ² ³ ³~ ² ³~¸ ¹ 9ker ker
isomorphism and so is a basis. :
In general, a basis for a submodule of a free module over a principal ideal
domain cannot be extended to a basis for the entire module. For example, the set
¸¹ is a basis for the submodule of the -module , but this set cannot be {{ {
extended to a basis for itself. We state without proof the following result {
along these lines.
Theorem 6.7 Let be a free -module of rank , where is a principal ideal49 9
domain. Let be a submodule of that is free of rank . Then there is a 54
basis for that contains a subset for which8 4: ~ ¸ # Á Ã Á # ¹
¸ # ÁÃÁ # ¹ 5 ÁÃÁ 9 is a basis for , for some nonzero elements of .
Torsion-Free and Free Modules
Let us explore the relationship between the concepts of torsion-free and free. It
is not hard to see that any free module over an integral domain is torsion-free.
The converse does not hold, unless we st rengthen the hypothe ses by requiring
that the module be finitely generated.
Theorem 6.8 A finitely generated module over a principal ideal domain is free
if and only if it is torsion-free.
Proof. We leave proof that a free module over an integral domain is torsion-free
to the reader. Let be a generating set for . Consider first the .~¸# ÁÃÁ# ¹ 4
case , whence . Then is a basis for since singleton sets are~ .~¸ # ¹ . 4
linearly independent in a torsion-free module. Hence, is free. 4
Now suppose that is a generating set with . If is linearly .~¸"Á#¹ "Á#£ .
independent, we are done. If not, then there exist nonzero for which Á 9
"~ # 4~ º º " Á# » »º º " » » 4 . It follows that and so is a submodule of a
free module and is therefore free by Theorem 6.5. But the map ¢4¦ 4
defined by is an isomorphism because is torsion-free. Thus is #~ # 4 4
also free.
Now we can do the general case. Write
.~¸" ÁÃÁ" Á# ÁÃÁ# ¹ c
where is a maximal linearly independent subset of . Note:~¸" ÁÃÁ" ¹ . (
that is nonempty because singleton sets are linearly independent.: )
For each , the set is linearly dependent and so there exist #¸ " Á Ã Á " Á # ¹
9 ÁÃÁ 9 and for which
146 Advanced Linear Algebra
#b" bÄb" ~
If , then~Ä c
4~º º" ÁÃÁ" Á# ÁÃÁ# » »º º" ÁÃÁ" » » c
and since the latter is a free module, so is , and therefore so is . 4 4
The Primary Cyclic Decomposition Theorem
The first step in the decomposition of a finitely generated module over a 4
principal ideal domain is an easy one. 9
Theorem 6.9 Any finitely generated module over a principal ideal domain 49
is the direct sum of a finitely generated free -module and a finitely generated 9
torsion -module9
4~4 l4 free tor
The torsion part is unique, since it must be the set of all torsion elements of 4tor
44, whereas the free part is unique only up to isomorphism, that is, the free
rank of the free part is unique.
Proof. It is easy to see that the set of all torsion elements is a submodule of 4tor
44 ° 4 4 and the quotient is torsion-fr ee. Moreover, since is finitely tor
generated, so is . Hence, Theorem 6.8 implies that is free. 4°4 4°4 tor tor
Hence, Theorem 5.6 implies that
4~4 l- tor
where is free.-4 ° 4 tor
As to the uniqueness of the torsion part, suppose that where is 4~;l. ;
torsion and is free. Then . But if for and .; 4 # ~ ! b 4 ! ; tor tor
. ~#c!4 ~ #; ;~4 , then and so and . Thus, . tor tor
For the free part, since , the submodules and 4 ~ 4l - ~ 4l . - . tor tor
are both complements of and hence are isomorphic. 4tor
Note that if is a basis for we can write ¸$ ÁÃÁ$ ¹ 4 free
4 ~ ºº$ »» l Ä l ºº$ »» l 4 tor
where each cyclic submodule has zero annihilator. This is a partial ºº$ »»
decomposition of into a direct sum of cyclic submodules. 4
The Primary Decomposition
In view of Theorem 6.9, we turn our attention to the decomposition of finitely
generated torsion modules over a princip al ideal domain. The first step is to 4
decompose into a direct sum of submodules, defined as follows. 4 primary
Modules Over a Principal Ideal Domain 147
Definition Let be a prime in . A - or just module is a9 primary primary ()
module whose order is a power of .
Theorem 6.10 The primary decomposition theorem() Let be a torsion4
module over a principal ideal domain , with order 9
~ Ä
where the 's are distinct nonassociate primes in . 9
1 is the direct sum)4
4~4 lÄl4
where
4 ~ 4 ~ ¸# 4 # ~ ¹
is a primary submodule of order . This decomposition of into primary 4
submodules is called the of . primary decomposition 4
2 The primary decomposition of is unique up to order of the summands.) 4
That is, if
4~5 lÄl5
where is primary of order and are distinct nonassociate5 Á Ã Á
primes, then and, after a possible reindexing, . Hence, ~ 5 ~4
~ ~ Á Ã Á and , for .
3 Two -modules and are isomorphic if and only if the summands in)94 5
their primary decompositions are pairwise isomorphic, that is, if
4~4 lÄl4
and
5~5 lÄl5
are primary decompositions, then and, after a possible reindexing, ~
4 5 ~ÁÃÁ for .
Proof. Let us write and show first that ~°
4~ 4 ~ ¸# # 4 ¹
Since , we have . On the other hand, since ² 4³ ~ 4 ~ ¸¹ 4 4
and are relatively prime, there exist for which Á 9
b ~
and so if then %4
148 Advanced Linear Algebra
%~² b ³ %~ % 4
Hence .4~ 4
For part 1 , since , there exist scalars for which )g c d ²Á à Á³ ~
b Ä b ~
and so for any , %4
%~² bÄb ³% 4
~
Moreover, since the and the 's are pairwise relatively prime, it ² 4³
follows that the sum of the submodules is direct, that is, 4
4~ 4lÄl 4~4 lÄl4
As to the annihilators, it is clear that . For the reverse º » ² 4³
ann
inclusion, if , then and so , that is, ² 4³ ² 4³ ann ann
º » ²4 ³ ~ º »
and so . Thus . ann
As to uniqueness, we claim that is an order of . It is clear that ~ Ä 4
annihilates and so . On the other hand, contains an element of 4 5 "
order and so the sum has order , which implies that #~" bÄb"
. Hence, and are associates.
Unique factorization in now imp lies that and, after a suitable 9 ~
reindexing, that and and are associates. Hence, is primary of ~ 5
order . For convenience, we can write as . Hence,5 5
5 ¸# 4 # ~ ¹ ~ 4
But if
5 lÄl5 ~4 lÄl4
and for all , we must have for all .5 4 5~ 4
For part 3), if and , then the map defined by ~ ¢4 5 ¢4¦5
² bÄb ³~ ²³bÄb ² ³
is an isomorphism and so . Conversely, suppose that . Then 45 ¢45
45 and have the same annihilators and therefore the same order
~ Ä
Hence, part 1) and part 2) imply th at and after a suitable reindexing,~
Modules Over a Principal Ideal Domain 149
~ . Moreover, since
4 ¯ ~¯ ² ³~¯ ~¯ 5
it follows that . ¢4 5
The Cyclic Decomposition of a Primary Module
The next step in the decomposition proce ss is to show that a primary module
can be decomposed into a direct sum of cyclic submodules. While this
decomposition is not unique see the exerci ses , the set of annihilators is unique, ()
as we will see. To establish this uniqueness, we use the following result.
Lemma 6.11 Let be a module over a principal ideal domain and let49
9 be a prime.
1 If , then is a vector space over the field with scalar)4 ~ ¸¹ 4 9°º»
multiplication defined by
²bº»³# ~ #
for all .#4
2 For any submodule of the set) :4
: ~ ¸# : # ~ ¹²³
is also a submodule of and if , then 44 ~ : l ;
4~ :l ;²³ ²³ ²³
Proof. For part 1 , since is prime, the ideal is maximal and so is a )º »9 ° º »
field. We leave the proof that is a vector space over to the reader. For 49 ° º »
part 2 , it is straightforward to s how that is a submodule of . Since ) :4²³
: : ; ; : q; ~ ¸¹ # 4²³ ²³ ²³ ²³ ²³ and we see that . Also, if , then
#~ #~ b! : !; ~ #~ b ! . But for some and and so .
Since and we deduce that , whence . : ! ; ~ ! ~ # : l;²³ ²³
Thus, . But the reverse inequality is manifest.4 :l ;²³ ²³ ²³
Theorem 6.12 The cyclic decomposition theorem of a primary module () Let
4 be a primary finitely generated torsi on module over a principal ideal domain
9, with order .
1 is a direct sum)4
4 ~ ºº# »» l Ä l ºº# »» ()6.1
of cyclic submodules with annihilators , which can be ann²ºº# »»³ ~ º »
arranged in ascending order
ann ann²²ºº# »»³ Ä ºº# »»³
150 Advanced Linear Algebra
or equivalently,
~ Ä
2 As to uniqueness, suppose that is also the direct sum) 4
4 ~ ºº" »» l Ä l ºº" »»
of cyclic submodules with annihilators , arranged in ann²ºº" »»³ ~ º »
ascending order
ann ann²²ºº" »»³ Ä ºº" »»³
or equivalently
Ä
Then the two chains of annihila tors are identical, that is, and ~
ann ann²²ºº" »»³ ~ ºº# »»³
for all . Thus, and for all . ~
3 Two -primary -modules)9
4 ~ ºº# »» l Ä l ºº# »»
and
5 ~ ºº" »» l Ä l ºº" »»
are isomorphic if and only if they have the same annihilator chains, that is,
if and only if and, after a possible reindexing, ~
ann ann²²ºº" »»³ ~ ºº# »»³
Proof . Let have order equal to the order of , that is,# 4 4
ann ann²# ³ ~ ²4³ ~ º »
Such an element must exist since for all and if this inequality ²# ³ # 4
is strict, then will annihilate . 4c
If we show that is complemented, that is, for some ºº# »» 4 ~ ºº# »» l :
submodule , then : since is also a finitely ge nerated primary torsion module :
over , we can repeat the process to get9
4 ~ #l #l :ºº »» ºº »»
where . We can continue this decomposition: ann²# ³ ~ º »
4 ~ #l #l Ä l #l :ºº »» ºº »» ºº »»
as long as . But the ascending sequence of submodules : £ ¸¹
Modules Over a Principal Ideal Domain 151
ºº »» ºº »» ºº »»# #l # Ä
must terminate since is Noetherian and so there is an integer for which 4
eventually , giving 6.1 . : ~ ¸¹ ()
Let . The direct sum clearly exists. Suppose that the# ~ # 4 ~ ºº#»»l¸¹
direct sum
4~ º º # » » l :
exists. We claim that if , then it is possible to find a submodule 4 4 : b
for which and for which the direct sum also : : 4 ~ º º # » » l : b b b
exists. This process must also stop af ter a finite number of steps, giving
4~º º # » »l: as desired.
If and let4 4 " 4 ± 4
:~ º º : Á " c # » »b
for . Then since . We wish to show that for some9 : : "¤4 b
9, the direct sum
ºº#»» l : b
exists, that is,
% ºº#»» q ºº: Á " c #»» ¬ % ~
Now, there exist scalars and for which
%~ #~ b ² "c # ³
for and so if we find a scalar for which :
²"c #³ : (6.2)
then implies that and th e proof of existence will be ºº#»»q: ~ ¸¹ % ~
complete.
Solving for gives "
" ~ ²b ³#c ºº#»»l: ~ 4
so let us consider the ideal of all such scalars:
?~¸ 9 "4¹
Since and is principal, we have??
?~º »
for some . Also, since implies that . "¤4 ¤ ?
152 Advanced Linear Algebra
Since , we have and there exist and for which ~ 9 !:?
"~ # b !
Hence,
" ~ " ~² # b ! ³ ~ # b!
Now we need more information about . Multiplying the expression for by "
c gives
~"~ ² " ³~ #b ! c c c
and since , it follows that . Hence, , that ºº#»»q: ~ ¸¹ # ~ c c
is, and so for some . Now we can write ~ 9
" ~ #b !
and so
²"c #³ ~ ! :
Thus, we take to get (6.2) and th at completes the proof of existence. ~
For uniqueness, note first that has orders and and so and are 4
associates and . Next we show that . According to part 2 of ~ ~ )
Lemma 6.10,
4 ~ ºº# »» l Ä l ºº# »»²³ ²³ ²³
and
4 ~ ºº" »» l Ä l ºº" »»²³ ²³ ²³
where all summands are nonzero. Since , it follows from Lemma 4 ~ ¸¹²³
6.10 that is a vector space over and so each of the preceding 49 ° º »²³
decompositions expresses as a direct sum of one-dimensional vector 4²³
subspaces. Hence, . ~ ² 4 ³~ dim²³
Finally, we show that the exponents and are equal using induction on . If
~ ~ ~ ~ , then for all and since , we also have for all .
Suppose the result is true whenever and let . Write c ~
² ÁÃÁ ³~² ÁÃÁÁÁÃÁ³Á
and
² ÁÃÁ ³~² ÁÃÁÁÁÃÁ³Á ! !
Modules Over a Principal Ideal Domain 153
Then
4 ~ ºº# »»lÄlºº# »»
and
4 ~ ºº" »»lÄlºº" »» !
But is a cyclic submodule of with annihilator and soºº# »» ~ ºº# »» 4 º »c
by the induction hypothesis
~! ~ ÁÃÁ ~ and
which concludes the proof of uniqueness.
For part 3), suppose that and has annihilator chain ¢45 4
ann ann²²ºº# »»³ Ä ºº# »»³
and has annihilator chain5
ann ann²²ºº" »»³ Ä ºº" »»³
Then
5 ~ 4 ~ ºº # »» l Ä l ºº # »»
and so and after a suitable reindexing,~
ann ann ann²ºº# »»³ ~ ²ºº # »»³ ~ ²ºº" »»³
Conversely, suppose that
4 ~ ºº# »» l Ä l ºº# »»
and
5 ~ ºº" »» l Ä l ºº" »»
have the same annihilator chains, that is, and ~
ann ann²²ºº" »»³ ~ ºº# »»³
Then
ºº" »» ~ ºº# »»99
²ºº" »»³ ²ºº# »»³
ann ann
The Primary Cyclic Decomposition
Now we can combine the various decompositions.
Theorem 6.13 The primary cyclic decomposition theorem () L e t b e a4
finitely generated torsion module over a principal ideal domain . 9
154 Advanced Linear Algebra
1 If has order)4
~ Ä
where the 's are distinct nonassociate primes in , then can be 9 4
uniquely decomposed up to the order of the summands into the direct sum ()
4~4 lÄl4
where
4 ~ 4 ~ ¸# 4 # ~ ¹
is a primary submodule with annihilator . Finally, each primary º »
submodule can be written as a direct sum of cyclic submodules, so that 4
4~´º º# » »lÄlº º# » »µlÄl´º º# » »lÄlº º# » »µ Á Á Á Á
44
where and the terms in each cyclic decomposition can ann²ºº# »»³ ~ º »Á Á
be arranged so that, for each ,
ann ann²²ºº# »»³ Ä ºº# »»³Á Á
or, equivalently,
~ Ä Á Á Á
2 As for uniqueness, suppose that)
4~´º º" » »lÄlº º" » »µlÄl´º º" » »lÄlº º" » »µ Á Á Á Á
55
is also a primary cyclic decomposition of . Then, 4
a The number of summands is the same in both decompositions; in fact, )
~ ~ " and after possible reindexing, for all . ""
b The primary submodules are the same; that is, after possible )
reindexing, and 5 ~ 4
c For each primary submodule pair , the cyclic submodules ) 5~ 4
have the same annihilator chains; t hat is, after possible reindexing,
ann ann²ºº" »»³ ~ ²ºº# »»³Á Á
for all .Á
In summary, the primary submodules and annihilator chains are uniquely
determined by the module . 4
3 Two -modules and are isomorphic if and only if they have the same)94 5
annihilator chains.
Modules Over a Principal Ideal Domain 155
Elementary Divisors
Since the chain of annihilators
ann²ºº# »»³ ~ º »Á Á
is unique except for order, the multiset of generators is uniquely ¸ ¹Á
determined up to associate. The generators are called the Áelementary
divisors of . Note that for each prime , the elementary divisor of largest4 Á
exponent is precisely the factor of associated to . ²4³
Let us write to denote the multiset of elementary divisors of ElemDiv ²4³ all
4 ²4³ ²4³. Thus, if , then any associate of is also in . ElemDiv ElemDiv
We can now say that is a complete invariant for isomorphism. ElemDiv ²4³
Technically, the function is the complete invariant, but this 4ª ² 4 ³ ElemDiv
hair is not worth splitting. Also, we c ould work with a system of distinct
representatives for the associate classes of the elementary divisors, but in
general, there is no way to single out a special representative.
Theorem 6.14 Let be a principal ideal domain. The multiset is9² 4 ³ ElemDiv
a complete invariant for isomorphism of finitely generated torsion -modules, 9
that is,
4 5 ¯ ²4³ ~ ²5³ ElemDiv ElemDiv
We have seen (Theorem 6.2) that if
4~(l)
then
²4³ ~ ²²(³Á²)³³ lcm
Let us now compare the elementary divisors of to those of and . 4( )
Theorem 6.15 Let be a finitely generated torsion module over a principal4
ideal domain and suppose that
4~(l)
1 The primary cyclic decomposition of is the direct sum of the primary) 4
cyclic decompositons of and ; that is, if ()
( ~ ºº »» ) ~ ºº »» Á Á and
are the primary cyclic decompositions of and , respectively, then ()
M~ ºº »» l ºº »»45 45Á Á
is the primary cyclic decomposition of M.
156 Advanced Linear Algebra
2 The elementary divisors of are ) 4
ElemDiv ElemDiv ElemDiv ²4³ ~ ²(³r ²)³
where the union is a multiset union; that is, we keep all duplicate
members.
The Invariant Factor Decomposition
According to Theorem 6.4, if and are cyclic submodules with relatively :;
prime orders, then is a cyclic submodule whose order is the product of :l;
the orders of and . Accordingly, in the primary cyclic decomposition of , :; 4
4~´º º# » »lÄlº º# » »µlÄl´º º# » »lÄlº º# » »µ Á Á Á Á
44
with elementary divisors satisfying Á
~ Ä Á Á Á ()6.3
we can combine cyclic summands with relatively prime orders. One judicious
way to do this is to take the left most highest-order cyclic submodules from ()
each group to get
+ ~ ºº# »» l Ä l ºº# »» Á Á
and repeat the process
+ ~ ºº# »» l Ä l ºº# »»
+ ~ ºº# »» l Ä l ºº# »»
Å Á Á
Á Á
Of course, some summands may be missing here since different primary
modules do not necessarily have the same number of summands. In any 4
case, the result of this regrouping and combining is a decomposition of the form
4~+lÄl+
which is called an of . invariant factor decomposition 4
For example, suppose that
4 ~ ´ºº# »» l ºº# »»µ l ´ºº# »»µ l ´ºº# »» l ºº# »» l ºº# »»µ Á Á Á Á Á Á
Then the resulting regrouping and combining gives
4 ~ ´ ºº# »» l ºº# »» l ºº# »» µ l ´ ºº# »» l ºº# »» µ l ´ ºº# »» µ Á Á Á Á Á Á
++ +
As to the orders of the summands, referring to 6.3 , if has order , then () +
since the highest powers of each prime are taken for , the second–highest
for and so on, we conclude that
Modules Over a Principal Ideal Domain 157
Ä c ()6.4
or equivalently,
ann ann²+ ³ ²+ ³ Ä
The numbers are called of the decomposition. invariant factors
For instance, in the example above suppos e that the elementary divisors are
Á Á Á Á Á
Then the invariant factors are
~
~
~
The process described above that passes from a sequence of elementary Á
divisors in order 6.3 to a sequence of invariant factors in order 6.4 is () ()
reversible. The inverse process takes a sequence satisfying 6.4 , Á Ã Á ()
factors each into a product of distinct nonassociate prime powers with the
primes in the same order and then “peels off” like prime powers from the left.
()The reader may wish to try it on the example above.
This fact, together with Theore m 6.4, implies that primary cyclic
decompositions and invariant factor deco mpositions are essentially equivalent.
Therefore, since the multiset of elementary divisors of is unique up to 4
associate, the multiset of invariant facto rs of is also unique up to associate.4
Furthermore, the multiset of invarian t factors is a complete invariant for
isomorphism.
Theorem 6.16 The invariant factor decomposition theorem () Let be a4
finitely generated torsion module over a principal ideal domain . Then 9
4~+lÄl+
where D is a cyclic submodule of , with order , where 4
Ä c
This decomposition is called an of and the invariant factor decomposition 4
scalars are called the of .4 invariant factors
1 The multiset of invariant factors is uniquely determined up to associate by )
the module . 4
2 The multiset of invariant factors is a complete invariant for isomorphism. )
The annihilators of an invariant factor decomposition are called the invariant
ideals of . The chain of invariant ideals is unique, as is the chain of4
158 Advanced Linear Algebra
annihilators in the primary cyclic decompos ition. Note that is an order of , 4
that is,
ann²4³ ~ º »
Note also that the product
~Ä
of the invariant factors of has some nice properties. For example, is the 4
product of all the elementary divisors of . We will see in a later chapter that 4
in the context of a linear operator on a vector space, is the characteristic
polynomial of .
Characterizing Cyclic Modules
The primary cyclic decomposition can be used to characterize cyclic modules
via their elementary divisors.
Theorem 6.17 Let be a finitely generated torsion module over a principal4
ideal domain, with order
~ Ä
The following are equivalent:
1 is cyclic.)4
2 is the direct sum)4
4 ~ ºº# »» l Ä l ºº# »»
of primary cyclic submodules of order . ºº# »»
3 The elementary divisors of are precisely the prime power factors of :) 4
ElemDiv ²4³~¸ ÁÃÁ ¹
Proof. Suppose that is cyclic. Then the primary decomposition of is a 44
primary decomposition, since any subm odule of a cyclic module is cyclic. cyclic
Hence, 1) implies 2). Conversely, if 2) holds, then since the orders are relatively
prime, Theorem 6.4 implies that is cy clic. We leave the rest of the proof to4
the reader.
Indecomposable Modules
The primary cyclic decomposition of is a decomposition of into a direct 44
sum of submodules that cannot be further decomposed. In fact, this
characterizes the primary cyclic decomposition of . Before justifying these 4
statements, we make the following definition.
Definition A module is if it cannot be written as a direct 4 indecomposable
sum of proper submodules.
Modules Over a Principal Ideal Domain 159
We leave proof of the following as an exercise.
Theorem 6.18 Let be a finitely generated torsion module over a principal4
ideal domain. The following are equivalent:
1 is indecomposable)4
2 is primary cyclic)4
3 has only one elementary divisor:)4
ElemDiv ²4³ ~ ¸ ¹
Thus, the primary cyclic decomposition of is a decomposition of into a 44
direct sum of indecomposable modules. Conversely, if
4~(lÄl(
is a decomposition of into a direct sum of indecomposable submodules, then 4
each submodule is primary cyclic and so this is the primary cyclic (
decomposition of . 4
Indecomposable Submodules of Prime Order
Readers acquainted with group theory know that any group of prime order is
cyclic. However, as mentioned earlier, the order of a module corresponds to the
smallest exponent of a group, not to the order of a group. Indeed, there are
modules of prime order that are not cyc lic. Nevertheless, cyclic modules of
prime order are important.
Indeed, if is a finitely generated torsion module over a principal ideal 4
domain, with order , then each prime factor of gives rise to a cyclic
submodule of whose order is and so is also indecomposable. >4 >
Unfortunately, need not be complemented and so we cannot use it to >
decompose . Nevertheless, the theorem is still useful, as we will see in a later 4
chapter.
Theorem 6.19 Let be a finitely generated torsion module over a principal4
ideal domain, with order . If is a prime divisor of , then has a cyclic 4
()equivalently, indecomposable submodule of prime order . >
Proof. If , then there is a for which but .~ #4 $~ #£ $~
Then is annihilated by and so . But is prime and> ~ ºº$»» ²$³
²$³ £ ²$³ ~ > and so . Since has prime order, Theorem 6.18 implies
that is cyclic if and only if it is indecomposable.>
Exercises
1. Show that any free module over an integral domain is torsion-free.
2. Let be a finitely generated tors ion module over a principal ideal domain. 4
Prove that the following are equivalent:
a) is indecomposable 4
b) has only one elementary divisor (including multiplicity) 4
160 Advanced Linear Algebra
c) is cyclic of prime power order. 4
3. Let be a principal ideal domain a nd the field of quotients. Then is 99 9bb
an -module. Prove that any nonzero finitely generated submodule of 99b
is a free module of rank .
4. Let be a principal ideal domain. Let be a finitely generated torsion- 94
free -module. Suppose that is a submodule of for which is a free95 4 5
9 4 ° 5 4-module of rank and is a torsion module. Prove that is a free
9-module of rank .
5. Show that the primary cyclic d ecomposition of a torsion module over a
principal ideal domain is not unique even though the elementary divisors(
are .)
6. Show that if is a finitely generated -module where is a principal 49 9
ideal domain, then the free summand in the decomposition 4~-l4 tor
need not be unique.
7. If is a cyclic -module of order show that the map ºº#»» 9 ¢ 9 ¦ ºº#»»
defined by is a surjective -homomorphism with kernel and so~ # 9 º »
ºº#»» 9
º»
8. If is an integral domain with the property that all submodules of cyclic9
99-modules are cyclic, show that is a principal ideal domain.
9. Suppose that is a finite field and let be the set of all nonzero elements --i
of .-
a Show that if is a nonconstant polynomial over and if ) ²%³ -´%µ -
- ² % ³ %c ² % ³ is a root of , then is a factor of .
b Prove that a nonconstant polynomial of degree can have ) ²%³ -´%µ
at most distinct roots in .-
c Use the invariant factor or primary cyclic decomposition of a finite - ) {
module to prove that is cyclic. -i
10. Let be a principal ideal dom ain. Let be a cyclic -module 94 ~ º º # » » 9
with order . We have seen that any submodule of is cyclic. Prove that 4
for each such that there is a unique submodule of of order 9 4
.
11. Suppose that is a free module of finite rank over a principal ideal 4
domain . Let be a submodule of . If is torsion, prove that95 4 4 ° 5
rk rk²5³ ~ ²4³ .
12. Let be the ring of polynomials ove r a field and let be the ring -´%µ - - ´%µZ
of all polynomials in that have coefficient of equal to . Then -´%µ % -´%µ
is an -module. Show that is finitely generated and torsion-free- ´%µ -´%µZ
but not free. Is a principal ideal domain? -´ % µZ
13. Show that the rational numbers form a torsion-free -module that is not r{
free.
More on Complemented Submodules
14. Let be a principal ideal domain and let be a free -module.94 9
Modules Over a Principal Ideal Domain 161
a Prove that a submodule of is complemented if and only if ) 54 4 ° 5
is free.
b If is also finitely generated, prove that is complemented if and )45
only if is torsion-free.4°5
15. Let be a free module of finite rank over a principal ideal domain .49
a Prove that if is a complemented submodule of , then ) 54
rk rk²5³ ~ ²4³ 5 ~ 4 if and only if .
b Show that this need not hold if is not complemented. ) 5
c Prove that is complemented if and only if any basis for can be ) 55
extended to a basis for . 4
16. Let and be free modules of finite rank over a principal ideal domain45
9¢ 4 ¦ 59. Let be an -homomorphism.
a Prove that is complemented. )k e r ²³
b What about ? )i m ²³
c Prove that )
rk rk rk im rk rk² 4³~ ² ² ³ ³b ² ² ³ ³~ ² ² ³ ³b4
²³ker kerker 67
d If is surjective, then is an isomorphism if and only if )
rk rk²4³ ~ ²5³ .
e If is a submodule of and if is free, then )34 4 ° 3
rk rk rk674
3~² 4 ³ c² 3 ³
17. A submodule of a module is said to be if whenever 54 4 pure in
#¤4±5 #¤5 9 , then for all nonzero .
a Show that is pure if and only if and for implies ) 5# 5 # ~ $ 9
$5 .
b Show that is pure if and only if is torsion-free. ) 54 ° 5
c If is a principal ideal domain a nd is finitely generated, prove that )94
54 ° 5 is pure if and only if is free.
d If and are pure submodules of , then so are and . )35 4 3 q 53 r 5
What about ? 3b5
e If is pure in , then show that is pure in for any )54 3 q 53
submodule of . 34
18. Let be a free module of finite ra nk over a principal ideal domain . Let 49
35 4 3 4 and be submodules of with complemented in . Prove that
rk rk rk rk²3b5³b ²3q5³ ~ ²3³b ²5³
Chapter 7
The Structure of a Linear Operator
In this chapter, we study the structure of a linear operator on a finite-
dimensional vector space, using the powerful module decomposition theorems
of the previous chapter. Unless otherwise noted, all vector spaces will be
assumed to be finite-dimensional.
Let be a finite-dimensional vector space. Let us recall two earler theorems=
(Theorem 2.19 and Theorem 2.20).
Theorem 7.1 Let be a vector space of dimension .=
1 Two matrices and are similar written if and only if)( )d ( ) ()
they represent the same linea r operator , but possibly with B² = ³
respect to different ordered bases. In this case, the matrices and ()
represent exactly the same set of linear operators in . B²= ³
2 Then two linear operators and on are similar written if and)( ) =
only if there is a matrix that represents both operators, but with (C
respect to possibly different ordered bases. In this case, and are
represented by exactly the same set of matrices in . C
Theorem 7.1 implies that the matrices that represent a given linear operator are
precisely the matrices that lie in one similar ity class. Hence, in order to uniquely
represent all linear operators on , we would like to find a set consisting of one =
simple representative of each similarity class, that is, a set of simple canonical
forms for similarity.
One of the simplest types of matrix is the diagonal matrix. However, these are
too simple, since some operators cannot be represented by a diagonal matrix. A
less simple type of matrix is the upper triangular matrix. However, these are not
simple enough: Every operator (over an algebraically closed field) can be
represented by an upper triangular matrix but some operators can be represented
by more than one upper triangular matrix.
164 Advanced Linear Algebra
This gives rise to two different direc tions for further study. First, we can search
for a characterization of those linear operators that can be represented by
diagonal matrices. Such operators are called . Second, we can diagonalizable
search for a different type of “simple” matrix that does provide a set of
canonical forms for similarity. We will pursue both of these directions.
The Module Associated with a Linear Operator
If , we will think of not only as a vector space over a field butB² = ³ = -
also as a module over , with scalar multiplication defined by -´%µ
²%³# ~ ² ³²#³
We will write to indicate the dependence on . Thus, and are modules == =
with the same ring of scalars , although with different scalar multiplication -´%µ
if .£
Our plan is to interpret the concepts of the previous chapter for the module . =
First, if , then . This implies that is a torsion dim dim²= ³ ~ ² ²= ³³ ~ = B
module. In fact, the vectors b
ÁÁ Á Ã Á
are linearly dependent in , which implies that for some nonzero B²= ³ ² ³ ~
polynomial . Hence, and so is a nonzero ²%³ -´%µ ²%³ ²= ³ ²= ³ ann ann
principal ideal of . -´%µ
Also, since is finitely generated as a vector space, it is, a fortiori, finitely =
generated as an -module. Thus, is a finitely generated torsion module -´%µ =
over a principal ideal domain and so we may apply the decomposition -´%µ
theorems of the previous chapter. In the first part of this chapter, we embark on
a “translation project” to translate the powerful results of the previous chapter
into the language of the modules . =
Let us first characterize when two modules and are isomorphic. ==
Theorem 7.2 If , then BÁ² = ³
= = ¯
In particular, is a module isomorphism if and only if is a vector ¢= ¦=
space automorphism of satisfying =
~c
Proof. Suppose that is a module isomorphism. Then for , ¢= ¦= #=
²%#³ ~ %² #³
which is equivalent to
The Structure of a Linear Operator 165
²# ³ ~ ²# ³
and since is bijective, this is equivalent to
²³ # ~ # c
that is, . Since a module isomorphism from to is a vector space ~= =c
isomorphism as well, the result follows.
For the converse, suppose that is a vector space automorphism of and =
~~c, that is, . Then
²% #³ ~ ² #³ ~ ² #³ ~ % ² #³
and the -linearity of implies that for any polynomial ,- ² % ³ - ´ % µ
²² ³#³ ~ ² ³ #
Hence, is a module isomorphism from to . ==
Submodules and Invariant Subspaces
There is a simple connection between the submodules of the -module -´%µ =
and the subspaces of the vector space . Recall that a subspace of is - =: =
invariant if .::
Theorem 7.3 A subset is a submodule of if and only if is a - := = :
invariant subspace of . =
Orders and the Minimal Polynomial
We have seen that the annihilator of , =
ann²= ³ ~ ¸²%³ -´%µ ²%³= ~ ¸¹¹
is a nonzero principal ideal of , say -´%µ
ann²= ³ ~ º²%³»
Since the elements of the base ring of are polynomials, for the first time -´%µ =
in our study of modules there is a logi cal choice among all scalars in a given
associate class: Each associate class contains exactly one polynomial. monic
Definition Let . The unique monic order of is called the B² = ³ = minimal
polynomial for and is denoted by or . Thus, ² % ³ ²³ min
ann²= ³ ~ º ²%³»
In treatments of linear algebra that do not emphasize the role of the module , =
the minimal polynomial of a linear operato r is simply defined as the unique
166 Advanced Linear Algebra
monic smallest degree polynomial of for which . This ² % ³ ²³~
definition is equivalent to our definition.
The concept of minimal polynomial is also defined for matrices. The minimal
polynomial of matrix is defined as the minimal polynomial² % ³ ( ² - ³A C
of the multiplication operator . Equivalently, is the unique monic (( ² % ³
polynomial of smallest degree for which . ²%³ -´%µ ²(³ ~
Theorem 7.4
1 If are similar linear operators on , then . Thus, the)= ² % ³ ~ ² % ³
minimal polynomial is an invariant under similarity of operators.
2 If are similar matrices, then . Thus, the minimal)() ² % ³~ ² % ³ ( )
polynomial is an invariant under similarity of matrices.
3 The minimal polynomial of is the same as the minimal) B² = ³
polynomial of any matrix that represents .
Cyclic Submodules and Cyclic Subspaces
Let us now look at the cyclic submodules of : =
º º # » »~-´ % µ #~¸ ² ³ ² # ³ ² % ³-´ % µ ¹
which are -invariant subspaces of . Let be the minimal polynomial of = ² % ³
O ²²%³³ ~ ²%³# ºº#»»ºº#»» and suppose that . If , then writing deg
²%³ ~ ²%³²%³b²%³
where gives deg deg²%³ ²%³
²%³# ~ ´²%³²%³b²%³µ# ~ ²%³#
and so
ºº#»» ~ ¸²%³# ²%³ ¹ deg
Hence, the set
8 ~ ¸#Á%#ÁÃÁ% #¹ ~ ¸#Á #ÁÃÁ #¹c c
spans the . To see that is a basis for , note that any linear vector space ºº#»» ºº#»» 8
combination of the vectors in has the form for and so is 8 ²%³# ²²%³³ deg
equal to if and only if . Thus, is an ordered basis for . ²%³ ~ ºº#»» 8
Definition Let . A -invariant subspace of is - if has aB ² = ³ : = : cyclic
basis of the form
8~ ¸#Á #ÁÃÁ #¹c
for some and . The basis is called a - for . #= = 8 cyclic basis
The Structure of a Linear Operator 167
Thus, a cyclic submodule of with order of degree is a -cyclic ºº#»» = ²%³
subspace of of dimension . The converse is also true, for if =
8~ ¸#Á #ÁÃÁ #¹c
is a basis for a -invariant subspace of , then is a submodule of . := : =
Moreover, the minimal polynomial of has degree , since if O:
c
c #~c#c #cÄc #
then satisfies the polynomialO:
²%³~ b%bÄb % b% c c
but none of smaller degree since is linearly independent. 8
Theorem 7.5 Let be a finite-dimenional vector space and let . The=: =
following are equivalent:
1 is a cyclic submodule of with order of degree ):= ² % ³
2 is a -cyclic subspace of of dimension .):=
We will have more to say about cyclic modules a bit later in the chapter.
Summary
The following table summarizes the connection between the module concepts
and the vector space concepts that we have discussed so far.
-´%µ = - =
²%³# ² ³ ² ³²#³
==- -
Scalar multiplication: Action of :
Submodule of -Invariant subspace of
AnnihilModule Vector Space
ator: Annihilator:
Monic order of : Minimal polynomial of :ann ann
ann²= ³ ~ ¸²%³ ²%³= ~ ¸¹¹ ²= ³ ~ ¸²%³ ² ³²= ³ ~ ¸¹¹
²%³ =
²= ³ ~ º²%³»
²%³ ² ³ ~
==
ºº#»» ~ ¸²%³# ²%³ ²%³¹ º#Á #ÁÃÁ ²#³»Á ~ has smallest deg with
Cyclic submodule of : -cyclic subspace of :
deg deg dceg²²%³³
The Primary Cyclic Decomposition of =
We are now ready to translate the cyclic decomposition theorem into the
language of . =
Definition Let .B² = ³
1 The and of are the ) elementary divisors invariant factors monic
elementary divisors and invariant fact ors, respectively, of the module . =
We denote the multiset of elementary divisors of by and the ElemDiv ²³
multiset of invariant factors of by . InvFact²³
168 Advanced Linear Algebra
2 The and of a matrix are the) elementary divisors invariant factors (
elementary divisors and invariant factors, respectively, of the multiplication
operator :(
ElemDiv ElemDiv InvFact InvFact ²(³ ~ ² ³ ²(³ ~ ² ³ (( and
We emphasize that the elementary divisors and invariant factors of an operator
or matrix are by definition. Thus, we no longer need to worry about monic
uniqueness up to associate.
Theorem 7.6 Let be(The primary cyclic decomposition theorem for =³ =
finite-dimensional and let have minimal polynomial B² = ³
²%³ ~ ²%³Ä ²%³
where the polynomials are distinct monic primes. ² % ³
1 The -module is the direct sum)( )Primary decomposition -´%µ =
=~ =l Ä l =
where
= ~ = ~ ¸# = ² ³²#³ ~ ¹² % ³
² % ³
is a primary submodule of of order . In vector space terms, is a = ² % ³ =
-invariant subspace of and th e minimal polynomial of is =O =
min²O ³ ~ ² % ³=
2 Each primary summand can be decomposed)( )Cyclic decomposition =
into a direct sum
= ~ ºº# »» l Ä l ºº# »» Á Á
of -cyclic submodules of order with ºº# »» ²%³Á Á
~ Ä Á Á Á
In vector space terms, is a -cyclic subspace of and the minimal ºº# »» =Á
polynomial of is Oºº# »»Á
min²O ³ ~ ² % ³ºº# »»
ÁÁ
3 This yields the decomposition of into a)( )The complete decomposition =
direct sum of -cyclic subspaces
= ~²º º# » »lÄlº º# » »³lÄl²º º# » »lÄlº º# » »³ Á Á Á Á
4 The multiset of elementary divisors)( )Elementary divisors and dimensions
¸ ²%³¹ ² ²%³³ ~
ÁÁ Á is uniquely determined by . If , then the - deg
The Structure of a Linear Operator 169
cyclic subspace has -cyclic basis ºº# »»Á
8 Á Á Á Ác ~#Á#Á Ã Á #23Á
and . Hence,dim deg²ºº# »»³ ~ ² ³Á Á
dim deg²= ³ ~ ² ³
~
Á
We will call the basis
H8~
ÁÁ
for the for .== elementary divisor basis
Recall that if and if both and are -invariant subspaces of , =~ ( l ) ( ) =
the pair is said to . In module language, the pair reduces²(Á)³ ²(Á)³ reduce
if and are submodules of and() =
=~ (l )
We can now translate Theorem 6.15 into the current context.
Theorem 7.7 Let and letB² = ³
=~ (l )
1 The minimal polynomial of is)
²%³ ~ ² ²%³Á ²%³³ lcm OO( )
2 The primary cyclic decomposition of is the direct sum of the primary) =
cyclic decompositons of and ; that is, if ()
( ~ ºº »» ) ~ ºº »»Á Á and
are the primary cyclic decompositions of and , respectively, then ()
= ~ ºº »» l ºº »»45 45Á Á
is the primary cyclic decomposition of . =
3 The elementary divisors of are )
ElemDiv ElemDiv ElemDiv ²³ ~ ²O³ r ²O³ ( )
where the union is a multiset union; that is, we keep all duplicate
members.
170 Advanced Linear Algebra
The Characteristic Polynomial
To continue our translation project, we need a definition. Recall that in the
characterization of cyclic modules in Theorem 6.17, we made reference to the
product of the elementary divisors, one from each associate class. Now that we
have singled out a special representative from each associate class, we can make
a useful definition.
Definition Let . The of is theB ² = ³ ² % ³ characteristic polynomial
product of all of the elementary divisors of :
² % ³~ ² % ³
ÁÁ
Hence,
deg dim² ²%³³ ~ ²= ³
Similarly, the of a matrix is the product of characteristic polynomial ² % ³ 44
the elementary divisors of . 4
The following theorem describes the re lationship between the minimal and
characteristic polynomials.
Theorem 7.8 Let .B² = ³
1 The minimal polynomial of divides the)( )The Cayley–Hamilton theorem
characteristic polynomial of :
² % ³² % ³
Equivalently, satisfies its own c haracteristic polynomial, that is,
²³~
2 The minimal polynomial)
²%³ ~ ²%³Ä ²%³
Á Á
and characteristic polynomial
² % ³~ ² % ³
ÁÁ
of have the same set of prime factors and hence the same set of ² % ³
roots not counting multiplicity . ()
We have seen that the multiset of elementary divisors forms a complete
invariant for similarity. The reader should construct an example to show that the
pair is a complete invariant fo r similarity, that is, this pair of ² ²%³Á ²%³³ not
The Structure of a Linear Operator 171
polynomials does not uniquely determine the multiset of elementary divisors of
the operator .
In general, the minimal polynomial of a linear operator is hard to find. One of
the virtues of the characteristic polynom ial is that it is comparatively easy to
find and we will discuss this in detail a bit later in the chapter.
Note that since and both poly nomials are monic, it follows that ² % ³² % ³
²%³ ~ ²%³ ¯ ² ²%³³ ~ ² ²%³³ deg deg
Definition A linear operator is if its minimal B² = ³ nonderogatory
polynomial is equal to its characteristic polynomial:
² % ³~² % ³
or equivalently, if
deg deg² ²%³³ ~ ² ²%³³
or if
deg dim² ²%³³ ~ ²= ³
Similar statements hold for matrices.
Cyclic and Indecomposable Modules
We have seen (Theorem 6.17) that cyclic submodules can be characterized by
their elementary divisors. Let us tran slate this theorem into the language of =
(and add one more equivalence related to the characteristic polynomial).
Theorem 7.9 Let have minimal polynomialB² = ³
²%³ ~ ²%³Ä ²%³
where are distinct monic primes. The following are equivalent:² % ³
1 is cyclic.)=
2 is the direct sum)=
= ~ ºº# »» l Ä l ºº# »»
of -cyclic submodules of order . ºº# »» ²%³
3 The elementary divisors of are)
ElemDiv ² ³~¸ ²%³ÁÃÁ ²%³¹
4 is nonderogatory, that is,)
² % ³~² % ³
172 Advanced Linear Algebra
Indecomposable Modules
We have also seen (Theorem 6.19) that, in the language of , each prime factor =
²%³ ²%³ > = of the minimal polynomial gives rise to a cyclic submodule of
of prime order . ²%³
Theorem 7.10 Let and let be a prime factor of . Then B ²=³ ²%³ ²%³ =
has a cyclic submodule of prime order . > ² % ³
For a module of prime order, we have the following.
Theorem 7.11 For a module of prime order , the following are > ² % ³
equivalent:
1 is cyclic)>
2 is indecomposable)>
3 is irreducible)² % ³
4 is nonderogatory, that is, ) ² % ³~² % ³
5 .)d i m d e g²> ³ ~ ²²%³³
Our translation project is now complete a nd we can begin to look at issues that
are specific to the modules . =
Companion Matrices
We can also characterize the cyclic modul es via the matrix representations of=
the operator , which is obviously someth ing that we could not do for arbitrary
modules. Let be a cyclic module, with order =~ º º # » »
²%³~ b%bÄb % b% c c
and ordered -cyclic basis
8~# Á# Á Ã Á #23c
Then
²# ³ ~ # b
for andc
²# ³ ~ #
~c² b bÄb ³#
~c #c #cÄc #c
c c
c c
and so
The Structure of a Linear Operator 173
´µ ~Ä c
Ä c
Æ Å
ÅÅÆc
Äc 8vy
x{x{x{x{
wz
c
c2
This matrix is known as the for the polynomial . companion matrix ² % ³
Definition The of a monic polyomial companion matrix
²%³~ b%bÄb % b% c c
is the matrix
*´²%³µ~Ä c
Ä c
Æ Å
ÅÅÆc
Äc vy
x{x{x{x{
wz
c
c2
Note that companion matrices are defined only for polynomials. monic
Companion matrices are nonderogatory. Al so, companion matrices are precisely
the matrices that represent operators on -cyclic subspaces.
Theorem 7.12 Let .²%³ -´%µ
1 A companion matrix is nonderogatory; in fact,) (~*´ ² % ³ µ
²%³ ~ ²%³ ~ ²%³((
2 is cyclic if and only if can be represented by a companion matrix, in)=
which case the representing basis is -cyclic.
Proof. For part 1), let be the standard basis for . Since ;~² ÁÃÁ³ -
~ ( ² % ³c for , it follows that for any polynomial ,
²(³ ~ ¯ ²(³ ~ ¯ ²(³ ~ for all
If , then²%³~ b%bÄb % b% c c
² ( ³ ~ ( b( ~ c ~ b b
~ ~ ~c c c
and so , whence . Also, if²(³ ~ ²(³ ~
²%³~ b%bÄb % b % c c
is nonzero and has degree , then
²(³ ~ b bÄb b £ c b
174 Advanced Linear Algebra
since is linearly independent. Hence, has smallest degree among all; ²%³
polynomials satisfied by and so . Finally, ( ² % ³ ~ ² % ³ (
deg deg deg² ²%³³ ~ ²²%³³ ~ ² ²%³³((
For part 2), we have already proved that if is cyclic with -cyclic basis ,= 8
then . For the converse, if , then part 1) implies´ µ ~ *´²%³µ ´ µ ~ *´²%³µ88
that is nonderogatory. Hence, Theorem 7.11 implies that is cyclic. It is =
clear from the form of that is a -cyclic basis for . *´²%³µ = 8
The Big Picture
If , then Theorem 7.2 and the fact that the elementary divisors form BÁ² = ³
a complete invariant for isomorphism imply that
¯ = = ¯ ² ³ ~ ² ³ ElemDiv ElemDiv
Hence, the multiset of elementary divisors is a complete invariant for similarity
of operators. Of course, the same is true for matrices:
( ) ¯ - - ¯ ²(³ ~ ²)³
( ) ElemDiv ElemDiv
where we write in place of . --
( (
The connection between the elementary divisors of an operator and the
elementary divisors of the matrix represe ntations of is described as follows. If
(~´ µ ¢=-88, then the coordinate map is also a isomorphismmodule
8¢= ¦-
(. Specifically, we have
88 8 8 8²² ³#³ ~ ´² ³#µ ~ ²´ µ ³´#µ ~ ²(³ ²#³
and so preserves -scalar multiplication. Hence,8 -´%µ
(~´ µ ¬ = -88 for some
(
For the converse, suppose that . If we define by , ¢= - = ~
(
where is the th standard basis vector, then is an ordered ~ ² Á Ã Á ³ 8
basis for and is the coordinate map for . Hence, is a module =~ 8 88
isomorphism and so
88²# ³ ~ ² # ³ (
for all , that is,#=
´# µ ~ ² ´ # µ³88(
which shows that . (~´ µ8
Theorem 7.13 Let be a finite-dimensional vector space over . Let=-
B CÁ² = ³ ( Á ) ² - ³ and let .
The Structure of a Linear Operator 175
1 The multiset of elementary divisors or invariant factors is a complete )( )
invariant for similarity of operators, that is,
¯ = =
¯² ³ ~² ³
¯² ³ ~² ³
ElemDiv ElemDiv
InvFact InvFact
A similar statement holds for matrices:
()¯- -
¯² ( ³ ~² ) ³
¯² ( ³ ~² ) ³
( )
ElemDiv ElemDiv
InvFact InvFact
2 The connection between operators and their representing matrices is )
(~´ µ ¯= -
¯² ³ ~² ( ³
¯² ³ ~² ( ³8
8 for some
(
ElemDiv ElemDiv
InvFact InvFact
Theorem 7.13 can be summarized in Figure 7.1, which shows the big picture.
WVsimilarity classes
of L(V)
VWisomorphism classes
of F[x]-modulesVV
{ED1}Multisets of
elementary divisors{ED2}
[W]B[V]B
[W]R[V]RSimilarity classes
of matrices
Figure 7.1
Figure 7.1 shows that the similarity classes of are in one-to-one B²= ³
correspondence with the isomorphism classes of -modules and that these -´%µ =
are in one-to-one correspondence with the multisets of elementary divisors,
which, in turn, are in one-to-one corresponde nce with the similarity classes of
matrices.
We will see shortly that any multiset of prime power polynomials is the multiset
of elementary divisors for some operator (or matrix) and so the third family in
176 Advanced Linear Algebra
the figure could be replaced by the family of all multisets of prime power
polynomials.
The Rational Canonical Form
We are now ready to determine a set of canonical forms for similarity. Let
B H² = ³ = . The elementary divisor basis for that gives the primary cyclic
decomposition of , =
= ~ ²ºº# »»lÄlºº# »»³lÄl²ºº# »»lÄlºº# »»³ Á Á Á Á
is the union of the bases
8 Á Á Á Ác ~ ² #Á#Á à Á #³Á
and so the matrix of with respect to is the block diagonal matrix H
´ µ ~ ²*´ ²%³µÁÃÁ*´ ²%³µÁÃÁ*´ ²%³µÁÃÁ*´ ²%³µ³H diag
Á Á Á Á
with companion matrices on the block di agonal. This matrix has the following
form.
Definition A matrix is in the of ( elementary divisor form rational canonical
form if
(~ *´ ²%³µÁÃÁ*´ ²%³µ diag45
where the are monic prime polynomials. ² % ³
Thus, as shown in Figure 7.1, each similarity class contains at least one matrix I
in the elementary divisor form of rational canonical form.
On the other hand, suppose that is a rational canonical matrix 4
4 ~ ²*´ ²%³µÁÃÁ*´ ²%³µÁÃÁ*´ ²%³µÁÃÁ*´ ²%³µ³ diag
Á Á Á Á
of size . Then represents the matrix multiplication operator underd 4 4
the standard basis on . The basis can be partitioned into blocks ;; ;-Á
corresponding to the position of each of the companion matrices on the block
diagonal of . Since 4
´O µ~ * ´ ² % ³ µ4º»
;;ÁÁÁ
it follows from Theorem 7.12 that each subspace is -cyclic with monic º»;Á 4
order and so Theorem 7.9 imp lies that the multiset of elementary ² % ³Á
divisors of is .4 ¸ ²%³¹Á
This shows two important things. Fi rst, any multiset of prime power
polynomials is the multiset of elementary divisors for some matrix. Second, 4
The Structure of a Linear Operator 177
lies in the similarity class that is associated with the elementary divisors
¸ ²%³¹Á. Hence, two matrices in the elementary divisor form of rational
canonical form lie in the same similarity class if and only if they have the same
multiset of elementary divisors. In other words, the elementary divisor form of
rational canonical form is a set of canoni cal forms for similarity, up to order of
blocks on the block diagonal.
Theorem 7.14 Let()The rational canonical form: elementary divisor version
= ² = ³ be a finite-dimensional vector space and let have minimal B
polynomial
²%³ ~ ²%³Ä ²%³
where the 's are distinct monic prime polynomials. ² % ³
1 If is an elementary divisor basis for , then is in the elementary )H =´ µH
divisor form of rational canonical form:
´ µ ~ *´ ²%³µÁÃÁ*´ ²%³µÁÃÁ*´ ²%³µÁÃÁ*´ ²%³µH diag45
Á Á Á Á
where are the elementary divisors of . This block diagonal matrix² % ³Á
is called an of a of . elementary divisor version rational canonical form
2 Each similarity class of matrices contains a matrix in the elementary ) I 9
divisor form of rational canonical form. Moreover, the set of matrices in I
that have this form is the set of matr ices obtained from by reordering the 4
block diagonal matrices. Any such matrix is called an elementary divisor
verison rational canonical form of a of . (
3 The dimension of is the sum of th e degrees of the elementary divisors of ) =
, that is,
dim deg²= ³ ~ ² ³
~ ~
Á
Example 7.1 Let be a linear operator on the vector space and suppose thats7
has minimal polynomial
² % ³~² %c ³ ² %b ³
Noting that and are elementary divisors and that the sum of the %c ²% b³
degrees of all elementary divisors must equal , we have two possibilities:
1 1 )%cÁ²% b ³Á% b
2 1)%cÁ%cÁ%cÁ²% b ³
These correspond to the following rational canonical forms:
178 Advanced Linear Algebra
1)vy
x{x{x{x{x{x{x{x{
wz
c
c
c
2)vy
x{x{x{x{x{x{x{x{
wz
c
c
The rational canonical form may be far from the ideal of simplicity that we had
in mind for a set of simple canonical forms. Indeed, the rational canonical form
can be important as a theoretical tool, more so than a practical one.
The Invariant Factor Version
There is also an invariant factor version of the rational canonical form. We
begin with the following simple result.
Theorem 7.15 If are relatively prime polynomials, then²%³Á²%³ -´%µ
*´²%³²%³µ *´²%³µ
*´²%³µ67
block
Proof. Speaking in general terms, if an matrix has minimal d (
polynomial
²%³ ~ ²%³Ä ²%³
of degree equal to the size of the ma trix, then Theorem 7.14 implies that the
elementary divisors of are precisely (
²%³ÁÃÁ ²%³
Since the matrices and have the same size *´²%³²%³µ ²*´²%³µÁ*´²%³µ³ diag
d ²%³²%³ and the same minimal polynomial of degree , it follows that
they have the same multiset of elementary divisors and so are similar.
Definition A matrix is in the of ( invariant factor form rational canonical
form if
The Structure of a Linear Operator 179
(~ *´ ²%³µÁÃÁ*´ ²%³µ diag45
where for . ²%³ ²%³ ~ÁÃÁcb
Theorem 7.15 can be used to rearrange and combine the companion matrices in
an elementary divisor version of a rational canonical form to produce an 9
invariant factor version of rational canonical form that is similar to . Also, this 9
process is reversible.
Theorem 7.16 The rational canonical form: invariant factor version () Let
dim²= ³ B ²= ³ and suppose that has minimal polynomial B
²%³ ~ ²%³Ä ²%³
where the monic polynomials are di stinct prime irreducible polynomials ² % ³ ()
1 has an , that is, a basis for which)= invariant factor basis 8
´ µ ~ *´ ²%³µÁÃÁ*´ ²%³µ8diag45
where the polynomials are the invariant factors of and ² % ³
² % ³ ² % ³b . This block diagonal matrix is called an invariant factor
version rational canonical form of a of .
2 Each similarity class of matrices contains a matrix in the invariant ) I 9
factor form of rational canonical form. Moreover, the set of matrices in I
that have this form is the set of matr ices obtained from by reordering the 4
block diagonal matrices. Any such matrix is called an invariant factor
verison rational canonical form of a of . (
3 The dimension of is the sum of the degrees of the invariant factors of , ) =
that is,
dim deg²= ³ ~ ² ³
~
The Determinant Form of the Characteristic Polynomial
In general, the minimal polynomial of an operator is hard to find. One of the
virtues of the characteristic polynomial is that it is comparatively easy to find.
This also provides a nice example of the theoretical value of the rational
canonical form.
Let us first take the case of a companion matrix. If is the (~*´ ² % ³ µ
companion matrix of a monic polynomial
² % Â Á Ã Á ³ ~ b % b Ä b %b % c c c
then how can we recover from by arithmetic operations? ²%³ ~ ²%³ *´²%³µ (
180 Advanced Linear Algebra
When , we can write as~ ² % ³
² % Â Á ³~b%b% ~% ² %b³b
which looks suspiciously like a determinant:
² % Â Á ³~%
c %b
~% 0 cc
c
~ ²%0 c*´ ²%³µ³
det
det
det>?
67 >?
So, let us define
(²%Â ÁÃÁ ³ ~ %0 c*´ ²%³µ
~% Ä
c % Ä
c Æ Å
ÅÅ Æ %
Ä c % b c
c
cvy
x{x{x{x{
wz
where is an independent variable. Th e determinant of this matrix is a %
polynomial in whose degree equals the number of parameters . % ÁÃÁ c
We have just seen that
det²(²%Â Á ³³~ ²%Â Á ³
and this is also true for . As a basis for induction, if ~
det²(²%Â ÁÃÁ ³³~ ²%Â ÁÃÁ ³ c c
then expanding along the first row gives
det
det det
det²(²%Á ÁÃÁ ³³
~% ²(²%Á ÁÃÁ ³³b²c³ c % Ä
c Æ
ÅÅ Æ %
Ä c
~% ²(²%Á ÁÃÁ ³³b
~%² % ÂÁÃÁ³b
~%b% bÄb
d
vy
x{x{
wz
% b% b
~ ²%Â ÁÃÁ ³ b
b
We have proved the following.
The Structure of a Linear Operator 181
Lemma 7.17 For any , ²%³ -´%µ
det²%0 c*´²%³µ³ ~ ²%³
Now suppose that is a matrix in th e elementary divisor form of rational 9
canonical form. Since the determinant of a block diagonal matrix is the product
of the determinants of the blocks on the diagonal, it follows that
det²%0 c9³ ~ ²%³ ~ ²%³
Á
9Á
Moreover, if , say , then (9 (~79 7c
det det
det
det det det
det²%0 c(³ ~ ²%0 c797 ³
~´ 7 ² % 0 c 9 ³ 7 µ
~² 7 ³ ² % 0 c 9 ³ ² 7 ³
~² % 0 c 9 ³c
c
c
and so
det det²%0 c(³ ~ ²%0 c9³ ~ ²%³ ~ ²%³ 9 (
Hence, the fact that all matrices have a rational canonical form allows us to
deduce the following theorem.
Theorem 7.18 Let . If is any matrix that represents , thenB ² = ³ (
²%³ ~ ²%³ ~ ²%0 c(³ ( det
Changing the Base Field
A change in the base field will gener ally change the primeness of polynomials
and therefore has an effect on the multiset of elementary divisors. It is perhaps a
surprising fact that a change of base field has on the invariant factors— no effect
hence the adjective . invariant
Theorem 7.19 Let and be fields with . Suppose that the elementary-2 - 2
divisors of a matrix are ( ² -³C
7~¸ ÁÃÁ ÁÃÁ ÁÃÁ ¹
Á Á Á Á
Suppose also that the polynomials can be further factored over , say 2
~ Ä Á Á Á
Á
where is prime over . Then the prime powers2Á
8~¸ ÁÃÁ ÁÃÁÃÁ ÁÃÁ ¹Á Á Á
ÁÁ Á Á Á Á Á
Á Á
are the elementary divisors of over . (2
182 Advanced Linear Algebra
Proof. Consider the companion matrix in the rational canonical form *´ ²%³µÁ
of over . This is a matrix over as well and Theorem 7.15 implies that(- 2
*´ ²%³µ ²*´ µÁÃÁ*´ µ³ Á Á Á Á Á
Á Ádiag
Hence, is an elementary divisor basis for over .8 (2
As mentioned, unlike the elementary divisors, the invariant factors are field
independent . This is equivalent to saying that the invariant factors of a matrix
(4² -³ - are polynomials over the s ubfield of that contains the smallest
entries of (À
Theorem 7.20 Let and let be the smallest subfield of ( ² -³ ,- -C
that contains the entries of . (
1 The invariant factors of are polynomials over .) (,
2 Two matrices are similar over if and only if they are) (Á) ²-³ -C
similar over . ,
Proof. Part 1 follows immediately from Theorem 7.19, since using either or ) 7
8 to compute invariant factors gives the same result. Part 2) follows from the
fact that two matrices are similar over a given field if and only if they have the
same multiset of over that field. invariant factors
Example 7.2 Over the real field, the matrix
(~c
67
is the companion matrix for the polynomial , and so %b
ElemDiv InvFact ss²(³ ~ ¸% b¹ ~ ²(³
However, as a complex matrix, the rational canonical form for is (
(~
c 67
and so
ElemDiv InvFact dd²(³~¸%cÁ%b¹ ²(³~¸% b¹ and
Exercises
1. We have seen that any can be used to make into an - B² = ³ = - ´ % µ
module. Does every module over come from some ? =- ´ % µ ² = ³ B
Explain.
2. Let have minimal polynomialB² = ³
²%³ ~ ²%³Ä ²%³
The Structure of a Linear Operator 183
where are distinct monic primes. Prove that the following are² % ³
equivalent:
a is -cyclic. )=
b . )d e g d i m² ²%³³ ~ ²= ³
c The elementary divisors of are the prime power factors and so ) ² % ³
= ~ ºº# »» l Ä l ºº# »»
is a direct sum of -cyclic submodules of order . ºº# »» ²%³
3. Prove that a matrix is nonderogato ry if and only if it is similar ( ² -³C
to a companion matrix.
4. Show that if and are block diagonal matrices with the same blocks, but ()
in possibly different order, then and are similar. ()
5. Let . Justify the statement that the entries of any invariant ( ² -³C
factor version of a rational canonical form for are “rational” expressions (
in the coefficients of , hence the origin of the term ( rational canonical
form . Is the same true for the elementary divisor version?
6. Let where is finite-dimensional. If is irreducibleB² = ³ = ² % ³ - ´ % µ
and if is not one-to-one, prove that divides the minimal ² ³ ²%³
polynomial of .
7. Prove that the minimal polynomial of is the least common B² = ³
multiple of its elementary divisors.
8. Let where is finite-dimensional. Describe conditions on theB² = ³ =
minimal polynomial of that are equivalent to the fact that the elementary
divisor version of the rational canoni cal form of is diagonal. What can
you say about the elementary divisors?
9. Verify the statement that the multiset of elementary divisors or invariant (
factors is a complete invariant for similarity of matrices. )
10. Prove that given any multiset of monic prime power polynomials
4~¸ ²%³ÁÃÁ ²%³ÁÃÁÃÁ ²%³ÁÃÁ ²%³¹
Á Á Á Á
and given any vector space of dimension equal to the sum of the degrees =
of these polynomials, there is an operator whose multiset of B² = ³
elementary divisors is . 4
11. Find all rational canonical forms up to the order of the blocks on the ²
diagonal for a linear operator on having minimal polynomial ) s6
²%c ³ ²%b ³ 11 .
12. How many possible rational canoni cal forms up to order of blocks are ()
there for linear operators on with minimal polynomial 1 1 ? s6²%c ³²%b ³
13. a Show that if and are matrices, at least one of which is ) () d
invertible, then and are similar. () )(
184 Advanced Linear Algebra
b What do the matrices )
(~ )~
>? >? and
have to do with this issue?
c Show that even without the assumption on invertibility the matrices )
() )( and have the same characteristic polynomial. : Write Hint
(~70 8 Á
where and are invertible and is an matrix that has the78 0 d Á
d identity in the upper left-hand corner and 's elsewhere. Write
)~ 8 ) 7 ( ) ) (Z. Compute and and find their characteristic
polynomials.
14. Let be a linear operator on with minimal polynomial -
² % ³ ~ ² %b³ ² %c³1 2 . Find the rational canonical form for if
-~ -~ -~rs d, or .
15. Suppose that the minimal polynomial of is irreducible. What can B² = ³
you say about the dimension of ? =
16. Let where is finite-dimensional. Suppose that is anB² = ³ = ² % ³
irreducible factor of the minima l polynomial of . Suppose further ²%³
that have the property that . Prove that"Á# = ²"³ ~ ²#³ ~ ²%³
"~² ³# ²%³ #~² ³" for some polyjomial if and only if for some
polynomial . ²%³
Chapter 8
Eigenvalues and Eigenvectors
Unless otherwise noted, we will assume throughout this chapter that all vector
spaces are finite-dimensional.
Eigenvalues and Eigenvectors
We have seen that for any , the minimal and characteristic B² = ³
polynomials have the same set of root s (but not generally the same of multiset
roots). These roots are of vital importance.
Let be a matrix that represents . A scalar is a root of the(~´ µ - 8
characteristic polynomial if and only if ²%³ ~ ²%³ ~ ²%0 c(³ ( det
det²0 c ( ³ ~ ()8.1
that is, if and only if the matrix is singular. In particular, if ,0c( ² =³~ dim
then 8.1 holds if and only if there exists a nonzero vector for which () %-
²0 c ( ³ % ~
or equivalently,
(%~ %
If , then this is equivalent to´#µ ~ %8
´ µ ´#µ ~ ´#µ88 8
or in operator language,
#~ #
This prompts the following definition.
Definition Let be a vector space over a field and let .=- ² = ³ B
1 A scalar is an or of if there)( ) - eigenvalue characteristic value
exists a vector for which nonzero #=
186 Advanced Linear Algebra
#~ #
In this case, is called an or of # eigenvector characteristic vector ()
associated with .
2 A scalar is an for a matrix if there exists a ) - ( eigenvalue nonzero
column vector for which %
(% ~ %
In this case, is called an or for %( eigenvector characteristic vector ()
associated with .
3 The set of all eigenvectors associated with a given eigenvalue , together)
with the zero vector, forms a subspace of , called the of and = eigenspace
denoted by . This applies to both linear operators and matrices.;
4 The set of all eigenvalues of an operator or matrix is called the ) spectrum
of the operator or matrix. We denote the spectrum of by . Spec²³
Theorem 8.1 Let have minimal polynomial and characteristicB² = ³ ² % ³
polynomial . ² % ³
1 The spectrum of is the set of all roots of or of , not counting) ²%³ ²%³
multiplicity.
2 The eigenvalues of a matrix are invariants under similarity.)
3 The eigenspace of the matrix is the solution space to the homogeneous) ; (
system of equations
² 0 c(³²%³ ~
One way to compute the eigenvalues of a linear operator is to first represent
by a matrix and then solve the ( characteristic equation
det²%0 c(³ ~
Unfortunately, it is quite likely that this equation cannot be solved when
dim²= ³ . As a result, the art of approximati ng the eigenvalues of a matrix is
a very important area of applied linear algebra.
The following theorem describes the relationship between eigenspaces and
eigenvectors of distinct eigenvalues.
Theorem 8.2 Suppose that are distinct eigenvalues of a linear ÁÃÁ
operator .B² = ³
1 Eigenvectors associated with distinct eigenvalues are linearly independent;)
that is, if , then the set is linearly independent. # ¸# ÁÃÁ# ¹ ;
2 The sum is direct; that is, exists.) ;; ;; bÄb lÄl
Proof. For part 1), if is linearly dependent, then by renumbering if ¸# ÁÃÁ# ¹
necessary, we may assume that among all nontrivial linear combinations of
Eigenvalues and Eigenvectors 187
these vectors that equal , the equation
# bÄb# ~ ()8.2
has the fewest number of terms. Applying gives
# bÄb # ~ ()8.3
Multiplying (8.2) by and subtracting from (8.3) gives
²c³ # b Ä b ²c³ # ~
But this equation has fewer terms than (8.2) and so all of its coefficients must
equal . Since the 's are distinct, for and so as well. This ~ ~
contradiction implies that the 's are linearly independent. #
The next theorem describes the spectrum of a polynomial in . ² ³
Theorem 8.3 The Let be a vector space over() spectral mapping theorem =
an algebraically closed field . Let and let . Then - ²= ³ ²%³ -´%µB
Spec Spec Spec² ²³ ³ ~ ² ²³ ³ ~ ¸ ²³ ²³ ¹
Proof. We leave it as an exercise to show that if is an eigenvalue of , then
²³ ²³ ² ²³ ³ ² ²³ ³ is an eigenvalue of . Hence, . For the reverse Spec Spec
inclusion, let , that is, ² ² ³ ³Spec
²² ³c ³# ~
for . If#£
²%³c ~²%c³ IJ%c ³
where , then writing this as a product of (not necessarily distinct) linear -
factors, we have
²c ³ Ä ²c ³ Ä ²c ³ Ä ²c ³ # ~
(The operator is written for convenience.) We can remove factors from
the left end of this equation one by one until we arrive at an operator (perhaps
the identity) for which but . Then is an eigenvector #£ ² c³ #~ #
for with eigenvalue . But since , it follows that ² ³ c ~
~ ² ³ ² ² ³³ ²² ³³ ² ² ³³ Spec Spec Spec . Hence, .
The Trace and the Determinant
Let be algebraically closed and let have characteristic-( ² - ³ C
polynomial
²%³~% b % bÄb%b
~ ²%c ³Ä²%c ³( c
c
188 Advanced Linear Algebra
where are the eigenvalues of . ThenÁÃÁ (
²%³ ~ ²%0 c(³( det
and setting gives %~
det²(³ ~ c ~ ²c³ Ä c
Hence, if is algebraically closed then, , is the constant term -² ( ³ up to sign det
of and the product of the eigenvalues of , including multiplicity.² % ³ ((
The of the eigenvalues of a matrix over an algebraically closed field is alsosum
an interesting quantity. Like the determi nant, this quantity is one of the
coefficients of the characteristic polynomial (up to sign) and can also be
computed directly from the entrie s of the matrix, without knowing the
eigenvalues explicitly.
Definition The of a matrix , denoted by , is the sum of trace ( ² -³ ² ( ³C tr
the elements on the main diagonal of . (
Here are the basic propeties of the trace. Proof is left as an exercise.
Theorem 8.4 Let .(Á) ²-³C
1 A , f o r .)t r t r² ³ ~ ²(³ -
2 .)t r t r t r²(b)³ ~ ²(³b ²)³
3 .)t r t r²()³ ~ ²)(³
4 . However, may not equal) t rt rt r t r²()*³ ~ ²*()³ ~ ²)*(³ ²()*³
tr²(*)³ .
5 The trace is an invariant under similarity.)
6 If is algebraically closed, then is the sum of the eigenvalues of , )t r-² ( ³ (
including multiplicity, and so
tr²(³ ~ c c
where . ²%³~% b % bÄb%b( c c
Since the trace is invariant under similarity, we can make the following
definition.
Definition The of a linear operator is the trace of any matrix trace B² = ³
that represents .
As an aside, the reader who is familar with symmetric polynomials knows that
the coefficients of any polynomial
²%³~% b % bÄb%b
~ ²%c ³Ä²%c ³ c
c
Eigenvalues and Eigenvectors 189
are the of the roots: elementary symmetric functions
~ ² c ³
~ ² c ³
~ ² c ³
Å
~ ² c ³c
c
c
~
The most important elementary symmetric functions of the eigenvalues are the
first and last ones:
~ c bÄb ~ ²(³ ~ ²c³ Ä ~ ²(³c tr and det
Geometric and Algebraic Multiplicities
Eigenvalues actually have two forms of multiplicity, as described in the next
definition.
Definition Let be an eigenvalue of a linear operator . B ² = ³
1 The of is the multiplicity of as a root of the) algebraic multiplicity
characteristic polynomial . ² % ³
2 The of is the dimension of the eigenspace .) geometric multiplicity ;
Theorem 8.5 The geometric multiplicity of an eig envalue of is lessB² = ³
than or equal to its algebraic multiplicity.
Proof. We can extend any basis of to a basis for . 8; 8 ~¸ #ÁÃÁ#¹ =
Since is invariant under , the matrix of with respect to has the block ; 8
form
´µ ~0(
)
867
block
where and are matrices of the appropriate sizes and so()
²%³ ~ ²%0 c´ µ ³
~ ²%0 c 0 ³ ²%0 c)³
~² %c ³ ² % 0 c) ³8 det
det det
det
c
c
()Here is the dimension of . Hence, the algebraic multiplicity of is at least =
equal to the the geometric multiplicity of .
190 Advanced Linear Algebra
The Jordan Canonical Form
One of the virtues of the rational canonical form is that every linear operator on
a finite-dimensional vector space has a rational canonical form. However, as
mentioned earlier, the rational canonical form may be far from the ideal of
simplicity that we had in mind for a set of simple canonical forms and is really
more of a theoretical tool than a practical tool.
When the minimal polynomial of splits over , ² % ³ -
² % ³~² %c ³ Ä ² %c ³
there is another set of canoncial forms that is arguably simpler than the set of
rational canonical forms.
In some sense, the complexity of the rational canonical form comes from the
choice of basis for the cyclic submodules . Recall that the -cyclic bases ºº# »»Á
have the form
8 Á Á Á Ác ~#Á#Á Ã Á #23Á
where . With this basis, all of the complexity comes at the end, ~ ² ³Á degÁ
so to speak, when we attempt to express
²² # ³ ³ ~ ² # ³c
Á ÁÁ Á
as a linear combination of the basis vectors.
However, since has the form 8Á
23#Á #Á #ÁÃÁ # c
any ordered set of the form
²³ # Á ²³ # Á Ã Á ²³ # c
where will also be a basis for . In particular, when deg² ²%³³ ~ ºº# »» ²%³ Á
splits over , the elementary divisors are -
² % ³ ~ ² % c³
Á Á
and so the set
9 Á Á Á Ác ~#Á ²c ³ #Á Ã Á ²c ³ #23Á
is also a basis for . ºº# »»Á
If we temporarily denote the th basis vector in by , then for 9Á
~ ÁÃÁ c Á ,
Eigenvalues and Eigenvectors 191
~ ´² c ³ ²# ³µ
~² c b ³ ´ ² c ³² # ³ µ
~² c ³ ² # ³b ² c ³² # ³
~ b Á
Á
Á Á b
b
For , a similar computation, using the fact that~ cÁ
²c ³ ² #³ ~ ²c ³² #³ ~ Á Áb Á
gives
² ³ ~ c c Á Á
Thus, for this basis, the complexity is more or less spread out evenly, and the
matrix of with respect to is the matrix9O d ºº# »» Á Á Á Á
@
²Á ³ ~Ä Ä
ÆÅ
Æ ÆÅ
Å ÆÆÆ
Ä Á
vy
x{x{x{x{
wz
which is called a associated with the scalar . Note that a Jordan Jordan block
block has 's on the main diagonal, 's on the subdiagonal and 's elsewhere.
Let us refer to the basis
99~Á
as a for .Jordan basis
Theorem 8.6 The Jordan canonical form() Suppose that the minimal
polynomial of splits over the base field , that is, B² = ³ -
² % ³~² %c ³ Ä ² %c ³
where .-
1 The matrix of with respect to a Jordan basis is) 9
diag@ @ @ @² Á ³ÁÃÁ ² Á ³ÁÃÁ ² Á ³ÁÃÁ ² Á ³ Á Á Á Á
where the polynomials are the elementary divisors of . This ²%c ³Á
block diagonal matrix is said to be in and is called Jordan canonical form
the .Jordan canonical form of
2 If is algebraically closed, th en up to order of the block diagonal )-
matrices, the set of matrices in Jordan canonical form constitutes a set of
canonical forms for similarity.
Proof. For part 2), the companion matrix and corresponding Jordan block are
similar:
192 Advanced Linear Algebra
*´²%c ³ µ ² Á ³@ Á Á
since they both represent the same operator on the subspace . It follows ºº# »»Á
that the rational canonical matrix a nd the Jordan canonical matrix for are
similar.
Note that the diagonal elements of the Jordan canonical form of are @
precisely the eigenvalues of , each appearing a number of times equal to its
algebraic multiplicity. In general, the rational canonical form does not “expose”
the eigenvalues of the matrix, even when these eigenvalues lie in the base field.
Triangularizability and Schur's Lemma
We have discussed two different ca nonical forms for similarity: the rational
canonical form, which applies in all cases and the Jordan canonical form, which
applies only when the base field is algebraically closed. Moreover, there is an
annoying sense in which these sets of canoncial forms leave something to be
desired: One is too complex and the other does not always exist.
Let us now drop the rather strict re quirements of canonical forms and look at
two classes of matrices that are too large to be canonical forms (the upper
triangular matrices and the almost uppe r triangular matrices) and one class of
matrices that is too small to be a canonical form (the diagonal matrices).
The upper triangular matrices or lower triangular matrices have some nice ()
algebraic properties and it is of interest to know when an arbitrary matrix is
similar to a triangular matrix. We confine our attention to upper triangular
matrices, since there are direct analogs for lower triangular matrices as well.
Definition A linear operator is if there is an B² = ³ upper triangularizable
ordered basis of for which the matrix is upper 8~² #ÁÃÁ#³ = ´ µ 8
triangular, or equivalently, if
# º# ÁÃÁ#»
for all .~ ÁÃÁ
As we will see next, when the base field is algebraically closed, all operators are
upper triangularizable. However, since tw o distinct upper triangular matrices
can be similar, the class of upper triangul ar matrices is not a canonical form for
similarity. Simply put, there are just too many upper triangular matrices.
Theorem 8.7 ()Schur's theorem Let be a finite-dimensional vector space=
over a field . -
1 If the characteristic polynomial or minimal polynomial of splits )( ) B² = ³
over , then is upper triangularizable.-
2 If is algebraically closed, then all operators are upper triangularizable. )-
Eigenvalues and Eigenvectors 193
Proof. Part 2) follows from part 1). The proof of part 1) is most easily
accomplished by matrix means, namely, we prove that every square matrix
(4² -³ - whose characteristic polynomial splits over is similar to an upper
triangular matrix. If there is nothing to prove, since all matrices are ~ d
upper triangular. Assume the result is true for and let . c (4 ²-³
Let be an eigenvector associated with the eigenvalue of and extend# - (
¸# ¹ ~²# ÁÃÁ# ³ ( to an ordered basis for . The matrix of with respect 8s
to has the form8
´µ~i
(
(
8>?
block
for some . Since and are similar, we have ( 4 ² - ³ ´ µ ( c (8
det det det²%0 c(³ ~ ²%0 c´ µ ³ ~ ²%c ³ ²%0 c( ³ ( 8
Hence, the characteristic polynomial of also splits over and the induction (-
hypothesis implies that there is an invertible matrix for which 74 ² - ³ c
<~7 (7 c
is upper triangular. Hence, if
8~
7>?
block
then is invertible and8
8´(µ 8 ~ ~ i i
7 ( <
78c
c >? > ? > ? > ?
is upper triangular.
The Real Case
When the base field is , an operator is upper triangularizable if and -~s
only if its characteristic polynomial splits over . (Why?) We can, however, s
always achieve a form that is close to triangular by permitting values on the first
subdiagonal.
Before proceeding, let us recall Theorem 7.11, which says that for a module >
of prime order , the following are equivalent: ² % ³
1 is cyclic)>
2 is indecomposable)>
3 is irreducible)² % ³
4 is nonderogatory, that is, ) ² % ³~² % ³
5 .)d i m d e g²> ³ ~ ²²%³³
194 Advanced Linear Algebra
Now suppose that and is an irreducible quadratic. -~ ² % ³~%b %b!s
If is a -cyclic basis for , then8 >
´µ ~c !
c 8>?
However, there is a more appealing matrix representation of . To this end, let
(( be the matrix above. As a complex ma trix, has two distinct eigenvalues:
~c f ! c
j
Now, a matrix of the form
)~c
>?
has characteristic polynomial and eigenvalues . So ²%³ ~ ²%c³ b f
if we set
~c ~c ! c
andj
then has the same two distinct eigenvalues as and so and have the)( ( )
same Jordan canonical form over . It follows that and are similar over dd ()
and therefore also over , by Theorem 7.20. Thus, there is an ordered basis s9
for which . ´µ~ )9
Theorem 8.8 If and is cyclic and , then there is an-~ > ² ² % ³ ³~s deg
ordered basis for which 9
´µ~c
9>?
Now we can proceed with the real version of Schur's theorem. For the sake of
the exposition, we make the following definition.
Definition A matrix is if it has the form (4² -³ almost upper triangular
(~(i
(
Æ
(vy
x{x{
wz
block
where
Eigenvalues and Eigenvectors 195
(~ ´ µ (~c
or >?
for . A linear operator is ifÁ - ²= ³ B almost upper triangularizable
there is an ordered basis for which is almost upper triangular. 8 ´µ8
To see that every real linear operator is almost upper triangularizable, we use
Theorem 7.19, which states that if is a prime factor of , then has a ²%³ ²%³ =
cyclic submodule of order . Hence, is a -cyclic subspace of > ² % ³ >
dimension and has characteristic polynomial . deg²²%³³ O ²%³>
Now, the minimal polynomial of a real operator factors into a product B² = ³
of linear and irreducible quadratic factors. If has a linear factor over , ² % ³ -
then has a one-dimensional -invariant subspace . If has an=> ² % ³
irreducible quadratic factor , then has a cyclic submodule of order ²%³ = >
²%³ > and so a matrix representation of on is given by the matrix
(~c
>?
This is the basis for an inductive proof, as in the complex case.
Theorem 8.9 If is a real vector space, then()Schur's theorem: real case =
every linear operator on is almost upper triangularizable. =
Proof. As with the complex case, it is simpler to proceed using matrices, by
showing that any real matrix is similar to an almost upper triangular d (
matrix. The result is clear if . Assume for the purposes of induction that ~
any square matrix of size less than is almost upper triangularizable. d
We have just seen that has a one-dimensional -invariant subspace or a ->(
two-dimensional -cyclic subspace , where has irreducible characteristic (( >
polynomial on . Hence, we may choose a basis for for which the first >- 8
one or first two vectors are a basis for . Then >
´µ~(i
((
8>?
block
where
(~ ´ µ (~c
or >?
and has size . The induction hy pothesis applied to gives an ( d (
invertible matrix for which 74
<~7 (7 c
196 Advanced Linear Algebra
is almost upper triangular. Hence, if
8~0
7>?c
block
then is invertible and8
8´(µ 8 ~ ~0 ( i ( i
7 ( <0
78c c
c
c >? > ? > ? > ?
is almost upper triangular.
Unitary Triangularizability
Although we have not yet discussed inner product spaces and orthonormal
bases, the reader may very well be fam iliar with these concepts. For those who
are, we mention that when is a real or complex inner product space, then if an =
operator on can be triangularized (or almost triangularized) using an =
ordered basis , it can also be triangular ized (or almost triangularized) using an 8
orthonormal ordered basis . E
To see this, suppose we apply the Gram–Schmidt orthogonalization process to a
basis that triangularizes (or almost triangularizes) . The 8~² #ÁÃÁ#³
resulting ordered orthonormal basis has the property that E~² "ÁÃÁ"³
º# ÁÃÁ#»~º" ÁÃÁ"»
for all . Since is (almost) upper triangular, that is, ´ µ 8
# º# ÁÃÁ#»
for all , it follows that
" º # ÁÃÁ #»º# ÁÃÁ#»~º" ÁÃÁ"»
and so the matrix is also (almost) upper triangular. ´µE
A linear operator is if there is an ordered unitarily upper triangularizable
orthonormal basis with respect to which is upper triangular. Accordingly,
when is an inner product space, we can replace the term “upper=
triangularizable” with “unitarily upper triangularizable” in Schur's theorem. (A
similar statement holds for almo st upper triangular matrices.)
Diagonalizable Operators
Definition A linear operator is if there is an ordered B² = ³ diagonalizable
basis of for which the matrix is diagonal, or8~² #ÁÃÁ#³ = ´ µ 8
equivalently, if
Eigenvalues and Eigenvectors 197
#~ #
for all .~ ÁÃÁ
The previous definition leads imme diately to the following simple
characterization of diagonalizable operators.
Theorem 8.10 Let . The following are equivalent:B² = ³
1 is diagonalizable.)
2 has a basis consisting entirely of eigenvectors of .)=
3 has the form)=
=~ l Ä l;;
where are the distinct eigenvalues of . ÁÃÁ
Diagonalizable operators can also be ch aracterized in a simple way via their
minimal polynomials.
Theorem 8.11 A linear operator on a finite-dimensional vector space B² = ³
is diagonalizable if and only if its mini mal polynomial is the product of distinct
linear factors.
Proof. If is diagonalizable, then
=~ l Ä l;;
and Theorem 7.7 implies that is the least common multiple of the ² % ³
minimal polynomials of restricted to . Hence, is a product of %c ²%³ ;
distinct linear factors. Conversely, if is a product of distinct linear² % ³
factors, then the primary decomposition of has the form =
=~ =l Ä l =
where
= ~ ¸# = ² c ³# ~ ¹ ~ ;
and so is diagonalizable.
Spectral Resolutions
We have seen (Theorem 2.25) that reso lutions of the identity on a vector space
== correspond to direct sum decompositions of . We can do something similar
for any linear operator on ( not just the identity operator). diagonalizable =
Suppose that has the form
~b Ä b
where is a resolution of the identity and the are bÄb ~ -
distinct. This is referred to as a of . spectral resolution
198 Advanced Linear Algebra
We claim that the 's are the eigenvalues of and . Theorem 2.25 ; im²³ ~
implies that
=~ ² ³ l Ä l ² ³ im im
If , then# ² ³im
²# ³ ~ ² b Ä b ³# ~ ²# ³
and so . Hence, and so; ;# ² ³ im
= ~ ² ³lÄl ² ³ lÄl = im im ; ;
which implies that and im²³ ~;
=~ l Ä l;;
The converse also holds, for if and if is projection onto =~ l Ä l;;
; along the direct sum of the other eigenspaces, then
bÄb ~
and since , it follows that ~
~² b Ä b ³ ~ b Ä b
Theorem 8.12 A linear operator is diagonalizable if and only if it B² = ³
has a spectral resolution
~b Ä b
In this case, is the spectrum of and ¸Á à Á¹
im²³ ~ ²³ ~; ;
£and ker
Exercises
1. Let be the matrix all of whose entries are equal to . Find the 1 d
minimal polynomial and character istic polynomial of and the 1
eigenvalues.
2. Prove that the eigenvalues of a matrix do not form a complete set of
invariants under similarity.
3. Show that is invertible if and only if is not an eigenvalue of . B ² = ³
4. Let be an matrix over a field that contains all roots of the ( d -
characteristic polynomial of . Prove that is the product of the (² ( ³ det
eigenvalues of , counting multiplicity. (
5. Show that if is an eigenvalue of , then is an eigenvalue of , for ² ³ ² ³
any polynomial . Also, if , then is an eigenvalue for . ²%³ £ c c
6. An operator is if for some positive . B o² = ³ ~ nilpotent
Eigenvalues and Eigenvectors 199
a Show that if is nilpotent, then the spectrum of is . ) ¸¹
b Find a nonnilpotent operator with spectrum . ) ¸¹
7. Show that if and one of and is invertible, then B Á² = ³
and so and have the same eigenvalues, counting multiplicty.
8. Halmos()
a Find a linear operator that is not idempotent but for which )
²c³ ~ .
b Find a linear operator that is not idempotent but for which )
²c³~ .
c Prove that if , then is idempotent. ) ²c³ ~²c³~
9. An is a linear operator for which . If is idempotent involution ~
what can you say about ? Construct a one-to-one correspondence c
between the set of idempotents on and the set of involutions. =
10. Let and suppose that but ( Á)4² ³ ( ~) ~0Á( ) (~) c d
(£0 )£0 *4² ³ ( ) and . Show that if commutes with both and , d
then for some scalar .*~ 0 d
11. Let and letB² = ³
:~º # Á # ÁÃÁ # »c
be a -cyclic submodule of with minimal polynomial where = ²%³ ²%³
is prime of degree . Let restricted to . Show that is the ~ ² ³ º # » :
direct sum of -cyclic submodules each of dimension , that is,
:~; lÄl;
Hint: For each , consider the set
8 c ~¸ #Á² ³ #ÁÃÁ² ³ #»
12. Fix . Show that any complex matrix is similar to a matrix that looks
just like a Jordan matrix except that the entries that are equal to are
replaced by entries with value , where is any complex number. Thus, any
complex matrix is similar to a matr ix that is “almost” diagonal. : Hint
consider the fact that
vy vy v y vy
wz wz w z wz
~
c
c
13. Show that the Jordan canonical form is not very robust in the sense that a
small change in the entries of a matrix may result in a large jump in the(
entries of the Jordan form . : consider the matrix 1Hint
(~
>?
What happens to the Jordan form of as ? (¦
200 Advanced Linear Algebra
14. Give an example of a complex nonreal matrix all of whose eigenvalues are
real. Show that any such matrix is si milar to a real matrix. What about the
type of the invertible matrices that are used to bring the matrix to Jordan
form?
15. Let be the Jordan form of a linear operator . For a given1~´µ ² =³ B8
Jordan block of let be the subspace of spanned by the basis 1² Á³ < =
vectors of associated with that block.8
a Show that has a single eigenvalue with geometric multiplicity . ) O<
In other words, there is essentially only one eigenvector up to scalar (
multiple associated with each Jordan block. Hence, the geometric )
multiplicity of for is the number of Jordan blocks for . Show that
the algebraic multiplicity is the sum of the dimensions of the Jordan
blocks associated with .
b Show that the number of Jordan blocks in is the maximum number ) 1
of linearly independent eigenvectors of .
c What can you say about the Jordan blocks if the algebraic multiplicity )
of every eigenvalue is equal to its geometric multiplicity?
16. Assume that the base field is algebraically closed. Then assuming that the -
eigenvalues of a matrix are known, it is possible to determine the Jordan (
form of by looking at the rank of various matrix powers. A matrix is1( )
nilpotent if for some . The smallest such exponent is called)~
the .index of nilpotence
a Let be a single Jordan block of size . Show that )1~1 ²Á ³ d
1c 0 is nilpotent of index . Thus, is the smallest integer for
which . rk²1 c 0³ ~
Now let be a matrix in Jordan form but possessing only one eigenvalue 1
.
b Show that is nilpotent. Let be its index of nilpotence. Show ) 1c 0
that is the maximum size of the Jordan blocks of and that1
rk²1 c 0³ 1c is the number of Jordan blocks in of maximum size.
c Show that is equal to times the number of Jordan )r k ²1 c 0³ c
blocks of maximum size plus the number of Jordan blocks of size one
less than the maximum.
d Show that the sequence for uniquely )r k ²1c 0³ ~ÁÃÁ
determines the number and size of all of the Jordan blocks in , that is, 1
it uniquely determines up to the order of the blocks. 1
e Now let be an arbitrary Jordan matrix. If is an eigenvalue for ) 11
show that the sequence for where is the rk²1c 0³ ~ÁÃÁ
first integer for which uniquely rk rk²1 c 0³ ~ ²1 c 0³ b
determines up to the order of the blocks. 1
f Prove that for any matrix with spectrum the sequence ) (¸ Á Ã Á ¹
rk²(c 0³ ~ÁÃÁ ~ÁÃÁ for and where is the first
integer for which uniquely rk rk²(c 0³ ~ ²(c 0³ b
determines the Jordan matrix for up to the order of the blocks. 1(
17. Let .( ² -³C
Eigenvalues and Eigenvectors 201
a If all the roots of the characteristic polynomial of lie in prove that ) (-
(( ) is similar to its transpose . Hint: Let be the matrix!
)~Ä
ÅÇ
ÇÇÅ
Ä vy
x{x{
wz
with 's on the diagonal that moves up from left to right and 's
elsewhere. Let be a Jordan block of the same size as . Show that 1)
)1) ~ 1c !.
b Let . Let be a field cont aining . Show that if and )(Á) ²-³ 2 - (C
)2 ) ~ 7 ( 7 7 ² 2 ³ are similar over , that is, if where , thencC
() - 8 ² - ³ and are also similar over , that is, there exists for C
which .)~8 ( 8c
c Show that any matrix is similar to its transpose. )
The Trace of a Matrix
18. Let . Verify the following statements.( ² -³C
a) A , for . tr tr² ³ ~ ²(³ -
b) . tr tr tr²(b)³ ~ ²(³b ²)³
c) . tr tr²()³ ~ ²)(³
d) . Find an example to show that tr tr tr²()*³ ~ ²*()³ ~ ²)*(³
tr tr²()*³ ²(*)³ may not equal .
e) The trace is an invariant under similarity.
f) If is algebraically closed, then the trace of is the sum of the -(
eigenvalues of . (
19. Use the concept of the trace of a matrix, as defined in the previous exercise,
to prove that there are no matrices , for which () ²³Cd
() c)( ~ 0
20. Let be a function with the following properties. For all ;¢ ²-³¦-C
matrices and , (Á ) ²-³ -C
1 A );² ³~;²(³
2 );²(b)³~;²(³b;²)³
3 );²()³~;²)(³
Show that there exists for which tr , for all - ;² ( ³~ ² ( ³
( ² -³C .
Commuting Operators
Let
< B ?~¸ ² =³ ¹
be a family of operators on a vector space . Then is a if =< commuting family
every pair of operators commutes, that is, for all . A subspace <~Á
202 Advanced Linear Algebra
<= of is if it is -invariant fo r every . It is often of interest <-invariant <
to know whether a family of linear operators on has a < = common
eigenvector , that is, a single vector that is an eigenvector for every #=
< (the corresponding eigenvalues may be different for each operator,
however).
21. A pair of linear operators is if BÁ² = ³ simultaneously diagonalizable
there is an ordered basis for for which and are both diagonal, 8 =´ µ ´ µ 88
that is, is an ordered basis of eigenvectors for both and . Prove that8
two diagonalizable operators and are simultaneously diagonalizable if
and only if they commute, that is, . : If , then the ~~ Hint
eigenspaces of are invariant under .
22. Let . Prove that if and commute, then every eigenspace of B Á² = ³
< is -invariant. Thus, if is a commuting family, then every eigenspace
of any member of is -invariant. <<
23. Let be a family of operators in with the property that each operator<B ²= ³
in has a full set of eigenvalues in the base field , that is, the < -
characteristic polynomial splits over . Prove that if is a commuting - <
family, then has a common eigenvector . < #=
24. What do the real matrices
(~ )~
c c >? >? and
have to do with the issue of common eigenvectors?
Geršgorin Disks
It is generally impossible to determin e precisely the eigenvalues of a given
complex operator or matrix , for if , then the characteristic ( ² ³ Cd
equation has degree and cannot in general be solved. As a result, the
approximation of eigenvalues is big business. Here we consider one aspect of
this approximation problem, which also has some interesting theoretical
consequences.
Let and suppose that where . Comparing( ² ³ (#~ # #~² ÁÃÁ ³Cd !
th rows gives
~
(~
which can also be written in the form
² c( ³~ (
~
£
If has the property that for all , we have (( ( (
Eigenvalues and Eigenvectors 203
(( ( ( ( ( ( ( (( ( ( c ( ( (
~ ~
£ £
and thus
(( ( ( c( (
~
£
()8.7
The right-hand side is the sum of the abso lute values of all entries in the th row
of except the diagonal entry . This sum is the th (( 9 ² ( ³ deleted absolute
row sum of . The inequality 8.7 says th at, in the complex plane, the ( ()
eigenvalue lies in the disk centered at the diagonal entry with radius equal (
to . This disk9² ( ³
GR ²(³ ~ ¸' ' c( 9 ²(³¹ d((
is called the for the th row of . The union of all of the Geršgorin row disk (
Geršgorin row disks is called the for . Geršgorin row region (
Since there is no way to know in general which is the index for which
(( ( ( ( , the best we can say in general is that the eigenvalues of lie in the
union of all Geršgorin row disks, that is, in the Geršgorin row region of . (
Similar definitions can be made for columns and since a matrix has the same
eigenvalues as its transpose, we can say that the eigenvalu es of lie in the(
Geršgorin column region of . The of a matrix (. ² ( ³ Geršgorin region
(4² -³ is the intersection of the Geršgorin row region and the Geršgorin
column region and we can say that a ll eigenvalues of lie in the Geršgorin (
region of . In symbols, . (( . (
25. Find and sketch the Geršgorin region and the eigenvalues for the matrix
(~
vy
wz
26. A matrix is if for each , (4 ² ³ ~ÁÃÁ d diagonally dominant
((( 9 ² ( ³
and it is if strict inequality holds. Prove that strictly diagonally dominant
if is strictly diagonally dom inant, then it is invertible. (
27. Find a matrix that is diagona lly dominant but not invertible. (4² ³ d
28. Find a matrix that is inver tible but not strictly diagonally (4² ³ d
dominant.
Chapter 9
Real and Complex Inner Product Spaces
We now turn to a discussion of real and complex vector spaces that have an
additional function defined on them, ca lled an , as described in the inner product
following definition. In this chapter, will denote either the real or complex -
field. Also, the complex conjugate of is denoted by . d
Definition Let be a vector space over or . An =- ~ - ~ sd inner product
on is a function with the following properties:= ºÁ»¢= d= ¦-
1 For all ,)( )Positive definiteness #=
º#Á#» º#Á#» ~ ¯ # ~ and
2 F o r )( )-~d:Conjugate symmetry
º"Á#» ~ º#Á"»
For-~s:( )Symmetry
º"Á#» ~ º#Á"»
3 For all and )( )Linearity in the first coordinate "Á# = Á -
º"b #Á$» ~ º"Á$»b º#Á$»
A real or complex vector space , together with an inner product, is called a () =
real complex inner product space or .()
If , then we let?Á@ =
º?Á@»~¸º%Á&»%?Á&@¹
and
º#Á?» ~ ¸º#Á%» % ?¹
Note that a vector subspace of an inner product space is also an inner :=
product space under the restriction of the inner product of to . =:
206 Advanced Linear Algebra
We will study bilinear forms (also called ) on vector spaces over inner products
fields other than or in Chapter 11. Note that property 1) implies that sd º#Á#»
is always real, even if is a complex vector space. =
If , then properties 2) and 3) imply that the inner product is linear in both-~s
coordinates, that is, the inner product is . However, if , then bilinear -~d
º$Á"b #»~º"b #Á$»~º"Á$»b º#Á$»~º$Á"»b º$Á#»
This is referred to as in the second coordinate. Specifically, conjugate linearity
a function between complex vector spaces is if ¢= ¦> conjugate linear
²"b#³~²"³b²#³
and
²"³ ~ ²"³
for all and . Thus, a complex inner product is linear in its first"Á# = d
coordinate and conjugate linear in its s econd coordinate. This is often described
by saying that a complex inner product is . (Sesqui means “one and sesquilinear
a half times.”)
Example 9.1
1) The vector space is an inner product space under the sstandard inner
product dot product , or , defined by
º² ÁÃÁ ³Á² ÁÃÁ ³» ~ bÄb
The inner product space is often called s-dimensional Euclidean
space .
2) The vector space is an inner product space under the dstandard inner
product defined by
º² ÁÃÁ ³Á² ÁÃÁ ³» ~ bÄb
This inner product space is often called . -dimensional unitary space
3) The vector space of all continuous complex-valued functions on the *´Áµ
closed interval is a complex inner product space under the inner ´Áµ
product
ºÁ» ~ ²%³²%³ %
Example 9.2 One of the most important inner product spaces is the vector space
M² ³ of all real (or complex) sequences with the property that
(( B
Real and Complex Inner Product Spaces 207
under the inner product
º² ³Á²! ³» ~ !
~B
Such sequences are called . Of course, for this inner product square summable
to make sense, the sum on the right must converge. To see this, note that if
² ³Á²! ³ M, then
² c! ³ ~ c ! b! (( (( (( (( (( ((
and so
! b!(( ( (( (
which implies that . We leave it to the reader to verify that is an ² ! ³ M M
inner product space.
The following simple result is quite useful.
Lemma 9.1 If is an inner product space and for all ,= º"Á%» ~ º#Á%» % =
then ."~#
The next result points out one of the main differences between real and complex
inner product spaces and will play a key role in later work.
Theorem 9.2 Let be an inner product space and let .= ² = ³ B
1)
º #Á$»~ #Á$= ¬ ~ for all
2 If is a complex inner product space, then)=
º# Á # » ~ # = ¬ ~ for all
but this does not hold in general for real inner product spaces.
Proof. Part 1) follows directly from Lemma 9.1. As for part 2), let , #~ %b&
for and . Then%Á& = -
~ º ²%b&³Á%b&»
~º% Á % » b º& Á & » b º% Á & » b º& Á % »
~ º % Á& »b º & Á% »
((
Setting gives~
º% Á & » b º& Á % » ~
and setting gives ~
208 Advanced Linear Algebra
º% Á & » c º& Á % » ~
These two equations imply that for all and so part 1) º %Á&»~ %Á&=
implies that . For the last statement, rotation by degrees in the real ~
plane has the property that for all .sº# Á # » ~ #
Norm and Distance
If is an inner product space, the , or of is defined by=# = norm length
)) j#~ º # Á # » ()9.1
A vector is a if . Here are the basic properties of the norm. ## ~ unit vector ))
Theorem 9.3
1 and if and only if .))) ))# #~ # ~
2 For all and ,) - #=
)) ( ( ) )# ~ #
3 For all ,)( )The Cauchy–Schwarz inequality "Á# =
(( ) ) ) )º"Á#» " #
with equality if and only if one of and is a scalar multiple of the other. "#
4 For all ,)( )The triangle inequality "Á# =
)) ) ) ) )"b# " b #
with equality if and only if one of and is a scalar multiple of the other. "#
5 For all ,) "Á#Á% =
)) )) ))"c# "c% b %c#
6 For all ,) "Á# =
(( ) ))) ))"c# " c #
7 For all ,)( )The parallelogram law "Á# =
)) ))) ) ) )"b# b "c# ~ " b #
Proof. We prove only Cauchy–Schwarz and the triangle inequality. For
Cauchy–Schwarz, if either or is zero the result follows, so assume that "#
"Á# £ - . Then, for any scalar ,
"c #
~º "c # Á"c # »
~ º"Á"»cº"Á#»c´º#Á"»cº#Á#»µ))
Choosing makes the value in the square brackets equal to ~ º#Á"»°º#Á#»
Real and Complex Inner Product Spaces 209
and so
º " Á" »c ~ " cº#Á"»º"Á#» º"Á#»
º#Á#» #))((
))
which is equivalent to the Cauchy–Sch warz inequality. Furthermore, equality
holds if and only if , that is, if and only if , which is ))"c# ~ "c#~
equivalent to and being s calar multiples of one another. "#
To prove the triangle inequality, th e Cauchy–Schwarz inequality gives
))
)) )) )) ))
)) ))"b# ~º"b#Á"b#»
~ º"Á"»bº"Á#»bº#Á"»bº#Á#»
" b "#b#
~²" b #³
from which the triangle inequality fo llows. The proof of the statement
concerning equality is left to the reader.
Any vector space , together with a function that satisfies =h ¢ = ¦ ))s
properties 1), 2) and 4) of Theorem 9.3, is called a and the normed linear space
function is called a . Thus, any inner product space is a normed linear ))h norm
space, under the norm given by 9.1 . ()
It is interesting to observe that the inner product on can be recovered from the =
norm. Thus, knowing the length of all vectors in is equivalent to knowing all=
inner products of vectors in . =
Theorem 9.4 ()The polarization identities
1 If is a real inner product space, then)=
º"Á#»~ ² "b# c "c# ³
)) ))
2 If is a complex inner product space, then)=
º"Á#»~ ² "b# c "c# ³b ² "b# c "c# ³
)) )) ) ) ) )
The norm can be used to define the distance between any two vectors in an
inner product space.
Definition Let be an inner product space. The between any= ² " Á # ³ distance
two vectors and in is "# =
²"Á#³ ~ "c# )) ()9.2
Here are the basic properties of distance.
210 Advanced Linear Algebra
Theorem 9.5
1 and if and only if )²"Á#³ ²"Á#³ ~ " ~ #
2)( )Symmetry
²"Á#³ ~ ²#Á"³
3)( )The triangle inequality
²"Á#³ ²"Á$³b²$Á#³
Any nonempty set , together with a function that satisfies the = ¢ = d = ¦ s
properties of Theorem 9.5, is called a and the function is called metric space
a on . Thus, any inner product space is a metric space under the metricmetric =
()9.2 .
Before continuing, we should make a few remarks about our goals in this and
the next chapter. The presence of an inner product, and hence a metric, permits
the definition of a topology on , and in particular, convergence of infinite =
sequences. A sequence of vectors in to if ²# ³ = # = converges
lim
¦B))#c #~
Some of the more important concepts related to convergence are closedness and
closures, completeness and the continuity of linear operators and linear
functionals.
In the finite-dimensional case, the situation is very straightforward: All
subspaces are closed, all inner product spaces are complete and all linear
operators and functionals are continuous. However, in the infinite-dimensional
case, things are not as simple.
Our goals in this chapter and the next are to describe some of the basic
properties of inner product spaces—both finite and infinite-dimensional—and
then discuss certain special types of ope rators (normal, unitary and self-adjoint)
in the finite-dimensional case only. To achieve the latter goal as rapidly as
possible, we will postpone a discussion of convergence-related properties until
Chapter 12. This means that we must state some results only for the finite-
dimensional case in this chapter.
Isometries
An isomorphism of vector spaces preserves the vector space operations. The
corresponding concept for inner product spaces is the . isometry
Definition Let and be inner product spaces and let .=> ² = Á > ³ B
Real and Complex Inner Product Spaces 211
1 is an if it preserves the inner product, that is, if) isometry
º" Á# » ~ º " Á # »
for all ."Á# =
2 A bijective isometry is called an . When ) isometric isomorphism ¢= ¦>
is an isometric isomorphism, we say that and are => isometrically
isomorphic .
It is clear that an isometry is injective and so it is an isometric isomorphism
provided it is surjective. Moreover, if
dim dim²= ³ ~ ²>³ B
injectivity implies su rjectivity and is an isometry if and only if is an
isometric isomorphism. On the other ha nd, the following simple example shows
that this is not the case for infinite-dimensional inner product spaces.
Example 9.3 The map defined by¢M ¦M
²% Á% Á% Áó~²Á% Á% Áó
is an isometry, but it is clearly not surjective.
Since the norm determines the inner product, the following should not come as a
surprise.
Theorem 9.6 A linear transformation is an isometry if and only if B² = Á > ³
it preserves the norm, that is, if and only if
)) ) )#~#
for all .#=
Proof. Clearly, an isometry preserves the norm. The converse follows from the
polarization identities. In the real case, we have
º" Á# » ~ ² " b# c " c#³
~ ² ²"b#³ c ²"c#³ ³
~ ² "b# c "c# ³
~º " Á# »
)) ))
)) ))
)) ))
and so is an isometry. The complex case is similar.
Orthogonality
The presence of an inner product allows us to define the concept of
orthogonality.
212 Advanced Linear Algebra
Definition Let be an inner product space.=
1 Two vectors are , written , if) "Á# = " # orthogonal
º"Á#» ~
2 Two subsets are , written , if ,) ?Á@ = ? @ º?Á@» ~ ¸¹ orthogonal
that is, if for all and . We write in place of %& %? &@ #?
¸#¹ ? .
3 The of a subset is the set) orthogonal complement ?=
?~ ¸ # = # ? ¹
The following result is easily proved.
Theorem 9.7 Let be an inner product space.=
1 The orthogonal complement of any subset is a subspace of . ) ?? = =
2 For any subspace of ,) :=
: q: ~ ¸¹
Definition An inner product space is the of = orthogonal direct sum
subspaces and if :;
=~ :l ; Á : ;
In this case, we write
:p;
More generally, is the of the subspaces , = :ÁÃÁ: orthogonal direct sum
written
:~: pÄp:
if
=~ :l Ä l : : : £ and for
Theorem 9.8 Let be an inner product space. The following are equivalent.=
1)=~ :p ;
2 and )=~ :l ; ;~ :
Proof. If , then by definition, . However, if , then=~ :p ; ; : # :
#~ b! : !; ! # where and . Then is orthogonal to both and and so
is orthogonal to itself, which implies that and so . Hence, . ~ #; ;~:
The converse is clear.
Orthogonal and Orthonormal Sets
Definition A nonempty set of vectors in an inner product E~¸ " 2¹
space is said to be an if for all . If, in orthogonal set " " £ 2
addition, each vector is a unit vector, then is an . Thus, a " E orthonormal set
Real and Complex Inner Product Spaces 213
set is orthonormal if
º" Á" » ~ Á
for all , where is the Kronecker delta function.Á 2 Á
Of course, given any nonzero vector , we may obtain a unit vector by #= "
multiplying by the reciprocal of its norm: #
"~ #
#))
This process is referred to as the vector . Thus, it is a simple normalizing #
matter to construct an orthonormal set from an orthogonal set of nonzero
vectors.
Note that if , then "#
)) ) ) ) )"b# ~ " b #
and the converse holds if . -~s
Orthogonality is stronger than linear independence.
Theorem 9.9 Any orthogonal set of nonzero vectors in is linearly =
independent.
Proof. If is an orthogonal set of nonzero vectors andE~¸ " 2¹
" bÄb " ~
then
~º "bÄb"Á "»~º "Á "»
and so , for all . Hence, is linearly independent.~ E
Gram–Schmidt Orthogonalization
The Gram–Schmidt process can be used to transform a sequence of vectors into
an orthogonal sequence. We begin with the following.
Theorem 9.10 Let be an inner product()Gram–Schmidt augmentation =
space and let be an orthogonal set of vectors in . If E~¸ "ÁÃÁ"¹ =
#¤º" ÁÃÁ" » "= ¸" ÁÃÁ" Á"¹ , then there is a nonzero for which is
orthogonal and
º" ÁÃÁ" Á"»~º" ÁÃÁ" Á#»
In particular,
214 Advanced Linear Algebra
"~#c "
~
where
~" ~
"£
º#Á" »
º" Á" » Hif
if
Proof. We simply set
"~#c" cÄc "
and force for all , that is, ""
~º"Á"»~º#c " cÄc " Á"»~º#Á"»cº"Á"»
Thus, if , take and if , take"~ ~ "£
~º#Á" »
º" Á" »
The Gram–Schmidt augmentation is traditionally applied to a sequence of
linearly independent vectors, but it also applies to any sequence of vectors.
Theorem 9.11 The Gram–Schmidt orthogonalization process () Let
8~² #Á #Á Ã ³ = be a sequence of vectors in an inner product space . Define a
sequence by repeated Gram–Schmidt augmentation, that is,E~² "Á "Á Ã ³
"~ #c " Á
~c
where and"~ #
~" ~
"£ Á
º# Á" »
º" Á" » Hif
if
Then is an orthogonal sequence in with the property thatE =
º" ÁÃÁ" »~º# ÁÃÁ# »
for all . Also, if and only if . " ~ # º# ÁÃÁ# » c
Proof. The result holds for . Assume it holds for . If ~ c
# º# ÁÃÁ# » c , then
# º# ÁÃÁ# »~º" ÁÃÁ" » c c
Writing
Real and Complex Inner Product Spaces 215
#~ "
~c
we have
º# Á" » ~" ~
º "Á"» " £
Fif
if
Therefore, when and so . Hence, ~ "£ "~ Á
º" ÁÃÁ" »~º" ÁÃÁ" Á»~º# ÁÃÁ# »~º# ÁÃÁ# » c c
If then# ¤º# ÁÃÁ# » c
º" ÁÃÁ" »~º# ÁÃÁ# Á" »~º# ÁÃÁ# Á# » c c
Example 9.4 Consider the inner product space of real polynomials, with s´%µ
inner product defined by
º²%³Á²%³» ~ ²%³²%³%
c
Applying the Gram–Schmidt process to the sequence 8~² Á % Á %Á %Á Ã ³
gives
"² % ³~
"² % ³~%c h~%% %
%
"² % ³~% c hc h%~% c% % % %
% %%
"² % ³~%c hc h%c% % % %
% %%
c
c
c c
c c
c c c
c c
3
4
c
%² %c³ %
²% c ³ %h%c
~% c %
3
3>?
and so on. The polynomials in this sequence are at least up to multiplicative (
constants the . ) Legendre polynomials
The QR Factorization
The Gram–Schmidt process can be used to factor any real or complex matrix
into a product of a matrix with orthogonal columns and an upper triangular
matrix. Suppose that is an matrix with columns (~² # # Ä#³ d
# , where . The Gram–Schmidt process applied to these columns gives
orthogonal vectors for which 6~² " " Ä"³
216 Advanced Linear Algebra
º" ÁÃÁ" »~º# ÁÃÁ# »
for all . In particular,
#~ "b " Á
~c
where
~" ~
"£ Á
º# Á" »
º" Á" » Hif
if
In matrix terms,
² # # Ä#³~² " " Ä"³ Ä
Ä
Æ
Á Á
Ávy
x{x{
wz
that is, where has orthogonal columns and is upper triangular.(~6 ) 6 )
We may normalize the nonzero columns of and move the positive "6
constants to . In particular, if for and for , then ) ~" "£ ~ "~ ))
² # # Ä#³~ Ä"" "
Ä
Ä
Æ
Á Á
Á
67 cc c x {vy
x{
wz
and so
(~8 9
where the columns of are orthogonal and each column is either a unit vector 8
or the zero vector and is upper triangular with positive entries on the main 9
diagonal. Moreover, if the vectors are linearly independent, then the #Á Ã Á #
columns of are nonzero. Also, if and is nonsingular, then is 8 ~ ( 8
unitary/orthogonal.
If the columns of are not linearly independent, we can make one final (
adjustment to this matrix factorization. If a column is zero, then we may "°
replace this column by any vector as long as we replace the th entry in ²Á³ 9
by . Therefore, we can take nonzero columns of , extend to an orthonormal8
basis for the span of the columns of and replace the zero columns of by the 88
additional members of this orthonormal basis. In this way, is replaced by a 8
unitary/orthogonal matrix and is replaced by an upper triangular matrix 89 9ZZ
that has nonnegative entries on the main diagonal.
Real and Complex Inner Product Spaces 217
Theorem 9.12 Let , where or . There exists a( ² -³ -~ -~Cd sÁ
matrix with orthonormal columns and an upper triangular8 ² -³CÁ
matrix with nonnegative real entries on the main diagonal for9 ² -³C
which
(~8 9
Moreover, if , then is unitary/orthogonal. If is nonsingular, then ~ 8 ( 9
can be chosen to have positive entries on the main diagonal, in which case the
factors and are unique. The factorization is called the 89 ( ~ 8 9 89
factorization of the matrix . If is real, then and may be taken to be (( 8 9
real.
Proof. As to uniqueness, if is nonsingular and then (8 9 ~ 8 9
88 ~ 9 9c c
and the right side is upper triangular w ith nonzero entries on the main diagonal
and the left side is unitary. But an upper triangular matrix with positive entries
on the main diagonal is unitary if and only if it is the identity and so 8~ 8
and . Finally, if is real, then all computations take place in the real9~ 9 (
field and so and are real. 89
The decomposition has important applications. For example, a system of89
linear equations can be written in the form (% ~ "
89% ~ "
and since , we have 8~ 8c i
9% ~ 8 "i
This is an upper triangular system, which is easily solved by back substitution;
that is, starting from the bottom and working up.
We mention also that the factorization is associated with an algorithm for 89
approximating the eigenvalues of a matrix, called the . 89 algorithm
Specifically, if is an matrix, define a sequence of matrices as (~( d
follows:
1) Let be the factorization of and let .(~ 8 9 8 9 ( (~ 9 8
2) Once has been defined, let be the factorization of (( ~ 8 9 8 9 (
and let .(~ 9 8b
Then is unitarily/orthogonally similar to , since((
8( 8 ~ 8² 98³ 8 ~ 89 ~ (c c c c c c c c cii
For complex matrices, it can be shown that under certain circumstances, such as
when the eigenvalues of have distinct norms, the sequence converges ((
218 Advanced Linear Algebra
(entrywise) to an upper triangular matrix , which therefore has the eigenvalues<
of on its main diagonal. Results can be obtained in the real case as well. For(
more details, we refer the reader to [48], page 115.
Hilbert and Hamel Bases
Definition A in an inner product space is called amaximal orthonormal set =
Hilbert basis for .=
Zorn's lemma can be used to show that any nontrivial inner product space has a
Hilbert basis. We leave the details to the reader.
Some care must be taken not to confuse the concepts of a basis for a vector
space and a Hilbert basis for an inner product space. To avoid confusion, a
vector space basis, that is, a maximal linearly independent set of vectors, is
referred to as a . We will refer to an orthonormal Hamel basis as an Hamel basis
orthonormal basis .
To be perfectly clear, there are maximal linearly independent sets called
(Hamel) bases and maximal orthonormal sets (called Hilbert bases). If a
maximal linearly independent set (basis) is orthonormal, it is called an
orthonormal basis.
Moreover, since every orthonormal set is linearly independent, it follows that an
orthonormal basis is a Hilbert basis, si nce it cannot be properly contained in an
orthonormal set. For inner product spaces, the two types of finite-dimensional
bases are the same.
Theorem 9.13 Let be an inner product space. A finite subset=
E~¸ "ÁÃÁ"¹ = = of is an orthonormal Hamel basis for if and only if it is ()
a Hilbert basis for . =
Proof. We have seen that any orthonormal basis is a Hilbert basis. Conversely,
if is a finite maximal orthonormal set and , where is linearlyEE F F
independent, then we may apply part 1) to extend to a strictly larger E
orthonormal set, in contradiction to the maximality of . Hence, is maximal EE
linearly independent.
The following example shows that the pr evious theorem fails for infinite-
dimensional inner product spaces.
Example 9.5 Let and let be the set of all vectors of the form=~ M 4
~²ÁÃÁÁÁÁó
where has a in the th coordinate and 's elsewhere. Clearly, is an 4
orthonormal set. Moreover, it is maximal. For if has the property #~² %³M
that , then#4
Real and Complex Inner Product Spaces 219
%~ º # Á » ~
for all and so . Hence, no nonzero vector is orthogonal to .# ~ # ¤ 4 4
This shows that is a Hilbert basis for the inner product space . 4M
On the other hand, the vector space span of is the subspace of all 4:
sequences in that have finite support, that is, have only a finite number of M
nonzero terms and since , we see that is not a Hamel span²4³ ~ : £ M 4
basis for the vector space . M
The Projection Theorem and Best Approximations
Orthonormal bases have a great practical advantage over arbitrary bases. From a
computational point of view, if is a basis for , then each 8~¸ #ÁÃÁ#¹ =
#= has the form
#~# bÄb #
In general, determining the coordinates requires solving a system of linear
equations of size . d
On the other hand, if is an orthonormal basis for and E~¸ "ÁÃÁ"¹ =
#~" bÄb "
then the coefficients are quite easily computed:
º#Á"»~º " bÄb " Á"»~º"Á"»~
Even if is not a basis (but just an orthonormal set), we canE~¸ "ÁÃÁ"¹
still consider the expansion
# ~ º#Á" »" bÄbº#Á" »"V
Theorem 9.14 Let be an orthonormal subset of an innerE~¸ "ÁÃÁ"¹
product space and let . The with respect to of =: ~ º » EE Fourier expansion
a vector is#=
#~º#Á"»" bÄbº#Á"»"V
Each coefficient is called a of with respect to . º#Á" » # Fourier coefficient E
The vector can be characterized as follows: #V
1 is the unique vector for which .)# : ² # c ³ :V
2 is the to from within , that is, is the unique)## : #VV best approximation
vector that is closest to , in the sense that : #
)) ))#c# #c V
for all . :±¸ # ¹ V
220 Advanced Linear Algebra
3 holds for all , that is)Bessel's inequality #=
)) ))##V
Proof. For part 1), since
º#c#Á"»~º#Á"»cº#Á"»~VV
it follows that . Also, if for , then and #c#: #c : : c#:VV
c# ~ ²#c#³c²#c ³ :VV
and so . For part 2), if , then implies that ~# : #c#:VV
²#c#³ ²#c ³VV and so
)) ) ) )) ))#c ~ #c#b#c ~ #c# b #c VV V V
Hence, is smallest if and only if and the smallest value is ))#c ~# V
))#c#V. We leave proof of Bessel's inequality as an exercise.
Theorem 9.15 The If is a finite-dimensional subspace() projection theorem :
of an inner product space , then =
:~:p:
In particular, if , then #=
#~#b² #c# ³:p:VV
It follows that
dim dim dim²= ³ ~ ²:³b ²: ³
Proof. We have seen that and so . But #c# : = ~ : b: : q: ~ ¸¹V
and so .=~ :p :
The following example shows that the pr ojection theorem may fail if is not :
finite-dimensional. Indeed, in the infinite-dimensional case, must be a :
complete subspace, but we postpone a discussion of this case until Chapter 13.
Example 9.6 As in Example 9.5, let and let be the subspace of all =~ M :
sequences with finite support, that is, is spanned by the vectors :
~²ÁÃÁÁÁÁó
where has a in the th coordinate and 's elsewhere. If , then % ~ ² % ³ :
% ~ º%Á » ~ % ~ : ~ ¸¹ for all and so . Therefore, . However,
:p: ~:£M
The projection theorem has a variety of uses.
Real and Complex Inner Product Spaces 221
Theorem 9.16 Let be an inner product space and let be a finite-=:
dimensional subspace of . =
1):~ :
2 If and , then)? = ²º?»³ B dim
?~ º ? »
Proof. For part 1), it is clear that . On the other hand, if , then :: #:
the projection theorem implies that where and . Then #~ b : :ZZ
# ~ ZZ Z is orthogonal to both and and so is orthogonal to itself. Hence,
and and so . We leave the proof of part 2) as an exercise.#~ : :~:
Characterizing Orthonormal Bases
We can characterize orthonormal ba ses using Fourier expansions.
Theorem 9.17 Let be an orthonormal subset of an innerE~¸ "ÁÃÁ"¹
product space and let . The following are equivalent: =: ~ º » E
1 is an orthonormal basis for .)E =
2)º » ~ ¸¹E
3 Every vector is equal to its Fourier expansion, that is, for all ,) #=
#~#V
4 holds for all , that is,)Bessel's identity #=
)) ))#~#V
5 holds for all , that is,)Parseval's identity #Á$ =
º#Á$» ~ ´#µ h´$µ VVEE
where
´#µ h´$µ ~ º#Á" »º$Á" »bÄbº#Á" »º$Á" »VVEE
is the standard dot product in . -
Proof. To see that 1) implies 2), if is nonzero, then is#º » r¸ # °#¹EE))
orthonormal and so is not maximal. Conve rsely, if is not maximal, there is EE
an orthonormal set for which . Then any nonzero is in FE F F E # ±
º»E. Hence, 2) implies 1). We leave the rest of the proof as an exercise.
The Riesz Representation Theorem
We have been dealing with linear maps for some time. We now have a need for
conjugate linear maps.
Definition A function on complex vector spaces is ¢= ¦> conjugate linear
if it is additive,
²# b# ³ ~ # b #
222 Advanced Linear Algebra
and
²#³ ~ #
for all . A is a bijective conjugate linear map.d conjugate isomorphism
If , then the inner product function defined by%= ºhÁ% » ¢= ¦-
ºhÁ%»#~º#Á%»
is a linear functional on . Thus, the linear map defined by =¢ = ¦ = i
%~ºhÁ% »
is conjugate linear. Moreover, since implies , it follows ºhÁ%»~ºhÁ&» %~&
that is injective and therefore a conj ugate isomorphism (since is finite- =
dimensional).
Theorem 9.18 The Riesz representation theorem() Let be a finite-=
dimensional inner product space.
1 The map defined by) ¢= ¦=i
%~ºhÁ% »
is a conjugate isomorphism. In particular, for each , there exists a =i
unique vector for which , that is, %= ~ºhÁ% »
#~º#Á%»
for all . We call the for and denote it by .#= % 9 Riesz vector
2 The map defined by) 9¢= ¦ =i
9 ~ 9
is also a conjugate isomorphism, being the inverse of . We will call this
map the . Riesz map
Proof. Here is the usual proof that is surjective. If , then , so let ~ 9 ~
us assume that . Then has codimension and so £ 2~ ² ³ ker
=~ º $ » p 2
for . Letting for , we require that$2 %~ $ -
²#³ ~ º#Á $»
and since this clearly holds for any , it is sufficient to show that it holds#2
for , that is,#~$
²$³ ~ º$Á $» ~ º$Á$»
Thus, and~² $ ³ °$ ))
Real and Complex Inner Product Spaces 223
9~ $²$³
$ ))
For part 2), we have
º#Á9 » ~ ² b ³²#³
~ ²#³b ²#³
~º # Á 9»bº # Á 9»
~º # Á 9 b 9»b
for all and so#=
9~ 9 b 9b
Note that if , then , where is the = ~ 9 ~ ²² ³ÁÃÁ² ³³ ² ÁÃÁ ³s
standard basis for . s
Exercises
1. Prove that if a matrix is unitary, upper triangular and has positive entries 4
on the main diagonal, must be the identity matrix.
2. Use the QR factorization to show that any triangularizable matrix is
unitarily (orthogonally) triangularizable.
3. Verify the statement concerning equality in the triangle inequality.
4. Prove the parallelogram law.
5. Prove the Apollonius identity
)) ))) ) hh $c" b $c# ~ "c# b $c ²"b#³
6. Let be an inner product space with basis . Show that the inner product= 8
is uniquely defined by the values , for all . º"Á#» "Á# 8
7. Prove that two vectors and in a real inner product space are "# =
orthogonal if and only if
)) ) ) ) )"b# ~ " b #
8. Show that an isometry is injective.
9. Use Zorn's lemma to show that any nontrivial inner product space has a
Hilbert basis.
10. Prove Bessel's inequality.
11. Prove that an orthonormal set is a Hilbert basis for a finite-dimensionalE
vector space if and only if , for all . =# ~ # # = V
12. Prove that an orthonormal set is a Hilbert basis for a finite-dimensionalE
vector space if and only if Bessel's identity holds for all , that is, if =# =
and only if
224 Advanced Linear Algebra
)) ))#~#V
for all .#=
13. Prove that an orthonormal set is a Hilbert basis for a finite-dimensionalE
vector space if and only if Parseval's identity holds for all , that =# Á $ =
is, if and only if
º#Á$» ~ ´#µ h´$µ VVEE
for all .#Á$ =
14. Let and be in . The Cauchy–Schwarz"~² ÁÃÁ ³ #~² ÁÃÁ ³ s
inequality states that
(( bÄb ² bÄb ³² bÄb ³
Prove that we can do better:
² bÄb ³ ² bÄb ³² bÄb ³(( ((
15. Let be a finite-dimensional inner product space. Prove that for any subset=
?= ?~ ² ? ³ of , we have .span
16. Let be the inner product space of all polynomials of degree at most 3,F3
under the inner product
º²%³Á²%³» ~ ²%³²%³ %
cBB
c%
Apply the Gram–Schmidt process to the basis , thereby ¸ Á % Á %Á %¹
computing the first four at least up to a Hermite polynomials (
multiplicative constant . )
17. Verify uniqueness in the Riesz representation theorem.
18. Let be a complex inner product space and let be a subspace of . =: =
Suppose that is a vector for which for all # = º#Á »bº Á#» º Á »
: #: . Prove that .
19. If and are inner product spaces, consider the function on => = > ^
defined by
º²# Á$ ³Á²# Á$ ³» ~ º# Á# »bº$ Á$ »
Is this an inner product on ? =>^
20. A over or is a vector space over or normed vector space sd sd ()
together with a function for which for all and scalars ))¢= ¦ "Á#= s
we have
a ))) ( ( ) )# ~ #
b ))) ) ) ) )"b# " b #
c if and only if )))#~ # ~
If is a real normed space over and if the norm satisfies the= ()s
parallelogram law
Real and Complex Inner Product Spaces 225
)) ))) ) ) )"b# b "c# ~ " b #
prove that the polarization identity
º"Á#»~ ² "b# c "c# ³
)) ))
defines an inner product on . : Evaluate to show = º " Á % » b º # Á % »Hint
that and . Then complete theº"Á%» ~ º"Á%» º"Á%»bº#Á%» ~ º"b#Á%»
proof that . º"Á%»~º"Á%»
21. Let be a subspace of a finite-dimensional inner product space . Prove:=
that each coset in contains =°: exactly one vector that is orthogonal to . :
Extensions of Linear Functionals
22. Let be a linear functional on a subspace of a finite-dimensional inner:
product space . Let . Suppose that is an extension = ² # ³ ~ º # Á 9 » = i
of , that is, . What is the relationship between the Riesz vectors O ~ 9 :
and ?9
23. Let be a nonzero linear functional on a subspace of a finite-dimensional:
inner product space and let . Show that if is an =2 ~ ² ³ = keri
extension of , then . Moreover, for each vector 9 2 ± :
"2 ±: there is exactly one scalar for which the linear functional
²?³~º?Á "» is an extension of .
Positive Linear Functionals on s
A vector in is also called , written #~² ÁÃÁ ³ s nonnegative positive ()
# # # # , if for all . The vector is , written , if is strictly positive
nonnegative but not . The set of all strictly positive vectors in is called ssb
the in The vector is , writtennonnegative orthant strongly positive sÀ#
# , if for all . The set , of all strongly positive vectors in is bbss
the in strongly positive orthant sÀ
Let be a linear functional on a subspace of . Then is¢:¦ : ss
nonnegative also called , written , if() positive
#¬² # ³
for all and is , written , if#: strictly positive
#¬² # ³
for all #:À
24. Prove that a linear functional on is positive if and only if and 9 s
strictly positive if and only if . If is a subspace of is it true 9 :s
that a linear functional on is nonnegative if and only if ? : 9
226 Advanced Linear Algebra
25. Let be a strictly positive linear functional on a subspace of .¢:¦ :ss
Prove that has a strictly positive extension to . Use the fact that if s
< q ~ ¸¹s
b , where
sb
~¸² ÁÃÁ ³ ¹ all
and is a subspace of , then contains a strongly positive vector.<< s
26. If is a real inner product space, then we can define an inner product on its=
complexification as follows this is th e same formula as for the ordinary =d(
inner product on a complex vector space : )
º"b#Á%b&» ~ º"Á%»bº#Á&»b²º#Á%»cº"Á&»³
Show that
)) ) ) ) )²"b#³ ~ " b #
where the norm on the left is induced by the inner product on and the =d
norm on the right is induced by the inner product on . =
Chapter 10
Structure Theory for Normal Operators
Throughout this chapter, all vector spaces are assumed to be finite-dimensional
unless otherwise noted. Also, the field is either or . -sd
The Adjoint of a Linear Operator
The purpose of this chapter is to study the structure of certain special types of
linear operators on finite-dimensional real and complex inner product spaces. In
order to define these operators, we introduce another type of adjoint (different
from the operator adjoint of Chapter 3).
Theorem 10.1 Let and be finite-dimensional inner product spaces over => -
and let . Then there is a unique function , defined byB ² = Á > ³ ¢ > ¦ =i
the condition
º# Á $ » ~ º # Á $ »i
for all and . This function is in and is called the #= $> ² >Á=³ B adjoint
of .
Proof. If exists, then it is unique, for ifi
º# Á $ » ~ º # Á$ »
then for all and and so .º#Á $» ~ º#Á $» # $ ~ ii
We seek a linear map for which i¢> ¦=
º#Á $» ~ º #Á$»i
By way of motivation, the vector , if it exists, looks very much like a lineari$
map sending to . The only problem is that is supposed to be a #º # Á $ » #i
vector, not a linear map. But the Riesz representation theorem tells us that linear
maps can be represented by vectors.
228 Advanced Linear Algebra
Specifically, for each , the linear functional defined by $> = $i
# ~ º# Á $ »$
has the form
# ~ º # Á 9»$ $
where is the Riesz vector for . If is defined by9 = ¢ > ¦ =$i
$
i
$$~9 ~9 ² ³$
where is the Riesz map, then9
º#Á $» ~ º#Á9 » ~ # ~ º #Á$»i
$$
Finally, since is the composition of the Riesz map and the map i~9k 9
¢$ª $ and since both of these maps are conjugate linear, their composition
is linear.
Here are some of the basic properties of the adjoint.
Theorem 10.2 Let and be finite-dimensional inner product spaces. For=>
every and , BÁ² = Á > ³ -
1)²b³~ b ii i
2)² ³ ~ ii
3 and so)ii~
º# Á $ » ~ º # Á $ »i
4 I f , t h e n )=~ > ² ³~ ii i
5 If is invertible, then ) ²³ ~ ² ³c i i c
6 If and , then .)=~ > ² % ³ ´ % µ ²³~ ² ³ s ii
Moreover, if and is a subspace of , then B² = ³ : =
7 is -invariant if and only if is -invariant.)::i
8 reduces if and only if is bot h -invariant and -invariant, in )²:Á: ³ :i
which case
²O³~ ² ³ O::ii
Proof. For part 7), let and and write : ':
º' Á » ~ º ' Á »i
Now, if is -invariant, then for all and so and :º ' Á » ~ : ' : ii
:: º ' Á » ~ i i is -invariant. Conversely, if is -invariant, then for all
': : ~: : and so , whence is -invariant.
The first statement in part 8) follows from part 7) applied to both and . For ::
the second statement, since is bot h -invariant and -invariant, if ,: Á ! :i
Structure Theory for Normal Operators 229
then
º Á² ³O ²!³» ~ º Á !» ~ º Á!» ~ º O ² ³Á!» ii
::
Hence, by definition of adjoint, . ²³ O~ ² O ³ii::
Now let us relate the kernel and image of a linear transformation to those of its
adjoint.
Theorem 10.3 Let , where and are finite-dimensional innerB² = Á > ³ = >
product spaces.
1)
ker ker²³ ~ ² ³ ²³ ~ ² ³ i i im im and
and so
surjective injective
injective surjective¯
¯i
i
2)
ker ker ker ker²³ ~² ³ ²³ ~² ³ ii iand
3)
im im im im²³ ~² ³ ²³ ~² ³ ii iand
4)
²³ ~:Á;i
;Á :
Proof. For part 1),
" ² ³¯ "~
¯ º "Á=» ~ ¸¹
¯ º"Á =» ~ ¸¹
¯" ² ³ker
ii
i
im
and so . The second equation in part 1) follows by replacing ker²³ ~ ² ³ iim
by and taking complements.i
For part 2), it is clear that . For the reverse inclusion, we have ker ker²³ ² ³ i
ii" ~ ¬ º "Á"» ~ ¬ º "Á "» ~ ¬ " ~
and so . The second equation follows from the first by ker ker²³ ² ³ i
replacing with . We leave the rest of the proof for the reader.i
230 Advanced Linear Algebra
The Operator Adjoint and the Hilbert Space Adjoint
We should make some remarks about the relationship between the operator
adjoint of , as defined in Chapter 3 and the adjoint that we have just di
defined, which is sometimes called the . In the first place, Hilbert space adjoint
if , then and have different domains and ranges: ¢= ¦>di
di i i¢> ¦= ¢> ¦= and
The two maps are shown in Figure 10.1, along with the conjugate Riesz
isomorphisms and . 9¢ = ¦ = 9¢ >¦ >=i > i
V*
WWVWWx
W**
RV RW
Figure 10.1
The composite map defined by ¢> ¦=ii
~² 9 ³ k k9=c i >
is linear. Moreover, for all and , > #=i
²² ³ ³ # ~ ² # ³
~º # Á9 ² ³ »
~º # Á 9 ² ³ »
~ ´²9 ³ ² 9 ²³³µ²#³
~² ³ #
d
>
i>
=c i >
and so . Hence, the relationship between and is ~dd i
d= c i >~² 9 ³ k k9
Loosely speaking, the Riesz functions are like “change of variables” functions
from linear functionals to vectors, and we can say that does to Riesz vectors i
what does to the corresponding linear functionals. Put another way (and justd
as loosely), and are the same, up to conjugate Riesz isomorphism. i
In Chapter 3, we showed that the matr ix of the operator adjoint is the d
transpose of the matrix of the map . For Hilbert space adjoints, the situation is
slightly different (due to the conjuga te linearity of the inner product). Suppose
that and are ordered orthonormal bases for 89~² ÁÃÁ³ ~² ÁÃÁ ³ =
and , respectively. Then>
Structure Theory for Normal Operators 231
² ´ µ ³ ~º Á»~º Á »~º Á»~² ´ µ ³ ii
Á Á Á Á98 89
and so and are conjugate transposes. The conjugate transpose of a´µ ´ µiÁÁ98 89
matrix is(~² ³ Á
(~ ² ³i!
Á
and is called the of . adjoint (
Theorem 10.4 Let , where and are finite-dimensional innerB² = Á > ³ = >
product spaces.
1 The operator adjoint and the Hilbert space adjoint are related by) di
d= c i >~² 9 ³ k k9
where and are the conjugate Riesz isomorphisms on and ,99 = >=>
respectively.
2 If and are ordered for and , respectively, then)89 orthonormal bases =>
´µ ~ ² ´ µ³ii
ÁÁ98 89
In words, the matrix of the adjoint is the adjoint conjugate transpose ofi()
the matrix of .
Orthogonal Projections
In an inner product space, we can single out some special projection operators.
Definition A projection of the form is said to be . :Á: orthogonal
Equivalently, a projection is orthogonal if . ker²³ ²³ im
Some care must be taken to avoid confusion between orthogonal projections and
two projections that are orthogonal to each other, that is, for which
~~ .
We have seen that an operator is a projection operator if and only if it is
idempotent. Here is the analogous characterization of orthogonal projections.
Theorem 10.5 Let be a finite-dimensional inner product space. The following=
are equivalent for an operator on : =
1 is an orthogonal projection)
2 is idempotent and self-adjoint)
3 is idempotent and does not expand lengths, that is)
)) ) )##
for all .#=
232 Advanced Linear Algebra
Proof. Since
²³ ~:Á;i
;Á :
it follows that if and only if , that is, if and only if is ~: ~ ;i
orthogonal. Hence, 1) and 2) are equivalent.
To prove that 1) implies 3), let . Then if for and ~# ~ b ! ::Á:
!:, it follows that
)) )) )) )) ) )#~ b ! ~#
Now suppose that 3) holds. Then
im²³ l ²³ ~ =~ ²³ p ²³ ker ker ker
and we wish to show that the first sum is orthogonal. If , then $ ² ³ im
$ ~ % b & % ² ³ & ² ³ , where and . Hence, ker ker
$~ $~ %b &~ &
and so the orthogonality of and implies that %&
)) )) ) ) ) ) ))%b &~ $~& &
Hence, and so , which implies that .% ~ ²³ ²³ ²³ ~ ²³ im im ker ker
Orthogonal Resolutions of the Identity
We have seen (Theorem 2.25) that resolutions of the identity
bÄb ~
on correspond to direct sum decompositions of . If, in addition, the==
projections are orthogonal, then the direct sum is an orthogonal sum.
Definition An is a resolution of the orthogonal resolution of the identity
identity in which each projection is orthogonal. bÄb ~
The following theorem displays a correspondence between orthogonal direct
sum decompositions of and orthogonal resolutions of the identity. =
Theorem 10.6 Let be an inner product space. Orthogonal resolutions of the=
identity on correspond to orthogonal direct sum decompositions of as ==
follows:
1 If is an orthogonal resolution of the identity, then) bÄb ~
=~ ² ³ p Ä p ² ³ im im
and is orthogonal projection onto . im²³
Structure Theory for Normal Operators 233
2 Conversely, if)
=~ :p Ä p :
and if is orthogonal projection onto , then is an :b Ä b ~
orthogonal resolution of the identity.
Proof. To prove 1), if is an orthogonal resolution of the bÄb ~
identity, Theorem 2.25 implies that
=~ ² ³ l Ä l ² ³ im im
However, since the 's are pairwise orthogonal and self-adjoint, it follows that
º # Á $ »~º # Á $ »~º # Á »~
and so
=~ ² ³ p Ä p ² ³ im im
For the converse, Theorem 2.25 implies that is a resolution of bÄb ~
the identity where is projection onto along im²³
ker²³ ~ ²³ ~ ²³
£im im
Hence, is orthogonal.
Unitary Diagonalizability
We have seen (Theorem 8.10) that a linear operator on a finite- B² = ³
dimensional vector space is diagonalizable if and only if =
=~ l Ä l;;
Of course, each eigenspace has an orthonormal basis , but the union of ;E
these bases need not be an basis for . orthonormal =
Definition A linear operator is when is B² = ³ = unitarily diagonalizable (
complex and when is real if there is an )( ) orthogonally diagonalizable =
ordered orthonormal basis of for which the matrix is E~² "ÁÃÁ"³ = ´ µ E
diagonal, or equivalently, if
"~ "
for all .~ ÁÃÁ
Here is the counterpart of Theorem 8.10 for inner product spaces.
Theorem 10.7 Let be a finite-dimensional inner product space and let=
B² = ³ . The following are equivalent:
1 is unitarily orthogonally diagonalizable.)( )
2 has an orthonormal basis that consists entirely of eigenvectors of .)=
234 Advanced Linear Algebra
3 has the form)=
=~ p Ä p;;
where are the distinct eigenvalues of . ÁÃÁ
For simplicity in exposition, we will tend to use the term unitarily
diagonalizable for both cases. Since unita rily diagonalizable operators are so
well behaved, it is natural to seek a characterization of such operators.
Remarkably, there is a simple one, as we will see next.
Normal Operators
Operators that commute with their own adjonts are very special.
Definition
1 A linear operator on an inner product space is if it commutes) = normal
with its adjoint:
ii~
2 A matrix is if commutes with its adjoint .) ( ² -³ ( (Cinormal
If is normal and is an ordered orthonormal basis of , thenE =
´µ´µ ~ ´µ´ µ ~´ µ EE E EEii i
and
´µ´µ ~ ´ µ´µ ~´ µ EEE E Eii i
and so is normal if and only if is normal for some, and hence all, ´µE
orthonormal bases for . Note that this does not hold for bases that are not =
orthonormal.
Normal operators have some very special properties.
Theorem 10.8 Let be normal.B² = ³
1 The following are also normal:)
a , if reduces )O² : Á : ³:
b )i
c , if is invertible )c
d , for any polynomial )² ³ ²%³ -´%µ
2 For any , ) #Á$ =
º# Á$ » ~ º # Á $ » ii
and, in particular,
)) ) )#~ #i
Structure Theory for Normal Operators 235
and so
ker ker²³ ~ ² ³i
3 For any integer , )
ker ker²³ ~ ² ³
4 The minimal polynomial is a product of distinct prime monic ) ² % ³
polynomials.
5)
#~ # ¯ #~ #i
6 If and are submodules of with relatively prime orders, then .):; = : ;
7 If and are distinct eigenvalues of , then .) ;;
Proof. We leave part 1) for the reader. For part 2), normality implies that
º #Á $» ~ º #Á#» ~ º #Á#» ~ º #Á #» ii i i
We prove part 3) first for the operator , which is , that is, ~iself-adjoint
ii i i~² ³ ~ ~
If for , then#~
~ º # Á# » ~ º# Á# » c c c
and so . Continuing in this way gives . Now, if for c #~ #~ #~
, then
i i #~² ³#~² ³ #~
and so . Hence,#~
~º #Á#»~º #Á#»~º #Á #» i
and so .#~
For part 4 , suppose that )
²%³ ~ ²%³²%³
where is monic and prime. Then for any ,²%³ # =
²³ ´ ²³ # µ~
and since is also normal, part 3) implies that ² ³
² ³´² ³#µ ~
for all . Hence, , which implies that . Thus, the prime#= ² ³ ² ³~ ~
factors of appear only to the first power. ² % ³
236 Advanced Linear Algebra
Part 5) follows from part 2):
ker ker ker²c³ ~ ´ ²c³ µ ~ ² c³ ii
For part 6), if and , then there are polynomials ²:³ ~ ²%³ ²;³ ~ ²%³ ²%³
and for which and so²%³ ²%³²%³b²%³²%³ ~
²³ ²³ b ²³ ²³ ~
Now, annihilates and annihilates . Therefore ~ ² ³ ² ³ : ~ ² ³ ² ³ ;
i also annihilates and so ;
º:Á;»~º² b ³:Á;»~º :Á;»~º:Á ;»~¸¹ i
Part 7) follows from part 6), since and are ² ³ ~ %c ² ³ ~ %c; ;
relatively prime when . Alternatively, for and , we have ; ;£# $
º # Á$ »~º # Á$ »~º # Á $ »~º # Á $ »~ º # Á$ »i
and so implies that .£º # Á $ » ~
The Spectral Theorem for Normal Operators
Theorem 10.8 implies that when , the minimal polynomial splits-~ ² % ³d
into distinct linear factors and so The orem 8.11 implies that is diagonalizable,
that is,
=~ l Ä l ;;
Moreover, since distinct eigenspaces of a normal operator are orthogonal, we
have
=~ p Ä p ;;
and so is unitarily diagonalizable.
The converse of this is also true. If has an orthonormal basis =~ E
¸# ÁÃÁ# ¹ ´ µ ´ µ ~´ µii of eigenvectors for , then since and are EE E
diagonal, these matrices commute and therefore so do and . i
Theorem 10.9 ()The spectral theorem for normal operators: complex case
Let be a finite-dimensional complex inner product space and let .= ² = ³ B
The following are equivalent:
1 is normal.)
2 is unitarily diagonalizable, that is,)
=~ p Ä p ;;
Structure Theory for Normal Operators 237
3 has an ) orthogonal spectral resolution
~b Ä b (10.1)
where and is orthogonal for all , in which case, bÄb ~
¸Á à Á¹ is the spectrum of and
im²³ ~ ²³ ~; ;
£and ker
Proof. We have seen that 1) and 2) are equivalent. To see that 2) and 3) are
equivalent, Theorem 8.12 says that
=~ l Ä l ;;
if and only if
~b Ä b
and in this case,
im²³ ~ ²³ ~; ;
£and ker
But for if and only if;; £
im²³ ²³ ker
that is, if and only if each is orthogonal. Hence, the direct sum =~
;;lÄl is an orthogonal sum if and only if each projection is
orthogonal.
The Real Case
If , then has the form-~ ² % ³s
²%³ ~ ²%c ³Ä²%c ³ ²%³Ä ²%³
where each is an irreducible monic quadratic. Hence, the primary cyclic ² % ³
decomposition of gives =
=~ p Ä p p >p Ä p > ;;
where is cyclic with prime quadratic order . Therefore, Theorem 8.8> ² % ³
implies that there is an ordered basis for which 8
´O µ ~c
>
8>?
Theorem 10.10 A()The spectral theorem for normal operators: real case
linear operator on a finite-dimensional real inner product space is normal if
and only if
238 Advanced Linear Algebra
=~ p Ä p p >p Ä p >;;
where is the spectrum of and each is an indecomposable two-¸Á à Á¹ >
dimensional -invariant subspace w ith an ordered basis for which 8
´µ ~c
8>?
Proof. We need only show that if has such a decomposition, then is normal. =
But
´µ´µ ~ ² b ³ 0~ ´µ´µ 8888 ! !
and so is normal. It follows easily that is normal.´µ8
Special Types of Normal Operators
We now want to introduce some special types of normal operators.
Definition Let be an inner product space.=
1 is also called in the complex case and)(B² = ³ self-adjoint Hermitian
symmetric in the real case if )
i~
2 is also called in the)(B² = ³ skew self-adjoint skew-Hermitian
complex case and in the real case if skew-symmetric )
i~c
3 is in the complex case and in the real case if)B² = ³ unitary orthogonal
is invertible and
ic ~
There are also matrix versions of these definitions, obtained simply by replacing
the operator by a matrix . Moreover, the operator is self-adjoint if and only (
if any matrix that represents with respect to an ordered basis isE orthonormal
self-adjoint. Similar statements hold for the other types of operators in the
previous definition.
In some sense, square complex matrices are a generalization of complex
numbers and the adjoint (conjugate transpose) is a generalization of the complex
conjugate. In looking for a bette r analogy, we could consider just the diagonal
matrices, but this is a bit too restrictiv e. The next logical choice is the set of D
normal matrices.
Indeed, among the complex numbers, there are some special subsets: the real
numbers, the positive numbers and the numbers on the unit circle. We will soon
see that a complex matrix is self-adjoi nt if and only if its complex eigenvalues (
Structure Theory for Normal Operators 239
are real. This would suggest that the analog of the set of real numbers is the set
of self-adjoint matrices. Also, we will see that a complex matrix is unitary if and
only if its eigenvalues have norm , so numbers on the unit circle seem to
correspond to the set of unitary matrices. This leaves open the question of which
normal matrices correspond to the positive real numbers. These are the positive
definite matrices, which we will discuss later in the chapter.
Self-Adjoint Operators
Let us consider the basic properties of self-adjoint operators. The quadratic
form associated with the linear operator is the function defined 8¢ =¦-
by
8 ²#³ ~ º #Á#»
We have seen (Theorem 9.2) that in a inner product space, if and complex ~
only if but this does not hold, in general, for real inner product spaces.8~
However, it does hold for symmetric operators on a real inner product space.
Theorem 10.11 Let be a finite-dimensional inner product space and let=
BÁ² = ³ .
1 If and are self-adjoint, then so are the following:)
a )b
b , if is invertible )c
c , for any real polynomial )² ³ ²%³ ´%µs
2 A complex operator is Hermitian if and only if is real for all) 8² # ³
#= .
3 If is a complex operator or a real symmetric operator, then)
~ ¯ 8 ~
4 The characteristic polynomial of a self-adjoint operator splits over ) ² % ³
s, that is, all complex roots of are real. Hence, the minimal² % ³
polynomial of is the product of distinct monic linear factors over ² % ³
s.
Proof. For part 2), if is Hermitian, then
º #Á#» ~ º#Á #» ~ º #Á#»
and so is real. Conversely, if , then8 ²#³ ~ º #Á#» º #Á#» s
º#Á #» ~ º #Á#» ~ º#Á #» i
and so .~i
For part 3), we need only prove that implies when . But if 8~ ~ - ~ s
8~ , then
240 Advanced Linear Algebra
~º ² %b& ³ Á%b& »
~ º %Á%»bº &Á&»bº %Á&»bº &Á%»
~ º %Á&»bº &Á%»
~ º %Á&»bº%Á &»
~ º %Á&»bº %Á&»
~ º % Á& »
and so .~
For part 4), if is Hermitian ( ) and , then d -~ #~ #
#~ #~ #~ #i
and so is real. If is symmetric ( ) , we must be a bit careful, since s~- ~
a nonreal root of is an eigenvalue of . However, matrix techniques ² % ³ not
can come to the rescue here. If for any ordered orthonormal basis (~´ µEE
for , then . Now, is a real symmetric matrix, but can be= ² % ³ ~ ² % ³ ( (
thought of as a complex Hermitian matrix with real entries. As such, it
represents a Hermitian linear operator on the complex space and so, by what d
we have just shown, all (complex) roots of its characteristic polynomial are real.
But the characteristic polynomial of is th e same, whether we think of as a((
real or a complex matrix and so the result follows.
Unitary Operators and Isometries
We now turn to the basic properties of unitary operators. These are the
workhorse operators, in that a unitary operator is precisely a normal operator
that maps orthonormal bases to orthonormal bases.
Note that is unitary if and only if
º# Á $ » ~ º # Á $ »c
for all .#Á$ =
Theorem 10.12 Let be a finite-dimensional inner product space and let=
BÁ² = ³ .
1 If and are unitary/orthogonal, then so are the following:)
a , f o r ) Á ~ d ((
b )
c , if is invertible. )c
2 is unitary/orthogonal if and only it is an isometric isomorphism.)
3 is unitary/orthogonal if and only if it takes some orthonormal basis to an )
orthonormal basis, in which case it takes all orthonormal bases to
orthonormal bases.
4 If is unitary/orthogonal, then the ei genvalues of have absolute value . )
Structure Theory for Normal Operators 241
Proof. We leave the proof of part 1) to the reader. For part 2), a
unitary/orthogonal map is injective and si nce is finite-dimensional, it is=
bijective. Moreover, for a bijective linear map , we have
is an isometry for all
for all
is unitary/orthogonal¯ º #Á $» ~ º#Á$» #Á$ =
¯ º#Á $» ~ º#Á$» #Á$ =
¯~
¯~
¯i
i
ic
For part 3), suppose that is unitary/orthogonal and that is an E ~¸ "ÁÃÁ"¹
orthonormal basis for . Then =
º "Á "»~º "Á"»~ Á
and so is an orthonormal basis for . Conversely, suppose that and E E E =
are orthonormal bases for . Then =
º" Á"» ~ ~ º " Á "» Á
which implies that for all and so is º #Á $»~º#Á$» #Á$=
unitary/orthogonal.
For part 4), if is unitary and , then #~ #
º#Á#» ~ º #Á #» ~ º #Á #» ~ º#Á#»
and so , which implies that . (( (( ~~ ~
We also have the following theorem concerning unitary and orthogonal ()
matrices.
Theorem 10.13 Let be an matrix over or .( d -~ -~ ds
1 The following are equivalent:)
a is unitary/orthogonal. )(
b The columns of form an orthonormal set in . ) (-
c The rows of form an orthonormal set in . ) (-
2 If is unitary, then . If is orthogonal, then .)( ²(³ ~ ( ²(³ ~ f ((det det
Proof. The matrix is unitary if and only if , which is equivalent to (( ( ~ 0i
the rows of being orthonormal. Similarl y, is unitary if and only if ((
((~0 (i, which is equivalent to the columns of being orthonormal. As for
part 2),
(( ~ 0 ¬ ²(³ ²( ³ ~ ¬ ²(³ ²(³ ~ iidet det det det
from which the result follows.
242 Advanced Linear Algebra
Unitary/orthogonal matrices play the role of change of basis matrices when we
restrict attention to orthonormal ba ses. Let us first note that if 8~² "ÁÃÁ"³
is an ordered orthonormal basis and
#~" bÄb "
$~" bÄb "
then
º#Á$»~ bÄb ~´#µ h´$µ 88
where the right hand side is the standard inner product in and so if -# $
and only if . We can now state the analog of Theorem 2.9. ´#µ ´$µ88
Theorem 10.14 If we are given any two of the following:
1 A unitary/orthogonal matrix ,) d (
2 An ordered orthonormal basis for ,) 8-
3 An ordered orthonormal basis for ,) 9-
then the third is uniquely determined by the equation
(~489Á
Proof. Let be a basis for . If is an orthonormal basis for , then89~¸ ¹ = =
º Á » ~ ´ µ h´ µ 99
where is the th column of . Hence, is unitary if and only if ´ µ ( ~ 4 (Á98 9 8
is orthonormal. We leave the rest of the proof to the reader.
Unitary Similarity
We have seen that the change of basis formula for operators is given by
´µ ~ 7 ´µ788Zc
where is an invertible matrix. What happens when the bases are orthonormal?7
Definition
1 Two complex matrices and are also called)( () unitarily similar
unitarily equivalent ) if there exists a unitary matrix for which <
)~<(< ~<(<c i
The equivalence classes associated with unitary similarity are called
unitary similarity classes .
2 Similarly, two real matri ces and are also called )( () orthogonally similar
orthogonally equivalent ) if there exists an orthogonal matrix for which 6
) ~ 6(6 ~ 6(6c !
The equivalence classes associated with orthogonal similarity are called
orthogonal similarity classes .
Structure Theory for Normal Operators 243
The analog of Theorem 2.19 is the following.
Theorem 10.15 Let be an inner product space of dimension . Then two=
d ( ) matrices and are unitarily/orthogonally similar if and only if they
represent the same linear operator with respect to possibly differentB² = ³ ()
ordered orthonormal bases. In this case, and represent exactly the same ()
set of linear operators in with respect to ordered bases. B²= ³ orthonormal
Proof. If and represent , that is, if() ² = ³ B
( ~ ´µ )~ ´µ89and
for ordered orthonormal bases and , then 89
)~4 ( 489 98ÁÁ
and according to Theorem 10.14, is unitary/orthogonal. Hence, and 4( )89Á
are unitarily/orthogonally similar.
Now suppose that and are unitarily/orthogonally similar, say ()
)~<( <c
where is unitary/orthogonal. Suppose also that represents a linear operator<(
B 8² = ³ for some ordered orthonormal basis , that is,
(~´ µ8
Theorem 10.14 implies that there is a unique ordered orthonormal basis for 9=
for which . Hence <~489Á
)~4 ´µ4 ~´µ89 8 9 89 Á Ác
and so also represents . By symmetry, we see that and represent the)( )
same set of linear operators, under all possible ordered orthonormal bases.
We have shown (see the discussion of Schur's theorem) that any complex matrix
(( is unitarily similar to an upper triangul ar matrix, that is, that is unitarily
upper triangularizable. However, upper tria ngular matrices do not form a set of
canonical forms under unitary similarity . Indeed, the subject of canonical forms
for unitary similarity is rather complicated and we will not discuss it in this
book, but instead refer the reader to the survey article [28].
Reflections
The following defines a very sp ecial type of unitary operator.
244 Advanced Linear Algebra
Definition For a nonzero , the unique operator for which #= / #
/#~c # Á/$~$ $º # »## for all
is called a or a . reflection Householder transformation
It is easy to verify that
/%~%c #º%Á#»
º#Á#»#
Moreover, for if and only if for some and so /%~c % %£ %~ # -#
we can uniquely identify by the behavior of the reflection on . #=
If is a reflection and if we extend to an ordered orthonormal basis for , /# =# 8
then is the matrix obtained from the identity matrix by replacing the upper´/ µ#8
left entry by , c
´/ µ ~c
Æ
#8vy
x{x{
wz
Thus, a reflection is both unitary and Hermitian, that is,
/~ / ~ /##ic
#
Given two nonzero vectors of equal length, there is precisely one reflection that
interchanges these vectors.
Theorem 10.16 Let be distinct nonzero vectors of equal length. Then#Á$ =
/# $ $ ##c$ is the unique reflection sending to and to .
Proof. If , then and so)) ) )# ~ $ ²#c$³²#b$³
/² # c $ ³ ~ $ c #
/² # b $ ³ ~ # b $#c$
#c$
from which it follows that and . As to uniqueness, /² # ³ ~ $ /² $ ³ ~ ##c$ #c$
suppose is a reflection for which . Since , we have // ² # ³ ~ $ / ~ /%% % %c
/² $ ³~#% and so
/² #c$ ³~c ² #c$ ³%
which implies that . /~ /%# c $
Reflections can be used to characterize unitary operators.
Theorem 10.17 Let be a finite-dimensional inner product space. The=
following are equivalent for an operator : B² = ³
Structure Theory for Normal Operators 245
1 is unitary/orthogonal)
2 is a product of reflections.)
Proof. Since reflections are unitary /orthogonal and the product of unitary/
orthogonal operators is unitary, it follows that 2) implies 1). For the converse,
let be unitary. Let be an orthonormal basis for . Then8 ~² "ÁÃÁ"³ =
/² " ³ ~ " "c "
and so if then %~"c "
²/ ³" ~ "%
that is, is the identity on . Suppose that we have found% / º "»
reflections for which is the identity on /Á Ã Á / /Ä /%% c % %c c
º" ÁÃÁ" » c . Then
/² " ³ ~ "c "c " c
Moreover, we claim that for , since ²" c " ³ " c
º " c" Á" » ~ º²/ Ä/ ³" Á" »
~º "Á/ Ä / "»
~º "Á "»
~º "Á"»
~
c % %
% %
c
c
Hence, if , then %~ "c " c
²/ Ä/ ³" ~ / " ~ "%% %
and so is the identity on . Thus, for we% % / Ä/ º" ÁÃÁ" » ~
have and so , as desired./Ä / ~ ~ /Ä /%% %%
The Structure of Normal Operators
The following theorem includes the spectra l theorems stated above for real and
complex normal operators, along with some further refinements related to self-
adjoint and unitary/orthogonal operators.
Theorem 10.18 ()The structure theorem for normal operators
1 Let be a finite-dimensional complex inner product)( )Complex case =
space.
a The following are equivalent for : ) B² = ³
i is normal )
ii is unitarily diagonalizable )
iii has an orthogonal spectral resolution )
~b Ä b
246 Advanced Linear Algebra
b Among the normal operators, the Hermitian operators are precisely )
those for which all complex eigenvalues are real.
c Among the normal operators, the unitary operators are precisely those )
for which all complex eigenvalues have norm .
2 Let be a finite-dimensional real inner product space.)( )Real case =
a is normal if and only if )B² = ³
=~ p Ä p p >p Ä p >;;
where is the spectrum of and each is a two-¸Á à Á¹ >
dimensional indecompos able -invariant subspace with an ordered
basis for which8
´µ ~c
8>?
b Among the real normal operators, the symmetric operators are those )
for which there are no subspaces in the decomposition of part 2a . > )
Hence, the following are equivalent for : B² = ³
i is symmetric. )
ii is orthogonally diagonalizable. )
iii has the orthogonal spectral resolution )
~b Ä b
c Among the real normal operators, the orthogonal operators are )
precisely those for which the eigenvalues are equal to and the f
matrices described in part 2a have rows and columns of norm ´µ8 )( )
, that is,
´µ ~c
8>?sin cos
cos sin
for some .s
Proof. We have proved part 1a). As to part 1b), it is only necessary to look at a
diagonal matrix representing . This matrix has the eigenvalues of on its (
main diagonal and so it is Hermitian if and only if the eigenvalues of are real.
Similarly, is unitary if and only if th e eigenvalues of have absolute value (
equal to .
We have proved part 2a). Parts 2b) and 2c) follow by looking at the matrix
(~´ µ ~ (8 88 where . This matrix is symmetric if and only if is diagonal,
and is orthogonal if and only if and the matrices have(~ f ´ µ 8
orthonormal rows.
Matrix Versions
We can formulate matrix versions of the structure theorem for normal operators.
Structure Theory for Normal Operators 247
Theorem 10.19 ()The structure theorem for normal matrices
1)( )Complex case
a A complex matrix is normal if and only if it is unitarily ) (
diagonalizable, that is, if and only if there is a unitary matrix for <
which
<(< ~ ² ÁÃÁ ³i
diag
b A complex matrix is Hermitian if and only if 1a holds, where all )) (
eigenvalues are real.
c A complex matrix is unitary if and only if 1a holds, where all )) (
eigenvalues have norm .
2)( )Real case
a A real matrix is normal if and only if there is an orthogonal matrix ) (
6 for which
6(6 ~ ÁÃÁ Á ÁÃÁc c
!
diag67 >? > ?
b A real matrix is symmetric if and only if it is orthogonally ) (
diagonalizable, that is, if and only if there is an orthogonal matrix 6
for which
6(6 ~ ² ÁÃÁ ³!
diag
c A real matrix is orthogonal if and only if there is an orthogonal ) (
matrix for which6
6(6
~ ÁÃÁ Á ÁÃÁcc!
diag67 >? > ?
sin cos sin cos
cos sin cos sin
for some . sÁÃÁ
Functional Calculus
Let be a normal operator on a finite-dimensional inner product space and let =
have spectral resolution
~b Ä b
Since each is idempotent, we have for all . The pairwise ~
orthogonality of the projections implies that
~² bÄb ³ ~ bÄb
More generally, for any polynomial over , ²%³ -
² ³ ~ ² ³ bÄb² ³
Note that a polynomial of degree is uniquely determined by specifying anc
248 Advanced Linear Algebra
arbitrary set of of its values at the distinct points . This follows from Á Ã Á
the Lagrange interpolation formula
²%³ ~ ² ³%c
cvy
wz ~c
£
Therefore, we can define a unique pol ynomial by specifying the values ²%³
² ³ ~ÁÃÁ, for .
For example, for a given , if is a polynomial for which ² % ³
² ³~ Á
for , then~ ÁÃÁ
²³~
and so each projection is a polynomial function of . As another example, if
is invertible and , then ² ³ ~ c
² ³ ~ bÄb ~ c c c
as can easily be verified by direct calculation. Finally, if , then since ² ³ ~
each is self-adjoint, we have
² ³ ~ bÄb ~ i
and so is a polynomial in .i
We can extend this idea further by , for function defining any
¢¸ ÁÃÁ ¹¦-
the linear operator by ² ³
² ³~² ³ bÄb² ³
For example, we may define and so on. Notice, however, that jÁÁ c
since the spectral resolution of is a finite sum, we gain nothing (but
convenience) by using functions other th an polynomials, for we can always find
a polynomial for which for and so ²%³ ² ³ ~ ² ³ ~ ÁÃÁ
² ³~² ³ . The study of the properties of functions of an operator is
referred to as the of . functional calculus
According to the spectral theorem, if is complex and is normal, then is= ² ³
a normal operator whose eigenvalues are . Similarly, if is real and is ² ³ =
symmetric, then is symmetric, with eigenvalues . ² ³ ² ³
Structure Theory for Normal Operators 249
Commutativity
The functional calculus can be applied to the study of the commutativity
properties of operators. Here are two simple examples.
Theorem 10.20 Let be a finite-dimensional complex inner product space.=
For , we write to denote the fact that and commute. Let B Á² = ³ ©
and have spectral resolutions
~b Ä b
~b Ä b
Then
1 For any ,) B² = ³
©¯© for all
2)
©¯© Á , for all
3 If and are injective functions,)¢¸ ÁÃÁ ¹¦- ¢¸ ÁÃÁ ¹¦-
then
² ³©² ³ ¯ ©
Proof. For 1), if for all , then and the converse follows from the © ©
fact that is a polynomial in . Part 2) is similar. For part 3), clearly ©
implies . For the converse, let . Since is ² ³©² ³ ~¸ ÁÃÁ ¹ $
injective, the inverse function is well-defined and ¢ ² ³ ¦c$$
² ² ³ ³ ~ ² ³ ² ³c . Thus, is a function of . Similarly, is a function of .
It follows that implies . ² ³©² ³ ©
Theorem 10.21 Let and be normal operators on a finite-dimensional
complex inner product space . Then and commute if and only if they have =
the form
~ ² ² Á ³ ³
~ ²² Á ³³
where and are polynomials.²%³Á²%³ ²%Á&³
Proof. If and are polynomials in , then they clearly commute. ~ ² Á ³
For the converse, suppose that and let ~
~b Ä b
and
~b Ä b
be the orthogonal spectral resolutions of and .
250 Advanced Linear Algebra
Then Theorem 10.20 implies that . Hence, ~
Á
~² bÄb ³² bÄb ³
~² bÄb ³ ² bÄb ³
~
It follows that for any polynomial in two variables, ²%Á&³
² Á ³ ~ ² Á ³
Á
So if we choose with the property that are distinct, then ²%Á&³ ~ ² Á ³ Á
² Á ³ ~
ÁÁ
and we can also choose and so that for all and ²%³ ²%³ ² ³ ~ Á
² ³ ~ Á for all . Then
²² Á ³³ ~ ² ³ ~
~~ ~
89 8 9 Á ÁÁ
and similarly, . ²² Á ³³ ~
Positive Operators
One of the most important cases of the functional calculus is . ²%³~ % j
Recall that the quadratic form associated with a linear operator is
8 ²#³ ~ º #Á#»
Definition A self-adjoint linear operator is B² = ³
1 if for all )positive 8² # ³ #=
2 if for all .)positive definite 8² # ³ #£
Theorem 10.22 A self-adjoint operator on a finite-dimensional inner product
space is
1 positive if and only if all of its eigenvalues are nonnegative)
2 positive definite if and only if all of its eigenvalues are positive.)
Proof. If and , then8² # ³ #~ #
º #Á#»~ º#Á#»
Structure Theory for Normal Operators 251
and so . Conversely, if all eigenvalues of are nonnegative, then
~b Ä bÁ
and since , ~b Ä b
º # Á# »~ º # Á # »~ # ))
Á
and so is positive. Part 2) is proved similarly.
If is a positive operator, with spectral resolution
~b Ä bÁ
then we may take the of , positive square root
j jj ~b Ä b
where is the nonnegative square root of . It is clear that j
²³ ~j
and it is not hard to see that is th e only positive operator whose square is . j
In other words, every positive operato r has a unique positive square root.
Conversely, if has a positive square root, that is, if , for some positive ~
operator , then is positive. Hence, an operator is positive if and only if it
has a positive square root.
If is positive, then is self-adjoint and so j
²³ ~jj i
Conversely, if for some operator , then is positive, since it is clearly ~i
self-adjoint and
º # Á# »~º # Á# »~º # Á # » i
Thus, is positive if and only if it has the form for some operator . ~i
(A complex number is nonnegative if and only if has the form for '' ~ $ $
some complex number .) $
Theorem 10.23 Let .B² = ³
1 is positive if and only if it has a positive square root.)
2 is positive if and only if it has the form for some operator .) ~i
Here is an application of square roots.
Theorem 10.24 If and are positive operators and , then is ~
positive.
252 Advanced Linear Algebra
Proof. Since is a positive operator, it has a positive square root , which is j
a polynomial in . A similar statement holds for . Therefore, since and
commute, so do and . Hence, jj
²³ ~ ² ³ ² ³ ~jj j j
Since and are self-adjoint and commute, their product is self-adjoint jj
and so is positive.
The Polar Decomposition of an Operator
It is well known that any nonzero complex number can be written in the ' polar
form , where is a positive number and is real. We can do the same'~
for any nonzero linear operator on a finite-dimensional complex inner product
space.
Theorem 10.25 Let be a nonzero linear operator on a finite-dimensional
complex inner product space . =
1 There exist a positive operator and a unitary operator for which)
~ . Moreover, is unique and if is invertible, then is also unique.
2 Similarly, there exist a positive ope rator and a unitary operator for )
which . Moreover, is unique and if is invertible, then is also ~
unique.
Proof. Let us suppose for a moment that . Then ~
ii i i c ~² ³ ~ ~
and so
ic ~~
Also, if , then#=
#~ ² # ³
These equations give us a clue as to how to define and .
Let us define to be the unique positiv e square root of the positive operator
i. Then
)) )) # ~ º #Á #» ~ º #Á#» ~ º #Á#» ~ # i()10.2
Define on by im²³
²# ³ ~ #
for all . Equation 10.2 shows that implies that and so#= %~ & %~ & ()
this definition of on is well-defined. im²³
Moreover, is an isometry on , since 10.2 gives im ( )²³
Structure Theory for Normal Operators 253
) ) )) )) ²# ³ ~ # ~ #
Thus, if is an orthonormal basis for , then 8~¸ ÁÃÁ¹ ² ³ im
8 ~¸ ÁÃÁ ¹ ² ² ³ ³~ ² ³ is an orthonormal basis for . Finally, we im im
may extend both orthonormal bases to ort honormal bases for and then extend =
the definition of to an isometry on for which . =~
As for the uniqueness, we have seen that must satisfy and since i ~
has a unique positive square root, we dedu ce that is uniquely defined. Finally,
if is invertible, then so is since . Hence, is ker ker²³ ²³ ~c
uniquely determined by .
Part 2 can be proved by applying the previous theorem to the map , to get ) i
~² ³ ~² ³ ~ ~ii i c
where is unitary.
We leave it as an exercise to show that any unitary operator has the form
~, where is a self-adjoint operato r. This gives the following corollary.
Corollary 10.26 Let be a nonzero linear operator on()Polar decomposition
a finite-dimensional complex inner product space. Then there is a positive
operator and a self-adjoint operator for which has the polar
decomposition
~
Moreover, is unique and if is invertible, then is also unique.
Normal operators can be characterized using the polar decomposition.
Theorem 10.27 Let be a polar decomposition of a nonzero linear~
operator . Then is normal if and only if . ~
Proof. Since
i c ~ ~
and
ic c ~ ~
we see that is normal if and only if
~c
254 Advanced Linear Algebra
or equivalently,
~
Now, is a polynomial in and is a polynomial in and so this holds if
and only if . ~
Exercises
1. Let . If is surjective, find a formula for the right inverse of B ² < Á = ³
in terms of . If is injective, find a formula for a left inverse of in terms i
of . : Consider and . ii iHint
2. Let where is a complex vector space and letB² = ³ =
ii~² b³ ~² c³
and
Show that and are self-adjoint and that
~b ~c i and
What can you say about the uniqueness of these representations of and
i?
3. Prove that all of the roots of the characteristic polynomial of a skew-
Hermitian matrix are pure imaginary.
4. Give an example of a normal operato r that is neither self-adjoint nor
unitary.
5. Prove that if for all , where is complex, then is )) ) ) #~ ² # ³ # = =i
normal.
6. Let be a normal operator on a co mplex finite-dimensional inner product
space or a self-adjoint operator on a real finite-dimensional inner product=
space.
a Show that , for some polynomial . ) di~ ² ³ ²%³ ´%µ
b Show that for any , implies . In other ) B ² = ³ ~ ~ii
words, commutes with all operators that commute with .i
7. Show that a linear operator on a finite-dimensional complex inner product
space is normal if and only if whenever is an invariant subspace under=:
, so is .:
8. Let be a finite-dimensional inner product space and let be a normal =
operator on . =
a Prove that if is idempotent, then it is also self-adjoint. )
b Prove that if is nilpotent, then . ) ~
c Prove that if , then is idempotent. ) ~
9. Show that if is a normal operato r on a finite-dimensional complex inner
product space, then the algebraic multiplicity is equal to the geometric
multiplicity for all eigenvalues of .
10. Show that two orthogonal projections and are orthogonal to each other
if and only if . im im²³ ²³
Structure Theory for Normal Operators 255
11. Let be a normal operator and let be any operator on . If the =
eigenspaces of are -invariant, show that and commute.
12. Prove that if and are normal operators on a finite-dimensional complex
inner product space and if for some operator then . ~~ii
13. Prove that if two normal complex matrices are similar, then they are d
unitarily similar , that is, similar via a unitary matrix.
14. If is a unitary operator on a complex inner product space, show that there
exists a self-adjoint operator for which . ~
15. Show that a positive operator ha s a unique positive square root.
16. Prove that if has a square root, that is, if , for some positive ~
operator , then is positive.
17. Prove that if (that is, is positive) and if is a positive operator c
that commutes with both and then .
18. Using the factorization, prove the following result, known as the 89
Cholsky decomposition . An invertible linear operator is positive B² = ³
if and only if it has the form where is upper triangularizable. ~i
Moreover, can be chosen with pos itive eigenvalues, in which case the
factorization is unique.
19. Does every self-adjoint operator on a finite-dimensional real inner product
space have a square root?
20. Let be a linear operator on and let be the eigenvalues of ,d ÁÃÁ
each one written a number of times equal to its algebraic multiplicity. Show
that
((
i ² ³tr
where is the trace. Show also that equality holds if and only if is tr
normal.
21. If where is a real inner product space, show that the HilbertB² = ³ =
space adjoint satisfies . ²³~ ²³iidd
Part II—Topics
Chapter 11
Metric Vector Spaces: The Theory of
Bilinear Forms
In this chapter, we study vector spaces over arbitrary fields that have a bilinear
form defined on them.
Unless otherwise mentioned, all vector spaces are assumed to be finite-
dimensional. The symbol denotes an arbitrary field and denotes a finite --
field of size .
Symmetric, Skew-Symmetric and Alternate Forms
We begin with the basic definition.
Definition Let be a vector space over . A mapping is= - ºÁ»¢= d= ¦ -
called a if it is linear in each coordinate, that is, if bilinear form
º %b &Á'»~ º%Á'»b º&Á'»
and
º'Á %b &» ~ º'Á%»b º'Á&»
A bilinear form is
1 i f)symmetric
º%Á&» ~ º&Á%»
for all .%Á & =
2 o r i f)( )skew-symmetric antisymmetric
º%Á&» ~ cº&Á%»
for all .%Á& =
260 Advanced Linear Algebra
3 o r i f)( )alternate alternating
º%Á%» ~
for all .%=
A bilinear form that is either symmetric, skew-symmetric, or alternate is
referred to as an and a pair , where is a vector space inner product ²=ÁºÁ»³ =
and is an inner product on , is called a or ºÁ» = metric vector space inner
product space . As usual, we will refer to as a metric vector space when the =
form is understood.
4 A metric vector space with a symmetric form is called an ) = orthogonal
geometry over .-
5 A metric vector space with an alternate form is called a ) = symplectic
geometry over .-
The term , from the Greek for “intertwined,” was introduced in 1939 symplectic
by the famous mathematician Hermann Weyl in his book , The Classical Groups
as a substitute for the term . According to the dictionary, symplectic complex
means “relating to or being an intergrowth of two different minerals.” An
example is , which is marble spotted with green serpentine. ophicalcite
Example 11.1 is the four-dimensional real orthogonalMinkowski space 44
geometry with inner product defined bys
º Á »~º Á »~º Á »~
º Á » ~ c
º Á » ~ £
33
44
for
where is the standard basis for .Á Ã Á 4 s
As is traditional, when the inner product is understood, we will use the phrase
“let be a metric vector space.”=
The real inner products discussed in Chapter 9 are inner products in the present
sense and have the additional property of being —a notion that positive definite
does not even make sense if the base field is not ordered. Thus, a real inner
product space is an orthogonal geometry. On the other hand, the complex inner
products of Chapter 9, being sesquilinear, are not inner products in the present
sense. For this reason, we use the term in this chapter, rather metric vector space
than . inner product space
If is a vector subspace of a metric vector space , then inherits the metric:= :
structure from . With this structure, we refer to as a of . =: = subspace
The concepts of being symmetric, skew-symmetric and alternate are not
independent. However, their relationship depends on the characteristic of the
base field , as do many other properties of metric vector spaces. In fact, the -
Metric Vector Spaces: The Theory of Bilinear Forms 261
next theorem tells us that we do not n eed to consider skew-symmetric forms per
se, since skew-symmetry is always equivalent to either symmetry or
alternateness.
Theorem 11.1 Let be a vector space over a field .=-
1 I f , t h e n)c h a r ²-³ ~
alternate symmetric skew-symmetric ¬¯
2 I f , t h e n)c h a r ²-³ £
alternate skew-symmetric ¯
Also, the only form that is both alter nate and symmetric is the zero form:
º%Á&»~ %Á&= for all .
Proof. First note that for an alternating form over any base field,
~º%b&Á%b&»~º%Á&»bº&Á%»
and so
º%Á&» ~ cº&Á%»
which shows that the form is skew-s ymmetric. Thus, alternate always implies
skew-symmetric.
If , then and so the definitions of symmetric and skew-char²-³ ~ c ~
symmetric are equivalent, which proves 1 . If and the form is )c h a r ²-³ £
skew-symmetric, then for any , we have or , % = º%Á%» ~ cº%Á%» º%Á%» ~
which implies that . Hence, the form is alternate. Finally, if the form º%Á%» ~
is alternate and symmetric, then it is also skew-symmetric and so
º"Á#»~cº"Á#» "Á#= º"Á#»~ "Á#= for all , that is, for all .
Example 11.2 The standard inner product on , defined by =² Á³
²% ÁÃÁ% ³h²& ÁÃÁ& ³~% & bÄb% &
is symmetric, but not alternate, since
²ÁÁÃÁ³h²ÁÁÃÁ³~£
The Matrix of a Bilinear Form
If is an ordered basis for a metric vector space , then a8~² ÁÃÁ³ =
bilinear form is completely determined by the matrix of values d
4 ~ ² ³ ~ ²º Á »³8 Á
This is referred to as the (o r the matrix of ) with respect to matrix of the form =
the ordered basis . Moreover, any matrix over is the matrix of some 8 d -
bilinear form on . =
262 Advanced Linear Algebra
Note that if then %~ '
4´ % µ ~º Á%»
Å
º Á%»88vy
wz
and
´%µ 4 ~ º%Á » Ä º%Á »88!
It follows that if , then &~
´%µ 4 ´&µ ~ ~ º%Á&» º%Á » Ä º%Á » Å
888!
vy
wz
and this uniquely defines the matrix , that is, if for all 4 ´%µ (´&µ ~ º%Á&»88 8!
%Á& = ( ~ 4 , then . 8
A matrix is if it is skew-symmetr ic and has 's on the main diagonal. alternate
Thus, we can say that a form is symmetric (skew-symmetric, alternate) if and
only if the matrix is symmetric (skew-symmetric, alternate). 48
Now let us see how the matrix of a form behaves with respect to a change of
basis. Let be an ordered basis for . Recall from Chapter 2 that9~² ÁÃÁ³ =
the change of basis matrix , whose th column is , satisfies 4 ´ µ98 8Á
´#µ ~ 4 ´#µ89 8 9 Á
Hence,
º%Á&» ~ ´%µ 4 ´&µ
~ ²´%µ 4 ³4 ²4 ´&µ ³
~´ % µ² 4 4 4 ³ ´ & µ888
99 8 89 8 9
99 8 89 8 9!
!!
Á Á
!!
Á Á
and so
4~ 4 4498 9 898Á!
Á
This prompts the following definition.
Definition Two matrices are if there exists an (Á) ²-³C congruent
invertible matrix for which 7
(~7) 7!
The equivalence classes under congruence are called . congruence classes
Metric Vector Spaces: The Theory of Bilinear Forms 263
Thus, if two matrices represent the same bilinear form on , they must be =
congruent. Conversely, if represents a bilinear form on and )~4 =8
(~7) 7!
where is invertible, then there is an ordered basis for for which7= 9
7~498Á
and so
(~4 4 49889 8Á!
Á
Thus, represents the same form with respect to .(~49 9
Theorem 11.2 Let be an ordered basis for an inner product8~² ÁÃÁ³
space , with matrix=
4~ ² º Á » ³8
1 The form can be recovered from the matrix by the formula)
º%Á&» ~ ´%µ 4 ´&µ888!
2 If is also an ordered basis for , then)9~² ÁÃÁ³ =
4~ 4 4498 9 898Á!
Á
where is the change of basis matrix from to .498Á 98
3 Two matrices and represent the same bilinear form on a vector space) ()
= if and only if they are congruent, in which case they represent the same
set of bilinear forms on . =
In view of the fact that congruent matrices have the same rank, we may define
the rank of a bilinear form (or of ) to be the rank of any matrix that represents=
that form.
The Discriminant of a Form
If and are congruent matrices, then()
det det det det²(³ ~ ²7 )7³ ~ ²7³ ²)³!
and so and differ by a square factor. The of a det det²(³ ²)³ discriminant "
bilinear form is the set of determinants of all of the matrices that represent the
form. Thus, if is an ordered basis for , then 8 =
"~- ² 4 ³~¸ ² 4 ³£-¹det det88
264 Advanced Linear Algebra
Quadratic Forms
There is a close link between symmetric bilinear forms on and quadratic =
forms on . =
Definition A on a vector space is a map with thequadratic form =8 ¢ = ¦ -
following properties:
1 For all ,) -Á#=
8²#³ ~ 8²#³
2 T h e m a p)
º"Á#» ~ 8²"b#³c8²"³c8²#³8
is a symmetric bilinear form.()
Thus, every quadratic form on defines a symmetric bilinear form 8= º " Á # » 8
on . Conversely, if and if is a symmetric bilinear form on ,=² - ³ £ º Á » = char
then the function
8²%³ ~ º%Á%»
is a quadratic form . Moreover, the bilin ear form associated with is the 88
original bilinear form:
º"Á#» ~ 8²"b#³c8²"³c8²#³
~ º"b#Á"b#»c º"Á"»c º#Á#»
~ º"Á#»b º#Á"» ~ º"Á#»
8
Thus, the maps and are inverses and so there is a one-to-one ºÁ» ¦ 8 8 ¦ ºÁ» 8
correspondence between symmetric bilinear forms on and quadratic forms on =
=. Put another way, knowing the quadratic form is equivalent to knowing the
corresponding bilinear form.
Again assuming that , if is an ordered basis for an char²-³ £ ~ ²# ÁÃÁ# ³ 8
orthogonal geometry and if the matrix of the symmetric form on is ==
4~ ² ³ % ~% #8 Á , then for , '
8 ² % ³~ º % Á% »~ ´ % µ 4 ´ % µ ~ %%
888!
ÁÁ
and so is a homogeneous polynomial of degree 2 in the coordinates .8²%³ %
(The term “form” means —hence the term quadratic homogeneous polynomial
form .)
Metric Vector Spaces: The Theory of Bilinear Forms 265
Orthogonality
As we will see, not all metric vector spaces behave as nicely as real inner
product spaces and this necessitates the introduction of a new set of terminology
to cover various types of behavior. (The ba se field is the culprit, of course.)-
The most striking differences stem from the possibility that for a º%Á%» ~
nonzero vector . %=
The following terminology should be familiar.
Definition Let be a metric vector space. A vector is to a vector=% orthogonal
&% & º % Á & » ~ % = :, written , if . A vector is to a subset of orthogonal
=% : º % Á » ~ : : =, written , if for all . A subset of is to a orthogonal
subset of , written , if for all and . The;= : ;º Á ! » ~ : ! ;
orthogonal complement of a subset of is the subspace?? =
?~ ¸ # = # ? ¹
Note that regardless of whether the fo rm is symmetric or alternate and hence (
skew-symmetric , orthogonality is a symmetric relation, that is, implies ) %&
&% . Indeed, this is precisely why we restrict attention to these two types of
bilinear forms.
There are two types of degenerate behaviors that a vector may possess: It may
be orthogonal to itself or, worse yet, it may be orthogonal to vector in . every =
With respect to the former, we have the following terminology.
Definition Let be a metric vector space.=
1 A nonzero is or if ; otherwise it is)( ) % = º%Á%» ~ isotropic null
nonisotropic .
2 is if it contains at least one isotropic vector. Otherwise, is)== isotropic
nonisotropic anisotropic or .()
3 is that is, symplectic if all vectors in are)( )== totally isotropic
isotropic.
Note that if is an isotropic vector, then so is for all . This can be # # -
expressed by saying that the set of isotropic vectors in is a in . (A 0= = cone
cone in is a nonempty subset that is closed under scalar multiplication.)=
With respect to the more severe forms of degeneracy, we have the following
terminology.
Definition Let be a metric vector space.=
266 Advanced Linear Algebra
1 A vector is if . The set of all degenerate) #= #= = degenerate
vectors is called the of and denoted by . Thus, radical =² = ³ rad
rad²= ³ ~ =
2 is , or , if .)r a d= ²=³ ~ ¸¹ nonsingular nondegenerate
3 is , or , if .)r a d= ²=³ £ ¸¹ singular degenerate
4 is , or , if .)r a d=² = ³ ~ = totally singular totally degenerate
Some of the above terminology is not entirely standard, so care should be
exercised in reading the literature.
Theorem 11.3 A metric vector space is nonsingular if and only if all =
representing matrices are nonsingular. 48
A note of caution is in order. If is a subspace of a metric vector space , then :=
rad rad²:³ : : ²:³ denotes the set of vectors in that are degenerate in , that is, is
the radical of , as a metric vector space in its own right. However, denotes ::
the set of all vectors in that are orthogonal to . Thus, =:
rad²:³ ~ : q:
Note also that
rad rad² : ³ ~ : q : : q :~ ² : ³
and so if is singular, then so is . ::
Example 11.3 Recall that is the set of all ordered -tuples whose =² Á³
components come from the finite field . (See Example 11.2.) It is easy to see -
that the subspace
: ~ ¸ÁÁÁ¹
of has the property that . Note also that is nonsingular= ²Á³ : ~ : = ²Á³
and yet the subspace is singular. :totally
The following result explains why we re strict attention to symmetric or alternate
forms (which includes skew-symmetric forms).
Theorem 11.4 Let be a vector space with a bilinear form. The following are=
equivalent:
1 Orthogonality is a symmetric relation, that is,)
%& ¬ &%
2 The form on is symmetric or alternate, that is, is a metric vector) ==
space.
Metric Vector Spaces: The Theory of Bilinear Forms 267
Proof. It is clear that orthogonality is symmetric if the form is symmetric or
alternate, since in the latter case, the form is also skew-symmetric.
For the converse, assume that orthogonality is symmetric. For convenience, let
% & º%Á&» ~ º&Á%» % = º%Á#» ~ º#Á%» mean that and let mean that for all
#= %= %= = . If for all , then is orthogonal and we are done. So let us
examine vectors with the property that . %% = \
We wish to show that
%= ¬ % ² %&¬%& ³\ is isotropic and (11.1)
Note that if the second conclusion hol ds, then since , it follows that is %% %
isotropic. So suppose that . Since , there is a for which %& %= '= \
º%Á'»£º'Á%» %& and so if and only if
º%Á&»²º%Á'»cº'Á%»³~
Now,
º%Á&»²º%Á'»cº'Á%»³ ~ º%Á&»º%Á'»cº%Á&»º'Á%»
~ º&Á%»º%Á'»cº%Á&»º'Á%»
~ º%Áº&Á%»' c&º'Á%»»
But reversing the coordinates in the last expression gives
ºº&Á%»'c&º'Á%»Á%»~º&Á%»º'Á%»cº&Á%»º'Á%»~
and so the symmetry of orthogonality implies that the last expression is and so
we have proven (11.1).
Let us assume that is not orthogonal and show that all vectors in are ==
isotropic, whence is symplectic. Si nce is not orthogonal, there exist ==
" Á#= "# "= #= " # \\ \ for which and so and . Hence, the vectors and
are isotropic and for all , &=
&" ¬ &"
&# ¬ &#
Since all vectors for which are isotropic, let . Then and $ $= $= $" \
$# $" $# and so and . Now write
$~² $c" ³b"
where , since is isotropic. Since the sum of two orthogonal$c"" "
isotropic vectors is isotropic, it follows th at is isotropic if is isotropic.$$ c "
But
º$b"Á#»~º"Á#»£º#Á"»~º#Á$b"»
268 Advanced Linear Algebra
and so , which implies that is isotropic. Thus, is also²$b"³= $b" $ \
isotropic and so all vectors in are isotropic. =
Orthogonal and Symplectic Geometries
If a metric vector space is both orthogonal and symplectic, then the form is both
symmetric and skew-symmetric and so
º"Á#» ~ º#Á"» ~ cº"Á#»
Therefore, when , is orthogonal and symplectic if and only if char²-³ £ = =
is totally degenerate.
However, if , then there are orthogonal symplectic geometries that char²-³ ~
are not totally degenerate. For example, let be a two- =~ ² " Á # ³ span
dimensional vector space and define a form on whose matrix is =
4~
>?
Since is both symmetric and alternate, so is the form.4
Linear Functionals
The Riesz representation theorem says that every linear functional on a finite-
dimensional real or complex inner product space is represented by a Riesz =
vector , in the sense that9 =
²#³ ~ º#Á9 »
for all . A similar result holds for metric vector spaces.#= nonsingular
Let be a metric vector space over . Let and define the inner product=- % =
map byºhÁ%»¢= ¦-
ºhÁ%»#~º#Á%»
This is easily seen to be a linear func tional and so we can define a linear map
¢= ¦=i by
%~ºhÁ% »
The bilinearity of the form ensures th at is linear and the kernel of is
ker² ³ ~ ¸% = º= Á%» ~ ¸¹¹ ~ = ~ ²= ³rad
Hence, is injective (and therefore an isomorphism) if and only if is =
nonsingular.
Theorem 11.5 The Riesz representation theorem () Let be a finite-=
dimensional nonsingular metric vector space. The map defined by ¢= ¦=i
Metric Vector Spaces: The Theory of Bilinear Forms 269
%~ºhÁ% »
is an isomorphism from to . It follo ws that for each there exists a == =ii
unique vector for which %=
#~º#Á%»
for all .#=
The requirement that be nonsingular is necessary. As a simple example, if ==
is totally singular, then no nonze ro linear functional could possibly be
represented by an inner product.
The Riesz representation theorem applies to nonsingular metric vector spaces.
However, we can also achieve something useful for subspaces of a singular :
nonsingular metric vector space. The reason is that any linear functional :i
can be extended to a linear functional on , where it has a Riesz vector, that =
is,
#~º # Á9»~ºhÁ9» #
Hence, also has this form, where its “Riesz vector” is an element of , but is=
not necessarily in . :
Theorem 11.6 The Riesz representation theorem for subspaces () Let be a:
subspace of a metric vector space . If either or is nonsingular, the linear == :
map defined by¢= ¦:i
%~ºhÁ% » O :
is surjective and has kernel . Hence, for any linear functional , there : :i
is a not necessarily unique vector for which for all .() %= ~º Á% » :
Moreover, if is nonsingular, then can be taken from , in which case it is :% :
unique.
Orthogonal Complements and Orthogonal Direct Sums
Definition A metric vector space is the of the = orthogonal direct sum
subspaces and , written :;
=~ :p ;
if and .=~ :l ; : ;
If is a subspace of a real inner product space, the projection theorem says that:
the orthogonal complement of is a true vector space complement of , :: :
that is,
=~ :p :
270 Advanced Linear Algebra
However, in general metric vector spaces, an orthogonal complement may not
be a vector space complement. In fact, Example 11.3 shows that in some cases
:~ : # º # »~ = . In other cases, for example, if is degenerate, then .
However, as we will see, the orthogonal complement of is a vector space :
complement if and only if either the sum is correct, , or the =~ :b :
intersection is correct, . Note that the latter is equivalent to the : q: ~ ¸¹
nonsingularity of . :
Many nice properties of orthogonality in real inner product spaces do carry over
to metric vector spaces. Moreover, the next result shows that thenonsingular
restriction to nonsingular spaces is not that severe.
Theorem 11.7 Let be a metric vector space. Then=
=~ ² = ³ p : rad
where is nonsingular and is totally singular.:² = ³ rad
Proof. If is any vector space complement of , then and so:² = ³ ² = ³ : rad rad
=~ ² = ³ p : rad
Also, is nonsingular since .:² : ³ ² = ³ rad rad
Here are some properties of orthogonality in nonsingular metric vector spaces.
In particular, if either or is nonsingular, then the orthogonal complement of =:
: always has the expected dimension,
dim dim dim²: ³ ~ ²= ³c ²:³
even if is not well behaved with respect to its intersection with .::
Theorem 11.8 Let be a subspace of a finite-dimensional metric vector space:
=.
1)If either or is nonsingular, then=:
dim dim dim²:³b ²: ³ ~ ²= ³
Hence, the following are equivalent:
a )=~ :b :
b is nonsingular, that is, ): : q: ~ ¸¹
c . )=~ :p :
2 If is nonsingular, then)=
a ):~ :
b )r a d r a d²:³ ~ ²: ³
c is nonsingular if and only if is nonsingular. )::
Proof. For part 1), the map of Theorem 11.6 is surjective and has ¢= ¦:i
kernel . Thus, the rank-plus-nullity theorem implies that:
Metric Vector Spaces: The Theory of Bilinear Forms 271
dim dim dim²: ³b ²: ³ ~ ²= ³i
However, and so part 1) follows. For part 2), since dim dim²: ³ ~ ²:³i
rad rad² : ³ ~ : q : : q :~ ² : ³
the nonsingularity of implies the nonsi ngularity of . Then part 1) implies ::
that
dim dim dim²:³b ²: ³ ~ ²= ³
and
dim dim dim²: ³b ²: ³ ~ ²= ³
Hence, and .:~ : ² : ³ ~² : ³ rad rad
The previous theorem cannot in genera l be strengthened. Consider the two-
dimensional metric vector space =~ span²"Á#³ where
º"Á"» ~ Áº"Á#» ~ Áº#Á#» ~
If , then . Now, is nonsingular but is singular:~ ² " ³ : ~ ² # ³ : : span span
and so 2c) does not hold. Also, and and so 2b) rad rad²:³ ~ ¸¹ ²: ³ ~ :
fails. Finally, and so 2a) fails. :~ = £ :
Isometries
We now turn to a discussion of struc ture-preserving maps on metric vector
spaces.
Definition Let and be metric vector spaces. We use the same notation => º Á »
for the bilinear form on each spac e. A bijective linear map is called ¢= ¦>
an ifisometry
º" Á# » ~ º " Á # »
for all vectors and in . If an isometry exists from to , we say that "# = = > =
and are and write . It is evi dent that the set of all >= > isometric
isometries from to forms a group under composition. ==
If is a nonsingular orthogonal geomet ry, an isometry of is called an ==
orthogonal transformation . The set of all orthogonal transformationsE²= ³
on is a group under composition, known as the of .== orthogonal group
If is a nonsingular symplectic geom etry, an isometry of is called a ==
symplectic transformation . The set of all symplectic transformations on Sp²= ³
== is a group under composition, known as the of . symplectic group
272 Advanced Linear Algebra
Note that, in contrast to the case of real inner product spaces, we must include
the requirement that be bijective sin ce this does not follow automatically if =
is singular. Here are a few of the basic properties of isometries.
Theorem 11.9 Let be a linear transformation between finite-B² = Á > ³
dimensional metric vector spaces and . =>
1 Let be a basis for . Then is an isometry if and only if )8 ~¸ #ÁÃÁ#¹ =
is bijective and
º# Á#» ~ º # Á #»
for all .Á
2 If is orthogonal and , then is an isometry if and only if it is)c h a r=² - ³ £
bijective and
º# Á# » ~ º # Á # »
for all .#=
3 Suppose that is an isometry and) ¢= >
=~ :p : >~ ;p ;and
If , then .:~; ² : ³~;
Proof. We prove part 3 only. To see that , if and , ) ²: ³ ~ ; ' : ! ;
then since , we can write for some and so ;~ : !~ :
º 'Á! »~º 'Á »~º 'Á »~
whence . But since the dimensions are equal, it follows that²: ³ ;
²: ³ ~ ;.
Hyperbolic Spaces
A special type of two-dimensional metric vector space plays an important role in
the structure theory of metric vector spaces.
Definition Let be a metric vector space. A is a pair of= hyperbolic pair
vectors for which"Á# =
º"Á"»~º#Á#»~Á º"Á#»~
Note that if is orthogonal and if is symplectic. In º#Á"»~ = º#Á"»~c =
either case, the subspace is called a and any /~ ² " Á# ³ span hyperbolic plane
space of the form
>~/ pÄp/
where each is a hyperbolic plane, is called a . If is /² " Á # ³ hyperbolic space
a hyperbolic pair for , then we refer to the basis /
²" Á# ÁÃÁ" Á# ³
Metric Vector Spaces: The Theory of Bilinear Forms 273
for as a . In the symplectic case, the usual term is> hyperbolic basis (
symplectic basis .)
Note that any hyperbolic space is nonsingular. >
In the orthogonal case, hyperbolic planes can be characterized by their degree of
isotropy, so to speak. In the symplectic case, all spaces are totally isotropic by (
definition. Indeed, we leave it as an exercise to prove that a two-dimensional )
nonsingular orthogonal geometry is a hyperbolic plane if and only if ==
contains exactly two one-dimensional totally isotropic equivalently, totally (
degenerate subspaces. Put another way, the cone of isotropic vectors is the )
union of two one-dimensional subspaces of . =
Nonsingular Completions of a Subspace
Let be a subspace of a nonsingular metric vector space . If is singular, it<= <
is of interest to find a nonsingular subspace of containing . minimal =<
Definition Let be a nonsingular metric vector space and let be a subspace=<
of . A subspace of for which is called an of . A=: = < : < extension
nonsingular completion of is an extension of that is minimal in the family <<
of all nonsingular extensions of . <
Theorem 11.10 Let be a nonsingular finite-dimensional metric vector space=
over . We assume that when is orthogonal.-² - ³ £ = char
1 Let be a subspace of . If is isotropic and the orthogonal direct sum ):= #
span²#³p:
exists, then there is a hyperbolic plane for which /~ ² # Á' ³ span
/p:
exists. In particular, if is isotropic, then there is a hyperbolic plane #
containing . #
2 Let be a subspace of and let)<=
< ~ ²# ÁÃÁ# ³p> span
where is nonsingular and are linearly independent in>¸ # Á Ã Á # ¹
rad²<³ ~ / pÄp/. Then there is a hyperbolic space with >
hyperbolic basis for which ²# Á'ÁÃÁ# Á' ³
<~ p>>
is a nonsingular proper extension of . If is a basis for<¸ # Á Ã Á # ¹
rad²<³, then
dim dim dim²<³ ~ ²<³b ² ²<³³ rad
274 Advanced Linear Algebra
and we refer to as a of . If is nonsingular, we << < hyperbolic extension
say that is a hyperbolic extension of itself.<
Proof. For part 1 , the nonsingularity of implies that . Hence, ) =: ~ :
#¤:~: %: º # Á% »£ = and so there is an for which . If is
symplectic, then all vectors are isotropic and so we can take . If ' ~ ²°º#Á%»³%
=' ~ # b % ² # Á ' ³ is orthogonal, let . The conditions defining as a hyperbolic
pair are since is isotropic ()#
~ º#Á'» ~ º#Á#b %» ~ º#Á%»
and
~º'Á'»~º#b %Á#b %»~ º#Á%»b º%Á%»~b º%Á%»
Since , the first of these equations can be solved for and sinceº#Á%» £
char²-³ £ , the second equation can then be solved for . Thus, in either case,
there is a vector for which is hyperbolic. Hence, ': /~ ² # Á' ³:span
: : / / / q/ ~ ¸¹ and since is nonsingular, that is, , we have
/q:~¸ ¹ /p: and so exists.
Part 2) is proved by induction on . Note first that all of the vectors are #
isotropic. If , then exists and so part 1 implies that there is ~ ² #³p> span )
a hyperbolic plane for which exists. /~ ² #Á' ³ /p> span
Assume that the result is true for independent sets of size less than . Since
span span²# ³p ²# ÁÃÁ# ³p>
exists, part 1) implies that ther e exists a hyperbolic plane for /~ ² # Á ' ³ span
which
/ p ²# ÁÃÁ# ³p> span
exists. Since are in the radical of , the inductive #Á Ã Á # ² #Á Ã Á #³p> span
hypothesis implies that there is a hyperbolic space with /p Ä p /
hyperbolic basis for which the orthogonal direct sum ²# Á'ÁÃÁ# Á' ³
/p Ä p /p >
exists. Hence, also exists. /p Ä p /p >
We can now prove that the hyperbolic exte nsions of are precisely the minimal <
nonsingular extensions of . <
Theorem 11.11 Let be a subspace of a()Nonsingular extension theorem <
nonsingular finite-dimensional metric vector space . The following are =
equivalent:
1 is a hyperbolic extension of );~ p> <>
2 is a minimal nons ingular extension of );<
Metric Vector Spaces: The Theory of Bilinear Forms 275
3 is a nonsingular extension of and);<
dim dim dim²;³ ~ ²<³b ² ²<³³ rad
Thus, any two nonsingular comp letions of are isometric. <
Proof. If where is nonsingular, then we may apply Theorem<?= ?
11.10 to as a subspace of , to obtain a hyperbolic extension of <? p > < A
for which
< p>?A
Thus, every nonsingular extension of contains a hyperbolic extension of .<<
Moreover, all hyperbolic extensions of have the same dimension: <
dim dim dim²p > ³ ~ ² < ³ b ²² < ³ ³> rad
and so no hyperbolic extension of is properly contained in another hyperbolic<
extension of . This proves that 1)–3) ar e equivalent. The final statement <
follows from the fact that hyperbolic spaces of the same dimension are
isometric.
Extending Isometries to Nonsingular Completions
Let and be isometric nonsingular metric vector spaces and let==Z
<~ ² < ³p> = rad be a subspace of , with nonsingular completion
<~ p>> .
If is an isometry, then it is a simple matter to extend to an ¢< ¦ <
isometry from onto a nonsingular completion of . To see this, let <<
²" Á'ÁÃÁ" Á' ³ ²" ÁÃÁ" ³ be a hyperbolic basis for . Since is a basis for >
rad rad²<³ ² " ÁÃÁ " ³ ² <³, it follows that is a basis for .
Hence, we can hyperbolically extend to get <~ ²> ³p > rad
> <~ p >Z
where has hyperbolic basis . To extend , simply set> Z ² " Á% ÁÃÁ " Á% ³
' ~% ~ÁÃÁ for all .
Theorem 11.12 Let and be isometric nonsingular metric vector spaces==Z
and let be a subspace of , with nonsingular completion . Any isometry <= <
¢< ¦ < < can be extended to an isom etry from onto a nonsingular
completion of . <
The Witt Theorems: A Preview
There are two important theorems that are quite easy to prove in the case of real
inner product spaces, but require more work in the case of metric vector spaces
in general. Let and be isometric nonsingular metric vector spaces over a ==Z
field . We assume that if is orthogonal.-² - ³ £ = char
276 Advanced Linear Algebra
The says that if is a subspace of , then any isometryWitt extension theorem :=
¢:¦ :=Z
can be extended to an isometry from to . The ==ZWitt cancellation theorem
says that if
=~ :p : = ~ ;p ;Z and
then
:;¬: ;
We will prove these theorems in both the orthogonal and symplectic cases a bit
later in the chapter. For now, we simply want to show that it is easy to prove
one Witt theorem using the other.
Suppose that the Witt extension theorem holds and assume that
=~ :p : = ~ ;p ;Z and
and . Then any isometry can be extended to an isometry from:; ¢:¦;
== ² : ³ ~ ; : ; to . According to Theorem 11.9, we have and so .Z
Hence, the Witt cancellation theorem holds.
Conversely, suppose that the Witt cancellation theorem holds and let
¢:¦ :=Z be an isometry. Since can be extended to a nonsingular
completion of , we may assume that is nonsingular. Then ::
=~ :p :
Since is an isometry, is also nonsingular and we can write :
=~: p ²: ³Z
Since , Witt's cancellation theorem implies that . If: : : ²: ³
¢: ¦² :³ ¢= ¦= Z is an isometry, then the map defined by
²"b#³ ~ "b #
for and is an isometry that extends . Hence Witt's extension": #:
theorem holds.
The Classification Problem for Metric Vector Spaces
The for a class of metric vector spaces such as the classification problem (
orthogonal or symplectic spaces is the problem of determining when two metric )
vector spaces in the class are isometric. The classification problem is considered
“solved,” at least in a theoretical se nse, by finding a set of canonical forms or a
complete set of invariants for matrices under congruence.
Metric Vector Spaces: The Theory of Bilinear Forms 277
To see why, suppose that is an isometry and is an 8¢= ¦> ~²# ÁÃÁ# ³
ordered basis for . Then is an ordered basis for and =~ ² # Á Ã Á # ³ >9
4 ² =³~² º #Á#» ³~² º #Á #» ³~4² >³89
Thus, the congruence class of matrices representing is identical to the =
congruence class of matrices representing . >
Conversely, suppose that and are metric vector spaces with the same =>
congruence class of representing matrices. Then if is an 8~² #ÁÃÁ#³
ordered basis for , there is an ordered basis for for which = ~²$ ÁÃÁ$ ³ > 9
²º# Á# »³ ~ 4 ²= ³ ~ 4 ²>³ ~ ²º$ Á$ »³ 89
Hence, the map defined by is an isometry from to . ¢= ¦> # ~$ = >
We have shown that two metric vector spaces are isometric if and only if they
have the same congruence class of representing matrices. Thus, we can
determine whether any two metric vector spaces are isometric by representing
each space with a matrix and determining whether these matrices are congruent,
using a set of canonical forms or a set of complete invariants.
Symplectic Geometry
We now turn to a study of the structure of orthogonal and symplectic geometries
and their isometries. Since the study of the structure and the structure itself of ()
symplectic geometries is simpler than that of orthogonal geometries, we begin
with the symplectic case. The reader who is interested only in the orthogonal
case may omit this section.
Throughout this section, let be a nonsingular symplectic geometry.=
The Classification of Symplectic Geometries
Among the simplest types of metric vector spaces are those that possess an
orthogonal basis. However, it is easy to see that a symplectic geometry has an =
orthogonal basis if and only if it is tota lly degenerate and so no “interesting”
symplectic geometries have orthogonal bases.
Thus, in searching for an orthogonal decomposition of , we turn to two- =
dimensional subspaces and this puts us in mind of hyperbolic spaces. Let be <
the family of all hyperbolic subspaces of , which is nonempty since the zero =
subspace is singular and so has a nonzero hyperbolic extension. Since is ¸¹ =
finite-dimensional, has a maximal member . Since is nonsingular, if <> >
>£=, then
=~ p>>
where . But then if is nonzero, there is a hyperbolic extension >>£ ¸¹ #
278 Advanced Linear Algebra
/p #>> > of containing , which contra dicts the maximality of . Hence,
=~>.
This proves the following structure th eorem for symplectic geometries.
Theorem 11.13
1 A symplectic geometry has an orthogonal basis if and only if it is totally)
degenerate.
2 Any nonsingular symplectic geometry is a hyperbolic space, that is, ) =
=~ /p /p Ä p /
where each is a hyperbolic plane. Thus, there is a hyperbolic basis for /
=, that is, a basis for which the matrix of the form is 8
@~
c
c
Æ
c vy
x{x{x{x{x{x{x{x{
wz
In particular, the dimension of is even. =
3 Any symplectic geometry has the form) =
=~ ² = ³ p rad>
where is a hyperbolic space and is a totally degenerate space.> rad²= ³
The rank of the form is and is uniquely determined up to dim²³ =>
isometry by its rank and its dimension. Put another way, up to isometry,
there is precisely one symplectic ge ometry of each rank and dimension.
Symplectic forms are represented by alternate matrices, that is, skew-symmetric
matrices with zero diagonal. Moreover, according to Theorem 11.13, each
d alternate matrix is congruent to a matrix of the form
?~@
Ác
c>?
block
Since the rank of is , no two such matrices are congruent. ? Ác
Theorem 11.14 The set of matrices of the form is a set of d ? Ác
canonical forms for alternate matrices under congruence.
The previous theorems solve the classification problem for symplectic
geometries by stating that the rank a nd dimension of form a complete set of =
Metric Vector Spaces: The Theory of Bilinear Forms 279
invariants under congruence and that the set of all matrices of the form ?Ác
is a set of canonical forms.
Witt's Extension and Cancellation Theorems
We now prove the Witt theorems for symplectic geometries.
Theorem 11.15 Witt's extension theorem () Let and be isometric==Z
nonsingular symplectic geometries over a field . Then any isometry -
¢:¦ :=Z
on a subspace of can be extended to an isometry from to . := ==Z
Proof. According to Theorem 11.12, we can extend to a nonsingular
completion of , so we may simply assume that and are nonsingular. :: :
Hence,
=~ :p :
and
=~: p ²: ³Z
To complete the extension of to , we need only choose a hyperbolic basis =
² Á ÁÃÁ Á ³
for and a hyperbolic basis:
² Á ÁÃÁ Á ³ ZZ ZZ
for and define the extension by setting and .²: ³ ~ ~ Z Z
As a corollary to Witt's extension theore m, we have Witt's cancellation theorem.
Theorem 11.16 Witt's cancellation theorem () Let and be isometric==Z
nonsingular symplectic geometries over a field . If -
=~ :p : = ~ ;p ;Z and
then
:;¬: ;
The Structure of the Symplectic Group: Symplectic Transvections
Let us examine the nature of symp lectic transformations (isometries) on a
nonsingular symplectic geometry . Recall that for a real vector space, an =
isometric isomorphism, which corresponds to an isometry in the present context,
is the same as an orthogonal map and orthogonal maps are products of
reflections (Theorem 10.17). Recall also th at a reflection is defined as an /#
operator for which
280 Advanced Linear Algebra
/#~c # Á/$~$ $º # »## for all
and that
/%~%c #º%Á#»
º#Á#»#
In the present context, we do not dare divide by , since all vectors are º#Á#»
isotropic. So here is the next-best thing.
Definition Let be a nonsingular symplectic geometry over . Let be=- # =
nonzero and let . The map defined by - ¢=¦= #Á
#Á²%³ ~ %bº%Á#»#
is called the determined by and . symplectic transvection #
Note that if , then and if , th en is the identity precisely ~ ~ £ #Á #Á
on the subspace of codimension . In the case of a reflection, is the span²#³ /#
identity precisely on and span²#³
=~ ² # ³ p ² # ³ span span
However, for a symplectic transvection, is the identity precisely on#Á
span span span²#³ £ ²#³ ²#³ (for ) but . Here are the basic properties of
symplectic transvections.
Theorem 11.17 Let be a symplectic transvection on . Then#Á =
1 is a symplectic transformation isometry .)( )#Á
2 if and only if .)#Á~ ~
3 If , then . For , if and only if .)%# ² % ³~% £%# ² % ³~% #Á #Á
4 .) #Á #Á #Áb~
5 .)#Ác#Ác~
6 For any symplectic transformation ,)
#Á #Ác~
7 F o r ,)-i
#Á #Á~
Note that if is a subspace of and if is a symplectic transvection on , <= < "Á
then, by definition, . However, the formula "<
"Á²%³ ~ %bº%Á"»"
also defines a symplectic transvection on , where ranges over . Moreover, =% =
for any , we have and so is the identity on .'< '~' <"Á "Á
Metric Vector Spaces: The Theory of Bilinear Forms 281
We now wish to prove that any symp lectic transformation on a nonsingular
symplectic geometry is the product of symplectic transvections. The proof is =
not difficult, but it is a bit lengthy, so we break it up into parts. Our first goal is
to show that we can get from any hype rbolic pair to any other hyperbolic pair
using a product of symplectic transvections.
Let us say that two and are if there is a hyperbolic pairs ²%Á&³ ²$Á'³ connected
product of symplectic transvections that carries to and to and write %$ &'
¢²%Á&³ª²$Á'³
or . It is clear that connectedness is an equivalence relation on²%Á&³ © ²$Á'³
the set of hyperbolic pairs.
Theorem 11.18 In a nonsingular symplectic geometry , every pair of =
hyperbolic pairs are connected.
Proof. Note first that if , then and so º Á!» £ £ !
!cÁ ! ! ! !~b º Ác » ² c ³ ~c º Á » ² c ³
Taking gives . Therefore, ~ °º Á » ~ ! ! ² Á " ³ c Á! if is hyperbolic, then we
can always find a vector for which %
² Á"³ © ²!Á%³
namely, and are hyperbolic, then%~ ² Á" ³ ² ! Á" ³ c!Á". Also, if both
² Á"³ © ²!Á"³
since and so .º c!Á"» ~ " ~ " c!Á
Actually, these statements are still true if º Á!» ~ . For in this case, there is a
nonzero vector for which and . This follows from the fact &º Á & » £ º ! Á & » £
that there is an for which and and so the Riesz vector = £ £ 9i !
is such a vector. Therefore, if is hyperbolic, then ² Á"³
² Á" ³©² & Á " ³©² ! Á " ³ c&Á &c!Á c&ÁZ
and if both ² Á"³ ²!Á"³ and are hyperbolic, then
² Á"³©²&Á"³©²!Á"³
Hence, transitivity gives the same result as in the case . º Á!» £
Finally, if ² "Á "³ ² #Á #³ & and are hyperbolic, then there is a for which
²" Á" ³ © ²# Á&³ © ²# Á# ³
and so transitivity shows that ²" Á" ³ © ²# Á# ³ .
282 Advanced Linear Algebra
We can now show that the symplectic transvections generate the symplectic
group.
Theorem 11.19 Every symplectic transformation on a nonsingular symplectic
geometry is the product of symplectic transvections. =
Proof. Let be a symplectic transformation on . We proceed by induction on =
~ ²= ³ ~ = ~ / ~ ²"Á'³ dim . If , then is a hyperbolic plane and span
Theorem 11.18 implies that there is a product of symplectic transvections on
= for which
¢²"Á'³ ª ² "Á '³
This proves the result if . Assume th at the result holds for all dimensions ~
less than and let . ² = ³ ~ dim
Now,
=~ /pA
where and is a symplectic geometry of dimension less than/~ ² " Á' ³ span A
that of . As before, there is a product of symplectic transvections on for==
which
¢²"Á'³ ª ² "Á '³
and so
O~O//
Note that and so Theorem 11.9 implies that . c c /~/ ² / ³~/
Since , the inductive hypothesis applied to the symplectic dim dim²/ ³ ²/³
transformation on implies that there is a product of symplectic c /
transvections on for which . As remarked earlier, is also a /~c
product of symplectic transvections on that is the identity on and so =/
O~ ~ ///and on
Thus, on both and on and so is a product of symplectic ~/ / ~
transvections on . =
The Structure of Orthogonal Geometries: Orthogonal Bases
We have seen that no interesting that is, not totally degenerate symplectic ()
geometries have orthogonal bases. By c ontrast, almost all interesting orthogonal
geometries have orthogonal bases. =
To understand why, it is convenient to group the orthogonal geometries into two
classes: those that are also symplectic and those that are not. The reason is that
all orthogonal geometries have orthogonal bases, as we will see. nonsymplectic
However, an orthogonal geometry has an orthogonal basis if and symplectic
Metric Vector Spaces: The Theory of Bilinear Forms 283
only if it is totally degenerate. Furthermore, we have seen that if , char²-³ £
then all orthogonal symplectic geometries are totally degenerate and so all such
geometries have orthogonal bases. But if , then there are orthogonal char²-³ ~
symplectic geometries that are not tota lly degenerate and therefore do not have
orthogonal bases.
Thus, if we exclude orthogonal symplectic geometries when , we char²-³ ~
can say that every orthogonal geometry has an orthogonal basis.
If a metric vector space has an orthogonal basis, the natural next step is to =
look for an orthonormal basis. However, if is singular, then there is a nonzero=
vector and such a vector can never be a linear combination of vectors#=
from an orthonormal basis , since the coefficients in such a linear ¸" ÁÃÁ" ¹
combination are . º#Á" » ~
However, even if is nonsingular, ort honormal bases do not always exist and =
the question of how close we can come to such an orthonormal basis depends on
the nature of the base field. We w ill examine this issue in three cases:
algebraically closed fields, the fiel d of real numbers and finite fields.
We should also mention that even when has an orthogonal basis, the Gram– =
Schmidt orthogonalization process may not apply to produce such a basis,
because even nonsingular orthogonal geometries may have isotropic vectors,
and so division by is problematic. º"Á"»
For example, consider an orthogonal hyperbolic plane and /~ ² " Á# ³ span
assume that . Thus, and are isotropic and . The char²-³ £ " # º"Á#» ~
vector cannot be extended to an orthogonal basis using the Gram–Schmidt"
process, since is orthogonal if and only if . However, does ¸"Á"b#¹ ~ /
have an orthogonal basis, namely, . ¸"b#Á"c#¹
Orthogonal Bases
Let be an orthogonal geometry. As we have discussed, if is also==
symplectic, then has an orthogonal basis if and only if it is totally degenerate. =
Moreover, when , all orthogonal symplectic geometries are totally char²-³ £
degenerate and so all orthogonal symplectic geometries have an orthogonal
basis.
If is orthogonal but not symplectic, th en contains a nonisotropic vector , == "
the subspace is nonsingular and span²" ³
=~ ² "³ p = span
where . If is not symplectic, then we may decompose it to get=~ ² " ³ = span
= ~ ²" ³p ²" ³p= span span
284 Advanced Linear Algebra
This process may be continued until we reach a decomposition
= ~ ²" ³pÄp ²" ³p< span span
where is symplectic as well as orthogonal. (This includes the case .)< < ~ ¸¹
Let .8~² "ÁÃÁ"³
If , then is totally degenerate. Thus, if is a basis for , then thechar²-³ £ < < 9
union is an orthogonal basis for . If , then89r= ² - ³ ~ char
<~ p ² < ³>> rad , where is hyperbolic and so
= ~ ²"³pÄp ²"³p p ²<³ span span rad >
where is totally degenerate and the are nonisotropic. If rad²<³ "
9> :~² %Á&ÁÃÁ% Á& ³ ~² 'ÁÃÁ' ³ is a hyperbolic basis for and is an
ordered basis for , then the union rad²<³
;89:~ r r ~²" ÁÃÁ" Á% Á& ÁÃÁ% Á& Á'ÁÃÁ' ³
is an ordered orthogonal basis for . However, we can do better (in some =
sense).
The following lemma says that when , a pair of isotropic basis char²-³ ~
vectors, such as , can be replaced by a pair of nonisotropic basis vectors, %Á&
when coupled with a nonisotropic basis vector, such as . "
Lemma 11.20 Suppose that . Let be a three-dimensional char²-³ ~ >
orthogonal geometry of the form
> ~ ²#³p ²%Á&³ span span
where is nonisotropic and is a hyperbolic plane. Then#/ ~ ² % Á & ³ span
> ~ ²# ³p ²# ³p ²# ³ span span span
where each is nonisotropic. #
Proof. It is straightforward to check that if , then the vectors º#Á#» ~
# ~"b%b&
#~ " b %
# ~ "b²c³%b&
are linearly independent and mutually orthogonal. Details are left to the
reader.
Using the previous lemma, we can replace the vectors with the ¸" Á% Á& ¹
nonisotropic vectors , while retaining orthogonality. Moreover, ¸# Á# Á# ¹ b b
the replacement process can continue until the isotropic vectors are absorbed,
leaving an orthogonal basi s of nonisotropic vectors.
Metric Vector Spaces: The Theory of Bilinear Forms 285
Let us summarize.
Theorem 11.21 Let be an orthogonal geometry.=
1 If is also symplectic, then has an orthogonal basis if and only if it is)==
totally degenerate. When , all orthogonal symplectic char²-³ £
geometries have an orthogonal basis, but this is not the case when
char²-³ ~ .
2 If is not symplectic, then has an ordered orthogonal basis)==
8~²" ÁÃÁ" Á'ÁÃÁ' ³ º"Á"»~ £ º'Á'»~ for which and .
Hence, has the diagonal form48
4~
Æ
Æ
8vy
x{x{x{x{x{x{
wz
with nonzero entries on the diagonal.~ ² 4³ rk8
As a corollary, we get a nice theorem about symmetric matrices.
Corollary 11.22 Let be a symmetric matrix and assume that is not44
alternate if . Then is congruent to a diagonal matrix. char²-³ ~ 4
The Classification of Orthogonal Geometries: Canonical
Forms
We now want to consider the question of improving upon Theorem 11.21. The
diagonal matrices of this theorem do not form a set of canonical forms for
congruence. In fact, if are nonzero scalars, then the matrix of with Á Ã Á =
respect to the basis is 9~² " ÁÃÁ " Á'ÁÃÁ' ³
4~
Æ
Æ
9vy
x{x{x{x{x{x{
wz
()11.2
Hence, and are congruent diagonal matrices. Thus, by a simple change4489
of basis, we can multiply any diagonal entry by a nonzero square in . -
The determination of a set of canonical forms for symmetric nonalternate when (
char )²-³ ~ matrices under congruence depends on the properties of the base
field. Our plan is to consider three types of base fields: algebraically closed
286 Advanced Linear Algebra
fields, the real field and finite fields. Here is a preview of the forthcoming s
results.
1 When the base field is algebraically closed, there is an ordered basis ) - 8
for which
4~ A ~
Æ
Æ
8 Ávy
x{x{x{x{x{x{
wz
If is nonsingular, then is an identity matrix and has an=4 = 8
orthonormal basis.
2 Over the real base field, there is an ordered basis for which) 8
4~ ~
Æ
c
Æ
c
Æ
8ZÁÁvy
x{x{x{x{x{x{x{x{x{x{x{x{
wz
3 If is a finite field, there is an ordered basis for which)- 8
4~ ² ³ ~
Æ
Æ
8ZÁvy
x{x{x{x{x{x{x{x{
wz
where is unique up to multiplication by a square and if , then² - ³ ~ char
we can take . ~
Now let us turn to the details.
Algebraically Closed Fields
If is algebraically closed, then for every , the polynomial has a- - % c
root in , that is, every element of has a square root in . Therefore, we may-- -
choose in 11.2 , which leads to the following result.~ ° j ()
Metric Vector Spaces: The Theory of Bilinear Forms 287
Theorem 11.23 Let be an orthogonal geometry over an algebraically closed=
field . Provided that is not symplectic as well when , then -= ² - ³ ~ = char
has an ordered orthogonal basis for which 8~²" ÁÃÁ" Á'ÁÃÁ' ³
º "Á"»~ º 'Á'»~ 4 and . Hence, has the diagonal form 8
4~ A ~
Æ
Æ
8 Ávy
x{x{x{x{x{x{
wz
with ones and zeros on the diagonal . In particular, if is nonsingular, =
then has an orthonormal basis.=
The matrix version of Theorem 11.23 follows.
Theorem 11.24 Let be the set of all symmetric matrices over anI d
algebraically closed field . If , we restrict to the set of all -² - ³ ~ char I
symmetric matrices with at least one nonzero entry on the main diagonal.
1 Any matrix in is congruent to a unique matrix of the form Z , in) 4I Á
fact, and .~ ² 4³ ~c ² 4³ rk rk
2 The set of all matrices of the form Z for is a set of canonical) Áb~
forms for congruence on . I
3 The rank of a matrix is a complete invariant for congruence on .) I
The Real Field s
If , we can choose , so that all nonzero diagonal elements in-~ ~ ° s j((
()11.2 will be either , or . c
Theorem 11.25 Sylvester's law of inertia () Any real orthogonal geometry =
has an ordered orthogonal basis
8~²" ÁÃÁ" Á# ÁÃÁ# Á'ÁÃÁ' ³
for which , and . Hence, the matrix has º "Á"»~º #Á#»~c º 'Á'»~ 4 8
the diagonal form
288 Advanced Linear Algebra
4~ ~
Æ
c
Æ
c
Æ
8ZÁÁvy
x{x{x{x{x{x{x{x{x{x{x{x{
wz
with ones, negative ones and zeros on the diagonal.
Here is the matrix version of Theorem 11.25.
Theorem 11.26 Let be the set of all symmetric matrices over the realI d
field .s
1 Any matrix in is congruent to a unique matrix of the form Z for) I ÁÁ Á
some and .Á ~cc
2 The set of all matrices of the form Z for is a set of) ÁÁ bb~
canonical forms for congruence on . I
3 Let and let be congruent to . Then is the rank of)4 4 A bI ÁÁ
4 c 4 ² Á Á ³ and is the of and the triple is the signature inertia
of . The pair , or equivalently the pair , is a4 ²Á³ ²bÁc³
complete invariant under congruence on . I
Proof. We need only prove the uniqueness statement in part 1 . Let )
8~²" ÁÃÁ" Á# ÁÃÁ# Á'ÁÃÁ' ³
and
9~²" ÁÃÁ" Á# ÁÃÁ# 'ÁÃÁ' ³ZZ ZZ ZZ
ZZ Z
be ordered bases for which the matrices and have the form shown in 4489
Theorem 11.25. Since the rank of these matrices must be equal, we have
b~ b ~ZZ Z and so .
If and , then% ²" ÁÃÁ" ³ %£span
º % Á% »~ "Á " ~ º "Á"»~ ~ LM Á
Á Á
On the other hand, if and , then & ²# ÁÃÁ# ³ &£spanZZ
Z
º & Á& »~ #Á # ~ º #Á#»~c ~c LM Á ZZ Z Z
Á Á
Hence, if then . It follows that & ²# ÁÃÁ# Á'ÁÃÁ' ³ º&Á&»spanZZ ZZ
Z Z
Metric Vector Spaces: The Theory of Bilinear Forms 289
span span²" ÁÃÁ" ³q ²# ÁÃÁ# Á'ÁÃÁ' ³~¸¹ZZ ZZ
ZZ
and so
b² c³Z
that is, . By symmetry, and so . Finally, since , it ~ ~ZZ Z Z
follows that . ~Z
Finite Fields
To deal with the case of finite fiel ds, we must know something about the
distribution of squares in finite fields , as well as the possible values of the
scalars .º#Á#»
Theorem 11.27 Let be a finite field with elements.-
1 If , then every element of is a square.)c h a r ²- ³ ~ -
2 If , then exactly half of the nonzero elements of are)c h a r ²- ³ £ -
squares, that is, there are nonzero squares in . Moreover, if ² c³° - %
is any nonsquare in , then all nonsquares have the form , for some - %
-.
Proof. Write , let be the subgroup of all nonzero elements in and-~- - - i
let
²- ³ ~ ¸ - ¹i i
be the subgroup of all nonzero squares in . The - Frobenius map
¢- ¦²- ³ ²³~ii defined by is a surjective group homomorphism, with
kernel
ker² ³ ~ ¸ - ~ ¹ ~ ¸cÁ¹
If , then and so is bijective and ,char² -³~ ² ³~¸ ¹ - ~ ² - ³ ker (( ( (ii
which proves part 1 . If , then and so , )c h a r²-³ £ ² ³ ~ - ~ ²- ³ (( ( ( ( (kerii
which proves the first part of part 2 . We leave proof of the last statement to the )
reader.
Definition A bilinear form on is if for any nonzero there = - universal
exists a vector for which . # = º#Á#» ~
Theorem 11.28 Let be an orthogonal geometry over a finite field with=-
char²-³ £ = and assume that has a nonsi ngular subspace of dimension at
least . Then the bilinea r form of is universal. =
Proof. Theorem 11.21 implies that contai ns two linearly independent vectors =
"# and for which
º"Á"»~£Á º#Á#»~£Á º"Á#»~
290 Advanced Linear Algebra
Given any , we want to find and for which -
~ º" b# Á" b# » ~ b
or
~ c
If , then , since there are nonzero( ~ ¸ -¹ ( ~ ² b³° ² c³°((
squares , along with . If , then for the same ~ )~¸ c -¹
reasons . It follows that cannot be the empty set and so (() ~ ² b³° (q)
there exist and for which . ~ c
Now we can proceed with the business at hand.
Theorem 11.29 Let be an orthogonal geometry over a finite field and=-
assume that is not symplectic if . If , then let be a =² - ³ ~ ² - ³ £ char char
fixed nonsquare in . For any nonzero , write - -
?² ³~
Æ
Æ
vy
x{x{x{x{x{x{x{x{
wz
where . rk²? ²³³ ~
1 If , then there is an ordered basis for which .)c h a r ²-³ ~ 4 ~ ? ²³ 8 8
2 If , then there is an ordered basis for which equals)c h a r ²-³ £ 4 8 8
?² ³ ?² ³ or .
Proof. We can dispose of the case quite easily: Referring to 11.2 , char ( )²-³ ~
since every element of has a square root, we may take . - ~ ² ³ cj
If , then Theorem 11.21 implies that there is an ordered orthogonalchar²-³ £
basis
8~²" ÁÃÁ" Á'ÁÃÁ' ³
for which and . Hence, has the diagonal form º "Á"»~ £ º 'Á'»~ 4 8
4~
Æ
Æ
8vy
x{x{x{x{x{x{
wz
Metric Vector Spaces: The Theory of Bilinear Forms 291
Now consider the nonsingular orthogonal geometry . =~ ² " Á " ³ span
According to Theorem 11.28, the form is universal when restricted to . =
Hence, there exists a for which . # = º #Á #»~
Now, for not both , and we may swap and if#~ "b " Á - " "
necessary to ensure that . Hence, £
8 ~²# Á" ÁÃÁ" Á'ÁÃÁ' ³
is an ordered basis for for which the matrix is diagonal and has a in the =4 8
upper left entry. We can repeat the process with the subspace . =~ ² # Á # ³ span
Continuing in this way, we can find an ordered basis
9~² #Á#ÁÃÁ#Á'ÁÃÁ' ³
for which for some nonzero . Now, if is a square in , 4~ ? ² ³ - -9
then we can replace by to get a basis for which . If #² ° ³ # 4 ~ ? ² ³ j : :
- ~ - # is not a square in , then for some and so replacing by
²°³# 4 ~ ? ²³ gives a basis for which . : :
Theorem 11.30 Let be the set of all symmetric matrices over a finiteI d
field . If , we restrict to the set of all symmetric matrices with-² - ³ ~ char I
at least one nonzero entry on the main diagonal.
1 If , then any matrix in is c ongruent to a unique matrix of the )c h a r ²-³ ~ I
form and the matrices form a set of? ²³ ¸? ²³ ~ ÁÃÁ¹
canonical forms for under congruence. Also, the rank is a complete I
invariant.
2 If , let be a fixed nonsquare in . Then any matrix is)c h a r ²-³ £ - I
congruent to a unique matrix of the form or . The set ?² ³ ?² ³
¸? ²³Á? ²³ ~ ÁÃÁ¹ is a set of canonical forms for congruence
on . Thus, there are exactly two congruence classes for each rank .I()
The Orthogonal Group
Having “settled” the classification question for orthogonal geometries over
certain types of fields, let us turn to a discussion of the structure-preserving
maps, that is, the isometries.
Rotations and Reflections
We begin by examining the matrix of an orthogonal transformation. If is an 8
ordered basis for , then for any , =% Á & =
º%Á&» ~ ´%µ 4 ´&µ888!
and so if , thenB² = ³
º% Á& » ~ ´% µ4´& µ ~ ´ % µ² ´µ4´µ³ ´ & µ 88 888 8 88!! !
Hence, is an orthogonal transformation if and only if
292 Advanced Linear Algebra
´µ4´µ ~ 4888 8!
Taking determinants gives
det det det²4 ³ ~ ²´ µ ³ ²4 ³88 8
Therefore, if is nonsingular, then =
det²´ µ ³ ~ f8
Since the determinant is an invariant under similarity, we have the following
theorem.
Theorem 11.31 Let be an orthogonal tr ansformation on a nonsingular
orthogonal geometry . =
1 is the same for all ordered bases for and)d e t²´ µ ³ =88
det²´ µ ³ ~ f8
This determinant is called the of and denoted by . determinant det²³
2 If , then is called a and if , then is)d e t d e t ²³ ~ ²³ ~ c rotation
called a . reflection
3 The set of rotations is a subgroup of the orthogonal group )EEb²= ³ ²= ³
and the determinant map is an epimorphism with det¢² = ³ ¦ ¸ c Á ¹E
kernel . Hence, if , then is a normal subgroupEEbb²= ³ ²-³ £ ²= ³ char
of of index .E²= ³
Symmetries
Recall again that for a real inner product space, a reflection is defined as an /"
operator for which
/"~c " Á/$~$ $º " »"" for all
and that
/%~%c "º%Á"»
º"Á"»"
In particular, if and is nonisotropic, then is char span²-³ £ " = ²"³
nonsingular and so
= ~ ²"³p ²"³ span span
Then the reflection is well-defined and, in the context of general orthogonal /"
geometries, is called the determined by and we will denote it by symmetry "
"". We can also write , that is, ~c p
"²%b&³~c%b&
for all and .% ² " ³ & ² " ³span span
Metric Vector Spaces: The Theory of Bilinear Forms 293
For real inner product spaces, Theorem 10.16 says that if , then )) ) )#~$£
/# $ / ² # ³ ~ $#c$ #c$ is the unique reflection sending to , that is, . In the
present context, we must be careful, since symmetries are defined for
nonisotropic vectors only. Here is what we can say.
Theorem 11.32 Let be a nonsingular orthogonal geometry over a field ,=-
with . If are nonisotropic vectors with the same nonzero char²-³ £ "Á# = ()
“length,” that is, if
º"Á"»~º#Á#»£
then there exists a symmetry for which
"~# "~c # or
Proof. Since and are nonisotropic, one of or must also be"# " c # " b #
nonisotropic, for otherwise, since and are orthogonal, their sum "c# "b# "
would also be isotropic. If is nonisotropic, then"b#
"b#²"b#³~c²"b#³
and
"b#²"c#³ ~ "c#
and so . On the other hand, if is nonisotropic, then"b#"~c # "c#
"c#²"c#³~c²"c#³
and
"c#²"b#³ ~ "b#
and so ."c#"~#
Recall that an operator on a real inner product space is unitary if and only if it is
a product of reflections. Here is the generalization to nonsingular orthogonal
geometries.
Theorem 11.33 Let be a nonsingular orthogonal geometry over a field =-
with . A linear transformation on is an orthogonal char²-³ £ =
transformation if and only if is the product of symmetries on . =
Proof. The proof is by induction on . If , then ~ ²= ³ ~ = ~ ²#³ dim span
where . Let for . Since is unitaryº#Á#» £ # ~ # -
º#Á#» ~ º #Á #» ~ º #Á #» ~ º#Á#»
and so . If , then is the iden tity, which is equal to . On the ~f ~#
other hand, if then . In either case, is a product of symmetries. ~c ~ #
294 Advanced Linear Algebra
Assume now that the theorem is true for dimensions less than and let
dim² =³~ #= º # Á # »~º # Á# »£ . Let be nonisotropic. Since , Theorem
11.32 implies the existence of a symmetry on for which =
²# ³ ~#
where . Thus, on . Sinc implies that ~f ~f ² # ³ span e Theorem 11.9
span²#³ is -invariant, we may apply the induction hypothesis to on
span²#³ to get
O~ Ä ~span²#³ $$
where and each is a symmetry on . But each can$ ² # ³ ² # ³$ $span span
be extended to a symmetry on by setting . Assume that is the =# ~ # $
extension of to , where on . Hence, on and =~ ~² # ³ span²#³ span
~ on span²#³.
If , then on and so , which completes the proof. If ~ ~ = ~
~c ~ ² # ³ , then on since is the identity on and ##span span ²#³
~~ = ~ =## # on on and so on span²#³. Hence, .
The Witt Theorems for Orthogonal Geometries
We are now ready to consider the Witt theorems for orthogonal geometries.
Theorem 11.34 Witt's cancellation theorem () Let and be isometric=>
nonsingular orthogonal geometries over a field with . Suppose -² - ³ £ char
that
=~ :p : >~ ;p ;and
Then
:;¬: ;
Proof. First, we prove that it is sufficient to consider the case . Suppose =~ >
that the result holds when and that is an isometry. Then =~ > ¢ =¦ >
²:³p ²: ³ ~ ²: p: ³ ~ = ~ > ~ ; p;
Furthermore, . We can therefore apply the theorem to to get ::; >
:² : ³ ;
as desired. To prove the theorem when , assume that =~ >
=~ :p : ~ ;p ;
where and are nonsingular and . Let be an isometry. We: ; :; ¢:¦;
proceed by induction on . dim²:³
Metric Vector Spaces: The Theory of Bilinear Forms 295
Suppose first that and that . Since dim²:³ ~ : ~ ² ³ span
º Á »~º Á »£
Theorem 11.32 implies that there is a symmetry for which where ~
~f = ;~ : . Hence, is an isometry of for which and Theorem 11.9
implies that . Thus, is the desired isometry. ;~² : ³ O:
Now suppose the theorem is true for and let . Let dim dim²:³ ²:³ ~
¢:¦; : be an isometry. Since is nons ingular, we can choose a nonisotropic
vector and write , where is nonsingular. It follows : :~ ² ³p< < span
that
=~ :p : ~ ² ³ p <p :span
and
=~ ;p ; ~ ² ² ³ ³ p <p ;span
Now we may apply the one-dimensional case to deduce that
<p: <p;
If is an isometry, then¢<p: ¦ <p;
<p ² : ³~ ² <p: ³~ <p;
But and since , the induction hypothesis < < ² <³~ ²<³ dim dim
implies that . :² : ³ ;
As we have seen, Witt's extension theo rem is a corollary of Witt's cancellation
theorem.
Theorem 11.35 Witt's extension theorem () Let and be isometric==Z
nonsingular orthogonal geometries over a field , with . Suppose -² - ³ £ char
that is a subspace of and<=
¢< ¦ < =Z
is an isometry. Then can be extended to an isometry from to . ==Z
Maximal Hyperbolic Subspaces of an Orthogonal Geometry
We have seen that any orthogonal geometry can be written in the form =
=~ <p ² = ³ rad
where is nonsingular. Nonsingular spaces are better behaved than singular<
ones, but they can still possess isotropic vectors.
296 Advanced Linear Algebra
We can improve upon the preceding decomposition by noticing that if is "<
isotropic, then Theorem 11.10 implies that can be “captured” in a span²"³
hyperbolic plane . Then we can write /~ ² " Á% ³ span
=~ /p / p ² = ³<rad
where is the orthogonal complement of in and has “one fewer”// <<
isotropic vector. In order to generalize this process, we first discuss maximal
totally degenerate subspaces.
Maximal Totally Degenerate Subspaces
Let be a nonsingular orthogonal ge ometry over a field , with . =- ² - ³ £ char
Suppose that and are maximal totally degenerate subspaces of . We << =Z
claim that . For if , then there is a vector dim dim dim dim²<³ ~ ²< ³ ²<³ ²< ³ZZ
space isomorphism , which is also an isometry, since and ¢< ¦ < < <Z
<Z are totally degenerate. Thus, Witt's extension theorem im plies the existence
of an isometry that extends . In particular, is a totally ¢= ¦= ²<³cZ
degenerate space that contains and so , which shows that <² < ³ ~ < cZ
dim dim²<³ ~ ²< ³Z.
Theorem 11.36 Let be a nonsingular orthogonal geometry over a field ,=-
with . char²-³ £
1 All maximal totally degenerate subs paces of have the same dimension, ) =
which is called the of and is denoted by . Witt index =$ ² = ³
2 Any totally degenerate subspace of of dimension is maximal.) =$ ² = ³
Maximal Hyperbolic Subspaces
We can prove by a similar argument that all maximal hyperbolic subspaces of =
have the same dimension. Let
> ~/ pÄp/
and
A ~2 pÄp2
be maximal hyperbolic subspaces of and suppose that and =/ ~ ² " Á # ³ span
2~ ² % Á & ³ ² ³ ²³ span . We may assume that . dim dim>A
The linear map defined by > A¢¦
"~ % Á #~ &
is clearly an isometry from to . Thus , Witt's extension theorem implies the > >
existence of an isometry that extends . In particular, is a A¢= ¦= ² ³c
hyperbolic space that contains and so . It follows that > A > Ac²³ ~ ²³ dim
~² ³dim>.
Metric Vector Spaces: The Theory of Bilinear Forms 297
It is not hard to see that the maxi mum dimension of a hyperbolic subspace ²=³
of is , where is the Witt index of . First, the nonsingular= $²= ³ $²= ³ =
extension of a maximal totally degenerate subspace of is a hyperbolic <=$
space of dimension and so . On the other hand, there is a $²= ³ ²= ³ $²= ³
totally degenerate subspace contained in any hyperbolic space and so < >
$ ² =³ ² ³$ ² =³ ² =³$ ² =³ , that is, . Hence and so dim>
² =³~$ ² =³ .
Theorem 11.37 Let be a nonsingular orthogonal geometry over a field ,=-
with . char²-³ £
1 All maximal hyperbolic subspaces of have dimension .) = $ ² = ³
2 Any hyperbolic subspace of dimension must be maximal.) $²= ³
3 The Witt index of a hyperbolic space is .) >
The Anisotropic Decomposition of an Orthogonal Geometry
If is a maximal hyperbolic subspace of , then> =
=~ p>>
Since is maximal, is anisotropic, for if were isotropic, then the >> >"
nonsingular extension of would be a hyperbolic space strictly >p² " ³span
larger than .>
Thus, we arrive at the following decomposition theorem for orthogonal
geometries.
Theorem 11.38 The anisotropic decomposition of an orthogonal geometry ()
Let be an orthogonal geometry over , with . Let=~ <p ² = ³ - ² - ³ £ rad char
>> be a maximal hyperbolic subspace of , where if has no < ~ ¸¹ <
isotropic vectors. Then
=~ :p p ² = ³> rad
where is anisotropic, is hyperbolic of dimension and is: $ ² = ³ ² = ³ > rad
totally degenerate.
Exercises
1. Let be subspaces of a metric vector space . Show that<Á> =
a )<>¬> <
b )<<
c )<~ <
2. Let be subspaces of a metric vector space . Show that<Á> =
a )²< b>³ ~ < q>
b )²< q>³ ~ < b>
3. Prove that the following are equivalent:
a is nonsingular )=
b for all implies )º"Á%» ~ º#Á%» % = " ~ #
298 Advanced Linear Algebra
4. Show that a metric vector space is nonsingular if and only if the matrix =
48 of the form is nonsingular, for every ordered basis . 8
5. Let be a finite-dimensional vector space with a bilinear form . We do=º Á »
not assume that the form is symmetric or alternate. Show that the following
are equivalent:
a for all )¸ #= º # Á$ »~ $=¹~
b for all )¸ #= º $ Á# »~ $=¹~
: Consider the singularity of the matrix of the form.Hint
6. Find a diagonal matrix congruent to
vy
wz
c
7. Prove that the matrices
0~ 4 ~
>? >? and
are congruent over the base field of rational numbers. Find an -~r
invertible matrix such that . 77 0 7 ~ 4!
8. Let be an orthogonal geometry over a field with . We =- ² - ³ £ char
wish to construct an orthogonal basis for , starting with E~² "ÁÃÁ"³ =
any generating set . Justify th e following steps, essentially =~² #ÁÃÁ#³
due to Lagrange. We may assume that is not totally degenerate. =
a If for some , then let . Otherwise, there are indices )º# Á# » £ " ~ #
£ º #Á#»£ " ~#b# for which . Let .
b Assume we have found an ordered set of vectors ) E ~² "ÁÃÁ"³
that form an orthogonal basis for a subspace of and that none of ==
the 's are isotropic. Then ."= ~ = p =
c For each , let ) #=
$~ #c "º# Á" »
º" Á" »
~
Then the vectors span . If is totally degenerate, take any $= =
basis for and append it to . Otherwise, repeat step a on to ==E )
get another vector and let . Eventually, we " ~²" ÁÃÁ" ³b b b E
arrive at an orthogonal basis for . E=
9. Prove that orthogonal hyperbolic planes may be characterized as two-
dimensional nonsingular orthogonal geomet ries that have exactly two one-
dimensional totally isotropic equivalently: totally degenerate subspaces. ()
10. Prove that a two-dimensional nons ingular orthogonal geometry is a
hyperbolic plane if and only if its discriminant is . -² c ³
11. Does Minkowski space contain any isotropic vectors? If so, find them.
12. Is Minkowski space isometric to Euclidean space ? s
Metric Vector Spaces: The Theory of Bilinear Forms 299
13. If is a symmetric bilinear form on and , show thatºÁ» = ²-³ £ char
8²%³ ~ º%Á%»° is a quadratic form.
14. Let be a vector space over a field , with ordered basis .=- ~ ² # Á Ã Á # ³ 8
Let be a polynomial of degree over , that is,²% ÁÃÁ% ³ - homogeneous
a polynomial each of whose terms has degree . The defined by -form
is the function from to defined as follows. If , then =- # ~ # '
²#³~² ÁÃÁ ³
()We use the same notation for the form and the polynomial. Prove that -
forms are the same as quadratic forms.
15. Show that is an isometry on if and only if where is =8 ² # ³ ~ 8 ² # ³ 8
the quadratic form associated with th e bilinear form on . Assume that =(
char²-³ £ ³ .
16. Show that a quadratic form on satisfies the parallelogram law: 8=
8²%b&³b8²%c&³~´8²%³b8²&³µ
17. Show that if is a nonsingular orthogonal geometry over a field , with =-
char²-³ £ = , then any totally isotropic subspace of is also a totally
degenerate space.
18. Is it true that ? = ~ ²= ³p ²= ³ rad rad
19. Let be a nonsingular symplectic geometry and let be a symplectic = #Á
transvection. Prove that
a ) #Á #Á #Áb~
b For any symplectic transformation , )
#Á #Ác~
c F o r , )-i
#Á #Á~
d For a fixed , the map is an isomorphism from the ) #£ ª #Á
additive group of onto the group Sp . -¸ - ¹ ² = ³ #Á
20. Prove that if is any nonsquare in a finite field , then all nonsquares %-
have the form , for some . Hence, the product of any two % -
nonsquares in is a square. -
21. Formulate Sylvester's law of inertia in terms of quadratic forms on . =
22. Show that a two-dimensional space is a hyperbolic plane if and only if it is
nonsingular and contains an isotropic vector. Assume that . char²-³ £
23. Prove directly that a hyperbolic plane in an orthogonal geometry cannot
have an orthogonal basis when . char²-³ ~
24. a Let be a subspace of . Show that the inner product ) <=
º%b<Á&b<»~º%Á&» =°< on the quotient space is well-defined if
and only if . < ² =³ rad
b If , when is nonsingular? )r a d< ² =³ =° <
25. Let , where is a totally degenerate space.=~ 5p : 5
300 Advanced Linear Algebra
a Prove that if and only if is nonsingular. )r a d 5~ ² =³ :
b If is nonsingular, prove that . )r a d:: = ° ² = ³
26. Let . Prove that implies dim dim²= ³ ~ ²>³ = ° ²= ³ >° ²>³ rad rad
= > .
27. Let . Prove that=~ :p ;
a ) rad rad rad²= ³ ~ ²:³p ²;³
b ) rad rad rad= ° ²= ³ :° ²:³p;° ²;³
c ) rad rad raddim dim dim²² = ³ ³ ~ ²² : ³ ³ b ²² ; ³ ³
d is nonsingular if and only if and are both nonsingular. )=: ;
28 Because the Riesz. Let be a nonsingular metric vector space. =
representation theorem is valid in , we can define the adjoint of a linear= i
map exactly as in the case of real inner product spaces. ProveB² = ³
that is an isometry if and only if it is bijective and unitary that is, (
i~).
29. If , prove that is an isometry if and only if it is char²-³ £ ²= Á>³ B
bijective and for all . ºÁ» ~ º # Á # » # =##
30. Let be a basis for . Prove that is an8 B~¸ #ÁÃÁ#¹ = ² =Á>³
isometry if and only if it is bijective and for all . º #Á #»~º#Á#» Á
31. Let be a linear operator on a metric vector space . Let 8 = ~²# ÁÃÁ# ³
be an ordered basis for and let be the matrix of the form relative to =4 8
8. Prove that is an isometry if and only if
´µ 4´µ ~ 4888 8!
32. Let be a nonsingular orthogona l geometry and let be an = ² = ³ B
isometry.
a Show that . )i m dim ker dim² ²c³ ³ ~ ² ²c³³
b Show that . How would you describe )i m ker²c³ ~ ²c³
ker²c³ in words?
c If is a symmetry, what is ? ) dim ker²² c ³ ³
d Can you characterize symmetries by means of ? )d i m k e r ²² c ³ ³
33. A linear transformation is called if is nilpotent. B ² = ³ c unipotent
Suppose that is a nonisotropic metric vector space and that is unipotent =
and isometric. Show that . ~
34. Let be a hyperbolic space of dimension and let be a hyperbolic= <
subspace of of dimension . Show that for each , there is a =
hyperbolic subspace of for which . >> =< =
35. Let . Prove that if is a totally degenerate subspace of an char²-³ £ ?
orthogonal geometry , then . =² ? ³ ² = ³ ° dim dim
36. Prove that an orthogonal geometry of dimension is a hyperbolic space=
if and only if is nonsingular, is even and contains a totally = =
degenerate subspace of dimension . °
37. Prove that a symplectic transformation has determinant equal to .
Chapter 12
Metric Spaces
The Definition
In Chapter 9, we studied the basic properties of real and complex inner product
spaces. Much of what we did does not depend on whether the space in question
is finite-dimensional or infinite-dime nsional. However, as we discussed in
Chapter 9, the presence of an inner product and hence a metric, on a vector
space, raises a host of new issues related to convergence. In this chapter, we
discuss briefly the concept of a metric space. This will enable us to study the
convergence properties of real and complex inner product spaces.
A metric space is not an algebraic structure. Rather it is designed to model the
abstract properties of distance.
Definition A is a pair , where is a nonempty set andmetric space ²4Á³ 4
¢4 d4 ¦ 4 s is a real-valued function, called a on , with the metric
following properties. The expression is read “the distance from to .” ²%Á&³ % &
1 For all ,)( )Positive definiteness %Á& 4
²%Á&³
and if and only if .²%Á&³ ~ % ~ &
2 For all ,)( )Symmetry %Á& 4
²%Á&³ ~ ²&Á%³
3 For all ,)( )Triangle inequality %Á&Á' 4
²%Á&³ ²%Á'³b²'Á&³
As is customary, when there is no cause for confusion, we simply say “let be 4
a metric space.”
302 Advanced Linear Algebra
Example 12.1 Any nonempty set is a metric space under the 4 discrete
metric , defined by
²%Á&³ ~% ~ &
% £ &Fif
if
Example 12.2
1 The set is a metric space, under the metric defined for )s%~²% ÁÃÁ% ³
and by&~²& ÁÃÁ& ³
²%Á&³~ ²% c&³ bÄb²% c& ³ j
This is called the on . We note that is also a metric Euclidean metric ss
space under the metric
²%Á&³~ % c& bÄb % c& (( ((
Of course, and are different metric spaces. ²Á ³ ²Á ³ss
2 The set is a metric space under the )dunitary metric
²%Á&³~ % c& bÄb % c& k(( ((
where and are in . %~²% ÁÃÁ% ³ &~²& ÁÃÁ& ³ d
Example 12.3
1 The set of all real-valued or complex-valued continuous functions)( ) *´Áµ
on is a metric space, under the metric´Áµ
²Á³ ~ ²%³c²%³ sup
%´Áµ((
We refer to this metric as the . sup metric
2 The set of all real-valued or complex-valued continuous functions)) *´Áµ ²
on is a metric space, under the metric´Áµ
²²%³Á²%³³ ~ ²%³c²%³ %
((
Example 12.4 Many important sequence spaces are metric spaces. We will
often use boldface italic letters to denote sequences, as in and %~² %³
&~² &³.
1 The set of all bounded sequences of real numbers is a metric space) MsB
under the metric defined by
² Á ³ ~ % c&%& sup
((
The set of all bounded complex sequences, with the same metric, is alsoMdB
a metric space. As is customary, we will usually denote both of these spaces
by .MB
Metric Spaces 303
2 For , let be the set of all sequences of real or complex)) M ~² %³ ²%
numbers for which
((
~B
% B
We define the of by -norm%
)) ( (89%
~B
°
~%
Then is a metric space, under the metricM
² Á ³ ~ c ~ % c&%& % & )) ( (89
~B
°
The fact that is a metric follows from some rather famous results about M
sequences of real or complex numbers, whose proofs we leave as well- (
hinted exercises. )
Let and . If and ,Holder's inequality¨ Á b ~ M M %&
then the product sequence is in and %&~² %&³ M
)) ) ) ) )%& % &
that is,
(( ( ( ( (89 89
~ ~ ~BB B
° °
%& % &
A special case of this with 2 is the ()~~ Cauchy–Schwarz inequality
(( ( ( ( ( mm
~ ~ ~BB B
%& % &
Minkowski's inequality For , if then the sum Á M b %& % &
~² % b&³ M is in and
)) ) ) ) )%& % &b b
that is,
89 8 9 8 9 (( ( ( ( (
~ ~ ~BB B
° ° °
%b & % b &
304 Advanced Linear Algebra
If is a metric space under a metric , then any nonempty subset of is4 : 4
also a metric under the restriction of to . The metric space thus : d : :
obtained is called a of . subspace 4
Open and Closed Sets
Definition Let be a metric space. Let and let be a positive real4% 4
number.
1 The centered at , with radius , is) open ball %
) ² %Á ³~¸ %4 ² % Á %³ ¹
2 The centered at , with radius , is) closed ball %
) ² %Á ³~¸ %4 ² % Á %³ ¹
3 The centered at , with radius , is) sphere %
: ² %Á ³~¸ %4 ² % Á %³~ ¹
Definition A subset of a metric space is said to be if each point of :4 : open
is the center of an open ball that is contained completely in . More :
specifically, is open if for all , there exists an such that :% :
)²%Á³ : ; 4 . Note that the empty set is open. A set is if its closed
complement in is open. ;4
It is easy to show that an open ball is an open set and a closed ball is a closed
set. If , we refer to any open set containing as an %4 : % open
neighborhood of . It is also easy to see that a set is open if and only if it%
contains an open neighborhood of each of its points.
The next example shows that it is possi ble for a set to be both open and closed,
or neither open nor closed.
Example 12.5 In the metric space with the usual Euclidean metric, the open s
balls are just the open intervals
)²% Á³ ~ ²% cÁ% b³
and the closed balls are the closed intervals
)²% Á³ ~ ´% cÁ% bµ
Consider the half-open interval , for . This set is not open, since :~² Á µ
it contains no open ball centered at and it is not closed, since its :
complement is not open, since it contains no open ball : ~²cBÁµr²ÁB³
about .
Metric Spaces 305
Observe also that the empty set is both open and closed, as is the entire space . s
(Although we will not do so, it is possible to show that these are the only two
sets that are both open and closed in . s³
It is not our intention to enter into a detailed discussion of open and closed sets,
the subject of which belongs to the branch of mathematics known as . topology
In order to put these concepts in perspective, however, we have the following
result, whose proof is left to the reader.
Theorem 12.1 The collection of all open subsets of a metric space has the E 4
following properties:
1 , )J 4EE
2 If , then ):; : q ;EE
3 If is any collection of open sets, then .)¸: 2¹ : 2E
These three properties form the basis for an axiom system that is designed to
generalize notions such as convergence and continuity and leads to the
following definition.
Definition Let be a nonempty set. A collection of subsets of is called a?? E
topology for if it has the following properties:?
1)J Á?EE
2 I f t h e n ): Á; :q;EE
3 If is any collection of sets in , then .)¸: 2¹ :
2EE
We refer to subsets in as and the pair as a EE open sets topological ²?Á ³
space .
According to Theorem 12.1, the open sets as we defined them earlier in a ()
metric space form a topology for , called the topology by the 44 induced
metric.
Topological spaces are the most general setting in which we can define concepts
such as convergence and continuity, which is why these concepts are called
topological concepts. However, since the topologies with which we will be
dealing are induced by a metric, we will generally phrase the definitions of the
topological properties that we will need directly in terms of the metric.
Convergence in a Metric Space
Convergence of sequences in a metric space is defined as follows.
Definition A sequence in a metric space to , written ²% ³ 4 % 4 converges
²% ³ ¦ % , if
lim
¦B²% Á%³ ~
Equivalently, if for any , there exists an such that ²% ³ ¦ % 5
306 Advanced Linear Algebra
5¬ ² %Á% ³
or equivalently,
5¬% ) ² % Á ³
In this case, is called the of the sequence . %² % ³ limit
If is a metric space and is a subset of , by a , we mean a4: 4 : sequence in
sequence whose terms all lie in . We next characterize closed sets and :
therefore also open sets, using convergence.
Theorem 12.2 Let be a metric space. A subset is closed if and only if4: 4
whenever is a sequence in and , then . In loose terms, a ²% ³ : ²% ³ ¦ % % :
subset is closed if it is closed under the taking of sequential limits. :
Proof. Suppose that is closed and let , where for all . :² % ³ ¦ % % :
Suppose that . Then since and is open, there exists an for %¤: %: :
which . But this implies that%) ² % Á ³:
)²%Á ³q¸% ¹ ~ J
which contradicts the fact that . Hence, . ²% ³ ¦ % % :
Conversely, suppose that is closed under the taking of limits. We show that :
:% : % is open. Let and suppose to the contrary that no open ball about is
contained in . Consider the open balls , for all . Since none of : )²%Á°³
these balls is contained in , for each , there is an . It is : % : q)²%Á°³
clear that and so . But cannot be in both and . This ²% ³ ¦ % % : % : :
contradiction implies that is open. Thus, is closed. ::
The Closure of a Set
Definition Let be any subset of a metric space . The of , denoted:4 : closure
by , is the smallest closed set containing .cl²:³ :
We should hasten to add that, since the entire space is closed and since the 4
intersection of any collection of closed sets is closed exercise , the closure of ()
any set does exist and is the inters ection of all closed sets containing . The ::
following definition will allow us to ch aracterize the closure in another way.
Definition Let be a nonempty subset of a metric space . An element :4 % 4
is said to be a , or , of if every open ball limit point accumulation point :
centered at meets at a point other than itself. Let us denote the set of all %: %
limit points of by . :M ² : ³
Here are some key facts concerning limit points and closures.
Metric Spaces 307
Theorem 12.3 Let be a nonempty subset of a metric space .:4
1 if and only if there is a sequence in for which for)%M ² :³ ² %³ : % £%
all and .² % ³ ¦ %
2 is closed if and only if . In words, is closed if and only if it):M ² : ³ : :
contains all of its limit points.
3 .)c l²:³ ~ : rM²:³
4 if and only if there is a sequence in for which .)c l% ² : ³ ² %³ : ² %³¦%
Proof. For part 1 , assume first that . For each , there exists a point ) %M ² :³
% £ % % )²%Á°³q: such that . Thus, we have
²% Á%³ °
and so . For the converse, suppose that , where .²% ³ ¦ % ²% ³ ¦ % % £ % :
If is any ball centered at , then there is some such that )²%Á³ % 5 5
implies . Hence, for any ball centered at , there is a point% )²%Á³ )²%Á³ %
%£ % % : q ) ² % Á ³ % : such that . Thus, is a limit point of .
As for part 2 , if is closed, then by part 1 , any is the limit of a )):% M ² : ³
sequence in and so must be in . Hence, . Conversely, if ²% ³ : : M²:³ :
M ² : ³: : ² %³ : ² %³¦% , then is closed. For if is any sequence in and , then
there are two possibilities. First, we might have for some , in which %~ %
case . Second, we might have for all , in which case%~% : % £%
²% ³ ¦ % % M²:³ : % : : implies that . In either case, and so is closed
under the taking of limits, which implies that is closed. :
For part 3 , let . Clearly, . To show that is closed, we );~:rM ² : ³ :; ;
show that it contains all of its limit poi nts. So let . Hence, there is a %M ² ;³
sequence for which and . Of course, each is ²% ³ ; % £ % ²% ³ ¦ % %
either in , or is a limit point of . We must show that , that is, that is :: % ; %
either in or is a limit point of . ::
Suppose for the purposes of contradiction that and . Then there %¤: %¤M ² :³
is a ball for which . However, since , there )²%Á³ )²%Á³q: £ J ²% ³ ¦ %
must be an . Since cannot be in , it must be a limit point of . % ) ² % Á ³ % : :
Referring to Figure 12.1, if , then consider the ball ² %Á% ³~
)²% Á²c³°³ )²%Á³ . This ball is completely contained in and must contain
an element of , since its center is a limit point of . But then &: % :
&:q) ² % Á ³ %: %M ² :³ , a contradiction. Hence, or . In either case,
%;~:rM ² :³ ; and so is closed.
Thus, is closed and contains and so . On the other hand,;: ² : ³ ; cl
; ~:rM²:³ ²:³ ²:³~; cl cl and so .
308 Advanced Linear Algebra
Figure 12.1
For part 4 , if , then there are two possibilities. If , then the )c l% ² :³ %:
constant sequence , with for all , is a sequence in that converges ²% ³ % ~ % % :
to . If , then and so there is a sequence in for which%% ¤ : % M ² : ³ ² % ³:
%£ % ² % ³ ¦ % : % and . In either case, there is a sequence in converging to .
Conversely, if there is a sequence in for which , then either ²% ³ : ²% ³ ¦ %
%~ % % : ² : ³ %£ % for some , in which case , or else for all , in cl
which case . %M²:³ ²:³ cl
Dense Subsets
The following concept is meant to convey the idea of a subset being :4
“arbitrarily close” to every point in . 4
Definition A subset of a metric space is in if . A :4 4 ² : ³ ~ 4 dense cl
metric space is said to be if it contains a dense subset. separable countable
Thus, a subset of is dense if every open ball about any point :4 % 4
contains at least one point of . :
Certainly, any metric space contains a dense subset, namely, the space itself.
However, as the next examples show, not every metric space contains a
countable dense subset.
Example 12.6
1 The real line is separable, since the rational numbers form a countable) sr
dense subset. Similarly, is separable, since the set is countable and sr
dense.
2 The complex plane is separable, as is for all .) dd
3 A discrete metric space is separable if and only if it is countable. We leave)
proof of this as an exercise.
Metric Spaces 309
Example 12.7 The space is not separable. Recall that is the set of all MMBB
bounded sequences of real numbers or complex numbers with metric ()
² Á ³ ~ % c&%& sup
((
To see that this space is not separable, consider the set of all binary sequences :
:~¸ ² %³%~ ¹ or for all
This set is in one-to-one correspondence with the set of all subsets of and so o
is uncountable. It has cardinality 2 . Now, each sequence in is (LL ³ :
certainly bounded and so lies in . Moreover, if , then the two M£ MBB%&
sequences must differ in at least one position and so . ²%Á&³ ~
In other words, we have a subset of that is uncountable and for which the :MB
distance between any two distinct elements is . This implies that the balls in the
uncountable collection are mu tually disjoint. Hence, no ¸)² Á°³ :¹
countable set can meet every ball, which implies that no countable set can be
dense in . MB
Example 12.8 The metric spaces are separable, for . The set of all M :
sequences of the form
~² ÁÃÁ ÁÁó
for all , where the 's are rational, is a countable set. Let us show that it is
dense in . Any satisfies M% M
((
~B
% B
Hence, for any , there exists an such that 5
((
~5bB
%
Since the rational numbers are dense in , we can find rational numbers for s
which
((%c 5
for all . Hence, if , then~ÁÃÁ5 ~² ÁÃÁ ÁÁó 5
² % Á ³ ~ % c b %b~
~5B
~5b(( ( (
which shows that there is an element of arbitrarily close to any element of . :M
Thus, is dense in and so is separable.:M M
310 Advanced Linear Algebra
Continuity
Continuity plays a central role in the study of linear operators on infinite-
dimensional inner product spaces.
Definition Let be a function from the metric space to the¢4 ¦4 ²4Á³Z
metric space . We say that is if for any , ²4 Á ³ % 4 ZZ continuous at
there exists a such that
²%Á% ³ ¬ ²²%³Á²% ³³ Z
or, equivalently,
)²% Á ³ )²²% ³Á ³45
()See Figure 12.2. A function is if it is continuous at every continuous
% 4 .
Figure 12.2
We can use the notion of convergence to characterize continuity for functions
between metric spaces.
Theorem 12.4 A function is continuous if and only if whenever ¢4 ¦4Z
²% ³ 4 % 4 ²²% ³³ is a sequence in that converges to , then the sequence
converges to , in short, ²% ³
²% ³ ¦ % ¬ ²²% ³³ ¦ ²% ³
Proof. Suppose first that is continuous at and let . Then, given % ² % ³ ¦ %
, the continuity of implies the existence of a such that
²)²% Á ³³ )²²% ³Á ³
Since , there exists an such that for and² %³¦% 5 % ) ² %Á ³ 5
so
5 ¬ ²% ³ )²²% ³Á ³
Thus, .²% ³¦²% ³
Conversely, suppose that implies . Suppose, for the ² %³¦% ² ² %³ ³¦² %³
purposes of contradiction, that is not continuous at . Then there exists an %
Metric Spaces 311
such that for all ,
) ² % Á³ \ )²²% ³Á ³ 45
Thus, for all ,
)% Á \ )²²% ³Á ³
4567
and so we may construct a sequence by choosing each term with the ²% ³ %
property that
% )% Á ² % ³ ¤ ) ² ² % ³ Á³
67 , but
Hence, , but does not converge to . This contradiction²% ³ ¦ % ²% ³ ²% ³
implies that must be continuous at . %
The next theorem says that the distan ce function is a continuous function in both
variables.
Theorem 12.5 Let be a metric space. If and , then²4Á³ ²% ³ ¦ % ²& ³ ¦ &
²% Á& ³ ¦ ²%Á&³ .
Proof. We leave it as an exercise to show that
((²% Á& ³c²%Á&³ ²% Á%³b²& Á&³
But the right side tends to as and so . ¦B ² %Á&³¦ ² % Á& ³
Completeness
The reader who has studied analysis will recognize the following definitions.
Definition A sequence in a metric space is a if for ²% ³ 4 Cauchy sequence
any , there exists an for which 5
Á 5 ¬ ²% Á% ³
We leave it to the reader to show that any convergent sequence is a Cauchy
sequence. When the converse holds, the space is said to be . complete
Definition Let be a metric space.4
1 is said to be if every Cauchy sequence in converges in .)44 4 complete
2 A subspace of is if it is complete as a metric space. Thus, ) :4 : complete
is complete if every Cauchy sequence in converges to an element in ² ³ :
:.
Before considering examples, we prove a very useful result about completeness
of subspaces.
312 Advanced Linear Algebra
Theorem 12.6 Let be a metric space.4
1 Any complete subspace of is closed.) 4
2 If is complete, then a subspace of is complete if and only if it is)4: 4
closed.
Proof. To prove 1 , assume that is a complete subspace of . Let be a ) :4 ² % ³
sequence in for which . Then is a Cauchy sequence in :² % ³ ¦ % 4 ² % ³ :
and since is complete, must converge to an element of . Since limits of :² % ³ :
sequences are unique, we have . Hence, is closed. %: :
To prove part 2 , first assume that is complete. Then part 1 shows that is )) ::
closed. Conversely, suppose that is closed and let be a Cauchy sequence :² % ³
in . Since is also a Cauchy sequence in the complete space , it must:² % ³ 4
converge to some . But since is closed, we have . Hence, %4 : ² %³¦%:
: is complete.
Now let us consider some examples of complete and incomplete metric spaces. ()
Example 12.9 It is well known that the metric space is complete. However, a s (
proof of this fact would lead us outside the scope of this book. Similarly, the ³
complex numbers are complete. d
Example 12.10 The Euclidean space and the unitary space are complete. sd
Let us prove this for . Suppose that is a Cauchy sequence in , where ss²% ³
% ~²% ÁÃÁ% ³ Á Á
Thus,
²% Á% ³ ~ ²% c% ³ ¦ Á ¦ B Á Á
~
as
and so, for each coordinate position ,
²% c% ³ ²% Á% ³ ¦ Á Á
which shows that the sequence of th coordinates is a Cauchy ²% ³ Á ~Á ÁÃ 2
sequence in . Since is complete, we must have ss
²% ³ ¦ & ¦ BÁ as
If , then&~²& ÁÃÁ& ³
²% Á&³ ~ ²% c& ³ ¦ ¦ B Á
~
as
and so . This proves that is complete.²% ³ ¦ & ss
Metric Spaces 313
Example 12.11 The metric space of all real-valued or complex- ²*´ÁµÁ³ (
valued continuous functions on , with metric ) ´Áµ
²Á³ ~ ²%³c²%³ sup
%´Áµ((
is complete. To see this, we first observe that the limit with respect to is the
uniform limit on , that is if and only if for any , there is ´Áµ ² Á³¦
an for which5
5 ¬ ² % ³c² % ³ %´ Á µ for all ((
Now let be a Cauchy sequence in . Thus, for any , there is² ³ ²*´ÁµÁ³
an for which5
Á 5 ¬ ²%³c ²%³ % ´Áµ (( for all 12.1 ()
This implies that, for each , the sequence is a Cauchy sequence %´ Á µ ² ² % ³ ³
of real or complex numbers and so it converges. We can therefore define a ()
function on by ´ Á µ
²%³~ ²%³ lim
¦B
Letting in 12.1 , we get¦B ()
5¬ ² % ³c² % ³ %´ Á µ (( for all
Thus, converges to uniformly. It is well known that the uniform ²%³ ²%³
limit of continuous functions is continuous and so . Thus, ²%³*´Áµ
² ²%³³ ¦ ²%³ *´Áµ ²*´ÁµÁ³ and so is complete.
Example 12.12 The metric space of all real-valued or complex- ²*´ÁµÁ ³ (
valued continuous functions on , with metric ) ´Áµ
²²%³Á²%³³ ~ ²%³c²%³ %
((
is not complete. For convenience, we take and leave the general ´Áµ ~ ´Áµ
case for the reader. Consider the sequence of functions whose graphs are ² % ³
shown in Figure 12.3. The definition of should be clear from the graph. ( ² % ³ ³
314 Advanced Linear Algebra
Figure 12.3
We leave it to the reader to show that the sequence is Cauchy, but does ² ²%³³
not converge in . The sequence converges to a function that is not ²*´ÁµÁ ³ (
continuous.)
Example 12.13 The metric space is complete. To see this, suppose that M² % ³B
is a Cauchy sequence in , where MB
% ~²% Á% Áó Á Á 2
Then, for each coordinate position , we have
(( ((%c % %c % ¦ Á ¦ BÁ Á Á Á
sup as 12.2 ()
Hence, for each , the sequence of th coordinates is a Cauchy sequence in ² % ³ Á
sd sd or . Since or is complete, we have() ()
²% ³ ¦ & ¦ BÁ as
for each coordinate position . We want to show that and that & ~ ² & ³ M B
²% ³ ¦ & .
Letting in 12.2 gives¦B ² )
sup
Á as 12.3((%c & ¦ ¦ B ()
and so, for some ,
((%c & Á for all
and so
(( ( (& b % Á for all
But since , it is a bounded sequence and therefore so is . That is, % M ² & ³ B
& ~ ²& ³ M ² ²% ³ ¦ & M B B. Since 12.3 implies that , we see that is )
complete.
Metric Spaces 315
Example 12.14 The metric space is complete. To prove this, let be a M² % ³
Cauchy sequence in , where M
% ~²% Á% Áó Á Á 2
Then, for each coordinate position ,
(( (( %c % %c % ~ ² % Á % ³ ¦ Á Á Á Á
~B
which shows that the sequence of th coordinates is a Cauchy sequence in ²% ³ Á
sd sd or . Since or is complete, we have() ()
²% ³ ¦ & ¦ BÁ as
We want to show that and that . &~² &³M ² %³¦&
To this end, observe that for any , there is an for which 5
Á 5 ¬ % c% ((
~
Á Á
for all . Now we let , to get ¦B
5¬ % c& ((
~
Á
for all . Letting , we get, for any , ¦B 5
((
~B
Á %c &
which implies that and so and in ²% ³c& M & ~ & c²% ³b²% ³ M
addition, . ²% ³ ¦ &
As we will see in the next chapter, th e property of completeness plays a major
role in the theory of inner product spaces. Inner product spaces for which the
induced metric space is complete are called . Hilbert spaces
Isometries
A function between two metric spaces that preserves distance is called an
isometry. Here is the formal definition.
Definition Let and be metric spaces. A function is²4Á³ ²4 Á ³ ¢4 ¦ 4ZZ Z
called an if isometry
²²%³Á²&³³ ~ ²%Á&³Z
316 Advanced Linear Algebra
for all . If is a bijective isometry from to , we say%Á& 4 ¢4 ¦ 4 4 4ZZ
that and are and write .44 4 4ZZisometric
Theorem 12.7 Let be an isometry. Then¢²4Á³ ¦ ²4 Á ³ZZ
1 is injective)
2 is continuous)
3 is also an isometry and hence also continuous.)¢ ² 4 ³ ¦ 4c
Proof. To prove 1 , we observe that )
²%³ ~ ²&³ ¯ ²²%³Á²&³³ ~ ¯ ²%Á&³ ~ ¯ % ~ &Z
To prove 2 , let in . Then )²% ³ ¦ % 4
²²% ³Á²%³³ ~ ²% Á%³ ¦ ¦ BZ
as
and so , which proves that is continuous. Finally, we have²²% ³³ ¦ ²%³
² ²²%³³Á ²²&³³ ~ ²%Á&³ ~ ²²%³Á²&³³c c Z
and so is an isometry.¢ ² 4 ³ ¦ 4c
The Completion of a Metric Space
While not all metric spaces are complete, any metric space can be embedded in
a complete metric space. To be more specific, we have the following important
theorem.
Theorem 12.8 Let be any metric space. Then there is a complete metric²4Á³
space and an isometry for which is dense in² 4Á³ ¢4¦ 44 4ZZ Z
4² 4 Á ³ ² 4 Á ³ZZ Z. The metric space is called a of . Moreover, completion
²4 Á ³ZZ is unique, up to bijective isometry.
Proof. The proof is a bit lengthy, so we divide it into various parts. We can
simplify the notation considerably by thinking of sequences in as ²% ³ 4
functions , where . ¢ ¦4 ²³~%o
Cauchy Sequences in 4
The basic idea is to let the elements of be equivalence classes of Cauchy 4Z
sequences in . So let denote the set of all Cauchy sequences in . If 4² 4 ³ 4 CS
Á ²4³ ²³ CS , then, intuitively sp eaking, the terms get closer together as
¦B ² ³ and so do the terms . Therefore, it seems reasonable that
²²³Á²³³ ¦ B should approach a finite limit as . Indeed, since
((²²³Á²³³c²²³Á²³³ ²²³Á²³³b²²³Á²³³ ¦
as it follows that is a Cauchy sequence of realÁ ¦ B ²²³Á²³³
numbers, which implies that
Metric Spaces 317
lim
¦B²²³Á²³³ B ()12.4
(That is, the limit exists and is finite. ³
Equivalence Classes of Cauchy Sequences in 4
We would like to define a metric on the set by ² 4 ³ZCS
²Á³ ~ ²²³Á²³³Z
¦Blim
However, it is possible that
lim
¦B²²³Á²³³ ~
for distinct sequences and , so this does not define a metric. Thus, we are led
to define an equivalence relation on by CS²4³
¯ ²²³Á²³³ ~ lim
¦B
Let be the set of all equivalence classes of Cauchy sequences andCS²4³
define, for , Á ²4³ CS
²Á³ ~ ²²³Á²³³Z
¦Blim ( ) 12.5
where and .
To see that is well-defined, suppose that and . Then since ZZ Z
ZZ and , we have
((² ²³Á ²³³c²²³Á²³³ ² ²³Á²³³b² ²³Á²³³ ¦ ZZ Z Z
as . Thus,¦B
¬ ² ²³Á ²³³ ~ ²²³Á²³³
¬ ² Á ³ ~ ² Á ³ZZ Z Z
¦B ¦B
ZZZ Z and lim lim
which shows that is well-defined. To see that is a metric, we verify the ZZ
triangle inequality, leaving the rest to the reader. If and are Cauchy Á
sequences, then
²²³Á²³³ ²²³Á²³³b²²³Á²³³
Taking limits gives
lim lim lim
¦B ¦B ¦B²²³Á²³³ ²²³Á²³³b ²²³Á²³³
318 Advanced Linear Algebra
and so
² Á ³² Á ³b² Á ³ZZZ
Embedding in ²4Á³ ²4 Á ³ZZ
For each , consider the constant Cauchy sequence , where % 4 ´%µ ´%µ²³ ~ %
for all . The map defined by¢ 4 ¦ 4 Z
%~´ % µ
is an isometry, since
² %Á &³ ~ ²´%µÁ´&µ³ ~ ²´%µ²³Á´&µ²³³ ~ ²%Á&³ZZ
¦B lim
Moreover, is dense in . This follows from the fact that we can 44Z
approximate any Cauchy sequence in by a constant sequence. In particular, 4
let . Since is a Cauchy sequence, for any , there exists an 4 5Z
such that
Á 5 ¬ ²²³Á²³³
Now, for the constant sequence we have ´²5³µ
´²5³µÁ ~ ²²5³Á²³³Z
¦B45 lim
and so is dense in .44Z
²4 Á ³ZZ Is Complete
Suppose that
Á Á ÁÃ 3
is a Cauchy sequence in . We wish to find a Cauchy sequence in for 4 4Z
which
² Á³ ~ ² ²³Á²³³ ¦ ¦ BZ
¦Blim as
Since and since is dense in , there is a constant sequence 4 4 4ZZ
´ µ~² Á Áó
for which
² Á´ µ³
Z
Metric Spaces 319
We can think of as a constant approximation to , with error at most . °
Let be the sequence of these constant approximations:
²³ ~
This is a Cauchy sequence in . Intu itively speaking, since the 's get closer4
to each other as , so do the constant approximations. In particular, we ¦B
have
² Á ³~² ´ µ Á ´ µ ³
²´ µÁ ³b ² Á ³b ² Á´ µ³
b ² Á ³ b¦
Z
ZZ Z
Z
as . To see that converges to , observe thatÁ ¦B
² Á³ ² Á´ µ³b ²´ µÁ³ b ² Á²³³
~b ² Á ³
ZZ Z
¦B
¦Blim
lim
Now, since is a Cauchy sequence, for any , there is an such that 5
Á5 ¬² Á ³
In particular,
5¬ ² Á³ lim
¦B
and so
5¬² Á ³ b
Z
which implies that , as desired. ¦
Uniqueness
Finally, we must show that if and are both completions of ²4 Á ³ ²4 Á ³Z Z ZZ ZZ
²4Á³ 4 4 , then . Note that we have bijective isometriesZZ Z
¢4¦ 44 ¢4¦ 44ZZ Z and
Hence, the map
~¢ 4 ¦ 4c
is a bijective isometry from onto , where is dense in . See 44 4 4 ²Z
Figure 12.4. ³
320 Advanced Linear Algebra
Figure 12.4
Our goal is to show that can be ex tended to a bijective isometry from to 4Z
4ZZ.
Let . Then there is a sequence in for which . Since%4 ² ³ 4 ² ³¦%Z
² ³ 4 ² ² ³³ 4 4ZZ is a Cauchy sequence in , is a Cauchy sequence in
and since is complete, we have for some . Let us 4 ² ² ³³ ¦ & & 4ZZ ZZ
define .²%³ ~ &
To see that is well-defined, suppose that and , where both ² ³ ¦ % ² ³ ¦ %
sequences lie in . Then 4
² ² ³Á ² ³³ ~ ² Á ³ ¦ ¦ BZZ Z
as
and so and converge to the same element of , which implies²² ³ ³ ²² ³ ³ 4ZZ
that does not depend on the choice of sequence in converging to .²%³ 4 %
Thus, is well-defined. Moreover, if , then the constant sequence 4 ´ µ
converges to and so lim , which shows that is an ²³ ~ ²³ ~ ²³
extension of .
To see that is an isometry, suppose that and . Then ² ³ ¦ % ² ³ ¦ &
² ² ³³ ¦ ²%³ ² ² ³³ ¦ ²&³ ZZ and and since is continuous, we have
² ²%³Á ²&³³ ~ ² ² ³Á ² ³³ ~ ² Á ³ ~ ²%Á&³ZZ ZZ Z Z
¦B ¦B lim lim
Thus, we need only show that is surjective. Note first that
4~ ²³ ²³ ²³ im im im . Thus, if is closed, we can deduce from the fact
that is dense in that . So, suppose that is a 4 4 ²³ ~ 4 ²² %³ ³ZZ ZZ im
sequence in and . Then is a Cauchy sequence and im²³ ²² %³ ³ ¦ ' ²² %³ ³
therefore so is . Thus, . But is continuous and so ²% ³ ²% ³ ¦ % 4Z
²² %³ ³ ¦ ² % ³ ² % ³ ~ ' ' ²³ , which implies that and so . Hence, is im
surjective and . 4 4ZZ Z
Metric Spaces 321
Exercises
1. Prove the generalized triangle inequality
² %Á %³ ² %Á %³b ² %Á %³bÄb ² % Á %³ c 3
2. a Use the triangle inequality to prove that )
((²%Á&³c²Á³ ²%Á³b²&Á³
b Prove that )
((²%Á'³c²&Á'³ ²%Á&³
3. Let be the subspace of all binary sequences sequences of 's and:M B(
:'s . Describe the metric on .)
4. Let be the set of all binary -tuples. Define a function 4 ~ ¸Á¹
¢: d: ¦ ²%Á&³ % s by letting be the number of positions in which and
& ´²³Á²³µ ~ differ. For example, . Prove that is a metric. It (
is called the and plays an important role in Hamming distance function
the theory of error-correcting codes. ³
5. Let .B
a If show that )%~² %³M % ¦
b Find a sequence that converges to but is not an element of any for ) M
B .
6. a Show that if , then for all . ) %%~² %³M M
b Find a sequence that is in for , but is not in . ) %~² %³ M M
7. Show that a subset of a metric space is open if and only if contains :4 :
an open neighborhood of each of its points.
8. Show that the intersection of any co llection of closed sets in a metric space
is closed.
9. Let be a metric space. The of a nonempty subset ²4Á³ : 4 diameter
is
²:³ ~ ²%Á&³ sup
%Á&:
A set is if .:² : ³ Bbounded
a Prove that is bounded if and only if there is some and ) :% 4 s
for which . :) ² % Á ³
b Prove that if and only if consists of a single point. ) ²:³ ~ :
c Prove that implies . ) :; ² : ³ ² ;³
d If and are bounded, show that is also bounded. ):; : r ;
10. Let be a metric space. Let be the function defined by²4Á³ Z
² % Á& ³~²%Á&³
b²%Á&³Z
322 Advanced Linear Algebra
a Show that is a metric space and that is bounded under this ) ²4Á ³ 4Z
metric, even if it is not bounded under the metric .
b Show that the metric spaces and have the same open ) ²4Á³ ²4Á ³Z
sets.
11. If and are subsets of a metric space , we define the :; ² 4 Á ³ distance
between and by :;
²:Á;³ ~ ²%Á&³ inf
%:Á!;
a Is it true that if and only if ? Is a metric? ) ²:Á;³ ~ : ~ ;
b Show that if and only if . )c l % ²:³ ²¸%¹Á:³ ~
12. Prove that is a limit point of if and only if every %4 :4
neighborhood of meets in a poi nt other than itself. %: %
13. Prove that is a limit point of if and only if every open ball %4 :4
)²%Á³ : contains infinite ly many points of .
14. Prove that limits are unique, that is, , implies that ²% ³ ¦ % ²% ³ ¦ &
%~& .
15. Let be a subset of a metric space . Prove that if and only if:4 % ² : ³ cl
there exists a sequence in that converges to . ²% ³ : %
16. Prove that the closure has the following properties:
a )c l: ² : ³
b )c l c l²² : ³ ³ ~ :
c )c l c l c l²: r;³ ~ ²:³r ²;³
d )c l c l c l²: q;³ ²:³q ²;³
Can the last part be strengthened to equality?
17. a Prove that the closed ball is always a closed subset. ) )²%Á³
b Find an example of a metric space in which the closure of an open ball )
)²%Á³ )²%Á³ is not equal to the closed ball .
18. Provide the details to show that is separable. s
19. Prove that is separable. d
20. Prove that a discrete metric space is separable if and only if it is countable.
21. Prove that the metric space of all bounded functions on , with 8´Áµ ´Áµ
metric
²Á³ ~ ²%³c²%³ sup
%´Áµ((
is not separable.
22. Show that a function is continuous if and only if the ¢²4Á³ ¦ ²4 Á ³ZZ
inverse image of any open set is open, that is, if and only if
² < ³ ~ ¸ % 4 ² % ³ < ¹ 4 <c is open in whenever is an open set
in .4Z
23. Repeat the previous exercise, replacing the word open by the word closed.
24. Give an example to show that if is a continuous ¢²4Á³¦²4Á³ZZ
function and is an open set in , it need not be the case that is <4 ² < ³
open in .4Z
Metric Spaces 323
25. Show that any convergent sequence is a Cauchy sequence.
26. If in a metric space , show that any subsequence of ²% ³ ¦ % 4 ²% ³ ²% ³
also converges to . %
27. Suppose that is a Cauchy sequence in a metric space and that some ²% ³ 4
subsequence of converges. Prove that converges to the ²% ³ ²% ³ ²% ³
same limit as the subsequence.
28. Prove that if is a Cauchy sequence, then the set is bounded. What ²% ³ ¸% ¹
about the converse? Is a bounded sequence necessarily a Cauchy sequence?
29. Let and be Cauchy sequences in a metric space . Prove that the²% ³ ²& ³ 4
sequence converges. ~ ² % Á & ³
30. Show that the space of all convergent sequences of real numbers or ²
complex numbers is complete as a subspace of . ) MB
31. Let denote the metric space of all polynomials over , with metricFd
²Á³ ~ ²%³c²%³ sup
%´Áµ((
Is complete?F
32. Let be the subspace of all sequences with finite support that is, :MB(
with a finite number of nonzero terms . Is complete? ):
33. Prove that the metric space of all integers, with metric {
²Á³ ~ c (( , is complete.
34. Show that the subspace of the metric space under the sup metric :* ´ Á µ ()
consisting of all functions for which is complete. *´Áµ ²³ ~ ²³
35. If and is complete, show that is also complete.44 4 4ZZ
36. Show that the metric spaces and , under the sup metric, are *´Áµ *´Áµ
isometric.
37. Prove Ho ¨lder's inequality
(( ( ( ( (89 89
~ ~ ~BB B
° °
%& % &
as follows:
a Show that ) ~! ¬!~ c c
b Let and be positive real numbers and consider the rectangle in )"# 9
s with corners , , and , with area . Argue ²Á³ ²"Á³ ²Á#³ ²"Á#³ "#
geometrically that is, draw a picture to show that ()
"# ! !b
"#
c c
and so
"# b"#
c Now let and . Apply the results of ) ?~ % B @~ & B''(( ((
part b to )
324 Advanced Linear Algebra
"~ Á #~%&
?@(( ((
° °
and then sum on to deduce Ho ¨lder's inequality.
38. Prove Minkowski's inequality
89 8 9 8 9 (( ( ( ( (
~ ~ ~BB B
° ° °
%b & % b &
as follows:
a Prove it for first. ) ~
b Assume . Show that )
(( ( ( (( ( ( ((%b & % %b & b& %b & c c
c Sum this from to and apply Ho ¨lder's inequality to each sum on ) ~
the right, to get
((
H8 9 8 9 I8 9 (( (( ( (~
~ ~ ~
° ° °%b &
% b & % b &
Divide both sides of this by the last factor on the right and let to ¦B
deduce Minkowski's inequality.
39. Prove that is a metric space. M
Chapter 13
Hilbert Spaces
Now that we have the necessary background on the topological properties of
metric spaces, we can resume our study of inner product spaces without
qualification as to dimension. As in Chapter 9, we restrict attention to real and
complex inner product spaces. Hence will denote either or . - sd
A Brief Review
Let us begin by reviewing some of the results from Chapter 9. Recall that an
inner product space over is a vector space , together with an inner =- =
product . If , then the inner product is bilinear and ifºÁ»¢= d= ¦ - - ~ s
-~d, the inner product is sesquilinear.
An inner product induces a norm on , defined by =
)) j#~ º # Á # »
We recall in particular the following properties of the norm.
Theorem 13.1
1 For all ,)( )The Cauchy-Schwarz inequality "Á# =
(( ) ) ) )º"Á#» " #
with equality if and only if for some . "~ # -
2 For all ,)( )The triangle inequality "Á# =
)) ) ) ) )"b# " b #
with equality if and only if for some . "~ # -
3)( )The parallelogram law
)) ))) ) ) )"b# b "c# ~ " b #
We have seen that the inner product can be recovered from the norm, as follows.
326 Advanced Linear Algebra
Theorem 13.2
1 If is a real inner product space, then)=
º"Á#»~ ² "b# c "c# ³
)) ))
2 If is a complex inner product space, then)=
º"Á#»~ ² "b# c "c# ³b ² "b# c "c# ³
)) )) ) ) ) )
The inner product also induces a metric on defined by =
²"Á#³ ~ "c# ))
Thus, any inner product space is a metric space.
Definition Let and be inner product spaces and let .=> ² = Á > ³ B
1 is an if it preserves the inner product, that is, if) isometry
º" Á# » ~ º " Á # »
for all ."Á# =
2 A bijective isometry is called an . When ) isometric isomorphism ¢= ¦>
is an isometric isomorphism, we say that and are => isometrically
isomorphic .
It is easy to see that an isometry is always injective but need not be surjective,
even if .=~ >
Theorem 13.3 A linear transformation is an isometry if and only B² = Á > ³
if it preserves the norm, that is, if and only if
)) ) )#~#
for all .#=
The following result points out one of the main differences between real and
complex inner product spaces.
Theorem 13.4 Let be an inner product space and let .= ² = ³ B
1 If for all , then .)º# Á $ » ~ # Á$ = ~
2 If is a complex inner product space and for all)= 8² # ³~º # Á# »~
#= ~ , then .
3 Part 2 does not hold in general for real inner product spaces.))
Hilbert Spaces
Since an inner product space is a metric space, all that we learned about metric
spaces applies to inner product spaces. In particular, if is a sequence of ²% ³
Hilbert Spaces 327
vectors in an inner product space , then =
²% ³ ¦ % % c% ¦ ¦ B if and only if as ))
The fact that the inner product is continuous as a function of either of its
coordinates is extremely useful.
Theorem 13.5 Let be an inner product space. Then=
1)² %³¦% Á² &³¦&¬º %Á&»¦º % Á& »
2)²% ³ ¦ % ¬ % ¦ % )) ) )
Complete inner product spaces play an especially important role in both theory
and practice.
Definition An inner product space that is complete under the metric induced by
the inner product is said to be a . Hilbert space
Example 13.1 One of the most important examples of a Hilbert space is the
space . Recall that the inner product is defined byM
ºÁ» ~ %&%&
~B
(In the real case, the conjugate is unnecessary. The metric induced by this inner ³
product is
² Á ³ ~ c ~ % c&%& % & )) ( (89
~B
°2
which agrees with the definition of the metric space given in Chapter 12. In M
other words, the metric in Chapter 12 is induced by this inner product. As we
saw in Chapter 12, this inner product space is complete and so it is a Hilbert
space. In fact, it is the prototype of all Hilbert spaces, introduced by David (
Hilbert in 1912, even before the axiomatic definition of Hilbert space was given
by John von Neumann in 1927. ³
The previous example raises the question whether the other metric spaces M
² £ ), with distance given by
² Á ³ ~ c ~ % c&%& % & )) ( (89
~B
°
13.1 ()
are complete inner product spaces. The fact is that they are not even inner
product spaces! More specifically, there is no inner product whose induced
metric is given by 13.1 . To see this, observe that, according to Theorem 13.1, ()
328 Advanced Linear Algebra
any norm that comes from an inner product must satisfy the parallelogram law
)) )) ) ) ) )%& %& % &bb c~ b
But the norm in 13.1 does not satis fy this law. To see this, take ()
%&~ ²ÁÁÁó ~ ²ÁcÁÁó and . Then
)) ))%& %&b~ Ác~
and
)) ))%&° °~ Á ~
Thus, the left side of the parallelogram law is and the right side is , h 2°
which equals if and only if . ~
Just as any metric space has a completion, so does any inner product space.
Theorem 13.6 Let be an inner product space. Then there exists a Hilbert=
space and an isometry for which is dense in . Moreover, /¢ = ¦ / = / /
is unique up to isometric isomorphism.
Proof. We know that the metric space , where is induced by the inner ²= Á³
product, has a unique completion , which consists of equivalence classes ²= Á ³ZZ
of Cauchy sequences in . If and , then we = ² %³² %³= ² &³² &³= ZZ
set
²% ³b²& ³ ~ ²% b& ³Á ²% ³ ~ ²% ³
and
º²% ³Á²& ³» ~ º% Á& » ¦Blim
It is easy to see that since and are Cauchy sequences, so are ²% ³ ²& ³ ²% b& ³
and . In addition, these definitions are well-defined, that is, they are²% ³
independent of the choice of representative from each equivalence class. For
instance, if , then ²% ³ ²% ³V
lim
¦B))%c % ~ V
and so
(( ( ( ) ) ) )º% Á& »cº% Á& » ~ º% c% Á& » % c% & ¦ VVV
(The Cauchy sequence is bounded. Hence, ²& ³ ³
º²% ³Á²& ³» ~ º% Á& » ~ º% Á& » ~ º²% ³Á²& ³» VV ¦B ¦Blim lim
We leave it to the reader to show that is an inner product space under these =Z
operations.
Hilbert Spaces 329
Moreover, the inner product on induces the metric , since =ZZ
º ² %c &³ Á ² %c &³ » ~ º %c &Á %c &»
~ ² % Á & ³
~ ²²% ³Á²& ³³ ¦B
¦B
Z
lim
lim
Hence, the metric space isometry is an isometry of inner product ¢= ¦=Z
spaces, since
º %Á &» ~ ² %Á &³ ~ ²%Á&³ ~ º%Á&» Z
Thus, is a complete inner product space and is a dense subspace of == =Z Z
that is isometrically isomorphic to . We leave the issue of uniqueness to the=
reader.
The next result concerns subspaces of inner product spaces.
Theorem 13.7
1 Any complete subspace of an inner product space is closed.)
2 A subspace of a Hilbert space is a Hilbert space if and only if it is closed.)
3 Any finite-dimensional subspace of an inner product space is closed and )
complete.
Proof. Parts 1 and 2 follow from Theorem 12.6. Let us prove that a finite- ))
dimensional subspace of an inner product space is closed. Suppose that :=
²% ³ : ²% ³ ¦ % % ¤ : ~ ¸ ÁÃÁ ¹ is a sequence in , and . Let be an 8
orthonormal Hamel basis for . The Fourier expansion :
~ º % Á»
~
in has the property that but:% c £
º%c Á»~º%Á»cº Á»~
Thus, if we write and , the sequence , which is &~%c & ~% c : ² &³
in , converges to a vector that is orthogonal to . But this is impossible, :& :
because implies that& &
) ) ) ) )) ))& c &~ & b & &¦ °
This proves that is closed. :
To see that any finite-dimensional subspace of an inner product space is :
complete, let us embed as an inner product space in its own right in its :()
completion . Then or rather an isometric copy of is a finite-dimensional :: :Z()
330 Advanced Linear Algebra
subspace of a complete inner product space and as such it is closed. :Z
However, is dense in and so , which shows that is complete. :: : ~ : :ZZ
Infinite Series
Since an inner product space allows both addition of vectors and convergence of
sequences, we can define the concept of infinite sums, or infinite series.
Definition Let be an inner product space. The of the= th partial sum
sequence in is ²% ³ =
~ %b Ä b %
If the sequence of partial sums converges to a vector , that is, if ² ³ =
)) c ¦ ¦ B as
then we say that the series to and write % converges
~B
%~
We can also define absolute convergence.
Definition A series is said to be if the series % absolutely convergent
))
~B
%
converges.
The key relationship between convergence and absolute convergence is given in
the next theorem. Note that completeness is required to guarantee that absolute
convergence implies convergence.
Theorem 13.8 Let be an inner product space. Then is complete if and only==
if absolute convergence of a series implies convergence.
Proof. Suppose that is complete and that . Then the sequence =% B ))
of partial sums is a Cauchy sequence, for if , we have
)) ) )ii c ~ % % ¦
~b ~b
Hence, the sequence converges, that is, the series converges. ² ³ %
Conversely, suppose that absolute convergence implies convergence and let
²% ³ = be a Cauchy sequence in . We wish to show that this sequence
converges. Since is a Cauchy sequence, for each , there exists an ²% ³ 5
Hilbert Spaces 331
with the property that
Á 5 ¬ % c%
))
Clearly, we can choose , in which case 5 5 Ä
))%c %
55 b
and so
))
~ ~BB
55%c % B
b
Thus, according to hypothesis, the series
~B
55²% c% ³b
converges. But this is a telescoping series, whose th partial sum is
%c %55b
and so the subsequence converges. Since any Cauchy sequence that has a ²% ³5
convergent subsequence must itself converge, the sequence converges and ²% ³
so is complete.=
An Approximation Problem
Suppose that is an inner product space and that is a subset of . It is of =: =
considerable interest to be able to fi nd, for any , a vector in that is %= :
closest to in the metric induced by the inner product, should such a vector%
exist. This is the for . approximation problem =
Suppose that and let %=
~% c inf
:))
Then there is a sequence for which
~% c ¦))
as shown in Figure 13.1.
332 Advanced Linear Algebra
Figure 13.1
Let us see what we can learn about this sequence. First, if we let , &~ % c
then according to the parallelogram law,
)) )) ) ) ) )&b & b&c & ~ ²& b& ³
or
)) ) ) ) ) hh &c & ~ ²& b& ³ c &b &
()13.2
Now, if the set is , that is, if :convex
%Á& : ¬ %b²c³& : for all
()in words, contains the line segment between any two of its points , then :
² b ³° : and so
hh h h&b & b
~% c
Thus, 13.2 gives ()
)) ) ) ) )&c & ²& b& ³ c ¦
as . Hence, if is convex, then the sequence is aÁ ¦ B : ²& ³ ~ ²%c ³
Cauchy sequence and therefore so is . ² ³
If we also require that be complete, then the Cauchy sequence converges :² ³
to a vector and by the continuity of the norm, we must have . %: %c% ~VV ))
Let us summarize and add a remark about uniqueness.
Theorem 13.9 Let be an inner product space and let be a complete convex=:
subset of . Then for any , there exists a unique for which =% = % : V
)) ))%c% ~ %c V inf
:
The vector is called the to in . %% :V best approximation
Hilbert Spaces 333
Proof. Only the uniqueness remains to be established. Suppose that
)) ) )%c% ~ ~ %c%VZ
Then, by the parallelogram law,
)) ) )
)) ) ) ) )
)) ) ) hh%c% ~ ²%c% ³c²%c%³VV
~ %c% b %c% c %c%c% VV
~ %c% b %c% c %c V%b%V
b c ~ZZ
ZZ
ZZ
2
and so .%~%VZ
Since any subspace of an inner product space is convex, Theorem 13.9 :=
applies to complete subspaces. However, in this case, we can say more.
Theorem 13.10 Let be an inner product space and let be a complete=:
subspace of . Then for any , the best approximation to in is the=% = % :
unique vector for which . % : % c % :ZZ
Proof. Suppose that , where . Then for any , we have %c% : % : :ZZ
%c% c%ZZ and so
) ) )) )) ))%c ~ %c% b % c %c% ZZ Z
Hence is the best approximation to in . Now we need only show that%~ % % :VZ
%c%: % % : :VV , where is the best approximation to in . For any , a little
computation reminiscent of completing the square gives
))
)) ))
)) ))89)) ))
)) ))89 89)) )) ))((%c ~º%c Á%c »
~ % cº%Á »cº Á%»b
~ % b c cº%Á » º%Á »
~% b c c cº%Á » º%Á » º%Á »
2
~% b c cº%Á » º%Á »
)) ))ee)) ))((
Now, this is smallest when
~ º%Á »
))
334 Advanced Linear Algebra
in which case
)) ) )((
))%c ~ % cº%Á »
Replacing by gives %% c % V
)) ) )((
))%c%c ~ %c% cVVº%c%Á »V
But is the best approximation to in and since we must have%% : % c :VV
)) ) )%c%c %c%VV
Hence,
((
))º%c%Á »V
~
or equivalently,
º%c%Á » ~ V
Hence, .%c%:V
According to Theorem 13.9, if is a complete subspace of an inner product :
space , then for any , we may write=% =
%~%b² %c% ³VV
where and . Hence, and since ,% : %c% : = ~ : b: : q: ~ ¸¹VV
we also have . This is the projection theorem for arbitrary inner =~ :p :
product spaces.
Theorem 13.11 The projection theorem () If is a complete subspace of an:
inner product space , then =
=~ :p :
In particular, if is a closed subspace of a Hilbert space , then :/
/~:p:
Theorem 13.12 Let , and be subspaces of an inner product space .:; ; =Z
1 I f t h e n .)=~ :p ; ;~ :
2 I f t h e n .):p;~:p; ;~;ZZ
Proof. If , then by definition of orthogonal direct sum. On=~ :p ; ; :
the other hand, if , then , for some and . Hence, ': '~ b! : !;
~ º'Á » ~ º Á »bº!Á » ~ º Á »
Hilbert Spaces 335
and so , implying that . Thus, . Part 2 follows from part ~ '~!; : ;)
1.)
Let us denote the closure of the span of a set of vectors by . :² : ³ cspan
Theorem 13.13 Let be a Hilbert space./
1 If is a subset of , then)(/
cspan²(³ ~ (
2 If is a subspace of , then):/
cl²:³ ~ :
3 If is a closed subspace of , then)2/
2~2
Proof. We leave it as an exercise to show that . Hence ´² ( ³ µ ~ (cspan
/ ~ ²(³p´ ²(³µ ~ ²(³p( cspan cspan cspan
But since is closed, we also have (
/~( p(
and so by Theorem 13.12, . The rest follows easily from part cspan²(³ ~ (
1.)
In the exercises, we provide an example of a closed subspace of an inner 2
product space for which . Hence, we cannot drop the requirement =2 £ 2
that be a Hilbert space in Theorem 13.13./
Corollary 13.14 If is a of a Hilbert space , then is dense in(/ ² ( ³ subset span
/ ( ~ ¸¹ if and only if .
Proof. As in the previous proof,
/~ ² ( ³p( cspan
and so if and only if .( ~ ¸¹ / ~ ²(³cspan
Hilbert Bases
We recall the following definition from Chapter 9.
Definition A maximal orthonormal set in a Hilbert space is called a / Hilbert
basis for ./
Zorn's lemma can be used to show that any nontrivial Hilbert space has a Hilbert
basis. Again, we should mention that the concepts of Hilbert basis and Hamel
basis a maximal linearly independent set are quite different. We will show ()
336 Advanced Linear Algebra
later in this chapter that any two Hilbert bases for a Hilbert space have the same
cardinality.
Since an orthonormal set is maximal if and only if , Corollary EE~ ¸¹
13.14 gives the following characterization of Hilbert bases.
Theorem 13.15 Let be an orthonormal subset of a Hilbert space . TheE /
following are equivalent:
1 is a Hilbert basis)E
2)E~ ¸¹
3 is a of , that is, .)c s p a nEE total subset /² ³ ~ /
Part 3 of this theorem says that a subset of a Hilbert space is a Hilbert basis if )
and only if it is a total orthonormal set.
Fourier Expansions
We now want to take a cl oser look at best approximati ons. Our goal is to find an
explicit expression for the best approximation to any vector from within a %
closed subspace of a Hilbert space . We will find it convenient to consider :/
three cases, depending on whether has finite, countably infinite, or:
uncountable dimension.
The Finite-Dimensional Case
Suppose that is an orthonormal set in a Hilbert space . E~¸ "ÁÃÁ"¹ /
Recall that the Fourier expansion of any , with respect to , is given by %/ E
% ~ º%Á" »"V
~
where is the Fourier coefficient of with respect to . Observe thatº%Á" » % "
º%c%Á" »~º%Á" »cº%Á" »~VV
and so span . Thus, according to Theorem 13.9, the Fourier %c% ² ³VE
expansion is the best approximation to in . Moreover, since %% ² ³V spanE
%c%%VV , we have
)) )) ) ) ))%~ %c % c % %VV
and so
)) ))%%V
with equality if and only if , which happens if and only if .%~% % ² ³V spanE
Let us summarize.
Hilbert Spaces 337
Theorem 13.16 Let be a finite orthonormal set in a HilbertE~¸ "ÁÃÁ"¹
space . For any , the Fourier expansion of is the best/% / % % V
approximation to in . We also have %² ³spanE Bessel's inequality
)) ))%%V
or equivalently,
(( ) )
~
º%Á" » % ()13.3
with equality if and only if . % ² ³spanE
The Countably Infinite-Dimensional Case
In the countably infinite case, we will be dealing with infinite sums and so
questions of convergence will arise. Thus, we begin with the following.
Theorem 13.17 Let be a countably infinite orthonormal set inE~¸ "Á "Á Ã ¹
a Hilbert space . The series /
~B
" ()13.4
converges in if and only if the series /
((
~B
()13.5
converges in . If these series converge, then they converge unconditionally s
(that is, any series formed by rearranging the order of the terms also
converges . Finally, if the series 13.4 converges, then )( )
ii ((
~ ~BB
" ~
Proof. Denote the partial sums of the first series by and the partial sums of
the second series by . Then for
)) ( ( ((ii c ~ " ~ ~c
~b ~b
Hence is a Cauchy sequence in if and only if is a Cauchy sequence² ³ / ² ³
in . Since both and are complete, converges if and only if ss /² ³ ² ³
converges.
If the series 13.5 converges, then it converges absolutely and hence ()
unconditionally. A real series converges unconditionally if and only if it (
338 Advanced Linear Algebra
converges absolutely. But if 13.5 converges unconditionally, then so does ³ ()
()13.4 . The last part of the theorem fo llows from the continuity of the norm.
Now let be a countably infinite orthonormal set in . TheE~¸ "Á "Á Ã ¹ /
Fourier expansion of a vector is defined to be the sum %/
% ~ º%Á" »"V
~B
()13.6
To see that this sum converges, observe that for any , 13.3 gives ()
(( ) )
~
º%Á" » %
and so
(( ) )
~B
º%Á" » %
which shows that the series on the left converges. Hence, according to Theorem
13.17, the Fourier expansion 13.6 converges unconditionally. ()
Moreover, since the inner product is continuous,
º%c%Á" »~º%Á" »cº%Á" »~VV
and so . Hence, is the best approximation%c%´ ² ³µ ~´ ² ³µ %VV span cspanEE
to in . Finally, since , we again have%² ³ % c % % VV cspanE
)) )) ) ) ))%~ %c % c % %VV
and so
)) ))%%V
with equality if and only if , which happens if and only if .%~% % ² ³V cspanE
Thus, the following analog of Theorem 13.16 holds.
Theorem 13.18 Let be a countably infinite orthonormal set inE~¸ "Á "Á Ã ¹
a Hilbert space . For any , the Fourier expansion /% /
% ~ º%Á" »"V
~B
of converges unconditionally and is the best approximation to in .%% ² ³ cspanE
We also have Bessel's inequality
)) ))%%V
Hilbert Spaces 339
or equivalently,
(( ) )
~B
º%Á" » %
with equality if and only if . % ² ³cspanE
The Arbitrary Case
To discuss the case of an arbitrary orthonormal set , let us E~¸ " 2¹
first define and discuss the concept of the sum of an arbitrary number of terms.
(This is a bit of a digression, since we could proceed without all of the coming
details but they are interesting.c³
Definition Let be an arbitrary family of vectors in an innerA~¸ % 2¹
product space . The sum =
2%
is said to to a vector and we write converge %=
%~ %
2 ()13.7
if for any , there exists a finite set for which :2
;: Á; ¬ % c% finiteii
;
For those readers familiar with the language of convergence of nets, the set
F²2³ 2 of all finite subsets of is a under inclusion for every directed set (
(Á) ²2³ * ²2³ ( )FF there is a containing and and the function )
:¦ %
:
is a net in . Convergence of 13.7 is convergence of this net. In any case, we /² )
will refer to the preceding definition as the of convergence. net definition
It is not hard to verify the follo wing basic properties of net convergence for
arbitrary sums.
Theorem 13.19 Let be an arbitrary family of vectors in anA~¸ % 2¹
inner product space . If =
2 2%~ % &~ & and
then
340 Advanced Linear Algebra
1)( )Linearity
2²% b & ³ ~ %b &
for any Á -
2)( )Continuity
2 2º% Á&»~º%Á&» º&Á% »~º&Á%» and
The next result gives a useful “Cau chy-type” description of convergence.
Theorem 13.20 Let be an arbitrary family of vectors in anA~¸ % 2¹
inner product space . =
1 I f t h e s u m)
2%
converges, then for any , there exists a finite set such that 02
1q0~J Á1 ¬ % finiteii
1
2 If is a Hilbert space, then the converse of 1 also holds.))=
Proof. For part 1 , given , let , finite, be such that ) :2 :
;: Á; ¬ % c% finite2ii
;
If , finite, then1q:~J1
ii i i
ii ii11
::
1r: :%~ ²% b% c % ³ c ²% c % ³
% c% b % c% b ~22
As for part 2 , for each , let be a finite set for which ) 0 2
1q0 ~J Á1 ¬ %
1 finiteii
and let
&~ %
0
Hilbert Spaces 341
Then is a Cauchy sequence, since²& ³
))ii i i
ii ii&c & ~ %c % ~ %c %
% b% b ¦
00 0 c 0 0 c 0
0c 0 0c 0
Since is assumed complete, we have .=² & ³ ¦ &
Now, given , there exists an such that 5
5¬ & c& ~ % c& ))ii
0
2
Setting gives for finite,~ ¸ 5Á °¹ ;0Á; max
ii i i
ii i i;
0; c 0
0; c 0%c &~ %c & b %
% c & b% b
and so converges to . 2%&
The following theorem tells us that conve rgence of an arbitrary sum implies that
only countably many terms can be nonzero so, in some sense, there is no such
thing as a nontrivial sum. uncountable
Theorem 13.21 Let be an arbitrary family of vectors in anA~¸ % 2¹
inner product space . If the sum =
2%
converges, then at most a countable number of terms can be nonzero. %
Proof. According to Theorem 13.20, for each , we can let , 0 2 0
finite, be such that
1q0 ~J Á1 ¬ %
1 finiteii
Let . Then is countable and0~ 0 0
¤0¬¸ ¹q0 ~J ¬ % ¬% ~
for all for all ))
342 Advanced Linear Algebra
Here is the analog of Theorem 13.17.
Theorem 13.22 Let be an arbitrary orthonormal family ofE~¸ " 2¹
vectors in a Hilbert space . The two series /
((
2 2 " and
converge or diverge together. If these series converge, then
ii ((
2 2
" ~
Proof. The first series converges if and only if for every , there exists a
finite set such that 02
1q0~J Á1 ¬ " finiteii
1
or equivalently,
1q0~J Á1 ¬ finite ((
1
and this is precisely what it means for the second series to converge. We leave
proof of the remaining statement to the reader.
The following is a useful characterization of arbitrary sums of nonnegative real
terms.
Theorem 13.23 Let be a collection of nonnegative real numbers.¸ 2¹
Then
2 1~ sup
1
12finite()13.8
provided that either of the preceding expressions is finite.
Proof. Suppose that
sup
1
12 finite
1~ 9 B
Then, for any , there exists a finite set such that :2
9 9c
:
Hilbert Spaces 343
Hence, if is a finite set for which , then since , ;2 ;:
9 9 c
;
:
and so
ii9c
;
which shows that converges to . Finally, if the sum on the left of 13.8 9 ()
converges, then the supremum on the right is finite and so 13.8 holds. ()
The reader may have noticed that we have two definitions of convergence for
countably infinite series: the net version and the traditional version involving
the limit of partial sums. Let us write
~B
ob%% and
for the net version and the partial sum version, respectively. Here is the
relationship between these two definitions.
Theorem 13.24 Let be a Hilbert space. If , then the following are/% /
equivalent:
1 converges net version to )( )
ob%%
2 converges unconditionally to )
~B
%%
Proof. Assume that 1 holds. Suppose that is any permutation of . Given ) ob
any , there is a finite set for whicho :b
;: Á; ¬ % c% finiteii
;
Let us denote the set of integers by and choose a positive integer ¸ÁÃÁ¹ 0
such that . Then for we have²0 ³ :
²0 ³ ²0 ³ : ¬ % c% ~ % c%
~
²³
²0 ³ii
and so 2 holds. )
344 Advanced Linear Algebra
Next, assume that 2 holds, but that the series in 1 does not converge. Then ))
there exists an such that for any finite subset , there exists a finite o 0b
subset with for which11 q 0 ~ J
ii
1%
From this, we deduce the existence of a countably infinite sequence of 1
mutually disjoint finite subsets of with the property that ob
max min²1 ³ ~ 4 ~ ²1 ³ b b
and
ii
1
%
Now we choose any permutation with the following properties o o¢¦bb
1)²´ Á4 µ³ ´ Á4 µ
2 if , then)1 ~¸ ÁÃÁ ¹ Á Á "
² ³~ Á ² b³~ ÁÃÁ ² b" c³~ Á Á Á " 2
The intention in property 2 is that for each , takes a set of consecutive )
integers to the integers in . 1
For any such permutation , we have
ii i i
~ 1b "c
²³
%~%
which shows that the sequence of partial sums of the series
~B
²³%
is not Cauchy and so this series doe s not converge. This contradicts 2 and )
shows that 2 implies at least that 1 converges. But if 1 converges to , )) ) &/
then since 1 implies 2 and since unconditional limits are unique, we have ))
&~% . Hence, 2 implies 1 . ))
Now we can return to the discussion of Fourier expansions. Let
E~¸ " 2¹ / be an arbitrary orthonormal set in a Hilbert space . Given
any , we may apply Theorem 13.16 to all finite subsets of , to deduce%/ E
Hilbert Spaces 345
that
sup
1
12 finite(( ) )
1º%Á" » %
and so Theorem 13.23 tells us that the sum
((
2º%Á" »
converges. Hence, according to Theorem 13.22, the Fourier expansion
% ~ º%Á" »"V
2
of also converges and%
)) ( ( %~ º % Á " »V
2
Note that, according to Theorem 13.21, is a countably infinite sum of terms of %V
the form and so is in . º%Á" »" ² ³ cspanE
The continuity of infinite sums with respect to the inner product Theorem (
13.19 implies that )
º%c%Á" »~º%Á" »cº%Á" »~VV
and so span cspan . Hence, Theorem 3.9 tells us that %c%´ ² ³µ ~´ ² ³µ %VVEE
is the best approximation to in . Finally, since , we again %² ³ % c % % VV cspanE
have
)) )) ) ) ))%~ %c % c % %VV
and so
)) ))%%V
with equality if and only if , which happens if and only if .%~% % ² ³V cspanE
Thus, we arrive at the most general form of a key theorem about Hilbert spaces.
Theorem 13.25 Let be an orthonormal family of vectors inE~¸ " 2¹
a Hilbert space . For any , the Fourier expansion /% /
% ~ º%Á" »"V
2
of converges in and is the unique best approximation to in .%/ % ² ³ cspanE
Moreover, we have Bessel's inequality
)) ))%%V
346 Advanced Linear Algebra
or equivalently,
(( ) )
2º%Á" » %
with equality if and only if . % ² ³cspanE
A Characterization of Hilbert Bases
Recall from Theorem 13.15 that an orthonormal set in a E~¸ " 2¹
Hilbert space is a Hilbert basis if and only if /
cspan²³ ~ /E
Theorem 13.25, then leads to the following characterization of Hilbert bases.
Theorem 13.26 Let be an orthonormal family in a HilbertE~¸ " 2¹
space . The following are equivalent:/
1 is a Hilbert basis a maximal orthonormal set)( )E
2)E~ ¸¹
3 is total that is, )( c s p a n )EE ²³ ~ /
4 for all )%~% %/V
5 Equality holds in Bessel's inequality for all , that is,) %/
)) ))%~%V
for all %/
6)Parseval's identity
º%Á&» ~ º%Á&» VV
holds for all , that is, %Á& /
º%Á&» ~ º%Á" »º&Á" »
2
Proof. Parts 1 , 2 and 3 are equivalent by Theorem 13.15. Part 4 implies part )) ) )
3 , since and 3 implies 4 since the unique best approximation of)c s p a n ) )% ² ³VE
any is itself and so . Parts 3 and 5 are equivalent by% ² ³ %~% V cspan ) )E
Theorem 13.25. Parseval's identity fo llows from part 4 using Theorem 13.19. )
Finally, Parseval's identity for imp lies that equality holds in Bessel's &~%
inequality.
Hilbert Dimension
We now wish to show that all Hilbert bases for a Hilbert space have the same /
cardinality and so we can define the Hilbert dimension of to be that /
cardinality.
Hilbert Spaces 347
Theorem 13.27 All Hilbert bases for a Hilbert space have the same /
cardinality. This cardinality is called the of , which we Hilbert dimension /
denote by . hdim²/³
Proof. If has a finite Hilbert basis, then that set is also a Hamel basis and so/
all finite Hilbert bases have size . Suppose next that dim²/³ ~ ¸ 2¹ 8
and are infinite Hilbert bases for . Then for each , we have9~¸ 1¹ /
~ º Á »
1
where is the countable set . Moreover, since no can be1¸ º Á » £ ¹
orthogonal to every , we have . Thus, since each is countable, 1 ~ 1 1 2
we have
( ( (( ((ee1~ 1 L2~2
2
By symmetry, we also have and so the Schro ¨der–Bernstein theorem (( ( (21
implies that . (( ( (1~2
Theorem 13.28 Two Hilbert spaces are isometrically isomorphic if and only if
they have the same Hilbert dimension.
Proof. Suppose that . Let be a hdim hdim²/ ³ ~ ²/ ³ ~ ¸" 2¹ E
Hilbert basis for and a Hilbert basis for . We may /~ ¸ # 2 ¹ / E
define a map as follows: ¢/ ¦/
45
2 2 " ~ #
We leave it as an exercise to verify th at is a bijective isometry. The converse
is also left as an exercise.
A Characterization of Hilbert Spaces
We have seen that any vector space is isomorphic to a vector space of =² - ³)
all functions from to that have finite support. There is a corresponding )-
result for Hilbert spaces. Let be any nonempty set and let 2
M² 2 ³~ ¢ 2¦ ² ³ B
2Dc E ((d
The functions in are referred to as . We can M² 2 ³ ²square summable functions
also define a real version of this set by replacing by . We define an inner ds³
product on by M² 2 ³
ºÁ»~ ²³²³
2
The proof that is a Hilbert space is quite similar to the proof that M² 2 ³
348 Advanced Linear Algebra
M~ M ²³o is a Hilbert space and the details are left to the reader. If we define
M² 2³ by
Á ²³ ~ ~ ~
£ Fif
if
then the collection
E~¸ 2¹
is a Hilbert basis for , of cardinality . To see this, observe that M² 2 ³ 2((
ºÁ» ~ ² ³² ³ ~ Á
2
and so is orthonormal. Moreover, if , then for only aE M² 2 ³ ² ³£
countable number of , say . If we define by 2 ¸ Á Á Ã ¹ Z
~ ² ³Z
~B
then and for all , which implies that . ²³ ² ³ ~ ² ³ 2 ~ ZZ ZcspanE
This shows that and so is a total orthonormal set, that is, a M² 2 ³~ ² ³cspanEE
Hilbert basis for . M² 2 ³
Now let be a Hilbert space, with Hilbert basis . We define /~ ¸ " 2 ¹ 8
a map as follows. Since is a Hilbert basis, any has the8¢/¦M ²2³ %/
form
% ~ º%Á" »"
2
Since the series on the right converges, Theorem 13.22 implies that the series
((
2º%Á" »
converges. Hence, another application of Theorem 13.22 implies that the
following series converges:
²%³ ~ º%Á" »
2
It follows from Theorem 13.19 that is linear and it is not hard to see that it is
also bijective. Notice that and so takes the Hilbert basis for 8²" ³ ~ /
to the Hilbert basis for . EM² 2 ³
Hilbert Spaces 349
Notice also that
)) ( ( ) ) ii ²%³ ~ º ²%³Á ²%³» ~ º%Á" » ~ º%Á" »" ~ %
2 2
and so is an isometric isomorphism. We have proved the following theorem.
Theorem 13.29 If is a Hilbert space of Hilbert dimension and if is any/2
set of cardinality , then is isometrically isomorphic to . /M ² 2 ³
The Riesz Representation Theorem
We conclude our discussion of Hilbert spaces by discussing the Riesz
representation theorem. As it happens, not all linear functionals on a Hilbert
space have the form “take the inner product with ,” as in the finite- Ã
dimensional case. To see this, observe that if , then the function &/
²%³ ~ º%Á&»&
is certainly a linear functional on . However, it has a special property. In /
particular, the Cauchy–Schwarz inequality gives, for all , %/
(( (( ) ) ) ) ²%³ ~ º%Á&» % &&
or, for all , %£
((
))))² % ³
%&&
Noticing that equality holds if , we have %~&
sup
%£&((
))))² % ³
%~&
This prompts us to make the following definition, which we do for linear
transformations between Hilbert spaces this covers the case of linear (
functionals . )
Definition Let be a linear transformation from to . Then is¢/ ¦/ / /
said to be if bounded
sup
%£))
))%
%B
If the supremum on the left is finite, we denote it by and call it the of )) norm
.
350 Advanced Linear Algebra
Of course, if is a bounded linear functional on , then ¢/ ¦- /
))((
))~²%³
%sup
%£
The set of all bounded linear functionals on a Hilbert space is called the /
continuous dual space conjugate space , or , of and denoted by . Note //i
that this differs from the algebraic dua l of , which is the set of all linear/
functionals on . In the finite-dimensional case, however, since all linear /
functionals are bounded exercise , the two concepts agree. Unfortunately, () (
there is no universal agreement on the notation for the algebraic dual versus the
continuous dual. Since we will discuss only the continuous dual in this section,
no confusion should arise. ³
The following theorem gives some simp le reformulations of the definition of
norm.
Theorem 13.30 Let be a bounded linear transformation.¢/ ¦/
1))) ) )~%sup
))%~
2))) ) )~%sup
))%
3 for all ))) ) ) ))s ~¸ % % % / ¹inf
The following theorem explains the importance of bounded linear
transformations.
Theorem 13.31 Let be a linear transformation. The following are¢/ ¦/
equivalent:
1 is bounded)
2 is continuous at any point ) % /
3 is continuous.)
Proof. Suppose that is bounded. Then
)) ) ) ) ) ) ) %c % ~ ²%c%³ %c% ¦
as . Hence, is continuous at . Thus, 1 implies 2 . If 2 holds, then%¦% % )) )
for any , we have&/
)) ) ) %c & ~ ²%c&b%³c ²%³ ¦
as , since is continuous at and as . Hence, %¦& % %c&b% ¦% &¦%
is continuous at any and 3 holds. Finally, suppose that 3 holds. Thus, &/ ))
is continuous at and so there exists a such that
)) ) )% ¬ %
Hilbert Spaces 351
In particular,
))))
))%~ ¬ %
%
and so
)) ) ))) ) )
)) ) )%~ ¬ %~ ¬ ¬ ²% ³ %
%%
Thus, is bounded.
Now we can state and prove the Riesz representation theorem.
Theorem 13.32 The Riesz representation theorem () Let be a Hilbert/
space. For any bounded linear functional on , there is a unique / ' /
such that
²%³ ~ º%Á' »
for all . Moreover, .%/ ' ~ )) ) )
Proof. If , we may take , so let us assume that . Hence,~ ' ~ £
2~ ² ³£/ 2 ker and since is continuous, is closed. Thus
/~2p2
Now, the first isomorphism theorem, applied to the linear functional , ¢/ ¦-
implies that as vector spaces . In addition, Theorem 3.5 implies that /°2 - ()
/°2 2 2 - ²2 ³~ and so . In particular, . dim
For any , we have'2
%2¬² % ³~~º % Á'»
Since , all we need do is find for which dim²2 ³ ~ £ ' 2
²'³~º'Á'»
for then for all , showing that ²'³ ~ ²'³ ~ º'Á'» ~ º'Á'» -
²%³ ~ º%Á'» % 2 for as well.
But if , then£'2
'~ '²'³
º'Á'»
has this property, as can be easily checked. The fact that has )) ) )'~
already been established.
352 Advanced Linear Algebra
Exercises
1. Prove that the sup metric on the metric space of continuous *´Áµ
functions on does not come from an inner product. Hint: let ´Áµ ²!³ ~
and a a and consider the parallelogram law.²!³ ~ ²!c ³°² c ³
2. Prove that any Cauchy sequence that has a convergent subsequence must
itself converge.
3. Let be an inner product space and let and be subsets of . Show=( ) =
that
a )()¬) (
b is a closed subspace of )(=
c )c s p a n´² ( ³ µ ~ (
4. Let be an inner product space and . Under what conditions is =: =
:~ : ?
5. Prove that a subspace of a Hilbert space is closed if and only if :/
:~:.
6. Let be the subspace of consisting of all sequences of real numbers =M
with the property that each sequence has only a finite number of nonzero
terms. Thus, is an inner product space. Let be the subspace of =2 =
consisting of all sequences in with the property that %~² %³ =
'%° ~ 2 2 £2. Show that is closed, but that . Hint: For the latter,
show that by considering the sequences , 2 ~¸¹ "~²ÁÃÁcÁó
where the term is in the th coordinate position. c
7. Let be an orthonormal set in . If converges,E'~¸ "Á "Á Ã ¹ / %~ "
show that
)) ( ( %~
~B
8. Prove that if an infinite series
~B
%
converges absolutely in a Hilbert space , then it also converges in the /
sense of the “net” definition given in this section.
9. Let be a collection of nonnegative real numbers. If the sum¸ 2¹
on the left below converges, show that
2 1~ sup
1
12 finite
10. Find a countably infinite sum of real numbers that converges in the sense of
partial sums, but not in the sense of nets.
11. Prove that if a Hilbert space has infinite Hilbert dimension, then no /
Hilbert basis for is a Hamel basis. /
12. Prove that is a Hilbert space for any nonempty set . M² 2 ³ 2
Hilbert Spaces 353
13. Prove that any linear transformation between finite-dimensional Hilbert
spaces is bounded.
14. Prove that if , then is a closed subspace of . / ² ³ /iker
15. Prove that a Hilbert space is separable if and only if . hdim²/³ L
16. Can a Hilbert space have countably infinite Hamel dimension?
17. What is the Hamel dimension of ? M² ³o
18. Let and be bounded linear operators on . Verify the following: /
a ))) ( ( ) )~
b ))) ) ) ) ) b b
c ))) ) ) ) )
19. Use the Riesz representation theorem to show that for any Hilbert / /i
space ./
Chapter 14
Tensor Products
In the preceding chapters, we have seen several ways to construct new vector
spaces from old ones. Two of the most important such constructions are the
direct sum and the vector space of all linear transformations <l= ² <Á=³ B
from to . In this chapter, we consid er another very important construction, <=
known as the . tensor product
Universality
We begin by describing a general ty pe of that will help motivate the universality
definition of tensor product. Our descrip tion is strongly related to the formal
notion of a in category theory, but we will be somewhat less universal pair
formal to avoid the need to formally define categorical concepts. Accordingly,
the terminology that we shall introduce is not standard, but does not contradict
any standard terminology.
Referring to Figure 14.1, consider a set and two functions and , with (
domain .(
A
gWS
Xf
Figure 14.1
Suppose that there exists a function for which this diagram ¢:¦?
commutes , that is,
~ k
This is sometimes expressed by saying that can be . What factored through
does this say about the relationship between the functions and ?
356 Advanced Linear Algebra
Let us think of the “information” a bout contained in a function as( ¢ ( ¦ )
the way in which elements of using from . The ( )distinguishes labels
relationship above implies that
²³ £ ²³ ¬ ²³ £ ²³
and this can be phrased by saying that whatever ability has to distinguish
elements of is also possessed by . Put another way, except for labeling (
differences, any information about that is contained in is also contained in (
.
If happens to be injective, then the difference between and is the only
values of the labels. That is, the two functions have the same information about
(. However, in general, is not required to be injective and so may contain
more information than .
Now consider a family of sets and a family I
<I~¸ ¢(¦?? ¹
Assume that and . If the diagram in Figure 14.1 commutes : ¢(¦:I<
for all , then the information contained in every function in is also <<
contained in . Moreover, since , the function cannot contain more <
information than is contained in the en tire family and so we conclude that
contains exactly the same information as is contained in the entire family . In <
this sense, is among all functions in . ¢(¦: ¢(¦? universal <
In this way, a single function , or more precisely, a single pair , ¢(¦: ²:Á³
can capture a mathematical concept as described by a family of functions. Some
examples from linear algebra are basis for a vector space, quotient space, direct
sum and bilinearity (as we will see).
Let us make a formal definition.
Definition Referring to Figure 14.2, let be a set and let be a family of sets. ( I
Let
<I~¸ ¢(¦?? ¹
be a family of functions, all of whic h have domain and range a member of . ( I
Let
> I~¸ ¢?¦@ ?Á@ ¹
be a family of functions with domain and range in . We assume that has theI>
following structure:
1 contains the identity function for each member of .)> I :
Tensor Products 357
2 is closed under composition of functions, which is an associative)>
operation.
3 For any and , the composition is defined and belongs to) > < k
<.
A
S3f3S2f2S1
f1
W3W2W1
Figure 14.2
We refer to as the and its members as > measuring family measuring
functions .
A pair , where and has the for²:Á¢(¦:³ : I< universal property
the family , or is a for , if for every<> < > as measured by universal pair ²Á³
¢( ¦ ? ¢: ¦ ? in , there is a unique in for which the diagram in< >
Figure 14.1 commutes, that is, for which
~ k
or equivalently, any can be . The unique measuring < factored through
function is called the for . mediating morphism
Note the requirement that the mediati ng morphism be unique. Universal pairs
are essentially unique, as the following describes.
Theorem 14.1 Let and be universal pairs for²:Á¢(¦:³ ²;Á¢(¦;³
²Á³ : ~ ;<> > . Then there is a bijective measuring function for which .
In fact, the mediating morphism of with respect to and the mediating
morphism of with respect to are isomorphisms.
Proof. With reference to Figure 14.3, there are mediating morphisms ¢:¦;
and for which¢; ¦:
~ k
~ k
Hence,
~² k ³k
~² k ³k
However, referring to the third diagram in Figure 14.3, both and k¢ : ¦ :
the identity map are mediating morphisms for and so the uniqueness ¢:¦:
358 Advanced Linear Algebra
of mediating morphisms implies that . Similarly and so and k~ k~
are inverses of one another, making the desired bijection.
A
gWS
Tf
Ag
V
ST
fAf
VW L
SS
f
Figure 14.3
Examples of Universality
Now let us look at some examples of the universal property. Let denote Vect²-³
the family of all vector spaces over the base field . We use the term -( family
informally to represent what in set theory is formally referred to as a class. A
class is a “collection” that is too large to be considered a set. For example,
Vect )²-³ is a class.
Example 14.1 Let be a nonempty set and let()Bases8
1)I~² - ³Vect
2) set functions from to members of <8 <~
3) linear transformations>~
If is a vector space with basis , then the pair , where is=² = Á ¢ ¦ = ³ 88 8 88
the inclusion map , is universal for . To see this, note that the # ~ # ² Á ³ <>
condition that can be factored through , <
~ k
is equivalent to the statement that for each basis vector . But this 8#~ # #
uniquely defines a linear transformation .
In fact, the universality of the pa ir is the statement that a linear²= Á³8 precisely
transformation is uniquely determined by assigning its values arbitrarily on a
basis , the function doing the arbitrary a ssignment in this context. Note also 8
that Theorem 14.1 implies that if is also universal for , ²>Á¢ ¦ >³ ² Á ³8< >
then there is a bijective mediating mo rphism from to , that is, and => > =88
are isomorphic.
Example 14.2 Let be a vector()Quotient spaces and canonical projections =
space and let be a subspace of . Let 2=
1)I~² - ³Vect
Tensor Products 359
2) linear maps with domain , whose kernels contain <~= 2
3) linear transformations>~
Theorem 3.4 says precisely that the pair , where is the ²= °2Á ¢= ¦ = °2³
canonical projection map, has the univers al property for as measured by . <>
Example 14.3 Let and be vector spaces over . Let()Direct sums <= -
1)I~² - ³Vect
2) ordered pairs of linear transformations<~ ²¢< ¦>Á¢= ¦>³
3) linear transformations>~
Here we have a slight variation on the definition of universal pair: In this case,
< > < is a family of of functions. For and , we set pairs ² Á ³
k²Á³~² kÁ k³
Then the pair , where ²< = Á² Á ³¢²<Á= ³ ¦ < = ³^^
"~² " Á ³ #~² Á # ³ and
are called the , has the universal property for . To canonical injections ²Á³<>
see this, observe that for any pair in , the condition ²Á³¢²<Á= ³ ¦ > <
² Á ³~ k² Á ³
is equivalent to
²Á³ ~ ² k Á k ³
or
²"Á³~²"³ ²Á#³~²#³ and
But these conditions define a unique linear transformation . ^¢< = ¦>
Thus, bases, quotient spaces and direct sums are all examples of universal pairs
and it should be clear from these exampl es that the notion of universal property
is, well, universal. In fact, it happens th at the most useful definition of tensor
product is through a universal property, which we now explore.
Bilinear Maps
The universality that defines tensor products rests on the notion of a bilinear
map.
Definition Let , and be vector spaces over . Let be the<= > - < d =
cartesian product of and . A set function <= as sets
¢< d= ¦>
360 Advanced Linear Algebra
is if it is linear in both va riables separately, that is, if bilinear
²"b "Á#³~²"Á#³b ²"Á#³ZZ
and
²"Á#b #³~²"Á#³b ²"Á#³ZZ
The set of all bilinear functions from to is denoted by <d= >
hom-²<Á=Â>³ ¢<d= ¦- . A bilinear function w ith values in the base
field is called a on .-< d = bilinear form
Note that bilinearity can also be expressed in matrix language as follows: If
~² ÁÃÁ ³- Á ~² ÁÃÁ ³-
and
"~²" ÁÃÁ" ³< Á #~²# ÁÃÁ# ³=
then is bilinear if¢< d= ¦>
²"Á#³~-!! !
where .- ~ ´²" Á# ³µ Á
It is important to emphasize that, in th e definition of bilinear function, is <d=
the , not the direct product of vector spaces. In othercartesian product of sets
words, we do not consider any algebraic structure on when defining <d=
bilinear functions, so expressions like
²%Á&³b²'Á$³ ²%Á&³ and
are meaningless.
In fact, if is a vector space, there are two classes of functions from to == d =
> ² =d =Á > ³ =d =~= =: the linear maps , where is the direct B^
product of vector spaces, and the bilinear maps , where is hom²=Á=Â>³ = d=
just the cartesian product of sets. We leave it as an exercise to show that these
two classes have only the zero map in common. In other words, the only map
that is both linear and bilinear is the zero map.
We made a thorough study of bilinear forms on a finite-dimensional vector
space in Chapter 11 although this material is not assumed here . However,= ()
bilinearity is far more important and far-reaching than its application to metric
vector spaces, as the following examples show. Indeed, both multiplication and
evaluation are bilinear.
Example 14.4 If is an algebra, the product map()Multiplication is bilinear (
¢(d(¦( defined by
Tensor Products 361
²Á³ ~
is bilinear, that is, multiplication is linear in each position.
Example 14.5 If and are vector spaces, then the()Evaluation is bilinear =>
evaluation map defined byB¢² = Á > ³ d = ¦ >
²Á#³ ~ #
is bilinear. In particular, the evaluation map defined by ¢= d= ¦-i
²Á#³ ~ # = d= is a bilinear form on .i
Example 14.6 If and are vector spaces, and and , then the=> = >ii
product map defined by ¢= d> ¦-
²#Á$³ ~ ²#³²$³
is bilinear. Dually, if and , then the map #= $> ¢= d> ¦- ii
defined by
²Á³ ~ ²#³²$³
is bilinear.
It is precisely the tensor product that will allow us to generalize the previous
example. In particular, if and , then we would like B B² < Á > ³ ² = Á > ³
to consider a “product” map defined by ¢<d= ¦>
²"Á#³ ~ ²"³ ²#³ ?
The tensor product is just the thing to replace the question mark, because it n
has the desired bilinearity property, as we will see. In fact, the tensor product is
bilinear and nothing else, so it is what we need! exactly
Tensor Products
Let and be vector spaces. Our guide for the definition of the tensor product<=
<n= will be the desire to have a universal property for bilinear functions, as
measured by linearity. Referring to Figure 14.4, we want to define a vector
space and a bilinear map so that any bilinear map with;! ¢ < d = ¦ ;
domain can be factored through . Intuitively speaking, is the most<d= ! !
“general” or “universal” bilinear ma p with domain : It is bilinear <d= and
nothing more .
362 Advanced Linear Algebra
Wfbilinear
linearWT VUt
bilinear
Figure 14.4
Definition Let be the cartesian product of two vector spaces over . Let<d= -
I~² - ³Vect . Let
<I~ ¸ ²<Á=Â>³> ¹
>-hom
be the family of all bilinear maps from to any vector space . The <d= >
measuring family is the fam ily of all linear transformations. >
A pair is if it is universal for²;Á!¢<d= ¦;³ universal for bilinearity
²Á³ ¢ < d = ¦ ><> , that is, if for every bilinear map , there is a unique
linear transformation for which ¢; ¦>
~ k!
The map is called the for . mediating morphism
We can now define the tensor product via this universal property.
Definition Let and be vector spaces over a field . Any universal pair<= -
²;Á!¢<d= ¦;³ < = for bilinearity is called a of and . The tensor product
vector space is denoted by and sometimes referred to by itself as the ;< n =
tensor product. The map is called the and the elements of !< n = tensor map
are called . tensors
It is customary to use the symbol to denote the image of any ordered pair n
²"Á#³ under the tensor map, that is,
"n#~!²"Á#³
for any and . A tensor of the form is said to be "< #= "n#
decomposable , that is, the decomposable tensors are the images under the
tensor map.
Since universal pairs are unique up to isomorphism, we may refer to “the”
tensor product of vector spaces. Note also that the tensor product is not a n
product in the sense of a binary operation on a set. In fact, even when , =~ <
the tensor product is not in , but rather in . "n" < <n<
Tensor Products 363
As we will see, there are other, more c onstructive ways to define the tensor
product. Since we have adopted the universal pair definition, the other ways to
define tensor product are, for us, constructions rather than definitions. Let us
examine some of these constructions.
Construction I: Intuitive but Not Coordinate Free
The universal property for bilinearity capture s the essence of bilinearity and the
tensor map is the most “general” b ilinear function on . To see how this <d=
universality can be achieved in a constructive manner, let be a basis ¸ 0¹
for and let be a basis for . Then a bilinear map on is<¸ 1 ¹ = ! < d =
uniquely determined by assigning arb itrary values to the “basis” pairs ² Á ³
and extending by bilinearity, that is, if and , then "~ #~
!²"Á#³ ~ ! Á ~ !² Á ³ 45
Now, the tensor map , being the most general bilinear map, must do this ! and
nothing more . To achieve this goal, we define the tensor map on the pairs !
²Á³ !²Á³ in such a way that the images , and then extend do not interact
by bilinearity.
In particular, for each ordered pair , we invent a new formal symbol, ² Á ³
written , and define to be the vector space with basisn ;
:~¸ n 0Á1¹
The tensor map is defined by setting and extending by !² Á ³ ~ n
bilinearity. Thus,
!²"Á#³ ~ ! Á ~ ² n ³ 45
To see that the pair is the tensor product of and , if is ²;Á!³ < = ¢<d= ¦>
bilinear, the universality condition is equivalent to ~ k!
² n ³ ~ ² Á ³
which does indeed uniquely define a map . Hence, has linear ¢; ¦> ²;Á!³
the universal property for bilinearity and so we can write and refer ;~<n=
to as the tensor map.!
Note that while the set is a basi s for (by definition), the set :~¸ n¹ ;
¸ "n#"<Á#=¹
of decomposable tensors spans , but is not linearly independent. This does ;
cause some initial confusion during the learning process. For example, one
cannot define a linear map on by assigning values arbitrarily to the <n=
decomposable tensors, nor is it alw ays easy to tell when a tensor is "n #
364 Advanced Linear Algebra
equal to . We will consider the latter issue in some detail a bit later in the
chapter.
The fact that is a basis for gives the following. : <n=
Theorem 14.2 For finite-dimensional vector spaces and , <=
dim dim dim²< n= ³ ~ ²<³h ²= ³
Construction II: Coordinate Free
The previous construction of the tensor product is reasonably intuitive, but has
the disadvantage of not being coordinate free. The following approach does not
require the choice of a basis.
Let be the vector space over with basis . Let be the subspace-- < d = :<d=
of generated by all vectors of the form-<d=
²"Á$³b ²#Á$³c²"b #Á$³ ()14.1
and
²"Á#³b ²"Á$³c²"Á#b $³ ()14.2
where and and are in the appropriate spaces. Note that theseÁ - "Á# $
vectors are precisely what we must “identify” as the zero vector in order to
enforce bilinearity. Put another way, these vectors are if the ordered pairs are
replaced by tensors according to our previous construction.
Accordingly, the quotient space
<n=-
:<d=
is also sometimes taken as the definition of the tensor product of and . <=
(Strictly speaking, we should not be using the symbol until we have <n=
shown that this is the tensor product.) The elements of have the form <n=
45 ! ² "Á#³ b:~ ² "Á#³b:
However, since and , we can absorb ²"Á#³c²"Á#³: ²"Á#³c²"Á#³:
the scalar in either coordinate, that is,
´²"Á#³b:µ~²"Á#³b:~²"Á#³b:
and so the elements of can be written simply as <n=
´²" Á# ³b:µ
It is customary to denote the coset by , and so any element of ²"Á#³b: "n#
Tensor Products 365
<n= has the form
"n #
as in the previous construction.
The tensor map is defined by ! ¢<d=¦<n=
! ² " Á# ³~"n#~² " Á# ³b:
This map is bilinear, since
!²"b#Á$³~²"b #Á$³b:
~´²"Á$³b ²#Á$³µb:
~´ ² " Á$ ³b:µb´ ² # Á$ ³b:µ
~ !²"Á$³b !²#Á$³
and similarly for the second coordinate.
We next prove that the pair is universal for ²<n=Á!¢<d= ¦<n=³
bilinearity when is defined as a quotient space . <n= - ° : <d=
Theorem 14.3 Let and be vector spaces. The pair<=
²<n=Á!¢<d= ¦<n=³
is the tensor product of and . <=
Proof. Consider the diagram in Figure 14.5. Here is the vector space with -<d=
basis .<d=
FVU
VS j
Wf WVU VUt
Figure 14.5
Since
k²"Á#³ ~ ²"Á#³ ~ ²"Á#³b: ~ "n# ~ !²"Á#³
we have
!~ k
The universal property of vector spaces described in Example 14.1 implies that
366 Advanced Linear Algebra
there is a unique linear transformation for which ¢- ¦><d=
k~
Note that sends the vectors (14.1) and (14.2) that generate to the zero vector :
and so . For example,: ² ³ ker
´²"Á$³b ²#Á$³c²"b #Á$³µ
~ ´²"Á$³b ²#Á$³c²"b #Á$³µ
~ ²"Á$³b ²#Á$³c ²"b #Á$³
~²"Á$³b ²#Á$³c²"b #Á$³
~
and similarly for the second coordinate. Hence, Theorem 3.4 the universal (
property described in Example 14.2) implies that there exists a unique linear
transformation for which ¢<n= ¦>
k~
Hence,
k!~ k k~ k~
As to uniqueness, if , then Zk!~
Z´²"Á#³b:µ~²"Á#³~ ´²"Á#³b:µ
and since the cosets generate , we conclude that . ²"Á#³b: - °: ~ <d=Z
Thus, is the mediating morphism and is universal for bilinearity. ²< n= Á!³
Let us take a moment to compare the two previous constructions. Let
¸ 0¹ ¸ 1¹ < = ²; Á! ³ZZ and be bases for and , respectively. Let be
the tensor product as constructed using these two bases and let
²;Á!³~²- °:Á!³ <d= be the tensor product construction using quotient spaces.
Since both of these pairs are universal for bilinearity, Theorem 14.1 implies that
the mediating morphism for with respect to , that is, the map !! ¢ ; ¦ ;ZZ
defined by
² n ³ ~ ² Á ³b:
is a vector space isomorphism. Therefore, the basis of is sent to ¸² n ³¹ ;Z
the set , which is therefore a basis for .¸² Á ³b:¹ ;
In other words, given any two bases and for and , ¸ 0¹ ¸ 1¹ < =
respectively, the tensors form a basis for , regardless of which n <n =
construction of the tensor product we use. Therefore, we are free to think of
n <n = either as a formal symbol belonging to a basis for or as the coset
² Á ³b: < n= belonging to a basis for .
Tensor Products 367
Bilinearity on Equals Linearity on <d= <n=
The universal property for bilinearity says that to each function bilinear
¢<d=¦> ¢<n=¦> , there corresponds a unique function , linear
called the mediating morphism for . Thus, we can define the mediating
morphism map
B¢ ² <Á=Â>³¦ ² <n=Á>³hom
by setting . In other words, is the unique linear map for which ~
² ³²"n#³ ~ ²"Á#³
Observe that is itself linear, since if , then Á ²<Á=Â>³ hom
´ ²³b ²³µ²"n#³ ~ ²"Á#³b ²"Á#³ ~ ² b ³²"Á#³
and so is the mediating morphism for , that is,² ³ b ² ³ b
²³b ²³ ~ ² b ³
Also, is surjective, since if is any linear map, then ¢<n= ¦>
~ k! ¢<d=¦> is bilinear and has mediating morphism , that is,
~ ~ ~ k!~ . Finally, is injective, for if , then . We have
established the following result.
Theorem 14.4 Let , and be vector spaces over . Then the mediating<= > -
morphism map , where is the unique B ¢ ² <Á=Â>³¦ ² <n=Á>³ hom
linear map satisfying , is an isomorphism and so ~ k!
B¢ ² <Á=Â>³ ² <n=Á>³hom
When Is a Tensor Product Zero?
Armed with the universal property of b ilinearity, we can now discuss some of
the basic properties of tensor products. Let us first consider the question of
when a tensor is zero. "n #
The bilinearity of the tensor product gives
n#~²b³n#~n#bn#
and so . Similarly, . Now suppose thatn#~ "n~
"n #~
where we may assume that none of the vectors and are . Let "#
¢<d=¦> ¢<n=¦> be a bilinear map and let be its mediating
morphism, that is, . Then k!~
368 Advanced Linear Algebra
~ " n# ~ ² k! ³ ² "Á#³~ ² "Á#³45
The key point is that this holds for bilinear function . In any ¢< d= ¦>
particular, let and and define by < = ii
²"Á#³ ~ ²"³ ²#³
which is easily seen to be bilinear. Then the previous display becomes
²" ³ ²# ³ ~
If, for example, the vectors are linearly independent, we can take to be a "
dual vector to get "i
~ "² "³ ² #³~ ² #³
i
and since this holds for all linear functionals , it follows that . We = # ~i
have proved the following useful result.
Theorem 14.5 If are linearly independent vectors in and"Á Ã Á " <
#Á Ã Á # = are arbitrary vectors in , then
"n #~ ¬ #~ for all
In particular, if and only if or . "n#~ "~ #~
Coordinate Matrices and Rank
If is a basis for and is a basis for , then89~¸ " 0¹ < ~¸ # 1¹ =
any vector has a unique expression as a sum '<n=
'~ ² "n#³
0 1Á
where only a finite number of the coefficients are nonzero. In fact, for a Á
fixed , we may reindex the bases so that'<n=
'~ ² "n#³
~ ~
Á
where none of the rows or columns of the matrix consists only of 's. 9~² ³ Á
The matrix is called a of with respect to the 9~² ³ ' Á coordinate matrix
bases and .89
Note that a coordinate matrix is deter mined only up to the order of its rows 9
and columns. We could remove this ambiguity by considering ordered bases,
Tensor Products 369
but this is not necessary for our discu ssion and adds a complication, since the
bases may be infinite.
Suppose that and are also bases for and MN~¸ $ 0¹ ~¸ % 1¹ <
=, respectively, and that
'~ ² $n%³
~ ~
Á
where is a coordinate matrix of with respect to these bases. We:~² ³ ' Á
claim that the coordinate matrices and have the same rank, which can then 9:
be defined as the of the tensor . rank '<n=
Each is a finite linear combination of basis vectors in , perhaps$Á Ã Á $ 8
involving some of and perhaps i nvolving other vectors in . We can "Á Ã Á " 8
further reindex so that each is a linear combination of the vectors 8 $
8Z~² "ÁÃÁ"³ , where and set
< ~ ²" ÁÃÁ" ³ span
Next, extend to a basis for . ²$ ÁÃÁ$³ ~²$ ÁÃÁ$Á$ ÁÃÁ$ ³ < b ZM
(Since we no longer need the rest of the basis , we have commandeered the M
symbols , for simplicity. Hence $Á Ã Á $b )
$~ " ~ Á Ã Á Á
~
for
where is invertible of size .(~² ³ d Á
Now repeat this process on the second coordinate. Reindex the basis so that 9
the subspace =~ span²# ÁÃÁ# ³ % ÁÃÁ% contains and extend to a basis
NZ b ~²% ÁÃÁ% Á% ÁÃÁ% ³ = for . Then
%~ # ~ Á Ã Á Á
~
for
where is invertible of size .)~² ³ d Á
Next, write
'~ ² "n#³
~ ~
Á
by setting for or . Thus, the matrix comes ~ d 9 ~ ² ³Á Á
from by adding rows of 's to the bottom and then columns of9 c c
9 9's. In particular, and have the same rank.
370 Advanced Linear Algebra
The expression for in terms of the basis vectors and can ' $ ÁÃÁ$ % ÁÃÁ%
also be extended using coefficients to
'~ ² $n%³
~ ~
Á
where the matrix has the same rank as . d : ~² ³ : Á
Now at last, we can compute. First, bilinearity gives
$n %~ "n # Á Á
~~
and so
'~ ² $n%³~ " n#
~² ³ " n #
~ 89
89
89~ ~ ~ ~
Á Á Á Á
~~
~~
~ ~Á Á Á
~~
~
!
Á Á
~~
!
Á²( : ³ " n#
~( : ) " n #
23
Thus
~ ~
Á Á
~~!² " n # ³ ~ ' ~ ² ( : ) ³ ² "n # ³
and so . Since and are invertible, we deduce that9~ ( : ) ( )!
rk rk rk rk² 9 ³~ ² 9³~ ² :³~ ² :³
as desired. Moreover, in block matrix terms, we can write
9~ :~9 :
>? >?
block blockand
and if we write
(~ ) ~(i
ii)i
ii!!
Á Á>? >?
block blockand
then implies that9~ ( : )!
Tensor Products 371
9~( :)!
ÁÁ
We shall soon have use for the following special case. If
'~ "n# ~ $n%
~ ~
()14.3
then and so9~:~0
$~ " ~ Á Ã Á Á
~
for
and
%~ # ~ Á Ã Á Á
~
for
where if and , then (~ ² ³ )~ ² ³Á Á Á Á
0~ ( ) Á !
Á
The Rank of a Decomposable Tensor
Recall that a tensor of the form is said to be decomposable. If "n# ¸" 0¹
is a basis for and is a basis for , then any decomposable vector <¸ # 1 ¹ =
has the form
"n#~ ²" n#³
Á
Hence, the rank of a decomposable vector is , since the rank of a matrix whose
²Á³ th entry is is .
Characterizing Vectors in a Tensor Product
There are several useful representations of the tensors in . <n=
Theorem 14.6 Let be a basis for and let be a basis¸" 0¹ < ¸# 1¹
for . By an “essentially unique” sum, we mean unique up to order and=
presence of zero terms.
1 Every has an essentially unique expression as a finite sum of) '<n=
the form
ÁÁ " n #
where and the tensors are distinct. - " n #Á
372 Advanced Linear Algebra
2 Every has an essentially unique expression as a finite sum of) '<n=
the form
"n &
where and the 's are distinct.& = "
3 Every has an essentially unique expression as a finite sum of) '<n=
the form
%n #
where and the 's are distinct.% < #
4 Every nonzero has an expression of the form) '<n=
'~ %n&
~
where the 's are distinct, the 's are distinct and the sets and %& ¸ % ¹ <
¸& ¹ = ' are linearly independent. As to uniqueness, is the rank of and
so it is unique. Also, the equation
~ ~
%n &~ $n '
where the 's are distinct, the 's are distinct and and $'¸ $ ¹ <
¸' ¹ = are linearly independent, holds if and only if there exist invertible
d (~² ³ )~² ³ ()~0 matrices and for which and Á Á!
$~ % '~ & Á Á
~ ~
and
for .~ ÁÃÁ
Proof. Part 1) merely expresses the fact that is a basis for . ¸" n#¹ <n=
From part 2), we write
@A
Á Á Á " n #~ " n # ~ " n &
Uniqueness follows from Theorem 14.5. Part 3) is proved similarly. As to part
4), we start with the expression from part 2):
~
"n &
where we may assume that none of the 's are . If the set is linearly & ¸ & ¹
independent, we are done. If not, then we may suppose (after reindexing if
Tensor Products 373
necessary) that
&~ &
~c
Then
89
~ ~ ~ c c
~ ~c c
~c
"n &~ "n &b "n &
~" n & b " n &
~² " b " ³ n &
But the vectors are linearly independent. This ¸" b " c¹
reduction can be repeated until the second coordinates are linearly independent.
Moreover, the identity matrix is a coordinate matrix for and so0'
~ ²0 ³ ~ ²'³ rk rk . As to uniqueness, one direction was proved earlier; see
()14.3 and the other direction is left to the reader.
The proof of Theorem 14.6 shows that if and '£
'~ n!
0
where and , then if the multiset is not linearly < ! = ¸ 0¹
independent, we can rewrite in the form '
'~ n!
0Z
where is linearly independent. Then we can do the same for the¸ 0 ¹
second coordinate to arrive so at the representation
'~ %n&
~²%³
rk
where the multisets and are linearly independent sets. Therefore, ¸% ¹ ¸& ¹
rk²%³ 0 ' ' (( and so the rank of is the integer for which can be smallest
written as a sum of decomposable ten sors. This is often taken as the
definition of the rank of a tensor.
However, we caution the reader that there is another meaning to the word rank
when applied to a tensor, namely, it is the number of indices required to write
the tensor. Thus, a scalar has rank , a vector has rank , the tensor above has '
rank and a tensor of the form
374 Advanced Linear Algebra
'~ n!n"
0
has rank .
Defining Linear Transformations on a Tensor Product
One of the simplest and most useful ways to define a linear transformation on
the tensor product is through the universal property, for this property <n=
says precisely that a bilinear function on gives rise to a unique (and< d =
well-defined) linear transformation on . The proof of the following <n=
theorem illustrates this well.
Theorem 14.7 Let and be vector spaces. There is a unique linear<=
transformation
¢< n= ¦²<n=³ii i
defined by where² n³ ~ p
²p³²"n#³~²"³²#³
Moreover, is an embedding and is an isomorphism if and are finite- <=
dimensional. Thus, the tensor product of linear functionals is via thisn (
embedding a linear functional on tensor products. )
Proof. Informally, for fixed and , the function is bilinear ² " Á # ³ ¦ ² " ³ ² # ³
in and and so there is a unique linear map taking to ."# p " n # ² " ³ ² # ³
The function is bilinear in and since ²Á³ ¦ p
² b ³p~ ² p ³b ² p ³
and so there is a unique linear map taking to . n p
More formally, for fixed and , the map defined by - ¢ < d = ¦ - Á
- ²"Á#³ ~ ²"³²#³Á
is bilinear and so the universal property of tensor products implies that there
exists a unique for which p² <n=³i
²p³²"n#³~²"³²#³
Next, the map defined by .¢< d= ¦ ²< n=³ii i
.²Á³ ~ p
Tensor Products 375
is bilinear since, for example,
´²b ³pµ²"n#³~²b ³²"³h²#³
~ ²"³²#³b ²"³²#³
~ ´² p³b ²p³µ²"n#³
which shows that is linear in its first coordinate. Hence, the universal .
property implies that there exists a unique linear map
¢< n= ¦²<n=³ii i
for which
² n³ ~ p
To see that is an injection, if is nonzero, then we may write in < n= ii
the form
~ n
~
where the are nonzero and is linearly < ¸ ¹ =ii
independent. If , then for any and , we have ²³ ~ " < # =
~ ²³²"n#³ ~ ² n ³²"n#³ ~ ²"³ ²#³
~ ~
Hence, for each nonzero , the linear functional "<
~
² " ³
is the zero map and so the linear independence of implies that ¸ ¹ ²"³ ~
for all . Since is arbitrary, it follows that for all and so ." ~ ~
Finally, in the finite-dimensional cas e, the map is a bijection since
dim dim²< n= ³ ~ ²²< n= ³ ³ Bii i
Combining the isomorphisms of Theorem 14.4 and Theorem 14.7, we have, for
finite-dimensional vector spaces and , <=
< n= ²< n= ³ ²<Á= Â-³ii ihom
The Tensor Product of Linear Transformations
We wish to generalize Theorem 14.7 to arbitrary linear transformations. Let
B B ² < Á < ³ ² = Á = ³ ² " ³ ² # ³ZZ and . While the product does not make
sense, the product does and is bilinear in and , that is, the tensor "n # " #
following function is bilinear:
376 Advanced Linear Algebra
²"Á#³~ "n #
The same argument that we used in the proof of Theorem 14.7 will work here.
Namely, the map from to is bilinear in and ²"Á#³ ª "n # < d= < n= "ZZ
# ² p ³¢<n= ¦< n= and so there is a unique linear map for which ZZ
²p³ ² " n # ³ ~" n#
The function
B B B¢ ² <Á<³d ² =Á= ³¦ ² <n=Á< n= ³ZZ Z Z
defined by
²Á³ ~ p
is bilinear, since
²² b ³p ³²"n#³ ~ ² b ³²"³n #
~² "b " ³n #
~ ´ "n # µb ´ "n # µ
~ ² p ³²"n#³b² p ³²"n#³
~ ²² p ³b² p ³³²"n#³
and similarly for the second coordinate. Hence, there is a unique linear
transformation
B B B¢ ² <Á<³n ² =Á= ³¦ ² <n=Á< n= ³ZZ Z Z
satisfying
²n³ ~ p
that is,
´ ²n³ µ ² " n # ³ ~" n#
To see that is injective, if is nonzero, then we may B B ²<Á< ³n ²= Á= ³ZZ
write
~ n
~
where the are nonzero and the set is linearly ² < Á <³ ¸ ¹ ² = Á =³ZZBB
independent. If , then for all and we have ²³ ~ " < # =
~ ²³²"n#³ ~ ² n ³²"n#³ ~ ²"³n ²#³
~ ~
Since , it follows that for some and so we may choose a £ £ "<
such that for some . Moreover, we may assume, by reindexing if ² " ³£
Tensor Products 377
necessary, that the set is a maximal linearly independent ¸ ²"³ÁÃÁ ²"³¹
subset of . Hence, for each , we have ¸ ²"³ÁÃÁ ²"³¹
² " ³~ ² " ³ Á
~
and so
~ ² " ³n² # ³
~ ²"³n ²#³b ²"³ n ²#³
~ ²"³n²#³b ²"³n²#³
~ ² " ³ n ² #
@A
!
~
~ ~
Á
~b
~ ~
Á
~b
~
³b ²"³n ²#³
~ ² " ³ n ² # ³ b ² # ³@A
@A~
Á
~b
~
Á
~b
Thus, the linear independence of implies that for each ¸ ²"³ÁÃÁ ²"³¹
,
² # ³b ² # ³~ Á
~b
for all and so#=
b ~ Á
~b
But this contradicts the fact that the s et is linearly independent. Hence, it¸ ¹
cannot happen that for and so is injective. ²³ ~ £
The embedding of into means that BB B²<Á< ³n ²= Á= ³ ²< n= Á< n= ³ZZ Z Z
each can be thought of as the linear transformation from to np < n =
<n =ZZ, defined by
²p³ ² " n # ³ ~" n#
In fact, the notation is often us ed to denote both the tensor product of n
vectors linear transformations and the linear map , and we will do this as () p
well. In summary, we can say that the tensor product of linear n
transformations is (up to isomorphism) a linear transformation on tensor
products.
378 Advanced Linear Algebra
Theorem 14.8 There is a unique linear transformation
B B B¢ ² <Á<³n ² =Á= ³¦ ² <n=Á< n= ³ZZ Z Z
defined by where ²n³ ~ p
²p³ ² " n # ³ ~" n#
Moreover, is an embedding and is an isomorphism if all vector spaces are
finite-dimensional. Thus, the tensor product of linear transformations isn
()via this embedding a linear transformation on tensor products.
Let us note a few special cases of the previous theorem.
Corollary 14.9 Let us use the symbol to denote the fact that there is an ?@Æ
embedding of into that is an isomorphism if and are finite- ?@ ?@
dimensional.
1 Taking gives) <~ -Z
<n² = Á = ³ ² < n = Á = ³iZ ZBÆ B
where
² n ³²"n#³ ~ ²"³ ²#³
for .<i
2 Taking and gives) <~ - =~ -ZZ
<n = ² < n = ³ii iÆ
where
² n³²"n#³ ~ ²"³²#³
3 Taking and noting that and gives) =~ - ² - Á =³ = <n - < BZZ
()letting >~=Z
BÆ B²<Á< ³n> ²<Á< n>³ZZ
where
² n$³²"³ ~ "n$
4 Taking and gives letting )( ) <~ - =~ - >~ =ZZ
<n > ² < Á > ³iÆB
where
² n$³²"³ ~ ²"³$
Tensor Products 379
Change of Base Field
The tensor product provides a convenient way to extend the base field of a
vector space that is more general than the complexification of a real vector
space, discussed earlier in the book. We refer to a vector space over a field as -
an and write .-=-space -
Actually, there are several approaches to “upgrading” the base field of a vector
space. For instance, suppose that is an extension field of , that is, . 2- - 2
If is a basis for , then every has the form¸ ¹ = % =- -
%~
where . We can define a -space simply by taking all formal linear - 2 =2
combinations of the form
%~
where . Note that the dimension of as a -space is the same as the22 = 2
dimension of as an -space. Also, is an -space just restrict the scalars =- =--2 (
to and as such, the inclusion map sending to- ¢ = ¦ = % =) -2 -
²%³ ~ % = - 2 is an -monomorphism.
The approach described in the previous paragraph uses an arbitrarily chosen
basis for and is therefore not coordinate free. However, we can give a =-
coordinate-free approach using tensor products as follows. Since is a vector 2
space over , we can form the tensor product -
>~ 2 n=-- -
It is customary to include the subscr ipt on to denote the fact that the-n -
tensor product is taken with respect to the base field . (All relevant maps are -
-- = 2-bilinear and -linear.) However, since is not a -space, the only tensor -
product of and that makes sense is the -tensor product and so we will 2= - -
drop the subscript . -
The tensor product is an -space by definition of tensor product, but we >--
can make it into a -space as follows. For , the temptation is to “absorb” 2 2
the scalar into the first coordinate,
²n # ³ ~ ² ³ n #
but we must be certain that this is well-defined, that is,
n#~ n$ ¬ ² ³n#~² ³n$
But for a fixed , the map is bilinear and so the universal ²Á # ³ ª ² ³ n #
property of tensor products implies that there is a unique linear map
n#ª² ³n# , which we define to be scalar multiplication by .
380 Advanced Linear Algebra
To be absolutely clear, we have two distinct vector spaces: the -space -
> ~2n= 2 > ~2n=-- 2- defined by the tensor product and the -space
with scalar multiplication by elements of defined as absorption into the first2
coordinate. The spaces and are identical as sets and as abelian groups. >>-2
It is only the “permission to multiply by ” that is different. Accordingly, we can
recover from simply by restricting scalar multiplication to scalars from>>-2
-.
Thus, we can speak of “ -linear” maps from into , with the expected -= > -2
meaning, that is,
²"b #³ ~ "b #
for all scalars . Á -
If the dimension of as a vector space over is , then 2-
dim dim dim-- - - - -² > ³~ ² 2n= ³~h ² = ³
As to the dimension of , it is not hard to see that if is a basis for , >¸ ¹ =2 -
then is a basis for . Hence¸n ¹ > 2
dim dim22 - -²> ³ ~ ²= ³
The map defined by is easily seen to be injective and¢= ¦> #~n#--
-> =-linear and so contains an isom orphic copy of . We can also think of --
as mapping into , in which case is called the of . => 2 =-2 - -extension map
This map has a universal property of its own, as described in the next theorem.
Theorem 14.10 The -linear -extension map has the-2 ¢ = ¦ 2 n = --
universal property for the family of all -linear maps from into a -space,-= 2 -
as measured by -linear maps. Specifically, for any -linear map , 2- ¢ = ¦ @ -
where is a -space, there exists a unique -linear map for@2 2 ¢ 2 n = ¦ @ -
which the diagram in Figure 14.6 commutes, that is, for which
k~
Proof. If such a -linear map is to exist, then it must satisfy, 2¢ 2 n = ¦ @ -
for any ,2
² n#³ ~ ²n#³ ~ ²#³ ~ ²#³
This shows that if exists, it is uniquely determined by . As usual, when
searching for a linear map on a tensor product such as , we look for a 2n= -
bilinear map. The map defined by ¢²2 d= ³ ¦ @ -
² Á#³ ~ ²#³
Tensor Products 381
is bilinear and so there exists a unique -linear map for which -
²n # ³ ~ ² # ³
It is easy to see that is also -linear, since if , then 2 2
´ ² n#³µ~ ² n#³~ ²#³~ ² n#³
VF K VF
YfWP
Figure 14.6
Theorem 14.10 is the key to describing how to extend an -linear map to a - -2
linear map. Figure 14.7 shows an -linear map between -spaces -¢ = ¦ > - =
and . It also shows the -extensions for both spaces, where and>2 2 n =
2n> 2 are -spaces.
V W
PWW
WK
VPV
K
W
Figure 14.7
If there is a unique -linear map that makes the diagram in Figure 14.7 2
commute, then this would be the obvious choice for the extension of the - -
linear map to a -linear map.2
Consider the -linear map into the -space -~ ² k ³ ¢ = ¦ 2 n > 2 >
2n> 2 . Theorem 14.10 implies that there is a unique -linear map
¢2n= ¦2n> for which
k~=
that is,
k~k=>
Now, satisfies
382 Advanced Linear Algebra
² n#³ ~ ²n#³
~ ² k ³²#³
~ ² k ³²#³
~² n# ³
~n #
~² n ³ ² n# ³=
>
2
and so . ~n2
Theorem 14.11 Let and be -spaces, with -extension maps and=> - 2 =
>, respectively. See Figure 14.7. Then for any -linear map , the () -¢ = ¦ >
map is the unique -linear map that makes the2n ¢2n=¦2n> 2
diagram in Figure 14.7 commute, that is, for which
k~ ² n ³ k 2
Multilinear Maps and Iterated Tensor Products
The tensor product operation can easily be extended to more than two vector
spaces. We begin with the extension of the concept of bilinearity.
Definition If and are vector spaces over , a function=ÁÃÁ= > -
¢= dÄd= ¦> is said to be if it is linear in each coordinate multilinear
separately, that is, if
² " Á Ã Á "Á # b # Á "Á Ã Á " ³
~ ² " Á Ã Á "Á # Á "Á Ã Á " ³ b ² " Á Ã Á "Á # Á "Á Ã Á " ³ c b Z
c b c b Z
for all . A multilinear function of va riables is also referred to as ~ ÁÃÁ
an . The set of all -linear functions as defined above will be-linear function
denoted by . A multilinear function from to hom²=ÁÃÁ= Â>³ = dÄd=
the base field is called a or . - multilinear form -form
Example 14.7
1 If is an algebra, then the product map defined by)(¢ ( d Ä d ( ¦ (
² ÁÃÁ ³~ Ä is -linear.
2 The determinant function is an -linear form on the columns)d e t ¢¦ - C
of the matrices in . C
The tensor product is defined via its universal property.
Definition As pictured in Figure 14.8, let be the cartesian =d Ä d =
product of vector spaces over . A pair is - ²;Á!¢= dÄd= ¦;³ universal
for multilinearity if for every multilinear map , there is ¢= dÄd= ¦>
a unique linear transformation for which ¢; ¦>
Tensor Products 383
~ k!
The map is called the for . If is universal for mediating morphism ² ; Á ! ³
multilinearity, then is called the of and denoted by ; =ÁÃÁ= tensor product
=n Ä n = ! . The map is called the . tensor map
WfWV1
VntV1uuVn
Figure 14.8
As we have seen, the tensor product is unique up to isomorphism.
The basis construction and coordinate-free construction given earlier for the
tensor product of two vector spaces carry over to the multilinear case.
In particular, let be a basis for for . For each 8 Á ~¸ 1¹ = ~ ÁÃÁ
ordered -tuple , construct a new formal symbol ² Á Ã Á ³ Á Á
n Ä n ;Á Á and define to be the vector space with basis
:~¸ nÄn 1¹Á Á
The tensor map is defined by setting !¢= dÄd= ¦ ;
!² ÁÃÁ ³~ nÄnÁ Á Á Á
and extending by multilinearity. This unique ly defines a multilinear map that is !
universal for multilinear functions from . =d Ä d =
Indeed, if is multilinear, the condition is ¢= dÄd= ¦ > ~ k!
equivalent to
² nÄn ³~² ÁÃÁ ³Á Á Á Á
which uniquely defines a linear map . Hence, has the universal ¢; ¦> ²;Á!³
property for multilinearity.
Alternatively, we may take the coordinate-free quotient space approach as
follows.
Definition Let be vector spaces over and let be the vector space=ÁÃÁ= - <
with basis . Let be the subspace of generated by all vectors of =d Ä d = : <
the form
384 Advanced Linear Algebra
²# ÁÃÁ# Á"Á# ÁÃÁ# ³b ²# ÁÃÁ# Á"Á# ÁÃÁ# ³
c²#ÁÃÁ# Á"b "Á# ÁÃÁ# ³ c b c b Z
c b Z
for , and for . The quotient space is theÁ - "Á" = # = £ °:Z <
tensor product of and the tensor map is the map =ÁÃÁ=
! ² #Á Ã Á #³~² #Á Ã Á #³b:
As before, we denote the coset by and so any ²# ÁÃÁ# ³b: # nÄn#
element of is a sum of decomposable tensors, that is, =n Ä n =
#n Ä n #
where the vector space operations are linear in each variable.
Here are some of the basic properties of multiple tensor products. Proof is left to
the reader.
Theorem 14.12 The tensor product has the following properties. Note that all
vector spaces are over the same field . -
1 There exists an isomorphism)( )Associativity
¢²= nÄn= ³n²> nÄn> ³¦= nÄn= n> nÄn>
for which
´²# nÄn# ³n²$ nÄn$ ³µ~# nÄn# n$ nÄn$
In particular,
²<n=³n><n²= n>³<n= n>
2 Let be any permutation of the indices . Then)( )Commutativity ¸ÁÃÁ¹
there is an isomorphism
¢= nÄn= ¦= nÄn= ²³ ²³
for which
²# nÄn# ³~# nÄn# ²³ ²³
3 There is an isomorphism for which) ¢-n= ¦=
²n#³ ~ #
and similarly, there is an isomorphism for which ¢= n- ¦=
²#n³ ~ #
Hence, .-n n-===
The analog of Theorem 14.4 is the following.
Tensor Products 385
Theorem 14.13 Let and be vector spaces over . Then the=ÁÃÁ= > -
mediating morphism map
B¢ ² =Á à Á = > ³¦ ² =n Ä n =Á > ³hom
defined by the fact that is the unique mediating morphism for is an
isomorphism. Thus,
hom²=ÁÃÁ= Â>³ ²= nÄn= Á>³ B
Moreover, if all vector spaces are finite-dimensional, then
dim hom dim dim´² = Á Ã Á = Â > ³ µ ~² > ³ h ² = ³
~
Theorem 14.8 and its corollary can also be extended.
Theorem 14.14 The linear transformation
B B B¢ ² <Á <³ n Ä n ² <Á <³¦ ² <n Ä n <Á <n Ä n <³ ZZ Z Z
defined by
² nÄn ³²" nÄn" ³~ " nÄn "
is an embedding and is an isomorphism if all vector spaces are finite-
dimensional. Thus, the tensor product of linear transformations is nÄn
()via this embedding a linear transform ation on tensor products. Two important
special cases of this are
< nÄn< ²< nÄn< ³ii i
Æ
where
² nÄn ³²" nÄn" ³~²"³Ä ²" ³
and
<n Ä n <n = ² <n Ä n < Á = ³ii
ÆB
where
² nÄn n#³²" nÄn" ³~²"³Ä ²" ³#
Tensor Spaces
Let be a finite-dimensional vector space. For nonnegative integers and ,=
the tensor product
386 Advanced Linear Algebra
; ²=³~ = nÄn= n= nÄn= ~= n²= ³i i n i n
factors factors
is called the space of , where is the tensors of type contravariant type ²Á³
and is the . If , then , the base field. Here ~~ ; ² =³~- covariant type
we use the notation for the -fold tensor product of with itself. We will = =n
also write for the -fold cartesian product of with itself. = =d
Since , we have= =ii
; ²= ³ ~ = n²= ³ ²²= ³ n= ³ ²²= ³ d= Á-³ n in in n i id d
-hom
which is the space of all multilinear functionals on
= dÄd= d= dÄd=ii
factors factors
In fact, tensors of type are often defined as multilinear functionals in this ²Á³
way.
Note that
dim dim²; ²= ³³ ~ ´ ²= ³µ b
Also, the associativity and commutativ ity of tensor products allows us to write
; ² =³n; ² =³~; ² =³
b
b
at least up to isomorphism.
Tensors of type are called ²Á³ contravariant tensors
;² = ³ ~ ;² = ³ ~=n Ä n =
factors
and tensors of type are called ²Á³ covariant tensors
;² =³~;² =³~= nÄn=i i
factors
Tensors with both contravariant and covariant indices are called . mixed tensors
In general, a tensor can be interpreted in a variety of ways as a multilinear map
on a cartesian product, or a linear map on a tensor product. Indeed, the
interpretation we mentioned above that is sometimes used as the definition is
only one possibility. We simply need to decide how many of the contravariant
factors and how many of the covariant factors should be “active participants”
and how many should be “passive participants.”
Tensor Products 387
More specifically, consider a tensor of type , written ²Á³
# nÄn# nÄn# n nÄn nÄn ; ²=³
where and . Here we are choosing the first vectors and the first
linear functionals as active participants. This determines the number of
arguments of the map. In fact, we define a map from the cartesian product
= dÄd= d= dÄd=ii
factors factors
to the tensor product
= nÄn= n= nÄn=
c cii
factors factors
of the remaining factors by
²
~ ²# ³Ä ²# ³ ²% ³Ä ²% ³# nÄn# n nÄn³² ÁÃÁ Á% ÁÃÁ% ³
# nÄn# n nÄn
b b
In words, the first group of (active) vectors interacts with the first #n Ä n #
group of arguments to produce the scalar . The first ÁÃÁ ²# ³Ä ²# ³
group of (active) functionals interacts with the second groupn Ä n
%Á Ã Á % ² %³ Ä ² %³ of arguments to produce the scalar . The remaining
(passive) vectors and functionals are just # nÄn# nÄnb b
“copied” to the image tensor.
It is easy to see that this map is multilin ear and so there is a unique linear map
from the tensor product
= nÄn= n= nÄn=ii
factors factors
to the tensor product
= nÄn= n= nÄn=
c cii
factors factors
defined by
²
~ ²# ³Ä ²# ³ ²% ³Ä ²% ³# pÄp# p pÄp³² nÄn n% nÄn% ³
# nÄn# n nÄn
b b
Moreover, the map
B¢= n²= ³ ¦ ²²= ³ n= Á= n²= ³ ³n i n i n n n²c³ i n²c³
defined by
²# nÄn# n nÄn³~# pÄp# p pÄp
388 Advanced Linear Algebra
is an isomorphism, since if # pÄp# p pÄp is the zero map then
²# ³Ä ²# ³ ²% ³Ä ²% ³ # nÄn# n nÄn ~b b
for all and , which implies that = % =i
# nÄn# n nÄn ~
As usual, we denote the map by # pÄp# p pÄp
# nÄn# n nÄn
Theorem 11.15 For and ,
;² = ³
B²²= ³ n= Á= n²= ³ ³i n n n²c³ i n²c³
When and , we get~ ~
;² = ³
B²²= ³ n= Á-³ ~ ²²= ³ n= ³i n n i n n i
as before.
Let us look at some special cases. For we have ~
;² = ³
B²²= ³ Á= ³in n ² c ³
where
² ²# ³Ä ²# ³# nÄn#³² nÄn ³~ # nÄn# b
When , we get for and ,~~ ~ ~
;² = ³
BB²-n=Á= n-³ ²=³
where
²#n³²$³ ~ ²$³#
and for and ,~ ~
;² = ³
BB²= n-Á-n= ³ ²= Á= ³ii i i
where
²#n³²³ ~ ²#³
Finally, when , we get a multilinear form ~~
²#n³²Á$³ ~ ²#³²$³
Consider also a tensor of type . When we get a n ²Á³ ~ ~
multilinear functional defined by n ¢² =d=³¦-
Tensor Products 389
² n³²#Á$³ ~ ²#³²$³
This is just a bilinear form on . =
Contraction
Covariant and contravariant factors can be “combined” in the following way.
Consider the map
¢= d²= ³ ¦ ; ²=³d i d c
c
defined by
²#ÁÃÁ#ÁÁÃÁ³~²#³²# nÄn# n nÄn³
This is easily seen to be multilinear and so there is a unique linear map
¢; ²=³¦; ²=³
c
c
defined by
²# nÄn# n nÄn³~²#³²# nÄn# n nÄn³
This is called the in the contra variant index and covariant index contraction
. Of course, contraction in other indi ces (one contravariant and one covariant)
can be defined similarly.
Example 14.8 Let and consider the tensor space , which is dim²= ³ ; ²= ³
isomorphic to via the map B²= ³
²#n³²$³ ~ ²$³#
For a “decomposable” linear operator of the form as defined above with #n
#£ £ ²#n³~ ²³ and , we have , which has codimension . ker ker
Hence, if , then ²$³²#³ ~ ²#n³²$³ £
= ~ º$»l ²³ ~ º$»l ker ;
where is the eigenspace of associated with the eigenvalue .; #n
In particular, if , then ²#³£
²#n³²#³ ~ ²#³#
and so is an eigenvector for the nonzero eigenvalue . Hence,# ² # ³
=~ º # » l ~ l;; ;²#³
and so the trace of is #n
390 Advanced Linear Algebra
tr²#n³~²#³~ ²#n³
where is the contraction map.
The Tensor Algebra of =
Consider the contravariant tensor spaces
;² = ³ ~ ;² = ³ ~ =n
For we take . The external direct sum~ ; ² =³~-
;²=³~ ; ²=³
~B
of these tensor spaces is a vector space with the property that
; ² =³n; ² =³~; ² =³ b
This is an example of a , where are the elements of graded algebra grade ;² = ³
; ² = ³ =. The graded algebra is called the over . We will tensor algebra (
formally define graded structures a bit later in the chapter.)
Since
;²=³~= nÄn= ~; ²= ³ii i
factors
there is no need to look separately at . ;² =³
Special Multilinear Maps
The following definitions describe some special types of multilinear maps.
Definition
1 A multilinear map is if interchanging any two) ¢= ¦>dsymmetric
coordinate positions change s nothing, that is, if
²# ÁÃÁ#ÁÃÁ#ÁÃÁ# ³~²# ÁÃÁ#ÁÃÁ#ÁÃÁ# ³
for any .£
2 A multilinear map is or if) ¢= ¦>dantisymmetric skew-symmetric
interchanging any two coordinate positions introduces a factor of , that c
is, if
²# ÁÃÁ#ÁÃÁ#ÁÃÁ# ³~c²# ÁÃÁ#ÁÃÁ#ÁÃÁ# ³
for .£
Tensor Products 391
3 A multilinear map is or if) ¢= ¦>alternate alternating
# ~# £ ¬ ²# ÁÃÁ# ³~ for some
As in the case of bilinear forms, we have some relationships between these
concepts. In particular, if , then char²-³ ~
alternate symmetric skew-symmetric ¬¯
and if , then char²-³ £
alternate skew-symmetric ¯
A few remarks about permutations are in order. A of the set permutation
5 ~ ¸ÁÃÁ¹ ¢5 ¦ 5 is a bijective function . We denote the group under (
composition of all such permutations by . This is the on ) : symmetric group
symbols. A of length is a permutation of the form , which cycle ² ÁÁÃÁ³
sends to for and also sends to . All other elements " ~ Á Ã Á c "" b
of are left fixed. Every permutation is the product (composition) of disjoint5
cycles.
A is a cycle of length . Every cycle and therefore everytransposition ²Á³ (
permutation is the product of transpositions. In general, a permutation can be )
expressed as a product of transpositions in many ways. However, no matter how
one represents a given permutation as such a product, the number of
transpositions is either always even or always odd. Therefore, we can define the
parity of a permutation to be the parity of the number of transpositions :
in any decomposition of as a product of transpositions. The of a sign
permutation is defined by
sg has even parity
has odd parity²³ ~
c
F
If sg , then is an and if sg , then is an²³ ~ ²³ ~ c even permutation
odd permutation . The sign of is often written . ²c³
With these facts in mind, it is apparent that is symmetric if and only if
²# ÁÃÁ# ³~²# ÁÃÁ# ³ ²² ³1)
for all permutations and that is skew-symmetric if and only if :
²# ÁÃÁ# ³~²c³ ²# ÁÃÁ# ³ ²² ³
1)
for all permutations . :
A word of caution is in order with respect to the notation above, which is very
convenient albeit somewhat prone to confusi on. It is intended that a permutation
permutes the coordinate positions in , not the indices (despite appearances).
Suppose, for example, that and that is a basis for . ¢ d ¦? ¸ Á ¹ss s
392 Advanced Linear Algebra
If , then applied to gives and not , since~ ²³ ² Á ³ ² Á ³ ² Á ³
permutes the two coordinate positions in . ² #Á #³
Graded Algebras
We need to pause for a few definitions that are useful in discussing tensor
algebras. An algebra over is said to be a if as a vector (- graded algebra
space over , can be written in the form -(
(~ (
~B
for subspaces of , and where multiplication behaves nicely, that is, ((
(( ( b
The elements of are said to be . If is written ( ( homogeneous of degree
~ bÄb
for , , then is called the of of ( £ homogeneous component
degree .
The ring of polynomials provides a prime example of a graded algebra, -´%µ
since
-´%µ~ -´%µ
~B
where is the subspace of consisting of all scalar multiples of .- ´%µ -´%µ %
More generally, the ring of polynomials in several variables is a -´% ÁÃÁ% µ
graded algebra, since it is the direct sum of the subspaces of homogeneous
polynomials of degree . A polynomial is if each ( homogeneous of degree
term has degree . For example, is homogeneous of degree ~%%b%%%
.)
The Symmetric and Antisymmetric Tensor Algebras
Our discussion of symmetric and antisymmetric tensors will benefit by a
discussion of a few definitions and se tting a bit of notation at the outset.
Let denote the vector space of all homogeneous polynomials of-´ Á Ã Á µ
degree (together with the zero polynomial) in the independent variables
Á Ã Á . As is sometimes done in this context, we denote the product in
-´ Á Ã Á µ v vv by , for example, writing as . The algebra of
all polynomials in is denoted by . ÁÃÁ -´ ÁÃÁ µ
Tensor Products 393
We will also need the counterpart of in which multiplication acts -´ Á Ã Á µ
anticommutatively , that is, . ~c
Definition Let be a sequence of independent variables. For,~² ÁÃÁ ³
- ´ ÁÃÁ µ - , let be the vector space over with basisc
7 ² ,³~¸ Ä Ä¹
consisting of all words of length over that are in ascending order. Let,
-´ Á Ã Á µ ~ - - -c
, which we identify with by identifying with .
Define a product on the direct sum
-´ Á Ã Á µ ~ -cc
~
as follows. First, the product of monomials and w ~%Ä % - c
~&Ä & -c is defined as follows:
1 If has a repeated factor then .)%Ä %&Ä & w~
2 Otherwise, reorder in ascending order, say , via the) %Ä %&Ä & 'Ä ' b
permutation and set
w~² c ³'Ä '
b
Extend the product by distributivity to . The resulting product -´ Á Ã Á µc
makes into a noncommutative algebra over . This product is-´ Á Ã Á µ -c ()
called the or on . wedge product exterior product -´ Á Ã Á µc
For example, by definition of wedge product,
w w ~ c w w
Let be a basis for . It will be convenient to group the8~¸ ÁÃÁ¹ =
decomposable basis tensors according to their index multiset. n Ä n
Specifically, for each multiset with , let be the 4~¸ÁÃÁ ¹ . 4
set of all tensors
n Ä n
where is a permutation of . For example, if² ÁÃÁ ³ ¸ÁÃÁ ¹
4 ~ ¸ÁÁ¹ , then
. ~¸ n nÁ n nÁ n n¹4
If has the form#; ² =³
#~ nÄn
Á Ã Á Á Ã Á
where , then let be the subset of whose elements appearÁ Ã Á 4 4£ . ² # ³ .
394 Advanced Linear Algebra
in the sum for . For example, if #
# ~ n n b n n b n n
then
. ² # ³ ~ ¸ n n Á n n ¹¸ÁÁ¹
Let denote the sum of the terms of associated with . For:² # ³ # .² # ³4 4
example,
: ² # ³ ~ n n b n n ¸ÁÁ¹
Thus, can be written in the form#
#~ : ² # ³~ ! ps
qt 444!
!. ²#³4
where the sum is over a collection of multisets with . Note also 4: ² # ³ £ 4
that since . Finally, let!4£ !. ² # ³
"~ n Ä n 4
be the unique member of for which . . Ä4
Now we can get to the business at hand.
Symmetric and Antisymmetric Tensors
Let be the symmetric group on . For each , the multilinear: ¸ÁÃÁ¹ :
map defined by¢ = ¦;² =³d
²%ÁÃÁ%³~% nÄn%
determines a unique linear operator on for which ;² = ³
²% nÄn%³~% nÄn%
For example, if and , then ~ ~² ³
²³ ²# n# n#³~# n# n#
Let be a basis for . Since is a bijection of the basis¸ ÁÃÁ ¹ =
88~¸ nÄn ¹
it follows that is an isomorphism of . Note also that is a ;² = ³
permutation of each , that is, the sets are invariant under . ..44
Definition Let be a finite-dimensional vector space.=
Tensor Products 395
1 A tensor is if) !; ² =³symmetric
!~!
for all permutations . The set of all symmetric tensors :
:; ²= ³ ~ ¸! ; ²= ³ ! ~ ! : ¹
for all
is a subspace of , called the of degree ;² = ³ symmetric tensor space
over .=
2 A tensor is if) !; ² =³antisymmetric
!~² c ³!
The set of all antisymmetric tensors
(; ²= ³ ~ ¸! ; ²= ³ ! ~ ²c³ ! : ¹
for all
is a subspace of , called the or ;² = ³antisymmetric tensor space exterior
product space of degree over . =
We can develop the theory of symmetr ic and antisymmetric tensors in tandem.
Accordingly, let us write (anti)symmetr ic to denote a tensor that is either
symmetric or antisymmetrtic.
Since for any , there is a permutation taking to , an Á! . ! 4
(anti)symmetric tensor must have and so #. ² # ³ ~ . 44
#~ : ² # ³~ ! 89
444!
!.4
Since is a permutation of , it follows that is symmetric if and only if .#4
²: ²#³³ ~ : ²#³44
for all and this holds if and only if the coefficients of are: : ² # ³! 4
equal, say for all . Hence, the symmetric tensors are precisely!4 4~! .
the tensors of the form
#~ !89
44
!.
4
The tensor is antisymmetric if and only if #
²: ²#³³ ~ ²c³ : ²#³44 (14.4)
In this case, the coefficients of differ only by sign. Before examining !4:² # ³
this more closely, we observe that must be a set. For if has an element 44
of multiplicity greater than , we can split into two disjoint parts: . 4
396 Advanced Linear Algebra
.~ .r .4 44ZZ Z
where are the tensors that have in positions and :. 4Z
. ~¸ nÄn nÄn nÄn ¹4Z
position position
Then fixes each element of and sends the elements of to other² ³ 44ZZ Z..
elements of . Hence, applying to the corresponding decomposition of .4ZZ² ³
:² # ³4 :
:² # ³ ~ :b :4 44ZZ Z
gives
c²: b: ³ ~ c: ~ : ~ : b :44 4 4ZZ Z Z Z Z
44 ²³ ²³
and so , whence . Thus, is a set.:~ : ² # ³ ~ 44Z4
Now, since for any , :
.~ ¸! ! . ¹44
equation (14.4) implies that
² c ³ !~ ! ~ !~ !
89
!. !. !. !.!! ! !
44 4 4c
which holds if and only if , or equivalently,
c! ! ~² c ³
!!~² c ³
for all and . Choosing , where !. : "~" ~ nÄn4 4
Ä " Á ! , as standard-bearer, if denotes the permutation for which
"Á!²"³ ~ ! , then
!"~² c ³"Á!
Thus, is antisymmetric if and only if it has the form#
#~ ² c ³ !89
44
!.
4"Á!
where and the sum is over a family of .4"~£ sets
In summary, the symmetric tensors are
Tensor Products 397
#~ !89
44
!.
4
where is a multiset and the antisymmetric tensors are4
#~ ² c ³ !89
44
!.
4"Á!
where is a set.4
We can simplify these expressions considerably by representing the inside sums
more succinctly. In the symmetric case, define a surjective linear map
¢; ²=³¦- ´ ÁÃÁ µ
by
² nÄn ³~ vÄv
and extending by linearity. Since takes every member of to the same .4
monomial , where , we have"~ vÄv Ä
#~ ! ~ . "89 89((
4444 4 4
!.4
In the antisymmetric case, define a surjective linear map
¢; ²=³¦- ´ ÁÃÁ µc
by
² nÄn ³~ wÄw
and extending by linearity. Since
! ~ ²c³ ""Á!4
we have
#~ ² c ³ !
~"
~. "89
89
((44
!.
444
!.
444 44"Á!
4
398 Advanced Linear Algebra
Thus, in both cases,
#~ . " ((
444 4
where with and" ~ nÄn Ä4
" ~ vÄv " ~ wÄw4 4 or
depending on whether is symmetric or an tisymmetric. However, in either case, #
the monomials are linearly independent for distinct multisets/sets . "44
Therefore, if then for all multisets/sets . Hence, if #~ . ~ 4 44((
char²-³ ~ ~ # ~ , then and so . This shows that the restricted maps4
OO:; ²=³ (; ²=³ and are isomorphisms.
Theorem 14.16 Let be a finite-dimensional vector space over a field with=-
char²-³ ~ .
1 The symmetric tensor space is isomorphic to the algebra) :; ²=³
-´ Á Ã Á µ of homogeneous polynomials, via the isomorphism
45 ÁÃÁ ÁÃÁ nÄn ~ ² vÄv ³
2 For , the antisymmetric tensor space is isomorphic to the ) ( ; ² =³
algebra of anticommutative homogeneous polynomials of -´ Á Ã Á µc
degree , via the isomorphism
45 ÁÃÁ ÁÃÁ nÄn ~ ² wÄw ³
The direct sum
:;²=³~ :; ²=³-´ ÁÃÁ µ
~B
is called the of and the direct sum symmetric tensor algebra =
(;²=³ ~ (; ²=³ - ´ ÁÃÁ µ
~
c
is called the or the of . These antisymmetric tensor algebra exterior algebra =
vector spaces are graded algebras, where the product is defined using the vector
space isomorphisms described in the previous theorem to move the products of
- ´ Á Ã Á µ - ´ Á Ã Á µ : ; ² =³ ( ; ² =³ c and to and , respectively.
Thus, restricting the domains of the ma ps gives a nice description of the
symmetric and antisymmetric tensor algebras, when . However, char²-³ ~
there are many important fields, such as finite fields, that have nonzero
characteristic. We can proceed in a different, albeit somewhat less appealing,
Tensor Products 399
manner that holds regardless of the char acteristic of the base field. Namely,
rather than restricting the domain of in order to get an isomorphism, we can
factor out by the kernel of .
Consider a tensor
#~ : ² # ³~ ! ps
qt 444!
!. ²#³4
Since sends elements of different groups to different .² # ³ ~ ¸ ! Á Ã Á ! ¹4
monomials in or , it follows that if and -´ Á Ã Á µ - ´ Á Ã Á µ # ²³ c
ker
only if for all , that is, if and only if²: ²#³³ ~ 44
! !!b Ä b !~
In the symmetric case, is constant on and so if and only if .² # ³ # ² ³4 ker
!!bÄb ~
In the antisymmetric case, where and so !~ ² c ³ ! ² !³ ~ ! Á
Á
# ² ³ker if and only if
²c³ bÄb²c³ ~Á Á
!!
In both cases, we solve for and substitute into . In the symmetric case, !4 :² # ³
!! ! ~c cÄc
and so
: ² # ³ ~! b Ä b! ~² ! c ! ³ b Ä b² ! c ! ³4! ! ! !
In the antisymmetric case,
!! ! Á Á~c²c³ cÄc²c³
and so
:² # ³ ~ !b Ä b !
~ ²²c³ ! c! ³bÄb ²²c³ ! c! ³4! !
! !
Á Á
Since , it follows that and therefore , is in the span of tensors of ! : ² # ³ #48
the form in the symmetric case and in the ²!³c! ²c³ ²!³c!
antisymmetric case, where and . 8: !
Hence, in the symmetric case,
ker² ³0 º ² ! ³c!! Á :» 8
and since , it follows that . In the antisymmetric ² ² ! ³c! ³~ ² ³~0 ker
case,
400 Advanced Linear Algebra
ker² ³ 0 º²c³ ²!³c! ! Á : » 8
and since , it follows that . ² ² c ³ ² ! ³c! ³~ ² ³~0 ker
We now have quotient-space characterizations of the symmetric and
antisymmetric tensor spaces that do not place any restriction on the
characteristic of the base field.
Theorem 14.17 Let be a finite-dimensional vector space over a field .=-
1 The surjective linear map defined by) ¢; ²=³¦- ´ ÁÃÁ µ
45 ÁÃÁ ÁÃÁ n Ä n ~ v Ä v
has kernel
0~ º ² ! ³ c ! ! Á : »8
and so
;² = ³
0- ´ ÁÃÁ µ
The vector space is also referred to as the ;² = ³ ° 0symmetric tensor
space of degree of . =
2 The surjective linear map defined by) ¢; ²=³¦- ´ ÁÃÁ µc
45 ÁÃÁ ÁÃÁ n Ä n ~ w Ä w
has kernel
0 ~ º²c³ ²!³c! ! Á : »
8
and so
;² = ³
0- ´ ÁÃÁ µ
c
The vector space is also referred to as the ;² = ³ ° 0antisymmetric tensor
space exterior product space or of degree of . =
The isomorphic exterior spaces and are usually denoted by (; ²= ³ ; ²= ³°0
= (;²=³ ;²=³°0 and the isomorphic exterior algebras and are usually
denoted by . =
Theorem 14.18 Let be a vector space of dimension .=
Tensor Products 401
1 The dimension of the symmetric tensor space is equal to the) :; ²=³
number of monomials of degree in the variables and this is ÁÃÁ
dim²:; ²= ³³ ~bc
67
2 The dimension of the exterior tensor space is equal to the number of) ²= ³
words of length in ascending order over the alphabet , ~ ¸ Á Ã Á ¹
and this is
dim²² = ³ ³ ~
67
Proof. For part 1), the dimension is equal to the number of multisets of size
taken from an underlying set of size . Such multisets correspond ¸ ÁÃÁ ¹
bijectively to the solutions, in nonnegative integers, of the equation
%b Ä b %~
where is the multiplicity of in the multiset. To count the number of%
solutions, invent two symbols and . Then any solution to the %° % ~
previous equation can be described by a sequence of 's and 's consisting of %°
%° % °'s followed by one , followed by 's and another , and so on. For example,
if and , the solution corresponds to the sequence~
~ bbb~
%%%°%°°%%
Thus, the solutions correspond bijectively to sequences consisting of 's and %
c° 's. To count the number of such sequences, note that such a sequence can
be formed by considering “blanks” and selecting of these blanks for bc
the 's. This can be done in%
67bc
ways.
The Universal Property
We defined tensor products through a universal property, which as we have seen
is a powerful technique for determini ng the properties of tensor products. It is
easy to show that the symmetric tensor spaces are universal for symmetric
multilinear maps and the antisymmetric tensor spaces are universal for
antisymmetric multilinear maps.
Theorem 14.19 Let be a finite-dimensional vector space with basis=
¸ ÁÃÁ ¹ .
1 The pair , where is the) ²- ´% ÁÃÁ% µÁ!³ !¢= ¦- ´% ÁÃÁ% µ d
multilinear map defined by
402 Advanced Linear Algebra
!² ÁÃÁ ³~ vÄv
is universal for symmetric -linear m aps with domain ; that is, for any =d
symmetric -linear map where is a vector space, there is a ¢ = ¦ < <d
unique linear map for which ¢- ´% ÁÃÁ% µ¦<
² vÄv ³~² ÁÃÁ ³
2 The pair , where is the) ²- ´% ÁÃÁ% µÁ!³ !¢= ¦- ´% ÁÃÁ% µcd c
multilinear map defined by
!² ÁÃÁ ³~ wÄw
is universal for antisymmetric -linear maps with domain ; that is, for =d
any antisymmetric -linear map where is a vector space, ¢ = ¦ < <d
there is a unique linear map for which ¢- ´% ÁÃÁ% µ¦<c
² wÄw ³~² ÁÃÁ ³
Proof. For part 1), the property
² vÄv ³~² ÁÃÁ ³
does indeed uniquely define a linear tr ansformation , provided that it is well-
defined. However,
vÄv ~ vÄv
if and only if the multisets and are the same, which ¸ ÁÃÁ ¹ ¸ ÁÃÁ ¹
implies that , since is symmetric. ² ÁÃÁ ³~² ÁÃÁ ³
For part 2), since is antisymmetric, it is completely determined by the fact that
it is alternate and by its values on the basis of ascending words . w Ä w
Accordingly, the condition
² wÄw ³~² ÁÃÁ ³
uniquely defines a linear transformation .
The Symmetrization Map
When , we can define a linear map , called char² - ³ ~ : ¢ ;² = ³ ¦ : ;² = ³
the , bysymmetrization map
:!~ !
[
:
Since , we have ~
Tensor Products 403
²:!³ ~ ! ~ ! ~ ! ~ :!
[ [ [
: : :
and so is symmetric. The reason for the factor is that if is a symmetric:! °[ #
tensor, then and so #~#
:#~ #~ #~#
[ [
: :
that is, the symmetrization map fixes all symmetric tensors. It follows that for
any tensor , !; ² =³
:!~: !
Thus, is idempotent and is therefore the projection map of onto:; ² = ³
im²:³ ~ :; ²= ³.
The Determinant
The universal property for antisymmetr ic multilinear maps has the following
corollary.
Corollary 14.20 Let be a vector space of dimension over a field . Let= -
,~² ÁÃÁ ³ = be an ordered basis for . Then there is at most one
antisymmetric -linear form for which ¢ = ¦ -d
² ÁÃÁ ³ ~
Proof. According to the universal property for antisymmetric -linear forms, for
every antisymmetric -linear form satisfying , ¢ = ¦ - ² Á Ã Á ³ ~ d
there is a unique linear map for which ¢= ¦ -
² wÄw ³~²ÁÃÁ ³~
But has dimension and so there is only one linear map = ¢ = ¦ -
with . Therefore, if and are two such forms, then² wÄw ³~
~~ , from which it follows that
~ k!~ k!~
We now wish to construct an antisymmetric form , which is unique ¢= ¦ -d
by the previous theorem. Let be a basis for . For any , write for 8 =# = ´ # µ 8Á
the th coordinate of the coordinate matrix . Thus,´ # µ 8
#~ ´ # µ
Á 8
For clarity, and since we will not change the basis, let us write for . ´#µ ´#µÁ 8
404 Advanced Linear Algebra
Consider the map defined by ¢= ¦ -d
² #Á Ã Á #³~ ² c ³´ #µ Ä ´ #µ
:
Then is multilinear since
²# b" ÁÃÁ# ³ ~ ²c³ ´# b" µ Ä´# µ
~² c ³ ² ´ # µ b ´ " µ ³ Ä ´ # µ
~ ²c³ ´# µ Ä´# µ
b ²c³ ´" µ Ä´#
:
:
:
:
()
µ
~ ²# ÁÃÁ# ³b²" Á# ÁÃÁ# ³
and similarly for any coordinate position.
To see that is alternating, and therefore antisymmetric since , ² - ³ £ char
suppose for instance that . For any permutation , let #~ # :
Z~² ³
Then for andZ%~ % %£ Á
ZZ~ ~ and
Hence, . Also, since , if the sets and intersect, ZZ Z Z Z£² ³ ~ ¸ Á ¹ ¸ Á ¹
then they are identical. Thus, the dis tinct sets form a partition of . It ¸Á ¹ :Z
follows that
² #Á #Á #Á Ã Á #³~ ² c ³´ #µ ´ #µ Ä ´ #µ
~ ²c³ ´# µ ´# µ Ä´# µ b²c³ ´# µ ´# µ Ä´# µ
:
¸Á ¹
>?
ZZ
ZZ Z
pairs
But
´ #µ ´ #µ ~´ #µ ´ #µ ZZ
and since , the sum of the two terms involving the pair ²c³ ~ c²c³ ¸ Á ¹ZZ
is . Hence, . A similar argument holds for any coordinate ²# Á# ÁÃÁ# ³~
pair.
Tensor Products 405
Finally,
² Á Ã Á ³~ ² c ³´ µ Ä ´ µ
~² c ³ Ä
~
:
:Á Á
Thus, the map is the unique antisymmetric -linear form on for which =d
² ÁÃÁ ³ ~ .
Under the ordered basis , we can view as the space of ;~² ÁÃÁ³ = -
coordinate vectors and view as the space of matrices, via the =4 ² - ³ d d
isomorphism
²# ÁÃÁ# ³ª´# µ Ä ´# µ
ÅÅ
´# µ Ä ´# µ
vy
wz
where all coordinate matrices are with respect to . ;
With this viewpoint, becomes an an tisymmetric -form on the columns of a
matrix given by(~² ³ Á
²(³ ~ ²c³ Ä
:Á Á
This is called the of the matrix . determinant (
Properties of the Determinant
Let us explore some of the properties of the determinant function.
Theorem 14.21 If , then(4² -³
²(³ ~ ²( ³!
Proof. We have
²(³ ~ ²c³ Ä
~² c ³ Ä
~² c ³ Ä
~ ² (³
:Á Á
:Á Á
:Á Á
!
cc
c c
as desired.
406 Advanced Linear Algebra
Theorem 14.22 If , then(Á) 4 ²-³
²()³ ~ ²(³²)³
Proof. Consider the map defined by ¢ 4² - ³ ¦ -(
²?³~²(?³(
We can consider as a function on the columns of and think of it as a ?(
composition
¢²? ÁÃÁ? ³ª²(? ÁÃÁ(? ³ª²(?³(²³ ²³ ²³ ²³
Each step in this map is multilinear and so is multilinear. It is also clear that(
(( is antisymmetric and so is a scalar multiple of the determinant function,
say . Then ²?³~²?³(
²(?³ ~ ²?³ ~ ²?³ (
Setting gives and so?~0 ² ( ³~
²(?³ ~ ²(³²?³
as desired.
Theorem 14.23 A matrix is invertible if and only if . (4² -³ ² ( ³£
Proof. If is invertible, then and so74² - ³ 7 7 ~0 c
²7³²7 ³ ~ c
which shows that and . Conversely, any matrix ²7³ £ ²7 ³ ~ °²7³c
(4² -³ is equivalent to a diagonal matrix
(~7+ 8
where and are invertible and is diagonal with 's and 's on the main78 +
diagonal. Hence,
²(³ ~ ²7³²+³²8³
and so if , then , which happens if and only if , ²(³ £ ²+³ £ + ~ 0
whence is invertible.(
Exercises
1. Show that if is a linear map and is bilinear, ¢> ¦? ¢<d= ¦>
then is bilinear.k¢<d= ¦?
2. Show that the only map that is both linear and -linear for is the ()
zero map.
3. Find an example of a bilinear map whose image ¢= d= ¦>
im²³ ~ ¸² " Á # ³ " Á # = ¹ > is not a subspace of .
Tensor Products 407
4. Let be a basis for and let be a basis89~¸ " 0¹ < ~¸ # 1¹
for . Show that the set=
:~¸ " n# 0Á1¹
is a basis for by showing that it is linearly independent and spans. <n=
5. Prove that the following property of a pair with ²>Á¢<d= ¦>³
bilinear characterizes the tensor product up ²<n=Á!¢<d= ¦<n=³
to isomorphism, and thus could have been used as the definition of tensor
product: For a pair with bilinear if is a basis ²>Á¢<d= ¦>³ ¸"¹
for and is a basis for , then is a basis for .<¸ # ¹ = ¸ ² " Á # ³ ¹ >
6. Prove that . <n==n<
7. Let and be nonempty sets. Use the universal property of tensor ?@
products to prove that . << <?d@ ? @n
8. Let and . Assuming that , show that "Á" < #Á# = "n# £ ZZ
"n#~"n# " ~ " # ~ # £ZZ Z Zc if and only if and , for .
9. Let be a basis for and be a basis for . Show that any89~¸ ¹ < ~¸ ¹ =
function can be extended to a linear function ¢ d ¦>89
¢< n= ¦> . Deduce that the function can be extended in a unique
way to a bilinear map . Show that all bilinear maps are ¢< d= ¦>V
obtained in this way.
10. Let be subspaces of . Show that:Á: <
²: n= ³q²: n= ³ ²: q: ³n=
11. Let and be subspaces of vector spaces and , :< ;= < =
respectively. Show that
²:n=³q²<n;³:n;
12. Let and be subspaces of and , respectively. : Á : <; Á ; = <=
Show that
²: n; ³q²: n; ³ ²: q: ³n²; n; ³
13. Find an example of two vector spaces and and a nonzero vector <=
%<n= that has at least two distin ct not including order of the terms()
representations of the form
%~ "n#
~
where the 's are linearly independent and so are the 's. "#
14. Let denote the identity operator on a vector space . Prove that? ?
=>= n >p~ .
408 Advanced Linear Algebra
15. Suppose that and . 2 2ZZ¢< ¦=Á ¢= ¦> ¢< ¦= Á ¢= ¦>
Prove that
²k³ p ²k³ ~ ²p³ k ²p³
16. Connect the two approaches to extending the base field of an -space to -=
2 at least in the finite-dimensional case by showing that()
-n2 ² 2 ³- .
17. Prove that in a tensor product for which not all vectors <n< ² <³ dim
have the form for some . : Suppose that are "n# "Á#< "Á#< Hint
linearly independent and consider . "n#b#n"
18. Prove that for the block matrix
4~()
*>?
block
we have . ²4³ ~ ²(³²*³
19. Let . Prove that if either or is invertible, then the (Á) 4 ²-³ ( )
matrices are invertible except for a finite number of 's. (b )
The Tensor Product of Matrices
20. Let be the matrix of a linear operator with respect to(~² ³ ² =³ Á B
the ordered basis . Let be the matrix of a linear 7~² "ÁÃÁ"³ )~² ³ Á
operator with respect to the ordered basis .B 8² = ³ ~ ² # Á Ã Á # ³
Consider the ordered basis ordered lexicographically; that is 9~² " n#³
" n# " n# M ~M M if or and . Show that the matrix of
9n with respect to is
(n)~))Ä)
))Ä)
ÅÅ Å
))Ä)ps
ruru
qtÁ Á Á
Á Á Á
Á Á Áblock
This matrix is called the , or tensor product Kronecker product direct
product of the matrix with the matrix . ()
21. Show that the tensor product is not, in general, commutative.
22. Show that the tensor product is bilinear in both and . (n) ( )
23. Show that if and only if or . (n)~ (~ )~
24. Show that
a )²(n)³ ~ ( n)!!!
b w h e n )( )²(n)³ ~ ( n) - ~iiid
25. Show that if , then as row vectors . " Á#- "#~" n#! !()
26. Suppose that and are matrices of the given sizes. (Á ) Á * +Á Á Á Á
Prove that
²(n)³²* n+³ ~ ²(*³n²)+³
Discuss the case . ~~
Tensor Products 409
27. Prove that if and are nonsingular, then so is and () ( n )
²(n)³ ~ ( n)c c c
28. Prove that . tr tr tr²(n)³ ~ ²(³h ²)³
29. Suppose that is algebraically closed. Prove that if has eigenvalues -(
ÁÃÁ ) ÁÃÁ and has eigenvalues , both lists including
multiplicity, then has eigenvalues , again (n) ¸ Á ¹
counting multiplicity.
30. Prove that . det det det²( n) ³ ~ ² ²( ³³ ² ²) ³³Á Á Á Á
Chapter 15
Positive Solutions to Linear Systems:
Convexity and Separation
It is of interest to determine c onditions that guarantee the existence of positive
solutions to homogeneous systems of linear equations
(% ~
where .( ² ³CsÁ
Definition Let .#~² ÁÃÁ ³ s
1 is , written , if)## nonnegative
for all
()The term is also used for this property. The set of all nonnegative positive
vectors in is the in ssnonnegative orthant À
2 is , written , if is nonnegative but not , that is, if)## # strictly positive
for all and for some
The set of all strictly positive vectors in is the ss
b strictly positive
orthant in sÀ
3 is , written , if)## strongly positive
for all
The set of all strongly positive vectors in is the ssbbstrongly positive
orthant in sÀ
We are interested in conditions under which the system has strictly (% ~
positive or strongly positive solutions. Since the strictly and strongly positive
orthants in are not subspaces of , it is difficult to use strictly linear ss
methods in studying this issue: we must also use geometric methods, in
particular, methods of convexity.
412 Advanced Linear Algebra
Let us pause briefly to consider an important application of strictly positive
solutions to a system . If is a strictly positive solution (%~ ?~²% ÁÃÁ% ³
to this system, then so is the vector
& ''~? ~² % Á Ã Á % ³ ~ ² Á Ã Á ³
%%
which is a , that is, and . Moreover, probability distribution ~
if is a strongly positive solution, then has the property that each probability? &
is positive.
Now, the product is the expected valu e of the columns of with respect to ((&
the probability distribution . Hence, has a strictly (strongly) positive & (% ~
solution if and only if there is a strictly (strongly) positive probability
distribution for which the columns of have expected value . If each column (
of represents the possible payoffs from a game of chance, where each row is a(
different possible outcome of the game, then the game is fair when the expected
value of the columns is . Thus, has a strictly (strongly) positive ( % ~
solution if and only if the game with payoffs and probabilities is fair. ?( ?
As another (related) example, in discre te option pricing models of mathematical
finance, the absence of arbitrage opportunities in the model is equivalent to the
fact that a certain vector describing the gain s in a portfolio does not intersect the
strictly positive orthant in . As we will s ee in this chapter, this is equivalent s
to the existence of a strongly positive solution to a homogeneous system of
equations. This solution, when normalized to a probability distribution, is called
a .martingale measure
Of course, the equation has a strictly positive solution if and only if (% ~
ker²(³ contains a strictly positive vector, that is, if and only if
ker²(³ ~ ²(³ RowSpace
meets the strictly positive orthant in . Thus, we wish to characterize thes
subspaces of for which meets the strictly positive orthant in , in ::ss
symbols,
:q £ J
bs
for these are precisely the row spaces of the matrices for which has a (( % ~
strictly positive solution. A similar statement holds for strongly positive
solutions.
Looking at the real plane , we can di vine the answer with a picture. A one- s
dimensional subspace of has the property that its orthogonal complement :s
:: % meets the strictly positive orthant quadr ant in if and only if is the - () s
axis, the -axis or a line with negative slope. For the case of the strongly &
Positive Solutions to Linear Systems: Convexity and Separation 413
positive orthant, must have negative sl ope. Our task is to generalize this to :
s.
This will lead us to the following resu lts, which are quite intuitive in and : ss
:q £ J ¯ : q ~ J
bb bss (15.1)
and
:q £ J ¯ : q ~ J
bb bss (15.2)
Let us translate these statements in to the language of the matrix equation
(% ~ : ~ ²(³ : ~ ²(³ . If , then and so we have RowSpaceker
ker²(³q £ J ¯ ²(³q ~ Jss
bb b RowSpace
and
ker²(³q £ J ¯ ²(³q ~ Jss
bb b RowSpace
Now,
RowSpace ²(³q ~ ¸#( #( ¹s
b
and
RowSpace ²(³q ~ ¸#( #( ¹s
bb
and so these statements become
( %~ ¯ ¸ # (# ( ¹~J has a strongly positive solution
and
( %~ ¯ ¸ # (# ( ¹~J has a strictly positive solution
We can rephrase these results in the form of a , that theorem of the alternative
is, a theorem that says that exactly one of two conditions holds.
Theorem 15.1 Let .( ² ³CsÁ
1 Exactly one of the following holds:)
a for some strongly positive . )(" ~ " s
b for some . )#( # s
2 Exactly one of the following holds:)
a for some strictly positive . )(" ~ " s
b for some . )#( # s
Before proving Theorem 15.1, we require some background.
Convex, Closed and Compact Sets
We shall need the following concepts.
414 Advanced Linear Algebra
Definition
1 Let . Any linear combination of the form)%Á Ã Á % s
!% bÄb!%
where and is called a of! ! bÄb! ~ convex combination
the vectors . %Á Ã Á %
2 A subset is if whenever , then the line segment) ? % Á&?sconvex
between and also lies in , in symbols, %& ?
¸ ! %b² c! ³ &! ¹?
3 A subset is if whenever is a convergent sequence of) ? ² %³s closed
elements of , then the limit is also in . ??
4 A subset is if it is both closed and bounded.) ?scompact
5 A subset is a if implies that for all .) ? %? %? scone
We will also have need of the following facts from analysis.
1 A continuous function that is defined on a compact set in takes on) ?s
maximum and minimum values at some points within the set . ?
2 A subset of is compact if and only if every sequence in has a) ??s
subsequence that converges to a point in . ?
Theorem 15.2 Let and be subsets of . Define?@ s
?b@~¸ b?Á@¹
1 If and are convex, then so is )?@ ? b @
2 If is compact and is closed, then is closed.)?@? b @
Proof. For 1), let and be in . The line segment between %b & %b & ? b @
these two points is
!²% b& ³b²c!³²% b& ³
~ ´!% b²c!³% µb´!& b²c!³& µ ? b@
for and so is convex.! ?b@
For part 2), let be a convergent sequence in . Suppose that %b & ? b @
%b &¦ ' ' ? b @ % . We must show that . Since is a sequence in the
compact set , it has a convergent subsequence whose limit lies in . ?% % ?
Since and we can conclude that . Since %b &¦ ' %¦ % &¦ ' c % @
is closed, it follows that and so . 'c%@ '~%b² 'c% ³?b@
Convex Hulls
We will also have use for the notion of the smallest convex set containing a
given set.
Positive Solutions to Linear Systems: Convexity and Separation 415
Definition The of a set of vectors in is the convex hull :~¸% ÁÃÁ% ¹ s
smallest convex set in that contains . We will denote the convex hull of s::
by .9²:³
Here is a characterization of convex hulls.
Theorem 15.3 Let be a set of vectors in . Then the convex:~¸% ÁÃÁ% ¹ s
hull is the set of all convex combinations of vectors in , that is,9"²:³ :
9"² :³~ !% bÄb!% ! Á ! ~ DE
Proof. Clearly, if is a convex set that contains , then also contains . +: + "
Hence . To prove the reverse inclusion, we need only show that is"9 "² : ³
convex, since then implies that . So let : ² : ³"9 "
?~!% bÄb!%
@ ~ % bÄb %
be in . If and then"b~ Á
?b@ ~²!% bÄb!%³b² % bÄb %³
~ ²! b ³% bÄb²! b ³%
But this is also a convex combination of the vectors in , because :
! b ²b³h ² Á!³~ ² Á!³ max max
and
~ ~ ~
² !b ³~ !b ~b~
Thus, .? b@ "
Theorem 15.4 The convex hull of a set of vectors 9²:³ : ~ ¸% ÁÃÁ% ¹ finite
in is a compact set.s
Proof. The set
+~ ²!ÁÃÁ! ³! Á ! ~ DE
is closed and bounded in and therefore compact. Define a function s
¢+¦ !~²!ÁÃÁ! ³s as follows: If , then
²!³~! % bÄb! %
To see that is continuous, let and let . Given ~ ² Á Ã Á ³ 4 ~ ² % ³ max ))
c! ° 4, if then))
416 Advanced Linear Algebra
(( ) ) c ! c ! 4
and so
)) ) )
(( ) ) (( ) )
))² ³c²!³ ~ ² c!³% bÄb² c!³%
c ! % b Ä b c ! %
4 c!
~
Finally, since , it follows that is compact. ²+³~ ²:³ ²:³99
Linear and Affine Hyperplanes
We next discuss hyperplanes in . A in is an - sslinear hyperplane ²c³
dimensional subspace of . As such, it is the solution set of a linear equation s
of the form
% bÄb % ~
or
º5Á%» ~
where is nonzero and . Geometrically5~² ÁÃÁ ³ %~²% ÁÃÁ% ³
speaking, this is the set of all vectors in that are perpendicular (normal) tos
the vector . 5
An , or just , in is a linear hyperplane that has affine hyperplane hyperplane s
been translated by a vector. Thus, it is th e solution set to an equation of the form
²% c³bÄb ²% c ³~
or equivalently,
º5Á%» ~
where . We denote this hyperplane by~ bÄb
>s²5Á³~¸% º5Á%»~¹
Note that the hyperplane
>s²5Á 5 ³ ~ ¸% º5Á%» ~ 5 ¹)) ))
contains the point , which is th e point of closest to the origin, 5² 5 Á 5 ³ > ))
since Cauchy's inequality gives
)) )) ) )5~ º 5 Á % » 5%
and so for all . Moreover, we leave it as an )) ) ) ))5% % ² 5 Á5³ >
Positive Solutions to Linear Systems: Convexity and Separation 417
exercise to show that any hyperplane has the form for an >²5Á 5 ³))
appropriate vector . 5
A hyperplane defines two closed half-spaces
>s
>sb
c²5Á³ ~ ¸% º5Á%» ¹
²5Á³ ~ ¸% º5Á%» ¹
and two disjoint open half-spaces
>s
>sbk
ck²5Á³ ~ ¸% º5Á%» ¹
²5Á³ ~ ¸% º5Á%» ¹
It is clear that
>>>bc²5Á³q ²5Á³~ ²5Á³
and that the sets , and form a partition of . >> > sbckk ²5Á³ ²5Á³ ²5Á³
If and , we let5 ?ss
º5Á?» ~ ¸º5Á%» % ?¹
and write
º5Á?»
to denote the fact that for all . º5Á%» % ?
Definition Two subsets and of are by a hyperplane ?@ sstrictly separated
>>²5Á³ ? ²5Á³ @ if lies in one open half-space determined by and lies in
the other open half-space; in symbols, one of the following holds:
1)º5Á?»º5Á@»
2)º5Á@»º5Á?»
Note that 1) holds for and if and only if 2) holds for and , and so we 5 c5 c
need only consider one of the conditions to demonstrate that two sets and ?@
are strictly separated. Specifically, if 1) fails for all and , then thenot 5
condition
ºc5Á@» c ºc5Á?»
also fails for all and and so 2) also fails for all and , whence and 5 5 ?@
are not strictly separated.
Definition Two subsets and of are by a hyperplane ?@ sstrongly separated
>²5Á³ if there is an for which one of the following holds:
1)º 5Á?»cbº 5Á@»
2)º 5Á@»cbº 5Á?»
418 Advanced Linear Algebra
As before, we need only consider one of the conditions to show that two sets are
not strongly separated. Note also that if
º5Á%» º5Á@»
for , then and are stongly separated by the hyperplane % @sss
>675Ábº5Á%»
Separation
Now that we have the preliminaries out of the way, we can get down to some
theorems. The first is a well-known that is the basis for separation theorem
many other separation theorems. It says that if a closed convex set does *s
not contain a vector , then can be separated from . * strongly
Theorem 15.5 Let be a closed convex subset of .* s
1 contains a vector of minimum norm, that is, there is a unique)*5 unique
vector for which5*
)) ) )5%
for all .%*Á%£5
2 If , then lies in the closed half-space)¤* *
>b23))5Á 5 bº5Á»
that is,
º5Á*» 5 bº5Á» º5Á» ))
where is the unique vector of minimum norm in the closed convex set5
*c~¸ c*¹
Hence, and are strongly separated by the hyperplane*
>89))5Á5b º 5 Á »
Proof. For part 1), if then this is the unique vector of minimum norm, so *
we may assume that . It follows that no two distinct elements of can be ¤* *
negative scalar multiples of each other. For if and were in , where % % * Á
then taking gives !~c ° ² c ³
~ !%b²c!³% *
which is false.
Positive Solutions to Linear Systems: Convexity and Separation 419
We first show that contains a v ector of minimum norm. Recall that the *5
Euclidean norm (distance) is a con tinuous function. Although need not be *
compact, if we choose a real number such that the closed ball
)² ³~¸ ' ' ¹ s))
intersects , then that intersection is both closed and bounded ** ~ * q ) ² ³Z
and so is compact. The norm function therefore achieves its minimum on , *Z
say at the point . It is clear that if for some , then 5** # 5 #*Z)) ) )
#* 5 5Z, in contradiction to the minimality of . Hence, is a vector of
minimum norm in . *
We establish uniqueness first for closed line segments in . If ´"Á#µ " ~ #s
where , then
)) ( ( ) )!"b²c!³# ~ !b²c!³ #
is smallest when for and for . Assume that and are !~ !~ " #
not scalar multiples of each other and suppose that in have %£& ´ " Á# µ
minimum norm . If then since and are also not scalar '~² %b& ³ ° % &
multiples of each other, the Cauchy-Schwarz inequality is strict and so
P'P ~ %b&
~ ²P%P bº%Á&»bP&P ³
² b %& ³
~
))
)) ))
which contradicts the minimality of . Thus, has a unique point of´ " Á # µ
minimum norm.
Finally, if also has minimum norm, then and are points of minimum %* 5 %
norm in the line segment and so . Hence, has a unique ´5Á%µ * % ~ 5 *
element of minimum norm.
For part 2), suppose the result is true when . Then implies that ¤* ¤*
¤*c 5*c and so if has smallest norm, then
º5Á* c» 5 ))
Therefore,
º5Á*» 5 bº5Á» º5Á» ))
and so and are strongly separated by the hyperplane*
420 Advanced Linear Algebra
>23 )) 5Á²°³ 5 bº5Á»
Thus, we need only prove part 2) for , that is, we need only prove that ~
º5Á*» 5 ))
If there is a nonzero for which %*
º5Á%» 5 ))
then and)) ) )5%
º5Á%» ~ 5 c ))
for some . Then for the open line segment with ² ! ³~! 5b² c! ³ %
! , we have
P²!³P ~ !5 b²c!³%
~ ! P5P b!²c!³º5Á%»b²c!³ P%P
²!c!³P5P c!²c!³ b²c!³P%P
~ ²cP5P b b % ³! b² 5 c cP%P ³!b %
))
)) ) ) ))
Let denote the final expression above, which is a quadratic in . It is easy to²!³ !
see that has its minimum at the in terior point of the line segment ²!³ ´5Á%µ
corresponding to
!~ c5 b b%
c5 b b%
)) ) )
)) ) )
and so , which is a contradiction. )) ) )²! ³ ²! ³²³~ 5
The next result brings us closer to our goal by replacing a single vector with a
subspace disjoint from . However, we mu st also require that be bounded, :* *
and therefore compact.
Theorem 15.6 Let be a compact convex subset of and let be a subspace*: s
of such that . Then there exists a nonzero such thats *q:~J 5:
º5Á%» 5 ))
for all . Hence, the hyperplane strongly separates and% * ²5Á 5 °³ : > ))
*.
Proof. Theorem 15.2 implies that the set is closed and convex. :b*
Furthermore, implies that and so Theorem 15.5 implies *q:~J ¤:b*
that can be strongly separated from the origin. Hence, there is a nonzero:b*
5s such that
Positive Solutions to Linear Systems: Convexity and Separation 421
º5Á »bº5Á»~º5Á b» 5 ))
for all and . But if for some , then we can replace : * º 5Á »£ :
by an appropriate scalar multiple of in order to make the left side of this
inequality negative, which is im possible. Hence, for all , that º5Á » ~ :
is, and5:
º5Á*» 5 ))
We can now prove (15.1) and (15.2).
Theorem 15.7 Let be a subspace of .: s
1 if and only if ):q ~J : q £Jss
bb b
2 if and only if ):q ~J : q £Jss
bb b
Proof. In both cases, one direction is easy. It is clear that there cannot exist
vectors and that are orthogonal. Hence, and " # :qss s
bb b b
:q :q £ J
bb bbss cannot both be nonempty and so implies
:q ~J :q : qss s
bb b b . Also, and cannot both be nonempty and so
: q £J :q ~J
bb bss implies that .
For the converse in part 1), to prove that
:q ~J ¬ : q £Jss
bb b
a good candidate for an element of would be a normal to a :q
bbs
hyperplane that separates from a subset of . Note that our separation : s
b
theorems do not allow us to separate from , because is not compact. So :ss
bb
consider instead the convex hull of the standard basis vectors in " ÁÃÁ
s
b:
" '~ ¸ ! b Ä b ! ! Á !~ ¹
which is compact. Moreover, im plies that and so Theorem "s "q : ~ J
b
15.6 implies that there is a nonzero vector such that 5~² ÁÃÁ ³:
º5Á » 5 ))
for all Taking gives" À ~
~ º 5 Á » 5 ))
and so , which is therefore nonempty.5: q
bbs
To prove part 2 , again we note th at there cannot exist orthogonal vectors )
" # :q : qss s s
bb b bb b and and so and cannot both be nonempty.
Thus, implies that .: q £J :q ~J
bb bss
422 Advanced Linear Algebra
To finish the proof of part 2), we must prove that
:q ~J ¬ : q £Jss
bb b
Let be a basis for . Then if and8~¸ )ÁÃÁ)¹ : 5~² ÁÃÁ³:
only if for all . In matrix terms, if5)
4~² ³~² ) ) Ä)³ Á
has rows , then if and only if , that is, 9Á Ã Á 9 5: 5 4~
9 bÄb 9 ~
Now, contains a strictly positive vector if and only if this:5 ~ ² Á Ã Á ³
equation holds, where for all and for some . Moreover, we may
assume without loss of generality that , or equivalently, that is in the '~
convex hull of the row space of . Hence, 9 4
:q £ J ¯
bs9
Thus, we wish to prove that
:q ~J ¬ s9
bb
or equivalently,
¤ ¬ :q £J9s
bb
Now we have something to separate. Since is closed and convex, Theorem 9
15.5 implies that there is a nonzero vector for which )~² ÁÃÁ ³ s
º)Á » ) 9 ))
Consider the vector
#~) bÄb) :
The th coordinate of is#
bÄb ~º)Á9» ) Á Á ))
and so is strongly positive. Hence, , which is therefore ## : q s
bb
nonempty.
Inhomogeneous Systems
We now turn our attention to inhomogeneous systems
(% ~
The following lemma is required.
Positive Solutions to Linear Systems: Convexity and Separation 423
Lemma 15.8 Let . Then the set( ² ³CsÁ
9s~¸ ( && Á& ¹
is a closed, convex cone.
Proof. We leave it as an exercise to prove that is a convex cone and omit the 9
proof that is closed.9
Theorem 15.9 Let and let be()Farkas's lemma ( ² ³ Cs sÁ
nonzero. Then exactly one of the following holds:
1 There is a strictly positive solution to the system .) " ( %~s
2 There is a vector for which and .) # # ( º # Á »s
Proof. Suppose first that 1) holds. If 2) also holds, then
²#(³" ~ #²("³ ~ º#Á»
However, and imply that . This contradiction implies #( " ²#(³"
that 2) cannot hold.
Assume now that 1) fails to hold. By Lemma 15.8, the set
*~¸ ( && Á& ¹ ss
is closed and convex. The fact that 1) fails to hold is equivalent to . ¤*
Hence, there is a hyperplane that strongly separates and . All we require is *
that and be strictly separated, that is, for some and ,* # s s
º#Á%» º#Á» % * for all
Since , it follows that and so . Also, the first inequality is * º # Á »
equivalent to , that is, º#Á(&»
º( #Á&» !
for all . We claim that this implies that cannot have any& Á& (#s!
positive coordinates and thus . For if the th coordinate is #( ²( #³!
positive, then taking for we get &~
²( #³ !
which does not hold for large . Thus, 2) holds.
In the exercises, we ask the reader to show that the previous result cannot be
improved by replacing in statement 2) with . #( #(
Exercises
1. Show that any hyperplane has the form for an appropriate >²5Á 5 ³))
vector .5
424 Advanced Linear Algebra
2. If is an matrix prove that the set is a( d ¸ ( %% Á% ¹ s
convex cone in . s
3. If and are strictly separated subsets of and if is finite, prove that() ( s
() and are strongly separated as well.
4. Let be a vector space over a field with . Show that a =- ² - ³ £ char
subset of is closed under the taking of convex combinations of any?=
two of its points if and only if is closed under the taking of arbitrary ?
convex combinations, that is, for all ,
%Á Ã Á % ? Á ~ Á ¬ % ?
~ ~
5. Explain why an -dimensional subspace of is the solution set of a ²c³ s
linear equation of the form . % bÄb % ~
6 Show thatÀ
>>>bc²5Á³q ²5Á³~ ²5Á³
and that , and are pairwise disjoint and>> >bckk²5Á³ ²5Á³ ²5Á³
>>> sbckk ²5Á³r ²5Á³r ²5Á³ ~
7. A function is if it has the form for ;¢ ¦ ;²#³~ #bss affine
² Á ³ *s B s s s , where . Prove that if is convex, then so is
;²*³s.
8. Find a cone in that is not convex. Prove that a subset of is a ss?
convex cone if and only if implies that for all %Á& ? %b & ?
Á .
9. Prove that the convex hull of a set in is bounded, without ¸% ÁÃÁ% ¹s
using the fact that it is compact.
10. Suppose that a vector has two distinct representations as convex %s
combinations of the vectors . Prove that the vectors #Á Ã Á #
#c #Á Ã Á #c # are linearly dependent.
11. Suppose that is a nonempty convex subset of and that is a *² 5 Á ³ s>
hyperplane disjoint from . Prove that lies in one of the open half-spaces **
determined by . >²5Á³
12. Prove that the conclusion of Theorem 15.6 may fail if we assume only that
* is closed and convex.
13. Find two nonempty convex subsets of that are strictly separated but nots
strongly separated.
14. Prove that and are strongly separated by if and only if ?@ ² 5 Á ³ >
º5Á% » % ? º5Á& » & @ZZ ZZ for all and for all
where and and where is the? ~ ? b )²Á³ @ ~ @ b )²Á³ )²Á³
closed unit ball.
Positive Solutions to Linear Systems: Convexity and Separation 425
15. Show that Farkas's lemma cannot be improved by replacing in #(
statement 2 with . : A nice counterexample exists for )#( Hint
~ Á~ .
Chapter 16
Affine Geometry
In this chapter, we will study the geom etry of a finite-dimensional vector space
=, along with its structure-preserving maps. Throughout this chapter, all vector
spaces are assumed to be finite-dimensional.
Affine Geometry
The cosets of a quotient space have a special geometric name.
Definition Let be a subspace of a vector space . The coset:=
#b:~¸ #b :¹
is called a in with and . We also refer to flat base flat representative=: #
#b: : ²=³ = as a of . The set of all flats in is called the translate affine 7
geometry dimension of . The of is defined to be .= ² ²= ³³ ²= ³ ²= ³ dim dim77
While a flat may have many flat representatives, it only has one base since
%b:~&b; %&b; %b:~&b;~%b; implies that and so ,
whence .:~;
Definition The of a flat is . A flat of dimension is dimension %b: ²:³ dim
called a . A -flat is a , a -flat is a and a -flat is a . A flat -flat point line plane
of dimension is called a . dim²² = ³ ³ c 7 hyperplane
Definition Two flats and are said to be if ?~%b: @~&b; parallel
:; ;: ?@ or . This is denoted by .
We will denote subspaces of by the letters and flats in by = :Á;ÁÃ =
?Á@ÁÃ .
Here are some of the basic intersection properties of flats.
428 Advanced Linear Algebra
Theorem 16.1 Let and be subspaces of and let and:; = ? ~ % b :
@~& b ; = be flats in .
1 The following are equivalent:)
a some translate of is in : for some ) ?@ $ b ? @ $ =
b some translate of is in : for some ) :; # b : ; # =
c ):;
2 The following are equivalent:)
a and are translates: for some )?@ $ b ? ~ @ $ =
b and are translates: for some ):; # b : ~ ; # =
c ):~;
3)?q@£J Á:;¯?@
4)?q@£J Á:~;¯?~@
5 If then , or )?@ ?@ @? ?q@~J
6 if and only if some translation of one of these flats is contained in)?@
the other.
Proof. If 1a) holds, then and so 1b) holds. If 1b) holds, c&b$b%b:;
then and so and so 1c) holds. If 1c) holds, then#; :~² #b:³c#;
&c%b?~&b:&b;~@ and so 1a) holds. Part 2) is proved in a
similar manner.
For part 3), implies that for some and so if :; #b?@ #=
'?q@ #b'@ #@ ?@ then and so , which implies that .
Conversely, if then part 1) imp lies that . Part 4) follows similarly. ?@ :;
We leave proof of 5) and 6) to the reader.
Affine Combinations
Let be a nonempty subset of . It is well known that?=
1) is a subspace of if and only if is closed under linear combinations,?= ?
or equivalently, is closed under linear combinations of any two vectors ?
in .?
2) The smallest subspace of containing is the set of all linear =?
combinations of elements of . In different language, the of ?? linear hull
is equal to the of . linear span ?
We wish to establish the corresponding properties of affine subspaces of , =
beginning with the counterpart of a linear combination.
Definition Let be a vector space and let . A linear combination=% =
% bÄb %
where and is called an of the vectors . - ~ % affine combination
Let us refer to a nonempty subset of as if is closed under ?= ? affine closed
any affine combination of vectors in and if is closed ?? two-affine closed
Affine Geometry 429
under affine combinations of any two vectors in . These are not standard ?
terms.
The containing two distinct vectors is the setline %Á& =
%& ~ ¸%b²c³& -¹ ~ & bº%c&»
of all affine combinations of and . Thus, a subset of is two-affine closed %& ? =
if and only if contains the line through any two of its points. ?
Theorem 16.2 Let be a vector space over a field with . Then a=- ² - ³ £ char
subset of is affine closed if and only if it is two-affine closed.?=
Proof. The theorem is proved by induction on the number of terms in an
affine combination. The case holds by assumption. Assume the result true ~
for affine combinations with fewer than terms and consider the affine
combination
'~% bÄb %
where . There are two cases to consider. If either of and is not equal
to , say , write £
'~%b² c³ %bÄb %
c c
>?
and if , then since 2, we may write~ ~ ² - ³ £ char
'~ % b % b% bÄb%
>? 33
In either case, the inductive hypothesis applies to the expression inside the
square brackets and then to . '
The requirement is necessary, for if , then the subset char²-³ £ - ~ {
? ~ ¸²Á³Á²Á³Á²Á³¹
of is two-affine closed but not affi ne closed. We can now characterize flats. -
Theorem 16.3 A nonempty subset of a vector space is a flat if and only if ?=
?² - ³ £ ? ? is affine closed. Moreover, if , then is a flat if and only if is char
two-affine closed.
Proof. Let be a flat and let , where . If?~%b: % ~%b ? :
'~ , then
% ~ ² %b ³~%b %b:
and so is affine closed. Conversely, suppose that is affine closed, let ??
%? :~?c% - : and let . If and then
430 Advanced Linear Algebra
b ~²% c%³b²% c%³~% b% c² b³%
for . Since the sum of the coefficients of , and in the last% ? % % %
expression is , it follows that
b b%~% b% c² b c³%?
and so . Thus, is a subspace of and is b ?c%~: : = ?~%b:
a flat. The rest follows from Theorem 16.2.
Affine Hulls
The following definition is the analog of the subspace spanned by a collection
of vectors.
Definition Let be a nonempty set of vectors in .?=
1 The of , denoted by , is the smallest flat containing) affhull affine hull ?² ? ³
?.
2 The of , denoted by , is the set of all affine) affspan affine span ?² ? ³
combinations of vectors in . ?
Theorem 16.4 Let be a nonempty subset of . Then?=
affhull affspan span²?³ ~ ²?³ ~ %b ²? c%³
or equivalently, for a subspace of , :=
%b:~ ² ? ³ ¯ :~ ² ?c% ³ affspan span
Also,
dim dim²² ? ³ ³ ~ ² ² ? c % ³ ³affspan span
Proof. Theorem 16.3 implies that and so it is affspan affhull²?³ ²?³
sufficient to show that is a f lat, or equivalently, that for any (~ ² ? ³ affspan
&( :~(c& = , the set is a subspace of . To this end, let
&~ %
~
Á
Then any two elements of have the form and , where :& c & & c &
&~ % &~ % Á Á
~ ~
and 2
are in . But if , then( Á ! -
Affine Geometry 431
' ²& c&³b!²& c&³
~ b! c² b!³ %
~ b! c² b!c³ % c&
~
Á Á Á
~
Á Á Á 45
452
2
which is in , since the last sum is an affine sum. Hence, is a (c&~: :
subspace of . We leave the rest of the proof to the reader. =
The Lattice of Flats
The intersection of subspaces is a subspace, although it may be trivial. For flats,
if the intersection is not empty, then it is also a flat. However, since the
intersection of flats may be empty, the set does not form a lattice under 7²= ³
intersection. However, we can easily fix this.
Theorem 16.5 Let be a vector space. The set=
77²= ³ ~ ²= ³r¸J¹
of all flats in , together with the em pty set, is a complete lattice in which meet =
is intersection. In particular:
1 is closed under arbitrary intersection. In fact, if)7²= ³
<~¸ % b: 2¹ has nonempty intersection, then
2 2 2 <~² % b : ³ ~ % b:
for some . In other words, th e base of the intersection is the %<
intersection of the bases.
2 The join of the family is the intersection of all ) << ~¸ % b: 2¹
flats containing the members of . Also, <
45 <<~affhull
3 If and are flats in , then)?~%b: @~&b; =
? v@ ~ %b²º%c&»b: b;³
If , then?q@£J
?v@~%b² :b;³
Proof. For part 1), if
% ² %b:³
2
then for all and so%b :~ % b : 2
432 Advanced Linear Algebra
2 2 2 ²% b: ³ ~ ²%b: ³ ~ %b :
We leave proof of part 2) to the reader.
For part 3), since , it follows that %Á& ? v@
?v@~%b<~&b<
for some subspace of . Thus, . Also, implies that < = %c&< %b:%b<
:< ;< :b;< >~ and similarly , whence and so if
º%c&»b:b; >< %b>%b<~?v@ , then . Hence, . On the
other hand,
?~%b:%b>
and
@~&b;~%c² %c& ³b;%b>
and so . Thus, .?v@%b> ?v@~%b>
If , then we may take the flat representatives for and to be any?q@£J ? @
element , in which case part 1) gives'?q@
?v@~'b² º 'c'»b:b;³~'b:b;
and since , we also have . %?v@ ?v@ ~%b:b;
We can now describe the dimension of the join of two flats.
Theorem 16.6 Let and be flats in .?~%b: @~&b; =
1 I f , t h e n)?q@£J
dim dim dim dim dim²? v@³ ~ ²: b;³ ~ ²?³b ²@³c ²? q@³
2 I f , t h e n)?q@~J
dim dim²? v@³ ~ ²: b;³b
Proof. We have seen that if , then ?q@£J
?v@~%b:b;
and so
dim dim²? v@³ ~ ²: b;³
On the other hand, if , then ?q@~J
?v@ ~%b²º%c&»b:b;³
and since , we get dim²º%c&»³ ~
Affine Geometry 433
dim dim²? v@³ ~ ²: b;³b
Finally, we have
dim dim dim dim²: b;³ ~ ²:³b ²;³c ²: q;³
and Theorem 16.5 implies that
dim dim²? q@³ ~ ²: q;³
Affine Independence
We now discuss the affine counterpart of linear independence.
Theorem 16.7 Let be a nonempty set of vectors in . The following are?=
equivalent:
1 For all , the set) %?
²? c%³±¸¹
is linearly independent.
2 For all ,) %?
%¤ ² ?±¸ % ¹ ³affhull
3 For any vectors ,) % ?
% ~ Á ~ ¬ ~ for all
4 For affine combinations of vectors in ,) ?
% ~ % ¬ ~ for all
5 When is finite, ) ?~¸% ÁÃÁ% ¹
dim²² ? ³ ³ ~ c affhull
A set of vectors satisfying any any hence all of these conditions is said to be? ()
affinely independent .
Proof. If 1) holds but there is an affine combination equal to , %
%~ %
~
where for all , then%£ %
~
² %c% ³~
434 Advanced Linear Algebra
Since is nonzero for some , this contradicts 1). Hence, 1) implies 2). Suppose
that 2) holds and
% ~
where . If some , say , is nonzero then ~
%~ c ² ° ³ % ² ? ± ¸ % ¹ ³
affhull
which contradicts 2) and so for all . Hence, 2) implies 3). ~
If 3) holds and the affine combinations satisfy
% ~ %
then
² c ³% ~
and since , it follows that for all . Hence, 4) ² c ³ ~ c ~ ~
holds. Thus, it is clear that 3) a nd 4) are equivalent. If 3) holds and
² %c% ³~
for , then%£%
45%c %~
and so 3) implies that for all . ~
Finally, suppose that . Since ?~¸% ÁÃÁ% ¹
dim dim² ²?³³ ~ ²º? c% »³affhull
it follows that 5) holds if and only if , which has size , is ²? c% ³±¸¹ c
linearly independent.
Affinely independent sets enjoy some of the basic properties of linearly
independent sets. For example, a nonempty subset of an affinely independent set
is affinely independent. Also, any nonempty set contains an affinely ?
independent set.
Since the affine hull of an affinely independent set is not the /~ ² ? ³ ? affhull
affine hull of any proper subset of , we deduce that is a minimal affine ??
spanning set of its affine hull.
Affine Geometry 435
Affine Bases and Barycentric Coordinates
We have seen that a set is affinely independent if and only if the set ?
? ~ ²? c%³±¸¹%
is linearly independent. We have also seen that for a subsapce of , :=
%b:~ ² ? ³ ¯ :~ ² ?³ affspan span %
Therefore, if by analogy, we define a subset of a flat to be an 8 (~%b:
affine basis for if is affinely independent and , then is an(² ³ ~ (88 8 affspan
affine basis for if and only if is a basis for . %b: : 8%
Theorem 16.8 Let be a flat of dimension . Let be(~%b: ~² %ÁÃÁ%³ 8
an ordered basis for and let be an : ² b%³r¸%¹ ~ ²% b%ÁÃÁ% b%Á%³8
ordered affine basis for . Then every has a unique expression as an (# (
affine combination
#~% bÄb % b % b
The coefficients are called th e of with respect to # barycentric coordinates
the ordered affine basis . 8b%
For example, in , a plane is a flat of the form where s (~%bº #Á#»
8s~² #Á #³ is an ordered basis for a two-dimensional subspace of . Then
² b%³r¸%¹ ~ ²# b%Á# b%Á%³ ~ ² Á Á ³8
are barycentric coordinates for the plane, that is, any has the form #(
bb
where .b b ~
Affine Transformations
Now let us discuss some properties of maps that preserve affine structure.
Definition A function that preserves a ffine combinations, that is, for ¢= ¦=
which
45
~ ¬ % ~ ² % ³
is called an or , or . affine transformation affine map affinity ()
We should mention that some authors re quire that be bijective in order to be
an affine map. The following theorem is the analog of Theorem 16.2.
436 Advanced Linear Algebra
Theorem 16.9 If , then a function is an affinechar²-³ £ ¢= ¦ =
transformation if and only if it preserves affine combinations of every pair of its
points, that is, if and only if
²%b²c³&³ ~ ²%³b²c³²&³
Thus, if , then a map is an affine transformation if and only if it char²-³ £
sends the line through and to the line through and . It is clear that % & ²%³ ²&³
linear transformations are affine transformations. So are the following maps.
Definition Let . The affine map defined by#= ;¢= ¦= #
;² % ³~%b##
for all , is called by .%= # translation
It is not hard to see that any composition , where , is affine. ;k ² = ³# B
Conversely, any affine map must have this form.
Theorem 16.10 A function is an affine transformation if and only if ¢= ¦=
it is a linear operator followed by a translation,
~;k #
where and .#= ² =³B
Proof. We leave proof that is an affine transformation to the reader. Let ;k #
be an affine map and suppose that . Then . Moreover, ~c' ; k~ '
letting , we have~; k'
²"b #³~²"b #³b'
~ ²"b #b²cc ³³b'
~"b #c²cc ³'b'
~ "b #
and so is linear.
Corollary 16.11
1 The composition of two affine trans formations is an affine transformation. )
2 An affine transformation is bijective if and only if is bijective.) ~;k #
3 The set of all bijective affine transformations on is a group under)a f f ²= ³ =
composition of maps, called the of . affine group =
Let us make a few group-theoretic remarks about . The set of all aff trans²= ³ ²= ³
translations of is a subgroup of . We can define a function =² = ³ aff
B¢ ² =³¦ ² =³aff by
²; k ³ ~#
Affine Geometry 437
It is not hard to see that is a well-defined group homomorphism from aff²= ³
onto , with kernel . Hence, is a normal subgroup ofB²= ³ ²= ³ ²= ³ trans trans
aff²= ³ and
aff
trans²= ³
²= ³² = ³B
Projective Geometry
If , the join affine hull of any two distinct points in is a line. Ondim²= ³ ~ = ()
the other hand, it is not the case that the intersection of any two lines is a point,
since the lines may be parallel. Thus, there is a certain asymmetry between the
concepts of points and lines in . This asymmetry can be removed by =
constructing the . Our plan here is to very briefly describe one projective plane
possible construction of projective geometries of all dimensions.
By way of motivation, let us consider Figure 16.1.
Figure 16.1
Note that is a hyperplane in a 3-dimensional vector space and that . /= ¤ /
Now, the set of all flats of that lie in is an affine geometry of 7²/³ = /
dimension . According to our definition of affine geometry, must be a /(
vector space in order to define . However, we hereby extend the definition 7²/³
of affine geometry to include the collec tion of all flats contained in a flat of . =³
Figure 16.1 shows a one-dimensional flat and its linear span , as well as a ?º ? »
zero-dimensional flat and its span . Note that, for any flat in , we @º @ » ? /
have
dim dim²º?»³ ~ ²?³b
Note also that if and are any tw o distinct lines in , the corresponding 33 /
438 Advanced Linear Algebra
planes and have the property that their intersection is a line through º3 » º3 »
the origin, . We are now ready to define projective even if the lines are parallel
geometries.
Let be a vector space of any dimension and let be a hyperplane in not=/ =
containing the origin. To each flat in , we associate the subspace of ?/ º ? »=
generated by . Thus, the linear span function maps affine ?7 ¢ ² / ³ ¦ ² = ³ 7I
subspaces of to subspaces of . The span function is not surjective: ?/ º ? »=
Its image is the set of all subspaces that are contained in the base subspace not
2/ of the flat .
The linear span function is one-to-one and its inverse is intersection with , /
7< ~ < q /c
for any subspace not contained in . <2
The affine geometry is, as we have remarked, somewhat incomplete. In 7²/³
the case , every pair of points determines a line but not every pair dim²/³ ~
of lines determines a point.
Now, since the linear span function is injective, we can identify with7² / ³ 7
its image , which is the set of all subspaces of not contained in the 7² ²/³³ =7
base subspace . This view of allows us to “complete” by 2² / ³ ² / ³ 77
including the base subspace . In the three-dimensional case of Figure 16.1, the 2
base plane, in effect, adds a projective lin e at infinity. With th is inclusion, every
pair of lines intersects, parallel lines in tersecting at a point on the line at infinity.
This two-dimensional projective geometry is called the . projective plane
Definition Let be a vector space. The set of all subspaces of is=² = ³ = I
called the of . The of projective geometry projective dimension =² : ³ pdim
: ² =³I is defined as
pdim²:³ ~ ²:³c dim
The of is defined to be . Aprojective dimension F² =³ ² =³~ ² =³c pdim dim
subspace of projective dimension is called a and a subspace projective point
of projective dimension is called a . projective line
Thus, referring to Figure 16.1, a projective point is a line through the origin and,
provided that it is not contained in the base plane , it meets in an affine 2/
point. Similarly, a projective line is a plane through the origin and, provided that
it is not , it will meet in an affine line. In short,2/
span
span²³ ~ ~
²³ ~ ~affine point line through the origin projective point
affine line plane through the origin projective line
The linear span function has the following properties.
Affine Geometry 439
Theorem 16.12 The linear span function from the affine 7¢ ²/³ ¦ ²= ³7I
geometry to the projective geometry defined by 7I²/³ ²= ³ 7? ~ º?»
satisfies the following properties:
1 The linear span function is injective, with inverse given by)
7< ~ < q /c
for all subspaces not contained in the base subspace of . <2 /
2 The image of the span function is the set of all subspaces of that are not ) =
contained in the base subspace of . 2/
3 if and only if )?@ º ? »º @»
4 If are flats in with nonempty intersection, then)?/
span span45
2 2?~ ² ? ³
5 For any collection of flats in ,) /
span span89
2 2?~ ² ? ³
6 The linear span function preserves dimension, in the sense that)
pdim span²² ? ³ ³ ~² ? ³ dim
7 if and only if one of and is contained in the)?@ º ? »q2 º @»q2
other.
Proof. To prove part 1 , let be a flat in . Then and so )%b: / %/
/ ~ %b2 : 2 º%b:» ~ º%»b: , which implies that . Note also that and
'º%b:»q/~²º%»b:³q²%b2³¬'~%b ~%b
for some , and . This implies that , which : 2 - ² c ³ %2
implies that either or . But implies and so , %2 ~ %/ %¤2 ~
which implies that . In other words, '~%b %b:
º%b:»q/ %b:
Since the reverse inclusion is clear, we have
º%b:»q/ ~ %b:
This establishes 1 . )
To prove 2 , let be a subspace of that is not contained in . We wish to )<= 2
show that is in the imag e of the linear span functi on. Note first that since <
<2 ² 2 ³ ~ ² = ³ c <b 2~ = and , we have and so dim dim
dim dim dim dim dim²< q2³ ~ ²<³b ²2³c ²< b2³ ~ ²<³c
440 Advanced Linear Algebra
Now let . Then£%<c2
%¤2¬º % »b2~=
¬ %b/ £-Á2
¬ %/ for some
Thus, for some . Hence, the flat lies in %<q/ £- %b²<q2³ /
and
dim dim dim²%b²< q2³³ ~ ²< q2³ ~ ²<³c
which implies that lies in and has span²%b²< q2³³ ~ º%»b²< q2³ <
the same dimension as . In other words, <
span²%b²< q2³³ ~ º%»b²< q2³ ~ <
We leave proof of the remaining parts of the theorem as exercises.
Exercises
1. Show that if , then the set is a %ÁÃÁ% = :~¸ % ~ ¹ ''
subspace of . =
2. Prove that if is nonempty then ?=
affhull²?³ ~ %bº? c%»
3. Prove that the set in is closed under the ? ~ ¸²Á³Á ²Á³Á ²Á³¹ ² ³ {
formation of lines, but not affine hulls.
4. Prove that a flat contains the origin if and only if it is a subspace.
5. Prove that a flat is a subspace if and only if for some we have ?% ?
% ? £ - for some .
6. Show that the join of a co llection of flats in is the9~¸ % b: 2¹ =
intersection of all flats that contain all flats in . 9
7. Is the collection of all flats in a lattice under set inclusion? If not, how=
can you “fix” this?
8. Suppose that and . Prove that if ?~%b: @~&b; ² ? ³~ ² @³ dim dim
and , then .?@ :~;
9. Suppose that and are disjoint hyperplanes in . ?~%b: @~&b; =
Show that . :~;
10. (The parallel postulate) Let be a flat in and . Show that there is ?= # ¤ ?
exactly one flat containing , parallel to and having the same dimension #?
as .?
11. a Find an example to show that the join of two flats may not be ) ?v@
the set of all lines connecting all points in the union of these flats.
b Show that if and are flats with , then is the ) ?@ ? q @ £ J ? v @
union of all lines where and . %& % ? & @
12. Show that if and , then ?@ ?q@~J
dim max dim dim²? v@³ ~ ¸ ²?³Á ²@³¹b
Affine Geometry 441
13. Let . Prove the following: dim²= ³ ~
a The join of any two distinct points is a line. )
b The intersection of any two nonparallel lines is a point. )
14. Let . Prove the following: dim²= ³ ~
a The join of any two distinct points is a line. )
b The intersection of any two nonparallel planes is a line. )
c The join of any two lines whose intersection is a point is a plane. )
d The intersection of two coplanar nonparallel lines is a point. )
e The join of any two dis tinct parallel lines is a plane. )
f The join of a line and a point not on that line is a plane. )
g The intersection of a plane and a line not on that plane is a point. )
15. Prove that is a surjective affine transformation if and only if ¢= ¦=
~ k; $= ² =³ B$ for some and .
16. Verify the group-theoretic re marks about the group homomorphism
B¢ ² =³¦ ² =³ ² =³ ² =³aff trans aff and the subgroup of .
Chapter 17
Singular Values and the Moore–Penrose
Inverse
Singular Values
Let and be finite-dimensional inner product spaces over or and let<= ds
B ² < Á = ³ . The spectral theorem applied to can be of considerable helpi
in understanding the relationship between and its adjoint . This irelationship
is shown in Figure 17.1. Note that and can be decomposed into direct sums <=
<~(l) =~*l+ and
in such a manner that and act symmetrically in the sense ¢(¦* ¢*¦(i
that
¢" ª # ¢# ª " iand
Also, both and are zero on and , respectively.i)+
We begin by noting that is a positive Hermitian operator. Hence, if Bi² < ³
~ ² ³~ ² ³ <rk rk i, then has an ordered orthonormal basis
8~²" ÁÃÁ"Á" ÁÃÁ" ³ b
of eigenvectors for , where the corresponding eigenvalues can be arranged i
so that
b Ä ~ ~Ä~
The set is an ordered orthonormal basis for ²" ÁÃÁ" ³ ² ³~ ² ³b iker ker
and so is an ordered orthonormal basis for .²" ÁÃÁ"³ ² ³ ~ ² ³iker im
444 Advanced Linear Algebra
W(uk)=skvk
W
(vk)=skuku1 v1
ur+1 vr+1
un vmim(W*)
ker(W)k e r ( W*)im(W)
ur vr
W
W
ONB of
eigenvectors
for W*WONB of
eigenvectors
for WW*
Figure 17.1
For , the positive numbers are called the ~ ÁÃÁ ~ j singular values
of . If we set for , then ~
i
"~ "
for . We can achieve some “symmetry” here between and by~ ÁÃÁ i
setting for each , giving# ~ ²° ³ "
"~ #
F
and
i
#~ "
F
The vectors are orthonormal, since if , then #Á Ã Á # Á
º #Á#»~ º "Á "»~ º "Á"»~ º "Á"»~
Á
i
Hence, is an orthonormal basis for , which can be²# ÁÃÁ#³ ² ³~ ² ³iim ker
extended to an orthonormal basis for , the extension 9~² #ÁÃÁ# ³ =
²# ÁÃÁ# ³ ² ³b i being an orthonormal basis for . Moreover, since ker
i
#~ "~ #
the vectors are eigenvectors for with the same eigenvalues #Á Ã Á #i
i~ as for . This completes the picture in Figure 17.1.
445
Theorem 17.1 Let and be finite-dimensional inner product spaces over <= d
or and let s B ² < Á = ³ have rank . Then there are ordered orthonormal
bases and for and , respectively, for which89 <=
8~²" ÁÃÁ" Á " ÁÃÁ" ³ b
²³ ²³ ONB for ONB for im i ker
and
9~²# ÁÃÁ# Á # ÁÃÁ# ³ b
²³ ²³ ONB for ONB for im keri
Moreover, for ,
"~ #
#~ "
i
where are called the of , defined by singular values
i
"~ " Á
for . The vectors are called the for and " ÁÃÁ" right singular vectors
the vectors are called the for . #Á Ã Á # left singular vectors
The matrix version of the previous discussion leads to the well-known singular-
value decomposition of a matrix. Let and let ( ²-³ ~ ²" ÁÃÁ" ³C8Á
and be the orthonormal bases from and , respectively, in9~² #ÁÃÁ# ³ < =
Theorem 17.1, for the operator . Then (
´ µ ~ ~ ² Á ÁÃÁ ÁÁÃÁ³'89Á diag
A change of orthonormal bases from the standard bases to and gives 9:
( ~ ´µ ~ 4 ´µ4 ~ 78'((ÁÁ Á Ái
;; 9 ; 8 9 ;8
where and are unitary/orthogonal. This is the singular-7~4 8~49; 8;ÁÁ
value decomposition of . (
As to uniqueness, if , where and are unitary and is diagonal, (~7 8 7 8''i
with diagonal entries , then
((~² 7 8³7 8 ~8 8ii i i i i'' ' '
and since , it follows that the 's are eigenvalues of'' i
~² Á Ã Á ³diag
((i, that is, they are the squares of the singular values along with a sufficient
number of 's. Hence, is uniquely determined by , up to the order of the ( '
diagonal elements.Singular Values and the Moore–Penrose Inverse
446 Advanced Linear Algebra
We state without proof the following unique ness facts and refer the reader to
[48] for details. If and if the eigenvalues are distinct, then is 7
uniquely determined up to multiplication on the right by a diagonal matrix of the
form with . If , then is never uniquely+~ ²'ÁÃÁ' ³ ' ~ 8 diag ((
determined. If , then for any given there is a unique . Thus, we ~~ 7 8
see that, in general, the singular-value decomposition is not unique.
The Moore–Penrose Generalized Inverse
Singular values lead to a generalization of the inverse of an operator that applies
to all linear transformations. The setup is the same as in Figure 17.1. Referring
to that figure, we are prompted to define a linear transformation by b¢= ¦<
b
#~"
Ffor
for
since then
²³ O ~
²³ O ~
b
º" ÁÃÁ" »
b
º" ÁÃÁ" »
b
and
²³ O ~
²³ O ~
b
º# ÁÃÁ# »
b
º# ÁÃÁ# »
b
Hence, if , then . The transformation is called the ~~ ~ bc b
Moore–Penrose generalized inverse Moore–Penrose pseudoinverse or of .
We abbreviate this as MP inverse.
Note that the composition is the identity on the largest possible subspace of b
< on which any composition of the form could be the identity, namely, the
orthogonal complement of the kernel of . A similar statement holds for the
composition . Hence, is as “close” to an inverse for as is possible. bb
We have said that if is inver tible, then . More is true: If is bc ~
injective, then and so is a left inverse for . Also, if is surjective, bb~
then is a right inverse for . Hence the MP inverse generalizes the one- bb
sided inverses as well.
Here is a characterization of the MP inverse.
Theorem 17.2 Let . The MP inverse of is completelyB ² < Á = ³b
characterized by the following four properties:
1) b~
2) bb b~
3 is Hermitian)b
4 is Hermitian)b
447
Proof . We leave it to the reader to show that does indeed satisfy conditionsb
1)–4) and prove only the uniqueness. Suppose that and satisfy 1)–4) when
substituted for . Then b
~
~² ³
~
~² ³
~
~² ³
~
~
~
i
ii
ii
iiii
iii
ii
and
~
~² ³
~
~² ³
~
~² ³
~
~
~
i
ii
ii
iiii
ii i
ii
which shows that . ~
The MP inverse can also be defined for matrices. In particular, if , (4 ² -³ Á
then the matrix operator has an MP inverse . Since this is a linear ( (b
transformation from to , it is just multiplication by a matrix . -- ~
(b)
This matrix is the for and is denoted by . )( ( MP inverseb
Since and , the matrix version of Theorem 17.2 implies (b(( ) ( ) ~~ b
that is completely characterized by the four conditions(b
1)(( ( ~ (b
2)(( ( ~ (bb b
3 is Hermitian)((b
4 is Hermitian)((b
Moreover, if
(~< < i'
is the singular-value decomposition of , then (Singular Values and the Moore–Penrose Inverse
448 Advanced Linear Algebra
(~ < <bZ i
'
where is obtained from by repl acing all nonzero entries by their ''Z
multiplicative inverses. This follows from the characterization above and also
from the fact that for ,
<< # ~ < ~ < ~ " Zi Z c c
''
and for ,
<< # ~ < ~ Zi Z
''
Least Squares Approximation
Let us now discuss the most important use of the MP inverse. Consider the
system of linear equations
(% ~ #
where . As usual, or . This system has a solution(4 ² - ³ -~ -~ Á () ds
if and only if . If the system has no solution, then it is of considerable # ² ³im(
practical importance to be ab le to solve the system
(% ~ #V
where is the unique vector in that is closest to , as measured by the#² ³ #V im(
unitary or Euclidean distance. This problem is called the () linear least squares
problem. Any solution to the system is called a (% ~ #V least squares solution
to the system . Put another way, a least squares solution to is a (% ~ # (% ~ #
vector for which is minimized.%( % c # ))
Suppose that and are least squares solutions to . Then $' ( % ~ #
( $~#~( 'V
and so . We will write for . Thus, if is a particular least$c' ²(³ ( $ ker ( ) (
squares solution, then the set of all least squares solutions is . $b ²(³ ker
Among all solutions, the most interesting is the solution of minimum norm.
Note that if there is a least squares so lution that lies in , then for any$² ( ³ ker
' ² ( ³ker , we have
) ) )) ) ) ))$ b '~ $b ' $
and so will be the unique least squares solution of minimum norm.$
Before proceeding, we recall Theorem 9.14 that if is a subspace of a finite- () :
dimensional inner product space , then the best approximation to a vector =
#= : #: #c#: VV from within is the unique vector for which . Now we
can see how the MP inverse comes into play.
449
Theorem 17.3 Let . Among the least squares solutions to the(4 ² -³ Á
system
(% ~ #V
there is a unique solution of minimum norm, given by , where is the MP (# (bb
inverse of . (
Proof. A vector is a least squares solution if and only if . Using the $( $ ~ # V
characterization of the best approximation , we see that is a solution to #$V
($ ~ #V if and only if
($c# ²(³ im
Since this is equivalent to im²(³ ~ ²( ³iker
(² ( $c# ³~i
or
(( $~(#ii
This system of equations is called the for . Its normal equations (% ~ #
solutions are precisely the least squa res solutions to the system . (% ~ #
To see that is a least squares so lution, recall that, in the notation of $~( #b
Figure 17.1,
(( # ~#
b
F
and so
(( ² (#³~ ~(#(#
ib i
iF
and since is a basis for , we conclude that satisfies the9~² #ÁÃÁ# ³ = ( #b
normal equations. Finally, since , we deduce by the preceding (# ² ( ³bker
remarks that is the unique least squares solution of minimum norm. (#b
Exercises
1. Let . Show that the singular values of are the same as those ofB ² < ³i
.
2. Find the singular values and the singular value decomposition of the matrix
(~
>?
Find .(b
3. Find the singular values and the singular value decomposition of the matrixSingular Values and the Moore–Penrose Inverse
450 Advanced Linear Algebra
(~
>?
Find . : Is it better to work with or ?(( ( ( (bi iHint
4. Let be a column matrix over . Find a singular-value? ~ ² %%Ä %³ !d
decomposition of . ?
5. Let and let be the square matrix(4 ²-³ )4 ²-³ Á bÁb
)~(
(>?i
block
Show that, counting multiplicity, the nonzero eigenvalues of are )
precisely the singular values of t ogether with their negatives. : Let( Hint
(~< < ( ) i' be a singular-value decomposition of and try factoring
into a product where is unitary. Do not read the following second <:< <i
hint unless you get stuck. : Ve rify the block factorization Second Hint
)~< <
< <>? > ? > ?
ii
i'
'
What are the eigenvalues of the middle factor on the right? Try ( b b
and . b c )
6. Use the results of the previous exercise to show that a matrix
(4 ² -³ ( ( ( Ái!, its adjoint , its transpose and its conjugate all have
the same singular values. Show also that if and are unitary, then << (Z
and have the same singular values.<(<Z
7. Let be nonsingular. Show that the following procedure (4² -³
produces a singular-value decomposition of . (~< < ( i'
a Write where and the 's are )d i a g (~<+ < +~ ² ÁÃÁ ³i
positive and the columns of form an orthonormal basis of <
eigenvectors for . We never said that this was a practical procedure. (()
b Let where the square roots are nonnegative. )d i a g'~² Á à Á³° °
Also let and U .<~ < ~ ( <ic '
8. If is an matrix, then the of is(~² ³ d ( Á Frobenius norm
))89 (~ -
ÁÁ°
Show that is the sum of the squares of the singular values of )) (~ -
(.
Chapter 18
An Introduction to Algebras
Motivation
We have spent considerable time studying the structure of a linear operator
B² = ³ = -- on a finite-dimensional vector space over a field . In our
studies, we defined the -module and used the decomposition theorems -´%µ =
for modules over a principal ideal dom ain to dissect this module. We
concentrated on an individual operator , rather than the entire vector space
BB- - ²= ³ ²= ³. In fact, we have made relatively little use of the fact that is an
algebra under composition. In this chapter, we give a brief introduction to the
theory of algebras, of which is the most general, in the sense of Theorem B-²= ³
18.2 below.
Associative Algebras
An algebra is a combination of a ring and a vector space, with an axiom that
links the ring product with scalar multiplication.
Definition associative An over a field , or an , is a() algebra -algebra (- -
nonempty set , together with three operations, called denoted by ( addition (
b)( ) (, denoted by juxtaposition and alsomultiplication scalar multiplication
denoted by juxtaposition , for which the following properties hold: )
1 is a vector space over under addition and scalar multiplication.)(-
2 is a ring with identity under addition and multiplication.)(
3 If and , then)- Á(
²³ ~ ²³ ~ ²³
An algebra is if it is finite-dimensional as a vector space. An finite-dimensional
algebra is if is a commutative ring. An element is commutative ( (
invertible if there is for which . ( ~ ~
Our definition requires that have a multiplicative identity. Su ch algebras are (
called . Algebras without unit are also of great importance, but unital algebras
452 Advanced Linear Algebra
we will not study them here. Also, in this chapter, we will assume that all
algebras are associative. Nonassociative algebras, such as Lie algebras and
Jordan algebras, are important as well.
The Center of an Algebra
Definition The of an -algebra is the set center -(
A² ( ³~¸ ( %~% %( ¹ for all
of all elements of that commute with every element of . ((
The center of an algebra is never trivial since it contains a copy of : -
¸ -¹ A²(³
Definition An -algebra is if its center is as small as possible, that-( central
is, if
A²(³~¸-¹
From a Vector Space to an Algebra
If is a vector space over a field and if is a basis for ,=- ~ ¸ 0 ¹ = 8
then it is natural to wonder whether we can form an -algebra simply by -
defining a product for the basis elements a nd then using the distributive laws to
extend the product to . In particular, we choose a set of constants with the = Á
property that for each pair , only finitely many of the are nonzero. Then ²Á³ Á
we set
~
Á
and make multiplication bilinear, that is,
89
89~ ~
~ ~ ~
~
and
~ 89
~ ~
for . It is easy to see that this does define a nonunital associative algebra-
( provided that
An Introduction to Algebras 453
² ³ ~² ³
for all and that is commutative if and only ifÁÁ 0 (
~
for all . The constants are called the for theÁ 0 Ástructure constants
algebra . To get a unital algebra, we can take for a given , the structure( 0
constants to be
Á Á
Á~~
in which case is the mu ltiplicative identity. An alternativ e is to adjoin a new (
element to the basis and define its structure constants in this way.)
Examples
The following examples will make it clear why algebras are important.
Example 18.1 If are fields, then is a vector space over . This vector-, , -
space structure, along with the ring structure of , is an algebra over . ,-
Example 18.2 The ring of polynomials is an algebra over . -´%µ -
Example 18.3 The ring of all matrices over a field is anC²-³ d -
algebra over , where scalar multiplication is defined by -
4 ~ ² ³Á - ¬ 4 ~ ² ³ Á Á
Example 18.4 The set of all linear operators on a vector space over aB-²= ³ =
field is an -algebra, where additi on is addition of functions, multiplication is --
composition of functions and scalar multiplication is given by
² ³²#³ ~ ´ #µ
The identity map is the multip licative identity and the zero map B² = ³-
² =³ ² =³B- - is the additive identity. This algebra is also denoted by , End
since the linear operators on are also called endomorphisms of . ==
Example 18.5 If is a group and is a field, then we can form a vector space.-
-´.µ - - . over by taking all formal -linear combinations of elements of and
treating as a basis for . This vector space can be made into an -algebra.- ´ . µ -
where the structure constants are determined by the group product, that is, if
~ ~ ~ " Á " Á, then . The group identity is the algebra identity
since and so and similarly, .~ ~ ~ Á Á Á Á
The resulting associative algebra is called the over . -´.µ - group algebra
Specifically, the elements of have the form -´.µ
454 Advanced Linear Algebra
%~ bÄb
where and . If - .
&~ bÄb
then we can include additional terms with coefficients and reindex if
necessary so that we may assume that and for all . Then the sum ~ ~
in is given by-´.µ
89 89
~ ~ ~
b ~ ² b ³
Also, the product is given by
89 89
~ ~ Á
~
and the scalar product is
~ 89
~ ~
The Usual Suspects
Algebras have substructures and structure-preserving maps, as do groups, rings
and other algebraic structures.
Subalgebras
Definition Let be an -algebra. A of is a subset of that is(- ( ) ( subalgebra
a subring of with the same identity as and a subspace of . (( (()
The intersection of subalgebras is a s ubalgebra and so the family of all
subalgebras of is a complete lattice, wher e meet is intersection and the join of (
a family of subalgebras is the inter section of all subalgebras of that contain < (
the members of . <
The by a nonempty subset of an algebra is the subalgebra generated ?(
smallest subalgebra of that contains and is easily seen to be the set of all (?
linear combinations of finite products of elements of , that is, the subspace ?
spanned by the products of finite subsets of elements of : ?
º?» ~ º% Ä% % ?» alg
Alternatively, is the set of all pol ynomials in the variables in . In º?» ? alg
particular, the algebra generated by a single element is the set of all %(
polynomials in over . %-
An Introduction to Algebras 455
Ideals and Quotients
In defining the notion of an ideal of an alge bra , we must consider the fact that(
( may be noncommutative.
Definition two-sided A of an associative algebra is a nonempty() ideal (
subset of that is closed under addition and subtraction, that is,0(
Á0 ¬ bÁc0
and also left and right multiplication by elements of , that is, (
0Á Á ( ¬ 0
The by a nonempty subset of is the smallest ideal ideal generated ?(
containing and is equal to ?
º ? » ~ % % ? ÁÁ ( ideal HI
~
Definition An algebra is if ( simple
1 The product in is not trivial, that is, for at least one pair of) ( £
elements Á (
2 has no proper nonzero ideals.)(
Definition If is an ideal in , then the is the quotient 0( quotient algebra
ring/quotient space
(°0 ~ ¸b0 (¹
with operations
²b0³b²b0³~²b³b0
²b0³²b0³~b0
²b0³~b0
where and . These operations make an -algebra.Á ( - (°0 -
Homomorphisms
Definition If and are -algebras a map is an ()- ¢ ( ¦ ) algebra
homomorphism if it is a ring homomorphism as well as a linear
transformation, that is,
²b ³ ~ b Á ² ³ ~ ² ³² ³Á ~ ZZ ZZ
and
² ³ ~ ²³
for .-
456 Advanced Linear Algebra
The usual terms monomorphism, epimorphism, isomorphism, embedding,
endomorphism and automorphism apply to algebras with the analogous meaning
as for vector spaces and modules.
Example 18.6 Let be an -dimensional vector space over . Fix an ordered= -
basis for . Consider the map defined by8 B C=¢ ² = ³ ¦ ² - ³
²³ ~ ´µ 8
where is the matrix representation of with respect to the ordered basis . ´µ88
This map is a vector space isomorphism and since
´µ ~ ´ µ ´ µ 88 8
it is also an algebra isomorphism.
Another View of Algebras
If is an algebra over , then contains a copy of . Specifically, we define a (- (-
function by¢- ¦(
~
for all , where is the multiplicativ e identity. The elements are in - (
the center of , since for any , ( (
²³ ~ ²³ ~
and
²³ ~ ²³ ~
Thus, . To see that is a ring homomorphism, we have¢- ¦A²(³
² ³ ~ h ~
²b ³~²b ³~b ~ ²³b ² ³
² ³ ~ ² ³ ~ ² ³ ~ ² ³ ~ ²h ² ³³ ~ ²h³ ² ³ ~ ²³ ² ³--
Moreover, if and , then ~ £
~ ² ³~ h~c
-
and so provided that in , we have . Thus, is an embedding. £ ( ~
Theorem 18.1
1 If is an associative algebra over and if in , then the map)(- £ (
¢- ¦A²(³ defined by
~
is an embedding of the field into the center of the ring . Thus, -A ² ( ³ ( -
can be embedded as a subring of . A²(³
An Introduction to Algebras 457
2 Conversely, if is a ring with identity and if is a field, then ) 9- A ² 9 ³ 9
is an -algebra with scalar multiplication defined by the product in .-9
One interesting consequence of this theorem is that a ring whose center does 9
not contain a field is not an algebra over field . This happens, for example, any -
with the ring . {
The Regular Representation of an Algebra
An algebra homomorphism is called a of the B¢(¦ ²=³ - representation
algebra in . A representation is if it is injective, that is, if (² = ³B - faithful
is an embedding. In this case, is isomorphic to a subalgebra of . (² = ³ B-
Actually, the endomorphism algebras are the most general algebras B-²= ³
possible, in the sense that any algebra has a faithful representation in some (
endomorphism algebra.
Theorem 18.2 Any associative -algebra is isomorphic to a subalgebra of -(
the endomorphism algebra . In fact, if is the left multiplication map B-²(³
defined by
%~ %
then the map is an algebra embedding, called the B¢(¦ ²(³ left regular
representation of .(
When , we can select an ordered basis for and represent dim²(³ ~ B ( 8
the elements of by matrices. This gives an embedding of into the B-²(³ (
matrix algebra , called the of C²-³ ( left regular matrix representation
with respect to the ordered basis . 8
Example 18.7 Let be a finite cyclic group. Let. ~ ¸Á ÁÃÁ ¹c
8 ~² Á ÁÃÁ ³c
be an ordered basis for the group algebra . The multiplication map that -´.µ
is multiplcation by is a shifting of with wraparound and so the matrix 8()
representation of is the matrix whos e columns are obtain ed from the identity
matrix by shifting columns to the right with wrap around . For example, ()
´µ~Ä
Ä
Æ
ÅÅÆÅÅ
Ä8vy
x{x{x{x{
wz
These matrices are called . circulant matrices
458 Advanced Linear Algebra
Since the endomorphism algebras are of obvious importance, let us B-²= ³
examine them a bit more closely.
Theorem 18.3 Let be a vector space over a field .=-
1 The algebra has center) B-²= ³
A~¸ -¹
and so is central.B-²= ³
2 The set of all elements of that have finite rank is an ideal of) 0² = ³ B-
BB--²= ³ ²= ³ and is contained in all other ideals of .
3 is simple if and only if is finite-dimensional.)B-²= ³ =
Proof. We leave the proof of parts 1) and 3) as exercises. For part 2), we leave
it to the reader to show that is an ideal of . Let be a nonzero ideal of 0² = ³ 1 B-
BB 8 8 8-- ²= ³ ²= ³ ~ r. Let have rank . Then there is a basis (a
disjoint union) and a nonzero for which is a finite set, $ = ² ³ ~ ¸¹ 88
and for all . Thus, is a linear combination over of²³~ $ - 8
endomorphisms defined by
²³ ~ $Á ² ±¸¹³ ~ ¸¹ 8
Hence, we need only show that . 1
If is nonzero, then there is an for which . If 8 B1 ~"£ ² =³ -
is defined by
8 ~ Á ² ±¸¹³ ~ ¸¹
and is defined byB² = ³-
8" ~ $Á ² ±¸"¹³ ~ ¸¹
then
8²³ ~ $Á ² ±¸¹³ ~ ¸¹
and so .~ 1
Annihilators and Minimal Polynomials
If is an -algebra an , then it may happen that satisfies a nonzero (- (
polynomial . This always happens, in particular, if is finite- ²%³ -´%µ (
dimensional, since in this case the powers
ÁÁ ÁÃ
must be linearly dependent and so ther e is a nonzero polynomial in that is
equal to .
Definition Let be an -algebra. An element is if there is a(- ( algebraic
nonzero polynomial for which . If is algebraic, the ²%³ -´%µ ²³ ~
An Introduction to Algebras 459
monic polynomial of smallest degree that is satisfied by is called the ² % ³
minimal polynomial of .
If is algebraic over , then the subalgebra generated by over is( - -
-´µ ~ ¸²³ ²%³ -´%µÁ ²³ ² ³¹ deg deg
and this is isomorphic to the quotient algebra
-´µ-´%µ
º ²%³»
where is the ideal generated by the minimal polynomial of . We leaveº ²%³»
the details of this as an exercise.
The minimal polynomial can be used to tell when an element is invertible.
Theorem 18.4
1 The minimal polynomial of generates the of ,) ² % ³ ( annihilator
that is, the ideal
ann² ³~¸ ² % ³-´ % µ² ³~ ¹
of all polynomials that annihilate .
2 The element is invertible if and only if has nonzero constant) ( ² % ³
term.
Proof. We prove only the second statement. If is invertible but
²%³ ~ %²%³
then . Multiplying by gives , which contradicts ~ ²³ ~ ²³ ²³ ~ c
the minimality of . Conversely, if deg² ²%³³
² % ³~ b %bÄb % b% c c
where , then£
~ b bÄb b c c
and so
c² b bÄb b ³~
c c c
and so
~ ²b b Ä b b ³cc c c
c
460 Advanced Linear Algebra
Theorem 18.5 If is a finite-dimensional -algebra, then every element of (- (
is algebraic. There are infinite-dimensional algebras in which all elements are
algebraic.
Proof. The first statement has been proved. To prove the second, let us consider
the complex field as a -algebra. The set of algebraic elements of is a dr d (
field, known as the field of . These are the complex numbers algebraic numbers
that are roots of some nonzero polynomial with rational or integral ()
coefficients.
To see this, if , then the subalgeb ra is finite-dimensional. Also, ( ´ µ ´ µ rr
is a field. To prove this, first note that since is a field, the minimal polynomiald
of any nonzero is irreducible, for if , then ( ²%³ ~ ²%³²%³
~ ²³²³ ²³ ²³ and so one of and is , which implies that
²%³ ~ ²%³ ²%³ ~ ²%³ ²%³ or . Since is irreducible, it has nonzero
constant term and so the inverse of is a polynomial in , that is, . ´ µcr
Of course, is closed under addition and multiplication and so is a rr´µ ´µ
subfield of .d
Thus, is an algebra over . By similar reasoning, if , then thedr ´µ (
minimal polynomial of over is irreducible and so . Since ´µ ´µ´µrrc
rr´µ´µ ~ ´Áµ is the set of all polynomials in the “variables” and , it is
closed under addition and multiplication as well. Hence, is a finite- r´Áµ
dimensional algebra over , as well as a subfield of . Now, rd´µ
dim dim dimrr r ²´ Á µ ³ ~ ²´ Á µ ³ h ²´ µ ³rr r ´µ
and so is finite-dimensional over . Hence, the elements of arerr r´Áµ ´Áµ
algebraic over , that is, . But contains and rr r ´Áµ ( ´Áµ ÁbÁcc
( and so is a field.
We claim that is not finite-dimensi onal over . This follows from the fact ( r
that for every prime , the polynomial is irreducible over by % c r(
Eisenstein's criterion . Hence, if is a complex root of , then has ) % c
minimal polynomial over and so the dimension of over is . %c ´ µ rr r
Hence, cannot be finite-dimensional.(
The Spectrum of an Element
Let be an algebra. A nonzero element is a if ( ( ~ left zero divisor
for some and a if for some . In the £ ~ £ right zero divisor
exercises, we ask the reader to show that an element is a left zero algebraic
divisor if and only if it is a right zero divisor.
Theorem 18.6 Let be a algebra. An algebraic element is invertible if( (
and only if it is not a zero divisor.
Proof. If is invertible and , then multiplying by gives . ~ ~ c
Conversely, suppose that is not invertible but implies . Then ~ ~
An Introduction to Algebras 461
²%³ ~ %²%³ ²%³ ~ ²³ for some nonzero polynomial and so , which
implies that , a contradiction to the minimality of . ²³ ~ ²%³
We have seen that the eigenvalues of a linear operator on a finite-dimensional
vector space are the roots of the minimal polynomial of , or equivalently, the
scalars for which is not invertib le. By analogy, we can define the c
eigenvalues of an element of an algebra . (
Theorem 18.7 Let be an algebra and let be algebraic. An element( (
- ² % ³ c is a root of the minimal polynomial if and only if is not
invertible in . (
Proof. If is not invertible, thenc
²%³ ~ %²%³c
and since is satisfied by , it follows that ² %b ³ c
%²%³ ~ ²%³ ²%b³ c
Hence, . Alternatively, if is not invertible, then there is ²%c³ ²%³ c
a nonzero such that , that is, . Hence, for any ( ² c ³ ~ ~
polynomial we have . Setting gives ²%³ ²³ ~ ²³ ²%³ ~ ²%³
² ³~ .
Conversely, if , then and so ² ³~ ² % ³~² %c ³ ² % ³
~² c ³ ² ³ c , which shows that is a zero divisor and therefore not
invertible.
Definition Let A be an -algebra and let be algebraic. The roots of the - (
minimal polynomial of are called the of . The set of all eigenvalues
eigenvalues of
Spec² ³~¸ -² ³~ ¹
is called the of . spectrum
Note that is invertible if and only if . ( ¤ ² ³ Spec
Theorem 18.8 The Let be an algebra over an() spectral mapping theorem (
algebraically closed field . Let and let . Then - ( ² % ³-´ % µ
Spec Spec Spec²²³³ ~ ² ²³³ ~ ¸²³ ²³¹
Proof. We leave it as an exercise to show that . For ² ²³³ ²²³³Spec Spec
the reverse inclusion, let and suppose that ²²³³Spec
²%³c~²%c³ IJ%c ³
Then
462 Advanced Linear Algebra
²³c~²c³ IJc ³
and since the left-hand side is not i nvertible, neither is one of the factors
c ²³, whence . But Spec
² ³c ~
and so . Hence, . ~ ² ³ ² ²³³ ²²³³ ² ²³³ Spec Spec Spec
Theorem 18.9 Let be an algebra over an algebraically closed field . If(-
Á ( , then
Spec Spec²³ ~ ²³
Proof. If , then is invertible and a simple computation£¤ ² ³ c Spec
gives
² c³´²c³ cµ ~ c
and so is invertible and . If , then is c ¤ ²³ ¤ ²³ Spec Spec
invertible. We leave it as an exercise to show that this implies that is also
invertible and so . Thus, and by symmetry, ¤ ²³ ²³ ²³Spec Spec Spec
equality must hold.
Division Algebras
Some important associative algebras have the property that all nonzero(
elements are invertible and yet is not a field since it is not commutative. (
Definition An associative algebra over a field is a if +- division algebra
every nonzero element has a multiplicative inverse.
Our goal in this section is to classify all finite-dimensional division algebras
over the real field , over any algebraically closed field and over any finite s -
field. The classification of finite-dimensional division algebras over the rational
field is quite complicated and we will not treat it here.r
The Quaternions
Perhaps the most famous noncommutative division algebra is the following.
Define a real vector space with basis i
8~ ¸ÁÁÁ¹
To make into an -algebra, define the product of basis vectors as follows:i -
1 for all )% ~ % ~ % % 8
2)~ ~ ~ c
3) ~ Á ~ Á ~
4) ~c Á~c Á ~c
An Introduction to Algebras 463
Note that 3 can be stated as follows: The product of two consecutive elements )
ÁÁ &% ~ c%& is the next element with wraparound . Also, 4 says that for () )
%Á&¸ÁÁ¹ . This product is extended to all of by distributivity. i
We leave it to the reader to verify that is a division algebra, called Hamilton'si
quaternions , after their discoverer William Rowan Hamilton 1805-1865 . ()
(Readers familiar with group theory will recognize the quaternion group 8~
¸fÁfÁfÁf¹ . The quaternions have applications in geometry, computer)
science and physics.
Finite-Dimensional Division Algebras over an Algebraically Closed
Field
It happens that there are no interesti ng finite-dimensional division algebras over
an algebraically closed field.
Theorem 18.10 If is a finite-dimensional division algebra over an +
algebraically closed field then . -+ ~ -
Proof. Let have minimal polynomial . Since a division algebra has + ² % ³
no zero divisors, must be irreducible over and so must be linear. ² % ³ -
Hence, and so .² % ³~%c ~-
Finite-Dimensional Division Algebras over a Finite Field
The finite-dimensional division algebra s over a finite field are also easily
described: they are all commutative and so are finite fields. The proof, however,
is a bit more challenging. To understand the proof, we need two facts: the class
equation and some information about the complex roots of unity. So let us
briefly describe what we need.
The Class Equation
Those who have studied group theory have no doubt encountered the famous
class equation. Let be a finite group. Each can be thought of as a . .
permutation of defined by .
c%~ %
for all . The set of all conjugates of is denoted by and so% . % % ¸%¹c .
¸%¹ ~ ¸ % .¹.
This set is also called a in . Now, the following are conjugacy class .
equivalent:
464 Advanced Linear Algebra
c c
c c
c
.%~ %
% ~ %
% ~ %
* ² % ³
where
* ² % ³~¸ . %~% ¹.
is the of . But if and only if and are in the same centralizer % * ² % ³ c.
coset of . Thus, there is a one-to-one correspondence between the *² % ³.
conjugates of and the cosets of . Hence, %* ² % ³ .
bb¸%¹ ~ ². ¢ * ²%³³.
.
Since the distinct conjugacy classes form a partition of (because conjugacy is .
an equivalence relation), we have
(( bb . ~ ¸%¹ ~ ². ¢ * ²%³³
%: %:.
.
where is a set consisting of exactly one element from each conjugacy class:
¸%¹ ¸%¹ % ~ %.. c . Note that a conjugacy class has size if and only if for
all , that is, for all and these are precisely the elements in. % ~ % .
the center of . Hence, the previous equation can be written in the form A².³ .
(( ( ( .~A ² . ³b ² . ¢ *² % ³ ³
%:.
Z
where is a set consisting of exactly one element from each conjugacy class:Z
¸%¹ .. of size greater than . This is the for . class equation
The Complex Roots of Unity
If is a positive integer, then th e complex th are the complex roots of unity
solutions to the equation
%c ~
The set of complex th roots of unity is a cyclic group of order . To see<
this, note first that is an abelian group since implies that < Á< <
and . Also, since has no multiple roots, has order . < % c < c
Now, in any finite abelian group , if is the maximum order of all elements .
of , then for all . Thus, if no element of has order , then. ~ . <
. % c~ and every satisfies the equation , which has fewer
than solutions. This contradiction im plies that some element of must have <
order and so is cyclic.<
An Introduction to Algebras 465
The elements of that generate , that is, the elements of order are called <<
the th roots of unity. We denot e the set of primitive th roots of primitive
unity by . Hence, if , then++
+~¸ ² Á³~ ¹
has size , where is the Euler phi function. The value is defined to ²³ ²³ (
be the number of positive integers less than or equal to and relatively prime to
.)
The th is defined bycyclotomic polynomial
8² % ³~ ² % c $ ³
$
+
Thus,
deg²8 ²%³³ ~ ²³
Since every th root of unity is a prim itive th root of unity for some and
since every primitive th root of unity for is also an th root of unity, we
deduce that
<~
+
where the union is a disjoint one. It follows that
%c ~ 8 ² % ³
Finally, we show that is monic and has integer coefficients by induction 8² % ³
on . It is clear from the definition that is monic. Since ,8 ² % ³ 8 ² % ³ ~ % c
the result is true for . If is a prime, then all nonidentity th roots of unity ~
are primitive and so
8 ² % ³ ~ ~ %b %b Ä b % b %c
%c
c c 2
and the result holds for . Assume the result holds for all proper divisors of ~
. Then
% c ~ 8 ²%³ 8 ²%³ ~ 8 ²%³9²%³
By the induction hypothesis, has in teger coefficients and it follows that 9²%³
8² % ³ must also have integer coefficients.
Wedderburn's Theorem
Now we can prove Wedderburn's theorem.
466 Advanced Linear Algebra
Theorem 18.11 If is a finite division algebra, then()Wedderburn's theorem +
+ is a field.
Proof. We must show that is co mmutative. Let be the multiplicative+. ~ +i
group of all nonzero elements of . The class equation is +
((( ( +~ A ² + ³ b ² + ¢ * ² ³ ³ii i
where the sum is taken over one representative from each conjugacy class of
size greater than . If we assume for the purposes of contradiction that is not +
commutative, that is, that , then th e sum on the far right is not an A²+ ³£+ii
empty sum and so for some . (( ( (*²³+ +ii i
The sets and are subalgebras of and, in fact, is a A²+³ *² ³ + A²+³
commutative division algebra; th at is, a field. Let . Since ((A² + ³ ~'
A²+³ *² ³ *² ³ + A²+³ , we may view and as vector spaces over and so
(( ( (*² ³ ~' + ~'² ³ and
for integers . The class equation now gives ² ³
'c ~ ' c b'c
'c
² ³
² ³
and since , it follows that . 'c ' c ² ³ ² ³
If is the th cyclotomic poly nomial, then divides . But 8² % ³ 8² ' ³ ' c
8² ' ³ also divides each summand on the far right above, since its roots are not
roots of . It follows that . On the other hand,'c 8 ² ' ³ ' c ² ³
8² ' ³~ ² 'c ³
+
and since implies that , we have a contradiction. Hence + ' c ' c ((
A²+ ³~+ + +ii and is commutative, that is, is a field.
Finite-Dimensional Real Division Algebras
We now consider the finite-dimensional division algebras over the real field . s
In 1877, Frobenius proved that there are only three such division algebras.
Theorem 18.12 If is a finite-dime nsional division algbera ()Frobenius, 1877 +
over , thens
+~ Á +~ +~sd i or
Proof. Note first that the minimal polynomial of any is either ² % ³ +
linear, in which case or irreducible quadratic with ² % ³~% b %b s
c . Completing the square gives
An Introduction to Algebras 467
~²³~ bb ~ b b ² c³
67
Hence, any has the form +
~ b c ~ b!
67
where and either or . Hence, but . Thus, every! ~ ¤s s s
element of is the sum of an element of and an element of the set + s
+~ ¸ + ¹Z
that is, as sets:
+~ b+sZ
Also, . To see that is a subspace of , let . We wishsq+ ~ ¸¹ + + "Á# +ZZ Z
to show that . If for some , then "b#+ #~" Zs
"b# ~ ²b³" + " #Z. So assume that and are linearly independent. Then
"# and are nonzero and so also nonreal.
Now, and cannot both be real, si nce then and would be real. We "b# "c# " #
have seen that
"b#~b
and
"c#~ b
where , at least one of or is nonzero and . ThenÁ Á s
²"b#³ b²"c#³ ~²b ³ b² b ³
and so
" b# ~ b b b b b
Collecting the real part on one side gives
b ~ " b# c² b b b ³
Now, if we knew that and were linearly independent over we could sÁ
conclude that and so ~ ~
²"b#³ ~ ²"c#³ ~ and
which shows that and are in . "b# "c# +Z
To see that is linearly independe nt, it is equivalent to show that ¸ÁÁ ¹
¸"Á#Á¹ is linearly independent. But if
468 Advanced Linear Algebra
#~ "b
for , thenÁ s
#~ "b " b
and since , it follows that and so or . But since "¤ ~ ~ ~ £s
#¤ £ ¸ " Á# ¹s and since are linearly independent.
Thus, is a subspace of and++Z
+~ l+sZ
We now look at , which is a real vector space. If , then and + + ~ ¸¹ + ~ZZs
we are done, so assume otherwise. If is nonzero, then where + ~c Z
~ + ~c s. Hence, satisfies . Ifc Z
+ ~ ~¸ ¹Zss
then and we are done. If not, then is a proper subspace of+~ l ~ ss d s
+Z.
In the quaternion field, there is an element for which . So we seek a b ~
+± +Z Zs with this property. To this end, define a bilinear form on by
º"Á#» ~ c²"#b#"³
Then it is easy to see that this form is a real inner product on positive +Z(
definite, symmetric and bilinear . Hence, if is a proper subspace of , then ) s+Z
+~ p :Zs
where denotes the orthogonal direct sum. If is nonzero, thenp" :
"~ c ~ " c for and so if , thens
~ c b ~ and
Now, is a subspace of and sos:
+~ p p ;Zss
Setting , we have~
cºÁ» ~ b ~ b ~
and
c º Á»~ b~ b ~
and so and we can write;
+~ p p p <Zsss
An Introduction to Algebras 469
Now, if , then there is a for which and< £ ¸¹ " < " ~ c
" ~ c"
" ~ c"
" ~ c"
The third equation is and so "²³ ~ c²³"
"²³ ~ c²³" ~ " ~ c"
whence , which is false. Hence, and" ~ < ~ ¸¹
+ ~ l l l ~ss s s i
This completes the proof.
Exercises
1. Prove that the subalgebra generated by a nonempty subset of an algebra ?
( is the subspace spanned by the products of finite subsets of elements of
?:
º?» ~ º% Ä% % ?» alg
2. Verify that the group algebra is i ndeed an associative algebra over . -´.µ -
3. Show that the kernel of an algebra homomorphism is an ideal.
4. Let be a finite-dimensional algebra over and let be a subalgebra.(- )
Show that if is invertible, then . ) )c
5. If is an algebra and is nonempty, define the of(: ( * ² : ³ centralizer (
:( : to be the set of elements of that commute with all elements of . Prove
that is a subalgebra of .*² : ³ ((
6. Show that is not an algebra over any field. {
7. Let be the algebra generated ove r by a single algebraic element (~-´ µ -
( -´%µ°º²%³». Show that is isomorphic to the quotient algebra , where
º²%³» ²%³ -´%µ is the ideal generated by . What can you say about
²%³ ( ? What is the dimension of ? What happens if is not algebraic?
8. Let be a finite group. For of the form.~¸~ ÁÃÁ ¹ %-´.µ
%~ bÄb
let . Prove that is an algebra;²%³~ bÄb ;¢-´.µ¦-
homomorphism, where is an algebra over itself. -
9. Prove the of algebras: A homomorphism first isomorphism theorem
¢(¦) - ¢(° ² ³ ² ³ of -algebras induces an isomorphism ker im
defined by . ² ² ³³ ~ ker
10. Prove that the quaternion field is an -algebra and a field. : For- Hint
%~bbb£
470 Advanced Linear Algebra
()~ consider
%~ccc
11. Describe the left regular represent ation of the quaternions using the ordered
basis .8~² Á Á Á³
12. Let be the group of permutations bijective functions of the ordered set: ()
?~²% ÁÃÁ% ³ , under composition. Verify the following statements.
Each defines a linear isomorphism on the vector space with: =
basis over a field . This defines an algebra homomorphism?-
¢-´: µ¦ ²=³ ² ³~ -B with the property that . What does the matrix
representation of a look like? Is the representation faithful? :
13. Show that the center of the algebra is B-²= ³
A~¸ -¹
14. Show that is simple if and only if . B-²= ³ ²= ³ B dim
15. Prove that for , the matrix algebras are central and simple. ² -³ C
16. An element is if there is a for which , in ( ( ~ left-invertible
which case is called a of . Similarly, is ( left inverse right-
invertible if there is a for which , in which case is called a ( ~
right inverse one-sided inverses of . Left and right inverses are called
and an ordinary inverse is called a . Let be two-sided inverse (
algebraic over . -
a Prove that for some if and only if for some . ) ~ £ ~ £
Does necessarily equal ?
b Prove that if has a one-sided inverse , then is a two-sided inverse. )
Does this hold if is not algebraic? : Consider the algebra Hint
( ~ ²-´%µ³B- .
c Let be algebraic. Show that is invertible if and only if )Á (
and are invertible, in which case is also invertible.
Chapter 19
The Umbral Calculus
In this chapter, we give a brief introduction to an area called the umbral
calculus . This is a linear-algebraic theory used to study certain types of
polynomial functions that play an im portant role in app lied mathematics. We
give only a brief introduction to the s ubject, emphasizing the algebraic aspects
rather than the applications. For more on the umbral calculus, may we suggest
The Umbral Calculus , by Roman 1984 ? ´µ
One bit of notation: The are defined by lower factorial numbers
²³ ~²c³Ä²cb³
Formal Power Series
We begin with a few remarks concerning formal power series. Let denote the <
algebra of formal power series in the variable , with complex coefficients. !
Thus, is the set of all formal sums of the form<
²!³~ !
~B
()19.1
where the complex numbers . Addition and multiplication are purelyd()
formal:
~ ~ ~BBB
! b ! ~ ² b ³!
and
45 45 4 5
~ ~ ~BB B
c
~! ! ~ !
The of is the smallest exponent of that appears with a nonzero order²³ !
coefficient. The order of the zero series is defined to be . Note that a series bB
472 Advanced Linear Algebra
² ³ ~ has a multiplicative inverse, denoted by , if and only if . Wec
leave it to the reader to show that
²³ ~ ²³b²³
and
min² b³ ¸²³Á²³¹
If is a sequence in with as , then for any series ² ³ ¦ B ¦ <
²!³ ~ !
~B
we may substitute for to get the series !
²!³ ~ ²!³
~B
which is well-defined since the coefficient of each power of is a finite sum. In !
particular, if , then and so the ²³ ² ³ ¦ Bcomposition
²k³²!³ ~ ²²!³³ ~ ²!³
~B
is well-defined. It is easy to see that . ²k³ ~ ²³²³
If , then has a compositional inverse, denoted by and satisfying²³ ~
² k³²!³ ~ ² k³²!³ ~ !
A series with is called a . ² ³ ~ delta series
The sequence of powers of a delta series forms a for , in the pseudobasis <
sense that for any , there exists a unique sequence of constants for <
which
²!³ ~ ²!³
~B
Finally, we note that the formal deri vative of the series 19.1 is given by ()
C² ! ³~² ! ³~ !!Z c
~B
The operator is a derivation, that is, C!
C²³~C²³bC²³!! !
The Umbral Calculus 473
The Umbral Algebra
Let denote the algebra of polynomials in a single variable over the Fd~´ % µ %
complex field. One of the starting point s of the umbral calculus is the fact that
any formal power series in can play three different roles: as a formal power <
series, as a linear functional on and as a linear operator on . Let us first FF
explore the connection between formal power series and linear functionals.
Let denote the vector space of all linear functionals on . Note that is theFF Fi i
algebraic dual space of , as defined in Chapter 2. It will be convenient to F
denote the action of on by 3 ² % ³FFi
º3 ²%³»
()This is the “bra-ket” notation of Paul Dirac. The vector space operations on Fi
then take the form
º3b4 ²%³» ~ º3 ²%³»bº4 ²%³»
and
º 3 ² % ³ »~ º 3 ² % ³ » Á d
Note also that since any linear functiona l on is uniquely determined by itsF
values on a basis for the functional is uniquely determined by the FFÁ3 i
values for .º3 % »
Now, any formal series in can be written in the form <
²!³~ !
~B
!
and we can use this to define a linear functional by setting ²!³
º²!³ % » ~
for . In other words, the linear functional is defined by ² ! ³
²!³~ !º²!³ % »
~B
!
where the expression on the left is just a formal power series. Note in ²!³
particular that
º! % » ~
Á!
where is the Kronecker delta function. This implies thatÁ
º! ²%³» ~ ²³² ³
and so is the functional “ th derivativ e at .” Also, is evaluation at . ! !
474 Advanced Linear Algebra
As it happens, any linear functional on has the form . To see this, we 3 ² ! ³F
simply note that if
² ! ³ ~ !º3 % »
3
~B
!
then
º ²!³ % » ~ º3 % »3
for all and so as linear functionals, . 3~ ² ! ³ 3
Thus, we can define a map by . F < ¢¦ ² 3 ³ ~ ² ! ³i3
Theorem 19.1 The map defined by is a vector spaceF < ¢¦ ² 3 ³ ~ ² ! ³i3
isomorphism from onto . F<i
Proof. To see that is injective, note that
² ! ³~ ² ! ³¬º 3%»~º 4%» ¬3~434 for all
Moreover, the map is surjective, since for any , the linear functional <
3 ~ ²!³ ²3³ ~ ²!³ ~ ²!³ has the property that . Finally, 3
²3b 4³ ~ !º3b 4 % »
~ ! b !º3 % » º4 % »
~ ² 3 ³b ² 4³
~B
~ ~BB
!
!!
From now on, we shall identify the vector space with the vector space , F<i
using the isomorphism . Thus, we think of linear functionals on F < F¢¦i
simply as formal power series. The advantage of this approach is that is more <
than just a vector space—it is an algebra. Hence, we have automatically defined
a multiplication of linear functionals, namely, the product of formal power
series. The algebra , when thought of as both the algebra of formal power <
series and the algebra of linear functionals on , is called the . F umbral algebra
Let us consider an example.
Example 19.1 For , the is defined by d F evaluation functional i
º ²%³» ~ ²³
The Umbral Calculus 475
In particular, and so the formal power series representation for º % » ~
this functional is
² ! ³ ~ !~ !~ º % »
~ ~BB
!
!!
which is the exponential series. If is evaluation at , then !
~ ! ! ²b³!
and so the product of evaluation at a nd evaluation at is evaluation at
b .
When we are thinking of a delta series as a linear functional, we refer to <
it as a . Similarly, an invertible series is referred to as an delta functional <
invertible functional . Here are some simple consequences of the development
so far.
Theorem 19.2
1 For any ,) <
²!³~ !º²!³ % »
[
~B
2 For any ,) F
²%³ ~ %º! ²%³»
[
3 For any ,) Á<
º²!³²!³ % » ~ º²!³ % »º²!³ % »
c
~
45
4)²²!³³ ²%³¬º²!³²%³»~ deg
5 If for all , then)² ³ ~
Lc M
~ B
²!³ ²%³ ~ º ²!³ ²%³»
where the sum on the right is a finite one.
6 If for all , then)² ³ ~
º ² ! ³ ² % ³ »~º ² ! ³ ² % ³ » ¬ ² % ³~ ² % ³ for all
7 If for all , then)d e g ² % ³~
º ² ! ³² % ³ »~º ² ! ³² % ³ » ¬² ! ³~ ² ! ³ for all
476 Advanced Linear Algebra
Proof. We prove only part 3 . Let )
²!³ ~ ! ²!³ ~ !
~BB
~
!! and
Then
²!³²!³ ~ !
44 55
~B
~ c
!
and applying both sides of this as linear functionals to gives ²% )
º²!³²!³ % » ~
~
c 45
The result now follows from the fact that part 1 implies and ) ~ º²!³ % »
~ º ² ! ³ % »cc.
We can now present our first “umbral” result.
Theorem 19.3 For any and ,²!³ ²%³<F
º²!³ %²%³» ~ ºC ²!³ ²%³» !
Proof. By linearity, we need only establish this for . But if ²%³ ~ %
²!³~ !
~B
!
then
ºC ²!³ % » ~ ! %
² c³
~
² c³
~
~º ² ! ³% »! c
~B
~B
cÁ
b
bLc M
!
!
Let us consider a few examples of important linear functionals and their power
series representations.
Example 19.2
1 We have already encountered the , satisfying) evaluation functional !
º ²%³» ~ ²³!
The Umbral Calculus 477
2 The is the delta functional ,) forward difference functional c !
satisfying
º c ²%³» ~ ²³c²³!
3 The is the delta functional e , satisfying) Abel functional !!
º! ²%³» ~ ²³e! Z
4 The invertible functional satisfies) ²c!³c
º²c!³ ²%³» ~ ²"³ "c c"
B
as can be seen by setting and expanding the expression ²%³ ~ %
²c!³c.
5 To determine the linear functional satisfying)
º²!³ ²%³» ~ ²"³"
we observe that
²!³~ ! ~ ! ~º²!³ % » ! c
² b ³ [ !
~ ~BB b !
!
The inverse of this functiona l is associated with the Bernoulli !°² c³!
polynomials, which play a very important role in mathematics and its
applications. In fact, the numbers
)~ %!
c !Lc M
are known as the . Bernoulli numbers
Formal Power Series as Linear Operators
We now turn to the connection between formal power series and linear
operators on . Let us denote the th derivative operator on by . Thus, FF !
! ² % ³~ ² % ³² ³
We can then extend this to formal series in , !
²!³~ !
~B
!19.2()
478 Advanced Linear Algebra
by defining the linear operator by ²!³¢ ¦FF
²!³²%³ ~ ´! ²%³µ ~ ²%³
~ B
² ³
!!
the latter sum being a finite one. Note in particular that
²!³% ~ %
c
~
45 ()19.3
With this definition, we see that each formal power series plays three <
roles in the umbral calculus, namely, as a formal power series, as a linear
functional and as a linear operator. The two notations and º²!³ ²%³»
²!³²%³ will make it clear whether we are thinking of as a functional or as an
operator.
It is important to note that in if and only if as linear functionals, ~ ~<
which holds if and only if as linear operators. It is also worth noting that ~
´²!³²!³µ²%³ ~ ²!³´²!³²%³µ
and so we may write without ambiguity. In addition, ²!³²!³²%³
²!³²!³²%³ ~ ²!³²!³²%³
for all and .Á <F
When we are thinking of a delta series as an operator, we call it a delta
operator . The following theorem describes the key relationship between linear
functionals and linear operators of the form . ²!³
Theorem 19.4 If , thenÁ<
º²!³²!³ ²%³» ~ º²!³ ²!³²%³»
for all polynomials . ²%³ F
Proof. If has the form 19.2 , then by 19.3 , () ()
º ! ! ³ %»~ ! % ~ ~º ² ! ³%» ²
c
~
() Lc 45 M 19.4
By linearity, this holds for replaced by any polynomial . Hence, % ² % ³
applying this to the product gives
º²!³²!³ ²%³» ~ º! ²!³²!³²%³»
~ º! ²!³´²!³²%³µ» ~ º²!³ ²!³²%³»
Equation 19.4 shows that applying the linear functional is equivalent to () ( ! ³
applying the operator and then following by evaluation at . ²!³ %~
The Umbral Calculus 479
Here are the operator versions of the functionals in Example 19.2.
Example 19.3
1 The operator satisfies) !
%~ ! %~ % ~ ² % b ³
! c
~ ~B
45!
and so
² % ³ ~ ² % b ³!
for all . Thus is a . F!translation operator
2 The is the delta operator , where) forward difference operator c !
² c ²%³ ~ ²%b³c²³!)
3 The is the delta operator e , where) Abel operator !!
! ² % ³ ~ ² % b ³e! Z
4 The invertible operator satisfies) ²c!³c
²c!³ ²%³ ~ ²%b"³ "c c"
B
5 The operator is easily seen to satisfy)) ² c °!!
c
!²%³ ~ ²"³"!
%%b
We have seen that all linear functionals on have the form , for . F< ²!³
However, not all linear operators on have this form. To see this, observe thatF
deg deg´²!³²%³µ ²%³
but the linear operator defined by does not have F F ¢ ¦ ²²%³³ ~ %²%³
this property.
Let us characterize the linear operators of the form . First, we need a lemma. ²!³
Lemma 19.5 If is a linear operator on and for some delta;7 ; ² ! ³ ~ ² ! ³ ;
series , then .²!³ ²;²%³³ ²²%³³ deg deg
Proof. For any ,
deg deg deg²;% ³c ~ ²²!³;% ³ ~ ²;²!³% ³
and so
480 Advanced Linear Algebra
deg deg²;% ³ ~ ²;²!³% ³b
Since we have the basis for an induction. When deg²²!³% ³ ~ c ~
we get . Assume that the result is true for . Then deg²;³ ~ c
deg deg²;% ³ ~ ²;²!³% ³b cb ~
Theorem 19.6 The following are equivalent for a linear operator . ;¢ ¦FF
1 has the form , that is, there exists an for which , as); ²!³ ; ~ ²!³ <
linear operators.
2 commutes with the derivative operator, that is, .);; ! ~ ! ;
3 commutes with any delta operator , that is, .); ²!³ ;²!³ ~ ²!³;
4 commutes with any translation operator, that is, .);; ~ ;! !
Proof. It is clear that 1 implies 2 . For the converse, let ))
²!³ ~ !º! ;% »
~B
!
Then
º²!³ % » ~ º! ;% »
Now, since commutes with , we have ;!
º ! ; %»~º !!; %»
~º ! ;!%»
~² ³º ! ;% »
~ ²³ º! ²!³% »
~º ! ² ! ³ %»
c
c
and since this holds for all and we get . We leave the rest of the ; ~ ² ! ³
proof as an exercise.
Sheffer Sequences
We can now define the principal object of study in the umbral calculus. When
referring to a sequence in , we shall always assume that ² % ³ ² % ³~F deg
for all .
Theorem 19.7 Let be a delta series, let be an invertible series and consider
the geometric sequence
Á Á Á ÁÃ
in . Then there is a unique sequence in satisfying the <F ² % ³ orthogonality
conditions
The Umbral Calculus 481
º²!³ ²!³ ²%³» ~ [
Á (19.5)
for all .Á
Proof. The uniqueness follows from Theorem 19.2. For the existence, if we set
² % ³~ % Á
~
and
²!³ ²!³ ~ !
~B
Á
where , then 19.5 is£ Á ()
[ ~ ! %
~ º ! % »
~ [Á Á Á
~B
~
~B
~Á Á
~
Á ÁLc M
c
Taking we get~
~
Á
Á
For we have~c
~ ²c³[b [cÁc Ác cÁ Á
and using the fact that we can solve this for . By ~ ° Á Á Ác
successively taking we can solve the resulting ~ÁcÁcÁÃ
equations for the coefficients of the sequence . ² % ³Á
Definition The sequence in 19.5 is called the for the ² % ³ () Sheffer sequence
ordered pair . We shorten this by saying that is ²²!³Á²!³³ ²%³ Sheffer for
²²!³Á²!³³ .
Two special types of Sheffer sequences deserve explicit mention.
Definition The Sheffer sequence for a pair of the form is called the ²Á²!³³
associated sequence for . The Sheffer sequence for a pair of the form²!³
²²!³Á!³ ²!³ is called the for . Appell sequence
482 Advanced Linear Algebra
Note that the sequence is Sheffer for if and only if ²%³ ²²!³Á²!³³
º²!³ ²!³ ²%³» ~ [
Á
which is equivalent to
º ²!³ ²!³ ²%³» ~ [
Á
which, in turn, is equivalent to saying that the sequence is ²%³ ~ ²!³ ²%³
the associated sequence for . ²!³
Theorem 19.8 The sequence is Sheffer for if and only if the ²%³ ²²!³Á²!³³
sequence is the associated sequence for . ²%³ ~ ²!³ ²%³ ²!³
Before considering examples, we wish to describe several characterizations of
Sheffer sequences. First, we require a key result.
Theorem 19.9 The expansion theorems () Let be Sheffer for . ²%³ ²²!³Á²!³³
1 For any ,) <
²!³ ~ ²!³ ²!³º²!³ ²%³»
~B
!
2 For any ,) F
²%³ ~ ²%³º²!³ ²!³ ²%³
!
Proof. Part 1 follows from Theorem 19.2, since )
Lc M
~ ~BB
Á
º²!³ ²%³» º²!³ ²%³»
²!³ ²!³ ²%³ ~
~ º²!³ ²%³»!!!
Part 2 follows in a similar way from Theorem 19.2. )
We can now begin our characterization of Sheffer sequences, starting with the
generating function. The idea of a generating function is quite simple. If is ² % ³
a sequence of polynomials, we may define a formal power series of the form
²!Á%³ ~ !² % ³
~B
!
This is referred to as the for the sequence ()exponential generating function
² % ³ . The term exponential refers to the presence of ! in this series. When(
this is not present, we have an ordinary generating function. Since the series is ³
a formal one, knowing is equivalent in theory, if not always in practice ²!Á%³ ()
The Umbral Calculus 483
to knowing the polynomials . Moreover, a knowledge of the generating ² % ³
function of a sequence of polynomials can often lead to a deeper understanding
of the sequence itself, that might not be otherwise easily accessible. For this
reason, generating functions are studied quite extensively.
For the proofs of the following characterizations, we refer the reader to Roman
´µ1984 .
Theorem 19.10 Generating function ()
1 The sequence is the associated sequence for a delta series if and) ² % ³ ² ! ³
only if
~ !² & ³
&²!³
~B
!
where is the compositional inverse of .²!³ ²!³
2 The sequence is Sheffer for if and only if) ²%³ ²²!³Á²!³³
² & ³
²²!³³~ !&²!³
~B
!
The sum on the right is called the of . generating function ² % ³
Proof. Part 1 is a special case of part 2 . For part 2 , the expression above is )) )
equivalent to
² & ³
²!³ ~ ² ! ³&!
~B
!
which is equivalent to
~ ²!³ ²!³ ² & ³
&!
~B
!
But if is Sheffer for , then th is is just the expansion theorem ²%³ ²²!³Á²!³³
for . Conversely, this expression implies that&!
²&³ ~ º ²%³» ~ º²!³ ²!³ ²%³» ² & ³
&!
~B
!
and so , which says that is Sheffer for º²!³ ²!³ ²%³» ~ [ ²%³ Á
²Á³ .
We can now give a representation for Sheffer sequences.
Theorem 19.11 Conjugate representation ()
484 Advanced Linear Algebra
1 A sequence is the associated sequence for if and only if) ² % ³ ² ! ³
² % ³~ º ² ! ³ %» %
[
~
2 A sequence is Sheffer for if and only if) ²%³ ²²!³Á²!³³
² % ³~ º ² ² ! ³ ³ ² ! ³ %» %
[
~
c
Proof. We need only prove part 2 . We know that is Sheffer for ) ² % ³
²²!³Á²!³³ if and only if
² & ³
²²!³³~ !&²!³
~B
!
But this is equivalent to
NO Lc M ² & ³
²²!³³ % ~ ! % ~ ² & ³&²!³
~B
!
Expanding the exponential on the left gives
Lc M
~ ~BB c
º²²!³³ ²!³ % » ²&³
[ &~ !% ~ ² & ³!
Replacing by gives the result. &%
Sheffer sequences can also be characterized by means of linear operators.
Theorem 19.12 Operator characterization ² )
1 A sequence is the associated sequence for if and only if) ² % ³ ² ! ³
a )² ³~ Á
b f o r )²!³ ²%³ ~ ²%³ c
2 A sequence is Sheffer for for some invertible series if) ²%³ ²²!³Á²!³³ ²!³
and only if
²!³ ²%³ ~ ²%³ c
for all .
Proof. For part 1 , if is associated with , then )² % ³ ² ! ³
² ³~º ² % ³ »~º ² ! ³² % ³ »~ [ Á !
and
The Umbral Calculus 485
º²!³ ²!³ ²%³» ~ º²!³ ²%³»
~ [
~ ²c³[
~ º²!³ ²%³» b
Áb
cÁ
c
and since this holds for all we get 1b . Conversely, if 1a and 1b hold, )) )
then
º²!³ ²%³» ~ º! ²!³ ²%³»
~² ³ ² ³
~² ³
~ [
c
c Á
Á
and so is the associated sequence for .² % ³ ² ! ³
As for part 2 , if is Sheffer for , then ) ²%³ ²²!³Á²!³³
º²!³²!³ ²!³ ²%³» ~ º²!³²!³ ²%³»
~ [
~ ²c³[
~ º²!³²!³ ²%³» b
Áb
cÁ
c
and so , as desired. Conversely, suppose that²!³ ²%³ ~ ²%³ c
²!³ ²%³ ~ ²%³ c
and let be the associated sequence fo r . Let be the invertible linear ² % ³ ² ! ³ ;
operator on defined by =
; ²%³~ ²%³
Then
;²!³ ²%³ ~ ; ²%³ ~ ²%³ ~ ²!³ ²%³ ~ ²!³; ²%³ c c
and so Lemma 19.5 implies that for some invertible series . Then ; ~ ²!³ ²!³
º²!³²!³ ²%³» ~ º²!³ ²!³ ²%³»
~º ! ² ! ³² % ³ »
~² ³ ² ³
~² ³
~ [
c
c Á
Á
and so is Sheffer for . ²%³ ²²!³Á²!³³
We next give a formula for the action of a linear operator on a Sheffer ²!³
sequence.
486 Advanced Linear Algebra
Theorem 19.13 Let be a Sheffer sequence for and let ²%³ ²²!³Á²!³³ ²%³
be associated with . Then for any we have ²!³ ²!³
²!³ ²%³ ~ º²!³ ²%³» ²%³
c
~
45
Proof. By the expansion theorem
²!³ ~ ²!³ ²!³º²!³ ²%³»
~B
!
we have
²!³ ²%³ ~ ²!³ ²!³ ²%³º²!³ ²%³»
~² ³ ² % ³º²!³ ²%³»
~B
~B
c
!
!
which is the desired formula.
Theorem 19.14
1 A sequence is the associated sequence for a)( )The binomial identity ² % ³
delta series if and only if it is of , that is, if and only if it ²!³ binomial type
satisfies the identity
²%b&³ ~ ²&³ ²%³
c
~
45
for all .&d
2 A sequence is Sheffer for for)( )The Sheffer identity ²%³ ²²!³Á²!³³
some invertible if and only if ²!³
²%b&³ ~ ²&³ ²%³
c
~
45
for all , where is the associated sequence for . & ² % ³ ² ! ³d
Proof. To prove part 1 , if is an associated sequence, then taking )² % ³
²!³ ~ &! in Theorem 19.13 gives the binomial identity. Conversely, suppose
that the sequence is of binomial type. We will use the operator ² % ³
characterization to show that is an associated sequence. Taking ² % ³
%~&~ ~ we have for ,
²³ ~ ²³ ²³
and so . Also,² ³~
The Umbral Calculus 487
² ³~² ³ ² ³b² ³ ² ³~ ² ³
and so . Assuming that for we have ²³ ~ ²³ ~ ~ ÁÃÁc
²³ ~ ²³ ²³b ²³ ²³ ~ ²³
and so . Thus, . ²³ ~ ²³ ~ Á
Next, define a linear functional by ²!³
º²!³ ²%³» ~ Á
Since and we deduceº²!³ » ~ º²!³ ²%³» ~ º²!³ ²%³» ~ £
that is a delta series. Now, the binomial identity gives²!³
º²!³ ²%³» ~ ²&³º²!³ ²%³»
~ ² & ³
~ ² & ³&!
c
~
~
c Á
c45
45
and so
º ²!³ ²%³» ~ º ²%³»&! &!
c
and since this holds for all , we get . Thus, is the & ²!³ ²%³ ~ ²%³ ²%³ c
associated sequence for . ²!³
For part 2 , if is a Sheffer sequence, then taking in Theorem ) ² % ³ ² ! ³~&!
19.13 gives the Sheffer identity. Conversely, suppose that the Sheffer identity
holds, where is the associated sequence for . It suffices to show that ² % ³ ² ! ³
²!³ ²%³ ~ ²%³ ²!³ ; for some invertible . Define a linear operator by
; ²%³~ ²%³
Then
; ²%³ ~ ²%³ ~ ²%b&³&! &!
and by the Sheffer identity,
; ²%³ ~ ²&³; ²%³ ~ ²&³ ²%³
&!
c c
~ ~
45 45
and the two are equal by part 1 . Hence, commutes with and is therefore ) ;&!
of the form , as desired. ²!³
488 Advanced Linear Algebra
Examples of Sheffer Sequences
We can now give some examples of Sheffer sequences. While it is often a
relatively straightforward matter to verify that a given sequence is Sheffer for a
given pair , it is quite another matter to find the Sheffer sequence for ²²!³Á²!³³
a given pair. The umbral calculus provides two formulas for this purpose, one of
which is direct, but requires the usually very difficult computation of the series
²²!³°!³ ²%³c . The other is a recurrence relation that expresses each in terms
of previous terms in the Sheffer sequence. Unfortunately, space does not permit
us to discuss these formulas in detail. However, we will discuss the recurrence
formula for associated sequences later in this chapter.
Example 19.4 The sequence is the associated sequence for the delta ² % ³~%
series . The generating function for this sequence is²!³~!
~ !&
[&!
~B
and the binomial identity is the well-known binomial formula
²%b&³ ~ % &
c
~
45
Example 19.5 The lower factorial polynomials
²%³ ~%²%c³Ä²%cb³
form the associated sequence for the forward difference functional
²!³~ c!
discussed in Example 19.2. To see this, we simply compute, using Theorem
19.12. Since is defined to be , we have . Also, ²³ ²³ ~ Á
² c³²%³ ~²%b³ c²%³
~´²%b³%²%c³Ä²%cb³µc´%²%c³Ä²%cb³µ
~%²%c³Ä²%cb³´²%b³c²%cb³µ
~%²%c³Ä²%cb³
~ ² % ³!
c
The generating function for the lo wer factorial polynomials is
~ !²&³
[&² b ! ³
~B
log
The Umbral Calculus 489
which can be rewritten in the more familiar form
²b!³ ~ !&
&
~B
45
Of course, this is a formal identity, so there is no need to make any restrictions
on . The binomial identity in this case is!
²%b&³ ~ ²%³ ²&³
c
~
45
which can also be written in the form
45 4 5 4 5 %b& % &
c ~
~
This is known as the . Vandermonde convolution formula
Example 19.6 The Abel polynomials
(² % Â ³~% ² % c ³c
form the associated sequence for the Abel functional
²!³~! e!
also discussed in Example 19.2. We leave verification of this to the reader. The
generating function for the Abel polynomials is
~ !&²&c³
[&²!³
~B c
Taking the formal derivative of this with respect to gives &
²!³ ~ !²&c³²&c³
[&²!³
~B c
which, for , gives a formula for the compositional inverse of the series &~
²!³~!!,
²!³~ !²c³
²c³[
~B c
Example 19.7 The famous form the Appell Hermite polynomials /² % ³
sequence for the invertible functional
²!³ ~ !°
490 Advanced Linear Algebra
We ask the reader to show that is the Appell sequence for if and only ² % ³ ² ! ³
if . Using this fact, we get ² % ³~ ² ! ³ %c
/² % ³~ % ~ ² c ³ %² ³
[c! ° c
The generating function for the Hermite polynomials is
~ !/² & ³
[&!c! °
~B
and the Sheffer identity is
/² % b & ³~ /² % ³ &
~
c45
We should remark that the Hermite polynomials, as defined in the literature,
often differ from our definition by a multiplicative constant.
Example 19.8 The well-known and important Laguerre polynomials 3² % ³²³
of order form the Sheffer sequence for the pair
²!³ ~ ²c!³ Á ²!³ ~!
!ccc
It is possible to show although we will not do so here that ²³
3² % ³ ~ ² c % ³[ b
[ c²³
~
45
The generating function of the Laguerre polynomials is
²c!³ [~ !3² % ³
b&!°²!c³
~B
²³
As with the Hermite polynomials, some definitions of the Laguerre polynomials
differ by a multiplicative constant.
We presume that the few examples we have given here indicate that the umbral
calculus applies to a significant range of important polynomial sequences. In
Roman 1984 , we discuss approximately 30 different sequences of polynomials ´µ
that are or are closely related to Sheffer sequences. ()
Umbral Operators and Umbral Shifts
We have now established the basic framework of the umbral calculus. As we
have seen, the umbral algebra plays three roles: as the algebra of formal power
series in a single variable, as the algeb ra of all linear functionals on and as the F
The Umbral Calculus 491
algebra of all linear operators on th at commute with the derivative operator.F
Moreover, since is an algebra, we can consider geometric sequences <
Á Á Á ÁÃ
in , where and . We have seen by example that the< ²³ ~ ²³ ~
orthogonality conditions
º²!³ ²!³ ²%³» ~ [
Á
define important families of polynomial sequences.
While the machinery that we have developed so far does unify a number of
topics from the classical study of polynomial sequences for example, special (
cases of the expansion theorem include Taylor's expansion, the Euler–
MacLaurin formula and Boole's summati on formula , it does not provide much )
new insight into their study. Our plan now is to take a brief look at some of the
deeper results in the umbral calculus, which center on the interplay between
operators on and their adjoints, which are operators on the umbral algebra F
<F~i.
We begin by defining two important operators on associated with each F
Sheffer sequence.
Definition Let be Sheffer for . The linear operator ²%³ ²²!³Á²!³³
FFÁ¢¦ defined by
Á ²% ³ ~ ²%³
is called the for the pair , or for the sequence Sheffer operator ²²!³Á²!³³
²%³ ²%³ ²!³. If is the associated sequence for , the Sheffer operator
²% ³ ~ ²%³
is called the for , or for . umbral operator ²!³ ²%³
Definition Let be Sheffer for . The linear operator ²%³ ²²!³Á²!³³
FFÁ¢¦ defined by
Á b´ ²%³µ ~ ²%³
is called the for the pair , or for the sequence . If Sheffer shift ²²!³Á²!³³ ²%³
² % ³ ² ! ³ is the associated sequence for , the Sheffer operator
b ´ ²%³µ ~ ²%³
is called the for , or for . umbral shift ²!³ ²%³
492 Advanced Linear Algebra
It is clear that each Sheffer sequence uniquely determines a Sheffer operator and
vice versa. Hence, knowing the Sheffer operator of a sequence is equivalent to
knowing the sequence.
Continuous Operators on the Umbral Algebra
It is clearly desirable that a linear operator on the umbral algebra pass ; <
under infinite sums, that is, that
; ² ! ³ ~ ; ´ ² ! ³ µ45
~ ~BB
()19.6
whenever the sum on the left is defined, which is precisely when ² ²!³³ ¦ B
as . Not all operators on have this property, which leads to the ¦B <
following definition.
Definition A linear operator on the umbral algebra is if it ; < continuous
satisfies .²19.6)
The term continuous can be justifie d by defining a t opology on . However, <
since no additional topological concepts will be needed, we will not do so here.
Note that in order for 19.6 to make sense, we must have . It () ²;´ ²!³µ³ ¦ B
turns out that this condition is also sufficient.
Theorem 19.15 A linear operator on is continuous if and only if ;<
² ³ ¦ B ¬ ²;² ³³ ¦ B ()19.7
Proof. The necessity is clear. Suppose that 19.7 holds and that . () ² ³ ¦ B
For any , we have
Lc M Lc M Lc M; ² ! ³% ~ ; ² ! ³% b ; ² ! ³%
~ ~ B
()19.8
Since
² ! ³ ¦ B89
()19.7 implies that we may choose large enough that
; ² ! ³ 45
and
²;´ ²!³µ³ for
The Umbral Calculus 493
Hence, 19.8 gives ()
Lc M Lc M
Lc M
Lc M; ² ! ³ % ~ ; ² ! ³ %
~ ; ´ ² ! ³ µ %
~ ; ´ ² ! ³ µ %~ ~B
~
~B
which implies the desired result.
Operator Adjoints
If is a linear operator on , then its operator adjoint is anF F F ¢¦ ()d
operator on defined by F<i~
d´²!³µ ~ ²!³k
In the symbolism of th e umbral calculus, this is
º ²!³ ²%³» ~ º²!³ ²%³»d
(We have reduced the number of parentheses used to aid clarity. ³
Let us recall the basic properties of the adjoint from Chapter 3.
Theorem 19.16 For , BFÁ² ³
1)²b³~ b ddd
2 for any )² ³ ~ ddd
3)²³ ~ dd d
4 for any invertible )²³ ~ ² ³ ² ³ B Fc cdd
Thus, the map that sends to its adjoint BF B< F F < <¢²³ ¦²³ ¢ ¦ ¢ ¦d
is a linear transformation from to . Moreover, since implies BF B< ²³ ²³ ~ d
that for all and , which in turn impliesº ² ! ³ ² % ³ »~ ² ! ³ ² % ³< F
that , we deduce that is injective. The next theorem describes the range ~
of .
Theorem 19.17 A linear operator is the adjoint of a linear operator ; ² ³B<
BB F²³ ; if and only if is continuous.
Proof. First, suppose that for some and let . If ; ~ ² ³ ² ²!³³ ¦ B B Fd
, then for all we have
º ² ! ³ % » ~ º ² ! ³ % »d
and so it is only necessary to take large enough that for ² ²!³³ ²% ³ deg
all , whence
494 Advanced Linear Algebra
º ² ! ³ % » ~ d
for all and so . Thus, and is ² ²!³³ ² ²!³³ ¦ B dd d
continuous.
For the converse, assume that is continuous. If did have the form , then ;; d
º ; ! % » ~ º ! % » ~ º !% » d
and since
%~ %º! % »
[
we are prompted to by define
%~ %º;! % »
[
This makes sense since as and so the sum on the right is a ²;! ³ ¦ B ¦ B
finite sum. Then
º ! %»~º ! %»~ º ! %»~º ;! %»º;! % »
[d
which implies that for all . Finally, since and are both ;! ~ ! ;d d
continuous, we have . ;~d
Umbral Operators and Automorphisms of the Umbral Algebra
Figure 19.1 shows the map , which is an isomorphism from the vector space
BF <²³ onto the space of all continuous linear operators on . We are interested
in determining the images under this isomorphism of the set of umbral operators
and the set of umbral shifts, as pictured in Figure 19.1.
The Umbral Calculus 495
Figure 19.1
Let us begin with umbral operators. Suppose that is the umbral operator for
the associated sequence , with delta series . Then ² % ³ ² ! ³ <
º ²!³ % » ~ º²!³ % » ~ º²!³ ²%³» ~ [ ~ º! % » d
Á
for all and . Hence, and the continuity of implies that ² ! ³ ~ ! d d
d !~ ² ! ³
More generally, for any , ²!³ <
d²!³ ~ ²²!³³ ()19.9
In words, is composition by .d²!³
From 19.9 , we deduce that is a vector space isomorphism and that () d
dd d´²!³²!³µ ~ ²²!³³²²!³³ ~ ²!³ ²!³
Hence, is an automorphism of the umbral algebra . It is a pleasant fact that<d
this characterizes umbral operators. The firs t step in the proof of this is the
following, whose proof is left as an exercise.
Theorem 19.18 If is an automorphism of the umbral algebra, then ;;
preserves order, that is, . In particular, is continuous. ²;²!³³ ~ ²²!³³ ;
Theorem 19.19 A linear operator on is an umbral operator if and only if F
its adjoint is an automorphism of the umbral algebra . Moreover, if is an <
umbral operator, then
496 Advanced Linear Algebra
d²!³ ~ ²²!³³
for all . In particular, .²!³ ²!³ ~ !<d
Proof. We have already shown that the adjoint of is an automorphism
satisfying 19.9 . For the converse, suppose that is an automorphism of . () <d
Since is surjective, there is a unique series for which .d d²!³ ²!³~!
Moreover, Theorem 19.18 implies that is a delta series. Thus, ²!³
[ ~º ! %»~º ² ! ³ %»~º ² ! ³ %» Á d
which shows that is the associated sequence for and hence that is an % ² ! ³
umbral operator.
Theorem 19.19 allows us to fill in one of the boxes on the right side of Figure
19.1. Let us see how we might use Theorem 19.19 to advantage in the study of
associated sequences.
We have seen that the isomorphism maps the set of umbral operators Kªd
on onto the set of automorphisms of . But is a groupF< < F < aut aut²³ ~ ²³i
under composition. So if
¢% ¦ ²%³ ¢% ¦ ²%³ and
are umbral operators, then since
²k³ ~ k dd d
is an automorphism of , it follows that the composition is an umbral < k
operator. In fact, since
² k ³ ²²!³³ ~ k ²²!³³ ~ ²!³ ~ ! dd dd
we deduce that . Also, since k k~
k !k~ ~~
we have .c~
Thus, the set of umbral operators is a group under composition with K
k k~
and
c
~
Let us see how this plays out with respect to associated sequences. If the
The Umbral Calculus 497
associated sequence for is ²!³
² % ³~ % Á
~
then and so is the umbral operator for the k ¢% ¦ ²%³ ~ k
associated sequence
²k³ % ~ ² % ³ ~ % ~ ² % ³ Á Á
~ ~
This sequence, denoted by
² ²%³³ ~ ²%³ Á
~
()19.10
is called the of with . The umbral operator umbral composition ²%³ ²%³
'c Á ~ ² % ³ ~ % is the umbral operator for the associated sequence
where
c
%~ ² % ³
and so
%~ ² % ³
~
Á
Let us summarize.
Theorem 19.20
1 The set of umbral operators on is a group under composition, with)KF
k c
k~ ~ and
2 The set of associated sequences forms a group under umbral composition)
² ²%³³ ~ ²%³ Á
~
In particular, the umbral composition is the associated sequence ² ²%³³
for the composition , that is, k
k ¢% ¦ ² ²%³³
The identity is the sequence and the inverse of is the associated % ² % ³
sequence for the compositional inverse . ²!³
498 Advanced Linear Algebra
3 Let and . Then as operators,)K < ² ! ³
c²!³ ~ ²!³
4 Let and . Then)K < ² ! ³
²²!³³ ~ ²!³
Proof. We prove 3 as follows. For any and , ) ²!³ ²%³ 7<
º ² ! ³ ² ! ³ ² % ³ »~º ´ ² ! ³ µ ² ! ³ ² % ³ »
~ º ´²!³² ³ ²!³µ ²%³»
~ º²!³² ³ ²!³ ²%³»
~ º² ³ ²!³ ²!³ ²%³»
~ º²!³ ²!³ ²%³
dd
dc d
c d
c d
c»
which gives the desired result. Part 4 follows immediately from part 3 since ))
is composition by .
Sheffer Operators
If is Sheffer for , then the linear operator defined by ² % ³ ² Á ³ Á
Á ²% ³ ~ ²%³
is called a . Sheffer operators are closely related to umbral Sheffer operator
operators, since if is associated with , then ² % ³ ² ! ³
²%³ ~ ²!³ ²%³ ~ ²!³ %c c
and so
Á c~ ² ! ³
It follows that the Sheffer operators form a group with composition
Á Á c c
c c
c
k
h²k³Ákk ~ ²!³ ²!³
~ ²!³ ²²!³³
~ ´²!³²²!³³µ
~
and inverse
Ác
² ³ Á ~ c
From this, we deduce that the umbral composition of Sheffer sequences is a
Sheffer sequence. In particular, if is Sheffer for and ² % ³ ² Á ³
!² % ³~ ! % ² Á ³ Á ' is Sheffer for , then
The Umbral Calculus 499
Á Á Á Á
~
~
Á
k² % ³ ~! %
~! ² % ³
~!² ² % ³ ³
is Sheffer for . ²h² k³Á k³
Umbral Shifts and Derivations of the Umbral Algebra
We have seen that an operator on is an umbral operator if and only if itsF
adjoint is an automorphism of . Now suppose that is the umbral < B F ²³
shift for the associated sequence , associated with the delta series ² % ³
²!³<. Then
º ²!³ ²%³» ~ º²!³ ²%³»
~ º²!³ ²%³»
~² b [
~ ² c³[
~º ² ! ³ ² % ³ »
d
b
bÁ
Ác
c
1)
and so
d c ²!³ ~²!³ ()19.11
This implies that
d d d´²!³ ²!³ µ ~ ´²!³ µ²!³ b²!³ ´²!³ µ ()19.12
and further, by continuity, that
dd d´²!³²!³µ ~ ´ ²!³µ²!³b²!³´ ²!³µ ()19.13
Let us pause for a definition.
Definition Let be an algebra. A linear operator on is a if77 C derivation
C²³ ~ ²C³ bC b
for all .Á 7
Thus, we have shown that the adjoint of an umbral shift is a derivation of the
umbral algebra . Moreover, the expansion theorem and 19.11 show that < ()d
is surjective. This characterizes umbral sh ifts. First we need a preliminary result
on surjective derivations.
500 Advanced Linear Algebra
Theorem 19.21 Let be a surjective derivation on the umbral algebra . ThenC <
C ~ ²C ²!³³ ~ ²²!³³c ²²!³³ c for any and f , if . In constant <
particular, is continuous. C
Proof. We begin by noting that
C~C ~CbC~C
and so for all constants . Since is surjective, there mustC ~C~ C <
exist an for which²!³ <
C²!³~
Writing , we have²!³ ~ b! ²!³
~ C´ b! ²!³µ ~ ²C!³ ²!³b!C ²!³
which implies that . Fi nally, if , then , ²C!³ ~ ²²!³³ ~ ²!³ ~ ! ²!³
where and so² ²!³³ ~
´C²!³µ ~ ´C! ²!³µ ~ ´! C²!³b! ²!³C!µ ~ c c
Theorem 19.22 A linear operator on is an umbral shift if and only if its F
adjoint is a surjective derivation of th e umbral algebra . Moreover, if is an <
umbral shift, then is derivation with respect to , that is, d~C ² ! ³
c d²!³ ~²!³
for all . In particular, . ² ! ³~ d
Proof. We have already seen that is derivation with respect to . For the d²!³
converse, suppose that is a surjective derivation. Theorem 19.21 implies that d
there is a delta functional such that . If is the associated ²!³ ²!³~ ²%³ d
sequence for , then ²!³
º²!³ ²%³» ~ º ²!³ ²%³»
~ º²!³ ²!³ ²%³»
~º ² ! ³ ² % ³ »
~² b [
~ º²!³ ²%³»d
c d
c
bÁ
b
1)
Hence, , that is, is the umbral shift for . ² % ³~ ² % ³ ~ ² % ³ b
We have seen that the fact that the set of all automorphisms on is a group <
under composition shows that the set of all associated sequences is a group
under umbral composition. The set of all surjective derivations on does not <
form a group. However, we do have the chain rule for derivations [
The Umbral Calculus 501
Theorem 19.23 The chain rule () Let and be surjective derivations on .CC <
Then
C~ ² C ² ! ³ ³ C
Proof. This follows from
C ²!³ ~ ²!³ C ²!³ ~ ²C ²!³³C ²!³ c
and so continuity implies the result.
The chain rule leads to the following umbral result.
Theorem 19.24 If and are umbral shifts, then
~k C ² ! ³
Proof. Taking adjoints in the chain rule gives
d~ k²C ²!³³ ~ kC ²!³
We leave it as an exercise to show that . Now, by taking C ² ! ³~´ C ² ! ³ µc
²!³ ~ ! % ~ % in Theorem 19.24 and observing that and so is !! b
multiplication by , we get %
!c Z c~% C!~% ´ C² ! ³ µ ~% ´ ² ! ³ µ
Applying this to the associated sequence for gives the following ² % ³ ² ! ³
important recurrence relation for . ² % ³
Theorem 19.25 The recurrence formula () Let be the associated² % ³
sequence for . Then ²!³
1) ²%³ ~ %´ ²!³µ ²%³b Zc
2)² % ³ ~ % ´ ² ! ³ µ %b Z
Proof. The first part is proved. As to the second, using Theorem 19.20 we have
²%³ ~ %´ ²!³µ ²%³
~ %´ ²!³µ %
~ % ´ ²²!³³µ %
~ % ´²!³µ %b Zc
Zc
Zc
Z
Example 19.9 The recurrence relation can be used to find the associated
sequence for the forward difference functional . Since , ²!³ ~ c ²!³ ~ !Z !
the recurrence relation is
²%³ ~ % ²%³ ~ % ²%cb c!1)
502 Advanced Linear Algebra
Using the fact that , we have ² % ³~
²%³ ~ %Á ²%³ ~ %²%c³Á ²%³ ~ %²%c³²%c³ 3
and so on, leading easily to the lower factorial polynomials
²%³ ~ %²%c³Ä²%cb³ ~ ²%³
Example 19.10 Consider the delta functional
²!³ ~ ²b!³ log
Since is the forward difference functional, Theorem 19.20 implies²!³~ c!
that the associated sequence for is the inverse, under umbral ²%³ ²!³
composition, of the lower factorial polynomials. Thus, if we write
~
²%³ ~ :²Á³%
then
%~ : ² Á ³ ² % ³
~
The coefficients in this equation are known as the :²Á³ Stirling numbers of
the second kind and have great combinatorial significance. In fact, is :²Á³
the number of partitions of a set of size into blocks. The polynomials ² % ³
are called the . exponential polynomials
The recurrence relation for the exponential polynomials is
b Z²%³ ~ %²b!³ ²%³ ~ %² ²%³b ²%³³
Equating coefficients of on both side s of this gives the well-known formula %
for the Stirling numbers
:²bÁ³~:²Ác³b:²Á³
Many other properties of the Stirling numbers can be derived by umbral
means.
Now we have the analog of part 3 of Theorem 19.20. )
Theorem 19.26 Let be an umbral shift. Then
d
²!³ ~ ²!³ c ²!³
The Umbral Calculus 503
Proof. We have
º ² ! ³ ² ! ³ ² % ³ »~º ´ ² ! ³ µ ² ! ³² % ³ »
~ º ´²!³ ²!³µc²!³ ²!³ ²%³»
~ º ´²!³ ²!³µ ²%³»cº²!³ ²!³ ²%³»
~º ²d d
d d
d c
!³ ²!³ ²%³»cº ²!³ ²!³ ²%³»
~ º ²!³ ²!³ ²%³»cº ²!³ ²!³ ²%³»
~ º ²!³ ²!³ ²%³»cº ²!³ ²!³ ²%³»
c
d
from which the result follows.
If , then is multiplication by and is the derivative with respect to ²!³~! % d
! and so the previous result becomes
²!³ ~ ²!³%c%²!³Z
as operators on . The right side of this is called the of F Pincherle derivative
the operator . See [104]. ²!³ ()
Sheffer Shifts
Recall that the linear map
Á b´ ²%³µ ~ ²%³
where is Sheffer for is called a Sheffer shift. If is ²%³ ²²!³Á²!³³ ²%³
associated with , then and so ²!³ ²!³ ²%³ ~ ²%³
² ! ³ ² % ³ ~ ´ ² ! ³ ² % ³ µc c
b Á
and so
Á c~ ²!³ ²!³
From Theorem 19.26, the recurrence formula and the chain rule, we have
Á c
c d
c
c
!c
Zc c Zc Z~ ² ! ³ ² ! ³
~ ²!³´²!³ c ²!³µ
~c ² ! ³ C ² ! ³
~c ² ! ³ C ² ! ³
~c ² ! ³ C ! C ² ! ³
~ %´ ²!³µ c ²!³´ ²!³µ ²!³
~% c>?² ! ³
²!³ ²!³Z
Z
We have proved the following.
504 Advanced Linear Algebra
Theorem 19.27 Let be a Sheffer shift. ThenÁ
1)Á² ! ³
²!³ ²!³~% c<=Z
Z
2) ² % ³ ~ % c ² % ³b ² ! ³
²!³ ²!³<=Z
Z
The Transfer Formulas
We conclude with a pair of formulas for the computation of associated
sequences.
Theorem 19.28 The () transfer formulas Let be the associated sequence² % ³
for . Then²!³
1)² % ³~² ! ³ %Z²!³
!cc
45
2)² % ³~% %²!³
!cc45
Proof. First we show that 1 and 2 are equivalent. Write . Then )) ²!³ ~ ²!³°!
²!³²!³ % ~ ´!²!³µ ²!³ %
~ ²!³ % b! ²!³²!³ %
~ ²!³ % b ²!³²!³ %
~ ² ! ³ % b´ ² ! ³ µ%
~ ² ! ³ % c´ ² ! ³Z cc Z cc
c Z cc
c Z cc c
c c Z c
c c%c%²!³ µ%
~% ² ! ³ %c c
c c
To prove 1 , we verify the operation conditions for an associated sequence for )
the sequence . First, when the fourth equality ²%³ ~ ²!³²!³ % Zc c
above gives
º! ²%³» ~ º! ²!³²!³ % »
~º ! ² ! ³ % c´ ² ! ³ µ% »
~ º²!³ % »cº´²!³ µ % »
~ º²!³ % »cº²!³ % »
~ Z c c
c c Z c
c c Z c
c c
If , then , and so in general, we have as ~ º ! ² % ³ » ~ º ! ² % ³ » ~ Á
required.
For the second required condition,
²!³ ²%³ ~ ²!³ ²!³²!³ %
~ !²!³ ²!³²!³ %
~ ²!³²!³ %
~ ² % ³Zc c
Zc c
Zc c c
c
Thus, is the associated sequence for .² % ³ ² ! ³
The Umbral Calculus 505
A Final Remark
Unfortunately, space does not permit a detailed discussion of examples of
Sheffer sequences nor the application of the umbral calculus to various classical
problems. In [105], one can find a disc ussion of the following polynomial
sequences:
The lower factorial polynom ials and Stirling numbers
The exponential polynomials and Dobinski's formula
The Gould polynomials
The central factorial polynomials
The Abel polynomials
The Mittag-Leffler polynomials
The Bessel polynomials
The Bell polynomials
The Hermite polynomials
The Bernoulli polynomials and the Euler–MacLaurin expansion
The Euler polynomials
The Laguerre polynomials
The Bernoulli polynomials of the second kind
The Poisson–Charlier polynomials
The actuarial polynomials
The Meixner polynomials of the first and second kinds
The Pidduck polynomials
The Narumi polynomials
The Boole polynomials
The Peters polynomials
The squared Hermite polynomials
The Stirling polynomials
The Mahler polynomials
The Mott polynomials
and more. In [105], we also find a discussion of how the umbral calculus can be
used to approach the following types of problems:
The connection constants problem
Duplication formulas
The Lagrange inversion formula
Cross sequences
Steffensen sequences
Operational formulas
Inverse relations
Sheffer sequence solutions to recurrence relations
Binomial convolution
506 Advanced Linear Algebra
Finally, it is possible to generalize the classical umbral calculus that we have
described in this chapter to provide a context for studying polynomial sequences
such as those of the names Gegenbauer, Chebyshev and Jacobi. Also, there is a
q-version of the umbral calculus that involves the also q-binomial coefficients (
known as the Gaussian coefficients )
45² c ³ Ä ² c ³
²c³Ä²c ³²c³Ä²c ³~
c
in place of the binomial coefficients. There is also a logarithmic version of the
umbral calculus, which studies the and sequences of harmonic logarithms
logarithmic type . For more on these topics, please see [103], [106] and [107].
Exercises
1. Prove that , for any . ²³ ~ ²³b²³ Á <
2. Prove that min , for any . ² b³ ¸²³Á²³¹ Á <
3. Show that any delta series has a compositional inverse.
4. Show that for any delta series , the sequence is a pseudobasis.
5. Prove that is a derivation. C!
6. Show that is a delta functional if and only if and º »~<
º %» £ .
7. Show that is invertible if and only if . º »£<
8. Show that for any a , and º²!³ ²%³» ~ º²!³ ²%³» d<
F.
9. Show that e a for any polynomial . º! ²%³» ~ ² ³ ²%³ ! ZF
10. Show that in if and only if as linear functionals, which ~ ~<
holds if and only if as linear operators. ~
11. Prove that if is Sheffer for , then . ²%³ ²²!³Á²!³³ ²!³ ²%³ ~ ²%³ c
Hint: Apply the functionals to both sides. ²!³ ²!³
12. Verify that the Abel polynomials form the associated sequence for the Abel
functional.
13. Show that a sequence is the Appell sequence for if and only if ² % ³ ² ! ³
² % ³~ ² ! ³ %c .
14. If is a delta series, show that the adjoint of the umbral operator is a d
vector space isomorphism of . <
15. Prove that if is an automorphism of the umbral algebra, then preserves ;;
order, that is, . In particular, is continuous. ²;²!³³ ~ ²²!³³ ;
16. Show that an umbral operator maps associated sequences to associated
sequences.
17. Let and be associated sequences. Define a linear operator by ²%³ ²%³
¢ ²%³¦ ²%³ . Show that is an umbral operator.
18. Prove that if and are surjective derivations on , then CC <
C ² ! ³~´ C ² ! ³ µc.
References
General References
[1] Jacobson, N., , second edition, W.H. Freeman, 1985. Basic Algebra I
[2] Snapper, E. and Troyer, R., , Dover Publications, Metric Affine Geometry
1971.
General Linear Algebra
[3] Akivis, M., Goldberg, V., An Introduction To Linear Algebra and
Tensors , Dover, 1977.
[4] Blyth, T., Robertson, E., , Springer, 2002. Further Linear Algebra
[5] Brualdi, R., Friedland, S., Klee, V., Combinatorial and Graph-
Theoretical Problems in Linear Algebra , Springer, 1993.
[6] Curtis, M., Place, P., , Springer, 1990. Abstract Linear Algebra
[7] Fuhrmann, P., , Springer, 1996. A Polynomial Approach to Linear Algebra
[8] Gel'fand, I. M., , Dover, 1989. Lectures On Linear Algebra
[9] Greub, W., , Springer, 1995. Linear Algebra
[10] Halmos, P. R., , Mathematical Association Linear Algebra Problem Book
of America, 1995.
[11] Halmos, P. R., , Springer, 1974. Finite-Dimensional Vector Spaces
[12] Hamilton, A. G., , Cambridge University Press, 1990. Linear Algebra
[13] Jacobson, N., , Springer, Lectures in Abstract Algebra II: Linear Algebra
1953.
[14] Jänich, K., , Springer, 1994. Linear Algebra
[15] Kaplansky, I., , Dover, Linear Algebra and Geometry: A Second Course
2003.
[16] Kaye, R., Wilson, R., , Oxford University Press, 1998. Linear Algebra
[17] Kostrikin, A. and Manin, Y., , Gordon and Linear Algebra and Geometry
Breach Science Publishers, 1997.
[18] Lax, P., , John Wiley, 1996. Linear Algebra
[19] Lewis, J. G., Proceedings of the 5th SIAM Conference On Applied Linear
Algebra , SIAM, 1994.
508 Advanced Linear Algebra
[20] Marcus, M., Minc, H., , Dover, 1988. Introduction to Linear Algebra
[21] Mirsky, L., , Dover, 1990. An Introduction to Linear Algebra
[22] Nef, W., , Dover, 1988. Linear Algebra
[23] Nering, E. D., , John Wiley, 1976. Linear Algebra and Matrix Theory
[24] Pettofrezzo, A., , Dover, 1978. Matrices and Transformations
[25] Porter, G., Hill, D., Interactive Linear Algebra: A Laboratory Course
Using Mathcad , Springer, 1996.
[26] Schneider, H., Barker, G., , Dover, 1989. Matrices and Linear Algebra
[27] Schwartz, J., , Dover, 2001. Introduction to Matrices and Vectors
[28] Shapiro, H., A survey of canonical forms and invariants for unitary
similarity, 147:101–167 1991 . Linear Algebra and Its Applications ()
[29] Shilov, G., , Dover, 1977. Linear Algebra
[30] Wilkinson, J., , Oxford University The Algebraic Eigenvalue Problem
Press, 1988.
Matrix Theory
[31] Antosik, P., Swartz, C., , Springer, 1985. Matrix Methods in Analysis
[32] Bapat, R., Raghavan, T., , Nonnegative Matrices and Applications
Cambridge University Press, 1997.
[33] Barnett, S., , Oxford University Press, 1990. Matrices
[34] Bellman, R., , SIAM, 1997. Introduction to Matrix Analysis
[35] Berman, A., Plemmons, R., Non-negative Matrices in the Mathematical
Sciences , SIAM, 1994.
[36] Bhatia, R., , Springer, 1996. Matrix Analysis
[37] Bowers, J., , Oxford University Press, Matrices and Quadratic Forms
2000.
[38] Boyd, S., El Ghaoui, L., Feron, E.; Balakrishnan, V., Linear Matrix
Inequalities in System and Control Theory , SIAM, 1994.
[39] Chatelin, F., , John Wiley, 1993. Eigenvalues of Matrices
[40] Ghaoui, L., , Advances in Linear Matrix Inequality Methods in Control
SIAM, 1999.
[41] Coleman, T., Van Loan, C., , SIAM, Handbook for Matrix Computations
1988.
[42] Duff, I., Erisman, A., Reid, J., , Direct Methods for Sparse Matrices
Oxford University Press, 1989.
[43] Eves, H., , Dover, 1966. Elementary Matrix Theory
[44] Franklin, J., , Dover, 2000. Matrix Theory
[45] Gantmacher, F.R., , American Mathematical Society, Matrix Theory I
2000.
[46] Gantmacher, F.R., , American Mathematical Society, Matrix Theory II
2000.
[47] Gohberg, I., Lancaster, P., Rodman, L., Invariant Subspaces of Matrices
with Applications , John Wiley, 1986.
[48] Horn, R. and Johnson, C., , Cambridge University Press, Matrix Analysis
1985.
References 509
[49] Horn, R. and Johnson, C., , Cambridge Topics in Matrix Analysis
University Press, 1991.
[50] Jennings, A., McKeown, J. J., , John Wiley, 1992. Matrix Computation
[51] Joshi, A. W., , John Wiley, 1995. Matrices and Tensors in Physics
[52] Laub, A., , SIAM, 2004. Matrix Analysis for Scientists and Engineers
[53] Lütkepohl, H., , John Wiley, 1996. Handbook of Matrices
[54] Marcus, M., Minc, H., A Survey of Matrix Theory and Matrix
Inequalities , Dover, 1964.
[55] Meyer, C., , SIAM, 2000. Matrix Analysis and Applied Linear Algebra
[56] Muir, T., , Dover, 2003. A Treatise on the Theory of Determinants
[57] Perlis, S., , Dover, 1991. Theory of Matrices
[58] Serre, D., , Springer, 2002. Matrices: Theory and Applications
[59] Stewart, G., , SIAM, 1998. Matrix Algorithms
[60] Stewart, G., , SIAM, 2001. Matrix Algorithms Volume II: Eigensystems
[61] Watkins, D., , John Wiley, 1991. Fundamentals of Matrix Computations
Multilinear Algebra
[62] Marcus, M., , Marcel Finite Dimensional Multilinear Algebra, Part I
Dekker, 1971.
[63] Marcus, M., , Marcel Finite Dimensional Multilinear Algebra, Part II
Dekker, 1975.
[64] Merris, R., , Gordon & Breach, 1997. Multilinear Algebra
[65] Northcott, D. G., , Cambridge University Press, 1984. Multilinear Algebra
Applied and Numerical Linear Algebra
[66] Anderson, E., , SIAM, 1995. LAPACK User's Guide
[67] Axelsson, O., , Cambridge University Press, Iterative Solution Methods
1994.
[68] Bai, Z., Templates for the Solution of Al gebraic Eigenvalue Problems: A
Practical Guide , SIAM, 2000.
[69] Banchoff, T., Wermer, J., , Springer, Linear Algebra Through Geometry
1992.
[70] Blackford, L., , SIAM, 1997. ScaLAPACK User's Guide
[71] Ciarlet, P. G., Introduction to Numerical Linear Algebra and
Optimization , Cambridge University Press, 1989.
[72] Campbell, S., Meyer, C., Generalized Inverses of Linear
Transformations , Dover, 1991.
[73] Datta, B., Johnson, C., Kaashoek, M., Plemmons, R., Sontag, E., Linear
Algebra in Signals, Systems and Control , SIAM, 1988.
[74] Demmel, J., , SIAM, 1997. Applied Numerical Linear Algebra
[75] Dongarra, J., Bunch, J. R., Moler, C. B., Stewart, G. W., Linpack Users'
Guide , SIAM, 1979.
[76] Dongarra, J., Numerical Linear Algebra for High-Performance
Computers , SIAM, 1998.
510 Advanced Linear Algebra
[77] Dongarra, J., Templates for the Solution of Linear Systems: Building
Blocks For Iterative Methods , SIAM, 1993.
[78] Faddeeva, V. N., , Dover, Computational Methods of Linear Algebra
[79] Frazier, M., , An Introduction to Wavelets Through Linear Algebra
Springer, 1999.
[80] George, A., Gilbert, J., Liu, J., Graph Theory and Sparse Matrix
Computation , Springer, 1993.
[81] Golub, G., Van Dooren, P., Numerical Linear Algebra, Digital Signal
Processing and Parallel Algorithms , Springer, 1991.
[82] Granville S., , 2nd edition, Computational Methods of Linear Algebra
John Wiley, 2005.
[83] Greenbaum, A., , SIAM, Iterative Methods for Solving Linear Systems
1997.
[84] Gustafson, K., Rao, D., Numerical Range: The Field of Values of Linear
Operators and Matrices , Springer, 1996.
[85] Hackbusch, W., , Iterative Solution of Large Sparse Systems of Equations
Springer, 1993.
[86] Jacob, B., , Springer, 1995. Linear Functions and Matrix Theory
[87] Kuijper, M., , Birkhäuser, First-Order Representations of Linear Systems
1994.
[88] Meyer, C., Plemmons, R., Linear Algebra, Markov Chains, and Queueing
Models , Springer, 1993.
[89] Neumaier, A., , Cambridge Interval Methods for Systems of Equations
University Press, 1991.
[90] Nevanlinna, O., , Convergence of Iterations for Linear Equations
Birkhäuser, 1993.
[91] Olshevsky, V., Fast Algorithms for Structured Matrices: Theory and
Applications , SIAM, 2003.
[92] Plemmon, R.J., Gallivan, K.A., Sameh, A.H., Parallel Algorithms for
Matrix Computations , SIAM, 1990.
[93] Rao, K. N., , John Wiley, Linear Algebra and Group Theory for Physicists
1996.
[94] Reichel, L., Ruttan, A., Varga, R., , Walter de Numerical Linear Algebra
Gruyter, 1993.
[95] Saad, Y., , SIAM, 2003. Iterative Methods for Sparse Linear Systems
[96] Scharlau, W., , Springer, 1985. Quadratic and Hermitian Forms
[97] Snapper, E., Troyer, R., , Dover, Metric Affine Geometry
[98] Spedicato, E., Computer Algorithms for Sol ving Linear Algebraic
Equations , Springer, 1991.
[99] Trefethen, L., Bau, D., , SIAM, 1997. Numerical Linear Algebra
[100] Van Dooren, P., Wyman, B., , Linear Algebra for Control Theory
Springer, 1994.
[101] Vorst, H., , Cambridge Iterative Krylov Methods for Large Linear Systems
University Press, 2003.
[102] Young, D., , Dover, 2003. Iterative Solution of Large Linear Systems
References 511
The Umbral Calculus
[103] Loeb, D. and Rota, G.-C., Formal Power Series of Logarithmic Type,
Advances in Mathematics , Vol. 75, No. 1, May 1989 1–118. ()
[104] Pincherle, S. "Operatori linear i e coefficienti di fattoriali." Alti Accad.
Naz. Lincei, Rend. Cl. Fis. Mat. Nat . 6 18, 417–519, 1933.()
[105] Roman, S., , Pure and Applied Mathematics vol. The Umbral Calculus
111, Academic Press, 1984.
[106] Roman, S., The logarithmic binomial formula, American Mathematical
Monthly 99 1992 641–648.()
[107] Roman, S., The harmonic logarithms and the binomial formula, Journal
of Combinatorial Theory , series A, 63 1992 143–163. ()
Index of Symbols
*´²%³µ ²%³ : the companion matrix of
² % ³: characteristic polynomial of
crk²(³ (: column rank of
cs²(³ (: column space of
diag²( ÁÃÁ( ³ ( : a block diagonal matrix with 's on the block diagonal
ElemDiv ²³: the multiset of elementary divisors
InvFact²³: the multiset of invariant factors of
@²Á ³ Á : Jordan block
² % ³: minimal polynomial of
null²³: the nullity of
:: canonical projection modulo :
9 =i: Riesz vector for
rk²³: the rank of
rrk²(³ (: row rank of
rs²(³ (: row space of
(Á): projection onto along ()
(: the multiplication by operator (
supp²³: the support of a function
= - -´%µ ²%³# ~ ² ³#: the -vector space/ -module where
==d: the complexification of
"º:» " º:»: assignment, for example, means that stands for
: subspace or submodule
: proper subspace or proper submodule
º:» :: subspace/ideal spanned by
ºº: »» : : submodule spanned by
Æ¢an embedding that is an isomorphism when all is finite-dimensional.
: similarity of matrices or operators, associate in a ring.
d: cartesian product
p: orthogonal direct sum
`: external direct product
^: external direct sum
l: internal direct sum
%& º%Á&»~º&Á%»: means
514
w: wedge product
n: tensor product
n: -fold tensor product
d: -fold cartesion product
²Á³ ~ : and are relatively prime
: affine combinationIndex of Symbols
Index
Abel functional, 477, 489
Abel operator, 479
Abel polynomials, 489
abelian, 17
absolutely convergent, 330
accumulation point, 306
adjoint, 227, 231
affine basis, 435
affine closed, 428
affine combination, 428
affine geometry, 427
affine group, 436
affine hull, 430
affine hyperplane, 416
affine map, 435
affine span, 430
affine subspace, 57
affine transformation, 435
affine, 424
affinely independent, 433
affinity, 435
algebra homomorphism, 455
algebra, 31, 451
algebraic, 100, 458
algebraic closure, 30
algebraic dual space, 94
algebraic multiplicity, 189
algebraic numbers, 460
algebraically closed, 30
algebraically reflexive, 101
algorithm, 217
almost upper triangular, 194
along, 73
alternate, 260, 262, 391
alternating, 260, 391
ancestor, 14
anisotropic, 265
annihilator, 102, 115, 459
antisymmetric, 259, 390, 395antisymmetric tensor algebra, 398
antisymmetric tensor space, 395, 400
antisymmetry, 10
Apollonius identity, 223
Appell sequence, 481
approximation problem, 331
as measured by, 357
ascending chain condition, 26, 133
associate classes, 27
associated sequence, 481
associates, 26
automorphism, 60
barycentric coordinates, 435
base ring, 110
base, 427
basis, 47, 116
Bernoulli numbers, 477
Bernstein theorem, 13
Bessel's identity, 221
Bessel's inequality, 220, 337, 338, 345
best approximation, 219, 332
bijection, 6
bijective, 6
bilinear form, 259, 360
bilinear, 206, 360
binomial identity, 486
binomial type, 486
block diagonal matrix, 3
block matrix, 3
blocks, 7
bottom, 10
bounded, 321, 349
canonical form, 8
canonical injections, 359
canonical map, 100
canonical projection, 89
Cantor's theorem, 13
516 Index
cardinal number, 13
cardinality, 12, 13
cartesian product, 14
Cauchy sequence, 311
Cauchy–Schwarz inequality, 208, 303, 325
Cayley-Hamilton theorem, 170
center, 452
central, 452
centralizer, 464, 469
chain, 11
chain rule, 501
change of basis matrix, 65
change of basis operator, 65
change of coordinates operator, 65
characteristic, 30
characteristic equation, 186
characteristic polynomial, 170
characteristic value, 185
characteristic vector, 186
Cholsky decomposition, 255
circulant matrices, 457
class equation, 464
classification problem, 276
closed ball, 304
closed half-spaces, 417
closed interval, 143
closed, 304, 414
closure, 306
codimension, 93
coefficients, 36
column equivalent, 9
column rank, 52
column space, 52
common eigenvector, 202
commutative, 17, 19, 451
commutativity, 15, 35, 384
commuting family, 201
compact, 414
companion matrix, 173
complement, 42, 120
complemented, 120
complete, 40, 311
complete invariant, 8
complete system of invariants, 8
completion, 316
complex operator, 59complex vector space, 36
complexification, 53, 54, 82
complexification map, 54
composition, 472
cone, 265, 414
congruence classes, 262
congruence relation, 88
congruent modulo, 21, 87
congruent, 9, 262
conjugacy class, 463
conjugate isomorphism, 222
conjugate linear, 206, 221
conjugate linearity, 206
conjugate representation, 483
conjugate space, 350
conjugate symmetry, 205
connected, 281
continuity, 340
continuous, 310, 492
continuous dual space, 350
continuum, 16
contraction, 389
contravariant tensors, 386
contravariant type, 386
converge, 339, 210, 305, 330
convex combination, 414
convex, 332, 414
convex hull, 415
coordinate map, 51
coordinate matrix, 52, 368
correspondence theorem, 90, 118
coset, 22, 87, 118
coset representative, 22, 87
countable, 13
countably infinite, 13
covariant tensors, 386
covariant type, 386
cycle, 391
cyclic basis, 166
cyclic decomposition, 149, 168
cyclic group generated by, 18
cyclic group of order, 18
cyclic submodule, 113
cyclotomic polynomial, 465
decomposable, 362
Index 517
degenerate, 266
degree, 5
deleted absolute row sum, 203
delta functional, 475
delta operator, 478
delta series, 472
dense, 308
derivation, 499
descendants, 13
determinant, 292, 405
diagonal, 4
diagonalizable, 196
diagonally dominant, 203
diameter, 321
dimension, 50, 427
direct product, 41, 408
direct sum, 41, 73, 119
direct summand, 42, 120
discrete metric, 302
discriminant, 263
distance, 209, 322
divides, 5, 26
division algebra, 462
division algorithm, 5
domain, 6
dot product, 206
double, 100
dual basis, 96
dual space, 59, 100
eigenspace, 186
eigenvalue, 185, 186, 461
eigenvector, 186
elementary divisor basis, 169
elementary divisor form, 176
elementary divisor version, 177
elementary divisors, 155, 167, 168
elementary divisors and dimensions, 168
elementary matrix, 3
elementary symmetric functions, 189
embedding, 59, 117
endomorphism, 59, 117
epimorphism, 59, 117
equivalence class, 7
equivalence relation, 7
equivalent, 9, 69essentially unique, 45
Euclidean metric, 302
Euclidean space, 206
evaluation at, 96, 100
evaluation functional, 474, 476
even permutation, 391
even weight subspace, 38
exponential polynomials, 502
exponential, 482
extension by, 103
extension, 6, 273
exterior algebra, 398
exterior product, 393
exterior product space, 395, 400
external direct sum, 40, 41, 119
factored through, 355, 357
factorization, 217
faithful, 457
Farkas's lemma, 423
field of quotients, 24
field, 19, 29
finite support, 41
finite, 1, 12, 18
finite-dimensional, 50, 451
finitely generated, 113
first isomorphism theorem, 92, 118, 469
flat representative, 427
flat, 427
form, 299, 382
forward difference functional, 477
forward difference operator, 479
Fourier coefficient, 219
Fourier expansion, 219, 338, 345
free, 116
Frobenius norm, 450, 466
functional calculus, 248
functional, 94
Gaussian coefficients, 57, 506
generating function, 482, 483
geometric multiplicity, 189
Geršgorin region, 203
Geršgorin row disk, 203
Geršgorin row region, 203
graded algebra, 392
518 Index
Gram-Schmidt augmentation, 213
Gram-Schmidt orthogonalization process, 214
greatest common divisor, 5
greatest lower bound, 11
group algebra, 453
group, 17
Hamel basis, 218
Hamming distance function, 321
Hermite polynomials, 224, 489
Hermitian, 238
Hilbert basis theorem, 136
Hilbert basis, 218, 335
Hilbert dimension, 347
Hilbert space adjoint, 230
Hilbert space, 315, 327
Hölder's inequality, 303
homogeneous, 392
homomorphism, 59, 117
Householder transformation, 244
hyperbolic basis, 273
hyperbolic extension, 274
hyperbolic pair, 272
hyperbolic plane, 272
hyperbolic space, 272
hyperplane, 416, 427
ideal generated by, 21, 455
ideal, 20, 455
idempotent, 74, 125
identity, 17
image, 6, 61
imaginary part, 54
indecomposable, 158
index of nilpotence, 200
induced, 305
inertia, 288
infinite, 13
infinite-dimensional, 50
injection, 6, 117
inner product, 205, 260
inner product space, 205, 260
integral domain, 23
invariant, 8, 73, 83, 165
invariant factor basis, 179
invariant factor decomposition, 157invariant factor form, 178
invariant factor version, 179
invariant factors, 157, 167, 168
invariant factor decomposition theorem, 157
invariant ideals, 157
invariant under, 73
inverses, 17
invertible functional, 475
involution, 199
irreducible, 5, 26, 83
isometric isomorphism, 211, 326
isometric, 271, 316
isometrically isomorphic, 211, 326
isometry, 211, 271, 315, 326
isomorphic, 59, 62, 117
isotropic, 265
join, 40
Jordan basis, 191
Jordan block, 191
Jordan canonical form, 191
kernel, 61
Kronecker delta function, 96
Kronecker product, 408
Lagrange interpolation formula, 248
Laguerre polynomials, 490
largest, 10
lattice, 39, 40
leading coefficient, 5
leading entry, 3
least, 10
least squares solution, 448
least upper bound, 11
left inverse, 122, 470
left regular matrix representation, 457
left regular representation, 457
left singular vectors, 445
left zero divisor, 460
left-invertible, 470
Legendre polynomials, 215
length, 208
limit, 306
limit point, 306
line, 427, 429
Index 519
linear code, 38
linear combination, 36, 112
linear function, 382
linear functional, 59, 94
linear hyperplane, 416
linear least squares, 448
linear operator, 59
linear transformation, 59
linearity, 340
linearity in the first coordinate, 205
linearly dependent, 45, 114
linearly independent, 45, 114
linearly ordered set, 11
lower bound, 11
lower factorial numbers, 471
lower factorial polynomials, 488
lower triangular, 4
main diagonal, 2
matrix, 64
matrix of, 66
matrix of the form, 261
maximal element, 10
maximal ideal, 23
maximal orthonormal set, 218
maximum, 10
measuring family, 357
measuring functions, 357
mediating morphism map, 367
mediating morphism, 357, 362, 383
meet, 40
metric, 210, 301
metric space, 210, 301
metric vector space, 260
mimimum, 10
minimal element, 11
minimal polynomial, 165, 166, 459
Minkowski space, 260
Minkowski's inequality, 37, 303
mixed tensors, 386
modular law, 56
module, 109, 133, 167
modulo, 22, 87, 118
monic, 5
monomorphism, 59, 117
Moore-Penrose generalized inverse, 446Moore-Penrose pseudoinverse, 446
MP inverse, 447
multilinear, 382
multilinear form, 382
multiplicity, 1
multiset, 1
natural map, 100
natural projection, 89
natural topology, 80, 82
negative, 17
net definition, 339
nilpotent, 198, 200
Noetherian, 133
nondegenerate, 266
nonderogatory, 171
nonisotropic, 265
nonnegative orthant, 225, 411
nonnegative, 225, 411
nonsingular, 266
nonsingular completion, 273
nonsingular extension theorem, 274
nontrivial, 36
norm, 208, 209, 303, 349
normal equations, 449
normal, 234
normalizing, 213
normed linear space, 209, 224
null, 265
nullity, 61
odd permutation, 391
one-sided inverses, 122, 470
one-to-one, 6
onto, 6
open ball, 304
open half-spaces, 417
open neighborhood, 304
open rectangles, 79
open sets, 305
operator adjoint, 104
operator characterization, 484
order, 18, 101, 139, 471
order ideals, 115
ordered basis, 51
order-reversing, 102
520 Index
orthogonal complement, 212, 265
orthogonal direct sum, 212, 269
orthogonal geometry, 260
orthogonal group, 271
orthogonal resolution of the identity, 232
orthogonal set, 212
orthogonal similarity classes, 242
orthogonal spectral resolution, 237
orthogonal transformation, 271
orthogonal, 75, 212, 231, 238, 265
orthogonality conditions, 480
orthogonally diagonalizable, 233
orthogonally equivalent, 242
orthogonally similar, 242
orthonormal basis, 218
orthonormal set, 212
parallel, 427
parallelogram law, 208, 325
parity, 391
Parseval's identity, 221, 346
partial order, 10
partially ordered set, 10
partition, 7
permutation, 391
Pincherle derivative, 503
plane, 427
point, 427
polar decomposition, 253
polarization identities, 209
posets, 10
positive definite, 205, 250, 301
positive square root, 251
power of the continuum, 16
power set, 13
primary, 147
primary cyclic decomposition theorem, 153, 168
primary decomposition theorem, 147
primary decomposition, 147, 168
prime subfield, 97
prime, 26
primitive, 465
principal ideal domain, 24
principal ideal, 24
product, 15
projection modulo, 89projection theorem, 220, 334
projection, 73
projective dimension, 438
projective geometry, 438
projective line, 438
projective plane, 438
projective point, 438
proper subspace, 37
properly divides, 27
pseudobasis, 472
pure in, 161
q-binomial coefficients, 506
quadratic form, 239, 264
quaternions, 463
quotient algebra, 455
quotient field, 24
quotient module, 118
quotient ring, 22
quotient space, 87, 89
radical, 266
range, 6
rank, 53, 61, 129, 369
rank plus nullity theorem, 63
rational canonical form, 176–179
real operator, 59
real part, 54
real vector space, 36
real version, 53
recurrence formula, 501
reduce, 169
reduced row echelon form, 3, 4
reflection, 244, 292
reflexivity, 7, 10
relatively prime, 5, 27
representation, 457
resolution of the identity, 76
restriction, 6
retract, 122
retraction map, 122
Riesz map, 222
Riesz representation theorem, 222, 268, 351
Riesz vector, 222
right inverse, 122, 470
right singular vectors, 445
Index 521
right zero divisor, 460
right-invertible, 470
ring, 18
ring homomorphism, 19
ring with identity, 19
roots of unity, 464
rotation, 292
row equivalent, 4
row rank, 52
row space, 52
scalar multiplication, 31, 35, 451
scalars, 2, 35, 109
Schröder, 13
Schur's theorem, 192, 195
second isomorphism theorem, 93, 119
self-adjoint, 238
separable, 308
sesquilinear, 206
Sheffer for, 481
Sheffer identity, 486
Sheffer operator, 491, 498
Sheffer sequence, 481
Sheffer shift, 491
sign, 391
signature, 288
similar, 9, 70, 71
similarity classes, 70, 71
simple, 138, 455
simultaneously diagonalizable, 202
singular, 266
singular values, 444, 445
singular-value decomposition, 445
skew self-adjoint, 238
skew-Hermitian, 238
skew-symmetric, 2, 238, 259, 390
smallest, 10
span, 45, 112
spectral mapping theorem, 187, 461
spectral theorem for normal operators, 236, 237
spectral resolution, 197
spectrum, 186, 461
sphere, 304
split, 5
square summable, 207
square summable functions, 347standard basis, 47, 62, 131
standard inner product, 206
standard topology, 79
standard vector, 47
Stirling numbers of the second kind, 502
strictly diagonally dominant, 203
strictly positive orthant, 411
strictly positive, 225, 411
strictly separated, 417
strongly positive orthant, 225, 411
strongly positive, 56, 225, 411
strongly separated, 417
structure constants, 453
structure theorem for normal matrices, 247
structure theorem for normal operators, 245
subalgebra, 454
subfield, 57
subgroup, 18
submatrix, 2
submodule, 111
subring, 19
subspace spanned, 44
subspace, 37, 260, 304
sup metric, 302
support, 6, 41
surjection, 6
surjective, 6
Sylvester's law of inertia, 287
symmetric, 2, 238, 259, 390, 395
symmetric group, 391
symmetric tensor algebra, 398
symmetric tensor space, 395, 400
symmetrization map, 402
symplectic basis, 273
symplectic geometry, 260
symplectic group, 271
symplectic transformation, 271
symplectic transvection, 280
tensor algebra, 390
tensor map, 362, 383
tensor product, 362, 383, 408
tensors of type, 386
tensors, 362
theorem of the alternative, 413
third isomorphism theorem, 94, 119
522 Index
top, 10
topological space, 305
topological vector space, 79
topology, 305
torsion element, 115
torsion module, 115
torsion-free, 115
total subset, 336
totally degenerate, 266
totally isotropic, 265
totally ordered set, 11
totally singular, 266
trace, 188
transfer formulas, 504
transitivity, 7, 10
translate, 427
translation operator, 479
translation, 436
transpose, 2
transposition, 391
triangle inequality, 208, 210, 301, 325
trivial, 36
two-affine closed, 428
two-sided inverse, 122, 470
umbral algebra, 474
umbral composition, 497
umbral operator, 491
umbral shift, 491
uncountable, 13
underlying set, 1
unipotent, 300
unique factorization domain, 28
unit vector, 208
unit, 26
unital algebras, 451
unitarily diagonalizable, 233
unitarily equivalent, 242
unitarily similar, 242
unitarily upper triangularizable, 196
unitary, 238
unitary metric, 302
unitary similarity classes, 242
unitary space, 206
universal, 289
universal for bilinearity, 362universal for multilinearity, 382
universal pair, 357
universal property, 357
upper bound, 11
upper triangular, 4
upper triangularizable, 192
Vandermonde convolution formula, 489
Vector Space, 167
vector space, 35
vectors, 35
Wedderburn's Theorem, 465, 466
wedge product, 393
weight, 38
well ordering, 12
Well-ordering principle, 12
with respect to the bases, 66
Witt index, 296
Witt's cancellation theorem, 279, 294
Witt's extension theorem, 279, 295
zero divisor, 23
zero element, 17
zero subspace, 37
Zorn's lemma, 12
Graduate Texts in Mathematics
(continued from page ii)
76 I ITAKA . Algebraic Geometry.
77 H ECKE. Lectures on the Theory of Algebraic
Numbers.
78 B URRIS /SANKAPPANA V AR . A Course in
Universal Algebra.
79 W ALTERS . An Introduction to Ergodic
Theory.
80 R OBINSON . A Course in the Theory of
Groups. 2nd ed.
81 F ORSTER . Lectures on Riemann Surfaces.
82 B OTT/TU. Differential Forms in Algebraic
Topology.
83 W ASHINGTON . Introduction to Cyclotomic
Fields. 2nd ed.
84 I RELAND /ROSEN. A Classical Introduction to
Modern Number Theory. 2nd ed.
85 E DW ARDS . Fourier Series. V ol. II. 2nd ed.
86 VA NLINT. Introduction to Coding Theory.
2nd ed.
87 B ROWN . Cohomology of Groups.
88 P IERCE . Associative Algebras.
89 L ANG. Introduction to Algebraic and
Abelian Functions. 2nd ed.
90 B RØNDSTED . An Introduction to Convex
Polytopes.
91 B EARDON . On the Geometry of Discrete
Groups.
92 D IESTEL . Sequences and Series in Banach
Spaces.
93 D UBROVIN /FOMENKO /NOVIKOV . Modern
Geometry—Methods and Applications.
Part I. 2nd ed.
94 W ARNER . Foundations of Differentiable
Manifolds and Lie Groups.
95 S HIRY AEV . Probability. 2nd ed.
96 C ONW AY . A Course in Functional Analysis.
2nd ed.
97 K OBLITZ . Introduction to Elliptic Curves and
Modular Forms. 2nd ed.
98 B RÖCKER /TOMDIECK. Representations of
Compact Lie Groups.
99 G ROVE/BENSON . Finite Reflection Groups.
2nd ed.
100 B ERG/CHRISTENSEN /RESSEL . Harmonic
Analysis on Semigroups: Theory of
Positive Definite and Related Functions.
101 E DW ARDS . Galois Theory.
102 V ARADARAJAN . Lie Groups, Lie Algebras
and Their Representations.
103 L ANG. Complex Analysis. 3rd ed.
104 D UBROVIN /FOMENKO /NOVIKOV . Modern
Geometry—Methods and Applications.
Part II.
105 L ANG.SL2(R).
106 S IL VERMAN . The Arithmetic of Elliptic
Curves.107 O LV ER. Applications of Lie Groups to
Differential Equations. 2nd ed.
108 R ANGE . Holomorphic Functions and
Integral Representations in Several
Complex Variables.
109 L EHTO. Univalent Functions and
Teichmüller Spaces.
110 L ANG. Algebraic Number Theory.
111 H USEMÖLLER . Elliptic Curves. 2nd ed.
112 L ANG. Elliptic Functions.
113 K ARATZAS /SHREVE . Brownian Motion and
Stochastic Calculus. 2nd ed.
114 K OBLITZ . A Course in Number Theory and
Cryptography. 2nd ed.
115 B ERGER /GOSTIAUX . Differential Geometry:
Manifolds, Curves, and Surfaces.
116 K ELLEY /SRINIVASAN . Measure and Integral.
Vo l . I .
117 J.-P. S ERRE. Algebraic Groups and Class
Fields.
118 P EDERSEN .A n a l y s i sN o w .
119 R OTMAN . An Introduction to Algebraic
Topology.
120 Z IEMER . Weakly Differentiable Functions:
Sobolev Spaces and Functions of Bounded
Variation.
121 L ANG. Cyclotomic Fields I and II.
Combined 2nd ed.
122 R EMMERT . Theory of Complex Functions.
Readings in Mathematics
123 E BBINGHAUS /HERMES et al. Numbers.
Readings in Mathematics
124 D UBROVIN /FOMENKO /NOVIKOV . Modern
Geometry—Methods and Applications
Part III.
125 B ERENSTEIN /GAY. Complex Variables: An
Introduction.
126 B OREL. Linear Algebraic Groups. 2nd ed.
127 M ASSEY . A Basic Course in Algebraic
Topology.
128 R AUCH . Partial Differential Equations.
129 F ULTON /HARRIS . Representation Theory: A
First Course. Readings in Mathematics
130 D ODSON /POSTON . Tensor Geometry.
131 L AM. A First Course in Noncommutative
Rings. 2nd ed.
132 B EARDON . Iteration of Rational Functions.
133 H ARRIS . Algebraic Geometry: A First
Course.
134 R OMAN . Coding and Information Theory.
135 R OMAN . Advanced Linear Algebra. 3rd ed.
136 A DKINS /WEINTRAUB . Algebra: An Approach
via Module Theory.
137 A XLER/BOURDON /RAMEY . Harmonic
Function Theory. 2nd ed.
138 C OHEN . A Course in Computational
Algebraic Number Theory.
139 B REDON . Topology and Geometry.
140 A UBIN. Optima and Equilibria. An
Introduction to Nonlinear Analysis.
141 B ECKER /WEISPFENNING /KREDEL . Gröbner
Bases. A Computational Approach to
Commutative Algebra.
142 L ANG. Real and Functional Analysis. 3rd ed.
143 D OOB. Measure Theory.
144 D ENNIS /FARB. Noncommutative Algebra.
145 V ICK. Homology Theory. An Introduction
to Algebraic Topology. 2nd ed.
146 B RIDGES . Computability: A Mathematical
Sketchbook.
147 R OSENBERG . Algebraic K-Theory and Its
Applications.
148 R OTMAN . An Introduction to the Theory of
Groups. 4th ed.
149 R ATCLIFFE . Foundations of Hyperbolic
Manifolds. 2nd ed.
150 E ISENBUD . Commutative Algebra with a
View Toward Algebraic Geometry.
151 S IL VERMAN . Advanced Topics in the
Arithmetic of Elliptic Curves.
152 Z IEGLER . Lectures on Polytopes.
153 F ULTON . Algebraic Topology: A First
Course.
154 B ROWN /PEARCY . An Introduction to
Analysis.
155 K ASSEL . Quantum Groups.
156 K ECHRIS . Classical Descriptive Set Theory.
157 M ALLIAVIN . Integration and Probability.
158 R OMAN . Field Theory.
159 C ONW AY . Functions of One Complex
Variable II.
160 L ANG. Differential and Riemannian
Manifolds.
161 B ORWEIN /ERDÉLYI . Polynomials and
Polynomial Inequalities.
162 A LPERIN /BELL. Groups and Representations.
163 D IXON/MORTIMER . Permutation Groups.
164 N ATHANSON . Additive Number Theory: The
Classical Bases.
165 N ATHANSON . Additive Number Theory:
Inverse Problems and the Geometry of
Sumsets.
166 S HARPE . Differential Geometry: Cartan’s
Generalization of Klein’s Erlangen
Program.
167 M ORANDI . Field and Galois Theory.
168 E WA LD . Combinatorial Convexity and
Algebraic Geometry.
169 B HATIA . Matrix Analysis.
170 B REDON . Sheaf Theory. 2nd ed.
171 P ETERSEN . Riemannian Geometry. 2nd ed.
172 R EMMERT . Classical Topics in Complex
Function Theory.
173 D IESTEL . Graph Theory. 2nd ed.
174 B RIDGES . Foundations of Real and Abstract
Analysis.175 L ICKORISH . An Introduction to Knot Theory.
176 L EE. Riemannian Manifolds.
177 N EWMAN . Analytic Number Theory.
178 C LARKE /LEDYAEV /STERN/WOLENSKI .
Nonsmooth Analysis and Control Theory.
179 D OUGLAS . Banach Algebra Techniques in
Operator Theory. 2nd ed.
180 S RIVA STAVA . A Course on Borel Sets.
181 K RESS. Numerical Analysis.
182 W ALTER . Ordinary Differential Equations.
183 M EGGINSON . An Introduction to Banach
Space Theory.
184 B OLLOBAS . Modern Graph Theory.
185 C OX/LITTLE /O’S HEA. Using Algebraic
Geometry. 2nd ed.
186 R AMAKRISHNAN /VALENZA . Fourier Analysis
on Number Fields.
187 H ARRIS /MORRISON . Moduli of Curves.
188 G OLDBLATT . Lectures on the Hyperreals: An
Introduction to Nonstandard Analysis.
189 L AM. Lectures on Modules and Rings.
190 E SMONDE /MURTY. Problems in Algebraic
Number Theory. 2nd ed.
191 L ANG. Fundamentals of Differential
Geometry.
192 H IRSCH /LACOMBE . Elements of Functional
Analysis.
193 C OHEN . Advanced Topics in Computational
Number Theory.
194 E NGEL/NAGEL. One-Parameter Semigroups
for Linear Evolution Equations.
195 N ATHANSON . Elementary Methods in
Number Theory.
196 O SBORNE . Basic Homological Algebra.
197 E ISENBUD /HARRIS . The Geometry of
Schemes.
198 R OBERT . A Course in p-adic Analysis.
199 H EDENMALM /KORENBLUM /ZHU. Theory of
Bergman Spaces.
200 B AO/CHERN /SHEN. An Introduction to
Riemann–Finsler Geometry.
201 H INDRY /SIL VERMAN . Diophantine Geometry:
An Introduction.
202 L EE. Introduction to Topological Manifolds.
203 S AGAN . The Symmetric Group:
Representations, Combinatorial
Algorithms, and Symmetric Functions.
204 E SCOFIER . Galois Theory.
205 F ÉLIX/HALPERIN /THOMAS . Rational
HomotopyTheory.3rd ed.
206 M URTY. Problems in Analytic Number
Theory. Readings in Mathematics
207 G ODSIL /ROYLE. Algebraic Graph Theory.
208 C HENEY . Analysis for Applied Mathematics.
209 A RVESON . A Short Course on Spectral
Theory.
210 R OSEN. Number Theory in Function Fields.
211 L ANG. Algebra. Revised 3rd ed.
212 M ATOUŠEK . Lectures on Discrete Geometry.
213 F RITZSCHE /GRAUERT . From Holomorphic
Functions to Complex Manifolds.
214 J OST. Partial Differential Equations. 2nd ed.
215 G OLDSCHMIDT . Algebraic Functions and
Projective Curves.
216 D. S ERRE. Matrices: Theory and
Applications.
217 M ARKER . Model Theory: An Introduction.
218 L EE. Introduction to Smooth Manifolds.
219 M ACLACHLAN /REID. The Arithmetic of
Hyperbolic 3-Manifolds.
220 N ESTRUEV . Smooth Manifolds and
Observables.
221 G RÜNBAUM . Convex Polytopes. 2nd ed.
222 H ALL. Lie Groups, Lie Algebras, and
Representations: An Elementary
Introduction.
223 V RETBLAD . Fourier Analysis and Its
Applications.
224 W ALSCHAP . Metric Structures in Differential
Geometry.
225 B UMP. Lie Groups.
226 Z HU. Spaces of Holomorphic Functions in
the Unit Ball.
227 M ILLER /STURMFELS . Combinatorial
Commutative Algebra.
228 D IAMOND /SHURMAN . A First Course in
Modular Forms.
229 E ISENBUD . The Geometry of Syzygies.230 S TROOCK . An Introduction to Markov
Processes.
231 B JÖRNER /BRENTI . Combinatorics of Coxeter
Groups.
232 E VEREST /WARD. An Introduction to Number
Theory.
233 A LBIAC /KALTON . Topics in Banach Space
Theory.
234 J ORGENSON . Analysis and Probability.
235 S EPANSKI . Compact Lie Groups.
236 G ARNETT . Bounded Analytic Functions.
237 M ARTÍNEZ -AVENDAÑO /ROSENTHAL .A n
Introduction to Operators on the
Hardy-Hilbert Space.
238 A IGNER , A Course in Enumeration.
239 C OHEN , Number Theory, V ol. I.
240 C OHEN , Number Theory, V ol. II.
241 S IL VERMAN . The Arithmetic of Dynamical
Systems.
242 G RILLET . Abstract Algebra. 2nd ed.
243 G EOGHEGAN . Topological Methods in Group
Theory.
244 B ONDY /MURTY. Graph Theory.
245 G ILMAN /KRA/RODRIGUEZ . Complex
Analysis.
246 K ANIUTH . A Course in Commutative Banach
Algebras.