beezer linear algebra
PDF · 1033 pages · 3.6 MB
Open PDF file
Complete textbook by Robert A. Beezer of the University of Puget Sound, Version 2.30, dated December 23, 2011, released under the GNU Free Documentation License. The table of contents shows chapters on systems of linear equations, row reduction, vectors, orthogonality, matrices and matrix inverses, with reading questions, exercises and solutions in each section. This is a downloaded reference book by another author, kept in Phil's math book collection.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
A First Course in Linear Algebra
A First Course in Linear Algebra
by
Robert A. Beezer
Department of Mathematics and Computer Science
University of Puget Sound
Version 2.30
Robert A. Beezer is a Professor of Mathematics at the University of Puget Sound, where he has been
on the faculty since 1984. He received a B.S. in Mathematics (with an Emphasis in Computer Science)
from the University of Santa Clara in 1978, a M.S. in Statistics from the University of Illinois at Urbana-
Champaign in 1982 and a Ph.D. in Mathematics from the University of Illinois at Urbana-Champaign in
1984. He teaches calculus, linear algebra and abstract algebra regularly, while his research interests include
the applications of linear algebra to graph theory. His professional website is at http://buzzard.ups.edu .
Edition
Version 2.30.
December 23, 2011.
Publisher
Robert A. Beezer
Department of Mathematics and Computer Science
University of Puget Sound
1500 North Warner
Tacoma, Washington 98416-1043
USA
c
2004 by Robert A. Beezer.
Permission is granted to copy, distribute and/or modify this document under the terms of the GNU Free
Documentation License, Version 1.2 or any later version published by the Free Software Foundation; with
no Invariant Sections, no Front-Cover Texts, and no Back-Cover Texts. A copy of the license is included
in the appendix entitled \GNU Free Documentation License".
The most recent version of this work can always be found at http://linear.ups.edu .
Tomywife, Pat
Contents
Table of Contents vii
Contributors ix
Denitions xi
Theorems xiii
Notation xv
Diagrams xvii
Examples xix
Preface xxi
Acknowledgements xxvii
Part C Core
Chapter SLE Systems of Linear Equations 3
WILA What is Linear Algebra? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
LA \Linear" + \Algebra" . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
AA An Application . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9
SSLE Solving Systems of Linear Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
SLE Systems of Linear Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
PSS Possibilities for Solution Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
ESEO Equivalent Systems and Equation Operations . . . . . . . . . . . . . . . . . . . . 13
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
RREF Reduced Row-Echelon Form . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
MVNSE Matrix and Vector Notation for Systems of Equations . . . . . . . . . . . . . . 27
RO Row Operations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
RREF Reduced Row-Echelon Form . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44
vii
viii CONTENTS
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48
TSS Types of Solution Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
CS Consistent Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
FV Free Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 67
HSE Homogeneous Systems of Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71
SHS Solutions of Homogeneous Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . 71
NSM Null Space of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 75
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79
NM Nonsingular Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
NM Nonsingular Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
NSNM Null Space of a Nonsingular Matrix . . . . . . . . . . . . . . . . . . . . . . . . . 85
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 88
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90
SLE Systems of Linear Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95
Chapter V Vectors 97
VO Vector Operations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97
VEASM Vector Equality, Addition, Scalar Multiplication . . . . . . . . . . . . . . . . . 98
VSP Vector Space Properties . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 101
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106
LC Linear Combinations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 109
LC Linear Combinations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 109
VFSS Vector Form of Solution Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 113
PSHS Particular Solutions, Homogeneous Solutions . . . . . . . . . . . . . . . . . . . . . 124
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129
SS Spanning Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131
SSV Span of a Set of Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131
SSNS Spanning Sets of Null Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 142
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 145
LI Linear Independence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153
LISV Linearly Independent Sets of Vectors . . . . . . . . . . . . . . . . . . . . . . . . . 153
LINM Linear Independence and Nonsingular Matrices . . . . . . . . . . . . . . . . . . . 158
NSSLI Null Spaces, Spans, Linear Independence . . . . . . . . . . . . . . . . . . . . . . 159
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 162
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 163
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167
LDS Linear Dependence and Spans . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 175
LDSS Linearly Dependent Sets and Spans . . . . . . . . . . . . . . . . . . . . . . . . . . 175
Version 2.30
CONTENTS ix
COV Casting Out Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 177
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 184
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 185
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 187
O Orthogonality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 191
CAV Complex Arithmetic and Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . 191
IP Inner products . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 192
N Norm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 195
OV Orthogonal Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 196
GSP Gram-Schmidt Procedure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 199
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 202
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 203
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 204
V Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205
Chapter M Matrices 207
MO Matrix Operations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 207
MEASM Matrix Equality, Addition, Scalar Multiplication . . . . . . . . . . . . . . . . . 207
VSP Vector Space Properties . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 209
TSM Transposes and Symmetric Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . 210
MCC Matrices and Complex Conjugation . . . . . . . . . . . . . . . . . . . . . . . . . . 212
AM Adjoint of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 214
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 215
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 216
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 219
MM Matrix Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 223
MVP Matrix-Vector Product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 223
MM Matrix Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 226
MMEE Matrix Multiplication, Entry-by-Entry . . . . . . . . . . . . . . . . . . . . . . . 227
PMM Properties of Matrix Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . 229
HM Hermitian Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 233
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 235
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 236
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 239
MISLE Matrix Inverses and Systems of Linear Equations . . . . . . . . . . . . . . . . . . . . 243
IM Inverse of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 244
CIM Computing the Inverse of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . 245
PMI Properties of Matrix Inverses . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 250
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 252
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 253
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 256
MINM Matrix Inverses and Nonsingular Matrices . . . . . . . . . . . . . . . . . . . . . . . . . 259
NMI Nonsingular Matrices are Invertible . . . . . . . . . . . . . . . . . . . . . . . . . . . 259
UM Unitary Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 262
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 265
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 266
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 268
CRS Column and Row Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 271
CSSE Column Spaces and Systems of Equations . . . . . . . . . . . . . . . . . . . . . . 271
CSSOC Column Space Spanned by Original Columns . . . . . . . . . . . . . . . . . . . 274
Version 2.30
x CONTENTS
CSNM Column Space of a Nonsingular Matrix . . . . . . . . . . . . . . . . . . . . . . . 276
RSM Row Space of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 278
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 283
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 284
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 288
FS Four Subsets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 293
LNS Left Null Space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 293
CRS Computing Column Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 294
EEF Extended echelon form . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 297
FS Four Subsets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 299
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 307
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 308
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 310
M Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 315
Chapter VS Vector Spaces 317
VS Vector Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 317
VS Vector Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 317
EVS Examples of Vector Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 319
VSP Vector Space Properties . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 323
RD Recycling Denitions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 326
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 327
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 328
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 330
S Subspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 333
TS Testing Subspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 334
TSS The Span of a Set . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 338
SC Subspace Constructions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 343
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 344
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 345
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 347
LISS Linear Independence and Spanning Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . 351
LI Linear Independence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 351
SS Spanning Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 355
VR Vector Representation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 359
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 361
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 362
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 364
B Bases . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 371
B Bases . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 371
BSCV Bases for Spans of Column Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . 374
BNM Bases and Nonsingular Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . 376
OBC Orthonormal Bases and Coordinates . . . . . . . . . . . . . . . . . . . . . . . . . . 377
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 382
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 383
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 385
D Dimension . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 391
D Dimension . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 391
DVS Dimension of Vector Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 395
RNM Rank and Nullity of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 397
Version 2.30
CONTENTS xi
RNNM Rank and Nullity of a Nonsingular Matrix . . . . . . . . . . . . . . . . . . . . . 398
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 400
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 401
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 403
PD Properties of Dimension . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 407
GT Goldilocks' Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 407
RT Ranks and Transposes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 410
DFS Dimension of Four Subspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 412
DS Direct Sums . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 413
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 417
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 418
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 419
VS Vector Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 421
Chapter D Determinants 423
DM Determinant of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 423
EM Elementary Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 423
DD Denition of the Determinant . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 427
CD Computing Determinants . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 429
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 433
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 434
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 436
PDM Properties of Determinants of Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . 439
DRO Determinants and Row Operations . . . . . . . . . . . . . . . . . . . . . . . . . . . 439
DROEM Determinants, Row Operations, Elementary Matrices . . . . . . . . . . . . . . 443
DNMMM Determinants, Nonsingular Matrices, Matrix Multiplication . . . . . . . . . . 445
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 448
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 449
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 450
D Determinants . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 451
Chapter E Eigenvalues 453
EE Eigenvalues and Eigenvectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 453
EEM Eigenvalues and Eigenvectors of a Matrix . . . . . . . . . . . . . . . . . . . . . . . 453
PM Polynomials and Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 455
EEE Existence of Eigenvalues and Eigenvectors . . . . . . . . . . . . . . . . . . . . . . . 456
CEE Computing Eigenvalues and Eigenvectors . . . . . . . . . . . . . . . . . . . . . . . 460
ECEE Examples of Computing Eigenvalues and Eigenvectors . . . . . . . . . . . . . . . 463
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 470
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 471
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 473
PEE Properties of Eigenvalues and Eigenvectors . . . . . . . . . . . . . . . . . . . . . . . . . 479
ME Multiplicities of Eigenvalues . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 484
EHM Eigenvalues of Hermitian Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . 487
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 488
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 489
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 490
SD Similarity and Diagonalization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 493
SM Similar Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 493
PSM Properties of Similar Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 494
D Diagonalization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 496
Version 2.30
xii CONTENTS
FS Fibonacci Sequences . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 503
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 506
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 507
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 508
E Eigenvalues . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 513
Chapter LT Linear Transformations 515
LT Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 515
LT Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 515
LTC Linear Transformation Cartoons . . . . . . . . . . . . . . . . . . . . . . . . . . . . 519
MLT Matrices and Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . 520
LTLC Linear Transformations and Linear Combinations . . . . . . . . . . . . . . . . . . 524
PI Pre-Images . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 528
NLTFO New Linear Transformations From Old . . . . . . . . . . . . . . . . . . . . . . . 530
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 534
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 535
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 537
ILT Injective Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 541
EILT Examples of Injective Linear Transformations . . . . . . . . . . . . . . . . . . . . . 541
KLT Kernel of a Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . 545
ILTLI Injective Linear Transformations and Linear Independence . . . . . . . . . . . . . 549
ILTD Injective Linear Transformations and Dimension . . . . . . . . . . . . . . . . . . . 550
CILT Composition of Injective Linear Transformations . . . . . . . . . . . . . . . . . . . 551
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 551
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 552
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 555
SLT Surjective Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 559
ESLT Examples of Surjective Linear Transformations . . . . . . . . . . . . . . . . . . . . 559
RLT Range of a Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . 563
SSSLT Spanning Sets and Surjective Linear Transformations . . . . . . . . . . . . . . . 567
SLTD Surjective Linear Transformations and Dimension . . . . . . . . . . . . . . . . . . 569
CSLT Composition of Surjective Linear Transformations . . . . . . . . . . . . . . . . . . 570
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 570
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 571
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 574
IVLT Invertible Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 579
IVLT Invertible Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . 579
IV Invertibility . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 582
SI Structure and Isomorphism . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 586
RNLT Rank and Nullity of a Linear Transformation . . . . . . . . . . . . . . . . . . . . 588
SLELT Systems of Linear Equations and Linear Transformations . . . . . . . . . . . . . 591
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 592
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 593
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 596
LT Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 601
Chapter R Representations 603
VR Vector Representations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 603
CVS Characterization of Vector Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . 608
CP Coordinatization Principle . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 609
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 612
Version 2.30
CONTENTS xiii
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 613
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 614
MR Matrix Representations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 615
NRFO New Representations from Old . . . . . . . . . . . . . . . . . . . . . . . . . . . . 621
PMR Properties of Matrix Representations . . . . . . . . . . . . . . . . . . . . . . . . . 625
IVLT Invertible Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . 630
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 634
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 635
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 638
CB Change of Basis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 647
EELT Eigenvalues and Eigenvectors of Linear Transformations . . . . . . . . . . . . . . 647
CBM Change-of-Basis Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 648
MRS Matrix Representations and Similarity . . . . . . . . . . . . . . . . . . . . . . . . . 654
CELT Computing Eigenvectors of Linear Transformations . . . . . . . . . . . . . . . . . 660
READ Reading Questions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 668
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 669
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 670
OD Orthonormal Diagonalization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 675
TM Triangular Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 675
UTMR Upper Triangular Matrix Representation . . . . . . . . . . . . . . . . . . . . . . 676
NM Normal Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 680
OD Orthonormal Diagonalization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 681
NLT Nilpotent Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 685
NLT Nilpotent Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . . 685
PNLT Properties of Nilpotent Linear Transformations . . . . . . . . . . . . . . . . . . . 690
CFNLT Canonical Form for Nilpotent Linear Transformations . . . . . . . . . . . . . . . 694
IS Invariant Subspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 703
IS Invariant Subspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 703
GEE Generalized Eigenvectors and Eigenspaces . . . . . . . . . . . . . . . . . . . . . . . 706
RLT Restrictions of Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . 711
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 720
JCF Jordan Canonical Form . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 721
GESD Generalized Eigenspace Decomposition . . . . . . . . . . . . . . . . . . . . . . . . 721
JCF Jordan Canonical Form . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 727
CHT Cayley-Hamilton Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 740
R Representations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 743
Appendix CN Computation Notes 745
MMA Mathematica . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 745
ME.MMA Matrix Entry . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 745
RR.MMA Row Reduce . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 745
LS.MMA Linear Solve . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 746
VLC.MMA Vector Linear Combinations . . . . . . . . . . . . . . . . . . . . . . . . . . . 746
NS.MMA Null Space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 747
VFSS.MMA Vector Form of Solution Set . . . . . . . . . . . . . . . . . . . . . . . . . . 747
GSP.MMA Gram-Schmidt Procedure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 748
TM.MMA Transpose of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 749
MM.MMA Matrix Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 749
MI.MMA Matrix Inverse . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 749
TI86 Texas Instruments 86 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 750
Version 2.30
xiv CONTENTS
ME.TI86 Matrix Entry . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 750
RR.TI86 Row Reduce . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 750
VLC.TI86 Vector Linear Combinations . . . . . . . . . . . . . . . . . . . . . . . . . . . . 750
TM.TI86 Transpose of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 751
TI83 Texas Instruments 83 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 751
ME.TI83 Matrix Entry . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 751
RR.TI83 Row Reduce . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 751
VLC.TI83 Vector Linear Combinations . . . . . . . . . . . . . . . . . . . . . . . . . . . . 752
SAGE SAGE: Open Source Mathematics Software . . . . . . . . . . . . . . . . . . . . . . . . 752
R.SAGE Rings . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 752
ME.SAGE Matrix Entry . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 753
RR.SAGE Row Reduce . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 753
LS.SAGE Linear Solve . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 754
VLC.SAGE Vector Linear Combinations . . . . . . . . . . . . . . . . . . . . . . . . . . . 755
MI.SAGE Matrix Inverse . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 755
TM.SAGE Transpose of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 755
E.SAGE Eigenspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 755
Appendix P Preliminaries 757
CNO Complex Number Operations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 757
CNA Arithmetic with complex numbers . . . . . . . . . . . . . . . . . . . . . . . . . . . 757
CCN Conjugates of Complex Numbers . . . . . . . . . . . . . . . . . . . . . . . . . . . . 759
MCN Modulus of a Complex Number . . . . . . . . . . . . . . . . . . . . . . . . . . . . 760
SET Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 761
SC Set Cardinality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 762
SO Set Operations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 763
PT Proof Techniques . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 765
D Denitions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 765
T Theorems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 766
L Language . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 766
GS Getting Started . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 767
C Constructive Proofs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 768
E Equivalences . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 768
N Negation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 769
CP Contrapositives . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 769
CV Converses . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 769
CD Contradiction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 770
U Uniqueness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 771
ME Multiple Equivalences . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 771
PI Proving Identities . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 771
DC Decompositions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 772
I Induction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 772
P Practice . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 774
LC Lemmas and Corollaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 774
Appendix A Archetypes 777
A . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 781
B . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 786
C . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 791
D . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 795
E . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 799
Version 2.30
CONTENTS xv
F . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 803
G . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 808
H . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 812
I . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 816
J . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 820
K . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 825
L . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 829
M . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 833
N . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 836
O . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 839
P . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 842
Q . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 844
R . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 848
S . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 851
T . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 854
U . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 856
V . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 858
W . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 860
X . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 862
Appendix GFDL GNU Free Documentation License 865
1. APPLICABILITY AND DEFINITIONS . . . . . . . . . . . . . . . . . . . . . . . . . . . . 865
2. VERBATIM COPYING . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 866
3. COPYING IN QUANTITY . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 866
4. MODIFICATIONS . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 867
5. COMBINING DOCUMENTS . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 868
6. COLLECTIONS OF DOCUMENTS . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 869
7. AGGREGATION WITH INDEPENDENT WORKS . . . . . . . . . . . . . . . . . . . . . 869
8. TRANSLATION . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 869
9. TERMINATION . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 869
10. FUTURE REVISIONS OF THIS LICENSE . . . . . . . . . . . . . . . . . . . . . . . . . 869
ADDENDUM: How to use this License for your documents . . . . . . . . . . . . . . . . . . . 870
Part T Topics
F Fields . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 873
F Fields . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 873
FF Finite Fields . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 874
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 879
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 881
T Trace . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 883
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 887
SOL Solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 888
HP Hadamard Product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 889
DMHP Diagonal Matrices and the Hadamard Product . . . . . . . . . . . . . . . . . . . 891
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 894
VM Vandermonde Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 895
PSM Positive Semi-denite Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 899
PSM Positive Semi-Denite Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 899
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 902
Version 2.30
xvi CONTENTS
Chapter MD Matrix Decompositions 903
ROD Rank One Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 903
TD Triangular Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 909
TD Triangular Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 909
TDSSE Triangular Decomposition and Solving Systems of Equations . . . . . . . . . . . 912
CTD Computing Triangular Decompositions . . . . . . . . . . . . . . . . . . . . . . . . 913
SVD Singular Value Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 917
MAP Matrix-Adjoint Product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 917
SVD Singular Value Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 920
SR Square Roots . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 923
SRM Square Root of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 923
POD Polar Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 927
Part A Applications
CF Curve Fitting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 931
DF Data Fitting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 932
EXC Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 935
SAS Sharing A Secret . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 937
Version 2.30
Contributors
Beezer, David. Belarmine Preparatory School, Tacoma
Beezer, Robert. University of Puget Sound http://buzzard.ups.edu/
Black, Chris.
Braithwaite, David. Chicago, Illinois
Bucht, Sara. University of Puget Sound
Caneld, Steve. University of Puget Sound
Hubert, Dupont. Cr eteil, France
Fellez, Sarah. University of Puget Sound
Fickenscher, Eric. University of Puget Sound
Jackson, Martin. University of Puget Sound http://www.math.ups.edu/~martinj
Kessler, Ivan. University of Puget Sound
Kreher, Don. Michigan Technological University http://www.math.mtu.edu/~kreher/
Hamrick, Mark. St. Louis University
Linenthal, Jacob. University of Puget Sound
Million, Elizabeth. University of Puget Sound
Osborne, Travis. University of Puget Sound
Riegsecker, Joe. Middlebury, Indiana joepye (at) pobox (dot) com
Perkel, Manley. University of Puget Sound
Phelps, Douglas. University of Puget Sound
Shoemaker, Mark. University of Puget Sound
Toth, Zoltan. http://zoli.web.elte.hu
Zimmer, Andy. University of Puget Sound
xvii
xviii CONTRIBUTORS
Version 2.30
Denitions
Section WILA
Section SSLE
SLE System of Linear Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
SSLE Solution of a System of Linear Equations . . . . . . . . . . . . . . . . . . . . . . . . . 12
SSSLE Solution Set of a System of Linear Equations . . . . . . . . . . . . . . . . . . . . . . . 12
ESYS Equivalent Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
EO Equation Operations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
Section RREF
M Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
CV Column Vector . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
ZCV Zero Column Vector . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
CM Coecient Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
VOC Vector of Constants . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
SOLV Solution Vector . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
MRLS Matrix Representation of a Linear System . . . . . . . . . . . . . . . . . . . . . . . . 29
AM Augmented Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
RO Row Operations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
REM Row-Equivalent Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
RREF Reduced Row-Echelon Form . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33
RR Row-Reducing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
Section TSS
CS Consistent System . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
IDV Independent and Dependent Variables . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
Section HSE
HS Homogeneous System . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71
TSHSE Trivial Solution to Homogeneous Systems of Equations . . . . . . . . . . . . . . . . . 71
NSM Null Space of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73
Section NM
SQM Square Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
NM Nonsingular Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
IM Identity Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 84
Section VO
VSCV Vector Space of Column Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97
CVE Column Vector Equality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98
xix
xx DEFINITIONS
CVA Column Vector Addition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98
CVSM Column Vector Scalar Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . 99
Section LC
LCCV Linear Combination of Column Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . 109
Section SS
SSCV Span of a Set of Column Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131
Section LI
RLDCV Relation of Linear Dependence for Column Vectors . . . . . . . . . . . . . . . . . . . 153
LICV Linear Independence of Column Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . 153
Section LDS
Section O
CCCV Complex Conjugate of a Column Vector . . . . . . . . . . . . . . . . . . . . . . . . . 191
IP Inner Product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 192
NV Norm of a Vector . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 195
OV Orthogonal Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 196
OSV Orthogonal Set of Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 197
SUV Standard Unit Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 197
ONS OrthoNormal Set . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 201
Section MO
VSM Vector Space of mnMatrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 207
ME Matrix Equality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 207
MA Matrix Addition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 207
MSM Matrix Scalar Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 208
ZM Zero Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 210
TM Transpose of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 210
SYM Symmetric Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 211
CCM Complex Conjugate of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 212
A Adjoint . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 214
Section MM
MVP Matrix-Vector Product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 223
MM Matrix Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 226
HM Hermitian Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 234
Section MISLE
MI Matrix Inverse . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 244
Section MINM
UM Unitary Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 262
Section CRS
CSM Column Space of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 271
RSM Row Space of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 278
Section FS
Version 2.30
DEFINITIONS xxi
LNS Left Null Space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 293
EEF Extended Echelon Form . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 297
Section VS
VS Vector Space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 317
Section S
S Subspace . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 333
TS Trivial Subspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 337
LC Linear Combination . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 338
SS Span of a Set . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 339
Section LISS
RLD Relation of Linear Dependence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 351
LI Linear Independence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 351
TSVS To Span a Vector Space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 356
Section B
B Basis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 371
Section D
D Dimension . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 391
NOM Nullity Of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 397
ROM Rank Of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 397
Section PD
DS Direct Sum . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 413
Section DM
ELEM Elementary Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 423
SM SubMatrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 428
DM Determinant of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 428
Section PDM
Section EE
EEM Eigenvalues and Eigenvectors of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . 453
CP Characteristic Polynomial . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 460
EM Eigenspace of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 461
AME Algebraic Multiplicity of an Eigenvalue . . . . . . . . . . . . . . . . . . . . . . . . . . 463
GME Geometric Multiplicity of an Eigenvalue . . . . . . . . . . . . . . . . . . . . . . . . . . 463
Section PEE
Section SD
SIM Similar Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 493
DIM Diagonal Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 496
DZM Diagonalizable Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 496
Section LT
LT Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 515
PI Pre-Image . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 528
Version 2.30
xxii DEFINITIONS
LTA Linear Transformation Addition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 530
LTSM Linear Transformation Scalar Multiplication . . . . . . . . . . . . . . . . . . . . . . . 531
LTC Linear Transformation Composition . . . . . . . . . . . . . . . . . . . . . . . . . . . . 532
Section ILT
ILT Injective Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 541
KLT Kernel of a Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 545
Section SLT
SLT Surjective Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 559
RLT Range of a Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 563
Section IVLT
IDLT Identity Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 579
IVLT Invertible Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 579
IVS Isomorphic Vector Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 586
ROLT Rank Of a Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 588
NOLT Nullity Of a Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 588
Section VR
VR Vector Representation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 603
Section MR
MR Matrix Representation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 615
Section CB
EELT Eigenvalue and Eigenvector of a Linear Transformation . . . . . . . . . . . . . . . . . 647
CBM Change-of-Basis Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 648
Section OD
UTM Upper Triangular Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 675
LTM Lower Triangular Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 675
NRML Normal Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 680
Section NLT
NLT Nilpotent Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 685
JB Jordan Block . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 687
Section IS
IS Invariant Subspace . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 703
GEV Generalized Eigenvector . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 707
GES Generalized Eigenspace . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 707
LTR Linear Transformation Restriction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 711
IE Index of an Eigenvalue . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 717
Section JCF
JCF Jordan Canonical Form . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 727
Section CNO
CNE Complex Number Equality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 758
Version 2.30
DEFINITIONS xxiii
CNA Complex Number Addition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 758
CNM Complex Number Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 758
CCN Conjugate of a Complex Number . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 759
MCN Modulus of a Complex Number . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 760
Section SET
SET Set . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 761
SSET Subset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 761
ES Empty Set . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 761
SE Set Equality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 762
C Cardinality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 762
SU Set Union . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 763
SI Set Intersection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 763
SC Set Complement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 763
Section PT
Section F
F Field . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 873
IMP Integers Modulo a Prime . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 874
Section T
T Trace . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 883
Section HP
HP Hadamard Product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 889
HID Hadamard Identity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 890
HI Hadamard Inverse . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 890
Section VM
VM Vandermonde Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 895
Section PSM
PSM Positive Semi-Denite Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 899
Section ROD
Section TD
Section SVD
SV Singular Values . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 921
Section SR
SRM Square Root of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 926
Section POD
Section CF
LSS Least Squares Solution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 932
Section SAS
Version 2.30
xxiv DEFINITIONS
Version 2.30
Theorems
Section WILA
Section SSLE
EOPSS Equation Operations Preserve Solution Sets . . . . . . . . . . . . . . . . . . . . . . . 14
Section RREF
REMES Row-Equivalent Matrices represent Equivalent Systems . . . . . . . . . . . . . . . . . 31
REMEF Row-Equivalent Matrix in Echelon Form . . . . . . . . . . . . . . . . . . . . . . . . . 34
RREFU Reduced Row-Echelon Form is Unique . . . . . . . . . . . . . . . . . . . . . . . . . . 35
Section TSS
RCLS Recognizing Consistency of a Linear System . . . . . . . . . . . . . . . . . . . . . . . 58
ISRN Inconsistent Systems, randn. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
CSRN Consistent Systems, randn. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
FVCS Free Variables for Consistent Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . 60
PSSLS Possible Solution Sets for Linear Systems . . . . . . . . . . . . . . . . . . . . . . . . . 60
CMVEI Consistent, More Variables than Equations, Innite solutions . . . . . . . . . . . . . . 61
Section HSE
HSC Homogeneous Systems are Consistent . . . . . . . . . . . . . . . . . . . . . . . . . . . 71
HMVEI Homogeneous, More Variables than Equations, Innite solutions . . . . . . . . . . . . 73
Section NM
NMRRI Nonsingular Matrices Row Reduce to the Identity matrix . . . . . . . . . . . . . . . . 84
NMTNS Nonsingular Matrices have Trivial Null Spaces . . . . . . . . . . . . . . . . . . . . . . 86
NMUS Nonsingular Matrices and Unique Solutions . . . . . . . . . . . . . . . . . . . . . . . . 86
NME1 Nonsingular Matrix Equivalences, Round 1 . . . . . . . . . . . . . . . . . . . . . . . . 87
Section VO
VSPCV Vector Space Properties of Column Vectors . . . . . . . . . . . . . . . . . . . . . . . . 100
Section LC
SLSLC Solutions to Linear Systems are Linear Combinations . . . . . . . . . . . . . . . . . . 112
VFSLS Vector Form of Solutions to Linear Systems . . . . . . . . . . . . . . . . . . . . . . . 118
PSPHS Particular Solution Plus Homogeneous Solutions . . . . . . . . . . . . . . . . . . . . . 124
Section SS
SSNS Spanning Sets for Null Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 137
Section LI
xxv
xxvi THEOREMS
LIVHS Linearly Independent Vectors and Homogeneous Systems . . . . . . . . . . . . . . . . 155
LIVRN Linearly Independent Vectors, randn. . . . . . . . . . . . . . . . . . . . . . . . . . 156
MVSLD More Vectors than Size implies Linear Dependence . . . . . . . . . . . . . . . . . . . 158
NMLIC Nonsingular Matrices have Linearly Independent Columns . . . . . . . . . . . . . . . 159
NME2 Nonsingular Matrix Equivalences, Round 2 . . . . . . . . . . . . . . . . . . . . . . . . 159
BNS Basis for Null Spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160
Section LDS
DLDS Dependency in Linearly Dependent Sets . . . . . . . . . . . . . . . . . . . . . . . . . 175
BS Basis of a Span . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 180
Section O
CRVA Conjugation Respects Vector Addition . . . . . . . . . . . . . . . . . . . . . . . . . . 191
CRSM Conjugation Respects Vector Scalar Multiplication . . . . . . . . . . . . . . . . . . . 191
IPVA Inner Product and Vector Addition . . . . . . . . . . . . . . . . . . . . . . . . . . . . 193
IPSM Inner Product and Scalar Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . 194
IPAC Inner Product is Anti-Commutative . . . . . . . . . . . . . . . . . . . . . . . . . . . . 194
IPN Inner Products and Norms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 195
PIP Positive Inner Products . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 196
OSLI Orthogonal Sets are Linearly Independent . . . . . . . . . . . . . . . . . . . . . . . . 198
GSP Gram-Schmidt Procedure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 199
Section MO
VSPM Vector Space Properties of Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . 209
SMS Symmetric Matrices are Square . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 211
TMA Transpose and Matrix Addition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 211
TMSM Transpose and Matrix Scalar Multiplication . . . . . . . . . . . . . . . . . . . . . . . 212
TT Transpose of a Transpose . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 212
CRMA Conjugation Respects Matrix Addition . . . . . . . . . . . . . . . . . . . . . . . . . . 213
CRMSM Conjugation Respects Matrix Scalar Multiplication . . . . . . . . . . . . . . . . . . . 213
CCM Conjugate of the Conjugate of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . 213
MCT Matrix Conjugation and Transposes . . . . . . . . . . . . . . . . . . . . . . . . . . . . 214
AMA Adjoint and Matrix Addition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 214
AMSM Adjoint and Matrix Scalar Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . 214
AA Adjoint of an Adjoint . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 215
Section MM
SLEMM Systems of Linear Equations as Matrix Multiplication . . . . . . . . . . . . . . . . . . 224
EMMVP Equal Matrices and Matrix-Vector Products . . . . . . . . . . . . . . . . . . . . . . . 225
EMP Entries of Matrix Products . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 227
MMZM Matrix Multiplication and the Zero Matrix . . . . . . . . . . . . . . . . . . . . . . . . 229
MMIM Matrix Multiplication and Identity Matrix . . . . . . . . . . . . . . . . . . . . . . . . 229
MMDAA Matrix Multiplication Distributes Across Addition . . . . . . . . . . . . . . . . . . . . 230
MMSMM Matrix Multiplication and Scalar Matrix Multiplication . . . . . . . . . . . . . . . . . 230
MMA Matrix Multiplication is Associative . . . . . . . . . . . . . . . . . . . . . . . . . . . 231
MMIP Matrix Multiplication and Inner Products . . . . . . . . . . . . . . . . . . . . . . . . 231
MMCC Matrix Multiplication and Complex Conjugation . . . . . . . . . . . . . . . . . . . . . 232
MMT Matrix Multiplication and Transposes . . . . . . . . . . . . . . . . . . . . . . . . . . . 232
MMAD Matrix Multiplication and Adjoints . . . . . . . . . . . . . . . . . . . . . . . . . . . . 233
AIP Adjoint and Inner Product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 233
Version 2.30
THEOREMS xxvii
HMIP Hermitian Matrices and Inner Products . . . . . . . . . . . . . . . . . . . . . . . . . . 234
Section MISLE
TTMI Two-by-Two Matrix Inverse . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 246
CINM Computing the Inverse of a Nonsingular Matrix . . . . . . . . . . . . . . . . . . . . . 248
MIU Matrix Inverse is Unique . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 250
SS Socks and Shoes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 250
MIMI Matrix Inverse of a Matrix Inverse . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 251
MIT Matrix Inverse of a Transpose . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 251
MISM Matrix Inverse of a Scalar Multiple . . . . . . . . . . . . . . . . . . . . . . . . . . . . 252
Section MINM
NPNT Nonsingular Product has Nonsingular Terms . . . . . . . . . . . . . . . . . . . . . . . 259
OSIS One-Sided Inverse is Sucient . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 260
NI Nonsingularity is Invertibility . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 261
NME3 Nonsingular Matrix Equivalences, Round 3 . . . . . . . . . . . . . . . . . . . . . . . . 261
SNCM Solution with Nonsingular Coecient Matrix . . . . . . . . . . . . . . . . . . . . . . . 261
UMI Unitary Matrices are Invertible . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 263
CUMOS Columns of Unitary Matrices are Orthonormal Sets . . . . . . . . . . . . . . . . . . . 263
UMPIP Unitary Matrices Preserve Inner Products . . . . . . . . . . . . . . . . . . . . . . . . 264
Section CRS
CSCS Column Spaces and Consistent Systems . . . . . . . . . . . . . . . . . . . . . . . . . . 272
BCS Basis of the Column Space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 274
CSNM Column Space of a Nonsingular Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . 277
NME4 Nonsingular Matrix Equivalences, Round 4 . . . . . . . . . . . . . . . . . . . . . . . . 277
REMRS Row-Equivalent Matrices have equal Row Spaces . . . . . . . . . . . . . . . . . . . . 279
BRS Basis for the Row Space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 280
CSRST Column Space, Row Space, Transpose . . . . . . . . . . . . . . . . . . . . . . . . . . . 282
Section FS
PEEF Properties of Extended Echelon Form . . . . . . . . . . . . . . . . . . . . . . . . . . . 298
FS Four Subsets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 299
Section VS
ZVU Zero Vector is Unique . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 324
AIU Additive Inverses are Unique . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 324
ZSSM Zero Scalar in Scalar Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . . 324
ZVSM Zero Vector in Scalar Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . . 325
AISM Additive Inverses from Scalar Multiplication . . . . . . . . . . . . . . . . . . . . . . . 325
SMEZV Scalar Multiplication Equals the Zero Vector . . . . . . . . . . . . . . . . . . . . . . . 326
Section S
TSS Testing Subsets for Subspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 334
NSMS Null Space of a Matrix is a Subspace . . . . . . . . . . . . . . . . . . . . . . . . . . . 337
SSS Span of a Set is a Subspace . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 339
CSMS Column Space of a Matrix is a Subspace . . . . . . . . . . . . . . . . . . . . . . . . . 343
RSMS Row Space of a Matrix is a Subspace . . . . . . . . . . . . . . . . . . . . . . . . . . . 344
LNSMS Left Null Space of a Matrix is a Subspace . . . . . . . . . . . . . . . . . . . . . . . . . 344
Version 2.30
xxviii THEOREMS
Section LISS
VRRB Vector Representation Relative to a Basis . . . . . . . . . . . . . . . . . . . . . . . . . 360
Section B
SUVB Standard Unit Vectors are a Basis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 371
CNMB Columns of Nonsingular Matrix are a Basis . . . . . . . . . . . . . . . . . . . . . . . . 376
NME5 Nonsingular Matrix Equivalences, Round 5 . . . . . . . . . . . . . . . . . . . . . . . . 377
COB Coordinates and Orthonormal Bases . . . . . . . . . . . . . . . . . . . . . . . . . . . . 378
UMCOB Unitary Matrices Convert Orthonormal Bases . . . . . . . . . . . . . . . . . . . . . . 380
Section D
SSLD Spanning Sets and Linear Dependence . . . . . . . . . . . . . . . . . . . . . . . . . . 391
BIS Bases have Identical Sizes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 394
DCM Dimension of Cm. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 395
DP Dimension of Pn. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 395
DM Dimension of Mmn . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 395
CRN Computing Rank and Nullity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 397
RPNC Rank Plus Nullity is Columns . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 398
RNNM Rank and Nullity of a Nonsingular Matrix . . . . . . . . . . . . . . . . . . . . . . . . 399
NME6 Nonsingular Matrix Equivalences, Round 6 . . . . . . . . . . . . . . . . . . . . . . . . 399
Section PD
ELIS Extending Linearly Independent Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . 407
G Goldilocks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 407
PSSD Proper Subspaces have Smaller Dimension . . . . . . . . . . . . . . . . . . . . . . . . 410
EDYES Equal Dimensions Yields Equal Subspaces . . . . . . . . . . . . . . . . . . . . . . . . 410
RMRT Rank of a Matrix is the Rank of the Transpose . . . . . . . . . . . . . . . . . . . . . . 411
DFS Dimensions of Four Subspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 412
DSFB Direct Sum From a Basis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 413
DSFOS Direct Sum From One Subspace . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 414
DSZV Direct Sums and Zero Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 414
DSZI Direct Sums and Zero Intersection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 415
DSLI Direct Sums and Linear Independence . . . . . . . . . . . . . . . . . . . . . . . . . . . 416
DSD Direct Sums and Dimension . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 416
RDS Repeated Direct Sums . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 417
Section DM
EMDRO Elementary Matrices Do Row Operations . . . . . . . . . . . . . . . . . . . . . . . . . 425
EMN Elementary Matrices are Nonsingular . . . . . . . . . . . . . . . . . . . . . . . . . . . 427
NMPEM Nonsingular Matrices are Products of Elementary Matrices . . . . . . . . . . . . . . . 427
DMST Determinant of Matrices of Size Two . . . . . . . . . . . . . . . . . . . . . . . . . . . 429
DER Determinant Expansion about Rows . . . . . . . . . . . . . . . . . . . . . . . . . . . . 429
DT Determinant of the Transpose . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 430
DEC Determinant Expansion about Columns . . . . . . . . . . . . . . . . . . . . . . . . . . 431
Section PDM
DZRC Determinant with Zero Row or Column . . . . . . . . . . . . . . . . . . . . . . . . . . 439
DRCS Determinant for Row or Column Swap . . . . . . . . . . . . . . . . . . . . . . . . . . 439
DRCM Determinant for Row or Column Multiples . . . . . . . . . . . . . . . . . . . . . . . . 440
DERC Determinant with Equal Rows or Columns . . . . . . . . . . . . . . . . . . . . . . . . 441
Version 2.30
THEOREMS xxix
DRCMA Determinant for Row or Column Multiples and Addition . . . . . . . . . . . . . . . . 441
DIM Determinant of the Identity Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . 443
DEM Determinants of Elementary Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . 444
DEMMM Determinants, Elementary Matrices, Matrix Multiplication . . . . . . . . . . . . . . . 445
SMZD Singular Matrices have Zero Determinants . . . . . . . . . . . . . . . . . . . . . . . . 445
NME7 Nonsingular Matrix Equivalences, Round 7 . . . . . . . . . . . . . . . . . . . . . . . . 446
DRMM Determinant Respects Matrix Multiplication . . . . . . . . . . . . . . . . . . . . . . . 447
Section EE
EMHE Every Matrix Has an Eigenvalue . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 457
EMRCP Eigenvalues of a Matrix are Roots of Characteristic Polynomials . . . . . . . . . . . . 461
EMS Eigenspace for a Matrix is a Subspace . . . . . . . . . . . . . . . . . . . . . . . . . . . 461
EMNS Eigenspace of a Matrix is a Null Space . . . . . . . . . . . . . . . . . . . . . . . . . . 462
Section PEE
EDELI Eigenvectors with Distinct Eigenvalues are Linearly Independent . . . . . . . . . . . . 479
SMZE Singular Matrices have Zero Eigenvalues . . . . . . . . . . . . . . . . . . . . . . . . . 480
NME8 Nonsingular Matrix Equivalences, Round 8 . . . . . . . . . . . . . . . . . . . . . . . . 480
ESMM Eigenvalues of a Scalar Multiple of a Matrix . . . . . . . . . . . . . . . . . . . . . . . 481
EOMP Eigenvalues Of Matrix Powers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 481
EPM Eigenvalues of the Polynomial of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . 481
EIM Eigenvalues of the Inverse of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . 482
ETM Eigenvalues of the Transpose of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . 483
ERMCP Eigenvalues of Real Matrices come in Conjugate Pairs . . . . . . . . . . . . . . . . . . 483
DCP Degree of the Characteristic Polynomial . . . . . . . . . . . . . . . . . . . . . . . . . . 484
NEM Number of Eigenvalues of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . 485
ME Multiplicities of an Eigenvalue . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 485
MNEM Maximum Number of Eigenvalues of a Matrix . . . . . . . . . . . . . . . . . . . . . . 487
HMRE Hermitian Matrices have Real Eigenvalues . . . . . . . . . . . . . . . . . . . . . . . . 487
HMOE Hermitian Matrices have Orthogonal Eigenvectors . . . . . . . . . . . . . . . . . . . . 488
Section SD
SER Similarity is an Equivalence Relation . . . . . . . . . . . . . . . . . . . . . . . . . . . 494
SMEE Similar Matrices have Equal Eigenvalues . . . . . . . . . . . . . . . . . . . . . . . . . 495
DC Diagonalization Characterization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 497
DMFE Diagonalizable Matrices have Full Eigenspaces . . . . . . . . . . . . . . . . . . . . . . 499
DED Distinct Eigenvalues implies Diagonalizable . . . . . . . . . . . . . . . . . . . . . . . . 501
Section LT
LTTZZ Linear Transformations Take Zero to Zero . . . . . . . . . . . . . . . . . . . . . . . . 519
MBLT Matrices Build Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . 522
MLTCV Matrix of a Linear Transformation, Column Vectors . . . . . . . . . . . . . . . . . . . 523
LTLC Linear Transformations and Linear Combinations . . . . . . . . . . . . . . . . . . . . 525
LTDB Linear Transformation Dened on a Basis . . . . . . . . . . . . . . . . . . . . . . . . 525
SLTLT Sum of Linear Transformations is a Linear Transformation . . . . . . . . . . . . . . . 530
MLTLT Multiple of a Linear Transformation is a Linear Transformation . . . . . . . . . . . . 531
VSLT Vector Space of Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . . . . 532
CLTLT Composition of Linear Transformations is a Linear Transformation . . . . . . . . . . 533
Section ILT
Version 2.30
xxx THEOREMS
KLTS Kernel of a Linear Transformation is a Subspace . . . . . . . . . . . . . . . . . . . . . 546
KPI Kernel and Pre-Image . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 547
KILT Kernel of an Injective Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . 548
ILTLI Injective Linear Transformations and Linear Independence . . . . . . . . . . . . . . . 549
ILTB Injective Linear Transformations and Bases . . . . . . . . . . . . . . . . . . . . . . . . 550
ILTD Injective Linear Transformations and Dimension . . . . . . . . . . . . . . . . . . . . . 550
CILTI Composition of Injective Linear Transformations is Injective . . . . . . . . . . . . . . 551
Section SLT
RLTS Range of a Linear Transformation is a Subspace . . . . . . . . . . . . . . . . . . . . . 564
RSLT Range of a Surjective Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . 565
SSRLT Spanning Set for Range of a Linear Transformation . . . . . . . . . . . . . . . . . . . 567
RPI Range and Pre-Image . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 568
SLTB Surjective Linear Transformations and Bases . . . . . . . . . . . . . . . . . . . . . . . 568
SLTD Surjective Linear Transformations and Dimension . . . . . . . . . . . . . . . . . . . . 569
CSLTS Composition of Surjective Linear Transformations is Surjective . . . . . . . . . . . . . 570
Section IVLT
ILTLT Inverse of a Linear Transformation is a Linear Transformation . . . . . . . . . . . . . 582
IILT Inverse of an Invertible Linear Transformation . . . . . . . . . . . . . . . . . . . . . . 582
ILTIS Invertible Linear Transformations are Injective and Surjective . . . . . . . . . . . . . 582
CIVLT Composition of Invertible Linear Transformations . . . . . . . . . . . . . . . . . . . . 585
ICLT Inverse of a Composition of Linear Transformations . . . . . . . . . . . . . . . . . . . 585
IVSED Isomorphic Vector Spaces have Equal Dimension . . . . . . . . . . . . . . . . . . . . . 587
ROSLT Rank Of a Surjective Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . 588
NOILT Nullity Of an Injective Linear Transformation . . . . . . . . . . . . . . . . . . . . . . 588
RPNDD Rank Plus Nullity is Domain Dimension . . . . . . . . . . . . . . . . . . . . . . . . . 588
Section VR
VRLT Vector Representation is a Linear Transformation . . . . . . . . . . . . . . . . . . . . 603
VRI Vector Representation is Injective . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 607
VRS Vector Representation is Surjective . . . . . . . . . . . . . . . . . . . . . . . . . . . . 608
VRILT Vector Representation is an Invertible Linear Transformation . . . . . . . . . . . . . . 608
CFDVS Characterization of Finite Dimensional Vector Spaces . . . . . . . . . . . . . . . . . . 608
IFDVS Isomorphism of Finite Dimensional Vector Spaces . . . . . . . . . . . . . . . . . . . . 609
CLI Coordinatization and Linear Independence . . . . . . . . . . . . . . . . . . . . . . . . 609
CSS Coordinatization and Spanning Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . 610
Section MR
FTMR Fundamental Theorem of Matrix Representation . . . . . . . . . . . . . . . . . . . . . 617
MRSLT Matrix Representation of a Sum of Linear Transformations . . . . . . . . . . . . . . . 621
MRMLT Matrix Representation of a Multiple of a Linear Transformation . . . . . . . . . . . . 621
MRCLT Matrix Representation of a Composition of Linear Transformations . . . . . . . . . . 622
KNSI Kernel and Null Space Isomorphism . . . . . . . . . . . . . . . . . . . . . . . . . . . . 625
RCSI Range and Column Space Isomorphism . . . . . . . . . . . . . . . . . . . . . . . . . . 628
IMR Invertible Matrix Representations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 630
IMILT Invertible Matrices, Invertible Linear Transformation . . . . . . . . . . . . . . . . . . 633
NME9 Nonsingular Matrix Equivalences, Round 9 . . . . . . . . . . . . . . . . . . . . . . . . 633
Section CB
Version 2.30
THEOREMS xxxi
CB Change-of-Basis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 649
ICBM Inverse of Change-of-Basis Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 649
MRCB Matrix Representation and Change of Basis . . . . . . . . . . . . . . . . . . . . . . . 654
SCB Similarity and Change of Basis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 656
EER Eigenvalues, Eigenvectors, Representations . . . . . . . . . . . . . . . . . . . . . . . . 659
Section OD
PTMT Product of Triangular Matrices is Triangular . . . . . . . . . . . . . . . . . . . . . . . 675
ITMT Inverse of a Triangular Matrix is Triangular . . . . . . . . . . . . . . . . . . . . . . . 676
UTMR Upper Triangular Matrix Representation . . . . . . . . . . . . . . . . . . . . . . . . . 676
OBUTR Orthonormal Basis for Upper Triangular Representation . . . . . . . . . . . . . . . . 679
OD Orthonormal Diagonalization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 681
OBNM Orthonormal Bases and Normal Matrices . . . . . . . . . . . . . . . . . . . . . . . . . 683
Section NLT
NJB Nilpotent Jordan Blocks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 689
ENLT Eigenvalues of Nilpotent Linear Transformations . . . . . . . . . . . . . . . . . . . . . 690
DNLT Diagonalizable Nilpotent Linear Transformations . . . . . . . . . . . . . . . . . . . . . 691
KPLT Kernels of Powers of Linear Transformations . . . . . . . . . . . . . . . . . . . . . . . 691
KPNLT Kernels of Powers of Nilpotent Linear Transformations . . . . . . . . . . . . . . . . . 692
CFNLT Canonical Form for Nilpotent Linear Transformations . . . . . . . . . . . . . . . . . . 694
Section IS
EIS Eigenspaces are Invariant Subspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . 705
KPIS Kernels of Powers are Invariant Subspaces . . . . . . . . . . . . . . . . . . . . . . . . 705
GESIS Generalized Eigenspace is an Invariant Subspace . . . . . . . . . . . . . . . . . . . . . 707
GEK Generalized Eigenspace as a Kernel . . . . . . . . . . . . . . . . . . . . . . . . . . . . 708
RGEN Restriction to Generalized Eigenspace is Nilpotent . . . . . . . . . . . . . . . . . . . . 716
MRRGE Matrix Representation of a Restriction to a Generalized Eigenspace . . . . . . . . . . 719
Section JCF
GESD Generalized Eigenspace Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . 721
DGES Dimension of Generalized Eigenspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . 727
JCFLT Jordan Canonical Form for a Linear Transformation . . . . . . . . . . . . . . . . . . . 728
CHT Cayley-Hamilton Theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 740
Section CNO
PCNA Properties of Complex Number Arithmetic . . . . . . . . . . . . . . . . . . . . . . . . 758
CCRA Complex Conjugation Respects Addition . . . . . . . . . . . . . . . . . . . . . . . . . 759
CCRM Complex Conjugation Respects Multiplication . . . . . . . . . . . . . . . . . . . . . . 760
CCT Complex Conjugation Twice . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 760
Section SET
Section PT
Section F
FIMP Field of Integers Modulo a Prime . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 875
Section T
TL Trace is Linear . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 883
TSRM Trace is Symmetric with Respect to Multiplication . . . . . . . . . . . . . . . . . . . 884
Version 2.30
xxxii THEOREMS
TIST Trace is Invariant Under Similarity Transformations . . . . . . . . . . . . . . . . . . . 884
TSE Trace is the Sum of the Eigenvalues . . . . . . . . . . . . . . . . . . . . . . . . . . . . 884
Section HP
HPC Hadamard Product is Commutative . . . . . . . . . . . . . . . . . . . . . . . . . . . . 889
HPHID Hadamard Product with the Hadamard Identity . . . . . . . . . . . . . . . . . . . . . 890
HPHI Hadamard Product with Hadamard Inverses . . . . . . . . . . . . . . . . . . . . . . . 890
HPDAA Hadamard Product Distributes Across Addition . . . . . . . . . . . . . . . . . . . . . 891
HPSMM Hadamard Product and Scalar Matrix Multiplication . . . . . . . . . . . . . . . . . . 891
DMHP Diagonalizable Matrices and the Hadamard Product . . . . . . . . . . . . . . . . . . . 891
DMMP Diagonal Matrices and Matrix Products . . . . . . . . . . . . . . . . . . . . . . . . . . 892
Section VM
DVM Determinant of a Vandermonde Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . 895
NVM Nonsingular Vandermonde Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 898
Section PSM
CPSM Creating Positive Semi-Denite Matrices . . . . . . . . . . . . . . . . . . . . . . . . . 899
EPSM Eigenvalues of Positive Semi-denite Matrices . . . . . . . . . . . . . . . . . . . . . . 900
Section ROD
ROD Rank One Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 904
Section TD
TD Triangular Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 909
TDEE Triangular Decomposition, Entry by Entry . . . . . . . . . . . . . . . . . . . . . . . . 913
Section SVD
EEMAP Eigenvalues and Eigenvectors of Matrix-Adjoint Product . . . . . . . . . . . . . . . . 917
SVD Singular Value Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 921
Section SR
PSMSR Positive Semi-Denite Matrices and Square Roots . . . . . . . . . . . . . . . . . . . . 923
EESR Eigenvalues and Eigenspaces of a Square Root . . . . . . . . . . . . . . . . . . . . . . 924
USR Unique Square Root . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 926
Section POD
PDM Polar Decomposition of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 927
Section CF
IP Interpolating Polynomial . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 931
LSMR Least Squares Minimizes Residuals . . . . . . . . . . . . . . . . . . . . . . . . . . . . 932
Section SAS
Version 2.30
Notation
M A: Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
MC [ A]ij: Matrix Components . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
CV v: Column Vector . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
CVC [ v]i: Column Vector Components . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
ZCV 0: Zero Column Vector . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
MRLSLS(A;b): Matrix Representation of a Linear System . . . . . . . . . . . . . . . . . . 29
AM [ Ajb]: Augmented Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
RO Ri$Rj,Ri,Ri+Rj: Row Operations . . . . . . . . . . . . . . . . . . . . . . . . 31
RREFA r,D,F: Reduced Row-Echelon Form Analysis . . . . . . . . . . . . . . . . . . . . . . 33
NSMN(A): Null Space of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73
IM Im: Identity Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 84
VSCV Cm: Vector Space of Column Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . 97
CVE u=v: Column Vector Equality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98
CVA u+v: Column Vector Addition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99
CVSM u: Column Vector Scalar Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . 99
SSVhSi: Span of a Set of Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131
CCCV u: Complex Conjugate of a Column Vector . . . . . . . . . . . . . . . . . . . . . . . . 191
IPhu;vi: Inner Product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 192
NVkvk: Norm of a Vector . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 195
SUV ei: Standard Unit Vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 197
VSM Mmn: Vector Space of Matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 207
ME A=B: Matrix Equality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 207
MA A+B: Matrix Addition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 208
MSM A: Matrix Scalar Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 208
ZMO: Zero Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 210
TM At: Transpose of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 210
CCM A: Complex Conjugate of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . 212
A A: Adjoint . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 214
MVP A u: Matrix-Vector Product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 223
MI A 1: Matrix Inverse . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 244
CSMC(A): Column Space of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 271
RSMR(A): Row Space of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 278
LNSL(A): Left Null Space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 293
D dim ( V): Dimension . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 391
NOM n(A): Nullity of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 397
ROM r(A): Rank of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 397
DS V=UW: Direct Sum . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 413
ELEM Ei;j,Ei(),Ei;j(): Elementary Matrix . . . . . . . . . . . . . . . . . . . . . . . . . 424
SM A(ijj): SubMatrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 428
xxxiii
xxxiv NOTATION
DMjAj, det (A): Determinant of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . 428
AME A(): Algebraic Multiplicity of an Eigenvalue . . . . . . . . . . . . . . . . . . . . . . 463
GME
A(): Geometric Multiplicity of an Eigenvalue . . . . . . . . . . . . . . . . . . . . . 463
LT T:U!V: Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 515
KLTK(T): Kernel of a Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . 545
RLTR(T): Range of a Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . 563
ROLT r(T): Rank of a Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . 588
NOLT n(T): Nullity of a Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . 588
VR B(w): Vector Representation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 603
MR MT
B;C: Matrix Representation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 615
JB Jn(): Jordan Block . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 687
GESGT(): Generalized Eigenspace . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 707
LTR TjU: Linear Transformation Restriction . . . . . . . . . . . . . . . . . . . . . . . . . . 711
IE T(): Index of an Eigenvalue . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 717
CNE =: Complex Number Equality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 758
CNA +: Complex Number Addition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 758
CNM : Complex Number Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . 758
CCN : Conjugate of a Complex Number . . . . . . . . . . . . . . . . . . . . . . . . . . . . 759
SETM x2S: Set Membership . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 761
SSET ST: Subset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 761
ES;: Empty Set . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 761
SE S=T: Set Equality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 762
CjSj: Cardinality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 762
SU S[T: Set Union . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 763
SI S\T: Set Intersection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 763
SC S: Set Complement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 763
T t(A): Trace . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 883
HP AB: Hadamard Product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 889
HID Jmn: Hadamard Identity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 890
HIbA: Hadamard Inverse . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 890
SRM A1=2: Square Root of a Matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 926
Version 2.30
Diagrams
DTSLS Decision Tree for Solving Linear Systems . . . . . . . . . . . . . . . . . . . . . . . . . 61
CSRST Column Space and Row Space Techniques . . . . . . . . . . . . . . . . . . . . . . . . 307
DLTA Denition of Linear Transformation, Additive . . . . . . . . . . . . . . . . . . . . . . 516
DLTM Denition of Linear Transformation, Multiplicative . . . . . . . . . . . . . . . . . . . 516
GLT General Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 520
NILT Non-Injective Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . 542
ILT Injective Linear Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 544
FTMR Fundamental Theorem of Matrix Representations . . . . . . . . . . . . . . . . . . . . 618
FTMRA Fundamental Theorem of Matrix Representations (Alternate) . . . . . . . . . . . . . 619
MRCLT Matrix Representation and Composition of Linear Transformations . . . . . . . . . . 625
xxxv
xxxvi DIAGRAMS
Version 2.30
Examples
Section WILA
TMP Trail Mix Packaging . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
Section SSLE
STNE Solving two (nonlinear) equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
NSE Notation for a system of equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
TTS Three typical systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
US Three equations, one solution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
IS Three equations, innitely many solutions . . . . . . . . . . . . . . . . . . . . . . . . 17
Section RREF
AM A matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
NSLE Notation for systems of linear equations . . . . . . . . . . . . . . . . . . . . . . . . . . 29
AMAA Augmented matrix for Archetype A . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
TREM Two row-equivalent matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
USR Three equations, one solution, reprised . . . . . . . . . . . . . . . . . . . . . . . . . . 32
RREF A matrix in reduced row-echelon form . . . . . . . . . . . . . . . . . . . . . . . . . . . 33
NRREF A matrix not in reduced row-echelon form . . . . . . . . . . . . . . . . . . . . . . . . 33
SAB Solutions for Archetype B . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39
SAA Solutions for Archetype A . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40
SAE Solutions for Archetype E . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
Section TSS
RREFN Reduced row-echelon form notation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
ISSI Describing innite solution sets, Archetype I . . . . . . . . . . . . . . . . . . . . . . . 56
FDV Free and dependent variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
CFV Counting free variables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60
OSGMD One solution gives many, Archetype D . . . . . . . . . . . . . . . . . . . . . . . . . . 61
Section HSE
AHSAC Archetype C as a homogeneous system . . . . . . . . . . . . . . . . . . . . . . . . . . 71
HUSAB Homogeneous, unique solution, Archetype B . . . . . . . . . . . . . . . . . . . . . . . 72
HISAA Homogeneous, innite solutions, Archetype A . . . . . . . . . . . . . . . . . . . . . . 72
HISAD Homogeneous, innite solutions, Archetype D . . . . . . . . . . . . . . . . . . . . . . 72
NSEAI Null space elements of Archetype I . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74
CNS1 Computing a null space, #1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74
CNS2 Computing a null space, #2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 75
Section NM
xxxvii
xxxviii EXAMPLES
S A singular matrix, Archetype A . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
NM A nonsingular matrix, Archetype B . . . . . . . . . . . . . . . . . . . . . . . . . . . . 84
IM An identity matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 84
SRR Singular matrix, row-reduced . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85
NSR Nonsingular matrix, row-reduced . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85
NSS Null space of a singular matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85
NSNM Null space of a nonsingular matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86
Section VO
VESE Vector equality for a system of equations . . . . . . . . . . . . . . . . . . . . . . . . . 98
VA Addition of two vectors in C4. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99
CVSM Scalar multiplication in C5. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100
Section LC
TLC Two linear combinations in C6. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 109
ABLC Archetype B as a linear combination . . . . . . . . . . . . . . . . . . . . . . . . . . . 110
AALC Archetype A as a linear combination . . . . . . . . . . . . . . . . . . . . . . . . . . . 111
VFSAD Vector form of solutions for Archetype D . . . . . . . . . . . . . . . . . . . . . . . . . 114
VFS Vector form of solutions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115
VFSAI Vector form of solutions for Archetype I . . . . . . . . . . . . . . . . . . . . . . . . . . 121
VFSAL Vector form of solutions for Archetype L . . . . . . . . . . . . . . . . . . . . . . . . . 122
PSHS Particular solutions, homogeneous solutions, Archetype D . . . . . . . . . . . . . . . . 125
Section SS
ABS A basic span . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131
SCAA Span of the columns of Archetype A . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133
SCAB Span of the columns of Archetype B . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135
SSNS Spanning set of a null space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 137
NSDS Null space directly as a span . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 138
SCAD Span of the columns of Archetype D . . . . . . . . . . . . . . . . . . . . . . . . . . . . 139
Section LI
LDS Linearly dependent set in C5. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153
LIS Linearly independent set in C5. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 154
LIHS Linearly independent, homogeneous system . . . . . . . . . . . . . . . . . . . . . . . . 155
LDHS Linearly dependent, homogeneous system . . . . . . . . . . . . . . . . . . . . . . . . . 156
LDRN Linearly dependent, r<n . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157
LLDS Large linearly dependent set in C4. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157
LDCAA Linearly dependent columns in Archetype A . . . . . . . . . . . . . . . . . . . . . . . 158
LICAB Linearly independent columns in Archetype B . . . . . . . . . . . . . . . . . . . . . . 158
LINSB Linear independence of null space basis . . . . . . . . . . . . . . . . . . . . . . . . . . 159
NSLIL Null space spanned by linearly independent set, Archetype L . . . . . . . . . . . . . . 161
Section LDS
RSC5 Reducing a span in C5. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 176
COV Casting out vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 177
RSC4 Reducing a span in C4. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 182
RES Reworking elements of a span . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 182
Section O
Version 2.30
EXAMPLES xxxix
CSIP Computing some inner products . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 192
CNSV Computing the norm of some vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . 195
TOV Two orthogonal vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 196
SUVOS Standard Unit Vectors are an Orthogonal Set . . . . . . . . . . . . . . . . . . . . . . 197
AOS An orthogonal set . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 197
GSTV Gram-Schmidt of three vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 200
ONTV Orthonormal set, three vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 201
ONFV Orthonormal set, four vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 202
Section MO
MA Addition of two matrices in M23. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 208
MSM Scalar multiplication in M32. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 208
TM Transpose of a 3 4 matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 210
SYM A symmetric 5 5 matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 211
CCM Complex conjugate of a matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 212
Section MM
MTV A matrix times a vector . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 223
MNSLE Matrix notation for systems of linear equations . . . . . . . . . . . . . . . . . . . . . . 224
MBC Money's best cities . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 224
PTM Product of two matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 226
MMNC Matrix multiplication is not commutative . . . . . . . . . . . . . . . . . . . . . . . . . 227
PTMEE Product of two matrices, entry-by-entry . . . . . . . . . . . . . . . . . . . . . . . . . . 228
Section MISLE
SABMI Solutions to Archetype B with a matrix inverse . . . . . . . . . . . . . . . . . . . . . 243
MWIAA A matrix without an inverse, Archetype A . . . . . . . . . . . . . . . . . . . . . . . . 244
MI Matrix inverse . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 245
CMI Computing a matrix inverse . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 247
CMIAB Computing a matrix inverse, Archetype B . . . . . . . . . . . . . . . . . . . . . . . . 249
Section MINM
UM3 Unitary matrix of size 3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 262
UPM Unitary permutation matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 262
OSMC Orthonormal set from matrix columns . . . . . . . . . . . . . . . . . . . . . . . . . . . 263
Section CRS
CSMCS Column space of a matrix and consistent systems . . . . . . . . . . . . . . . . . . . . 271
MCSM Membership in the column space of a matrix . . . . . . . . . . . . . . . . . . . . . . . 272
CSTW Column space, two ways . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 274
CSOCD Column space, original columns, Archetype D . . . . . . . . . . . . . . . . . . . . . . 275
CSAA Column space of Archetype A . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 276
CSAB Column space of Archetype B . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 276
RSAI Row space of Archetype I . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 278
RSREM Row spaces of two row-equivalent matrices . . . . . . . . . . . . . . . . . . . . . . . . 280
IAS Improving a span . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 281
CSROI Column space from row operations, Archetype I . . . . . . . . . . . . . . . . . . . . . 282
Section FS
LNS Left null space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 293
Version 2.30
xl EXAMPLES
CSANS Column space as null space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 294
SEEF Submatrices of extended echelon form . . . . . . . . . . . . . . . . . . . . . . . . . . . 297
FS1 Four subsets, #1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 303
FS2 Four subsets, #2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 304
FSAG Four subsets, Archetype G . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 305
Section VS
VSCV The vector space Cm. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 319
VSM The vector space of matrices, Mmn . . . . . . . . . . . . . . . . . . . . . . . . . . . . 319
VSP The vector space of polynomials, Pn. . . . . . . . . . . . . . . . . . . . . . . . . . . . 319
VSIS The vector space of innite sequences . . . . . . . . . . . . . . . . . . . . . . . . . . . 320
VSF The vector space of functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 321
VSS The singleton vector space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 321
CVS The crazy vector space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 322
PCVS Properties for the Crazy Vector Space . . . . . . . . . . . . . . . . . . . . . . . . . . . 326
Section S
SC3 A subspace of C3. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 333
SP4 A subspace of P4. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 335
NSC2Z A non-subspace in C2, zero vector . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 336
NSC2A A non-subspace in C2, additive closure . . . . . . . . . . . . . . . . . . . . . . . . . . 336
NSC2S A non-subspace in C2, scalar multiplication closure . . . . . . . . . . . . . . . . . . . 337
RSNS Recasting a subspace as a null space . . . . . . . . . . . . . . . . . . . . . . . . . . . . 338
LCM A linear combination of matrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 338
SSP Span of a set of polynomials . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 340
SM32 A subspace of M32. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 341
Section LISS
LIP4 Linear independence in P4. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 351
LIM32 Linear independence in M32. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 353
LIC Linearly independent set in the crazy vector space . . . . . . . . . . . . . . . . . . . . 355
SSP4 Spanning set in P4. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 356
SSM22 Spanning set in M22. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 357
SSC Spanning set in the crazy vector space . . . . . . . . . . . . . . . . . . . . . . . . . . 358
AVR A vector representation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 359
Section B
BP Bases for Pn. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 372
BM A basis for the vector space of matrices . . . . . . . . . . . . . . . . . . . . . . . . . . 372
BSP4 A basis for a subspace of P4. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 372
BSM22 A basis for a subspace of M22. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 373
BC Basis for the crazy vector space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 374
RSB Row space basis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 374
RS Reducing a span . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 375
CABAK Columns as Basis, Archetype K . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 376
CROB4 Coordinatization relative to an orthonormal basis, C4. . . . . . . . . . . . . . . . . . 378
CROB3 Coordinatization relative to an orthonormal basis, C3. . . . . . . . . . . . . . . . . . 379
Section D
LDP4 Linearly dependent set in P4. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 394
Version 2.30
EXAMPLES xli
DSM22 Dimension of a subspace of M22. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 395
DSP4 Dimension of a subspace of P4. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 396
DC Dimension of the crazy vector space . . . . . . . . . . . . . . . . . . . . . . . . . . . . 396
VSPUD Vector space of polynomials with unbounded degree . . . . . . . . . . . . . . . . . . . 396
RNM Rank and nullity of a matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 397
RNSM Rank and nullity of a square matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . 398
Section PD
BPR Bases for Pn, reprised . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 408
BDM22 Basis by dimension in M22. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 409
SVP4 Sets of vectors in P4. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 409
RRTI Rank, rank of transpose, Archetype I . . . . . . . . . . . . . . . . . . . . . . . . . . . 411
SDS Simple direct sum . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 413
Section DM
EMRO Elementary matrices and row operations . . . . . . . . . . . . . . . . . . . . . . . . . 424
SS Some submatrices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 428
D33M Determinant of a 3 3 matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 428
TCSD Two computations, same determinant . . . . . . . . . . . . . . . . . . . . . . . . . . . 432
DUTM Determinant of an upper triangular matrix . . . . . . . . . . . . . . . . . . . . . . . . 432
Section PDM
DRO Determinant by row operations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 442
ZNDAB Zero and nonzero determinant, Archetypes A and B . . . . . . . . . . . . . . . . . . . 446
Section EE
SEE Some eigenvalues and eigenvectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 453
PM Polynomial of a matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 455
CAEHW Computing an eigenvalue the hard way . . . . . . . . . . . . . . . . . . . . . . . . . . 458
CPMS3 Characteristic polynomial of a matrix, size 3 . . . . . . . . . . . . . . . . . . . . . . . 460
EMS3 Eigenvalues of a matrix, size 3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 461
ESMS3 Eigenspaces of a matrix, size 3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 462
EMMS4 Eigenvalue multiplicities, matrix of size 4 . . . . . . . . . . . . . . . . . . . . . . . . . 463
ESMS4 Eigenvalues, symmetric matrix of size 4 . . . . . . . . . . . . . . . . . . . . . . . . . . 464
HMEM5 High multiplicity eigenvalues, matrix of size 5 . . . . . . . . . . . . . . . . . . . . . . 465
CEMS6 Complex eigenvalues, matrix of size 6 . . . . . . . . . . . . . . . . . . . . . . . . . . . 466
DEMS5 Distinct eigenvalues, matrix of size 5 . . . . . . . . . . . . . . . . . . . . . . . . . . . 468
Section PEE
BDE Building desired eigenvalues . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 482
Section SD
SMS5 Similar matrices of size 5 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 493
SMS3 Similar matrices of size 3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 494
EENS Equal eigenvalues, not similar . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 496
DAB Diagonalization of Archetype B . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 496
DMS3 Diagonalizing a matrix of size 3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 498
NDMS4 A non-diagonalizable matrix of size 4 . . . . . . . . . . . . . . . . . . . . . . . . . . . 501
DEHD Distinct eigenvalues, hence diagonalizable . . . . . . . . . . . . . . . . . . . . . . . . . 501
HPDM High power of a diagonalizable matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . 502
Version 2.30
xlii EXAMPLES
FSCF Fibonacci sequence, closed form . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 503
Section LT
ALT A linear transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 516
NLT Not a linear transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 517
LTPM Linear transformation, polynomials to matrices . . . . . . . . . . . . . . . . . . . . . 518
LTPP Linear transformation, polynomials to polynomials . . . . . . . . . . . . . . . . . . . 518
LTM Linear transformation from a matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . 520
MFLT Matrix from a linear transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 522
MOLT Matrix of a linear transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 524
LTDB1 Linear transformation dened on a basis . . . . . . . . . . . . . . . . . . . . . . . . . 526
LTDB2 Linear transformation dened on a basis . . . . . . . . . . . . . . . . . . . . . . . . . 527
LTDB3 Linear transformation dened on a basis . . . . . . . . . . . . . . . . . . . . . . . . . 527
SPIAS Sample pre-images, Archetype S . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 528
STLT Sum of two linear transformations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 531
SMLT Scalar multiple of a linear transformation . . . . . . . . . . . . . . . . . . . . . . . . . 532
CTLT Composition of two linear transformations . . . . . . . . . . . . . . . . . . . . . . . . 533
Section ILT
NIAQ Not injective, Archetype Q . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 541
IAR Injective, Archetype R . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 542
IAV Injective, Archetype V . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 544
NKAO Nontrivial kernel, Archetype O . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 545
TKAP Trivial kernel, Archetype P . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 546
NIAQR Not injective, Archetype Q, revisited . . . . . . . . . . . . . . . . . . . . . . . . . . . 548
NIAO Not injective, Archetype O . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 549
IAP Injective, Archetype P . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 549
NIDAU Not injective by dimension, Archetype U . . . . . . . . . . . . . . . . . . . . . . . . . 550
Section SLT
NSAQ Not surjective, Archetype Q . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 559
SAR Surjective, Archetype R . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 560
SAV Surjective, Archetype V . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 561
RAO Range, Archetype O . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 563
FRAN Full range, Archetype N . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 564
NSAQR Not surjective, Archetype Q, revisited . . . . . . . . . . . . . . . . . . . . . . . . . . . 566
NSAO Not surjective, Archetype O . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 566
SAN Surjective, Archetype N . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 567
BRLT A basis for the range of a linear transformation . . . . . . . . . . . . . . . . . . . . . 568
NSDAT Not surjective by dimension, Archetype T . . . . . . . . . . . . . . . . . . . . . . . . 569
Section IVLT
AIVLT An invertible linear transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 579
ANILT A non-invertible linear transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . 580
CIVLT Computing the Inverse of a Linear Transformations . . . . . . . . . . . . . . . . . . . 583
IVSAV Isomorphic vector spaces, Archetype V . . . . . . . . . . . . . . . . . . . . . . . . . . 586
Section VR
VRC4 Vector representation in C4. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 604
VRP2 Vector representations in P2. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 606
Version 2.30
EXAMPLES xliii
TIVS Two isomorphic vector spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 609
CVSR Crazy vector space revealed . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 609
ASC A subspace characterized . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 609
MIVS Multiple isomorphic vector spaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 609
CP2 Coordinatizing in P2. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 610
CM32 Coordinatization in M32. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 611
Section MR
OLTTR One linear transformation, three representations . . . . . . . . . . . . . . . . . . . . . 615
ALTMM A linear transformation as matrix multiplication . . . . . . . . . . . . . . . . . . . . . 619
MPMR Matrix product of matrix representations . . . . . . . . . . . . . . . . . . . . . . . . . 622
KVMR Kernel via matrix representation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 626
RVMR Range via matrix representation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 629
ILTVR Inverse of a linear transformation via a representation . . . . . . . . . . . . . . . . . . 632
Section CB
ELTBM Eigenvectors of linear transformation between matrices . . . . . . . . . . . . . . . . . 647
ELTBP Eigenvectors of linear transformation between polynomials . . . . . . . . . . . . . . . 648
CBP Change of basis with polynomials . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 649
CBCV Change of basis with column vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . 652
MRCM Matrix representations and change-of-basis matrices . . . . . . . . . . . . . . . . . . . 654
MRBE Matrix representation with basis of eigenvectors . . . . . . . . . . . . . . . . . . . . . 657
ELTT Eigenvectors of a linear transformation, twice . . . . . . . . . . . . . . . . . . . . . . 660
CELT Complex eigenvectors of a linear transformation . . . . . . . . . . . . . . . . . . . . . 665
Section OD
ANM A normal matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 680
Section NLT
NM64 Nilpotent matrix, size 6, index 4 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 685
NM62 Nilpotent matrix, size 6, index 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 686
JB4 Jordan block, size 4 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 687
NJB5 Nilpotent Jordan block, size 5 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 688
NM83 Nilpotent matrix, size 8, index 3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 689
KPNLT Kernels of powers of a nilpotent linear transformation . . . . . . . . . . . . . . . . . . 693
CFNLT Canonical form for a nilpotent linear transformation . . . . . . . . . . . . . . . . . . . 698
Section IS
TIS Two invariant subspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 703
EIS Eigenspaces as invariant subspaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 705
ISJB Invariant subspaces and Jordan blocks . . . . . . . . . . . . . . . . . . . . . . . . . . 706
GE4 Generalized eigenspaces, dimension 4 domain . . . . . . . . . . . . . . . . . . . . . . . 708
GE6 Generalized eigenspaces, dimension 6 domain . . . . . . . . . . . . . . . . . . . . . . . 709
LTRGE Linear transformation restriction on generalized eigenspace . . . . . . . . . . . . . . . 711
ISMR4 Invariant subspaces, matrix representation, dimension 4 domain . . . . . . . . . . . . 714
ISMR6 Invariant subspaces, matrix representation, dimension 6 domain . . . . . . . . . . . . 715
GENR6 Generalized eigenspaces and nilpotent restrictions, dimension 6 domain . . . . . . . . 717
Section JCF
JCF10 Jordan canonical form, size 10 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 729
Version 2.30
xliv EXAMPLES
Section CNO
ACN Arithmetic of complex numbers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 757
CSCN Conjugate of some complex numbers . . . . . . . . . . . . . . . . . . . . . . . . . . . 759
MSCN Modulus of some complex numbers . . . . . . . . . . . . . . . . . . . . . . . . . . . . 760
Section SET
SETM Set membership . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 761
SSET Subset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 761
CS Cardinality and Size . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 762
SU Set union . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 763
SI Set intersection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 763
SC Set complement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 764
Section PT
Section F
IM11 Integers mod 11 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 875
VSIM5 Vector space over integers mod 5 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 875
SM2Z7 Symmetric matrices of size 2 over Z7. . . . . . . . . . . . . . . . . . . . . . . . . . . 876
FF8 Finite eld of size 8 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 876
Section T
Section HP
HP Hadamard Product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 889
Section VM
VM4 Vandermonde matrix of size 4 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 895
Section PSM
Section ROD
ROD2 Rank one decomposition, size 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 905
ROD4 Rank one decomposition, size 4 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 906
Section TD
TD4 Triangular decomposition, size 4 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 911
TDSSE Triangular decomposition solves a system of equations . . . . . . . . . . . . . . . . . 912
TDEE6 Triangular decomposition, entry by entry, size 6 . . . . . . . . . . . . . . . . . . . . . 915
Section SVD
Section SR
Section POD
Section CF
PTFP Polynomial through ve points . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 931
Section SAS
SS6W Sharing a secret 6 ways . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 938
Version 2.30
Preface
This textbook is designed to teach the university mathematics student the basics of linear algebra and
the techniques of formal mathematics. There are no prerequisites other than ordinary algebra, but it is
probably best used by a student who has the \mathematical maturity" of a sophomore or junior. The text
has two goals: to teach the fundamental concepts and techniques of matrix algebra and abstract vector
spaces, and to teach the techniques associated with understanding the denitions and theorems forming
a coherent area of mathematics. So there is an emphasis on worked examples of nontrivial size and on
proving theorems carefully.
This book is copyrighted. This means that governments have granted the author a monopoly | the
exclusive right to control the making of copies and derivative works for many years (too many years in
some cases). It also gives others limited rights, generally referred to as \fair use," such as the right to
quote sections in a review without seeking permission. However, the author licenses this book to anyone
under the terms of the GNU Free Documentation License (GFDL), which gives you more rights than most
copyrights (see Appendix GFDL [865]). Loosely speaking, you may make as many copies as you like at no
cost, and you may distribute these unmodied copies if you please. You may modify the book for your own
use. The catch is that if you make modications and you distribute the modied version, or make use of
portions in excess of fair use in another work, then you must also license the new work with the GFDL. So
the book has lots of inherent freedom, and no one is allowed to distribute a derivative work that restricts
these freedoms. (See the license itself in the appendix for the exact details of the additional rights you
have been given.)
Notice that initially most people are struck by the notion that this book is free (the French would say
gratuit , at no cost). And it is. However, it is more important that the book has freedom (the French
would say libert e , liberty). It will never go \out of print" nor will there ever be trivial updates designed
only to frustrate the used book market. Those considering teaching a course with this book can examine
it thoroughly in advance. Adding new exercises or new sections has been purposely made very easy, and
the hope is that others will contribute these modications back for incorporation into the book, for the
benet of all.
Depending on how you received your copy, you may want to check for the latest version (and other
news) at http://linear.ups.edu/ .
Topics The rst half of this text (through Chapter M [207]) is basically a course in matrix algebra,
though the foundation of some more advanced ideas is also being formed in these early sections. Vectors
are presented exclusively as column vectors (since we also have the typographic freedom to avoid writing
a column vector inline as the transpose of a row vector), and linear combinations are presented very early.
Spans, null spaces, column spaces and row spaces are also presented early, simply as sets, saving most of
their vector space properties for later, so they are familiar objects before being scrutinized carefully.
You cannot do everything early, so in particular matrix multiplication comes later than usual. However,
with a denition built on linear combinations of column vectors, it should seem more natural than the
more frequent denition using dot products of rows with columns. And this delay emphasizes that linear
algebra is built upon vector addition and scalar multiplication. Of course, matrix inverses must wait for
matrix multiplication, but this does not prevent nonsingular matrices from occurring sooner. Vector space
xlv
xlvi PREFACE
properties are hinted at when vector and matrix operations are rst dened, but the notion of a vector
space is saved for a more axiomatic treatment later (Chapter VS [317]). Once bases and dimension have
been explored in the context of vector spaces, linear transformations and their matrix representation follow.
The goal of the book is to go as far as Jordan canonical form in the Core (Part C [3]), with less central
topics collected in the Topics (Part T [873]). A third part contains contributed applications (Part A [931]),
with notation and theorems integrated with the earlier two parts.
Linear algebra is an ideal subject for the novice mathematics student to learn how to develop a topic
precisely, with all the rigor mathematics requires. Unfortunately, much of this rigor seems to have escaped
the standard calculus curriculum, so for many university students this is their rst exposure to careful
denitions and theorems, and the expectation that they fully understand them, to say nothing of the
expectation that they become procient in formulating their own proofs. We have tried to make this text
as helpful as possible with this transition. Every denition is stated carefully, set apart from the text.
Likewise, every theorem is carefully stated, and almost every one has a complete proof. Theorems usually
have just one conclusion, so they can be referenced precisely later. Denitions and theorems are cataloged
in order of their appearance in the front of the book (Denitions [xi], Theorems [xiii]), and alphabetical
order in the index at the back. Along the way, there are discussions of some more important ideas relating
to formulating proofs (Proof Techniques [ ??]), which is part advice and part logic.
Origin and History This book is the result of the con
uence of several related events and trends.
At the University of Puget Sound we teach a one-semester, post-calculus linear algebra course to
students majoring in mathematics, computer science, physics, chemistry and economics. Between
January 1986 and June 2002, I taught this course seventeen times. For the Spring 2003 semester, I
elected to convert my course notes to an electronic form so that it would be easier to incorporate the
inevitable and nearly-constant revisions. Central to my new notes was a collection of stock examples
that would be used repeatedly to illustrate new concepts. (These would become the Archetypes,
Appendix A [777].) It was only a short leap to then decide to distribute copies of these notes and
examples to the students in the two sections of this course. As the semester wore on, the notes began
to look less like notes and more like a textbook.
I used the notes again in the Fall 2003 semester for a single section of the course. Simultaneously, the
textbook I was using came out in a fth edition. A new chapter was added toward the start of the
book, and a few additional exercises were added in other chapters. This demanded the annoyance
of reworking my notes and list of suggested exercises to conform with the changed numbering of the
chapters and exercises. I had an almost identical experience with the third course I was teaching
that semester. I also learned that in the next academic year I would be teaching a course where my
textbook of choice had gone out of print. I felt there had to be a better alternative to having the
organization of my courses bueted by the economics of traditional textbook publishing.
I had used T EX and the Internet for many years, so there was little to stand in the way of typesetting,
distributing and \marketing" a free book. With recreational and professional interests in software
development, I had long been fascinated by the open-source software movement, as exemplied by
the success of GNU and Linux, though public-domain T EX might also deserve mention. Obviously,
this book is an attempt to carry over that model of creative endeavor to textbook publishing.
As a sabbatical project during the Spring 2004 semester, I embarked on the current project of creating
a freely-distributable linear algebra textbook. (Notice the implied nancial support of the University
of Puget Sound to this project.) Most of the material was written from scratch since changes in
notation and approach made much of my notes of little use. By August 2004 I had written half the
material necessary for our Math 232 course. The remaining half was written during the Fall 2004
semester as I taught another two sections of Math 232.
Version 2.30
PREFACE xlvii
While in early 2005 the book was complete enough to build a course around and Version 1.0 was
released. Work has continued since, lling out the narrative, exercises and supplements.
However, much of my motivation for writing this book is captured by the sentiments expressed by H.M.
Cundy and A.P. Rollet in their Preface to the First Edition of Mathematical Models (1952), especially the
nal sentence,
This book was born in the classroom, and arose from the spontaneous interest of a Mathematical
Sixth in the construction of simple models. A desire to show that even in mathematics one could
have fun led to an exhibition of the results and attracted considerable attention throughout the
school. Since then the Sherborne collection has grown, ideas have come from many sources, and
widespread interest has been shown. It seems therefore desirable to give permanent form to the
lessons of experience so that others can benet by them and be encouraged to undertake similar
work.
How To Use This Book Chapters, Theorems, etc. are not numbered in this book, but are instead
referenced by acronyms. This means that Theorem XYZ will always be Theorem XYZ, no matter if
new sections are added, or if an individual decides to remove certain other sections. Within sections,
the subsections are acronyms that begin with the acronym of the section. So Subsection XYZ.AB is the
subsection AB in Section XYZ. Acronyms are unique within their type, so for example there is just one
Denition B [371], but there is also a Section B [371]. At rst, all the letters
ying around may be confusing,
but with time, you will begin to recognize the more important ones on sight. Furthermore, there are lists
of theorems, examples, etc. in the front of the book, and an index that contains every acronym. If you
are reading this in an electronic version (PDF or XML), you will see that all of the cross-references are
hyperlinks, allowing you to click to a denition or example, and then use the back button to return. In
printed versions, you must rely on the page numbers. However, note that page numbers are not permanent!
Dierent editions, dierent margins, or dierent sized paper will aect what content is on each page. And
in time, the addition of new material will aect the page numbering.
Chapter divisions are not critical to the organization of the book, as Sections are the main organizational
unit. Sections are designed to be the subject of a single lecture or classroom session, though there is
frequently more material than can be discussed and illustrated in a fty-minute session. Consequently,
the instructor will need to be selective about which topics to illustrate with other examples and which
topics to leave to the student's reading. Many of the examples are meant to be large, such as using ve
or six variables in a system of equations, so the instructor may just want to \walk" a class through these
examples. The book has been written with the idea that some may work through it independently, so the
hope is that students can learn some of the more mechanical ideas on their own.
The highest level division of the book is the three Parts: Core, Topics, Applications (Part C [3], Part T
[873], Part A [931]). The Core is meant to carefully describe the basic ideas required of a rst exposure to
linear algebra. In the nal sections of the Core, one should ask the question: which previous Sections could
be removed without destroying the logical development of the subject? Hopefully, the answer is \none."
The goal of the book is to nish the Core with a very general representation of a linear transformation
(Jordan canonical form, Section JCF [721]). Of course, there will not be universal agreement on what
should, or should not, constitute the Core, but the main idea is to limit it to about forty sections. Topics
(Part T [873]) is meant to contain those subjects that are important in linear algebra, and which would
make protable detours from the Core for those interested in pursuing them. Applications (Part A [931])
should illustrate the power and widespread applicability of linear algebra to as many elds as possible. The
Archetypes (Appendix A [777]) cover many of the computational aspects of systems of linear equations,
matrices and linear transformations. The student should consult them often, and this is encouraged by
exercises that simply suggest the right properties to examine at the right time. But what is more important,
this a repository that contains enough variety to provide abundant examples of key theorems, while also
providing counterexamples to hypotheses or converses of theorems. The summary table at the start of this
appendix should be especially useful.
Version 2.30
xlviii PREFACE
I require my students to read each Section prior to the day's discussion on that section. For some
students this is a novel idea, but at the end of the semester a few always report on the benets, both for
this course and other courses where they have adopted the habit. To make good on this requirement, each
section contains three Reading Questions. These sometimes only require parroting back a key denition or
theorem, or they require performing a small example of a key computation, or they ask for musings on key
ideas or new relationships between old ideas. Answers are emailed to me the evening before the lecture.
Given the
avor and purpose of these questions, including solutions seems foolish.
Every chapter of Part C [3] ends with \Annotated Acronyms", a short list of critical theorems or
denitions from that chapter. There are a variety of reasons for any one of these to have been chosen,
and reading the short paragraphs after some of these might provide insight into the possibilities. An
end-of-chapter review might usefully incorporate a close reading of these lists.
Formulating interesting and eective exercises is as dicult, or more so, than building a narrative.
But it is the place where a student really learns the material. As such, for the student's benet, complete
solutions should be given. As the list of exercises expands, the amount with solutions should similarly
expand. Exercises and their solutions are referenced with a section name, followed by a dot, then a
letter (C,M, or T) and a number. The letter `C' indicates a problem that is mostly computational in
nature, while the letter `T' indicates a problem that is more theoretical in nature. A problem with a letter
`M' is somewhere in between (middle, mid-level, median, middling), probably a mix of computation and
applications of theorems. So Solution MO.T13 [221] is a solution to an exercise in Section MO [207] that
is theoretical in nature. The number `13' has no intrinsic meaning.
More on Freedom This book is freely-distributable under the terms of the GFDL, along with the
underlying T EX code from which the book is built. This arrangement provides many benets unavailable
with traditional texts.
No cost, or low cost, to students. With no physical vessel (i.e. paper, binding), no transportation
costs (Internet bandwidth being a negligible cost) and no marketing costs (evaluation and desk copies
are free to all), anyone with an Internet connection can obtain it, and a teacher could make available
paper copies in sucient quantities for a class. The cost to print a copy is not insignicant, but is
just a fraction of the cost of a traditional textbook when printing is handled by a print-on-demand
service over the Internet. Students will not feel the need to sell back their book (nor should there be
much of a market for used copies), and in future years can even pick up a newer edition freely.
Electronic versions of the book contain extensive hyperlinks. Specically, most logical steps in proofs
and examples include links back to the previous denitions or theorems that support that step. With
whatever viewer you might be using (web browser, PDF reader) the \back" button can then return
you to the middle of the proof you were studying. So even if you are reading a physical copy of this
book, you can benet from also working with an electronic version.
A traditional book, which the publisher is unwilling to distribute in an easily-copied electronic form,
cannot oer this very intuitive and
exible approach to learning mathematics.
The book will not go out of print. No matter what, a teacher can maintain their own copy and use the
book for as many years as they desire. Further, the naming schemes for chapters, sections, theorems,
etc. is designed so that the addition of new material will not break any course syllabi or assignment
list.
With many eyes reading the book and with frequent postings of updates, the reliability should become
very high. Please report any errors you nd that persist into the latest version.
For those with a working installation of the popular typesetting program T EX, the book has been
designed so that it can be customized. Page layouts, presence of exercises, solutions, sections or chap-
ters can all be easily controlled. Furthermore, many variants of mathematical notation are achieved
Version 2.30
PREFACE xlix
via T EX macros. So by changing a single macro, one's favorite notation can be re
ected throughout
the text. For example, every transpose of a matrix is coded in the source as \transpose{A} , which
when printed will yield At. However by changing the denition of \transpose{ } , any desired al-
ternative notation (superscript t, superscript T, superscript prime) will then appear throughout the
text instead.
The book has also been designed to make it easy for others to contribute material. Would you like
to see a section on symmetric bilinear forms? Consider writing one and contributing it to one of the
Topics chapters. Should there be more exercises about the null space of a matrix? Send me some.
Historical Notes? Contact me, and we will see about adding those in also.
You have no legal obligation to pay for this book. It has been licensed with no expectation that you
pay for it. You do not even have a moral obligation to pay for the book. Thomas Jeerson (1743 {
1826), the author of the United States Declaration of Independence, wrote,
If nature has made any one thing less susceptible than all others of exclusive property, it
is the action of the thinking power called an idea, which an individual may exclusively
possess as long as he keeps it to himself; but the moment it is divulged, it forces itself into
the possession of every one, and the receiver cannot dispossess himself of it. Its peculiar
character, too, is that no one possesses the less, because every other possesses the whole of it.
He who receives an idea from me, receives instruction himself without lessening mine; as he
who lights his taper at mine, receives light without darkening me. That ideas should freely
spread from one to another over the globe, for the moral and mutual instruction of man,
and improvement of his condition, seems to have been peculiarly and benevolently designed
by nature, when she made them, like re, expansible over all space, without lessening their
density in any point, and like the air in which we breathe, move, and have our physical
being, incapable of connement or exclusive appropriation.
Letter to Isaac McPherson
August 13, 1813
However, if you feel a royalty is due the author, or if you would like to encourage the author, or if you
wish to show others that this approach to textbook publishing can also bring nancial compensation,
then donations are gratefully received. Moreover, non-nancial forms of help can often be even more
valuable. A simple note of encouragement, submitting a report of an error, or contributing some
exercises or perhaps an entire section for the Topics or Applications are all important ways you can
acknowledge the freedoms accorded to this work by the copyright holder and other contributors.
Conclusion Foremost, I hope that students nd their time spent with this book protable. I hope that
instructors nd it
exible enough to t the needs of their course. And I hope that everyone will send me
their comments and suggestions, and also consider the myriad ways they can help (as listed on the book's
website at http://linear.ups.edu ).
Robert A. Beezer
Tacoma, Washington
July 2008
Version 2.30
l PREFACE
Version 2.30
Acknowledgements
Many people have helped to make this book, and its freedoms, possible.
First, the time to create, edit and distribute the book has been provided implicitly and explicitly by
the University of Puget Sound. A sabbatical leave Spring 2004 and a course release in Spring 2007 are two
obvious examples of explicit support. The latter was provided by support from the Lind-VanEnkevort Fund.
The university has also provided clerical support, computer hardware, network servers and bandwidth.
Thanks to Dean Kris Bartanen and the chair of the Mathematics and Computer Science Department,
Professor Martin Jackson, for their support, encouragement and
exibility.
My colleagues in the Mathematics and Computer Science Department have graciously taught our
introductory linear algebra course using preliminary versions and have provided valuable suggestions that
have improved the book immeasurably. Thanks to Professor Martin Jackson (v0.30), Professor David Scott
(v0.70) and Professor Bryan Smith (v0.70, 0.80, v1.00).
University of Puget Sound librarians Lori Ricigliano, Elizabeth Knight and Jeanne Kimura provided
valuable advice on production, and interesting conversations about copyrights.
Many aspects of the book have been in
uenced by insightful questions and creative suggestions from
the students who have labored through the book in our courses. For example, the
ashcards with theorems
and denitions are a direct result of a student suggestion. I will single out a handful of students have been
especially adept at nding and reporting mathematically signicant typographical errors: Jake Linenthal,
Christie Su, Kim Le, Sarah McQuate, Andy Zimmer, Travis Osborne, Andrew Tapay, Mark Shoemaker,
Tasha Underhill, Tim Zitzer, Elizabeth Million, and Steve Caneld.
I have tried to be as original as possible in the organization and presentation of this beautiful subject.
However, I have been in
uenced by many years of teaching from another excellent textbook, Introduction
to Linear Algebra by L.W. Johnson, R.D. Reiss and J.T. Arnold. When I have needed inspiration for
the correct approach to particularly important proofs, I have learned to eventually consult two other
textbooks. Sheldon Axler's Linear Algebra Done Right is a highly original exposition, while Ben Noble's
Applied Linear Algebra frequently strikes just the right note between rigor and intuition. Noble's excellent
book is highly recommended, even though its publication dates to 1969.
Conversion to various electronic formats have greatly depended on assistance from: Eitan Gurari,
author of the powerful L ATEX translator, tex4ht ; Davide Cervone, author of jsMath ; and Carl Witty, who
advised and tested the Sony Reader format. Thanks to these individuals for their critical assistance.
General support and encouragement of free and aordable textbooks, in addition to specic promotion
of this text, was provided by Nicole Allen, Textbook Advocate at Student Public Interest Research Groups.
Nicole was an early consumer of this material, back when it looked more like lecture notes than a textbook.
Finally, in every possible case, the production and distribution of this book has been accomplished with
open-source software. The range of individuals and projects is far too great to pretend to list them all.
The book's web site will someday maintain pointers to as many of these projects as possible.
li
lii ACKNOWLEDGEMENTS
Version 2.30
Part C
Core
Chapter SLE
Systems of Linear Equations
We will motivate our study of linear algebra by studying solutions to systems of linear equations. While
the focus of this chapter is on the practical matter of how to nd, and describe, these solutions, we will
also be setting ourselves up for more theoretical ideas that will appear later.
Section WILA
What is Linear Algebra?
Subsection LA
\Linear" + \Algebra"
The subject of linear algebra can be partially explained by the meaning of the two terms comprising
the title. \Linear" is a term you will appreciate better at the end of this course, and indeed, attaining
this appreciation could be taken as one of the primary goals of this course. However for now, you can
understand it to mean anything that is \straight" or \
at." For example in the xy-plane you might be
accustomed to describing straight lines (is there any other kind?) as the set of solutions to an equation
of the form y=mx+b, where the slope mand they-interceptbare constants that together describe
the line. In multivariate calculus, you may have discussed planes. Living in three dimensions, with
coordinates described by triples ( x; y; z ), they can be described as the set of solutions to equations of the
formax+by+cz=d, wherea; b; c; d are constants that together determine the plane. While we might
describe planes as \
at," lines in three dimensions might be described as \straight." From a multivariate
calculus course you will recall that lines are sets of points described by equations such as x= 3t 4,
y= 7t+ 2,z= 9t, wheretis a parameter that can take on any value.
Another view of this notion of \
atness" is to recognize that the sets of points just described are solutions
to equations of a relatively simple form. These equations involve addition and multiplication only. We
will have a need for subtraction, and occasionally we will divide, but mostly you can describe \linear"
equations as involving only addition and multiplication. Here are some examples of typical equations we
will see in the next few sections:
2x+ 3y 4z= 13 4 x1+ 5x2 x3+x4+x5= 0 9 a 2b+ 7c+ 2d= 7
What we will not see are equations like:
xy+ 5yz= 13 x1+x3
2=x4 x3x4x2
5= 0 tan( ab) + log(c d) = 7
The exception will be that we will on occasion need to take a square root.
3
4 Section WILA What is Linear Algebra?
You have probably heard the word \algebra" frequently in your mathematical preparation for this
course. Most likely, you have spent a good ten to fteen years learning the algebra of the real numbers,
along with some introduction to the very similar algebra of complex numbers (see Section CNO [757]).
However, there are many new algebras to learn and use, and likely linear algebra will be your second
algebra. Like learning a second language, the necessary adjustments can be challenging at times, but the
rewards are many. And it will make learning your third and fourth algebras even easier. Perhaps you have
heard of \groups" and \rings" (or maybe you have studied them already), which are excellent examples of
other algebras with very interesting properties and applications. In any event, prepare yourself to learn a
new algebra and realize that some of the old rules you used for the real numbers may no longer apply to
this newalgebra you will be learning!
The brief discussion above about lines and planes suggests that linear algebra has an inherently geomet-
ric nature, and this is true. Examples in two and three dimensions can be used to provide valuable insight
into important concepts of this course. However, much of the power of linear algebra will be the ability to
work with \
at" or \straight" objects in higher dimensions, without concerning ourselves with visualizing
the situation. While much of our intuition will come from examples in two and three dimensions, we will
maintain an algebraic approach to the subject, with the geometry being secondary. Others may wish to
switch this emphasis around, and that can lead to a very fruitful and benecial course, but here and now
we are laying our bias bare.
Subsection AA
An Application
We conclude this section with a rather involved example that will highlight some of the power and tech-
niques of linear algebra. Work through all of the details with pencil and paper, until you believe all the
assertions made. However, in this introductory example, do not concern yourself with how some of the
results are obtained or how you might be expected to solve a similar problem. We will come back to
this example later and expose some of the techniques used and properties exploited. For now, use your
background in mathematics to convince yourself that everything said here really is correct.
Example TMP
Trail Mix Packaging
Suppose you are the production manager at a food-packaging plant and one of your product lines is trail
mix, a healthy snack popular with hikers and backpackers, containing raisins, peanuts and hard-shelled
chocolate pieces. By adjusting the mix of these three ingredients, you are able to sell three varieties of this
item. The fancy version is sold in half-kilogram packages at outdoor supply stores and has more chocolate
and fewer raisins, thus commanding a higher price. The standard version is sold in one kilogram packages
in grocery stores and gas station mini-markets. Since the standard version has roughly equal amounts of
each ingredient, it is not as expensive as the fancy version. Finally, a bulk version is sold in bins at grocery
stores for consumers to load into plastic bags in amounts of their choosing. To appeal to the shoppers
that like bulk items for their economy and healthfulness, this mix has many more raisins (at the expense
of chocolate) and therefore sells for less.
Your production facilities have limited storage space and early each morning you are able to receive
and store 380 kilograms of raisins, 500 kilograms of peanuts and 620 kilograms of chocolate pieces. As
production manager, one of your most important duties is to decide how much of each version of trail mix
to make every day. Clearly, you can have up to 1500 kilograms of raw ingredients available each day, so
to be the most productive you will likely produce 1500 kilograms of trail mix each day. Also, you would
prefer not to have any ingredients leftover each day, so that your nal product is as fresh as possible and
so that you can receive the maximum delivery the next morning. But how should these ingredients be
allocated to the mixing of the bulk, standard and fancy versions?
Version 2.30
Subsection WILA.AA An Application 5
First, we need a little more information about the mixes. Workers mix the ingredients in 15 kilogram
batches, and each row of the table below gives a recipe for a 15 kilogram batch. There is some additional
information on the costs of the ingredients and the price the manufacturer can charge for the dierent
versions of the trail mix.
Raisins Peanuts Chocolate Cost Sale Price
(kg/batch) (kg/batch) (kg/batch) ($/kg) ($/kg)
Bulk 7 6 2 3.69 4.99
Standard 6 4 5 3.86 5.50
Fancy 2 5 8 4.45 6.50
Storage (kg) 380 500 620
Cost ($/kg) 2.55 4.65 4.80
As production manager, it is important to realize that you only have three decisions to make | the amount
of bulk mix to make, the amount of standard mix to make and the amount of fancy mix to make. Everything
else is beyond your control or is handled by another department within the company. Principally, you are
also limited by the amount of raw ingredients you can store each day. Let us denote the amount of each
mix to produce each day, measured in kilograms, by the variable quantities b,sandf. Your production
schedule can be described as values of b,sandfthat do several things. First, we cannot make negative
quantities of each mix, so
b0 s0 f0
Second, if we want to consume all of our ingredients each day, the storage capacities lead to three (linear)
equations, one for each ingredient,
7
15b+6
15s+2
15f= 380 (raisins)
6
15b+4
15s+5
15f= 500 (peanuts)
2
15b+5
15s+8
15f= 620 (chocolate)
It happens that this system of three equations has just one solution. In other words, as production manager,
your job is easy, since there is but one way to use up all of your raw ingredients making trail mix. This
single solution is
b= 300 kg s= 300 kg f= 900 kg:
We do not yet have the tools to explain why this solution is the only one, but it should be simple for you
to verify that this is indeed a solution. (Go ahead, we will wait.) Determining solutions such as this, and
establishing that they are unique, will be the main motivation for our initial study of linear algebra.
So we have solved the problem of making sure that we make the best use of our limited storage space,
and each day use up all of the raw ingredients that are shipped to us. Additionally, as production manager,
you must report weekly to the CEO of the company, and you know he will be more interested in the prot
derived from your decisions than in the actual production levels. So you compute,
300(4:99 3:69) + 300(5 :50 3:86) + 900(6 :50 4:45) = 2727:00
for a daily prot of $2,727 from this production schedule. The computation of the daily prot is also
beyond our control, though it is denitely of interest, and it too looks like a \linear" computation.
As often happens, things do not stay the same for long, and now the marketing department has
suggested that your company's trail mix products standardize on every mix being one-third peanuts.
Adjusting the peanut portion of each recipe by also adjusting the chocolate portion leads to revised recipes,
and slightly dierent costs for the bulk and standard mixes, as given in the following table.
Version 2.30
6 Section WILA What is Linear Algebra?
Raisins Peanuts Chocolate Cost Sale Price
(kg/batch) (kg/batch) (kg/batch) ($/kg) ($/kg)
Bulk 7 5 3 3.70 4.99
Standard 6 5 4 3.85 5.50
Fancy 2 5 8 4.45 6.50
Storage (kg) 380 500 620
Cost ($/kg) 2.55 4.65 4.80
In a similar fashion as before, we desire values of b,sandfso that
b0 s0 f0
and
7
15b+6
15s+2
15f= 380 (raisins)
5
15b+5
15s+5
15f= 500 (peanuts)
3
15b+4
15s+8
15f= 620 (chocolate)
It now happens that this system of equations has innitely many solutions, as we will now demonstrate.
Letfremain a variable quantity. Then if we make fkilograms of the fancy mix, we will make 4 f 3300
kilograms of the bulk mix and 5f+ 4800 kilograms of the standard mix. Let us now verify that, for any
choice off, the values of b= 4f 3300 ands= 5f+ 4800 will yield a production schedule that exhausts
all of the day's supply of raw ingredients (right now, do not be concerned about how you might derive
expressions like these for bands). Grab your pencil and paper and play along.
7
15(4f 3300) +6
15( 5f+ 4800) +2
15f= 0f+5700
15= 380
5
15(4f 3300) +5
15( 5f+ 4800) +5
15f= 0f+7500
15= 500
3
15(4f 3300) +4
15( 5f+ 4800) +8
15f= 0f+9300
15= 620
Convince yourself that these expressions for bandsallow us to vary fand obtain an innite number of
possibilities for solutions to the three equations that describe our storage capacities. As a practical matter,
there really are not an innite number of solutions, since we are unlikely to want to end the day with a
fractional number of bags of fancy mix, so our allowable values of fshould probably be integers. More
importantly, we need to remember that we cannot make negative amounts of each mix! Where does this
lead us? Positive quantities of the bulk mix requires that
b0) 4f 33000)f825
Similarly for the standard mix,
s0) 5f+ 48000)f960
So, as production manager, you really have to choose a value of ffrom the nite set
f825;826; :::; 960g
leaving you with 136 choices, each of which will exhaust the day's supply of raw ingredients. Pause now
and think about which youwould choose.
Version 2.30
Subsection WILA.READ Reading Questions 7
Recalling your weekly meeting with the CEO suggests that you might want to choose a production
schedule that yields the biggest possible prot for the company. So you compute an expression for the
prot based on your as yet undetermined decision for the value of f,
(4f 3300)(4:99 3:70) + ( 5f+ 4800)(5:50 3:85) + (f)(6:50 4:45) = 1:04f+ 3663
Sincefhas a negative coecient it would appear that mixing fancy mix is detrimental to your prot and
should be avoided. So you will make the decision to set daily fancy mix production at f= 825. This has
the eect of setting b= 4(825) 3300 = 0 and we stop producing bulk mix entirely. So the remainder of
your daily production is standard mix at the level of s= 5(825)+4800 = 675 kilograms and the resulting
daily prot is ( 1:04)(825) + 3663 = 2805. It is a pleasant surprise that daily prot has risen to $2,805,
but this is not the most important part of the story. What is important here is that there are a large
number of ways to produce trail mix that use all of the day's worth of raw ingredients andyou were able
to easily choose the one that netted the largest prot. Notice too how all of the above computations look
\linear."
In the food industry, things do not stay the same for long, and now the sales department says that
increased competition has led to the decision to stay competitive and charge just $5.25 for a kilogram of the
standard mix, rather than the previous $5.50 per kilogram. This decision has no eect on the possibilities
for the production schedule, but will aect the decision based on prot considerations. So you revisit just
the prot computation, suitably adjusted for the new selling price of standard mix,
(4f 3300)(4:99 3:70) + ( 5f+ 4800)(5:25 3:85) + (f)(6:50 4:45) = 0:21f+ 2463
Now it would appear that fancy mix is benecial to the company's prot since the value of fhas a positive
coecient. So you take the decision to make as much fancy mix as possible, setting f= 960. This leads
tos= 5(960) + 4800 = 0 and the increased competition has driven you out of the standard mix market
all together. The remainder of production is therefore bulk mix at a daily level of b= 4(960) 3300 = 540
kilograms and the resulting daily prot is 0 :21(960) + 2463 = 2664 :60. A daily prot of $2,664.60 is less
than it used to be, but as production manager, you have made the best of a dicult situation and shown
the sales department that the best course is to pull out of the highly competitive standard mix market
completely.
This example is taken from a eld of mathematics variously known by names such as operations research,
systems science, or management science. More specically, this is a perfect example of problems that are
solved by the techniques of \linear programming."
There is a lot going on under the hood in this example. The heart of the matter is the solution to
systems of linear equations, which is the topic of the next few sections, and a recurrent theme throughout
this course. We will return to this example on several occasions to reveal some of the reasons for its
behavior.
Subsection READ
Reading Questions
1. Is the equation x2+xy+ tan(y3) = 0 linear or not? Why or why not?
2. Find all solutions to the system of two linear equations 2 x+ 3y= 8,x y= 6.
3. Describe how the production manager might explain the importance of the procedures described in
the trail mix application (Subsection WILA.AA [4]).
Version 2.30
8 Section WILA What is Linear Algebra?
Subsection EXC
Exercises
C10 In Example TMP [4] the rst table lists the cost (per kilogram) to manufacture each of the three
varieties of trail mix (bulk, standard, fancy). For example, it costs $3.69 to make one kilogram of the bulk
variety. Re-compute each of these three costs and notice that the computations are linear in character.
Contributed by Robert Beezer
M70 In Example TMP [4] two dierent prices were considered for marketing standard mix with the
revised recipes (one-third peanuts in each recipe). Selling standard mix at $5.50 resulted in selling the
minimum amount of the fancy mix and no bulk mix. At $5.25 it was best for prots to sell the maximum
amount of fancy mix and then sell no standard mix. Determine a selling price for standard mix that allows
for maximum prots while still selling some of each type of mix.
Contributed by Robert Beezer Solution [9]
Version 2.30
Subsection WILA.SOL Solutions 9
Subsection SOL
Solutions
M70 Contributed by Robert Beezer Statement [8]
If the price of standard mix is set at $5.292, then the prot function has a zero coecient on the variable
quantityf. So, we can set fto be any integer quantity in f825;826; :::; 960g. All but the extreme values
(f= 825,f= 960) will result in production levels where some of every mix is manufactured. No matter
what value of fis chosen, the resulting prot will be the same, at $2,664.60.
Version 2.30
10 Section WILA What is Linear Algebra?
Version 2.30
Section SSLE Solving Systems of Linear Equations 11
Section SSLE
Solving Systems of Linear Equations
We will motivate our study of linear algebra by considering the problem of solving several linear equations
simultaneously. The word \solve" tends to get abused somewhat, as in \solve this problem." When talking
about equations we understand a more precise meaning: nd allof the values of some variable quantities
that make an equation, or several equations, true.
Subsection SLE
Systems of Linear Equations
Example STNE
Solving two (nonlinear) equations
Suppose we desire the simultaneous solutions of the two equations,
x2+y2= 1
x+p
3y= 0
You can easily check by substitution that x=p
3
2; y=1
2andx= p
3
2; y= 1
2are both solutions. We
need to also convince ourselves that these are the only solutions. To see this, plot each equation on the
xy-plane, which means to plot ( x; y) pairs that make an individual equation true. In this case we get
a circle centered at the origin with radius 1 and a straight line through the origin with slope1p
3. The
intersections of these two curves are our desired simultaneous solutions, and so we believe from our plot
that the two solutions we know already are indeed the only ones. We like to write solutions as sets, so in
this case we write the set of solutions as
S=np
3
2;1
2
;
p
3
2; 1
2o
In order to discuss systems of linear equations carefully, we need a precise denition. And before
we do that, we will introduce our periodic discussions about \Proof Techniques." Linear algebra is an
excellent setting for learning how to read, understand and formulate proofs. But this is a dicult step in
your development as a mathematician, so we have included a series of short essays containing advice and
explanations to help you along. These can be found back in Section PT [765] of Appendix P [757], and
we will reference them as they become appropriate. Be sure to head back to the appendix to read this as
they are introduced. With a denition next, now is the time for the rst of our proof techniques. Head
back to Section PT [765] of Appendix P [757] and study Technique D [765]. We'll be right here when you
get back. See you in a bit.
Denition SLE
System of Linear Equations
Asystem of linear equations is a collection of mequations in the variable quantities x1; x2; x3;:::;xn
of the form,
a11x1+a12x2+a13x3++a1nxn=b1
a21x1+a22x2+a23x3++a2nxn=b2
Version 2.30
12 Section SSLE Solving Systems of Linear Equations
a31x1+a32x2+a33x3++a3nxn=b3
...
am1x1+am2x2+am3x3++amnxn=bm
where the values of aij,biandxjare from the set of complex numbers, C. 4
Don't let the mention of the complex numbers, C, rattle you. We will stick with real numbers exclusively
for many more sections, and it will sometimes seem like we only work with integers! However, we want to
leave the possibility of complex numbers open, and there will be occasions in subsequent sections where
they are necessary. You can review the basic properties of complex numbers in Section CNO [757], but
these facts will not be critical until we reach Section O [191].
Now we make the notion of a solution to a linear system precise.
Denition SSLE
Solution of a System of Linear Equations
Asolution of a system of linear equations in nvariables,x1; x2; x3; :::; xn(such as the system given in
Denition SLE [11], is an ordered list of ncomplex numbers, s1; s2; s3; :::; snsuch that if we substitute
s1forx1,s2forx2,s3forx3, . . . ,snforxn, then for every equation of the system the left side will equal
the right side, i.e. each equation is true simultaneously. 4
More typically, we will write a solution in a form like x1= 12,x2= 7,x3= 2 to mean that s1= 12,
s2= 7,s3= 2 in the notation of Denition SSLE [12]. To discuss allof the possible solutions to a system
of linear equations, we now dene the set of all solutions. (So Section SET [761] is now applicable, and
you may want to go and familiarize yourself with what is there.)
Denition SSSLE
Solution Set of a System of Linear Equations
Thesolution set of a linear system of equations is the set which contains every solution to the system,
and nothing more. 4
Be aware that a solution set can be innite, or there can be no solutions, in which case we write the
solution set as the empty set, ;=fg(Denition ES [761]). Here is an example to illustrate using the
notation introduced in Denition SLE [11] and the notion of a solution (Denition SSLE [12]).
Example NSE
Notation for a system of equations
Given the system of linear equations,
x1+ 2x2+x4= 7
x1+x2+x3 x4= 3
3x1+x2+ 5x3 7x4= 1
we haven= 4 variables and m= 3 equations. Also,
a11= 1 a12= 2 a13= 0 a14= 1 b1= 7
a21= 1 a22= 1 a23= 1 a24= 1 b2= 3
a31= 3 a32= 1 a33= 5 a34= 7 b3= 1
Additionally, convince yourself that x1= 2,x2= 4,x3= 2,x4= 1 is one solution (Denition SSLE
[12]), but it is not the only one! For example, another solution is x1= 12,x2= 11,x3= 1,x4= 3,
and there are more to be found. So the solution set contains at least two elements.
We will often shorten the term \system of linear equations" to \system of equations" leaving the linear
aspect implied. After all, this is a book about linear algebra.
Version 2.30
Subsection SSLE.PSS Possibilities for Solution Sets 13
Subsection PSS
Possibilities for Solution Sets
The next example illustrates the possibilities for the solution set of a system of linear equations. We will
not be too formal here, and the necessary theorems to back up our claims will come in subsequent sections.
So read for feeling and come back later to revisit this example.
Example TTS
Three typical systems
Consider the system of two equations with two variables,
2x1+ 3x2= 3
x1 x2= 4
If we plot the solutions to each of these equations separately on the x1x2-plane, we get two lines, one with
negative slope, the other with positive slope. They have exactly one point in common, ( x1; x2) = (3; 1),
which is the solution x1= 3,x2= 1. From the geometry, we believe that this is the only solution to the
system of equations, and so we say it is unique.
Now adjust the system with a dierent second equation,
2x1+ 3x2= 3
4x1+ 6x2= 6
A plot of the solutions to these equations individually results in two lines, one on top of the other! There
are innitely many pairs of points that make both equations true. We will learn shortly how to describe
this innite solution set precisely (see Example SAA [40], Theorem VFSLS [118]). Notice now how the
second equation is just a multiple of the rst.
One more minor adjustment provides a third system of linear equations,
2x1+ 3x2= 3
4x1+ 6x2= 10
A plot now reveals two lines with identical slopes, i.e. parallel lines. They have no points in common, and
so the system has a solution set that is empty, S=;.
This example exhibits all of the typical behaviors of a system of equations. A subsequent theorem will
tell us that every system of linear equations has a solution set that is empty, contains a single solution
or contains innitely many solutions (Theorem PSSLS [60]). Example STNE [11] yielded exactly two
solutions, but this does not contradict the forthcoming theorem. The equations in Example STNE [11] are
not linear because they do not match the form of Denition SLE [11], and so we cannot apply Theorem
PSSLS [60] in this case.
Subsection ESEO
Equivalent Systems and Equation Operations
With all this talk about nding solution sets for systems of linear equations, you might be ready to begin
learning how to nd these solution sets yourself. We begin with our rst denition that takes a common
word and gives it a very precise meaning in the context of systems of linear equations.
Version 2.30
14 Section SSLE Solving Systems of Linear Equations
Denition ESYS
Equivalent Systems
Two systems of linear equations are equivalent if their solution sets are equal. 4
Notice here that the two systems of equations could lookvery dierent (i.e. not be equal), but still have
equal solution sets, and we would then call the systems equivalent. Two linear equations in two variables
might be plotted as two lines that intersect in a single point. A dierent system, with three equations in
two variables might have a plot that is three lines, all intersecting at a common point, with this common
point identical to the intersection point for the rst system. By our denition, we could then say these
two very dierent looking systems of equations are equivalent, since they have identical solution sets. It is
really like a weaker form of equality, where we allow the systems to be dierent in some respects, but we
use the term equivalent to highlight the situation when their solution sets are equal.
With this denition, we can begin to describe our strategy for solving linear systems. Given a system
of linear equations that looks dicult to solve, we would like to have an equivalent system that is easy to
solve. Since the systems will have equal solution sets, we can solve the \easy" system and get the solution
set to the \dicult" system. Here come the tools for making this strategy viable.
Denition EO
Equation Operations
Given a system of linear equations, the following three operations will transform the system into a dierent
one, and each operation is known as an equation operation .
1. Swap the locations of two equations in the list of equations.
2. Multiply each term of an equation by a nonzero quantity.
3. Multiply each term of one equation by some quantity, and add these terms to a second equation, on
both sides of the equality. Leave the rst equation the same after this operation, but replace the
second equation by the new one.
4
These descriptions might seem a bit vague, but the proof or the examples that follow should make it
clear what is meant by each. We will shortly prove a key theorem about equation operations and solutions
to linear systems of equations. We are about to give a rather involved proof, so a discussion about just
what a theorem really is would be timely. Head back and read Technique T [766]. In the theorem we are
about to prove, the conclusion is that two systems are equivalent. By Denition ESYS [14] this translates
to requiring that solution sets be equal for the two systems. So we are being asked to show that two sets
are equal . How do we do this? Well, there is a very standard technique, and we will use it repeatedly
through the course. If you have not done so already, head to Section SET [761] and familiarize yourself
with sets, their operations, and especially the notion of set equality, Denition SE [762] and the nearby
discussion about its use.
Theorem EOPSS
Equation Operations Preserve Solution Sets
If we apply one of the three equation operations of Denition EO [14] to a system of linear equations
(Denition SLE [11]), then the original system and the transformed system are equivalent.
Proof We take each equation operation in turn and show that the solution sets of the two systems are
equal, using the denition of set equality (Denition SE [762]).
1. It will not be our habit in proofs to resort to saying statements are \obvious," but in this case, it
should be. There is nothing about the order in which we write linear equations that aects their
solutions, so the solution set will be equal if the systems only dier by a rearrangement of the order
of the equations.
Version 2.30
Subsection SSLE.ESEO Equivalent Systems and Equation Operations 15
2. Suppose 6= 0 is a number. Let's choose to multiply the terms of equation ibyto build the new
system of equations,
a11x1+a12x2+a13x3++a1nxn=b1
a21x1+a22x2+a23x3++a2nxn=b2
a31x1+a32x2+a33x3++a3nxn=b3
...
ai1x1+ai2x2+ai3x3++ainxn=bi
...
am1x1+am2x2+am3x3++amnxn=bm
LetSdenote the solutions to the system in the statement of the theorem, and let Tdenote the
solutions to the transformed system.
(a) ShowST. Suppose ( x1; x2; x3; :::;xn) = (1; 2; 3; :::;n)2Sis a solution to the original
system. Ignoring the i-th equation for a moment, we know it makes all the other equations of
the transformed system true. We also know that
ai11+ai22+ai33++ainn=bi
which we can multiply by to get
ai11+ai22+ai33++ainn=bi
This says that the i-th equation of the transformed system is also true, so we have established
that (1; 2; 3; :::;n)2T, and therefore ST.
(b) Now show TS. Suppose ( x1; x2; x3; :::;xn) = (1; 2; 3; :::;n)2Tis a solution to the
transformed system. Ignoring the i-th equation for a moment, we know it makes all the other
equations of the original system true. We also know that
ai11+ai22+ai33++ainn=bi
which we can multiply by1
, since6= 0, to get
ai11+ai22+ai33++ainn=bi
This says that the i-th equation of the original system is also true, so we have established that
(1; 2; 3; :::;n)2S, and therefore TS. Locate the key point where we required that
6= 0, and consider what would happen if = 0.
3. Suppose is a number. Let's choose to multiply the terms of equation ibyand add them to
equationjin order to build the new system of equations,
a11x1+a12x2++a1nxn=b1
a21x1+a22x2++a2nxn=b2
a31x1+a32x2++a3nxn=b3
...
(ai1+aj1)x1+ (ai2+aj2)x2++ (ain+ajn)xn=bi+bj
Version 2.30
16 Section SSLE Solving Systems of Linear Equations
...
am1x1+am2x2++amnxn=bm
LetSdenote the solutions to the system in the statement of the theorem, and let Tdenote the
solutions to the transformed system.
(a) ShowST. Suppose ( x1; x2; x3; :::;xn) = (1; 2; 3; :::;n)2Sis a solution to the
original system. Ignoring the j-th equation for a moment, we know this solution makes all the
other equations of the transformed system true. Using the fact that the solution makes the i-th
andj-th equations of the original system true, we nd
(ai1+aj1)1+ (ai2+aj2)2++ (ain+ajn)n
= (ai11+ai22++ainn) + (aj11+aj22++ajnn)
=(ai11+ai22++ainn) + (aj11+aj22++ajnn)
=bi+bj:
This says that the j-th equation of the transformed system is also true, so we have established
that (1; 2; 3; :::;n)2T, and therefore ST.
(b) Now show TS. Suppose ( x1; x2; x3; :::;xn) = (1; 2; 3; :::;n)2Tis a solution to the
transformed system. Ignoring the j-th equation for a moment, we know it makes all the other
equations of the original system true. We then nd
aj11+aj22++ajnn
=aj11+aj22++ajnn+bi bi
=aj11+aj22++ajnn+ (ai11+ai22++ainn) bi
=aj11+ai11+aj22+ai22++ajnn+ainn bi
= (ai1+aj1)1+ (ai2+aj2)2++ (ain+ajn)n bi
=bi+bj bi
=bj
This says that the j-th equation of the original system is also true, so we have established that
(1; 2; 3; :::;n)2S, and therefore TS.
Why didn't we need to require that 6= 0 for this row operation? In other words, how does the
third statement of the theorem read when = 0? Does our proof require some extra care when
= 0? Compare your answers with the similar situation for the second row operation. (See Exercise
SSLE.T20 [22].)
Theorem EOPSS [14] is the necessary tool to complete our strategy for solving systems of equations.
We will use equation operations to move from one system to another, all the while keeping the solution set
the same. With the right sequence of operations, we will arrive at a simpler equation to solve. The next
two examples illustrate this idea, while saving some of the details for later.
Example US
Three equations, one solution
We solve the following system by a sequence of equation operations.
x1+ 2x2+ 2x3= 4
x1+ 3x2+ 3x3= 5
Version 2.30
Subsection SSLE.ESEO Equivalent Systems and Equation Operations 17
2x1+ 6x2+ 5x3= 6
= 1 times equation 1, add to equation 2:
x1+ 2x2+ 2x3= 4
0x1+ 1x2+ 1x3= 1
2x1+ 6x2+ 5x3= 6
= 2 times equation 1, add to equation 3:
x1+ 2x2+ 2x3= 4
0x1+ 1x2+ 1x3= 1
0x1+ 2x2+ 1x3= 2
= 2 times equation 2, add to equation 3:
x1+ 2x2+ 2x3= 4
0x1+ 1x2+ 1x3= 1
0x1+ 0x2 1x3= 4
= 1 times equation 3:
x1+ 2x2+ 2x3= 4
0x1+ 1x2+ 1x3= 1
0x1+ 0x2+ 1x3= 4
which can be written more clearly as
x1+ 2x2+ 2x3= 4
x2+x3= 1
x3= 4
This is now a very easy system of equations to solve. The third equation requires that x3= 4 to be true.
Making this substitution into equation 2 we arrive at x2= 3, and nally, substituting these values of x2
andx3into the rst equation, we nd that x1= 2. Note too that this is the only solution to this nal
system of equations, since we were forced to choose these values to make the equations true. Since we
performed equation operations on each system to obtain the next one in the list, all of the systems listed
here are all equivalent to each other by Theorem EOPSS [14]. Thus ( x1; x2; x3) = (2; 3;4) is the unique
solution to the original system of equations (and all of the other intermediate systems of equations listed
as we transformed one into another).
Example IS
Three equations, innitely many solutions
The following system of equations made an appearance earlier in this section (Example NSE [12]), where
we listed oneof its solutions. Now, we will try to nd all of the solutions to this system. Don't concern
yourself too much about why we choose this particular sequence of equation operations, just believe that
the work we do is all correct.
x1+ 2x2+ 0x3+x4= 7
Version 2.30
18 Section SSLE Solving Systems of Linear Equations
x1+x2+x3 x4= 3
3x1+x2+ 5x3 7x4= 1
= 1 times equation 1, add to equation 2:
x1+ 2x2+ 0x3+x4= 7
0x1 x2+x3 2x4= 4
3x1+x2+ 5x3 7x4= 1
= 3 times equation 1, add to equation 3:
x1+ 2x2+ 0x3+x4= 7
0x1 x2+x3 2x4= 4
0x1 5x2+ 5x3 10x4= 20
= 5 times equation 2, add to equation 3:
x1+ 2x2+ 0x3+x4= 7
0x1 x2+x3 2x4= 4
0x1+ 0x2+ 0x3+ 0x4= 0
= 1 times equation 2:
x1+ 2x2+ 0x3+x4= 7
0x1+x2 x3+ 2x4= 4
0x1+ 0x2+ 0x3+ 0x4= 0
= 2 times equation 2, add to equation 1:
x1+ 0x2+ 2x3 3x4= 1
0x1+x2 x3+ 2x4= 4
0x1+ 0x2+ 0x3+ 0x4= 0
which can be written more clearly as
x1+ 2x3 3x4= 1
x2 x3+ 2x4= 4
0 = 0
What does the equation 0 = 0 mean? We can choose anyvalues forx1; x2; x3; x4and this equation will
be true, so we only need to consider further the rst two equations, since the third is true no matter what.
We can analyze the second equation without consideration of the variable x1. It would appear that there
is considerable latitude in how we can choose x2; x3; x4and make this equation true. Let's choose x3and
x4to be anything we please, say x3=aandx4=b.
Now we can take these arbitrary values for x3andx4, substitute them in equation 1, to obtain
x1+ 2a 3b= 1
x1= 1 2a+ 3b
Version 2.30
Subsection SSLE.READ Reading Questions 19
Similarly, equation 2 becomes
x2 a+ 2b= 4
x2= 4 +a 2b
So our arbitrary choices of values for x3andx4(aandb) translate into specic values of x1andx2. The
lone solution given in Example NSE [12] was obtained by choosing a= 2 andb= 1. Now we can easily
and quickly nd many more (innitely more). Suppose we choose a= 5 andb= 2, then we compute
x1= 1 2(5) + 3( 2) = 17
x2= 4 + 5 2( 2) = 13
and you can verify that ( x1; x2; x3; x4) = ( 17;13;5; 2) makes all three equations true. The entire
solution set is written as
S=f( 1 2a+ 3b;4 +a 2b; a; b )ja2C; b2Cg
It would be instructive to nish o your study of this example by taking the general form of the solutions
given in this set and substituting them into each of the three equations and verify that they are true in
each case (Exercise SSLE.M40 [22]).
In the next section we will describe how to use equation operations to systematically solve any system
of linear equations. But rst, read one of our more important pieces of advice about speaking and writing
mathematics. See Technique L [766].
Before attacking the exercises in this section, it will be helpful to read some advice on getting started
on the construction of a proof. See Technique GS [767].
Subsection READ
Reading Questions
1. How many solutions does the system of equations 3 x+ 2y= 4, 6x+ 4y= 8 have? Explain your
answer.
2. How many solutions does the system of equations 3 x+ 2y= 4, 6x+ 4y= 2 have? Explain your
answer.
3. What do we mean when we say mathematics is a language?
Version 2.30
20 Section SSLE Solving Systems of Linear Equations
Subsection EXC
Exercises
C10 Find a solution to the system in Example IS [17] where x3= 6 andx4= 2. Find two other solutions
to the system. Find a solution where x1= 17 andx2= 14. How many possible answers are there to each
of these questions?
Contributed by Robert Beezer
C20 Each archetype (Appendix A [777]) that is a system of equations begins by listing some specic
solutions. Verify the specic solutions listed in the following archetypes by evaluating the system of
equations with the solutions listed.
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]
Archetype E [799]
Archetype F [803]
Archetype G [808]
Archetype H [812]
Archetype I [816]
Archetype J [820]
Contributed by Robert Beezer
C30 Find all solutions to the linear system:
x+y= 5
2x y= 3
Contributed by Chris Black Solution [23]
C31 Find all solutions to the linear system:
3x+ 2y= 1
x y= 2
4x+ 2y= 2
Contributed by Chris Black
C32 Find all solutions to the linear system:
x+ 2y= 8
x y= 2
x+y= 4
Contributed by Chris Black
C33 Find all solutions to the linear system:
x+y z= 1
Version 2.30
Subsection SSLE.EXC Exercises 21
x y z= 1
z= 2
Contributed by Chris Black
C34 Find all solutions to the linear system:
x+y z= 5
x y z= 3
x+y z= 0
Contributed by Chris Black
C50 A three-digit number has two properties. The tens-digit and the ones-digit add up to 5. If the
number is written with the digits in the reverse order, and then subtracted from the original number, the
result is 792. Use a system of equations to nd all of the three-digit numbers with these properties.
Contributed by Robert Beezer Solution [23]
C51 Find all of the six-digit numbers in which the rst digit is one less than the second, the third digit is
half the second, the fourth digit is three times the third and the last two digits form a number that equals
the sum of the fourth and fth. The sum of all the digits is 24. (From The MENSA Puzzle Calendar for
January 9, 2006.)
Contributed by Robert Beezer Solution [23]
C52 Driving along, Terry notices that the last four digits on his car's odometer are palindromic. A mile
later, the last ve digits are palindromic. After driving another mile, the middle four digits are palindromic.
One more mile, and all six are palindromic. What was the odometer reading when Terry rst looked at
it? Form a linear system of equations that expresses the requirements of this puzzle. ( Car Talk Puzzler,
National Public Radio, Week of January 21, 2008) (A car odometer displays six digits and a sequence is a
palindrome if it reads the same left-to-right as right-to-left.)
Contributed by Robert Beezer Solution [24]
M10 Each sentence below has at least two meanings. Identify the source of the double meaning, and
rewrite the sentence (at least twice) to clearly convey each meaning.
1. They are baking potatoes.
2. He bought many ripe pears and apricots.
3. She likes his sculpture.
4. I decided on the bus.
Contributed by Robert Beezer Solution [24]
M11 Discuss the dierence in meaning of each of the following three almost identical sentences, which
all have the same grammatical structure. (These are due to Keith Devlin.)
1. She saw him in the park with a dog.
2. She saw him in the park with a fountain.
3. She saw him in the park with a telescope.
Version 2.30
22 Section SSLE Solving Systems of Linear Equations
Contributed by Robert Beezer Solution [24]
M12 The following sentence, due to Noam Chomsky, has a correct grammatical structure, but is mean-
ingless. Critique its faults. \Colorless green ideas sleep furiously." (Chomsky, Noam. Syntactic Structures ,
The Hague/Paris: Mouton, 1957. p. 15.)
Contributed by Robert Beezer Solution [24]
M13 Read the following sentence and form a mental picture of the situation.
The baby cried and the mother picked it up.
What assumptions did you make about the situation?
Contributed by Robert Beezer Solution [24]
M30 This problem appears in a middle-school mathematics textbook: Together Dan and Diane have
$20. Together Diane and Donna have $15. How much do the three of them have in total? ( Transition
Mathematics , Second Edition, Scott Foresman Addison Wesley, 1998. Problem 5{1.19.)
Contributed by David Beezer Solution [25]
M40 Solutions to the system in Example IS [17] are given as
(x1; x2; x3; x4) = ( 1 2a+ 3b;4 +a 2b; a; b )
Evaluate the three equations of the original system with these expressions in aandband verify that each
equation is true, no matter what values are chosen for aandb.
Contributed by Robert Beezer
M70 We have seen in this section that systems of linear equations have limited possibilities for solution
sets, and we will shortly prove Theorem PSSLS [60] that describes these possibilities exactly. This exercise
will show that if we relax the requirement that our equations be linear, then the possibilities expand greatly.
Consider a system of two equations in the two variables xandy, where the departure from linearity involves
simply squaring the variables.
x2 y2= 1
x2+y2= 4
After solving this system of non-linear equations, replace the second equation in turn by x2+ 2x+y2= 3,
x2+y2= 1,x2 4x+y2= 3, x2+y2= 1 and solve each resulting system of two equations in two
variables. (This exercise includes suggestions from Don Kreher.)
Contributed by Robert Beezer Solution [25]
T10 Technique D [765] asks you to formulate a denition of what it means for a whole number to be
odd. What is your denition? (Don't say \the opposite of even.") Is 6 odd? Is 11 odd? Justify your
answers by using your denition.
Contributed by Robert Beezer Solution [25]
T20 Explain why the second equation operation in Denition EO [14] requires that the scalar be nonzero,
while in the third equation operation this restriction on the scalar is not present.
Contributed by Robert Beezer Solution [25]
Version 2.30
Subsection SSLE.SOL Solutions 23
Subsection SOL
Solutions
C30 Contributed by Chris Black Statement [20]
Solving each equation for y, we have the equivalent system
y= 5 x
y= 2x 3:
Setting these expressions for yequal, we have the equation 5 x= 2x 3, which quickly leads to x=8
3.
Substituting for xin the rst equation, we have y= 5 x= 5 8
3=7
3. Thus, the solution is x=8
3,y=7
3.
C50 Contributed by Robert Beezer Statement [21]
Letabe the hundreds digit, bthe tens digit, and cthe ones digit. Then the rst condition says that
b+c= 5. The original number is 100 a+ 10b+c, while the reversed number is 100 c+ 10b+a. So the
second condition is
792 = (100a+ 10b+c) (100c+ 10b+a) = 99a 99c
So we arrive at the system of equations
b+c= 5
99a 99c= 792
Using equation operations, we arrive at the equivalent system
a c= 8
b+c= 5
We can vary cand obtain innitely many solutions. However, cmust be a digit, restricting us to ten values
(0 { 9). Furthermore, if c>1, then the rst equation forces a>9, an impossibility. Setting c= 0, yields
850 as a solution, and setting c= 1 yields 941 as another solution.
C51 Contributed by Robert Beezer Statement [21]
Letabcdef denote any such six-digit number and convert each requirement in the problem statement into
an equation.
a=b 1
c=1
2b
d= 3c
10e+f=d+e
24 =a+b+c+d+e+f
In a more standard form this becomes
a b= 1
b+ 2c= 0
3c+d= 0
d+ 9e+f= 0
a+b+c+d+e+f= 24
Version 2.30
24 Section SSLE Solving Systems of Linear Equations
Using equation operations (or the techniques of the upcoming Section RREF [27]), this system can be
converted to the equivalent system
a+16
75f= 5
b+16
75f= 6
c+8
75f= 3
d+8
25f= 9
e+11
75f= 1
Clearly, choosing f= 0 will yield the solution abcde = 563910. Furthermore, to have the variables result
in single-digit numbers, none of the other choices for f(1;2; :::; 9) will yield a solution.
C52 Contributed by Robert Beezer Statement [21]
198888 is one solution, and David Braithwaite found 199999 as another.
M10 Contributed by Robert Beezer Statement [21]
1. Does \baking" describe the potato or what is happening to the potato?
Those are potatoes that are used for baking.
The potatoes are being baked.
2. Are the apricots ripe, or just the pears? Parentheses could indicate just what the adjective \ripe" is
meant to modify. Were there many apricots as well, or just many pears?
He bought many pears and many ripe apricots.
He bought apricots and many ripe pears.
3. Is \sculpture" a single physical object, or the sculptor's style expressed over many pieces and many
years?
She likes his sculpture of the girl.
She likes his sculptural style.
4. Was a decision made while in the bus, or was the outcome of a decision to choose the bus. Would the
sentence \I decided on the car," have a similar double meaning?
I made my decision while on the bus.
I decided to ride the bus.
M11 Contributed by Robert Beezer Statement [21]
We know the dog belongs to the man, and the fountain belongs to the park. It is not clear if the telescope
belongs to the man, the woman, or the park.
M12 Contributed by Robert Beezer Statement [22]
In adjacent pairs the words are contradictory or inappropriate. Something cannot be both green and
colorless, ideas do not have color, ideas do not sleep, and it is hard to sleep furiously.
M13 Contributed by Robert Beezer Statement [22]
Did you assume that the baby and mother are human?
Did you assume that the baby is the child of the mother?
Did you assume that the mother picked up the baby as an attempt to stop the crying?
Version 2.30
Subsection SSLE.SOL Solutions 25
M30 Contributed by Robert Beezer Statement [22]
Ifx,yandzrepresent the money held by Dan, Diane and Donna, then y= 15 zandx= 20 y=
20 (15 z) = 5 +z. We can let ztake on any value from 0 to 15 without any of the three amounts being
negative, since presumably middle-schoolers are too young to assume debt.
Then the total capital held by the three is x+y+z= (5+z)+(15 z)+z= 20+z. So their combined
holdings can range anywhere from $20 (Donna is broke) to $35 (Donna is
ush).
We will have more to say about this situation in Section TSS [55], and specically Theorem CMVEI
[61].
M70 Contributed by Robert Beezer Statement [22]
The equation x2 y2= 1 has a solution set by itself that has the shape of a hyperbola when plotted. Four
of the ve dierent second equations have solution sets that are circles when plotted individually (the last
is another hyperbola). Where the hyperbola and circles intersect are the solutions to the system of two
equations. As the size and location of the circles vary, the number of intersections varies from four to one
(in the order given). Teh last equation is a hyperbola that \opens" in the other direction. Sketching the
relevant equations would be instructive, as was discussed in Example STNE [11].
The exact solution sets are (according to the choice of the second equation),
x2+y2= 4 :( r
5
2;r
3
2!
;
r
5
2;r
3
2!
; r
5
2; r
3
2!
;
r
5
2; r
3
2!)
x2+ 2x+y2= 3 :n
(1;0);( 2;p
3);( 2; p
3)o
x2+y2= 1 :f(1;0);( 1;0)g
x2 4x+y2= 3 :f(1;0)g
x2+y2= 1 :fg
T10 Contributed by Robert Beezer Statement [22]
We can say that an integer is odd if when it is divided by 2 there is a remainder of 1. So 6 is not odd
since 6 = 32 + 0, while 11 is odd since 11 = 5 2 + 1.
T20 Contributed by Robert Beezer Statement [22]
Denition EO [14] is engineered to make Theorem EOPSS [14] true. If we were to allow a zero scalar to
multiply an equation then that equation would be transformed to the equation 0 = 0, which is true for
any possible values of the variables. Any restrictions on the solution set imposed by the original equation
would be lost.
However, in the third operation, it is allowed to choose a zero scalar, multiply an equation by this
scalar and add the transformed equation to a second equation (leaving the rst unchanged). The result?
Nothing. The second equation is the same as it was before. So the theorem is true in this case, the two
systems are equivalent. But in practice, this would be a silly thing to actually ever do! We still allow it
though, in order to keep our theorem as general as possible.
Notice the location in the proof of Theorem EOPSS [14] where the expression1
appears | this explains
the prohibition on = 0 in the second equation operation.
Version 2.30
26 Section SSLE Solving Systems of Linear Equations
Version 2.30
Section RREF Reduced Row-Echelon Form 27
Section RREF
Reduced Row-Echelon Form
After solving a few systems of equations, you will recognize that it doesn't matter so much what we call
our variables, as opposed to what numbers act as their coecients. A system in the variables x1; x2; x3
would behave the same if we changed the names of the variables to a; b; c and kept all the constants the
same and in the same places. In this section, we will isolate the key bits of information about a system of
equations into something called a matrix, and then use this matrix to systematically solve the equations.
Along the way we will obtain one of our most important and useful computational tools.
Subsection MVNSE
Matrix and Vector Notation for Systems of Equations
Denition M
Matrix
Anmnmatrix is a rectangular layout of numbers from Chavingmrows andncolumns. We will use
upper-case Latin letters from the start of the alphabet ( A; B; C;::: ) to denote matrices and squared-o
brackets to delimit the layout. Many use large parentheses instead of brackets | the distinction is not
important. Rows of a matrix will be referenced starting at the top and working down (i.e. row 1 is at the
top) and columns will be referenced starting from the left (i.e. column 1 is at the left). For a matrix A,
the notation [ A]ijwill refer to the complex number in row iand column jofA.
(This denition contains Notation M.)
(This denition contains Notation MC.) 4
Be careful with this notation for individual entries, since it is easy to think that [ A]ijrefers to the
whole matrix. It does not. It is just a number , but is a convenient way to talk about the individual entries
simultaneously. This notation will get a heavy workout once we get to Chapter M [207].
Example AM
A matrix
B=2
4 1 2 5 3
1 0 6 1
4 2 2 23
5
is a matrix with m= 3 rows and n= 4 columns. We can say that [ B]2;3= 6 while [B]3;4= 2.
Some mathematical software is very particular about which types of numbers (integers, rationals,
reals, complexes) you wish to work with. See: Computation R.SAGE [752] A calculator or computer
language can be a convenient way to perform calculations with matrices. But rst you have to enter the
matrix. See: Computation ME.MMA [745] Computation ME.TI86 [750] Computation ME.TI83
[751] Computation ME.SAGE [753] When we do equation operations on system of equations, the names
of the variables really aren't very important. x1,x2,x3, ora,b,c, orx,y,z, it really doesn't matter. In
this subsection we will describe some notation that will make it easier to describe linear systems, solve the
systems and describe the solution sets. Here is a list of denitions, laden with notation.
Denition CV
Column Vector
Acolumn vector ofsizemis an ordered list of mnumbers, which is written in order vertically, starting
at the top and proceeding to the bottom. At times, we will refer to a column vector as simply a vector .
Version 2.30
28 Section RREF Reduced Row-Echelon Form
Column vectors will be written in bold, usually with lower case Latin letter from the end of the alphabet
such as u,v,w,x,y,z. Some books like to write vectors with arrows, such as ~ u. Writing by hand, some
like to put arrows on top of the symbol, or a tilde underneath the symbol, as in u
. To refer to the entry
orcomponent that is number iin the list that is the vector vwe write [ v]i.
(This denition contains Notation CV.)
(This denition contains Notation CVC.) 4
Be careful with this notation. While the symbols [ v]imight look somewhat substantial, as an object
this represents just one component of a vector, which is just a single complex number.
Denition ZCV
Zero Column Vector
Thezero vector of sizemis the column vector of size mwhere each entry is the number zero,
0=2
6666640
0
0
...
03
777775
or dened much more compactly, [ 0]i= 0 for 1im.
(This denition contains Notation ZCV.) 4
Denition CM
Coecient Matrix
For a system of linear equations,
a11x1+a12x2+a13x3++a1nxn=b1
a21x1+a22x2+a23x3++a2nxn=b2
a31x1+a32x2+a33x3++a3nxn=b3
...
am1x1+am2x2+am3x3++amnxn=bm
thecoecient matrix is themnmatrix
A=2
666664a11a12a13::: a 1n
a21a22a23::: a 2n
a31a32a33::: a 3n
...
am1am2am3::: amn3
777775
4
Denition VOC
Vector of Constants
For a system of linear equations,
a11x1+a12x2+a13x3++a1nxn=b1
a21x1+a22x2+a23x3++a2nxn=b2
a31x1+a32x2+a33x3++a3nxn=b3
...
Version 2.30
Subsection RREF.MVNSE Matrix and Vector Notation for Systems of Equations 29
am1x1+am2x2+am3x3++amnxn=bm
thevector of constants is the column vector of size m
b=2
666664b1
b2
b3
...
bm3
777775
4
Denition SOLV
Solution Vector
For a system of linear equations,
a11x1+a12x2+a13x3++a1nxn=b1
a21x1+a22x2+a23x3++a2nxn=b2
a31x1+a32x2+a33x3++a3nxn=b3
...
am1x1+am2x2+am3x3++amnxn=bm
thesolution vector is the column vector of size n
x=2
666664x1
x2
x3
...
xn3
777775
4
The solution vector may do double-duty on occasion. It might refer to a list of variable quantities at
one point, and subsequently refer to values of those variables that actually form a particular solution to
that system.
Denition MRLS
Matrix Representation of a Linear System
IfAis the coecient matrix of a system of linear equations and bis the vector of constants, then we will
writeLS(A;b) as a shorthand expression for the system of linear equations, which we will refer to as the
matrix representation of the linear system.
(This denition contains Notation MRLS.) 4
Example NSLE
Notation for systems of linear equations
The system of linear equations
2x1+ 4x2 3x3+ 5x4+x5= 9
3x1+x2+x4 3x5= 0
2x1+ 7x2 5x3+ 2x4+ 2x5= 3
Version 2.30
30 Section RREF Reduced Row-Echelon Form
has coecient matrix
A=2
42 4 3 5 1
3 1 0 1 3
2 7 5 2 23
5
and vector of constants
b=2
49
0
33
5
and so will be referenced as LS(A;b).
Denition AM
Augmented Matrix
Suppose we have a system of mequations in nvariables, with coecient matrix Aand vector of constants
b. Then the augmented matrix of the system of equations is the m(n+ 1) matrix whose rst n
columns are the columns of Aand whose last column (number n+ 1) is the column vector b. This matrix
will be written as [ Ajb].
(This denition contains Notation AM.) 4
The augmented matrix represents all the important information in the system of equations, since the
names of the variables have been ignored, and the only connection with the variables is the location of
their coecients in the matrix. It is important to realize that the augmented matrix is just that, a matrix,
andnota system of equations. In particular, the augmented matrix does not have any \solutions," though
it will be useful for nding solutions to the system of equations that it is associated with. (Think about
your objects, and review Technique L [766].) However, notice that an augmented matrix always belongs
to some system of equations, and vice versa, so it is tempting to try and blur the distinction between the
two. Here's a quick example.
Example AMAA
Augmented matrix for Archetype A
Archetype A [781] is the following system of 3 equations in 3 variables.
x1 x2+ 2x3= 1
2x1+x2+x3= 8
x1+x2= 5
Here is its augmented matrix.
2
41 1 2 1
2 1 1 8
1 1 0 53
5
Subsection RO
Row Operations
An augmented matrix for a system of equations will save us the tedium of continually writing down the
names of the variables as we solve the system. It will also release us from any dependence on the actual
names of the variables. We have seen how certain operations we can perform on equations (Denition
EO [14]) will preserve their solutions (Theorem EOPSS [14]). The next two denitions and the following
theorem carry over these ideas to augmented matrices.
Version 2.30
Subsection RREF.RO Row Operations 31
Denition RO
Row Operations
The following three operations will transform an mnmatrix into a dierent matrix of the same size, and
each is known as a row operation .
1. Swap the locations of two rows.
2. Multiply each entry of a single row by a nonzero quantity.
3. Multiply each entry of one row by some quantity, and add these values to the entries in the same
columns of a second row. Leave the rst row the same after this operation, but replace the second
row by the new values.
We will use a symbolic shorthand to describe these row operations:
1.Ri$Rj: Swap the location of rows iandj.
2.Ri: Multiply row iby the nonzero scalar .
3.Ri+Rj: Multiply row iby the scalar and add to row j.
(This denition contains Notation RO.) 4
Denition REM
Row-Equivalent Matrices
Two matrices, AandB, arerow-equivalent if one can be obtained from the other by a sequence of row
operations. 4
Example TREM
Two row-equivalent matrices
The matrices
A=2
42 1 3 4
5 2 2 3
1 1 0 63
5 B=2
41 1 0 6
3 0 2 9
2 1 3 43
5
are row-equivalent as can be seen from
2
42 1 3 4
5 2 2 3
1 1 0 63
5R1$R3 !2
41 1 0 6
5 2 2 3
2 1 3 43
5 2R1+R2 !2
41 1 0 6
3 0 2 9
2 1 3 43
5
We can also say that any pair of these three matrices are row-equivalent.
Notice that each of the three row operations is reversible (Exercise RREF.T10 [47]), so we do not have
to be careful about the distinction between \ Ais row-equivalent to B" and \Bis row-equivalent to A."
(Exercise RREF.T11 [47]) The preceding denitions are designed to make the following theorem possible.
It says that row-equivalent matrices represent systems of linear equations that have identical solution sets.
Theorem REMES
Row-Equivalent Matrices represent Equivalent Systems
Suppose that AandBare row-equivalent augmented matrices. Then the systems of linear equations that
they represent are equivalent systems.
Proof If we perform a single row operation on an augmented matrix, it will have the same eect as if
we did the analogous equation operation on the corresponding system of equations. By exactly the same
Version 2.30
32 Section RREF Reduced Row-Echelon Form
methods as we used in the proof of Theorem EOPSS [14] we can see that each of these row operations will
preserve the set of solutions for the corresponding system of equations.
So at this point, our strategy is to begin with a system of equations, represent it by an augmented
matrix, perform row operations (which will preserve solutions for the corresponding systems) to get a
\simpler" augmented matrix, convert back to a \simpler" system of equations and then solve that system,
knowing that its solutions are those of the original system. Here's a rehash of Example US [16] as an
exercise in using our new tools.
Example USR
Three equations, one solution, reprised
We solve the following system using augmented matrices and row operations. This is the same system of
equations solved in Example US [16] using equation operations.
x1+ 2x2+ 2x3= 4
x1+ 3x2+ 3x3= 5
2x1+ 6x2+ 5x3= 6
Form the augmented matrix,
A=2
41 2 2 4
1 3 3 5
2 6 5 63
5
and apply row operations,
1R1+R2 !2
41 2 2 4
0 1 1 1
2 6 5 63
5 2R1+R3 !2
41 2 2 4
0 1 1 1
0 2 1 23
5
2R2+R3 !2
41 2 2 4
0 1 1 1
0 0 1 43
5 1R3 !2
41 2 2 4
0 1 1 1
0 0 1 43
5
So the matrix
B=2
41 2 2 4
0 1 1 1
0 0 1 43
5
is row equivalent to Aand by Theorem REMES [31] the system of equations below has the same solution
set as the original system of equations.
x1+ 2x2+ 2x3= 4
x2+x3= 1
x3= 4
Solving this \simpler" system is straightforward and is identical to the process in Example US [16].
Subsection RREF
Reduced Row-Echelon Form
The preceding example amply illustrates the denitions and theorems we have seen so far. But it still
leaves two questions unanswered. Exactly what is this \simpler" form for a matrix, and just how do we
get it? Here's the answer to the rst question, a denition of reduced row-echelon form.
Version 2.30
Subsection RREF.RREF Reduced Row-Echelon Form 33
Denition RREF
Reduced Row-Echelon Form
A matrix is in reduced row-echelon form if it meets all of the following conditions:
1. If there is a row where every entry is zero, then this row lies below any other row that contains a
nonzero entry.
2. The leftmost nonzero entry of a row is equal to 1.
3. The leftmost nonzero entry of a row is the only nonzero entry in its column.
4. Consider any two dierent leftmost nonzero entries, one located in row i, columnjand the other
located in row s, columnt. Ifs>i , thent>j .
A row of only zero entries will be called a zero row and the leftmost nonzero entry of a nonzero row will
be called a leading 1 . The number of nonzero rows will be denoted by r.
A column containing a leading 1 will be called a pivot column . The set of column indices for all of
the pivot columns will be denoted by D=fd1; d2; d3; :::; drgwhered1<d 2<d 3<<dr, while the
columns that are not pivot columns will be denoted as F=ff1; f2; f3; :::; fn rgwheref1<f2<f3<
<fn r.
(This denition contains Notation RREFA.) 4
The principal feature of reduced row-echelon form is the pattern of leading 1's guaranteed by conditions
(2) and (4), reminiscent of a
ight of geese, or steps in a staircase, or water cascading down a mountain
stream.
There are a number of new terms and notation introduced in this denition, which should make you
suspect that this is an important denition. Given all there is to digest here, we will mostly save the use
ofDandFuntil Section TSS [55]. However, one important point to make here is that all of these terms
and notation apply to a matrix. Sometimes we will employ these terms and sets for an augmented matrix,
and other times it might be a coecient matrix. So always give some thought to exactly which type of
matrix you are analyzing.
Example RREF
A matrix in reduced row-echelon form
The matrix Cis in reduced row-echelon form.
C=2
666641 3 0 6 0 0 5 9
0 0 0 0 1 0 3 7
0 0 0 0 0 1 7 3
0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 03
77775
This matrix has two zero rows and three leading 1's. So r= 3. Columns 1, 5, and 6 are pivot columns, so
D=f1;5;6gand thenF=f2;3;4;7;8g.
Example NRREF
A matrix not in reduced row-echelon form
The matrix Eis not in reduced row-echelon form, as it fails each of the four requirements once.
E=2
66666641 0 3 0 6 0 7 5 9
0 0 0 5 0 1 0 3 7
0 0 0 0 0 0 0 0 0
0 1 0 0 0 0 0 4 2
0 0 0 0 0 0 1 7 3
0 0 0 0 0 0 0 0 03
7777775
Version 2.30
34 Section RREF Reduced Row-Echelon Form
Our next theorem has a \constructive" proof. Learn about the meaning of this term in Technique C
[768].
Theorem REMEF
Row-Equivalent Matrix in Echelon Form
SupposeAis a matrix. Then there is a matrix Bso that
1.AandBare row-equivalent.
2.Bis in reduced row-echelon form.
Proof Suppose that Ahasmrows andncolumns. We will describe a process for converting Ainto
Bvia row operations. This procedure is known as Gauss{Jordan elimination . Tracing through this
procedure will be easier if you recognize that irefers to a row that is being converted, jrefers to a column
that is being converted, and rkeeps track of the number of nonzero rows. Here we go.
1. Setj= 0 andr= 0.
2. Increase jby 1. Ifjnow equals n+ 1, then stop.
3. Examine the entries of Ain columnjlocated in rows r+ 1 through m.
If all of these entries are zero, then go to Step 2.
4. Choose a row from rows r+ 1 through mwith a nonzero entry in column j.
Letidenote the index for this row.
5. Increase rby 1.
6. Use the rst row operation to swap rows iandr.
7. Use the second row operation to convert the entry in row rand column jto a 1.
8. Use the third row operation with row rto convert every other entry of column jto zero.
9. Go to Step 2.
The result of this procedure is that the matrix Ais converted to a matrix in reduced row-echelon form,
which we will refer to as B. We need to now prove this claim by showing that the converted matrix has the
requisite properties of Denition RREF [33]. First, the matrix is only converted through row operations
(Step 6, Step 7, Step 8), so AandBare row-equivalent (Denition REM [31]).
It is a bit more work to be certain that Bis in reduced row-echelon form. We claim that as we begin
Step 2, the rst jcolumns of the matrix are in reduced row-echelon form with rnonzero rows. Certainly
this is true at the start when j= 0, since the matrix has no columns and so vacuously meets the conditions
of Denition RREF [33] with r= 0 nonzero rows.
In Step 2 we increase jby 1 and begin to work with the next column. There are two possible outcomes
for Step 3. Suppose that every entry of column jin rowsr+ 1 through mis zero. Then with no changes
we recognize that the rst jcolumns of the matrix has its rst rrows still in reduced-row echelon form,
with the nal m rrows still all zero.
Suppose instead that the entry in row iof columnjis nonzero. Notice that since r+ 1im, we
know the rst j 1 entries of this row are all zero. Now, in Step 5 we increase rby 1, and then embark
on building a new nonzero row. In Step 6 we swap row rand rowi. In the rst jcolumns, the rst r 1
rows remain in reduced row-echelon form after the swap. In Step 7 we multiply row rby a nonzero scalar,
Version 2.30
Subsection RREF.RREF Reduced Row-Echelon Form 35
creating a 1 in the entry in column jof rowi, and not changing any other rows. This new leading 1 is the
rst nonzero entry in its row, and is located to the right of all the leading 1's in the preceding r 1 rows.
With Step 8 we insure that every entry in the column with this new leading 1 is now zero, as required for
reduced row-echelon form. Also, rows r+ 1 through mare now all zeros in the rst jcolumns, so we now
only have one new nonzero row, consistent with our increase of rby one. Furthermore, since the rst j 1
entries of row rare zero, the employment of the third row operation does not destroy any of the necessary
features of rows 1 through r 1 and rows r+ 1 through m, in columns 1 through j 1.
So at this stage, the rst jcolumns of the matrix are in reduced row-echelon form. When Step 2 nally
increasesjton+ 1, then the procedure is completed and the full ncolumns of the matrix are in reduced
row-echelon form, with the value of rcorrectly recording the number of nonzero rows.
The procedure given in the proof of Theorem REMEF [34] can be more precisely described using a
pseudo-code version of a computer program, as follows:
inputm,nandA
r 0
forj 1 ton
i r+ 1
whileimand [A]ij= 0
i i+ 1
ifi6=m+ 1
r r+ 1
swap rowsiandrofA(row op 1)
scale entry in row r, columnjofAto a leading 1 (row op 2)
fork 1 tom,k6=r
zero out entry in row k, columnjofA(row op 3 using row r)
outputrandA
Notice that as a practical matter the \and" used in the conditional statement of the while statement should
be of the \short-circuit" variety so that the array access that follows is not out-of-bounds.
So now we can put it all together. Begin with a system of linear equations (Denition SLE [11]), and
represent the system by its augmented matrix (Denition AM [30]). Use row operations (Denition RO
[31]) to convert this matrix into reduced row-echelon form (Denition RREF [33]), using the procedure
outlined in the proof of Theorem REMEF [34]. Theorem REMEF [34] also tells us we can always accomplish
this, and that the result is row-equivalent (Denition REM [31]) to the original augmented matrix. Since
the matrix in reduced-row echelon form has the same solution set, we can analyze the row-reduced version
instead of the original matrix, viewing it as the augmented matrix of a dierent system of equations. The
beauty of augmented matrices in reduced row-echelon form is that the solution sets to their corresponding
systems can be easily determined, as we will see in the next few examples and in the next section.
We will see through the course that almost every interesting property of a matrix can be discerned by
looking at a row-equivalent matrix in reduced row-echelon form. For this reason it is important to know
that the matrix Bguaranteed to exist by Theorem REMEF [34] is also unique.
Two proof techniques are applicable to the proof. First, head out and read two proof techniques:
Technique CD [770] and Technique U [771].
Theorem RREFU
Reduced Row-Echelon Form is Unique
Suppose that Ais anmnmatrix and that BandCaremnmatrices that are row-equivalent to A
and in reduced row-echelon form. Then B=C.
Proof We need to begin with no assumptions about any relationships between BandC, other than they
are both in reduced row-echelon form, and they are both row-equivalent to A.
Version 2.30
36 Section RREF Reduced Row-Echelon Form
IfBandCare both row-equivalent to A, then they are row-equivalent to each other. Repeated row
operations on a matrix combine the rows with each other using operations that are linear, and are identical
in each column. A key observation for this proof is that each individual row of Bis linearly related to the
rows ofC. This relationship is dierent for each row of B, but once we x a row, the relationship is the
same across columns. More precisely, there are scalars ik, 1i;kmsuch that for any 1 im,
1jn,
[B]ij=mX
k=1ik[C]kj
You should read this as saying that an entry of row iofB(in column j) is a linear function of the entries
of all the rows of Cthat are also in column j, and the scalars ( ik) depend on which row of Bwe are
considering (the isubscript on ik), but are the same for every column (no dependence on jinik). This
idea may be complicated now, but will feel more familiar once we discuss \linear combinations" (Denition
LCCV [109]) and moreso when we discuss \row spaces" (Denition RSM [278]). For now, spend some time
carefully working Exercise RREF.M40 [46], which is designed to illustrate the origins of this expression.
This completes our exploitation of the row-equivalence of BandC.
We now repeatedly exploit the fact that BandCare in reduced row-echelon form. Recall that a pivot
column is all zeros, except a single one. More carefully, if Ris a matrix in reduced row-echelon form, and
d`is the index of a pivot column, then [ R]kd`= 1 precisely when k=`and is otherwise zero. Notice also
that any entry of Rthat is both below the entry in row `andto the left of column d`is also zero (with
below and left understood to include equality). In other words, look at examples of matrices in reduced
row-echelon form and choose a leading 1 (with a box around it). The rest of the column is also zeros, and
the lower left \quadrant" of the matrix that begins here is totally zeros.
Assuming no relationship about the form of BandC, letBhavernonzero rows and denote the pivot
columns as D=fd1; d2; d3; :::; drg. ForCletr0denote the number of nonzero rows and denote the
pivot columns as D0=fd01; d02; d03; :::; d0r0g(Notation RREFA [33]). There are four steps in the proof,
and the rst three are about showing that BandChave the same number of pivot columns, in the same
places. In other words, the \primed" symbols are a necessary ction.
First Step. Suppose that d1<d0
1. Then
1 = [B]1d1Denition RREF [33]
=mX
k=11k[C]kd1
=mX
k=11k(0) d1<d0
1
= 0
The entries of Care all zero since they are left and below of the leading 1 in row 1 and column d0
1ofC.
This is a contradiction, so we know that d1d0
1. By an entirely similar argument, reversing the roles of
BandC, we could conclude that d1d0
1. Together this means that d1=d0
1.
Second Step. Suppose that we have determined that d1=d0
1,d2=d0
2,d3=d0
3, . . . ,dp=d0
p. Let's now
show thatdp+1=d0
p+1. Working towards a contradiction, suppose that dp+1<d0
p+1. For 1`p,
0 = [B]p+1;d`Denition RREF [33]
=mX
k=1p+1;k[C]kd`
=mX
k=1p+1;k[C]kd0
`
Version 2.30
Subsection RREF.RREF Reduced Row-Echelon Form 37
=p+1;`[C]`d0
`+mX
k=1
k6=`p+1;k[C]kd0
`Property CACN [758]
=p+1;`(1) +mX
k=1
k6=`p+1;k(0) Denition RREF [33]
=p+1;`
Now,
1 = [B]p+1;dp+1Denition RREF [33]
=mX
k=1p+1;k[C]kdp+1
=pX
k=1p+1;k[C]kdp+1+mX
k=p+1p+1;k[C]kdp+1Property AACN [758]
=pX
k=1(0) [C]kdp+1+mX
k=p+1p+1;k[C]kdp+1
=mX
k=p+1p+1;k[C]kdp+1
=mX
k=p+1p+1;k(0) dp+1<d0
p+1
= 0
This contradiction shows that dp+1d0
p+1. By an entirely similar argument, we could conclude that
dp+1d0
p+1, and therefore dp+1=d0
p+1.
Third Step. Now we establish that r=r0. Suppose that r0< r. By the arguments above, we know
thatd1=d0
1,d2=d0
2,d3=d0
3, . . . ,dr0=d0
r0. For 1`r0<r,
0 = [B]rd`Denition RREF [33]
=mX
k=1rk[C]kd`
=r0X
k=1rk[C]kd`+mX
k=r0+1rk[C]kd`Property AACN [758]
=r0X
k=1rk[C]kd`+mX
k=r0+1rk(0) Property AACN [758]
=r0X
k=1rk[C]kd`
=r0X
k=1rk[C]kd0
`
=r`[C]`d0
`+r0X
k=1
k6=`rk[C]kd0
`Property CACN [758]
Version 2.30
38 Section RREF Reduced Row-Echelon Form
=r`(1) +r0X
k=1
k6=`rk(0) Denition RREF [33]
=r`
Now examine the entries of row rofB,
[B]rj=mX
k=1rk[C]kj
=r0X
k=1rk[C]kj+mX
k=r0+1rk[C]kj Property CACN [758]
=r0X
k=1rk[C]kj+mX
k=r0+1rk(0) Denition RREF [33]
=r0X
k=1rk[C]kj
=r0X
k=1(0) [C]kj
= 0
So rowris a totally zero row, contradicting that this should be the bottommost nonzero row of B. So
r0r. By an entirely similar argument, reversing the roles of BandC, we would conclude that r0r
and therefore r=r0. Thus, combining the rst three steps we can say that D=D0. In other words, B
andChave the same pivot columns, in the same locations.
Fourth Step. In this nal step, we will not argue by contradiction. Our intent is to determine the
values of the ij. Notice that we can use the values of the diinterchangeably for BandC. Here we go,
1 = [B]idiDenition RREF [33]
=mX
k=1ik[C]kdi
=ii[C]idi+mX
k=1
k6=iik[C]kdiProperty CACN [758]
=ii(1) +mX
k=1
k6=iik(0) Denition RREF [33]
=ii
and for`6=i
0 = [B]id`Denition RREF [33]
=mX
k=1ik[C]kd`
=i`[C]`d`+mX
k=1
k6=`ik[C]kd`Property CACN [758]
Version 2.30
Subsection RREF.RREF Reduced Row-Echelon Form 39
=i`(1) +mX
k=1
k6=`ik(0) Denition RREF [33]
=i`
Finally, having determined the values of the ij, we can show that B=C. For 1im, 1jn,
[B]ij=mX
k=1ik[C]kj
=ii[C]ij+mX
k=1
k6=iik[C]kj Property CACN [758]
= (1) [C]ij+mX
k=1
k6=i(0) [C]kj
= [C]ij
SoBandChave equal values in every entry, and so are the same matrix.
We will now run through some examples of using these denitions and theorems to solve some systems
of equations. From now on, when we have a matrix in reduced row-echelon form, we will mark the leading
1's with a small box. In your work, you can box 'em, circle 'em or write 'em in a dierent color | just
identify 'em somehow. This device will prove very useful later and is a very good habit to start developing
right now.
Example SAB
Solutions for Archetype B
Let's nd the solutions to the following system of equations,
7x1 6x2 12x3= 33
5x1+ 5x2+ 7x3= 24
x1+ 4x3= 5
First, form the augmented matrix,
2
4 7 6 12 33
5 5 7 24
1 0 4 53
5
and work to reduced row-echelon form, rst with j= 1,
R1$R3 !2
41 0 4 5
5 5 7 24
7 6 12 333
5 5R1+R2 !2
41 0 4 5
0 5 13 1
7 6 12 333
5
7R1+R3 !2
41 0 4 5
0 5 13 1
0 6 16 23
5
Now, withj= 2,
1
5R2 !2
41 0 4 5
0 1 13
5 1
5
0 6 16 23
56R2+R3 !2
410 4 5
01 13
5 1
5
0 02
54
53
5
Version 2.30
40 Section RREF Reduced Row-Echelon Form
And nally, with j= 3,
5
2R3 !2
410 4 5
01 13
5 1
5
0 0 1 23
513
5R3+R2 !2
410 4 5
010 5
0 0 1 23
5
4R3+R1 !2
410 0 3
010 5
0 0 1 23
5
This is now the augmented matrix of a very simple system of equations, namely x1= 3,x2= 5,x3= 2,
which has an obvious solution. Furthermore, we can see that this is the only solution to this system, so we
have determined the entire solution set,
S=8
<
:2
4 3
5
23
59
=
;
You might compare this example with the procedure we used in Example US [16].
Archetypes A and B are meant to contrast each other in many respects. So let's solve Archetype A
now.
Example SAA
Solutions for Archetype A
Let's nd the solutions to the following system of equations,
x1 x2+ 2x3= 1
2x1+x2+x3= 8
x1+x2= 5
First, form the augmented matrix,
2
41 1 2 1
2 1 1 8
1 1 0 53
5
and work to reduced row-echelon form, rst with j= 1,
2R1+R2 !2
41 1 2 1
0 3 3 6
1 1 0 53
5 1R1+R3 !2
41 1 2 1
0 3 3 6
0 2 2 43
5
Now, withj= 2,
1
3R2 !2
41 1 2 1
0 1 1 2
0 2 2 43
51R2+R1 !2
410 1 3
0 1 1 2
0 2 2 43
5
2R2+R3 !2
410 1 3
01 1 2
0 0 0 03
5
The system of equations represented by this augmented matrix needs to be considered a bit dierently
than that for Archetype B. First, the last row of the matrix is the equation 0 = 0, which is always true, so
Version 2.30
Subsection RREF.RREF Reduced Row-Echelon Form 41
it imposes no restrictions on our possible solutions and therefore we can safely ignore it as we analyze the
other two equations. These equations are,
x1+x3= 3
x2 x3= 2:
While this system is fairly easy to solve, it also appears to have a multitude of solutions. For example,
choosex3= 1 and see that then x1= 2 andx2= 3 will together form a solution. Or choose x3= 0, and
then discover that x1= 3 andx2= 2 lead to a solution. Try it yourself: pick anyvalue ofx3you please,
and gure out what x1andx2should be to make the rst and second equations (respectively) true. We'll
wait while you do that. Because of this behavior, we say that x3is a \free" or \independent" variable. But
why do we vary x3and not some other variable? For now, notice that the third column of the augmented
matrix does not have any leading 1's in its column. With this idea, we can rearrange the two equations,
solving each for the variable that corresponds to the leading 1 in that row.
x1= 3 x3
x2= 2 +x3
To write the set of solution vectors in set notation, we have
S=8
<
:2
43 x3
2 +x3
x33
5x32C9
=
;
We'll learn more in the next section about systems with innitely many solutions and how to express their
solution sets. Right now, you might look back at Example IS [17].
Example SAE
Solutions for Archetype E
Let's nd the solutions to the following system of equations,
2x1+x2+ 7x3 7x4= 2
3x1+ 4x2 5x3 6x4= 3
x1+x2+ 4x3 5x4= 2
First, form the augmented matrix,
2
42 1 7 7 2
3 4 5 6 3
1 1 4 5 23
5
and work to reduced row-echelon form, rst with j= 1,
R1$R3 !2
41 1 4 5 2
3 4 5 6 3
2 1 7 7 23
53R1+R2 !2
41 1 4 5 2
0 7 7 21 9
2 1 7 7 23
5
2R1+R3 !2
41 1 4 5 2
0 7 7 21 9
0 1 1 3 23
5
Now, withj= 2,
R2$R3 !2
41 1 4 5 2
0 1 1 3 2
0 7 7 21 93
5 1R2 !2
411 4 5 2
0 1 1 3 2
0 7 7 21 93
5
Version 2.30
42 Section RREF Reduced Row-Echelon Form
1R2+R1 !2
410 3 2 0
0 1 1 3 2
0 7 7 21 93
5 7R2+R3 !2
410 3 2 0
011 3 2
0 0 0 0 53
5
And nally, with j= 4,
1
5R3 !2
410 3 2 0
011 3 2
0 0 0 0 13
5 2R3+R2 !2
410 3 2 0
011 3 0
0 0 0 0 13
5
Let's analyze the equations in the system represented by this augmented matrix. The third equation will
read 0 = 1. This is patently false, all the time. No choice of values for our variables will ever make it
true. We're done. Since we cannot even make the last equation true, we have no hope of making all of
the equations simultaneously true. So this system has no solutions, and its solution set is the empty set,
;=fg(Denition ES [761]).
Notice that we could have reached this conclusion sooner. After performing the row operation 7R2+
R3, we can see that the third equation reads 0 = 5, a false statement. Since the system represented by
this matrix has no solutions, none of the systems represented has any solutions. However, for this example,
we have chosen to bring the matrix fully to reduced row-echelon form for the practice.
These three examples (Example SAB [39], Example SAA [40], Example SAE [41]) illustrate the full
range of possibilities for a system of linear equations | no solutions, one solution, or innitely many
solutions. In the next section we'll examine these three scenarios more closely.
Denition RR
Row-Reducing
Torow-reduce the matrix Ameans to apply row operations to Aand arrive at a row-equivalent matrix
Bin reduced row-echelon form. 4
So the term row-reduce is used as a verb. Theorem REMEF [34] tells us that this process will always
be successful and Theorem RREFU [35] tells us that the result will be unambiguous. Typically, the analysis
ofAwill proceed by analyzing Band applying theorems whose hypotheses include the row-equivalence of
AandB.
After some practice by hand, you will want to use your favorite computing device to do the computations
required to bring a matrix to reduced row-echelon form (Exercise RREF.C30 [46]). See: Computation
RR.MMA [745] Computation RR.TI86 [750] Computation RR.TI83 [751] Computation RR.SAGE
[753]
Subsection READ
Reading Questions
1. Is the matrix below in reduced row-echelon form? Why or why not?
2
41 5 0 6 8
0 0 1 2 0
0 0 0 0 13
5
2. Use row operations to convert the matrix below to reduced row-echelon form and report the nal
matrix. 2
42 1 8
1 1 1
2 5 43
5
Version 2.30
Subsection RREF.READ Reading Questions 43
3. Find all the solutions to the system below by using an augmented matrix and row operations. Report
your nal matrix in reduced row-echelon form and the set of solutions.
2x1+ 3x2 x3= 0
x1+ 2x2+x3= 3
x1+ 3x2+ 3x3= 7
Version 2.30
44 Section RREF Reduced Row-Echelon Form
Subsection EXC
Exercises
C05 Each archetype below is a system of equations. Form the augmented matrix of the system of
equations, convert the matrix to reduced row-echelon form by using equation operations and then describe
the solution set of the original system of equations.
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]
Archetype E [799]
Archetype F [803]
Archetype G [808]
Archetype H [812]
Archetype I [816]
Archetype J [820]
Contributed by Robert Beezer
For problems C10{C19, nd all solutions to the system of linear equations. Use your favorite computing
device to row-reduce the augmented matrices for the systems, and write the solutions as a set, using correct
set notation.
C10
2x1 3x2+x3+ 7x4= 14
2x1+ 8x2 4x3+ 5x4= 1
x1+ 3x2 3x3= 4
5x1+ 2x2+ 3x3+ 4x4= 19
Contributed by Robert Beezer Solution [48]
C11
3x1+ 4x2 x3+ 2x4= 6
x1 2x2+ 3x3+x4= 2
10x2 10x3 x4= 1
Contributed by Robert Beezer Solution [48]
C12
2x1+ 4x2+ 5x3+ 7x4= 26
x1+ 2x2+x3 x4= 4
2x1 4x2+x3+ 11x4= 10
Contributed by Robert Beezer Solution [48]
C13
x1+ 2x2+ 8x3 7x4= 2
Version 2.30
Subsection RREF.EXC Exercises 45
3x1+ 2x2+ 12x3 5x4= 6
x1+x2+x3 5x4= 10
Contributed by Robert Beezer Solution [49]
C14
2x1+x2+ 7x3 2x4= 4
3x1 2x2+ 11x4= 13
x1+x2+ 5x3 3x4= 1
Contributed by Robert Beezer Solution [49]
C15
2x1+ 3x2 x3 9x4= 16
x1+ 2x2+x3= 0
x1+ 2x2+ 3x3+ 4x4= 8
Contributed by Robert Beezer Solution [49]
C16
2x1+ 3x2+ 19x3 4x4= 2
x1+ 2x2+ 12x3 3x4= 1
x1+ 2x2+ 8x3 5x4= 1
Contributed by Robert Beezer Solution [50]
C17
x1+ 5x2= 8
2x1+ 5x2+ 5x3+ 2x4= 9
3x1 x2+ 3x3+x4= 3
7x1+ 6x2+ 5x3+x4= 30
Contributed by Robert Beezer Solution [50]
C18
x1+ 2x2 4x3 x4= 32
x1+ 3x2 7x3 x5= 33
x1+ 2x3 2x4+ 3x5= 22
Contributed by Robert Beezer Solution [50]
Version 2.30
46 Section RREF Reduced Row-Echelon Form
C19
2x1+x2= 6
x1 x2= 2
3x1+ 4x2= 4
3x1+ 5x2= 2
Contributed by Robert Beezer Solution [51]
For problems C30{C33, row-reduce the matrix without the aid of a calculator, indicating the row
operations you are using at each step using the notation of Denition RO [31].
C30
2
42 1 5 10
1 3 1 2
4 2 6 123
5
Contributed by Robert Beezer Solution [51]
C31
2
41 2 4
3 1 3
2 1 73
5
Contributed by Robert Beezer Solution [51]
C32
2
41 1 1
4 3 2
3 2 13
5
Contributed by Robert Beezer Solution [52]
C33
2
41 2 1 1
2 4 1 4
1 2 3 53
5
Contributed by Robert Beezer Solution [52]
M40 Consider the two 3 4 matrices below
B=2
41 3 2 2
1 2 1 1
1 5 8 33
5 C=2
41 2 1 2
1 1 4 0
1 1 4 13
5
(a) Row-reduce each matrix and determine that the reduced row-echelon forms of BandCare
identical. From this argue that BandCare row-equivalent.
(b) In the proof of Theorem RREFU [35], we begin by arguing that entries of row-equivalent matrices
are related by way of certain scalars and sums. In this example, we would write that entries of Bfrom row
ithat are in column jare linearly related to the entries of Cin columnjfrom all three rows
[B]ij=i1[C]1j+i2[C]2j+i3[C]3j 1j4
Version 2.30
Subsection RREF.EXC Exercises 47
For each 1i3 nd the corresponding three scalars in this relationship. So your answer will be nine
scalars, determined three at a time.
Contributed by Robert Beezer Solution [52]
M45 You keep a number of lizards, mice and peacocks as pets. There are a total of 108 legs and 30 tails
in your menagerie. You have twice as many mice as lizards. How many of each creature do you have?
Contributed by Chris Black Solution [53]
M50 A parking lot has 66 vehicles (cars, trucks, motorcycles and bicycles) in it. There are four times
as many cars as trucks. The total number of tires (4 per car or truck, 2 per motorcycle or bicycle) is 252.
How many cars are there? How many bicycles?
Contributed by Robert Beezer Solution [53]
T10 Prove that each of the three row operations (Denition RO [31]) is reversible. More precisely, if
the matrix Bis obtained from Aby application of a single row operation, show that there is a single row
operation that will transform Bback intoA.
Contributed by Robert Beezer Solution [54]
T11 Suppose that A,BandCaremnmatrices. Use the denition of row-equivalence (Denition
REM [31]) to prove the following three facts.
1.Ais row-equivalent to A.
2. IfAis row-equivalent to B, thenBis row-equivalent to A.
3. IfAis row-equivalent to B, andBis row-equivalent to C, thenAis row-equivalent to C.
A relationship that satises these three properties is known as an equivalence relation , an important
idea in the study of various algebras. This is a formal way of saying that a relationship behaves like
equality, without requiring the relationship to be as strict as equality itself. We'll see it again in Theorem
SER [494].
Contributed by Robert Beezer
T12 Suppose that Bis anmnmatrix in reduced row-echelon form. Build a new, likely smaller, k`
matrixCas follows. Keep any collection of kadjacent rows, km. From these rows, keep columns 1
through`,`n. Prove that Cis in reduced row-echelon form.
Contributed by Robert Beezer
T13 Generalize Exercise RREF.T12 [47] by just keeping any krows, and not requiring the rows to be
adjacent. Prove that any such matrix Cis in reduced row-echelon form.
Contributed by Robert Beezer
Version 2.30
48 Section RREF Reduced Row-Echelon Form
Subsection SOL
Solutions
C10 Contributed by Robert Beezer Statement [44]
The augmented matrix row-reduces to
2
666410 0 0 1
010 0 3
0 0 10 4
0 0 0 1 13
7775
This augmented matrix represents the linear system x1= 1,x2= 3,x3= 4,x4= 1, which clearly has
only one possible solution. We can write this solution set then as
S=8
>><
>>:2
6641
3
4
13
7759
>>=
>>;
C11 Contributed by Robert Beezer Statement [44]
The augmented matrix row-reduces to
2
410 1 4 =5 0
01 1 1=10 0
0 0 0 0 13
5
Row 3 represents the equation 0 = 1, which is patently false, so the original system has no solutions. We
can express the solution set as the empty set, ;=fg.
C12 Contributed by Robert Beezer Statement [44]
The augmented matrix row-reduces to
2
412 0 4 2
0 0 1 3 6
0 0 0 0 03
5
In the spirit of Example SAA [40], we can express the innitely many solutions of this system compactly
with set notation. The key is to express certain variables in terms of others. More specically, each pivot
column number is the index of a variable that can be written in terms of the variables whose indices are
non-pivot columns. Or saying the same thing: for each iinD, we can nd an expression for xiin terms
of the variables without their index in D. HereD=f1;3g, so
x1= 2 2x2+ 4x4
x3= 6 3x4
As a set, we write the solutions precisely as
8
>><
>>:2
6642 2x2+ 4x4
x2
6 3x4
x43
775x2; x42C9
>>=
>>;
Version 2.30
Subsection RREF.SOL Solutions 49
C13 Contributed by Robert Beezer Statement [44]
The augmented matrix of the system of equations is
2
41 2 8 7 2
3 2 12 5 6
1 1 1 5 103
5
which row-reduces to2
410 2 1 0
013 4 0
0 0 0 0 13
5
Row 3 represents the equation 0 = 1, which is patently false, so the original system has no solutions. We
can express the solution set as the empty set, ;=fg.
C14 Contributed by Robert Beezer Statement [45]
The augmented matrix of the system of equations is
2
42 1 7 2 4
3 2 0 11 13
1 1 5 3 13
5
which row-reduces to 2
410 2 1 3
013 4 2
0 0 0 0 03
5
In the spirit of Example SAA [40], we can express the innitely many solutions of this system compactly
with set notation. The key is to express certain variables in terms of others. More specically, each pivot
column number is the index of a variable that can be written in terms of the variables whose indices are
non-pivot columns. Or saying the same thing: for each iinD, we can nd an expression for xiin terms of
the variables without their index in D. HereD=f1;2g, so rearranging the equations represented by the
two nonzero rows to gain expressions for the variables x1andx2yields the solution set,
S=8
>><
>>:2
6643 2x3 x4
2 3x3+ 4x4
x3
x43
775x3; x42C9
>>=
>>;
C15 Contributed by Robert Beezer Statement [45]
The augmented matrix of the system of equations is
2
42 3 1 9 16
1 2 1 0 0
1 2 3 4 83
5
which row-reduces to2
410 0 2 3
010 3 5
0 0 1 4 73
5
In the spirit of Example SAA [40], we can express the innitely many solutions of this system compactly
with set notation. The key is to express certain variables in terms of others. More specically, each pivot
column number is the index of a variable that can be written in terms of the variables whose indices are
non-pivot columns. Or saying the same thing: for each iinD, we can nd an expression for xiin terms
Version 2.30
50 Section RREF Reduced Row-Echelon Form
of the variables without their index in D. HereD=f1;2;3g, so rearranging the equations represented by
the three nonzero rows to gain expressions for the variables x1,x2andx3yields the solution set,
S=8
>><
>>:2
6643 2x4
5 + 3x4
7 4x4
x43
775x42C9
>>=
>>;
C16 Contributed by Robert Beezer Statement [45]
The augmented matrix of the system of equations is
2
42 3 19 4 2
1 2 12 3 1
1 2 8 5 13
5
which row-reduces to2
410 2 1 0
015 2 0
0 0 0 0 13
5
Row 3 represents the equation 0 = 1, which is patently false, so the original system has no solutions. We
can express the solution set as the empty set, ;=fg.
C17 Contributed by Robert Beezer Statement [45]
We row-reduce the augmented matrix of the system of equations,
2
664 1 5 0 0 8
2 5 5 2 9
3 1 3 1 3
7 6 5 1 303
775RREF !2
666410 0 0 3
010 0 1
0 0 10 2
0 0 0 1 53
7775
This augmented matrix represents the linear system x1= 3,x2= 1,x3= 2,x4= 5, which clearly has
only one possible solution. We can write this solution set then as
S=8
>><
>>:2
6643
1
2
53
7759
>>=
>>;
C18 Contributed by Robert Beezer Statement [45]
We row-reduce the augmented matrix of the system of equations,
2
41 2 4 1 0 32
1 3 7 0 1 33
1 0 2 2 3 223
5RREF !2
410 2 0 5 6
01 3 0 2 9
0 0 0 1 1 83
5
In the spirit of Example SAA [40], we can express the innitely many solutions of this system compactly
with set notation. The key is to express certain variables in terms of others. More specically, each pivot
column number is the index of a variable that can be written in terms of the variables whose indices are
non-pivot columns. Or saying the same thing: for each iinD, we can nd an expression for xiin terms
of the variables without their index in D. HereD=f1;2;4g, so
x1+ 2x3+ 5x5= 6!x1= 6 2x3 5x5
Version 2.30
Subsection RREF.SOL Solutions 51
x2 3x3 2x5= 9!x2= 9 + 3x3+ 2x5
x4+x5= 8!x4= 8 x5
As a set, we write the solutions precisely as
S=8
>>>><
>>>>:2
666646 2x3 5x5
9 + 3x3+ 2x5
x3
8 x5
x53
77775x3; x52C9
>>>>=
>>>>;
C19 Contributed by Robert Beezer Statement [46]
We form the augmented matrix of the system,
2
6642 1 6
1 1 2
3 4 4
3 5 23
775
which row-reduces to
2
66410 4
01 2
0 0 0
0 0 03
775
This augmented matrix represents the linear system x1= 4,x2= 2, 0 = 0, 0 = 0, which clearly has only
one possible solution. We can write this solution set then as
S=4
2
C30 Contributed by Robert Beezer Statement [46]
2
42 1 5 10
1 3 1 2
4 2 6 123
5R1$R2 !2
41 3 1 2
2 1 5 10
4 2 6 123
5
2R1+R2 !2
41 3 1 2
0 7 7 14
4 2 6 123
5 4R1+R3 !2
41 3 1 2
0 7 7 14
0 10 10 203
5
1
7R2 !2
41 3 1 2
0 1 1 2
0 10 10 203
53R2+R1 !2
41 0 2 4
0 1 1 2
0 10 10 203
5
10R2+R3 !2
410 2 4
011 2
0 0 0 03
5
C31 Contributed by Robert Beezer Statement [46]
2
41 2 4
3 1 3
2 1 73
53R1+R2 !2
41 2 4
0 5 15
2 1 73
5
Version 2.30
52 Section RREF Reduced Row-Echelon Form
2R1+R3 !2
41 2 4
0 5 15
0 5 153
51
5R2 !2
41 2 4
0 1 3
0 5 153
5
2R2+R1 !2
41 0 2
0 1 3
0 5 153
5 5R2+R3 !2
410 2
01 3
0 0 03
5
C32 Contributed by Robert Beezer Statement [46]
Following the algorithm of Theorem REMEF [34], and working to create pivot columns from left to right,
we have
2
41 1 1
4 3 2
3 2 13
54R1+R2 !2
41 1 1
0 1 2
3 2 13
5 3R1+R3 !2
41 1 1
0 1 2
0 1 23
5
1R2+R1 !2
41 0 1
0 1 2
0 1 23
51R2+R3 !2
410 1
01 2
0 0 03
5
C33 Contributed by Robert Beezer Statement [46]
Following the algorithm of Theorem REMEF [34], and working to create pivot columns from left to right,
we have
2
41 2 1 1
2 4 1 4
1 2 3 53
5 2R1+R2 !2
41 2 1 1
0 0 1 6
1 2 3 53
5
1R1+R3 !2
412 1 1
0 0 1 6
0 0 2 43
51R2+R1 !2
412 0 5
0 0 1 6
0 0 2 43
5
2R2+R3 !2
412 0 5
0 0 1 6
0 0 0 83
5 1
8R3 !2
412 0 5
0 0 16
0 0 0 13
5
6R3+R2 !2
412 0 5
0 0 10
0 0 0 13
5 5R3+R1 !2
412 0 0
0 0 10
0 0 0 13
5
M40 Contributed by Robert Beezer Statement [46]
(a) LetRbe the common reduced row-echelon form of BandC. A sequence of row operations converts
BtoRand a second sequence of row operations converts CtoR. If we \reverse" the second sequence's
order, and reverse each individual row operation (see Exercise RREF.T10 [47]) then we can begin with
B, convert to Rwith the rst sequence, and then convert to Cwith the reversed sequence. Satisfying
Denition REM [31] we can say BandCare row-equivalent matrices.
(b) We will work this carefully for the rst row of Band just give the solution for the next two rows.
For row 1 of Btakei= 1 and we have
[B]1j=11[C]1j+12[C]2j+13[C]3j 1j4
If we substitute the four values for jwe arrive at four linear equations in the three unknowns 11;12;13,
(j= 1) [B]11=11[C]11+12[C]21+13[C]31) 1 =11(1) +12(1) +13( 1)
Version 2.30
Subsection RREF.SOL Solutions 53
(j= 2) [B]12=11[C]12+12[C]22+13[C]32) 3 =11(2) +12(1) +13( 1)
(j= 3) [B]13=11[C]13+12[C]23+13[C]33) 2 =11(1) +12(4) +13( 4)
(j= 4) [B]14=11[C]14+12[C]24+13[C]34) 2 =11(2) +12(0) +13(1)
We form the augmented matrix of this system and row-reduce to nd the solutions,
2
6641 1 1 1
2 1 1 3
1 4 4 2
2 0 1 23
775RREF !2
66410 0 2
010 3
0 0 1 2
0 0 0 03
775
So the unique solution is 11= 2,12= 3,13= 2. Entirely similar work will lead you to
21= 1 22= 1 23= 1
and
31= 4 32= 8 33= 5
M45 Contributed by Chris Black Statement [47]
Letl;m;p denote the number of lizards, mice and peacocks. Then the statements from the problem yield
the equations:
4l+ 4m+ 2p= 108
l+m+p= 30
2l m= 0
We form the augmented matrix for this system and row-reduce
2
44 4 2 108
1 1 1 30
2 1 0 03
5RREF !2
410 0 8
010 16
0 0 163
5
From the row-reduced matrix, we see that we have an equivalent system l= 8,m= 16, andp= 6, which
means that you have 8 lizards, 16 mice and 6 peacocks.
M50 Contributed by Robert Beezer Statement [47]
Letc; t; m; b denote the number of cars, trucks, motorcycles, and bicycles. Then the statements from the
problem yield the equations:
c+t+m+b= 66
c 4t= 0
4c+ 4t+ 2m+ 2b= 252
We form the augmented matrix for this system and row-reduce
2
41 1 1 1 66
1 4 0 0 0
4 4 2 2 2523
5RREF !2
410 0 0 48
010 0 12
0 0 11 63
5
The rst row of the matrix represents the equation c= 48, so there are 48 cars. The second row of the
matrix represents the equation t= 12, so there are 12 trucks. The third row of the matrix represents the
Version 2.30
54 Section RREF Reduced Row-Echelon Form
equationm+b= 6 so there are anywhere from 0 to 6 bicycles. We can also say that bis a free variable,
but the context of the problem limits it to 7 integer values since you cannot have a negative number of
motorcycles.
T10 Contributed by Robert Beezer Statement [47]
If we can reverse each row operation individually, then we can reverse a sequence of row operations. The
operations that reverse each operation are listed below, using our shorthand notation. Notice how requiring
the scalarto be non-zero makes the second operation reversible.
Ri$RjRi$Rj
Ri; 6= 01
Ri
Ri+Rj Ri+Rj
Version 2.30
Section TSS Types of Solution Sets 55
Section TSS
Types of Solution Sets
We will now be more careful about analyzing the reduced row-echelon form derived from the augmented
matrix of a system of linear equations. In particular, we will see how to systematically handle the situation
when we have innitely many solutions to a system, and we will prove that every system of linear equations
has either zero, one or innitely many solutions. With these tools, we will be able to solve any system by
a well-described method.
Subsection CS
Consistent Systems
The computer scientist Donald Knuth said, \Science is what we understand well enough to explain to a
computer. Art is everything else." In this section we'll remove solving systems of equations from the realm
of art, and into the realm of science. We begin with a denition.
Denition CS
Consistent System
A system of linear equations is consistent if it has at least one solution. Otherwise, the system is called
inconsistent . 4
We will want to rst recognize when a system is inconsistent or consistent, and in the case of consistent
systems we will be able to further rene the types of solutions possible. We will do this by analyzing the
reduced row-echelon form of a matrix, using the value of r, and the sets of column indices, DandF, rst
dened back in Denition RREF [33].
Use of the notation for the elements of DandFcan be a bit confusing, since we have subscripted
variables that are in turn equal to integers used to index the matrix. However, many questions about
matrices and systems of equations can be answered once we know r,DandF. The choice of the letters D
andFrefer to our upcoming denition of dependent and free variables (Denition IDV [57]). An example
will help us begin to get comfortable with this aspect of reduced row-echelon form.
Example RREFN
Reduced row-echelon form notation
For the 59 matrix
B=2
66666415 0 0 2 8 0 5 1
0 0 10 4 7 0 2 0
0 0 0 13 9 0 3 6
0 0 0 0 0 0 14 2
0 0 0 0 0 0 0 0 03
777775
in reduced row-echelon form we have
r= 4
d1= 1 d2= 3 d3= 4 d4= 7
f1= 2 f2= 5 f3= 6 f4= 8 f5= 9
Notice that the sets
D=fd1; d2; d3; d4g=f1;3;4;7g F=ff1; f2; f3; f4; f5g=f2;5;6;8;9g
Version 2.30
56 Section TSS Types of Solution Sets
have nothing in common and together account for all of the columns of B(we say it is a partition of the
set of column indices).
The number ris the single most important piece of information we can get from the reduced row-
echelon form of a matrix. It is dened as the number of nonzero rows, but since each nonzero row has a
leading 1, it is also the number of leading 1's present. For each leading 1, we have a pivot column, so ris
also the number of pivot columns. Repeating ourselves, ris the number of nonzero rows, the number of
leading 1's andthe number of pivot columns. Across dierent situations, each of these interpretations of
the meaning of rwill be useful.
Before proving some theorems about the possibilities for solution sets to systems of equations, let's
analyze one particular system with an innite solution set very carefully as an example. We'll use this
technique frequently, and shortly we'll rene it slightly.
Archetypes I and J are both fairly large for doing computations by hand (though not impossibly large).
Their properties are very similar, so we will frequently analyze the situation in Archetype I, and leave you
the joy of analyzing Archetype J yourself. So work through Archetype I with the text, by hand and/or
with a computer, and then tackle Archetype J yourself (and check your results with those listed). Notice
too that the archetypes describing systems of equations each lists the values of r,DandF. Here we go. . .
Example ISSI
Describing innite solution sets, Archetype I
Archetype I [816] is the system of m= 4 equations in n= 7 variables.
x1+ 4x2 x4+ 7x6 9x7= 3
2x1+ 8x2 x3+ 3x4+ 9x5 13x6+ 7x7= 9
2x3 3x4 4x5+ 12x6 8x7= 1
x1 4x2+ 2x3+ 4x4+ 8x5 31x6+ 37x7= 4
This system has a 4 8 augmented matrix that is row-equivalent to the following matrix (check this!), and
which is in reduced row-echelon form (the existence of this matrix is guaranteed by Theorem REMEF [34]
and its uniqueness is guaranteed by Theorem RREFU [35]),
2
66414 0 0 2 1 3 4
0 0 10 1 3 5 2
0 0 0 12 6 6 1
0 0 0 0 0 0 0 03
775
So we nd that r= 3 and
D=fd1; d2; d3g=f1;3;4g F=ff1; f2; f3; f4; f5g=f2;5;6;7;8g
Letidenote one of the r= 3 non-zero rows, and then we see that we can solve the corresponding equation
represented by this row for the variable xdiand write it as a linear function of the variables xf1; xf2; xf3; xf4
(notice that f5= 8 does not reference a variable). We'll do this now, but you can already see how the
subscripts upon subscripts takes some getting used to.
(i= 1) xd1=x1= 4 4x2 2x5 x6+ 3x7
(i= 2) xd2=x3= 2 x5+ 3x6 5x7
(i= 3) xd3=x4= 1 2x5+ 6x6 6x7
Each element of the set F=ff1; f2; f3; f4; f5g=f2;5;6;7;8gis the index of a variable, except for
f5= 8. We refer to xf1=x2,xf2=x5,xf3=x6andxf4=x7as \free" (or \independent") variables since
they are allowed to assume any possible combination of values that we can imagine and we can continue on
Version 2.30
Subsection TSS.CS Consistent Systems 57
to build a solution to the system by solving individual equations for the values of the other (\dependent")
variables.
Each element of the set D=fd1; d2; d3g=f1;3;4gis the index of a variable. We refer to the variables
xd1=x1,xd2=x3andxd3=x4as \dependent" variables since they depend on the independent variables.
More precisely, for each possible choice of values for the independent variables we get exactly one set of
values for the dependent variables that combine to form a solution of the system.
To express the solutions as a set, we write
8
>>>>>>>><
>>>>>>>>:2
6666666644 4x2 2x5 x6+ 3x7
x2
2 x5+ 3x6 5x7
1 2x5+ 6x6 6x7
x5
x6
x73
777777775x2; x5; x6; x72C9
>>>>>>>>=
>>>>>>>>;
The condition that x2; x5; x6; x72Cis how we specify that the variables x2; x5; x6; x7are \free" to
assume any possible values.
This systematic approach to solving a system of equations will allow us to create a precise description
of the solution set for any consistent system once we have found the reduced row-echelon form of the
augmented matrix. It will work just as well when the set of free variables is empty and we get just a
single solution. And we could program a computer to do it! Now have a whack at Archetype J (Exercise
TSS.T10 [65]), mimicking the discussion in this example. We'll still be here when you get back.
Using the reduced row-echelon form of the augmented matrix of a system of equations to determine
the nature of the solution set of the system is a very key idea. So let's look at one more example like the
last one. But rst a denition, and then the example. We mix our metaphors a bit when we call variables
free versus dependent. Maybe we should call dependent variables \enslaved"?
Denition IDV
Independent and Dependent Variables
SupposeAis the augmented matrix of a consistent system of linear equations and Bis a row-equivalent
matrix in reduced row-echelon form. Suppose jis the index of a column of Bthat contains the leading 1
for some row (i.e. column jis a pivot column). Then the variable xjisdependent . A variable that is not
dependent is called independent orfree. 4
If you studied this denition carefully, you might wonder what to do if the system has nvariables and
columnn+ 1 is a pivot column? We will see shortly, by Theorem RCLS [58], that this never happens for
a consistent system.
Example FDV
Free and dependent variables
Consider the system of ve equations in ve variables,
x1 x2 2x3+x4+ 11x5= 13
x1 x2+x3+x4+ 5x5= 16
2x1 2x2+x4+ 10x5= 21
2x1 2x2 x3+ 3x4+ 20x5= 38
2x1 2x2+x3+x4+ 8x5= 22
Version 2.30
58 Section TSS Types of Solution Sets
whose augmented matrix row-reduces to
2
6666641 1 0 0 3 6
0 0 10 2 1
0 0 0 1 4 9
0 0 0 0 0 0
0 0 0 0 0 03
777775
There are leading 1's in columns 1, 3 and 4, so D=f1;3;4g. From this we know that the variables x1,
x3andx4will be dependent variables, and each of the r= 3 nonzero rows of the row-reduced matrix
will yield an expression for one of these three variables. The set Fis all the remaining column indices,
F=f2;5;6g. That 62Frefers to the column originating from the vector of constants, but the remaining
indices inFwill correspond to free variables, so x2andx5(the remaining variables) are our free variables.
The resulting three equations that describe our solution set are then,
(xd1=x1) x1= 6 +x2 3x5
(xd2=x3) x3= 1 + 2x5
(xd3=x4) x4= 9 4x5
Make sure you understand where these three equations came from, and notice how the location of the
leading 1's determined the variables on the left-hand side of each equation. We can compactly describe
the solution set as,
S=8
>>>><
>>>>:2
666646 +x2 3x5
x2
1 + 2x5
9 4x5
x53
77775x2; x52C9
>>>>=
>>>>;
Notice how we express the freedom for x2andx5:x2; x52C.
Sets are an important part of algebra, and we've seen a few already. Being comfortable with sets is
important for understanding and writing proofs. If you haven't already, pay a visit now to Section SET
[761].
We can now use the values of m,n,r, and the independent and dependent variables to categorize the
solution sets for linear systems through a sequence of theorems. Through the following sequence of proofs,
you will want to consult three proof techniques. See Technique E [768]. See Technique N [769]. See
Technique CP [769].
First we have an important theorem that explores the distinction between consistent and inconsistent
linear systems.
Theorem RCLS
Recognizing Consistency of a Linear System
SupposeAis the augmented matrix of a system of linear equations with nvariables. Suppose also that B
is a row-equivalent matrix in reduced row-echelon form with rnonzero rows. Then the system of equations
is inconsistent if and only if the leading 1 of row ris located in column n+ 1 ofB.
Proof (() The rst half of the proof begins with the assumption that the leading 1 of row ris located in
columnn+1 ofB. Then row rofBbegins with nconsecutive zeros, nishing with the leading 1. This is a
representation of the equation 0 = 1, which is false. Since this equation is false for any collection of values
we might choose for the variables, there are no solutions for the system of equations, and it is inconsistent.
()) For the second half of the proof, we wish to show that if we assume the system is inconsistent,
then the nal leading 1 is located in the last column. But instead of proving this directly, we'll form the
logically equivalent statement that is the contrapositive, and prove that instead (see Technique CP [769]).
Version 2.30
Subsection TSS.CS Consistent Systems 59
Turning the implication around, and negating each portion, we arrive at the logically equivalent statement:
If the leading 1 of row ris not in column n+ 1, then the system of equations is consistent.
If the leading 1 for row ris located somewhere in columns 1 through n, then every preceding row's
leading 1 is also located in columns 1 through n. In other words, since the last leading 1 is not in the
last column, no leading 1 for any row is in the last column, due to the echelon layout of the leading 1's
(Denition RREF [33]). We will now construct a solution to the system by setting each dependent variable
to the entry of the nal column for the row with the corresponding leading 1, and setting each free variable
to zero. That sentence is pretty vague, so let's be more precise. Using our notation for the sets DandF
from the reduced row-echelon form (Notation RREFA [33]):
xdi= [B]i;n+1;1ir x fi= 0;1in r
These values for the variables make the equations represented by the rst rrows ofBall true (convince
yourself of this). Rows numbered greater than r(if any) are all zero rows, hence represent the equation
0 = 0 and are also all true. We have now identied one solution to the system represented by B, and hence
a solution to the system represented by A(Theorem REMES [31]). So we can say the system is consistent
(Denition CS [55]).
The beauty of this theorem being an equivalence is that we can unequivocally test to see if a system
is consistent or inconsistent by looking at just a single entry of the reduced row-echelon form matrix. We
could program a computer to do it!
Notice that for a consistent system the row-reduced augmented matrix has n+ 12F, so the largest
element ofFdoes not refer to a variable. Also, for an inconsistent system, n+ 12D, and it then does not
make much sense to discuss whether or not variables are free or dependent since there is no solution. Take
a look back at Denition IDV [57] and see why we did not need to consider the possibility of referencing
xn+1as a dependent variable.
With the characterization of Theorem RCLS [58], we can explore the relationships between randn
in light of the consistency of a system of equations. First, a situation where we can quickly conclude the
inconsistency of a system.
Theorem ISRN
Inconsistent Systems, randn
SupposeAis the augmented matrix of a system of linear equations in nvariables. Suppose also that Bis a
row-equivalent matrix in reduced row-echelon form with rrows that are not completely zeros. If r=n+1,
then the system of equations is inconsistent.
Proof Ifr=n+ 1, thenD=f1;2;3; :::; n; n + 1gand every column of Bcontains a leading 1 and is
a pivot column. In particular, the entry of column n+ 1 for row r=n+ 1 is a leading 1. Theorem RCLS
[58] then says that the system is inconsistent.
Do not confuse Theorem ISRN [59] with its converse! Go check out Technique CV [769] right now.
Next, if a system is consistent, we can distinguish between a unique solution and innitely many
solutions, and furthermore, we recognize that these are the only two possibilities.
Theorem CSRN
Consistent Systems, randn
SupposeAis the augmented matrix of a consistent system of linear equations with nvariables. Suppose
also thatBis a row-equivalent matrix in reduced row-echelon form with rrows that are not zero rows.
Thenrn. Ifr=n, then the system has a unique solution, and if r<n , then the system has innitely
many solutions.
Proof This theorem contains three implications that we must establish. Notice rst that Bhasn+ 1
columns, so there can be at most n+ 1 pivot columns, i.e. rn+ 1. Ifr=n+ 1, then Theorem ISRN
[59] tells us that the system is inconsistent, contrary to our hypothesis. We are left with rn.
Version 2.30
60 Section TSS Types of Solution Sets
Whenr=n, we ndn r= 0 free variables (i.e. F=fn+ 1g) and any solution must equal the unique
solution given by the rst nentries of column n+ 1 ofB.
Whenr < n , we haven r >0 free variables, corresponding to columns of Bwithout a leading 1,
excepting the nal column, which also does not contain a leading 1 by Theorem RCLS [58]. By varying
the values of the free variables suitably, we can demonstrate innitely many solutions.
Subsection FV
Free Variables
The next theorem simply states a conclusion from the nal paragraph of the previous proof, allowing us
to state explicitly the number of free variables for a consistent system.
Theorem FVCS
Free Variables for Consistent Systems
SupposeAis the augmented matrix of a consistent system of linear equations with nvariables. Suppose
also thatBis a row-equivalent matrix in reduced row-echelon form with rrows that are not completely
zeros. Then the solution set can be described with n rfree variables.
Proof See the proof of Theorem CSRN [59].
Example CFV
Counting free variables
For each archetype that is a system of equations, the values of nandrare listed. Many also contain a few
sample solutions. We can use this information protably, as illustrated by four examples.
1. Archetype A [781] has n= 3 andr= 2. It can be seen to be consistent by the sample solutions given.
Its solution set then has n r= 1 free variables, and therefore will be innite.
2. Archetype B [786] has n= 3 andr= 3. It can be seen to be consistent by the single sample solution
given. Its solution set can then be described with n r= 0 free variables, and therefore will have
just the single solution.
3. Archetype H [812] has n= 2 andr= 3. In this case, r=n+ 1, so Theorem ISRN [59] says the
system is inconsistent. We should not try to apply Theorem FVCS [60] to count free variables, since
the theorem only applies to consistent systems. (What would happen if you did?)
4. Archetype E [799] has n= 4 andr= 3. However, by looking at the reduced row-echelon form of the
augmented matrix, we nd a leading 1 in row 3, column 5. By Theorem RCLS [58] we recognize the
system as inconsistent. (Why doesn't this example contradict Theorem ISRN [59]?)
We have accomplished a lot so far, but our main goal has been the following theorem, which is now
very simple to prove. The proof is so simple that we ought to call it a corollary, but the result is important
enough that it deserves to be called a theorem. (See Technique LC [774].) Notice that this theorem was
presaged rst by Example TTS [13] and further foreshadowed by other examples.
Theorem PSSLS
Possible Solution Sets for Linear Systems
A system of linear equations has no solutions, a unique solution or innitely many solutions.
Proof By its denition, a system is either inconsistent or consistent (Denition CS [55]). The rst case
describes systems with no solutions. For consistent systems, we have the remaining two possibilities as
Version 2.30
Subsection TSS.FV Free Variables 61
guaranteed by, and described in, Theorem CSRN [59].
Here is a diagram that consolidates several of our theorems from this section, and which is of practical
use when you analyze systems of equations.
Theorem RCLS
Consisten t Inconsisten tnoleading 1in
column n+1aleading 1in
column n+1
Theorem FVCS
Infinite solutions Unique solutionr<nr =n
Diagram DTSLS. Decision Tree for Solving Linear Systems
We have one more theorem to round out our set of tools for determining solution sets to systems of linear
equations.
Theorem CMVEI
Consistent, More Variables than Equations, Innite solutions
Suppose a consistent system of linear equations has mequations in nvariables. If n>m , then the system
has innitely many solutions.
Proof Suppose that the augmented matrix of the system of equations is row-equivalent to B, a matrix
in reduced row-echelon form with rnonzero rows. Because Bhasmrows in total, the number that are
nonzero rows is less. In other words, rm. Follow this with the hypothesis that n>m and we nd that
the system has a solution set described by at least one free variable because
n rn m> 0:
A consistent system with free variables will have an innite number of solutions, as given by Theorem
CSRN [59].
Notice that to use this theorem we need only know that the system is consistent, together with the
values ofmandn. We do not necessarily have to compute a row-equivalent reduced row-echelon form
matrix, even though we discussed such a matrix in the proof. This is the substance of the following
example.
Example OSGMD
One solution gives many, Archetype D
Archetype D is the system of m= 3 equations in n= 4 variables,
2x1+x2+ 7x3 7x4= 8
3x1+ 4x2 5x3 6x4= 12
x1+x2+ 4x3 5x4= 4
and the solution x1= 0,x2= 1,x3= 2,x4= 1 can be checked easily by substitution. Having been handed
this solution, we know the system is consistent. This, together with n > m , allows us to apply Theorem
CMVEI [61] and conclude that the system has innitely many solutions.
These theorems give us the procedures and implications that allow us to completely solve any system
of linear equations. The main computational tool is using row operations to convert an augmented matrix
Version 2.30
62 Section TSS Types of Solution Sets
into reduced row-echelon form. Here's a broad outline of how we would instruct a computer to solve a
system of linear equations.
1. Represent a system of linear equations by an augmented matrix (an array is the appropriate data
structure in most computer languages).
2. Convert the matrix to a row-equivalent matrix in reduced row-echelon form using the procedure from
the proof of Theorem REMEF [34].
3. Determine rand locate the leading 1 of row r. If it is in column n+ 1, output the statement that the
system is inconsistent and halt.
4. With the leading 1 of row rnot in column n+ 1, there are two possibilities:
(a)r=nand the solution is unique. It can be read o directly from the entries in rows 1 through
nof columnn+ 1.
(b)r<n and there are innitely many solutions. If only a single solution is needed, set all the free
variables to zero and read o the dependent variable values from column n+ 1, as in the second
half of the proof of Theorem RCLS [58]. If the entire solution set is required, gure out some nice
compact way to describe it, since your nite computer is not big enough to hold all the solutions
(we'll have such a way soon).
The above makes it all sound a bit simpler than it really is. In practice, row operations employ division
(usually to get a leading entry of a row to convert to a leading 1) and that will introduce round-o errors.
Entries that should be zero sometimes end up being very, very small nonzero entries, or small entries lead
to over
ow errors when used as divisors. A variety of strategies can be employed to minimize these sorts
of errors, and this is one of the main topics in the important subject known as numerical linear algebra.
Solving a linear system is such a fundamental problem in so many areas of mathematics, and its
applications, that any computational device worth using for linear algebra will have a built-in routine to
do just that. See: Computation LS.MMA [746] Computation LS.SAGE [754] In this section we've
gained a foolproof procedure for solving any system of linear equations, no matter how many equations
or variables. We also have a handful of theorems that allow us to determine partial information about a
solution set without actually constructing the whole set itself. Donald Knuth would be proud.
Subsection READ
Reading Questions
1. How do we recognize when a system of linear equations is inconsistent?
2. Suppose we have converted the augmented matrix of a system of equations into reduced row-echelon
form. How do we then identify the dependent and independent (free) variables?
3. What are the possible solution sets for a system of linear equations?
Version 2.30
Subsection TSS.EXC Exercises 63
Subsection EXC
Exercises
C10 In the spirit of Example ISSI [56], describe the innite solution set for Archetype J [820].
Contributed by Robert Beezer
For Exercises C21{C28, nd the solution set of the given system of linear equations. Identify the values
ofnandr, and compare your answers to the results of the theorems of this section.
C21
x1+ 4x2+ 3x3 x4= 5
x1 x2+x3+ 2x4= 6
4x1+x2+ 6x3+ 5x4= 9
Contributed by Chris Black Solution [67]
C22
x1 2x2+x3 x4= 3
2x1 4x2+x3+x4= 2
x1 2x2 2x3+ 3x4= 1
Contributed by Chris Black Solution [67]
C23
x1 2x2+x3 x4= 3
x1+x2+x3 x4= 1
x1+x3 x4= 2
Contributed by Chris Black Solution [67]
C24
x1 2x2+x3 x4= 2
x1+x2+x3 x4= 2
x1+x3 x4= 2
Contributed by Chris Black Solution [67]
C25
x1+ 2x2+ 3x3= 1
2x1 x2+x3= 2
3x1+x2+x3= 4
x2+ 2x3= 6
Version 2.30
64 Section TSS Types of Solution Sets
Contributed by Chris Black Solution [68]
C26
x1+ 2x2+ 3x3= 1
2x1 x2+x3= 2
3x1+x2+x3= 4
5x2+ 2x3= 1
Contributed by Chris Black Solution [68]
C27
x1+ 2x2+ 3x3= 0
2x1 x2+x3= 2
x1 8x2 7x3= 1
x2+x3= 0
Contributed by Chris Black Solution [68]
C28
x1+ 2x2+ 3x3= 1
2x1 x2+x3= 2
x1 8x2 7x3= 1
x2+x3= 0
Contributed by Chris Black Solution [68]
M45 Prove that Archetype J [820] has innitely many solutions without row-reducing the augmented
matrix.
Contributed by Robert Beezer Solution [69]
M46 Consider Archetype J [820], and specically the row-reduced version of the augmented matrix of the
system of equations, denoted as Bhere, and the values of r,DandFimmediately following. Determine
the values of the entries
[B]1;d1[B]3;d3[B]1;d3[B]3;d1[B]d1;1 [B]d3;3 [B]d1;3 [B]d3;1 [B]1;f1[B]3;f1
(See Exercise TSS.M70 [65] for a generalization.)
Contributed by Manley Perkel
For Exercises M51{M57 say as much as possible about each system's solution set. Be sure to make
it clear which theorems you are using to reach your conclusions.
M51 A consistent system of 8 equations in 6 variables.
Contributed by Robert Beezer Solution [69]
M52 A consistent system of 6 equations in 8 variables.
Contributed by Robert Beezer Solution [69]
Version 2.30
Subsection TSS.EXC Exercises 65
M53 A system of 5 equations in 9 variables.
Contributed by Robert Beezer Solution [69]
M54 A system with 12 equations in 35 variables.
Contributed by Robert Beezer Solution [69]
M56 A system with 6 equations in 12 variables.
Contributed by Robert Beezer Solution [69]
M57 A system with 8 equations and 6 variables. The reduced row-echelon form of the augmented matrix
of the system has 7 pivot columns.
Contributed by Robert Beezer Solution [69]
M60 Without doing any computations, and without examining any solutions, say as much as possible
about the form of the solution set for each archetype that is a system of equations.
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]
Archetype E [799]
Archetype F [803]
Archetype G [808]
Archetype H [812]
Archetype I [816]
Archetype J [820]
Contributed by Robert Beezer
M70 Suppose that Bis a matrix in reduced row-echelon form that is equivalent to the augmented matrix
of a system of equations with mequations in nvariables. Let r,DandFbe as dened in Notation RREFA
[33]. What can you conclude, in general, about the following entries?
[B]1;d1[B]3;d3[B]1;d3[B]3;d1[B]d1;1 [B]d3;3 [B]d1;3 [B]d3;1 [B]1;f1[B]3;f1
If you cannot conclude anything about an entry, then say so. (See Exercise TSS.M46 [64] for inspiration.)
Contributed by Manley Perkel
T10 An inconsistent system may have r > n . If we try (incorrectly!) to apply Theorem FVCS [60] to
such a system, how many free variables would we discover?
Contributed by Robert Beezer Solution [69]
T20 Suppose that Bis a matrix in reduced row-echelon form that is equivalent to the augmented matrix
of a system of equations with mequations in nvariables. Let r,DandFbe as dened in Notation RREFA
[33]. Prove that dkkfor all 1kr. Then suppose that r2 and 1k<`rand determine what
can you conclude, in general, about the following entries.
[B]k;dk[B]k;d`[B]`;dk[B]dk;k [B]dk;` [B]d`;k [B]dk;f`[B]d`;fk
If you cannot conclude anything about an entry, then say so. (See Exercise TSS.M46 [64] and Exercise
TSS.M70 [65].)
Contributed by Manley Perkel
T40 Suppose that the coecient matrix of a consistent system of linear equations has two columns that
are identical. Prove that the system has innitely many solutions.
Contributed by Robert Beezer Solution [69]
Version 2.30
66 Section TSS Types of Solution Sets
T41 Consider the system of linear equations LS(A;b), and suppose that every element of the vector of
constants bis a common multiple of the corresponding element of a certain column of A. More precisely,
there is a complex number , and a column index j, such that [ b]i=[A]ijfor alli. Prove that the system
is consistent.
Contributed by Robert Beezer Solution [69]
Version 2.30
Subsection TSS.SOL Solutions 67
Subsection SOL
Solutions
C21 Contributed by Chris Black Statement [63]
The augmented matrix for the given linear system and its row-reduced form are:
2
41 4 3 1 5
1 1 1 2 6
4 1 6 5 93
5RREF !2
410 7=5 7=5 0
012=5 3=5 0
0 0 0 0 13
5:
For this system, we have n= 4 andr= 3. However, with a leading 1 in the last column we see that the
original system has no solution by Theorem RCLS [58].
C22 Contributed by Chris Black Statement [63]
The augmented matrix for the given linear system and its row-reduced form are:
2
41 2 1 1 3
2 4 1 1 2
1 2 2 3 13
5RREF !2
41 2 0 0 3
0 0 10 2
0 0 0 1 23
5:
Thus, we see we have an equivalent system for any scalar x2:
x1= 3 + 2x2
x3= 2
x4= 2:
For this system, n= 4 andr= 3. Since it is a consistent system by Theorem RCLS [58], Theorem CSRN
[59] guarantees an innite number of solutions.
C23 Contributed by Chris Black Statement [63]
The augmented matrix for the given linear system and its row-reduced form are:
2
41 2 1 1 3
1 1 1 1 1
1 0 1 1 23
5RREF !2
410 1 1 0
010 0 0
0 0 0 0 13
5:
For this system, we have n= 4 andr= 3. However, with a leading 1 in the last column we see that the
original system has no solution by Theorem RCLS [58].
C24 Contributed by Chris Black Statement [63]
The augmented matrix for the given linear system and its row-reduced form are:
2
41 2 1 1 2
1 1 1 1 2
1 0 1 1 23
5RREF !2
410 1 1 2
010 0 0
0 0 0 0 03
5:
Thus, we see that an equivalent system is
x1= 2 x3+x4
x2= 0;
and the solution set is8
>><
>>:2
6642 x3+x4
0
x3
x43
775x3;x42C9
>>=
>>;. For this system, n= 4 andr= 2. Since it is a
consistent system by Theorem RCLS [58], Theorem CSRN [59] guarantees an innite number of solutions.
Version 2.30
68 Section TSS Types of Solution Sets
C25 Contributed by Chris Black Statement [63]
The augmented matrix for the given linear system and its row-reduced form are:
2
6641 2 3 1
2 1 1 2
3 1 1 4
0 1 2 63
775RREF !2
666410 0 0
010 0
0 0 10
0 0 0 13
7775:
Sincen= 3 andr= 4 =n+ 1, Theorem ISRN [59] guarantees that the system is inconsistent. Thus, we
see that the given system has no solution.
C26 Contributed by Chris Black Statement [64]
The augmented matrix for the given linear system and its row-reduced form are:
2
6641 2 3 1
2 1 1 2
3 1 1 4
0 5 2 13
775RREF !2
66410 0 4=3
010 1=3
0 0 1 1=3
0 0 0 03
775:
Sincer=n= 3 and the system is consistent by Theorem RCLS [58], Theorem CSRN [59] guarantees a
unique solution, which is
x1= 4=3
x2= 1=3
x3= 1=3:
C27 Contributed by Chris Black Statement [64]
The augmented matrix for the given linear system and its row-reduced form are:
2
6641 2 3 0
2 1 1 2
1 8 7 1
0 1 1 03
775RREF !2
66410 1 0
011 0
0 0 0 1
0 0 0 03
775:
For this system, we have n= 3 andr= 3. However, with a leading 1 in the last column we see that the
original system has no solution by Theorem RCLS [58].
C28 Contributed by Chris Black Statement [64]
The augmented matrix for the given linear system and its row-reduced form are:
2
6641 2 3 1
2 1 1 2
1 8 7 1
0 1 1 03
775RREF !2
66410 1 1
011 0
0 0 0 0
0 0 0 03
775:
For this system, n= 3 andr= 2. Since it is a consistent system by Theorem RCLS [58], Theorem CSRN
[59] guarantees an innite number of solutions. An equivalent system is
x1= 1 x3
x2= x3;
wherex3is any scalar. So we can express the solution set as
8
<
:2
41 x3
x3
x33
5x32C9
=
;
Version 2.30
Subsection TSS.SOL Solutions 69
M45 Contributed by Robert Beezer Statement [64]
Demonstrate that the system is consistent by verifying any one of the four sample solutions provided. Then
becausen= 9>6 =m, Theorem CMVEI [61] gives us the conclusion that the system has innitely many
solutions.
Notice that we only know the system will have at least 9 6 = 3 free variables, but very well could
have more. We do not know know that r= 6, only that r6.
M51 Contributed by Robert Beezer Statement [64]
Consistent means there is at least one solution (Denition CS [55]). It will have either a unique solution
or innitely many solutions (Theorem PSSLS [60]).
M52 Contributed by Robert Beezer Statement [64]
With 6 rows in the augmented matrix, the row-reduced version will have r6. Since the system is
consistent, apply Theorem CSRN [59] to see that n r2 implies innitely many solutions.
M53 Contributed by Robert Beezer Statement [65]
The system could be inconsistent. If it is consistent, then because it has more variables than equations
Theorem CMVEI [61] implies that there would be innitely many solutions. So, of all the possibilities in
Theorem PSSLS [60], only the case of a unique solution can be ruled out.
M54 Contributed by Robert Beezer Statement [65]
The system could be inconsistent. If it is consistent, then Theorem CMVEI [61] tells us the solution set
will be innite. So we can be certain that there is not a unique solution.
M56 Contributed by Robert Beezer Statement [65]
The system could be inconsistent. If it is consistent, and since 12 >6, then Theorem CMVEI [61] says
we will have innitely many solutions. So there are two possibilities. Theorem PSSLS [60] allows to state
equivalently that a unique solution is an impossibility.
M57 Contributed by Robert Beezer Statement [65]
7 pivot columns implies that there are r= 7 nonzero rows (so row 8 is all zeros in the reduced row-echelon
form). Then n+ 1 = 6 + 1 = 7 = rand Theorem ISRN [59] allows to conclude that the system is
inconsistent.
T10 Contributed by Robert Beezer Statement [65]
Theorem FVCS [60] will indicate a negative number of free variables, but we can say even more. If r>n ,
then the only possibility is that r=n+ 1, and then we compute n r=n (n+ 1) = 1 free variables.
T40 Contributed by Robert Beezer Statement [65]
Since the system is consistent, we know there is either a unique solution, or innitely many solutions
(Theorem PSSLS [60]). If we perform row operations (Denition RO [31]) on the augmented matrix of the
system, the two equal columns of the coecient matrix will suer the same fate, and remain equal in the
nal reduced row-echelon form. Suppose both of these columns are pivot columns (Denition RREF [33]).
Then there is single row containing the two leading 1's of the two pivot columns, a violation of reduced
row-echelon form (Denition RREF [33]). So at least one of these columns is not a pivot column, and the
column index indicates a free variable in the description of the solution set (Denition IDV [57]). With a
free variable, we arrive at an innite solution set (Theorem FVCS [60]).
T41 Contributed by Robert Beezer Statement [66]
The condition about the multiple of the column of constants will allow you to show that the following
values form a solution of the system LS(A;b),
x1= 0x2= 0::: xj 1= 0xj= xj+1= 0::: xn 1= 0xn= 0
With one solution of the system known, we can say the system is consistent (Denition CS [55]).
Version 2.30
70 Section TSS Types of Solution Sets
A more involved proof can be built using Theorem RCLS [58]. Begin by proving that each of the three
row operations (Denition RO [31]) will convert the augmented matrix of the system into another matrix
where column jistimes the entry of the same row in the last column. In other words, the \column
multiple property" is preserved under row operations. These proofs will get successively more involved as
you work through the three operations.
Now construct a proof by contradiction (Technique CD [770]), by supposing that the system is incon-
sistent. Then the last column of the reduced row-echelon form of the augmented matrix is a pivot column
(Theorem RCLS [58]). Then column jmust have a zero in the same row as the leading 1 of the nal
column. But the \column multiple property" implies that there is an in columnjin the same row as the
leading 1. So = 0. By hypothesis, then the vector of constants is the zero vector. However, if we began
with a nal column of zeros, row operations would never have created a leading 1 in the nal column. This
contradicts the nal column being a pivot column, and therefore the system cannot be inconsistent.
Version 2.30
Section HSE Homogeneous Systems of Equations 71
Section HSE
Homogeneous Systems of Equations
In this section we specialize to systems of linear equations where every equation has a zero as its constant
term. Along the way, we will begin to express more and more ideas in the language of matrices and begin
a move away from writing out whole systems of equations. The ideas initiated in this section will carry
through the remainder of the course.
Subsection SHS
Solutions of Homogeneous Systems
As usual, we begin with a denition.
Denition HS
Homogeneous System
A system of linear equations, LS(A;b) ishomogeneous if the vector of constants is the zero vector, in
other words, b=0. 4
Example AHSAC
Archetype C as a homogeneous system
For each archetype that is a system of equations, we have formulated a similar, yet dierent, homogeneous
system of equations by replacing each equation's constant term with a zero. To wit, for Archetype C [791],
we can convert the original system of equations into the homogeneous system,
2x1 3x2+x3 6x4= 0
4x1+x2+ 2x3+ 9x4= 0
3x1+x2+x3+ 8x4= 0
Can you quickly nd a solution to this system without row-reducing the augmented matrix?
As you might have discovered by studying Example AHSAC [71], setting each variable to zero will
always be a solution of a homogeneous system. This is the substance of the following theorem.
Theorem HSC
Homogeneous Systems are Consistent
Suppose that a system of linear equations is homogeneous. Then the system is consistent.
Proof Set each variable of the system to zero. When substituting these values into each equation, the
left-hand side evaluates to zero, no matter what the coecients are. Since a homogeneous system has zero
on the right-hand side of each equation as the constant term, each equation is true. With one demonstrated
solution, we can call the system consistent.
Since this solution is so obvious, we now dene it as the trivial solution.
Denition TSHSE
Trivial Solution to Homogeneous Systems of Equations
Suppose a homogeneous system of linear equations has nvariables. The solution x1= 0,x2= 0,. . . ,xn= 0
(i.e.x=0) is called the trivial solution . 4
Here are three typical examples, which we will reference throughout this section. Work through the
row operations as we bring each to reduced row-echelon form. Also notice what is similar in each example,
and what diers.
Version 2.30
72 Section HSE Homogeneous Systems of Equations
Example HUSAB
Homogeneous, unique solution, Archetype B
Archetype B can be converted to the homogeneous system,
11x1+ 2x2 14x3= 0
23x1 6x2+ 33x3= 0
14x1 2x2+ 17x3= 0
whose augmented matrix row-reduces to
2
410 0 0
010 0
0 0 103
5
By Theorem HSC [71], the system is consistent, and so the computation n r= 3 3 = 0 means the
solution set contains just a single solution. Then, this lone solution must be the trivial solution.
Example HISAA
Homogeneous, innite solutions, Archetype A
Archetype A [781] can be converted to the homogeneous system,
x1 x2+ 2x3= 0
2x1+x2+x3= 0
x1+x2 = 0
whose augmented matrix row-reduces to
2
410 1 0
01 1 0
0 0 0 03
5
By Theorem HSC [71], the system is consistent, and so the computation n r= 3 2 = 1 means the
solution set contains one free variable by Theorem FVCS [60], and hence has innitely many solutions. We
can describe this solution set using the free variable x3,
S=8
<
:2
4x1
x2
x33
5x1= x3; x2=x39
=
;=8
<
:2
4 x3
x3
x33
5x32C9
=
;
Geometrically, these are points in three dimensions that lie on a line through the origin.
Example HISAD
Homogeneous, innite solutions, Archetype D
Archetype D [795] (and identically, Archetype E [799]) can be converted to the homogeneous system,
2x1+x2+ 7x3 7x4= 0
3x1+ 4x2 5x3 6x4= 0
x1+x2+ 4x3 5x4= 0
whose augmented matrix row-reduces to
2
410 3 2 0
011 3 0
0 0 0 0 03
5
Version 2.30
Subsection HSE.NSM Null Space of a Matrix 73
By Theorem HSC [71], the system is consistent, and so the computation n r= 4 2 = 2 means the
solution set contains two free variables by Theorem FVCS [60], and hence has innitely many solutions.
We can describe this solution set using the free variables x3andx4,
S=8
>><
>>:2
664x1
x2
x3
x43
775x1= 3x3+ 2x4; x2= x3+ 3x49
>>=
>>;
=8
>><
>>:2
664 3x3+ 2x4
x3+ 3x4
x3
x43
775x3; x42C9
>>=
>>;
After working through these examples, you might perform the same computations for the slightly larger
example, Archetype J [820].
Notice that when we do row operations on the augmented matrix of a homogeneous system of linear
equations the last column of the matrix is all zeros. Any one of the three allowable row operations will
convert zeros to zeros and thus, the nal column of the matrix in reduced row-echelon form will also be
all zeros. So in this case, we may be as likely to reference only the coecient matrix and presume that we
remember that the nal column begins with zeros, and after any number of row operations is still zero.
Example HISAD [72] suggests the following theorem.
Theorem HMVEI
Homogeneous, More Variables than Equations, Innite solutions
Suppose that a homogeneous system of linear equations has mequations and nvariables with n > m .
Then the system has innitely many solutions.
Proof We are assuming the system is homogeneous, so Theorem HSC [71] says it is consistent. Then the
hypothesis that n>m , together with Theorem CMVEI [61], gives innitely many solutions.
Example HUSAB [72] and Example HISAA [72] are concerned with homogeneous systems where n=m
and expose a fundamental distinction between the two examples. One has a unique solution, while the
other has innitely many. These are exactly the only two possibilities for a homogeneous system and
illustrate that each is possible (unlike the case when n>m where Theorem HMVEI [73] tells us that there
is only one possibility for a homogeneous system).
Subsection NSM
Null Space of a Matrix
The set of solutions to a homogeneous system (which by Theorem HSC [71] is never empty) is of enough
interest to warrant its own name. However, we dene it as a property of the coecient matrix, not as a
property of some system of equations.
Denition NSM
Null Space of a Matrix
Thenull space of a matrix A, denotedN(A), is the set of all the vectors that are solutions to the
homogeneous system LS(A;0).
Version 2.30
74 Section HSE Homogeneous Systems of Equations
(This denition contains Notation NSM.) 4
In the Archetypes (Appendix A [777]) each example that is a system of equations also has a corre-
sponding homogeneous system of equations listed, and several sample solutions are given. These solutions
will be elements of the null space of the coecient matrix. We'll look at one example.
Example NSEAI
Null space elements of Archetype I
The write-up for Archetype I [816] lists several solutions of the corresponding homogeneous system. Here
are two, written as solution vectors. We can say that they are in the null space of the coecient matrix
for the system of equations in Archetype I [816].
x=2
6666666643
0
5
6
0
0
13
777777775y=2
666666664 4
1
3
2
1
1
13
777777775
However, the vector
z=2
6666666641
0
0
0
0
0
23
777777775
is not in the null space, since it is not a solution to the homogeneous system. For example, it fails to even
make the rst equation true.
Here are two (prototypical) examples of the computation of the null space of a matrix.
Example CNS1
Computing a null space, #1
Let's compute the null space of
A=2
42 1 7 3 8
1 0 2 4 9
2 2 2 1 83
5
which we write as N(A). Translating Denition NSM [73], we simply desire to solve the homogeneous
systemLS(A;0). So we row-reduce the augmented matrix to obtain
2
410 2 0 1 0
01 3 0 4 0
0 0 0 12 03
5
The variables (of the homogeneous system) x3andx5are free (since columns 1, 2 and 4 are pivot columns),
so we arrange the equations represented by the matrix in reduced row-echelon form to
x1= 2x3 x5
x2= 3x3 4x5
x4= 2x5
Version 2.30
Subsection HSE.READ Reading Questions 75
So we can write the innite solution set as sets using column vectors,
N(A) =8
>>>><
>>>>:2
66664 2x3 x5
3x3 4x5
x3
2x5
x53
77775x3; x52C9
>>>>=
>>>>;
Example CNS2
Computing a null space, #2
Let's compute the null space of
C=2
664 4 6 1
1 4 1
5 6 7
4 7 13
775
which we write as N(C). Translating Denition NSM [73], we simply desire to solve the homogeneous
systemLS(C;0). So we row-reduce the augmented matrix to obtain
2
66410 0 0
010 0
0 0 10
0 0 0 03
775
There are no free variables in the homogeneous system represented by the row-reduced matrix, so there is
only the trivial solution, the zero vector, 0. So we can write the (trivial) solution set as
N(C) =f0g=8
<
:2
40
0
03
59
=
;
Subsection READ
Reading Questions
1. What is always true of the solution set for a homogeneous system of equations?
2. Suppose a homogeneous system of equations has 13 variables and 8 equations. How many solutions
will it have? Why?
3. Describe in words (not symbols) the null space of a matrix.
Version 2.30
76 Section HSE Homogeneous Systems of Equations
Subsection EXC
Exercises
C10 Each Archetype (Appendix A [777]) that is a system of equations has a corresponding homogeneous
system with the same coecient matrix. Compute the set of solutions for each. Notice that these solution
sets are the null spaces of the coecient matrices.
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]/Archetype E [799]
Archetype F [803]
Archetype G [808]/ Archetype H [812]
Archetype I [816]
and Archetype J [820]
Contributed by Robert Beezer
C20 Archetype K [825] and Archetype L [829] are simply 5 5 matrices (i.e. they are not systems of
equations). Compute the null space of each matrix.
Contributed by Robert Beezer
For Exercises C21-C23, solve the given homogeneous linear system. Compare your results to the results
of the corresponding exercise in Section TSS [55].
C21
x1+ 4x2+ 3x3 x4= 0
x1 x2+x3+ 2x4= 0
4x1+x2+ 6x3+ 5x4= 0
Contributed by Chris Black Solution [79]
C22
x1 2x2+x3 x4= 0
2x1 4x2+x3+x4= 0
x1 2x2 2x3+ 3x4= 0
Contributed by Chris Black Solution [79]
C23
x1 2x2+x3 x4= 0
x1+x2+x3 x4= 0
x1+x3 x4= 0
Contributed by Chris Black Solution [79]
For Exercises C25-C27, solve the given homogeneous linear system. Compare your results to the results
of the corresponding exercise in Section TSS [55].
C25
x1+ 2x2+ 3x3= 0
Version 2.30
Subsection HSE.EXC Exercises 77
2x1 x2+x3= 0
3x1+x2+x3= 0
x2+ 2x3= 0
Contributed by Chris Black Solution [80]
C26
x1+ 2x2+ 3x3= 0
2x1 x2+x3= 0
3x1+x2+x3= 0
5x2+ 2x3= 0
Contributed by Chris Black Solution [80]
C27
x1+ 2x2+ 3x3= 0
2x1 x2+x3= 0
x1 8x2 7x3= 0
x2+x3= 0
Contributed by Chris Black Solution [80]
C30 Compute the null space of the matrix A,N(A).
A=2
6642 4 1 3 8
1 2 1 1 1
2 4 0 3 4
2 4 1 7 43
775
Contributed by Robert Beezer Solution [80]
C31 Find the null space of the matrix B,N(B).
B=2
4 6 4 36 6
2 1 10 1
3 2 18 33
5
Contributed by Robert Beezer Solution [81]
M45 Without doing any computations, and without examining any solutions, say as much as possible
about the form of the solution set for corresponding homogeneous system of equations of each archetype
that is a system of equations.
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]/Archetype E [799]
Archetype F [803]
Version 2.30
78 Section HSE Homogeneous Systems of Equations
Archetype G [808]/Archetype H [812]
Archetype I [816]
Archetype J [820]
Contributed by Robert Beezer
For Exercises M50{M52 say as much as possible about each system's solution set. Be sure to make
it clear which theorems you are using to reach your conclusions.
M50 A homogeneous system of 8 equations in 8 variables.
Contributed by Robert Beezer Solution [81]
M51 A homogeneous system of 8 equations in 9 variables.
Contributed by Robert Beezer Solution [81]
M52 A homogeneous system of 8 equations in 7 variables.
Contributed by Robert Beezer Solution [81]
T10 Prove or disprove: A system of linear equations is homogeneous if and only if the system has the
zero vector as a solution.
Contributed by Martin Jackson Solution [81]
T12 Give an alternate proof of Theorem HSC [71] that uses Theorem RCLS [58].
Contributed by Ivan Kessler
T20 Consider the homogeneous system of linear equations LS(A;0), and suppose that u=2
666664u1
u2
u3
...
un3
777775is one
solution to the system of equations. Prove that v=2
6666644u1
4u2
4u3
...
4un3
777775is also a solution to LS(A;0).
Contributed by Robert Beezer Solution [82]
Version 2.30
Subsection HSE.SOL Solutions 79
Subsection SOL
Solutions
C21 Contributed by Chris Black Statement [76]
The augmented matrix for the given linear system and its row-reduced form are:
2
41 4 3 1 0
1 1 1 2 0
4 1 6 5 03
5RREF !2
410 7=5 7=5 0
012=5 3=5 0
0 0 0 0 03
5:
Thus, we see that the system is consistent (as predicted by Theorem HSC [71]) and has an innite number
of solutions (as predicted by Theorem HMVEI [73]). With suitable choices of x3andx4, each solution can
be written as
2
664 7
5x3 7
5x4
2
5x3+3
5x4
x3
x43
775
C22 Contributed by Chris Black Statement [76]
The augmented matrix for the given linear system and its row-reduced form are:
2
41 2 1 1 0
2 4 1 1 0
1 2 2 3 03
5RREF !2
41 2 0 0 0
0 0 10 0
0 0 0 103
5:
Thus, we see that the system is consistent (as predicted by Theorem HSC [71]) and has an innite number
of solutions (as predicted by Theorem HMVEI [73]). With a suitable choice of x2, each solution can be
written as
2
6642x2
x2
0
03
775
C23 Contributed by Chris Black Statement [76]
The augmented matrix for the given linear system and its row-reduced form are:
2
41 2 1 1 0
1 1 1 1 0
1 0 1 1 03
5RREF !2
410 1 1 0
010 0 0
0 0 0 0 03
5:
Thus, we see that the system is consistent (as predicted by Theorem HSC [71]) and has an innite number
of solutions (as predicted by Theorem HMVEI [73]). With suitable choices of x3andx4, each solution can
be written as
2
664 x3+x4
0
x3
x43
775
Version 2.30
80 Section HSE Homogeneous Systems of Equations
C25 Contributed by Chris Black Statement [76]
The augmented matrix for the given linear system and its row-reduced form are:
2
6641 2 3 0
2 1 1 0
3 1 1 0
0 1 2 03
775RREF !2
66410 0 0
010 0
0 0 10
0 0 0 03
775:
An homogeneous system is always consistent (Theorem HSC [71]) and with n=r= 3 an application of
Theorem FVCS [60] yields zero free variables. Thus the only solution to the given system is the trivial
solution, x=0.
C26 Contributed by Chris Black Statement [77]
The augmented matrix for the given linear system and its row-reduced form are:
2
6641 2 3 0
2 1 1 0
3 1 1 0
0 5 2 03
775RREF !2
6641 0 0 0
0 1 0 0
0 0 1 0
0 0 0 03
775:
An homogeneous system is always consistent (Theorem HSC [71]) and with n=r= 3 an application of
Theorem FVCS [60] yields zero free variables. Thus the only solution to the given system is the trivial
solution, x=0.
C27 Contributed by Chris Black Statement [77]
The augmented matrix for the given linear system and its row-reduced form are:
2
6641 2 3 0
2 1 1 0
1 8 7 0
0 1 1 03
775RREF !2
66410 1 0
011 0
0 0 0 0
0 0 0 03
775:
An homogeneous system is always consistent (Theorem HSC [71]) and with n= 3,r= 2 an application of
Theorem FVCS [60] yields one free variable. With a suitable choice of x3each solution can be written in
the form
2
4 x3
x3
x33
5
C30 Contributed by Robert Beezer Statement [77]
Denition NSM [73] tells us that the null space of Ais the solution set to the homogeneous system LS(A;0).
The augmented matrix of this system is
2
6642 4 1 3 8 0
1 2 1 1 1 0
2 4 0 3 4 0
2 4 1 7 4 03
775
To solve the system, we row-reduce the augmented matrix and obtain,
2
66412 0 0 5 0
0 0 10 8 0
0 0 0 1 2 0
0 0 0 0 0 03
775
Version 2.30
Subsection HSE.SOL Solutions 81
This matrix represents a system with equations having three dependent variables ( x1,x3, andx4) and two
independent variables ( x2andx5). These equations rearrange to
x1= 2x2 5x5 x3= 8x5 x4= 2x5
So we can write the solution set (which is the requested null space) as
N(A) =8
>>>><
>>>>:2
66664 2x2 5x5
x2
8x5
2x5
x53
77775x2;x52C9
>>>>=
>>>>;
C31 Contributed by Robert Beezer Statement [77]
We form the augmented matrix of the homogeneous system LS(B;0) and row-reduce the matrix,
2
4 6 4 36 6 0
2 1 10 1 0
3 2 18 3 03
5RREF !2
410 2 1 0
01 6 3 0
0 0 0 0 03
5
We knew ahead of time that this system would be consistent (Theorem HSC [71]), but we can now see
there aren r= 4 2 = 2 free variables, namely x3andx4(Theorem FVCS [60]). Based on this analysis,
we can rearrange the equations associated with each nonzero row of the reduced row-echelon form into an
expression for the lone dependent variable as a function of the free variables. We arrive at the solution set
to the homogeneous system, which is the null space of the matrix by Denition NSM [73],
N(B) =8
>><
>>:2
664 2x3 x4
6x3 3x4
x3
x43
775x3; x42C9
>>=
>>;
M50 Contributed by Robert Beezer Statement [78]
Since the system is homogeneous, we know it has the trivial solution (Theorem HSC [71]). We cannot say
anymore based on the information provided, except to say that there is either a unique solution or innitely
many solutions (Theorem PSSLS [60]). See Archetype A [781] and Archetype B [786] to understand the
possibilities.
M51 Contributed by Robert Beezer Statement [78]
Since there are more variables than equations, Theorem HMVEI [73] applies and tells us that the solution
set is innite. From the proof of Theorem HSC [71] we know that the zero vector is one solution.
M52 Contributed by Robert Beezer Statement [78]
By Theorem HSC [71], we know the system is consistent because the zero vector is always a solution of a
homogeneous system. There is no more that we can say, since both a unique solution and innitely many
solutions are possibilities.
T10 Contributed by Robert Beezer Statement [78]
This is a true statement. A proof is:
()) Suppose we have a homogeneous system LS(A;0). Then by substituting the scalar zero for each
variable, we arrive at true statements for each equation. So the zero vector is a solution. This is the
content of Theorem HSC [71].
(() Suppose now that we have a generic (i.e. not necessarily homogeneous) system of equations,
LS(A;b) that has the zero vector as a solution. Upon substituting this solution into the system, we
discover that each component of bmust also be zero. So b=0.
Version 2.30
82 Section HSE Homogeneous Systems of Equations
T20 Contributed by Robert Beezer Statement [78]
Suppose that a single equation from this system (the i-th one) has the form,
ai1x1+ai2x2+ai3x3++ainxn= 0
Evaluate the left-hand side of this equation with the components of the proposed solution vector v,
ai1(4u1) +ai2(4u2) +ai3(4u3) ++ain(4un)
= 4ai1u1+ 4ai2u2+ 4ai3u3++ 4ainun Commutativity
= 4 (ai1u1+ai2u2+ai3u3++ainun) Distributivity
= 4(0) usolution toLS(A;0)
= 0
Sovmakes each equation true, and so is a solution to the system.
Notice that this result is not true if we change LS(A;0) from a homogeneous system to a non-
homogeneous system. Can you create an example of a (non-homogeneous) system with a solution u
such that vis not a solution?
Version 2.30
Section NM Nonsingular Matrices 83
Section NM
Nonsingular Matrices
In this section we specialize and consider matrices with equal numbers of rows and columns, which when
considered as coecient matrices lead to systems with equal numbers of equations and variables. We will
see in the second half of the course (Chapter D [423], Chapter E [453] Chapter LT [515], Chapter R [603])
that these matrices are especially important.
Subsection NM
Nonsingular Matrices
Our theorems will now establish connections between systems of equations (homogeneous or otherwise),
augmented matrices representing those systems, coecient matrices, constant vectors, the reduced row-
echelon form of matrices (augmented and coecient) and solution sets. Be very careful in your reading,
writing and speaking about systems of equations, matrices and sets of vectors. A system of equations is
not a matrix, a matrix is not a solution set, and a solution set is not a system of equations. Now would be
a great time to review the discussion about speaking and writing mathematics in Technique L [766].
Denition SQM
Square Matrix
A matrix with mrows andncolumns is square ifm=n. In this case, we say the matrix has sizen. To
emphasize the situation when a matrix is not square, we will call it rectangular . 4
We can now present one of the central denitions of linear algebra.
Denition NM
Nonsingular Matrix
SupposeAis a square matrix. Suppose further that the solution set to the homogeneous linear system
of equationsLS(A;0) isf0g, i.e. the system has only the trivial solution. Then we say that Ais a
nonsingular matrix. Otherwise we say Ais asingular matrix. 4
We can investigate whether any square matrix is nonsingular or not, no matter if the matrix is derived
somehow from a system of equations or if it is simply a matrix. The denition says that to perform this
investigation we must construct a very specic system of equations (homogeneous, with the matrix as
the coecient matrix) and look at its solution set. We will have theorems in this section that connect
nonsingular matrices with systems of equations, creating more opportunities for confusion. Convince
yourself now of two observations, (1) we can decide nonsingularity for any square matrix, and (2) the
determination of nonsingularity involves the solution set for a certain homogeneous system of equations.
Notice that it makes no sense to call a system of equations nonsingular (the term does not apply to a
system of equations), nor does it make any sense to call a 5 7 matrix singular (the matrix is not square).
Example S
A singular matrix, Archetype A
Example HISAA [72] shows that the coecient matrix derived from Archetype A [781], specically the
33 matrix,
A=2
41 1 2
2 1 1
1 1 03
5
Version 2.30
84 Section NM Nonsingular Matrices
is a singular matrix since there are nontrivial solutions to the homogeneous system LS(A;0).
Example NM
A nonsingular matrix, Archetype B
Example HUSAB [72] shows that the coecient matrix derived from Archetype B [786], specically the
33 matrix,
B=2
4 7 6 12
5 5 7
1 0 43
5
is a nonsingular matrix since the homogeneous system, LS(B;0), has only the trivial solution.
Notice that we will not discuss Example HISAD [72] as being a singular or nonsingular coecient
matrix since the matrix is not square.
The next theorem combines with our main computational technique (row-reducing a matrix) to make
it easy to recognize a nonsingular matrix. But rst a denition.
Denition IM
Identity Matrix
Themmidentity matrix ,Im, is dened by
[Im]ij=(
1i=j
0i6=j1i; jm
(This denition contains Notation IM.) 4
Example IM
An identity matrix
The 44 identity matrix is
I4=2
6641 0 0 0
0 1 0 0
0 0 1 0
0 0 0 13
775:
Notice that an identity matrix is square, and in reduced row-echelon form. So in particular, if we were
to arrive at the identity matrix while bringing a matrix to reduced row-echelon form, then it would have
all of the diagonal entries circled as leading 1's.
Theorem NMRRI
Nonsingular Matrices Row Reduce to the Identity matrix
Suppose that Ais a square matrix and Bis a row-equivalent matrix in reduced row-echelon form. Then
Ais nonsingular if and only if Bis the identity matrix.
Proof (() SupposeBis the identity matrix. When the augmented matrix [ Aj0] is row-reduced, the
result is [Bj0] = [Inj0]. The number of nonzero rows is equal to the number of variables in the linear
system of equations LS(A;0), son=rand Theorem FVCS [60] gives n r= 0 free variables. Thus, the
homogeneous system LS(A;0) has just one solution, which must be the trivial solution. This is exactly
the denition of a nonsingular matrix.
()) IfAis nonsingular, then the homogeneous system LS(A;0) has a unique solution, and has no
free variables in the description of the solution set. The homogeneous system is consistent (Theorem HSC
[71]) so Theorem FVCS [60] applies and tells us there are n rfree variables. Thus, n r= 0, and so
Version 2.30
Subsection NM.NSNM Null Space of a Nonsingular Matrix 85
n=r. SoBhasnpivot columns among its total of ncolumns. This is enough to force Bto be thenn
identity matrix In(see Exercise NM.T12 [89]).
Notice that since this theorem is an equivalence it will always allow us to determine if a matrix is
either nonsingular or singular. Here are two examples of this, continuing our study of Archetype A and
Archetype B.
Example SRR
Singular matrix, row-reduced
The coecient matrix for Archetype A [781] is
A=2
41 1 2
2 1 1
1 1 03
5
which when row-reduced becomes the row-equivalent matrix
B=2
410 1
01 1
0 0 03
5:
Since this matrix is not the 3 3 identity matrix, Theorem NMRRI [84] tells us that Ais a singular matrix.
Example NSR
Nonsingular matrix, row-reduced
The coecient matrix for Archetype B [786] is
A=2
4 7 6 12
5 5 7
1 0 43
5
which when row-reduced becomes the row-equivalent matrix
B=2
410 0
010
0 0 13
5:
Since this matrix is the 3 3 identity matrix, Theorem NMRRI [84] tells us that Ais a nonsingular matrix.
Subsection NSNM
Null Space of a Nonsingular Matrix
Nonsingular matrices and their null spaces are intimately related, as the next two examples illustrate.
Example NSS
Null space of a singular matrix
Given the coecient matrix from Archetype A [781],
A=2
41 1 2
2 1 1
1 1 03
5
Version 2.30
86 Section NM Nonsingular Matrices
the null space is the set of solutions to the homogeneous system of equations LS(A;0) has a solution set
and null space constructed in Example HISAA [72] as
N(A) =8
<
:2
4 x3
x3
x33
5x32C9
=
;
Example NSNM
Null space of a nonsingular matrix
Given the coecient matrix from Archetype B [786],
A=2
4 7 6 12
5 5 7
1 0 43
5
the homogeneous system LS(A;0) has a solution set constructed in Example HUSAB [72] that contains
only the trivial solution, so the null space has only a single element,
N(A) =8
<
:2
40
0
03
59
=
;
These two examples illustrate the next theorem, which is another equivalence.
Theorem NMTNS
Nonsingular Matrices have Trivial Null Spaces
Suppose that Ais a square matrix. Then Ais nonsingular if and only if the null space of A,N(A), contains
only the zero vector, i.e. N(A) =f0g.
Proof The null space of a square matrix ,A, is equal to the set of solutions to the homogeneous system ,
LS(A;0). A matrix is nonsingular if and only if the set of solutions to the homogeneous system ,LS(A;0),
has only a trivial solution. These two observations may be chained together to construct the two proofs
necessary for each half of this theorem.
The next theorem pulls a lot of big ideas together. Theorem NMUS [86] tells us that we can learn
much about solutions to a system of linear equations with a square coecient matrix by just examining a
similar homogeneous system.
Theorem NMUS
Nonsingular Matrices and Unique Solutions
Suppose that Ais a square matrix. Ais a nonsingular matrix if and only if the system LS(A;b) has a
unique solution for every choice of the constant vector b.
Proof (() The hypothesis for this half of the proof is that the system LS(A;b) has a unique solution
forevery choice of the constant vector b. We will make a very specic choice for b:b=0. Then we know
that the systemLS(A;0) has a unique solution. But this is precisely the denition of what it means for
Ato be nonsingular (Denition NM [83]). That almost seems too easy! Notice that we have not used the
full power of our hypothesis, but there is nothing that says we must use a hypothesis to its fullest.
()) We assume that Ais nonsingular of size nn, so we know there is a sequence of row operations that
will convert Ainto the identity matrix In(Theorem NMRRI [84]). Form the augmented matrix A0= [Ajb]
and apply this same sequence of row operations to A0. The result will be the matrix B0= [Injc], which is
in reduced row-echelon form with r=n. Then the augmented matrix B0represents the (extremely simple)
Version 2.30
Subsection NM.READ Reading Questions 87
system of equations xi= [c]i, 1in. The vector cis clearly a solution, so the system is consistent
(Denition CS [55]). With a consistent system, we use Theorem FVCS [60] to count free variables. We
nd that there are n r=n n= 0 free variables, and so we therefore know that the solution is unique.
(This half of the proof was suggested by Asa Scherer.)
This theorem helps to explain part of our interest in nonsingular matrices. If a matrix is nonsingular,
then no matter what vector of constants we pair it with, using the matrix as the coecient matrix will
always yield a linear system of equations with a solution, and the solution is unique. To determine if a
matrix has this property (non-singularity) it is enough to just solve one linear system, the homogeneous
system with the matrix as coecient matrix and the zero vector as the vector of constants (or any other
vector of constants, see Exercise MM.T10 [237]).
Formulating the negation of the second part of this theorem is a good exercise. A singular matrix has
the property that for some value of the vector b, the systemLS(A;b) does not have a unique solution
(which means that it has no solution or innitely many solutions). We will be able to say more about this
case later (see the discussion following Theorem PSPHS [124]). Square matrices that are nonsingular have
a long list of interesting properties, which we will start to catalog in the following, recurring, theorem. Of
course, singular matrices will then have all of the opposite properties. The following theorem is a list of
equivalences. We want to understand just what is involved with understanding and proving a theorem
that says several conditions are equivalent. So have a look at Technique ME [771] before studying the rst
in this series of theorems.
Theorem NME1
Nonsingular Matrix Equivalences, Round 1
Suppose that Ais a square matrix. The following are equivalent.
1.Ais nonsingular.
2.Arow-reduces to the identity matrix.
3. The null space of Acontains only the zero vector, N(A) =f0g.
4. The linear system LS(A;b) has a unique solution for every possible choice of b.
Proof ThatAis nonsingular is equivalent to each of the subsequent statements by, in turn, Theorem
NMRRI [84], Theorem NMTNS [86] and Theorem NMUS [86]. So the statement of this theorem is just a
convenient way to organize all these results.
Finally, you may have wondered why we refer to a matrix as nonsingular when it creates systems
of equations with single solutions (Theorem NMUS [86])! I've wondered the same thing. We'll have an
opportunity to address this when we get to Theorem SMZD [445]. Can you wait that long?
Subsection READ
Reading Questions
1. What is the denition of a nonsingular matrix?
2. What is the easiest way to recognize a nonsingular matrix?
3. Suppose we have a system of equations and its coecient matrix is nonsingular. What can you say
about the solution set for this system?
Version 2.30
88 Section NM Nonsingular Matrices
Subsection EXC
Exercises
In Exercises C30{C33 determine if the matrix is nonsingular or singular. Give reasons for your answer.
C30 2
664 3 1 2 8
2 0 3 4
1 2 7 4
5 1 2 03
775
Contributed by Robert Beezer Solution [90]
C31 2
6642 3 1 4
1 1 1 0
1 2 3 5
1 2 1 33
775
Contributed by Robert Beezer Solution [90]
C32 2
49 3 2 4
5 6 1 3
4 1 3 53
5
Contributed by Robert Beezer Solution [90]
C33 2
664 1 2 0 3
1 3 2 4
2 0 4 3
3 1 2 33
775
Contributed by Robert Beezer Solution [90]
C40 Each of the archetypes below is a system of equations with a square coecient matrix, or is itself
a square matrix. Determine if these matrices are nonsingular, or singular. Comment on the null space of
each matrix.
Archetype A [781]
Archetype B [786]
Archetype F [803]
Archetype K [825]
Archetype L [829]
Contributed by Robert Beezer
C50 Find the null space of the matrix Ebelow.
E=2
6642 1 1 9
2 2 6 6
1 2 8 0
1 2 12 123
775
Contributed by Robert Beezer Solution [90]
Version 2.30
Subsection NM.EXC Exercises 89
M30 LetAbe the coecient matrix of the system of equations below. Is Anonsingular or singular?
Explain what you could infer about the solution set for the system based only on what you have learned
aboutAbeing singular or nonsingular.
x1+ 5x2= 8
2x1+ 5x2+ 5x3+ 2x4= 9
3x1 x2+ 3x3+x4= 3
7x1+ 6x2+ 5x3+x4= 30
Contributed by Robert Beezer Solution [91]
For Exercises M51{M52 say as much as possible about each system's solution set. Be sure to make
it clear which theorems you are using to reach your conclusions.
M51 6 equations in 6 variables, singular coecient matrix.
Contributed by Robert Beezer Solution [91]
M52 A system with a nonsingular coecient matrix, not homogeneous.
Contributed by Robert Beezer Solution [91]
T10 Suppose that Ais a singular matrix, and Bis a matrix in reduced row-echelon form that is row-
equivalent to A. Prove that the last row of Bis a zero row.
Contributed by Robert Beezer Solution [91]
T12 Suppose that Ais a square matrix. Using the denition of reduced row-echelon form (Denition
RREF [33]) carefully, give a proof of the following equivalence: Every column of Ais a pivot column if and
only ifAis the identity matrix (Denition IM [84]).
Contributed by Robert Beezer
T30 Suppose that Ais a nonsingular matrix and Ais row-equivalent to the matrix B. Prove that Bis
nonsingular.
Contributed by Robert Beezer Solution [91]
T90 Provide an alternative for the second half of the proof of Theorem NMUS [86], without appealing
to properties of the reduced row-echelon form of the coecient matrix. In other words, prove that if Ais
nonsingular, then LS(A;b) has a unique solution for every choice of the constant vector b. Construct this
proof without using Theorem REMEF [34] or Theorem RREFU [35].
Contributed by Robert Beezer Solution [91]
Version 2.30
90 Section NM Nonsingular Matrices
Subsection SOL
Solutions
C30 Contributed by Robert Beezer Statement [88]
The matrix row-reduces to 2
666410 0 0
010 0
0 0 10
0 0 0 13
7775
which is the 44 identity matrix. By Theorem NMRRI [84] the original matrix must be nonsingular.
C31 Contributed by Robert Beezer Statement [88]
Row-reducing the matrix yields,2
66410 0 2
010 3
0 0 1 1
0 0 0 03
775
Since this is not the 4 4 identity matrix, Theorem NMRRI [84] tells us the matrix is singular.
C32 Contributed by Robert Beezer Statement [88]
The matrix is not square, so neither term is applicable. See Denition NM [83], which is stated for just
square matrices.
C33 Contributed by Robert Beezer Statement [88]
Theorem NMRRI [84] tells us we can answer this question by simply row-reducing the matrix. Doing this
we obtain,2
666410 0 0
010 0
0 0 10
0 0 0 13
7775
Since the reduced row-echelon form of the matrix is the 4 4 identity matrix I4, we know that Bis
nonsingular.
C50 Contributed by Robert Beezer Statement [88]
We form the augmented matrix of the homogeneous system LS(E;0) and row-reduce the matrix,
2
6642 1 1 9 0
2 2 6 6 0
1 2 8 0 0
1 2 12 12 03
775RREF !2
66410 2 6 0
01 5 3 0
0 0 0 0 0
0 0 0 0 03
775
We knew ahead of time that this system would be consistent (Theorem HSC [71]), but we can now see
there aren r= 4 2 = 2 free variables, namely x3andx4sinceF=f3;4;5g(Theorem FVCS [60]).
Based on this analysis, we can rearrange the equations associated with each nonzero row of the reduced
row-echelon form into an expression for the lone dependent variable as a function of the free variables. We
arrive at the solution set to this homogeneous system, which is the null space of the matrix by Denition
NSM [73],
N(E) =8
>><
>>:2
664 2x3+ 6x4
5x3 3x4
x3
x43
775x3; x42C9
>>=
>>;
Version 2.30
Subsection NM.SOL Solutions 91
M30 Contributed by Robert Beezer Statement [89]
We row-reduce the coecient matrix of the system of equations,
2
664 1 5 0 0
2 5 5 2
3 1 3 1
7 6 5 13
775RREF !2
666410 0 0
010 0
0 0 10
0 0 0 13
7775
Since the row-reduced version of the coecient matrix is the 4 4 identity matrix, I4(Denition IM [84]
byTheorem NMRRI [84], we know the coecient matrix is nonsingular. According to Theorem NMUS
[86] we know that the system is guaranteed to have a unique solution, based only on the extra information
that the coecient matrix is nonsingular.
M51 Contributed by Robert Beezer Statement [89]
Theorem NMRRI [84] tells us that the coecient matrix will not row-reduce to the identity matrix. So
if we were to row-reduce the augmented matrix of this system of equations, we would not get a unique
solution. So by Theorem PSSLS [60] the remaining possibilities are no solutions, or innitely many.
M52 Contributed by Robert Beezer Statement [89]
Any system with a nonsingular coecient matrix will have a unique solution by Theorem NMUS [86]. If
the system is not homogeneous, the solution cannot be the zero vector (Exercise HSE.T10 [78]).
T10 Contributed by Robert Beezer Statement [89]
Letndenote the size of the square matrix A. By Theorem NMRRI [84] the hypothesis that Ais singular
implies that Bis not the identity matrix In. IfBhasnpivot columns, then it would have to be In, soB
must have fewer than npivot columns. But the number of nonzero rows in B(r) is equal to the number
of pivot columns as well. So the nrows ofBhave fewer than nnonzero rows, and Bmust contain at least
one zero row. By Denition RREF [33], this row must be at the bottom of B.
T30 Contributed by Robert Beezer Statement [89]
SinceAandBare row-equivalent matrices, consideration of the three row operations (Denition RO [31])
will show that the augmented matrices, [ Aj0] and [Bj0], are also row-equivalent matrices. This says
that the two homogeneous systems, LS(A;0) andLS(B;0) are equivalent systems. LS(A;0) has only
the zero vector as a solution (Denition NM [83]), thus LS(B;0) has only the zero vector as a solution.
Finally, by Denition NM [83], we see that Bis nonsingular.
Form a similar theorem replacing \nonsingular" by \singular" in both the hypothesis and the conclu-
sion. Prove this new theorem with an approach just like the one above, and/or employ the result about
nonsingular matrices in a proof by contradiction.
T90 Contributed by Robert Beezer Statement [89]
We assume Ais nonsingular, and try to solve the system LS(A;b) without making any assumptions about
b. To do this we will begin by constructing a new homogeneous linear system of equations that looks very
much like the original. Suppose Ahas sizen(why must it be square?) and write the original system as,
a11x1+a12x2+a13x3++a1nxn=b1
a21x1+a22x2+a23x3++a2nxn=b2
a31x1+a32x2+a33x3++a3nxn=b3
... ( )
an1x1+an2x2+an3x3++annxn=bn
Form the new, homogeneous system in nequations with n+1 variables, by adding a new variable y, whose
coecients are the negatives of the constant terms,
a11x1+a12x2+a13x3++a1nxn b1y= 0
Version 2.30
92 Section NM Nonsingular Matrices
a21x1+a22x2+a23x3++a2nxn b2y= 0
a31x1+a32x2+a33x3++a3nxn b3y= 0
... ( )
an1x1+an2x2+an3x3++annxn bny= 0
Since this is a homogeneous system with more variables than equations ( m=n+1>n), Theorem HMVEI
[73] says that the system has innitely many solutions. We will choose one of these solutions, anyone of
these solutions, so long as it is notthe trivial solution. Write this solution as
x1=c1x2=c2x3=c3::: x n=cny=cn+1
We know that at least one value of the ciis nonzero, but we will now show that in particular cn+16= 0.
We do this using a proof by contradiction (Technique CD [770]). So suppose the ciform a solution as
described, and in addition that cn+1= 0. Then we can write the i-th equation of system ( ) as,
ai1c1+ai2c2+ai3c3++aincn bi(0) = 0
which becomes
ai1c1+ai2c2+ai3c3++aincn= 0
Since this is true for each i, we have that x1=c1; x2=c2; x3=c3;:::; xn=cnis a solution to the
homogeneous system LS(A;0) formed with a nonsingular coecient matrix. This means that the only
possible solution is the trivial solution, so c1= 0; c2= 0; c3= 0; :::; cn= 0. So, assuming simply that
cn+1= 0, we conclude that allof theciare zero. But this contradicts our choice of the cias not being the
trivial solution to the system ( ). Socn+16= 0.
We now propose and verify a solution to the original system ( ). Set
x1=c1
cn+1x2=c2
cn+1x3=c3
cn+1::: x n=cn
cn+1
Notice how it was necessary that we know that cn+16= 0 for this step to succeed. Now, evaluate the i-th
equation of system ( ) with this proposed solution, and recognize in the third line that c1throughcn+1
appear as if they were substituted into the left-hand side of the i-th equation of system ( ),
ai1c1
cn+1+ai2c2
cn+1+ai3c3
cn+1++aincn
cn+1
=1
cn+1(ai1c1+ai2c2+ai3c3++aincn)
=1
cn+1(ai1c1+ai2c2+ai3c3++aincn bicn+1) +bi
=1
cn+1(0) +bi
=bi
Since this equation is true for every i, we have found a solution to system ( ). To nish, we still need to
establish that this solution is unique .
With one solution in hand, we will entertain the possibility of a second solution. So assume system ( )
has two solutions,
x1=d1 x2=d2 x3=d3 ::: x n=dn
Version 2.30
Subsection NM.SOL Solutions 93
x1=e1 x2=e2 x3=e3 ::: x n=en
Then,
(ai1(d1 e1) +ai2(d2 e2) +ai3(d3 e3) ++ain(dn en))
= (ai1d1+ai2d2+ai3d3++aindn) (ai1e1+ai2e2+ai3e3++ainen)
=bi bi
= 0
This is the i-th equation of the homogeneous system LS(A;0) evaluated with xj=dj ej, 1jn.
SinceAis nonsingular, we must conclude that this solution is the trivial solution, and so 0 = dj ej,
1jn. That is,dj=ejfor alljand the two solutions are identical, meaning any solution to ( ) is
unique.
Notice that the proposed solution ( xi=ci
cn+1) appeared in this proof with no motivation whatsoever.
This is just ne in a proof. A proof should convince you that a theorem is true. It is your job to read the
proof and be convinced of every assertion. Questions like \Where did that come from?" or \How would I
think of that?" have no bearing on the validity of the proof.
Version 2.30
94 Section NM Nonsingular Matrices
Version 2.30
Annotated Acronyms NM.SLE Systems of Linear Equations 95
Annotated Acronyms SLE
Systems of Linear Equations
At the conclusion of each chapter you will nd a section like this, reviewing selected denitions and
theorems. There are many reasons for why a denition or theorem might be placed here. It might
represent a key concept, it might be used frequently for computations, provide the critical step in many
proofs, or it may deserve special comment.
These lists are not meant to be exhaustive, but should still be useful as part of reviewing each chapter.
We will mention a few of these that you might eventually recognize on sight as being worth memorization.
By that we mean that you can associate the acronym with a rough statement of the theorem | not that
the exact details of the theorem need to be memorized. And it is certainly not our intent that everything
on these lists is important enough to memorize.
Theorem RCLS [58]
We will repeatedly appeal to this theorem to determine if a system of linear equations, does, or doesn't,
have a solution. This one we will see often enough that it is worth memorizing.
Theorem HMVEI [73]
This theorem is the theoretical basis of several of our most important theorems. So keep an eye out for
it, and its descendants, as you study other proofs. For example, Theorem HMVEI [73] is critical to the
proof of Theorem SSLD [391], Theorem SSLD [391] is critical to the proof of Theorem G [407], Theorem
G [407] is critical to the proofs of the pair of similar theorems, Theorem ILTD [550] and Theorem SLTD
[569], while nally Theorem ILTD [550] and Theorem SLTD [569] are critical to the proof of an important
result, Theorem IVSED [587]. This chain of implications might not make much sense on a rst reading,
but come back later to see how some very important theorems build on the seemingly simple result that is
Theorem HMVEI [73]. Using the \nd" feature in whatever software you use to read the electronic version
of the text can be a fun way to explore these relationships.
Theorem NMRRI [84]
This theorem gives us one of simplest ways, computationally, to recognize if a matrix is nonsingular, or
singular. We will see this one often, in computational exercises especially.
Theorem NMUS [86]
Nonsingular matrices will be an important topic going forward (witness the NMEx series of theorems).
This is our rst result along these lines, a useful theorem for other proofs, and also illustrates a more
general concept from Chapter LT [515].
Version 2.30
96 Section NM Nonsingular Matrices
Version 2.30
Chapter V
Vectors
We have worked extensively in the last chapter with matrices, and some with vectors. In this chapter we
will develop the properties of vectors, while preparing to study vector spaces (Chapter VS [317]). Initially
we will depart from our study of systems of linear equations, but in Section LC [109] we will forge a
connection between linear combinations and systems of linear equations in Theorem SLSLC [112]. This
connection will allow us to understand systems of linear equations at a higher level, while consequently
discussing them less frequently.
Section VO
Vector Operations
In this section we dene some new operations involving vectors, and collect some basic properties of these
operations. Begin by recalling our denition of a column vector as an ordered list of complex numbers,
written vertically (Denition CV [27]). The collection of all possible vectors of a xed size is a commonly
used set, so we start with its denition.
Denition VSCV
Vector Space of Column Vectors
The vector space Cmis the set of all column vectors (Denition CV [27]) of size mwith entries from the
set of complex numbers, C.
(This denition contains Notation VSCV.) 4
When a set similar to this is dened using only column vectors where all the entries are from the real
numbers, it is written as Rmand is known as Euclidean m-space .
The term \vector" is used in a variety of dierent ways. We have dened it as an ordered list written
vertically. It could simply be an ordered list of numbers, and written as (2 ;3; 1;6). Or it could be
interpreted as a point in mdimensions, such as (3 ;4; 2) representing a point in three dimensions relative
tox,yandzaxes. With an interpretation as a point, we can construct an arrow from the origin to the
point which is consistent with the notion that a vector has direction and magnitude.
All of these ideas can be shown to be related and equivalent, so keep that in mind as you connect the
ideas of this course with ideas from other disciplines. For now, we'll stick with the idea that a vector is a
just a list of numbers, in some particular order.
97
98 Section VO Vector Operations
Subsection VEASM
Vector Equality, Addition, Scalar Multiplication
We start our study of this set by rst dening what it means for two vectors to be the same.
Denition CVE
Column Vector Equality
Suppose that u;v2Cm. Then uandvareequal , written u=vif
[u]i= [v]i 1im
(This denition contains Notation CVE.) 4
Now this may seem like a silly (or even stupid) thing to say so carefully. Of course two vectors are
equal if they are equal for each corresponding entry! Well, this is not as silly as it appears. We will see a
few occasions later where the obvious denition is notthe right one. And besides, in doing mathematics
we need to be very careful about making all the necessary denitions and making them unambiguous. And
we've done that here.
Notice now that the symbol `=' is now doing triple-duty. We know from our earlier education what it
means for two numbers (real or complex) to be equal, and we take this for granted. In Denition SE [762]
we dened what it meant for two sets to be equal. Now we have dened what it means for two vectors
to be equal, and that denition builds on our denition for when two numbers are equal when we use the
conditionui=vifor all 1im. So think carefully about your objects when you see an equal sign and
think about just which notion of equality you have encountered. This will be especially important when
you are asked to construct proofs whose conclusion states that two objects are equal.
OK, let's do an example of vector equality that begins to hint at the utility of this denition.
Example VESE
Vector equality for a system of equations
Consider the system of linear equations in Archetype B [786],
7x1 6x2 12x3= 33
5x1+ 5x2+ 7x3= 24
x1+ 4x3= 5
Note the use of three equals signs | each indicates an equality of numbers (the linear expressions are
numbers when we evaluate them with xed values of the variable quantities). Now write the vector
equality,2
4 7x1 6x2 12x3
5x1+ 5x2+ 7x3
x1+ 4x33
5=2
4 33
24
53
5:
By Denition CVE [98], this single equality (of two column vectors) translates into three simultaneous
equalities of numbers that form the system of equations. So with this new notion of vector equality we
can become less reliant on referring to systems ofsimultaneous equations. There's more to vector equality
than just this, but this is a good example for starters and we will develop it further.
We will now dene two operations on the set Cm. By this we mean well-dened procedures that
somehow convert vectors into other vectors. Here are two of the most basic denitions of the entire course.
Denition CVA
Column Vector Addition
Suppose that u;v2Cm. The sum ofuandvis the vector u+vdened by
[u+v]i= [u]i+ [v]i 1im
Version 2.30
Subsection VO.VEASM Vector Equality, Addition, Scalar Multiplication 99
(This denition contains Notation CVA.) 4
So vector addition takes two vectors of the same size and combines them (in a natural way!) to create a
new vector of the same size. Notice that this denition is required, even if we agree that this is the obvious,
right, natural or correct way to do it. Notice too that the symbol `+' is being recycled. We all know how
to add numbers , but now we have the same symbol extended to double-duty and we use it to indicate how
to add two new objects, vectors. And this denition of our new meaning is built on our previous meaning
of addition via the expressions ui+vi. Think about your objects, especially when doing proofs. Vector
addition is easy, here's an example from C4.
Example VA
Addition of two vectors in C4
If
u=2
6642
3
4
23
775v=2
664 1
5
2
73
775
then
u+v=2
6642
3
4
23
775+2
664 1
5
2
73
775=2
6642 + ( 1)
3 + 5
4 + 2
2 + ( 7)3
775=2
6641
2
6
53
775:
Our second operation takes two objects of dierent types, specically a number and a vector, and
combines them to create another vector. In this context we call a number a scalar in order to emphasize
that it is not a vector.
Denition CVSM
Column Vector Scalar Multiplication
Suppose u2Cmand2C, then the scalar multiple ofubyis the vector udened by
[u]i=[u]i 1im
(This denition contains Notation CVSM.) 4
Notice that we are doing a kind of multiplication here, but we are dening a new type, perhaps in what
appears to be a natural way. We use juxtaposition (smashing two symbols together side-by-side) to denote
this operation rather than using a symbol like we did with vector addition. So this can be another source
of confusion. When two symbols are next to each other, are we doing regular old multiplication, the kind
we've done for years, or are we doing scalar vector multiplication, the operation we just dened? Think
about your objects | if the rst object is a scalar, and the second is a vector, then it must be that we are
doing our new operation, and the result of this operation will be another vector.
Notice how consistency in notation can be an aid here. If we write scalars as lower case Greek letters
from the start of the alphabet (such as ,, . . . ) and write vectors in bold Latin letters from the end
of the alphabet ( u,v, . . . ), then we have some hints about what type of objects we are working with.
This can be a blessing anda curse, since when we go read another book about linear algebra, or read an
application in another discipline (physics, economics, . . . ) the types of notation employed may be very
dierent and hence unfamiliar.
Again, computationally, vector scalar multiplication is very easy.
Version 2.30
100 Section VO Vector Operations
Example CVSM
Scalar multiplication in C5
If
u=2
666643
1
2
4
13
77775
and= 6, then
u= 62
666643
1
2
4
13
77775=2
666646(3)
6(1)
6( 2)
6(4)
6( 1)3
77775=2
6666418
6
12
24
63
77775:
Vector addition and scalar multiplication are the most natural and basic operations to perform on
vectors, so it should be easy to have your computational device form a linear combination. See: Compu-
tation VLC.MMA [746] Computation VLC.TI86 [750] Computation VLC.TI83 [752] Computation
VLC.SAGE [755]
Subsection VSP
Vector Space Properties
With denitions of vector addition and scalar multiplication we can state, and prove, several properties of
each operation, and some properties that involve their interplay. We now collect ten of them here for later
reference.
Theorem VSPCV
Vector Space Properties of Column Vectors
Suppose that Cmis the set of column vectors of size m(Denition VSCV [97]) with addition and scalar
multiplication as dened in Denition CVA [98] and Denition CVSM [99]. Then
ACC Additive Closure, Column Vectors
Ifu;v2Cm, then u+v2Cm.
SCC Scalar Closure, Column Vectors
If2Candu2Cm, thenu2Cm.
CC Commutativity, Column Vectors
Ifu;v2Cm, then u+v=v+u.
AAC Additive Associativity, Column Vectors
Ifu;v;w2Cm, then u+ (v+w) = (u+v) +w.
ZC Zero Vector, Column Vectors
There is a vector, 0, called the zero vector , such that u+0=ufor all u2Cm.
AIC Additive Inverses, Column Vectors
Ifu2Cm, then there exists a vector u2Cmso that u+ ( u) =0.
SMAC Scalar Multiplication Associativity, Column Vectors
If; 2Candu2Cm, then(u) = ()u.
Version 2.30
Subsection VO.READ Reading Questions 101
DVAC Distributivity across Vector Addition, Column Vectors
If2Candu;v2Cm, then(u+v) =u+v.
DSAC Distributivity across Scalar Addition, Column Vectors
If; 2Candu2Cm, then (+)u=u+u.
OC One, Column Vectors
Ifu2Cm, then 1 u=u.
Proof While some of these properties seem very obvious, they all require proof. However, the proofs are
not very interesting, and border on tedious. We'll prove one version of distributivity very carefully, and
you can test your proof-building skills on some of the others. We need to establish an equality, so we will
do so by beginning with one side of the equality, apply various denitions and theorems (listed to the right
of each step) to massage the expression from the left into the expression on the right. Here we go with a
proof of Property DSAC [101]. For 1 im,
[(+)u]i= (+) [u]i Denition CVSM [99]
=[u]i+[u]i Distributivity in C
= [u]i+ [u]i Denition CVSM [99]
= [u+u]i Denition CVA [98]
Since the individual components of the vectors ( +)uandu+uare equal for alli, 1im,
Denition CVE [98] tells us the vectors are equal.
Many of the conclusions of our theorems can be characterized as \identities," especially when we are
establishing basic properties of operations such as those in this section. Most of the properties listed in
Theorem VSPCV [100] are examples. So some advice about the style we use for proving identities is
appropriate right now. Have a look at Technique PI [771].
Be careful with the notion of the vector u. This is a vector that we add to uso that the result is the
particular vector 0. This is basically a property of vector addition. It happens that we can compute u
using the other operation, scalar multiplication. We can prove this directly by writing that
[ u]i= [u]i= ( 1) [u]i= [( 1)u]i
We will see later how to derive this property as a consequence of several of the ten properties listed in
Theorem VSPCV [100].
Similarly, we will often write something you would immediately recognize as \vector subtraction." This
could be placed on a rm theoretical foundation | as you can do yourself with Exercise VO.T30 [104].
A nal note. Property AAC [100] implies that we do not have to be careful about how we \parenthesize"
the addition of vectors. In other words, there is nothing to be gained by writing ( u+v) + (w+ (x+y))
rather than u+v+w+x+y, since we get the same result no matter which order we choose to perform
the four additions. So we won't be careful about using parentheses this way.
Subsection READ
Reading Questions
1. Where have you seen vectors used before in other courses? How were they dierent?
2. In words, when are two vectors equal?
Version 2.30
102 Section VO Vector Operations
3. Perform the following computation with vector operations
22
41
5
03
5+ ( 3)2
47
6
53
5
Version 2.30
Subsection VO.EXC Exercises 103
Subsection EXC
Exercises
C10 Compute
42
666642
3
4
1
03
77775+ ( 2)2
666641
2
5
2
43
77775+2
66664 1
3
0
1
23
77775
Contributed by Robert Beezer Solution [106]
C11 Solve the given vector equation for x, or explain why no solution exists:
32
41
2
13
5+ 42
42
0
x3
5=2
411
6
173
5
Contributed by Chris Black Solution [106]
C12 Solve the given vector equation for , or explain why no solution exists:
2
41
2
13
5+ 42
43
4
23
5=2
4 1
0
43
5
Contributed by Chris Black Solution [106]
C13 Solve the given vector equation for , or explain why no solution exists:
2
43
2
23
5+2
46
1
23
5=2
40
3
63
5
Contributed by Chris Black Solution [106]
C14 Findandthat solve the vector equation.
1
0
+0
1
=3
2
Contributed by Chris Black Solution [107]
C15 Findandthat solve the vector equation.
2
1
+1
3
=5
0
Contributed by Chris Black Solution [107]
T5 Fill in each blank with an appropriate vector space property to provide justication for the proof of
the following proposition:
Proposition 1. For any vectors u;v;w2Cm, ifu+v=u+w, then v=w.
Proof : Let u;v;w2Cm, and suppose u+v=u+w.
Version 2.30
104 Section VO Vector Operations
1. Then u+ (u+v) = u+ (u+w), Additive Property of Equality
2. so ( u+u) +v= ( u+u) +w.
3. Thus, we have 0+v=0+w,
4. and it follows that v=w.
Thus, for any vectors u;v;w2Cm, ifu+v=u+w, then v=w.
Contributed by Chris Black Solution [107]
T6 Fill in each blank with an appropriate vector space property to provide justication for the proof of
the following proposition:
Proposition 2. For any vector u2Cm, 0u=0:
Proof : Let u2Cm.
1. Since 0 + 0 = 0, we have 0 u= (0 + 0) u. Substitution
2. We then have 0 u= 0u+ 0u.
3. It follows that 0 u+ [ (0u)] = (0 u+ 0u) + [ (0u)], Additive Property of Equality
4. so 0 u+ [ (0u)] = 0 u+ (0u+ [ (0u)]),
5. so that 0= 0u+0,
6. and thus 0= 0u.
Thus, for any vector u2Cm, 0u=0.
Contributed by Chris Black Solution [107]
T7 Fill in each blank with an appropriate vector space property to provide justication for the proof of
the following proposition:
Proposition 3. For any scalar c,c0=0.
Proof : Letcbe an arbitrary scalar.
1. Thenc0=c(0+0),
2. soc0=c0+c0.
3. We then have c0+ ( c0) = (c0+c0) + ( c0), Additive Property of Equality
4. so that c0+ ( c0) =c0+ (c0+ ( c0)).
5. It follows that 0=c0+0,
6. and nally we have 0=c0.
Thus, for any scalar c,c0=0.
Contributed by Chris Black Solution [107]
T13 Prove Property CC [100] of Theorem VSPCV [100]. Write your proof in the style of the proof of
Property DSAC [101] given in this section.
Contributed by Robert Beezer Solution [107]
T17 Prove Property SMAC [100] of Theorem VSPCV [100]. Write your proof in the style of the proof
of Property DSAC [101] given in this section.
Contributed by Robert Beezer
T18 Prove Property DVAC [101] of Theorem VSPCV [100]. Write your proof in the style of the proof of
Property DSAC [101] given in this section.
Contributed by Robert Beezer
T30 Suppose uandvare two vectors in Cm. Dene a new operation, called \subtraction," as the new
vector denoted u vand dened by
[u v]i= [u]i [v]i 1im
Prove that we can express the subtraction of two vectors in terms of our two basic operations. More
precisely, prove that u v=u+ ( 1)v. So in a sense, subtraction is not something new and dierent,
but is just a convenience. Mimic the style of similar proofs in this section.
Contributed by Robert Beezer
Version 2.30
Subsection VO.EXC Exercises 105
T31 Review the denition of vector subtraction in Exercise VO.T30 [104]. Prove, by using counterex-
amples, that vector subtraction is not commutative and not associative.
Contributed by Robert Beezer
T32 Review the denition of vector subtraction in Exercise VO.T30 [104]. Prove that vector subtraction
obeys a distributive property. Specically, prove that (u v) =u v.
Can you give two dierent proofs? Base one on the denition given in Exercise VO.T30 [104] and base
the other on the equivalent formulation proved in Exercise VO.T30 [104].
Contributed by Robert Beezer
Version 2.30
106 Section VO Vector Operations
Subsection SOL
Solutions
C10 Contributed by Robert Beezer Statement [103]2
666645
13
26
1
63
77775
C11 Contributed by Chris Black Statement [103]
Performing the indicated operations (Denition CVA [98], Denition CVSM [99]), we obtain the vector
equations
2
411
6
173
5= 32
41
2
13
5+ 42
42
0
x3
5=2
411
6
3 + 4x3
5
Since the entries of the vectors must be equal by Denition CVE [98], we have 3 + 4x= 17, which leads
tox= 5.
C12 Contributed by Chris Black Statement [103]
Performing the indicated operations (Denition CVA [98], Denition CVSM [99]), we obtain the vector
equations
2
4
2
3
5+2
412
16
83
5=2
4+ 12
2+ 16
+ 83
5=2
4 1
0
43
5
Thus, if a solution exists, by Denition CVE [98] then must satisfy the three equations:
+ 12 = 1
2+ 16 = 0
+ 8 = 4
which leads to = 13,= 8 and= 4. Since cannot simultaneously have three dierent values,
there is no solution to the original vector equation.
C13 Contributed by Chris Black Statement [103]
Performing the indicated operations (Denition CVA [98], Denition CVSM [99]), we obtain the vector
equations
2
43
2
23
5+2
46
1
23
5=2
43+ 6
2+ 1
2+ 23
5=2
40
3
63
5
Thus, if a solution exists, by Denition CVE [98] then must satisfy the three equations:
3+ 6 = 0
2+ 1 = 3
2+ 2 = 6
which leads to 3 = 6, 2= 4 and 2= 4. And thus, the solution to the given vector equation is
= 2.
Version 2.30
Subsection VO.SOL Solutions 107
C14 Contributed by Chris Black Statement [103]
Performing the indicated operations (Denition CVA [98], Denition CVSM [99]), we obtain the vector
equations
3
2
=1
0
+0
1
=+ 0
0 +
=
Since the entries of the vectors must be equal by Denition CVE [98], we have = 3 and= 2.
C15 Contributed by Chris Black Statement [103]
Performing the indicated operations (Denition CVA [98], Denition CVSM [99]), we obtain the vector
equations
5
0
=2
1
+1
3
=2+
+ 3
Since the entries of the vectors must be equal by Denition CVE [98], we obtain the system of equations
2+= 5
+ 3= 0:
which we can solve by row-reducing the augmented matrix of the system,
2 1 5
1 3 0
RREF !10 3
01 1
Thus, the only solution is = 3,= 1.
T5 Contributed by Chris Black Statement [103]
1. (Additive Property of Equality)
2. Additive Associativity Property AAC [100]
3. Additive Inverses Property AIC [100]
4. Zero Vector Property ZC [100]
T6 Contributed by Chris Black Statement [104]
1. (Substitution)
2. Distributive across Scalar Addition Property DSAC [101]
3. (Additive Property of Equality)
4. Additive Associativity Property AAC [100]
5. Additive Inverses Property AIC [100]
6. Zero Vector Property ZC [100]
T7 Contributed by Chris Black Statement [104]
1. Zero Vector Property ZC [100]
2. Distributive across Vector Addition Property DVAC [101]
3. (Additive Property of Equality)
4. Additive Associativity Property AAC [100]
5. Additive Inverses Property AIC [100]
6. Zero Vector Property ZC [100]
T13 Contributed by Robert Beezer Statement [104]
For all 1im,
[u+v]i= [u]i+ [v]i Denition CVA [98]
Version 2.30
108 Section VO Vector Operations
= [v]i+ [u]i Commutativity in C
= [v+u]i Denition CVA [98]
With equality of each component of the vectors u+vandv+ubeing equal Denition CVE [98] tells us
the two vectors are equal.
Version 2.30
Section LC Linear Combinations 109
Section LC
Linear Combinations
In Section VO [97] we dened vector addition and scalar multiplication. These two operations combine
nicely to give us a construction known as a linear combination, a construct that we will work with through-
out this course.
Subsection LC
Linear Combinations
Denition LCCV
Linear Combination of Column Vectors
Givennvectors u1;u2;u3; :::; unfromCmandnscalars1; 2; 3; :::; n, their linear combination
is the vector
1u1+2u2+3u3++nun
4
So this denition takes an equal number of scalars and vectors, combines them using our two new
operations (scalar multiplication and vector addition) and creates a single brand-new vector, of the same
size as the original vectors. When a denition or theorem employs a linear combination, think about the
nature of the objects that go into its creation (lists of scalars and vectors), and the type of object that
results (a single vector). Computationally, a linear combination is pretty easy.
Example TLC
Two linear combinations in C6
Suppose that
1= 1 2= 4 3= 2 4= 1
and
u1=2
66666642
4
3
1
2
93
7777775u2=2
66666646
3
0
2
1
43
7777775u3=2
6666664 5
2
1
1
3
03
7777775u4=2
66666643
2
5
7
1
33
7777775
then their linear combination is
1u1+2u2+3u3+4u4= (1)2
66666642
4
3
1
2
93
7777775+ ( 4)2
66666646
3
0
2
1
43
7777775+ (2)2
6666664 5
2
1
1
3
03
7777775+ ( 1)2
66666643
2
5
7
1
33
7777775
Version 2.30
110 Section LC Linear Combinations
=2
66666642
4
3
1
2
93
7777775+2
6666664 24
12
0
8
4
163
7777775+2
6666664 10
4
2
2
6
03
7777775+2
6666664 3
2
5
7
1
33
7777775=2
6666664 35
6
4
4
9
103
7777775
A dierent linear combination, of the same set of vectors, can be formed with dierent scalars. Take
1= 3 2= 0 3= 5 4= 1
and form the linear combination
1u1+2u2+3u3+4u4= (3)2
66666642
4
3
1
2
93
7777775+ (0)2
66666646
3
0
2
1
43
7777775+ (5)2
6666664 5
2
1
1
3
03
7777775+ ( 1)2
66666643
2
5
7
1
33
7777775
=2
66666646
12
9
3
6
273
7777775+2
66666640
0
0
0
0
03
7777775+2
6666664 25
10
5
5
15
03
7777775+2
6666664 3
2
5
7
1
33
7777775=2
6666664 22
20
1
1
10
243
7777775
Notice how we could keep our set of vectors xed, and use dierent sets of scalars to construct dierent
vectors. You might build a few new linear combinations of u1;u2;u3;u4right now. We'll be right here
when you get back. What vectors were you able to create? Do you think you could create the vector
w=2
666666413
15
5
17
2
253
7777775
with a \suitable" choice of four scalars? Do you think you could create anypossible vector from C6by
choosing the proper scalars? These last two questions are very fundamental, and time spent considering
them nowwill prove benecial later.
Our next two examples are key ones, and a discussion about decompositions is timely. Have a look at
Technique DC [772] before studying the next two examples.
Example ABLC
Archetype B as a linear combination
In this example we will rewrite Archetype B [786] in the language of vectors, vector equality and linear
combinations. In Example VESE [98] we wrote the system of m= 3 equations as the vector equality
2
4 7x1 6x2 12x3
5x1+ 5x2+ 7x3
x1+ 4x33
5=2
4 33
24
53
5
Version 2.30
Subsection LC.LC Linear Combinations 111
Now we will bust up the linear expressions on the left, rst using vector addition,
2
4 7x1
5x1
x13
5+2
4 6x2
5x2
0x23
5+2
4 12x3
7x3
4x33
5=2
4 33
24
53
5
Now we can rewrite each of these n= 3 vectors as a scalar multiple of a xed vector, where the scalar is
one of the unknown variables, converting the left-hand side into a linear combination
x12
4 7
5
13
5+x22
4 6
5
03
5+x32
4 12
7
43
5=2
4 33
24
53
5
We can now interpret the problem of solving the system of equations as determining values for the scalar
multiples that make the vector equation true. In the analysis of Archetype B [786], we were able to
determine that it had only one solution. A quick way to see this is to row-reduce the coecient matrix
to the 33 identity matrix and apply Theorem NMRRI [84] to determine that the coecient matrix is
nonsingular. Then Theorem NMUS [86] tells us that the system of equations has a unique solution. This
solution is
x1= 3 x2= 5 x3= 2
So, in the context of this example, we can express the fact that these values of the variables are a solution
by writing the linear combination,
( 3)2
4 7
5
13
5+ (5)2
4 6
5
03
5+ (2)2
4 12
7
43
5=2
4 33
24
53
5
Furthermore, these are the only three scalars that will accomplish this equality, since they come from a
unique solution.
Notice how the three vectors in this example are the columns of the coecient matrix of the system of
equations. This is our rst hint of the important interplay between the vectors that form the columns of
a matrix, and the matrix itself.
With any discussion of Archetype A [781] or Archetype B [786] we should be sure to contrast with the
other.
Example AALC
Archetype A as a linear combination
As a vector equality, Archetype A [781] can be written as
2
4x1 x2+ 2x3
2x1+x2+x3
x1+x23
5=2
41
8
53
5
Now bust up the linear expressions on the left, rst using vector addition,
2
4x1
2x1
x13
5+2
4 x2
x2
x23
5+2
42x3
x3
0x33
5=2
41
8
53
5
Rewrite each of these n= 3 vectors as a scalar multiple of a xed vector, where the scalar is one of the
unknown variables, converting the left-hand side into a linear combination
x12
41
2
13
5+x22
4 1
1
13
5+x32
42
1
03
5=2
41
8
53
5
Version 2.30
112 Section LC Linear Combinations
Row-reducing the augmented matrix for Archetype A [781] leads to the conclusion that the system is
consistent and has free variables, hence innitely many solutions. So for example, the two solutions
x1= 2 x2= 3 x3= 1
x1= 3 x2= 2 x3= 0
can be used together to say that,
(2)2
41
2
13
5+ (3)2
4 1
1
13
5+ (1)2
42
1
03
5=2
41
8
53
5= (3)2
41
2
13
5+ (2)2
4 1
1
13
5+ (0)2
42
1
03
5
Ignore the middle of this equation, and move all the terms to the left-hand side,
(2)2
41
2
13
5+ (3)2
4 1
1
13
5+ (1)2
42
1
03
5+ ( 3)2
41
2
13
5+ ( 2)2
4 1
1
13
5+ ( 0)2
42
1
03
5=2
40
0
03
5
Regrouping gives
( 1)2
41
2
13
5+ (1)2
4 1
1
13
5+ (1)2
42
1
03
5=2
40
0
03
5
Notice that these three vectors are the columns of the coecient matrix for the system of equations in
Archetype A [781]. This equality says there is a linear combination of those columns that equals the vector
of all zeros. Give it some thought, but this says that
x1= 1 x2= 1 x3= 1
is a nontrivial solution to the homogeneous system of equations with the coecient matrix for the original
system in Archetype A [781]. In particular, this demonstrates that this coecient matrix is singular.
There's a lot going on in the last two examples. Come back to them in a while and make some
connections with the intervening material. For now, we will summarize and explain some of this behavior
with a theorem.
Theorem SLSLC
Solutions to Linear Systems are Linear Combinations
Denote the columns of the mnmatrixAas the vectors A1;A2;A3; :::; An. Then xis a solution to
the linear system of equations LS(A;b) if and only if bequals the linear combination of the columns of
Aformed with the entries of x,
[x]1A1+ [x]2A2+ [x]3A3++ [x]nAn=b
Proof The proof of this theorem is as much about a change in notation as it is about making logical
deductions. Write the system of equations LS(A;b) as
a11x1+a12x2+a13x3++a1nxn=b1
a21x1+a22x2+a23x3++a2nxn=b2
a31x1+a32x2+a33x3++a3nxn=b3
...
am1x1+am2x2+am3x3++amnxn=bm
Version 2.30
Subsection LC.VFSS Vector Form of Solution Sets 113
Notice then that the entry of the coecient matrix Ain rowiand column jhas two names: aijas the
coecient of xjin equation iof the system and [ Aj]ias thei-th entry of the column vector in column
jof the coecient matrix A. Likewise, entry iofbhas two names: bifrom the linear system and [ b]i
as an entry of a vector. Our theorem is an equivalence (Technique E [768]) so we need to prove both
\directions."
(() Suppose we have the vector equality between band the linear combination of the columns of A.
Then for 1im,
bi= [b]i Notation CVC [28]
= [[x]1A1+ [x]2A2+ [x]3A3++ [x]nAn]iHypothesis
= [[x]1A1]i+ [[x]2A2]i+ [[x]3A3]i++ [[x]nAn]iDenition CVA [98]
= [x]1[A1]i+ [x]2[A2]i+ [x]3[A3]i++ [x]n[An]i Denition CVSM [99]
= [x]1ai1+ [x]2ai2+ [x]3ai3++ [x]nain Notation CVC [28]
=ai1[x]1+ai2[x]2+ai3[x]3++ain[x]n Property CMCN [758]
This says that the entries of xform a solution to equation iofLS(A;b) for all 1im, in other words,
xis a solution toLS(A;b).
()) Suppose now that xis a solution to the linear system LS(A;b). Then for all 1im,
[b]i=bi Notation CVC [28]
=ai1[x]1+ai2[x]2+ai3[x]3++ain[x]n Hypothesis
= [x]1ai1+ [x]2ai2+ [x]3ai3++ [x]nain Property CMCN [758]
= [x]1[A1]i+ [x]2[A2]i+ [x]3[A3]i++ [x]n[An]i Notation CVC [28]
= [[x]1A1]i+ [[x]2A2]i+ [[x]3A3]i++ [[x]nAn]iDenition CVSM [99]
= [[x]1A1+ [x]2A2+ [x]3A3++ [x]nAn]iDenition CVA [98]
Since the components of band the linear combination of the columns of Aagree for all 1im,
Denition CVE [98] tells us that the vectors are equal.
In other words, this theorem tells us that solutions to systems of equations are linear combinations of
thencolumn vectors of the coecient matrix ( Aj) which yield the constant vector b. Or said another way,
a solution to a system of equations LS(A;b) is an answer to the question \How can I form the vector b
as a linear combination of the columns of A?" Look through the archetypes that are systems of equations
and examine a few of the advertised solutions. In each case use the solution to form a linear combination
of the columns of the coecient matrix and verify that the result equals the constant vector (see Exercise
LC.C21 [127]).
Subsection VFSS
Vector Form of Solution Sets
We have written solutions to systems of equations as column vectors. For example Archetype B [786] has
the solution x1= 3; x2= 5; x3= 2 which we now write as
x=2
4x1
x2
x33
5=2
4 3
5
23
5
Now, we will use column vectors and linear combinations to express allof the solutions to a linear system
of equations in a compact and understandable way. First, here's two examples that will motivate our next
Version 2.30
114 Section LC Linear Combinations
theorem. This is a valuable technique, almost the equal of row-reducing a matrix, so be sure you get
comfortable with it over the course of this section.
Example VFSAD
Vector form of solutions for Archetype D
Archetype D [795] is a linear system of 3 equations in 4 variables. Row-reducing the augmented matrix
yields2
410 3 2 4
011 3 0
0 0 0 0 03
5
and we see r= 2 nonzero rows. Also, D=f1;2gso the dependent variables are then x1andx2.
F=f3;4;5gso the two free variables are x3andx4. We will express a generic solution for the system by
two slightly dierent methods, though both arrive at the same conclusion.
First, we will decompose (Technique DC [772]) a solution vector. Rearranging each equation represented
in the row-reduced form of the augmented matrix by solving for the dependent variable in each row yields
the vector equality,
2
664x1
x2
x3
x43
775=2
6644 3x3+ 2x4
x3+ 3x4
x3
x43
775
Now we will use the denitions of column vector addition and scalar multiplication to express this vector
as a linear combination,
=2
6644
0
0
03
775+2
664 3x3
x3
x3
03
775+2
6642x4
3x4
0
x43
775Denition CVA [98]
=2
6644
0
0
03
775+x32
664 3
1
1
03
775+x42
6642
3
0
13
775Denition CVSM [99]
We will develop the same linear combination a bit quicker, using three steps. While the method above is
instructive, the method below will be our preferred approach.
Step 1. Write the vector of variables as a xed vector, plus a linear combination of n rvectors, using
the free variables as the scalars.
x=2
664x1
x2
x3
x43
775=2
6643
775+x32
6643
775+x42
6643
775
Step 2. Use 0's and 1's to ensure equality for the entries of the the vectors with indices in F(corresponding
to the free variables).
x=2
664x1
x2
x3
x43
775=2
6640
03
775+x32
6641
03
775+x42
6640
13
775
Version 2.30
Subsection LC.VFSS Vector Form of Solution Sets 115
Step 3. For each dependent variable, use the augmented matrix to formulate an equation expressing the
dependent variable as a constant plus multiples of the free variables. Convert this equation into entries of
the vectors that ensure equality for each dependent variable, one at a time.
x1= 4 3x3+ 2x4) x=2
664x1
x2
x3
x43
775=2
6644
0
03
775+x32
664 3
1
03
775+x42
6642
0
13
775
x2= 0 1x3+ 3x4) x=2
664x1
x2
x3
x43
775=2
6644
0
0
03
775+x32
664 3
1
1
03
775+x42
6642
3
0
13
775
This nal form of a typical solution is especially pleasing and useful. For example, we can build solutions
quickly by choosing values for our free variables, and then compute a linear combination. Such as
x3= 2; x4= 5) x=2
664x1
x2
x3
x43
775=2
6644
0
0
03
775+ (2)2
664 3
1
1
03
775+ ( 5)2
6642
3
0
13
775=2
664 12
17
2
53
775
or,
x3= 1; x4= 3) x=2
664x1
x2
x3
x43
775=2
6644
0
0
03
775+ (1)2
664 3
1
1
03
775+ (3)2
6642
3
0
13
775=2
6647
8
1
33
775
You'll nd the second solution listed in the write-up for Archetype D [795], and you might check the rst
solution by substituting it back into the original equations.
While this form is useful for quickly creating solutions, its even better because it tells us exactly what
every solution looks like. We know the solution set is innite, which is pretty big, but now we can say that
a solution is some multiple of2
664 3
1
1
03
775plus a multiple of2
6642
3
0
13
775plus the xed vector2
6644
0
0
03
775. Period. So it only
takes us three vectors to describe the entire innite solution set, provided we also agree on how to combine
the three vectors into a linear combination.
This is such an important and fundamental technique, we'll do another example.
Example VFS
Vector form of solutions
Consider a linear system of m= 5 equations in n= 7 variables, having the augmented matrix A.
A=2
666642 1 1 2 2 1 5 21
1 1 3 1 1 1 2 5
1 2 8 5 1 1 6 15
3 3 9 3 6 5 2 24
2 1 1 2 1 1 9 303
77775
Version 2.30
116 Section LC Linear Combinations
Row-reducing we obtain the matrix
B=2
66666410 2 3 0 0 9 15
01 5 4 0 0 8 10
0 0 0 0 10 6 11
0 0 0 0 0 1 7 21
0 0 0 0 0 0 0 03
777775
and we see r= 4 nonzero rows. Also, D=f1;2;5;6gso the dependent variables are then x1; x2; x5;and
x6.F=f3;4;7;8gso then r= 3 free variables are x3; x4andx7. We will express a generic solution
for the system by two dierent methods: both a decomposition and a construction.
First, we will decompose (Technique DC [772]) a solution vector. Rearranging each equation represented
in the row-reduced form of the augmented matrix by solving for the dependent variable in each row yields
the vector equality,
2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
66666666415 2x3+ 3x4 9x7
10 + 5x3 4x4+ 8x7
x3
x4
11 + 6x7
21 7x7
x73
777777775
Now we will use the denitions of column vector addition and scalar multiplication to decompose this
generic solution vector as a linear combination,
=2
66666666415
10
0
0
11
21
03
777777775+2
666666664 2x3
5x3
x3
0
0
0
03
777777775+2
6666666643x4
4x4
0
x4
0
0
03
777777775+2
666666664 9x7
8x7
0
0
6x7
7x7
x73
777777775Denition CVA [98]
=2
66666666415
10
0
0
11
21
03
777777775+x32
666666664 2
5
1
0
0
0
03
777777775+x42
6666666643
4
0
1
0
0
03
777777775+x72
666666664 9
8
0
0
6
7
13
777777775Denition CVSM [99]
We will now develop the same linear combination a bit quicker, using three steps. While the method above
is instructive, the method below will be our preferred approach.
Step 1. Write the vector of variables as a xed vector, plus a linear combination of n rvectors, using
the free variables as the scalars.
x=2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
6666666643
777777775+x32
6666666643
777777775+x42
6666666643
777777775+x72
6666666643
777777775
Version 2.30
Subsection LC.VFSS Vector Form of Solution Sets 117
Step 2. Use 0's and 1's to ensure equality for the entries of the the vectors with indices in F(corresponding
to the free variables).
x=2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
6666666640
0
03
777777775+x32
6666666641
0
03
777777775+x42
6666666640
1
03
777777775+x72
6666666640
0
13
777777775
Step 3. For each dependent variable, use the augmented matrix to formulate an equation expressing the
dependent variable as a constant plus multiples of the free variables. Convert this equation into entries of
the vectors that ensure equality for each dependent variable, one at a time.
x1= 15 2x3+ 3x4 9x7) x=2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
66666666415
0
0
03
777777775+x32
666666664 2
1
0
03
777777775+x42
6666666643
0
1
03
777777775+x72
666666664 9
0
0
13
777777775
x2= 10 + 5x3 4x4+ 8x7) x=2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
66666666415
10
0
0
03
777777775+x32
666666664 2
5
1
0
03
777777775+x42
6666666643
4
0
1
03
777777775+x72
666666664 9
8
0
0
13
777777775
x5= 11 + 6x7 ) x=2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
66666666415
10
0
0
11
03
777777775+x32
666666664 2
5
1
0
0
03
777777775+x42
6666666643
4
0
1
0
03
777777775+x72
666666664 9
8
0
0
6
13
777777775
x6= 21 7x7 ) x=2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
66666666415
10
0
0
11
21
03
777777775+x32
666666664 2
5
1
0
0
0
03
777777775+x42
6666666643
4
0
1
0
0
03
777777775+x72
666666664 9
8
0
0
6
7
13
777777775
This nal form of a typical solution is especially pleasing and useful. For example, we can build solutions
quickly by choosing values for our free variables, and then compute a linear combination. For example
x3= 2; x4= 4; x7= 3)
Version 2.30
118 Section LC Linear Combinations
x=2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
66666666415
10
0
0
11
21
03
777777775+ (2)2
666666664 2
5
1
0
0
0
03
777777775+ ( 4)2
6666666643
4
0
1
0
0
03
777777775+ (3)2
666666664 9
8
0
0
6
7
13
777777775=2
666666664 28
40
2
4
29
42
33
777777775
or perhaps,
x3= 5; x4= 2; x7= 1)
x=2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
66666666415
10
0
0
11
21
03
777777775+ (5)2
666666664 2
5
1
0
0
0
03
777777775+ (2)2
6666666643
4
0
1
0
0
03
777777775+ (1)2
666666664 9
8
0
0
6
7
13
777777775=2
6666666642
15
5
2
17
28
13
777777775
or even,
x3= 0; x4= 0; x7= 0)
x=2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
66666666415
10
0
0
11
21
03
777777775+ (0)2
666666664 2
5
1
0
0
0
03
777777775+ (0)2
6666666643
4
0
1
0
0
03
777777775+ (0)2
666666664 9
8
0
0
6
7
13
777777775=2
66666666415
10
0
0
11
21
03
777777775
So we can compactly express allof the solutions to this linear system with just 4 xed vectors, provided
we agree how to combine them in a linear combinations to create solution vectors.
Suppose you were told that the vector wbelow was a solution to this system of equations. Could you
turn the problem around and write was a linear combination of the four vectors c,u1,u2,u3? (See
Exercise LC.M11 [128].)
w=2
666666664100
75
7
9
37
35
83
777777775c=2
66666666415
10
0
0
11
21
03
777777775u1=2
666666664 2
5
1
0
0
0
03
777777775u2=2
6666666643
4
0
1
0
0
03
777777775u3=2
666666664 9
8
0
0
6
7
13
777777775
Did you think a few weeks ago that you could so quickly and easily list allthe solutions to a linear
system of 5 equations in 7 variables?
We'll now formalize the last two (important) examples as a theorem.
Theorem VFSLS
Vector Form of Solutions to Linear Systems
Suppose that [ Ajb] is the augmented matrix for a consistent linear system LS(A;b) ofmequations in
Version 2.30
Subsection LC.VFSS Vector Form of Solution Sets 119
nvariables. Let Bbe a row-equivalent m(n+ 1) matrix in reduced row-echelon form. Suppose that
Bhasrnonzero rows, columns without leading 1's with indices F=ff1; f2; f3; :::; fn r; n+ 1g, and
columns with leading 1's (pivot columns) having indices D=fd1; d2; d3; :::; drg. Dene vectors c,uj,
1jn rof sizenby
[c]i=(
0 if i2F
[B]k;n+1ifi2D,i=dk
[uj]i=8
><
>:1 if i2F,i=fj
0 if i2F,i6=fj
[B]k;fjifi2D,i=dk:
Then the set of solutions to the system of equations LS(A;b) is
S=fc+1u1+2u2+3u3++n run rj1; 2; 3; :::; n r2Cg
Proof First,LS(A;b) is equivalent to the linear system of equations that has the matrix Bas its
augmented matrix (Theorem REMES [31]), so we need only show that Sis the solution set for the system
withBas its augmented matrix. The conclusion of this theorem is that the solution set is equal to the set
S, so we will apply Denition SE [762].
We begin by showing that every element of Sis indeed a solution to the system. Let 1; 2; 3; :::; n r
be one choice of the scalars used to describe elements of S. So an arbitrary element of S, which we will
consider as a proposed solution is
x=c+1u1+2u2+3u3++n run r
Whenr+ 1`m, row`of the matrix Bis a zero row, so the equation represented by that row is
always true, no matter which solution vector we propose. So concentrate on rows representing equations
1`r. We evaluate equation `of the system represented by Bwith the proposed solution vector x
and refer to the value of the left-hand side of the equation as `,
`= [B]`1[x]1+ [B]`2[x]2+ [B]`3[x]3++ [B]`n[x]n
Since [B]`di= 0 for all 1ir, except that [ B]`d`= 1, we see that `simplies to
`= [x]d`+ [B]`f1[x]f1+ [B]`f2[x]f2+ [B]`f3[x]f3++ [B]`fn r[x]fn r
Notice that for 1 in r
[x]fi= [c]fi+1[u1]fi+2[u2]fi+3[u3]fi++i[ui]fi++n r[un r]fi
= 0 +1(0) +2(0) +3(0) ++i(1) ++n r(0)
=i
So`simplies further, and we expand the rst term
`= [x]d`+ [B]`f11+ [B]`f22+ [B]`f33++ [B]`fn rn r
= [c+1u1+2u2+3u3++n run r]d`+
[B]`f11+ [B]`f22+ [B]`f33++ [B]`fn rn r
= [c]d`+1[u1]d`+2[u2]d`+3[u3]d`++n r[un r]d`+
[B]`f11+ [B]`f22+ [B]`f33++ [B]`fn rn r
Version 2.30
120 Section LC Linear Combinations
= [B]`;n+1+1( [B]`;f1) +2( [B]`;f2) +3( [B]`;f3) ++n r( [B]`;fn r)+
[B]`f11+ [B]`f22+ [B]`f33++ [B]`fn rn r
= [B]`;n+1
So`began as the left-hand side of equation `of the system represented by Band we now know it equals
[B]`;n+1, the constant term for equation `of this system. So the arbitrarily chosen vector from Smakes
every equation of the system true, and therefore is a solution to the system. So all the elements of Sare
solutions to the system.
For the second half of the proof, assume that xis a solution vector for the system having Bas its
augmented matrix. For convenience and clarity, denote the entries of xbyxi, in other words, xi= [x]i.
We desire to show that this solution vector is also an element of the set S. Begin with the observation
that a solution vector's entries makes equation `of the system true for all 1 `m,
[B]`;1x1+ [B]`;2x2+ [B]`;3x3++ [B]`;nxn= [B]`;n+1
When`r, the pivot columns of Bhave zero entries in row `with the exception of column d`, which will
contain a 1. So for 1 `r, equation`simplies to
1xd`+ [B]`;f1xf1+ [B]`;f2xf2+ [B]`;f3xf3++ [B]`;fn rxfn r= [B]`;n+1
This allows us to write,
[x]d`=xd`
= [B]`;n+1 [B]`;f1xf1 [B]`;f2xf2 [B]`;f3xf3 [B]`;fn rxfn r
= [c]d`+xf1[u1]d`+xf2[u2]d`+xf3[u3]d`++xfn r[un r]d`
=
c+xf1u1+xf2u2+xf3u3++xfn run r
d`
This tells us that the entries of the solution vector xcorresponding to dependent variables (indices in D),
are equal to those of a vector in the set S. We still need to check the other entries of the solution vector x
corresponding to the free variables (indices in F) to see if they are equal to the entries of the same vector
in the setS. To this end, suppose i2Fandi=fj. Then
[x]i=xi=xfj
= 0 + 0xf1+ 0xf2+ 0xf3++ 0xfj 1+ 1xfj+ 0xfj+1++ 0xfn r
= [c]i+xf1[u1]i+xf2[u2]i+xf3[u3]i++xfj[uj]i++xfn r[un r]i
=
c+xf1u1+xf2u2++xfn run r
i
So entries of xandc+xf1u1+xf2u2++xfn run rare equal and therefore by Denition CVE [98] they
are equal vectors. Since xf1; xf2; xf3; :::; xfn rare scalars, this shows us that xqualies for membership
inS. So the set Scontains all of the solutions to the system.
Note that both halves of the proof of Theorem VFSLS [118] indicate that i= [x]fi. In other words,
the arbitrary scalars, i, in the description of the set Sactually have more meaning | they are the values
of the free variables [ x]fi, 1in r. So we will often exploit this observation in our descriptions of
solution sets.
Theorem VFSLS [118] formalizes what happened in the three steps of Example VFSAD [114]. The
theorem will be useful in proving other theorems, and it it is useful since it tells us an exact procedure for
simply describing an innite solution set. We could program a computer to implement it, once we have
the augmented matrix row-reduced and have checked that the system is consistent. By Knuth's denition,
this completes our conversion of linear equation solving from art into science. Notice that it even applies
Version 2.30
Subsection LC.VFSS Vector Form of Solution Sets 121
(but is overkill) in the case of a unique solution. However, as a practical matter, I prefer the three-step
process of Example VFSAD [114] when I need to describe an innite solution set. So let's practice some
more, but with a bigger example.
Example VFSAI
Vector form of solutions for Archetype I
Archetype I [816] is a linear system of m= 4 equations in n= 7 variables. Row-reducing the augmented
matrix yields2
66414 0 0 2 1 3 4
0 0 10 1 3 5 2
0 0 0 12 6 6 1
0 0 0 0 0 0 0 03
775
and we see r= 3 nonzero rows. The columns with leading 1's are D=f1;3;4gso therdependent
variables are x1; x3; x4. The columns without leading 1's are F=f2;5;6;7;8g, so then r= 4 free
variables are x2; x5; x6; x7.
Step 1. Write the vector of variables ( x) as a xed vector ( c), plus a linear combination of n r= 4
vectors ( u1;u2;u3;u4), using the free variables as the scalars.
x=2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
6666666643
777777775+x22
6666666643
777777775+x52
6666666643
777777775+x62
6666666643
777777775+x72
6666666643
777777775
Step 2. For each free variable, use 0's and 1's to ensure equality for the corresponding entry of the the
vectors. Take note of the pattern of 0's and 1's at this stage, because this is the best look you'll have at it.
We'll state an important theorem in the next section and the proof will essentially rely on this observation.
x=2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
6666666640
0
0
03
777777775+x22
6666666641
0
0
03
777777775+x52
6666666640
1
0
03
777777775+x62
6666666640
0
1
03
777777775+x72
6666666640
0
0
13
777777775
Step 3. For each dependent variable, use the augmented matrix to formulate an equation expressing the
dependent variable as a constant plus multiples of the free variables. Convert this equation into entries of
the vectors that ensure equality for each dependent variable, one at a time.
x1= 4 4x2 2x5 1x6+ 3x7)
x=2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
6666666644
0
0
0
03
777777775+x22
666666664 4
1
0
0
03
777777775+x52
666666664 2
0
1
0
03
777777775+x62
666666664 1
0
0
1
03
777777775+x72
6666666643
0
0
0
13
777777775
x3= 2 + 0x2 x5+ 3x6 5x7)
Version 2.30
122 Section LC Linear Combinations
x=2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
6666666644
0
2
0
0
03
777777775+x22
666666664 4
1
0
0
0
03
777777775+x52
666666664 2
0
1
1
0
03
777777775+x62
666666664 1
0
3
0
1
03
777777775+x72
6666666643
0
5
0
0
13
777777775
x4= 1 + 0x2 2x5+ 6x6 6x7)
x=2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
6666666644
0
2
1
0
0
03
777777775+x22
666666664 4
1
0
0
0
0
03
777777775+x52
666666664 2
0
1
2
1
0
03
777777775+x62
666666664 1
0
3
6
0
1
03
777777775+x72
6666666643
0
5
6
0
0
13
777777775
We can now use this nal expression to quickly build solutions to the system. You might try to recreate
each of the solutions listed in the write-up for Archetype I [816]. (Hint: look at the values of the free
variables in each solution, and notice that the vector chas 0's in these locations.)
Even better, we have a description of the innite solution set, based on just 5 vectors, which we combine
in linear combinations to produce solutions.
Whenever we discuss Archetype I [816] you know that's your cue to go work through Archetype J [820]
by yourself. Remember to take note of the 0/1 pattern at the conclusion of Step 2. Have fun | we won't
go anywhere while you're away.
This technique is so important, that we'll do one more example. However, an important distinction
will be that this system is homogeneous.
Example VFSAL
Vector form of solutions for Archetype L
Archetype L [829] is presented simply as the 5 5 matrix
L=2
66664 2 1 2 4 4
6 5 4 4 6
10 7 7 10 13
7 5 6 9 10
4 3 4 6 63
77775
We'll interpret it here as the coecient matrix of a homogeneous system and reference this matrix as
L. So we are solving the homogeneous system LS(L;0) havingm= 5 equations in n= 5 variables. If
we built the augmented matrix, we would add a sixth column to Lcontaining all zeros. As we did row
operations, this sixth column would remain all zeros. So instead we will row-reduce the coecient matrix,
and mentally remember the missing sixth column of zeros. This row-reduced matrix is
2
66666410 0 1 2
010 2 2
0 0 1 2 1
0 0 0 0 0
0 0 0 0 03
777775
Version 2.30
Subsection LC.VFSS Vector Form of Solution Sets 123
and we see r= 3 nonzero rows. The columns with leading 1's are D=f1;2;3gso therdependent
variables are x1; x2; x3. The columns without leading 1's are F=f4;5g, so then r= 2 free variables
arex4; x5. Notice that if we had included the all-zero vector of constants to form the augmented matrix
for the system, then the index 6 would have appeared in the set F, and subsequently would have been
ignored when listing the free variables.
Step 1. Write the vector of variables ( x) as a xed vector ( c), plus a linear combination of n r= 2
vectors ( u1;u2), using the free variables as the scalars.
x=2
66664x1
x2
x3
x4
x53
77775=2
666643
77775+x42
666643
77775+x52
666643
77775
Step 2. For each free variable, use 0's and 1's to ensure equality for the corresponding entry of the the
vectors. Take note of the pattern of 0's and 1's at this stage, even if it is not as illuminating as in other
examples.
x=2
66664x1
x2
x3
x4
x53
77775=2
666640
03
77775+x42
666641
03
77775+x52
666640
13
77775
Step 3. For each dependent variable, use the augmented matrix to formulate an equation expressing the
dependent variable as a constant plus multiples of the free variables. Don't forget about the \missing"
sixth column being full of zeros. Convert this equation into entries of the vectors that ensure equality for
each dependent variable, one at a time.
x1= 0 1x4+ 2x5) x=2
66664x1
x2
x3
x4
x53
77775=2
666640
0
03
77775+x42
66664 1
1
03
77775+x52
666642
0
13
77775
x2= 0 + 2x4 2x5) x=2
66664x1
x2
x3
x4
x53
77775=2
666640
0
0
03
77775+x42
66664 1
2
1
03
77775+x52
666642
2
0
13
77775
x3= 0 2x4+ 1x5) x=2
66664x1
x2
x3
x4
x53
77775=2
666640
0
0
0
03
77775+x42
66664 1
2
2
1
03
77775+x52
666642
2
1
0
13
77775
The vector cwill always have 0's in the entries corresponding to free variables. However, since we are
solving a homogeneous system, the row-reduced augmented matrix has zeros in column n+ 1 = 6, and
hence allthe entries of care zero. So we can write
x=2
66664x1
x2
x3
x4
x53
77775=0+x42
66664 1
2
2
1
03
77775+x52
666642
2
1
0
13
77775=x42
66664 1
2
2
1
03
77775+x52
666642
2
1
0
13
77775
Version 2.30
124 Section LC Linear Combinations
It will always happen that the solutions to a homogeneous system has c=0(even in the case of a unique
solution?). So our expression for the solutions is a bit more pleasing. In this example it says that the
solutions are all possible linear combinations of the two vectors u1=2
66664 1
2
2
1
03
77775andu2=2
666642
2
1
0
13
77775, with no
mention of any xed vector entering into the linear combination.
This observation will motivate our next section and the main denition of that section, and after that
we will conclude the section by formalizing this situation.
Subsection PSHS
Particular Solutions, Homogeneous Solutions
The next theorem tells us that in order to nd all of the solutions to a linear system of equations, it is
sucient to nd just one solution, and then nd all of the solutions to the corresponding homogeneous
system. This explains part of our interest in the null space, the set of all solutions to a homogeneous
system.
Theorem PSPHS
Particular Solution Plus Homogeneous Solutions
Suppose that wis one solution to the linear system of equations LS(A; b). Then yis a solution toLS(A; b)
if and only if y=w+zfor some vector z2N(A).
Proof LetA1;A2;A3; :::; Anbe the columns of the coecient matrix A.
(() Suppose y=w+zandz2N(A). Then
b= [w]1A1+ [w]2A2+ [w]3A3++ [w]nAn Theorem SLSLC [112]
= [w]1A1+ [w]2A2+ [w]3A3++ [w]nAn+0 Property ZC [100]
= [w]1A1+ [w]2A2+ [w]3A3++ [w]nAn Theorem SLSLC [112]
+ [z]1A1+ [z]2A2+ [z]3A3++ [z]nAn
= ([w]1+ [z]1)A1+ ([w]2+ [z]2)A2++ ([w]n+ [z]n)An Theorem VSPCV [100]
= [w+z]1A1+ [w+z]2A2+ [w+z]3A3++ [w+z]nAn Denition CVA [98]
= [y]1A1+ [y]2A2+ [y]3A3++ [y]nAn Denition of y
Applying Theorem SLSLC [112] we see that the vector yis a solution toLS(A;b).
()) Suppose yis a solution toLS(A; b). Then
0=b b
= [y]1A1+ [y]2A2+ [y]3A3++ [y]nAn Theorem SLSLC [112]
([w]1A1+ [w]2A2+ [w]3A3++ [w]nAn)
= ([y]1 [w]1)A1+ ([y]2 [w]2)A2++ ([y]n [w]n)An Theorem VSPCV [100]
= [y w]1A1+ [y w]2A2+ [y w]3A3++ [y w]nAn Denition CVA [98]
By Theorem SLSLC [112] we see that the vector y wis a solution to the homogeneous system LS(A;0)
and by Denition NSM [73], y w2N (A). In other words, y w=zfor some vector z2N (A).
Rewritten, this is y=w+z, as desired.
After proving Theorem NMUS [86] we commented (insuciently) on the negation of one half of the the-
orem. Nonsingular coecient matrices lead to unique solutions for every choice of the vector of constants.
Version 2.30
Subsection LC.PSHS Particular Solutions, Homogeneous Solutions 125
What does this say about singular matrices? A singular matrix Ahas a nontrivial null space (Theorem
NMTNS [86]). For a given vector of constants, b, the systemLS(A; b) could be inconsistent, meaning
there are no solutions. But if there is at least one solution ( w), then Theorem PSPHS [124] tells us there
will be innitely many solutions because of the role of the innite null space for a singular matrix. So a
system of equations with a singular coecient matrix never has a unique solution. Either there are no
solutions, or innitely many solutions, depending on the choice of the vector of constants ( b).
Example PSHS
Particular solutions, homogeneous solutions, Archetype D
Archetype D [795] is a consistent system of equations with a nontrivial null space. Let Adenote the
coecient matrix of this system. The write-up for this system begins with three solutions,
y1=2
6640
1
2
13
775y2=2
6644
0
0
03
775y3=2
6647
8
1
33
775
We will choose to have y1play the role of win the statement of Theorem PSPHS [124], any one of the
three vectors listed here (or others) could have been chosen. To illustrate the theorem, we should be able
to write each of these three solutions as the vector wplus a solution to the corresponding homogeneous
system of equations. Since 0is always a solution to a homogeneous system we can easily write
y1=w=w+0:
The vectors y2andy3will require a bit more eort. Solutions to the homogeneous system LS(A;0) are
exactly the elements of the null space of the coecient matrix, which by an application of Theorem VFSLS
[118] is
N(A) =8
>><
>>:x32
664 3
1
1
03
775+x42
6642
3
0
13
775x3; x42C9
>>=
>>;
Then
y2=2
6644
0
0
03
775=2
6640
1
2
13
775+2
6644
1
2
13
775=2
6640
1
2
13
775+0
BB@( 2)2
664 3
1
1
03
775+ ( 1)2
6642
3
0
13
7751
CCA=w+z2
where
z2=2
6644
1
2
13
775= ( 2)2
664 3
1
1
03
775+ ( 1)2
6642
3
0
13
775
is obviously a solution of the homogeneous system since it is written as a linear combination of the vectors
describing the null space of the coecient matrix (or as a check, you could just evaluate the equations in
the homogeneous system with z2).
Again
y3=2
6647
8
1
33
775=2
6640
1
2
13
775+2
6647
7
1
23
775=2
6640
1
2
13
775+0
BB@( 1)2
664 3
1
1
03
775+ 22
6642
3
0
13
7751
CCA=w+z3
Version 2.30
126 Section LC Linear Combinations
where
z3=2
6647
7
1
23
775= ( 1)2
664 3
1
1
03
775+ 22
6642
3
0
13
775
is obviously a solution of the homogeneous system since it is written as a linear combination of the vectors
describing the null space of the coecient matrix (or as a check, you could just evaluate the equations in
the homogeneous system with z2).
Here's another view of this theorem, in the context of this example. Grab two new solutions of the
original system of equations, say
y4=2
66411
0
3
13
775y5=2
664 4
2
4
23
775
and form their dierence,
u=2
66411
0
3
13
775 2
664 4
2
4
23
775=2
66415
2
7
33
775:
It is no accident that uis a solution to the homogeneous system (check this!). In other words, the dierence
between any two solutions to a linear system of equations is an element of the null space of the coecient
matrix. This is an equivalent way to state Theorem PSPHS [124]. (See Exercise MM.T50 [238]).
The ideas of this subsection will appear again in Chapter LT [515] when we discuss pre-images of linear
transformations (Denition PI [528]).
Subsection READ
Reading Questions
1. Earlier, a reading question asked you to solve the system of equations
2x1+ 3x2 x3= 0
x1+ 2x2+x3= 3
x1+ 3x2+ 3x3= 7
Use a linear combination to rewrite this system of equations as a vector equality.
2. Find a linear combination of the vectors
S=8
<
:2
41
3
13
5;2
42
0
43
5;2
4 1
3
53
59
=
;
that equals the vector2
41
9
113
5.
Version 2.30
Subsection LC.READ Reading Questions 127
3. The matrix below is the augmented matrix of a system of equations, row-reduced to reduced row-
echelon form. Write the vector form of the solutions to the system.
2
413 0 6 0 9
0 0 1 2 0 8
0 0 0 0 1 33
5
Version 2.30
128 Section LC Linear Combinations
Subsection EXC
Exercises
C21 Consider each archetype that is a system of equations. For individual solutions listed (both for the
original system and the corresponding homogeneous system) express the vector of constants as a linear
combination of the columns of the coecient matrix, as guaranteed by Theorem SLSLC [112]. Verify this
equality by computing the linear combination. For systems with no solutions, recognize that it is then
impossible to write the vector of constants as a linear combination of the columns of the coecient matrix.
Note too, for homogeneous systems, that the solutions give rise to linear combinations that equal the zero
vector.
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]
Archetype E [799]
Archetype F [803]
Archetype G [808]
Archetype H [812]
Archetype I [816]
Archetype J [820]
Contributed by Robert Beezer Solution [129]
C22 Consider each archetype that is a system of equations. Write elements of the solution set in vector
form, as guaranteed by Theorem VFSLS [118].
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]
Archetype E [799]
Archetype F [803]
Archetype G [808]
Archetype H [812]
Archetype I [816]
Archetype J [820]
Contributed by Robert Beezer Solution [129]
C40 Find the vector form of the solutions to the system of equations below.
2x1 4x2+ 3x3+x5= 6
x1 2x2 2x3+ 14x4 4x5= 15
x1 2x2+x3+ 2x4+x5= 1
2x1+ 4x2 12x4+x5= 7
Contributed by Robert Beezer Solution [129]
C41 Find the vector form of the solutions to the system of equations below.
2x1 1x2 8x3+ 8x4+ 4x5 9x6 1x7 1x8 18x9= 3
Version 2.30
Subsection LC.EXC Exercises 129
3x1 2x2+ 5x3+ 2x4 2x5 5x6+ 1x7+ 2x8+ 15x9= 10
4x1 2x2+ 8x3+ 2x5 14x6 2x8+ 2x9= 36
1x1+ 2x2+ 1x3 6x4+ 7x6 1x7 3x9= 8
3x1+ 2x2+ 13x3 14x4 1x5+ 5x6 1x8+ 12x9= 15
2x1+ 2x2 2x3 4x4+ 1x5+ 6x6 2x7 2x8 15x9= 7
Contributed by Robert Beezer Solution [129]
M10 Example TLC [109] asks if the vector
w=2
666666413
15
5
17
2
253
7777775
can be written as a linear combination of the four vectors
u1=2
66666642
4
3
1
2
93
7777775u2=2
66666646
3
0
2
1
43
7777775u3=2
6666664 5
2
1
1
3
03
7777775u4=2
66666643
2
5
7
1
33
7777775
Can it? Can any vector in C6be written as a linear combination of the four vectors u1;u2;u3;u4?
Contributed by Robert Beezer Solution [130]
M11 At the end of Example VFS [115], the vector wis claimed to be a solution to the linear system
under discussion. Verify that wreally is a solution. Then determine the four scalars that express was a
linear combination of c,u1,u2,u3.
Contributed by Robert Beezer Solution [130]
Version 2.30
130 Section LC Linear Combinations
Subsection SOL
Solutions
C21 Contributed by Robert Beezer Statement [127]
Solutions for Archetype A [781] and Archetype B [786] are described carefully in Example AALC [111] and
Example ABLC [110].
C22 Contributed by Robert Beezer Statement [127]
Solutions for Archetype D [795] and Archetype I [816] are described carefully in Example VFSAD [114] and
Example VFSAI [121]. The technique described in these examples is probably more useful than carefully
deciphering the notation of Theorem VFSLS [118]. The solution for each archetype is contained in its
description. So now you can check-o the box for that item.
C40 Contributed by Robert Beezer Statement [127]
Row-reduce the augmented matrix representing this system, to nd
2
6641 2 0 6 0 1
0 0 1 4 0 3
0 0 0 0 1 5
0 0 0 0 0 03
775
The system is consistent (no leading one in column 6, Theorem RCLS [58]). x2andx4are the free variables.
Now apply Theorem VFSLS [118] directly, or follow the three-step process of Example VFS [115], Example
VFSAD [114], Example VFSAI [121], or Example VFSAL [122] to obtain
2
66664x1
x2
x3
x4
x53
77775=2
666641
0
3
0
53
77775+x22
666642
1
0
0
03
77775+x42
66664 6
0
4
1
03
77775
C41 Contributed by Robert Beezer Statement [127]
Row-reduce the augmented matrix representing this system, to nd
2
6666666410 3 2 0 1 0 0 3 6
012 4 0 3 0 0 2 1
0 0 0 0 1 2 0 0 1 3
0 0 0 0 0 0 10 4 0
0 0 0 0 0 0 0 1 2 2
0 0 0 0 0 0 0 0 0 03
77777775
The system is consistent (no leading one in column 10, Theorem RCLS [58]). F=f3;4;6;9;10g, so the
free variables are x3; x4; x6andx9. Now apply Theorem VFSLS [118] directly, or follow the three-step
process of Example VFS [115], Example VFSAD [114], Example VFSAI [121], or Example VFSAL [122]
Version 2.30
Subsection LC.SOL Solutions 131
to obtain the solution set
S=8
>>>>>>>>>>>><
>>>>>>>>>>>>:2
66666666666646
1
0
0
3
0
0
2
03
7777777777775+x32
6666666666664 3
2
1
0
0
0
0
0
03
7777777777775+x42
66666666666642
4
0
1
0
0
0
0
03
7777777777775+x62
66666666666641
3
0
0
2
1
0
0
03
7777777777775+x92
6666666666664 3
2
0
0
1
0
4
2
13
7777777777775x3; x4; x6; x92C9
>>>>>>>>>>>>=
>>>>>>>>>>>>;
M10 Contributed by Robert Beezer Statement [128]
No, it is not possible to create was a linear combination of the four vectors u1;u2;u3;u4. By creating the
desired linear combination with unknowns as scalars, Theorem SLSLC [112] provides a system of equations
that has no solution. This one computation is enough to show us that it is not possible to create all the
vectors of C6through linear combinations of the four vectors u1;u2;u3;u4.
M11 Contributed by Robert Beezer Statement [128]
The coecient of cis 1. The coecients of u1,u2,u3lie in the third, fourth and seventh entries of w.
Can you see why? (Hint: F=f3;4;7;8g, so the free variables are x3; x4andx7.)
Version 2.30
132 Section LC Linear Combinations
Version 2.30
Section SS Spanning Sets 133
Section SS
Spanning Sets
In this section we will describe a compact way to indicate the elements of an innite set of vectors, making
use of linear combinations. This will give us a convenient way to describe the elements of a set of solutions
to a linear system, or the elements of the null space of a matrix, or many other sets of vectors.
Subsection SSV
Span of a Set of Vectors
In Example VFSAL [122] we saw the solution set of a homogeneous system described as all possible linear
combinations of two particular vectors. This happens to be a useful way to construct or describe innite
sets of vectors, so we encapsulate this idea in a denition.
Denition SSCV
Span of a Set of Column Vectors
Given a set of vectors S=fu1;u2;u3; :::; upg, their span ,hSi, is the set of all possible linear combina-
tions of u1;u2;u3; :::; up. Symbolically,
hSi=f1u1+2u2+3u3++pupji2C;1ipg
=(pX
i=1iuii2C;1ip)
(This denition contains Notation SSV.) 4
The span is just a set of vectors, though in all but one situation it is an innite set. (Just when is it
not innite?) So we start with a nite collection of vectors S(pof them to be precise), and use this nite
set to describe an innite set of vectors, hSi. Confusing the nite setSwith the innite sethSiis one of
the most pervasive problems in understanding introductory linear algebra. We will see this construction
repeatedly, so let's work through some examples to get comfortable with it. The most obvious question
about a set is if a particular item of the correct type is in the set, or not.
Example ABS
A basic span
Consider the set of 5 vectors, S, from C4
S=8
>><
>>:2
6641
1
3
13
775;2
6642
1
2
13
775;2
6647
3
5
53
775;2
6641
1
1
23
775;2
664 1
0
9
03
7759
>>=
>>;
and consider the innite set of vectors hSiformed from all possible linear combinations of the elements of
S. Here are four vectors we denitely know are elements of hSi, since we will construct them in accordance
with Denition SSCV [131],
w= (2)2
6641
1
3
13
775+ (1)2
6642
1
2
13
775+ ( 1)2
6647
3
5
53
775+ (2)2
6641
1
1
23
775+ (3)2
664 1
0
9
03
775=2
664 4
2
28
103
775
Version 2.30
134 Section SS Spanning Sets
x= (5)2
6641
1
3
13
775+ ( 6)2
6642
1
2
13
775+ ( 3)2
6647
3
5
53
775+ (4)2
6641
1
1
23
775+ (2)2
664 1
0
9
03
775=2
664 26
6
2
343
775
y= (1)2
6641
1
3
13
775+ (0)2
6642
1
2
13
775+ (1)2
6647
3
5
53
775+ (0)2
6641
1
1
23
775+ (1)2
664 1
0
9
03
775=2
6647
4
17
43
775
z= (0)2
6641
1
3
13
775+ (0)2
6642
1
2
13
775+ (0)2
6647
3
5
53
775+ (0)2
6641
1
1
23
775+ (0)2
664 1
0
9
03
775=2
6640
0
0
03
775
The purpose of a set is to collect objects with some common property, and to exclude objects without that
property. So the most fundamental question about a set is if a given object is an element of the set or not.
Let's learn more about hSiby investigating which vectors are elements of the set, and which are not.
First, is u=2
664 15
6
19
53
775an element ofhSi? We are asking if there are scalars 1; 2; 3; 4; 5such that
12
6641
1
3
13
775+22
6642
1
2
13
775+32
6647
3
5
53
775+42
6641
1
1
23
775+52
664 1
0
9
03
775=u=2
664 15
6
19
53
775
Applying Theorem SLSLC [112] we recognize the search for these scalars as a solution to a linear system
of equations with augmented matrix
2
6641 2 7 1 1 15
1 1 3 1 0 6
3 2 5 1 9 19
1 1 5 2 0 53
775
which row-reduces to2
66410 1 0 3 10
01 4 0 1 9
0 0 0 1 2 7
0 0 0 0 0 03
775
At this point, we see that the system is consistent (Theorem RCLS [58]), so we know there isa solution
for the ve scalars 1; 2; 3; 4; 5. This is enough evidence for us to say that u2hSi. If we wished
further evidence, we could compute an actual solution, say
1= 2 2= 1 3= 2 4= 3 5= 2
This particular solution allows us to write
(2)2
6641
1
3
13
775+ (1)2
6642
1
2
13
775+ ( 2)2
6647
3
5
53
775+ ( 3)2
6641
1
1
23
775+ (2)2
664 1
0
9
03
775=u=2
664 15
6
19
53
775
Version 2.30
Subsection SS.SSV Span of a Set of Vectors 135
making it even more obvious that u2hSi.
Lets do it again. Is v=2
6643
1
2
13
775an element ofhSi? We are asking if there are scalars 1; 2; 3; 4; 5
such that
12
6641
1
3
13
775+22
6642
1
2
13
775+32
6647
3
5
53
775+42
6641
1
1
23
775+52
664 1
0
9
03
775=v=2
6643
1
2
13
775
Applying Theorem SLSLC [112] we recognize the search for these scalars as a solution to a linear system
of equations with augmented matrix
2
6641 2 7 1 1 3
1 1 3 1 0 1
3 2 5 1 9 2
1 1 5 2 0 13
775
which row-reduces to 2
666410 1 0 3 0
01 4 0 1 0
0 0 0 1 2 0
0 0 0 0 0 13
7775
At this point, we see that the system is inconsistent by Theorem RCLS [58], so we know there is not a
solution for the ve scalars 1; 2; 3; 4; 5. This is enough evidence for us to say that v62hSi. End of
story.
Example SCAA
Span of the columns of Archetype A
Begin with the nite set of three vectors of size 3
S=fu1;u2;u3g=8
<
:2
41
2
13
5;2
4 1
1
13
5;2
42
1
03
59
=
;
and consider the innite set hSi. The vectors of Scould have been chosen to be anything, but for reasons
that will become clear later, we have chosen the three columns of the coecient matrix in Archetype A
[781]. First, as an example, note that
v= (5)2
41
2
13
5+ ( 3)2
4 1
1
13
5+ (7)2
42
1
03
5=2
422
14
23
5
is inhSi, since it is a linear combination of u1;u2;u3. We write this succinctly as v2hSi. There is
nothing magical about the scalars 1= 5; 2= 3; 3= 7, they could have been chosen to be anything.
So repeat this part of the example yourself, using dierent values of 1; 2; 3. What happens if you
choose all three scalars to be zero?
So we know how to quickly construct sample elements of the set hSi. A slightly dierent question arises
when you are handed a vector of the correct size and asked if it is an element of hSi. For example, is
w=2
41
8
53
5inhSi? More succinctly, w2hSi?
Version 2.30
136 Section SS Spanning Sets
To answer this question, we will look for scalars 1; 2; 3so that
1u1+2u2+3u3=w
By Theorem SLSLC [112] solutions to this vector equation are solutions to the system of equations
1 2+ 23= 1
21+2+3= 8
1+2= 5
Building the augmented matrix for this linear system, and row-reducing, gives
2
410 1 3
01 1 2
0 0 0 03
5
This system has innitely many solutions (there's a free variable in x3), but all we need is one solution
vector. The solution,
1= 2 2= 3 3= 1
tells us that
(2)u1+ (3)u2+ (1)u3=w
so we are convinced that wreally is inhSi. Notice that there are an innite number of ways to answer
this question armatively. We could choose a dierent solution, this time choosing the free variable to be
zero,
1= 3 2= 2 3= 0
shows us that
(3)u1+ (2)u2+ (0)u3=w
Verifying the arithmetic in this second solution will make it obvious that wis in this span. And of course,
we now realize that there are an innite number of ways to realize was element ofhSi. Let's ask the same
type of question again, but this time with y=2
42
4
33
5, i.e. is y2hSi?
So we'll look for scalars 1; 2; 3so that
1u1+2u2+3u3=y
By Theorem SLSLC [112] solutions to this vector equation are the solutions to the system of equations
1 2+ 23= 2
21+2+3= 4
1+2= 3
Building the augmented matrix for this linear system, and row-reducing, gives
2
410 1 0
01 1 0
0 0 0 13
5
Version 2.30
Subsection SS.SSV Span of a Set of Vectors 137
This system is inconsistent (there's a leading 1 in the last column, Theorem RCLS [58]), so there are no
scalars1; 2; 3that will create a linear combination of u1;u2;u3that equals y. More precisely, y62hSi.
There are three things to observe in this example. (1) It is easy to construct vectors in hSi. (2) It is
possible that some vectors are in hSi(e.g.w), while others are not (e.g. y). (3) Deciding if a given vector
is inhSileads to solving a linear system of equations and asking if the system is consistent.
With a computer program in hand to solve systems of linear equations, could you create a program to
decide if a vector was, or wasn't, in the span of a given set of vectors? Is this art or science?
This example was built on vectors from the columns of the coecient matrix of Archetype A [781].
Study the determination that v2hSiand see if you can connect it with some of the other properties of
Archetype A [781].
Having analyzed Archetype A [781] in Example SCAA [133], we will of course subject Archetype B
[786] to a similar investigation.
Example SCAB
Span of the columns of Archetype B
Begin with the nite set of three vectors of size 3 that are the columns of the coecient matrix in Archetype
B [786],
R=fv1;v2;v3g=8
<
:2
4 7
5
13
5;2
4 6
5
03
5;2
4 12
7
43
59
=
;
and consider the innite set hRi. First, as an example, note that
x= (2)2
4 7
5
13
5+ (4)2
4 6
5
03
5+ ( 3)2
4 12
7
43
5=2
4 2
9
103
5
is inhRi, since it is a linear combination of v1;v2;v3. In other words, x2hRi. Try some dierent values
of1; 2; 3yourself, and see what vectors you can create as elements of hRi.
Now ask if a given vector is an element of hRi. For example, is z=2
4 33
24
53
5inhRi? Isz2hRi?
To answer this question, we will look for scalars 1; 2; 3so that
1v1+2v2+3v3=z
By Theorem SLSLC [112] solutions to this vector equation are the solutions to the system of equations
71 62 123= 33
51+ 52+ 73= 24
1+ 43= 5
Building the augmented matrix for this linear system, and row-reducing, gives
2
410 0 3
010 5
0 0 1 23
5
This system has a unique solution,
1= 3 2= 5 3= 2
Version 2.30
138 Section SS Spanning Sets
telling us that
( 3)v1+ (5)v2+ (2)v3=z
so we are convinced that zreally is inhRi. Notice that in this case we have only one way to answer the
question armatively since the solution is unique.
Let's ask about another vector, say is x=2
4 7
8
33
5inhRi? Isx2hRi?
We desire scalars 1; 2; 3so that
1v1+2v2+3v3=x
By Theorem SLSLC [112] solutions to this vector equation are the solutions to the system of equations
71 62 123= 7
51+ 52+ 73= 8
1+ 43= 3
Building the augmented matrix for this linear system, and row-reducing, gives
2
410 0 1
010 2
0 0 1 13
5
This system has a unique solution,
1= 1 2= 2 3= 1
telling us that
(1)v1+ (2)v2+ ( 1)v3=x
so we are convinced that xreally is inhRi. Notice that in this case we again have only one way to answer
the question armatively since the solution is again unique.
We could continue to test other vectors for membership in hRi, but there is no point. A question
about membership in hRiinevitably leads to a system of three equations in the three variables 1; 2; 3
with a coecient matrix whose columns are the vectors v1;v2;v3. This particular coecient matrix is
nonsingular, so by Theorem NMUS [86], the system is guaranteed to have a solution. (This solution is
unique, but that's not critical here.) So no matter which vector we might have chosen for z, we would have
been certain to discover that it was an element of hRi. Stated dierently, every vector of size 3 is in hRi,
orhRi=C3.
Compare this example with Example SCAA [133], and see if you can connect zwith some aspects of
the write-up for Archetype B [786].
Subsection SSNS
Spanning Sets of Null Spaces
We saw in Example VFSAL [122] that when a system of equations is homogeneous the solution set can be
expressed in the form described by Theorem VFSLS [118] where the vector cis the zero vector. We can
essentially ignore this vector, so that the remainder of the typical expression for a solution looks like an arbi-
trary linear combination, where the scalars are the free variables and the vectors are u1;u2;u3; :::; un r.
Which sounds a lot like a span. This is the substance of the next theorem.
Version 2.30
Subsection SS.SSNS Spanning Sets of Null Spaces 139
Theorem SSNS
Spanning Sets for Null Spaces
Suppose that Ais anmnmatrix, and Bis a row-equivalent matrix in reduced row-echelon form with r
nonzero rows. Let D=fd1; d2; d3; :::; drgbe the column indices where Bhas leading 1's (pivot columns)
andF=ff1; f2; f3; :::; fn rgbe the set of column indices where Bdoes not have leading 1's. Construct
then rvectors zj, 1jn rof sizenas
[zj]i=8
><
>:1 if i2F,i=fj
0 if i2F,i6=fj
[B]k;fjifi2D,i=dk
Then the null space of Ais given by
N(A) =hfz1;z2;z3; :::; zn rgi
Proof Consider the homogeneous system with Aas a coecient matrix, LS(A;0). Its set of solutions,
S, is by Denition NSM [73], the null space of A,N(A). LetB0denote the result of row-reducing the
augmented matrix of this homogeneous system. Since the system is homogeneous, the nal column of
the augmented matrix will be all zeros, and after any number of row operations (Denition RO [31]), the
column will still be all zeros. So B0has a nal column that is totally zeros.
Now apply Theorem VFSLS [118] to B0, after noting that our homogeneous system must be consistent
(Theorem HSC [71]). The vector chas zeros for each entry that corresponds to an index in F. For entries
that correspond to an index in D, the value is [B0]k;n+1, but forB0any entry in the nal column (index
n+ 1) is zero. So c=0. The vectors zj, 1jn rare identical to the vectors uj, 1jn r
described in Theorem VFSLS [118]. Putting it all together and applying Denition SSCV [131] in the nal
step,
N(A) =S
=fc+1u1+2u2+3u3++n run rj1; 2; 3; :::; n r2Cg
=f1u1+2u2+3u3++n run rj1; 2; 3; :::; n r2Cg
=hfz1;z2;z3; :::; zn rgi
Example SSNS
Spanning set of a null space
Find a set of vectors, S, so that the null space of the matrix Abelow is the span of S, that is,hSi=N(A).
A=2
6641 3 3 1 5
2 5 7 1 1
1 1 5 1 5
1 4 2 0 43
775
The null space of Ais the set of all solutions to the homogeneous system LS(A;0). If we nd the vector
form of the solutions to this homogeneous system (Theorem VFSLS [118]) then the vectors uj, 1jn r
in the linear combination are exactly the vectors zj, 1jn rdescribed in Theorem SSNS [137]. So
we can mimic Example VFSAL [122] to arrive at these vectors (rather than being a slave to the formulas
in the statement of the theorem).
Version 2.30
140 Section SS Spanning Sets
Begin by row-reducing A. The result is
2
66410 6 0 4
01 1 0 2
0 0 0 1 3
0 0 0 0 03
775
WithD=f1;2;4gandF=f3;5gwe recognize that x3andx5are free variables and we can express each
nonzero row as an expression for the dependent variables x1,x2,x4(respectively) in the free variables x3
andx5. With this we can write the vector form of a solution vector as
2
66664x1
x2
x3
x4
x53
77775=2
66664 6x3 4x5
x3+ 2x5
x3
3x5
x53
77775=x32
66664 6
1
1
0
03
77775+x52
66664 4
2
0
3
13
77775
Then in the notation of Theorem SSNS [137],
z1=2
66664 6
1
1
0
03
77775z2=2
66664 4
2
0
3
13
77775
and
N(A) =hfz1;z2gi=*8
>>>><
>>>>:2
66664 6
1
1
0
03
77775;2
66664 4
2
0
3
13
777759
>>>>=
>>>>;+
Example NSDS
Null space directly as a span
Let's express the null space of Aas the span of a set of vectors, applying Theorem SSNS [137] as econom-
ically as possible, without reference to the underlying homogeneous system of equations (in contrast to
Example SSNS [137]).
A=2
666642 1 5 1 5 1
1 1 3 1 6 1
1 1 1 0 4 3
3 2 4 4 7 0
3 1 5 2 2 33
77775
Theorem SSNS [137] creates vectors for the span by rst row-reducing the matrix in question. The row-
reduced version of Ais
B=2
66666410 2 0 1 2
011 0 3 1
0 0 0 1 4 2
0 0 0 0 0 0
0 0 0 0 0 03
777775
We will mechanically follow the prescription of Theorem SSNS [137]. Here we go, in two big steps.
Version 2.30
Subsection SS.SSNS Spanning Sets of Null Spaces 141
First, the non-pivot columns have indices F=f3;5;6g, so we will construct the n r= 6 3 = 3
vectors with a pattern of zeros and ones corresponding to the indices in F. This is the realization of the
rst two lines of the three-case denition of the vectors zj, 1jn r.
z1=2
66666641
0
03
7777775z2=2
66666640
1
03
7777775z3=2
66666640
0
13
7777775
Each of these vectors arises due to the presence of a column that is not a pivot column. The remaining
entries of each vector are the entries of the corresponding non-pivot column, negated, and distributed into
the empty slots in order (these slots have indices in the set Dand correspond to pivot columns). This is
the realization of the third line of the three-case denition of the vectors zj, 1jn r.
z1=2
6666664 2
1
1
0
0
03
7777775z2=2
66666641
3
0
4
1
03
7777775z3=2
6666664 2
1
0
2
0
13
7777775
So, by Theorem SSNS [137], we have
N(A) =hfz1;z2;z3gi=*8
>>>>>><
>>>>>>:2
6666664 2
1
1
0
0
03
7777775;2
66666641
3
0
4
1
03
7777775;2
6666664 2
1
0
2
0
13
77777759
>>>>>>=
>>>>>>;+
We know that the null space of Ais the solution set of the homogeneous system LS(A;0), but nowhere
in this application of Theorem SSNS [137] have we found occasion to reference the variables or equations
of this system. These details are all buried in the proof of Theorem SSNS [137].
More advanced computational devices will compute the null space of a matrix. See: Computation
NS.MMA [747] Here's an example that will simultaneously exercise the span construction and Theorem
SSNS [137], while also pointing the way to the next section.
Example SCAD
Span of the columns of Archetype D
Begin with the set of four vectors of size 3
T=fw1;w2;w3;w4g=8
<
:2
42
3
13
5;2
41
4
13
5;2
47
5
43
5;2
4 7
6
53
59
=
;
and consider the innite set W=hTi. The vectors of Thave been chosen as the four columns of the
coecient matrix in Archetype D [795]. Check that the vector
z2=2
6642
3
0
13
775
Version 2.30
142 Section SS Spanning Sets
is a solution to the homogeneous system LS(D;0) (it is the vector z2provided by the description of the
null space of the coecient matrix Dfrom Theorem SSNS [137]). Applying Theorem SLSLC [112], we can
write the linear combination,
2w1+ 3w2+ 0w3+ 1w4=0
which we can solve for w4,
w4= ( 2)w1+ ( 3)w2:
This equation says that whenever we encounter the vector w4, we can replace it with a specic linear
combination of the vectors w1andw2. So using w4in the setT, along with w1andw2, is excessive. An
example of what we mean here can be illustrated by the computation,
5w1+ ( 4)w2+ 6w3+ ( 3)w4= 5w1+ ( 4)w2+ 6w3+ ( 3) (( 2)w1+ ( 3)w2)
= 5w1+ ( 4)w2+ 6w3+ (6w1+ 9w2)
= 11w1+ 5w2+ 6w3
So what began as a linear combination of the vectors w1;w2;w3;w4has been reduced to a linear combi-
nation of the vectors w1;w2;w3. A careful proof using our denition of set equality (Denition SE [762])
would now allow us to conclude that this reduction is possible for any vector in W, so
W=hfw1;w2;w3gi
So the span of our set of vectors, W, has not changed, but we have described it by the span of a set of three
vectors, rather than four. Furthermore, we can achieve yet another, similar, reduction.
Check that the vector
z1=2
664 3
1
1
03
775
is a solution to the homogeneous system LS(D;0) (it is the vector z1provided by the description of the
null space of the coecient matrix Dfrom Theorem SSNS [137]). Applying Theorem SLSLC [112], we can
write the linear combination,
( 3)w1+ ( 1)w2+ 1w3=0
which we can solve for w3,
w3= 3w1+ 1w2
This equation says that whenever we encounter the vector w3, we can replace it with a specic linear
combination of the vectors w1andw2. So, as before, the vector w3is not needed in the description of W,
provided we have w1andw2available. In particular, a careful proof (such as is done in Example RSC5
[176]) would show that
W=hfw1;w2gi
SoWbegan life as the span of a set of four vectors, and we have now shown (utilizing solutions to a
homogeneous system) that Wcan also be described as the span of a set of just two vectors. Convince
yourself that we cannot go any further. In other words, it is not possible to dismiss either w1orw2in a
similar fashion and winnow the set down to just one vector.
What was it about the original set of four vectors that allowed us to declare certain vectors as surplus?
And just which vectors were we able to dismiss? And why did we have to stop once we had two vectors
remaining? The answers to these questions motivate \linear independence," our next section and next
denition, and so are worth considering carefully now.
It is possible to have your computational device crank out the vector form of the solution set to a linear
system of equations. See: Computation VFSS.MMA [747]
Version 2.30
Subsection SS.READ Reading Questions 143
Subsection READ
Reading Questions
1. Let S be the set of three vectors below.
S=8
<
:2
41
2
13
5;2
43
4
23
5;2
44
2
13
59
=
;
LetW=hSibe the span of S. Is the vector2
4 1
8
43
5inW? Give an explanation of the reason for your
answer.
2. UseSandWfrom the previous question. Is the vector2
46
5
13
5inW? Give an explanation of the
reason for your answer.
3. For the matrix Abelow, nd a set Sso thathSi=N(A), whereN(A) is the null space of A. (See
Theorem SSNS [137].)
A=2
41 3 1 9
2 1 3 8
1 1 1 53
5
Version 2.30
144 Section SS Spanning Sets
Subsection EXC
Exercises
C22 For each archetype that is a system of equations, consider the corresponding homogeneous system of
equations. Write elements of the solution set to these homogeneous systems in vector form, as guaranteed
by Theorem VFSLS [118]. Then write the null space of the coecient matrix of each system as the span
of a set of vectors, as described in Theorem SSNS [137].
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]/ Archetype E [799]
Archetype F [803]
Archetype G [808]/ Archetype H [812]
Archetype I [816]
Archetype J [820]
Contributed by Robert Beezer Solution [145]
C23 Archetype K [825] and Archetype L [829] are dened as matrices. Use Theorem SSNS [137] directly
to nd a set Sso thathSiis the null space of the matrix. Do not make any reference to the associated
homogeneous system of equations in your solution.
Contributed by Robert Beezer Solution [145]
C40 Suppose that S=8
>><
>>:2
6642
1
3
43
775;2
6643
2
2
13
7759
>>=
>>;. LetW=hSiand let x=2
6645
8
12
53
775. Isx2W? If so, provide
an explicit linear combination that demonstrates this.
Contributed by Robert Beezer Solution [145]
C41 Suppose that S=8
>><
>>:2
6642
1
3
43
775;2
6643
2
2
13
7759
>>=
>>;. LetW=hSiand let y=2
6645
1
3
53
775. Isy2W? If so, provide an
explicit linear combination that demonstrates this.
Contributed by Robert Beezer Solution [145]
C42 SupposeR=8
>>>><
>>>>:2
666642
1
3
4
03
77775;2
666641
1
2
2
13
77775;2
666643
1
0
3
23
777759
>>>>=
>>>>;. Isy=2
666641
1
8
4
33
77775inhRi?
Contributed by Robert Beezer Solution [146]
C43 SupposeR=8
>>>><
>>>>:2
666642
1
3
4
03
77775;2
666641
1
2
2
13
77775;2
666643
1
0
3
23
777759
>>>>=
>>>>;. Isz=2
666641
1
5
3
13
77775inhRi?
Contributed by Robert Beezer Solution [146]
Version 2.30
Subsection SS.EXC Exercises 145
C44 Suppose that S=8
<
:2
4 1
2
13
5;2
43
1
23
5;2
41
5
43
5;2
4 6
5
13
59
=
;. LetW=hSiand let y=2
4 5
3
03
5. Isy2W? If
so, provide an explicit linear combination that demonstrates this.
Contributed by Robert Beezer Solution [147]
C45 Suppose that S=8
<
:2
4 1
2
13
5;2
43
1
23
5;2
41
5
43
5;2
4 6
5
13
59
=
;. LetW=hSiand let w=2
42
1
33
5. Isw2W? If so,
provide an explicit linear combination that demonstrates this.
Contributed by Robert Beezer Solution [147]
C50 LetAbe the matrix below.
(a) Find a set Sso thatN(A) =hSi.
(b) If z=2
6643
5
1
23
775, then show directly that z2N(A).
(c) Write zas a linear combination of the vectors in S.
A=2
42 3 1 4
1 2 1 3
1 0 1 13
5
Contributed by Robert Beezer Solution [148]
C60 For the matrix Abelow, nd a set of vectors Sso that the span of Sequals the null space of A,
hSi=N(A).
A=2
41 1 6 8
1 2 0 1
2 1 6 73
5
Contributed by Robert Beezer Solution [149]
M10 Consider the set of all size 2 vectors in the Cartesian plane R2.
1. Give a geometric description of the span of a single vector.
2. How can you tell if two vectors span the entire plane, without doing any row reduction or calculation?
Contributed by Chris Black Solution [149]
M11 Consider the set of all size 3 vectors in Cartesian 3-space R3.
1. Give a geometric description of the span of a single vector.
2. Describe the possibilities for the span of two vectors.
3. Describe the possibilities for the span of three vectors.
Contributed by Chris Black Solution [149]
M12 Letu=2
41
3
23
5andv=2
42
2
13
5.
Version 2.30
146 Section SS Spanning Sets
1. Find a vector w1, dierent from uandv, so thathu;v;w1i=hu;vi.
2. Find a vector w2so thathu;v;w2i6=hu;vi.
Contributed by Chris Black Solution [150]
M20 In Example SCAD [139] we began with the four columns of the coecient matrix of Archetype D
[795], and used these columns in a span construction. Then we methodically argued that we could remove
the last column, then the third column, and create the same set by just doing a span construction with the
rst two columns. We claimed we could not go any further, and had removed as many vectors as possible.
Provide a convincing argument for why a third vector cannot be removed.
Contributed by Robert Beezer
M21 In the spirit of Example SCAD [139], begin with the four columns of the coecient matrix of
Archetype C [791], and use these columns in a span construction to build the set S. Argue that Scan be
expressed as the span of just three of the columns of the coecient matrix (saying exactly which three)
and in the spirit of Exercise SS.M20 [144] argue that no one of these three vectors can be removed and
still have a span construction create S.
Contributed by Robert Beezer Solution [150]
T10 Suppose that v1;v22Cm. Prove that
hfv1;v2gi=hfv1;v2;5v1+ 3v2gi
Contributed by Robert Beezer Solution [150]
T20 Suppose that Sis a set of vectors from Cm. Prove that the zero vector, 0, is an element of hSi.
Contributed by Robert Beezer Solution [151]
T21 Suppose that Sis a set of vectors from Cmandx;y2hSi. Prove that x+y2hSi.
Contributed by Robert Beezer
T22 Suppose that Sis a set of vectors from Cm,2C, and x2hSi. Prove that x2hSi.
Contributed by Robert Beezer
Version 2.30
Subsection SS.SOL Solutions 147
Subsection SOL
Solutions
C22 Contributed by Robert Beezer Statement [142]
The vector form of the solutions obtained in this manner will involve precisely the vectors described in
Theorem SSNS [137] as providing the null space of the coecient matrix of the system as a span. These
vectors occur in each archetype in a description of the null space. Studying Example VFSAL [122] may
be of some help.
C23 Contributed by Robert Beezer Statement [142]
Study Example NSDS [138] to understand the correct approach to this question. The solution for each is
listed in the Archetypes (Appendix A [777]) themselves.
C40 Contributed by Robert Beezer Statement [142]
Rephrasing the question, we want to know if there are scalars 1and2such that
12
6642
1
3
43
775+22
6643
2
2
13
775=2
6645
8
12
53
775
Theorem SLSLC [112] allows us to rephrase the question again as a quest for solutions to the system of
four equations in two unknowns with an augmented matrix given by
2
6642 3 5
1 2 8
3 2 12
4 1 53
775
This matrix row-reduces to2
66410 2
01 3
0 0 0
0 0 03
775
From the form of this matrix, we can see that 1= 2 and2= 3 is an armative answer to our question.
More convincingly,
( 2)2
6642
1
3
43
775+ (3)2
6643
2
2
13
775=2
6645
8
12
53
775
C41 Contributed by Robert Beezer Statement [142]
Rephrasing the question, we want to know if there are scalars 1and2such that
12
6642
1
3
43
775+22
6643
2
2
13
775=2
6645
1
3
53
775
Theorem SLSLC [112] allows us to rephrase the question again as a quest for solutions to the system of
Version 2.30
148 Section SS Spanning Sets
four equations in two unknowns with an augmented matrix given by
2
6642 3 5
1 2 1
3 2 3
4 1 53
775
This matrix row-reduces to2
66410 0
010
0 0 1
0 0 03
775
With a leading 1 in the last column of this matrix (Theorem RCLS [58]) we can see that the system of
equations has no solution, so there are no values for 1and2that will allow us to conclude that yis in
W. Soy62W.
C42 Contributed by Robert Beezer Statement [142]
Form a linear combination, with unknown scalars, of Rthat equals y,
a12
666642
1
3
4
03
77775+a22
666641
1
2
2
13
77775+a32
666643
1
0
3
23
77775=2
666641
1
8
4
33
77775
We want to know if there are values for the scalars that make the vector equation true since that is the
denition of membership in hRi. By Theorem SLSLC [112] any such values will also be solutions to the
linear system represented by the augmented matrix,
2
666642 1 3 1
1 1 1 1
3 2 0 8
4 2 3 4
0 1 2 33
77775
Row-reducing the matrix yields,2
66666410 0 2
010 1
0 0 1 2
0 0 0 0
0 0 0 03
777775
From this we see that the system of equations is consistent (Theorem RCLS [58]), and has a unique solution.
This solution will provide a linear combination of the vectors in Rthat equals y. Soy2R.
C43 Contributed by Robert Beezer Statement [142]
Form a linear combination, with unknown scalars, of Rthat equals z,
a12
666642
1
3
4
03
77775+a22
666641
1
2
2
13
77775+a32
666643
1
0
3
23
77775=2
666641
1
5
3
13
77775
Version 2.30
Subsection SS.SOL Solutions 149
We want to know if there are values for the scalars that make the vector equation true since that is the
denition of membership in hRi. By Theorem SLSLC [112] any such values will also be solutions to the
linear system represented by the augmented matrix,
2
666642 1 3 1
1 1 1 1
3 2 0 5
4 2 3 3
0 1 2 13
77775
Row-reducing the matrix yields,2
66666410 0 0
010 0
0 0 10
0 0 0 1
0 0 0 03
777775
With a leading 1 in the last column, the system is inconsistent (Theorem RCLS [58]), so there are no
scalarsa1; a2; a3that will create a linear combination of the vectors in Rthat equal z. Soz62R.
C44 Contributed by Robert Beezer Statement [143]
Form a linear combination, with unknown scalars, of Sthat equals y,
a12
4 1
2
13
5+a22
43
1
23
5+a32
41
5
43
5+a42
4 6
5
13
5=2
4 5
3
03
5
We want to know if there are values for the scalars that make the vector equation true since that is the
denition of membership in hSi. By Theorem SLSLC [112] any such values will also be solutions to the
linear system represented by the augmented matrix,
2
4 1 3 1 6 5
2 1 5 5 3
1 2 4 1 03
5
Row-reducing the matrix yields,2
410 2 3 2
011 1 1
0 0 0 0 03
5
From this we see that the system of equations is consistent (Theorem RCLS [58]), and has a innitely many
solutions. Any solution will provide a linear combination of the vectors in Rthat equals y. Soy2S, for
example,
( 10)2
4 1
2
13
5+ ( 2)2
43
1
23
5+ (3)2
41
5
43
5+ (2)2
4 6
5
13
5=2
4 5
3
03
5
C45 Contributed by Robert Beezer Statement [143]
Form a linear combination, with unknown scalars, of Sthat equals w,
a12
4 1
2
13
5+a22
43
1
23
5+a32
41
5
43
5+a42
4 6
5
13
5=2
42
1
33
5
Version 2.30
150 Section SS Spanning Sets
We want to know if there are values for the scalars that make the vector equation true since that is the
denition of membership in hSi. By Theorem SLSLC [112] any such values will also be solutions to the
linear system represented by the augmented matrix,
2
4 1 3 1 6 2
2 1 5 5 1
1 2 4 1 33
5
Row-reducing the matrix yields,2
410 2 3 0
011 1 0
0 0 0 0 13
5
With a leading 1 in the last column, the system is inconsistent (Theorem RCLS [58]), so there are no
scalarsa1; a2; a3; a4that will create a linear combination of the vectors in Sthat equal w. Sow62hSi.
C50 Contributed by Robert Beezer Statement [143]
(a) Theorem SSNS [137] provides formulas for a set Swith this property, but rst we must row-reduce A
ARREF !2
410 1 1
01 1 2
0 0 0 03
5
x3andx4would be the free variables in the homogeneous system LS(A;0) and Theorem SSNS [137]
provides the set S=fz1;z2gwhere
z1=2
6641
1
1
03
775z2=2
6641
2
0
13
775
(b) Simply employ the components of the vector zas the variables in the homogeneous system LS(A;0).
The three equations of this system evaluate as follows,
2(3) + 3( 5) + 1(1) + 4(2) = 0
1(3) + 2( 5) + 1(1) + 3(2) = 0
1(3) + 0( 5) + 1(1) + 1(2) = 0
Since each result is zero, zqualies for membership in N(A).
(c) By Theorem SSNS [137] we know this must be possible (that is the moral of this exercise). Find
scalars1and2so that
1z1+2z2=12
6641
1
1
03
775+22
6641
2
0
13
775=2
6643
5
1
23
775=z
Theorem SLSLC [112] allows us to convert this question into a question about a system of four equations
in two variables. The augmented matrix of this system row-reduces to
2
66410 1
012
0 0 0
0 0 03
775
Version 2.30
Subsection SS.SOL Solutions 151
A solution is 1= 1 and2= 2. (Notice too that this solution is unique!)
C60 Contributed by Robert Beezer Statement [143]
Theorem SSNS [137] says that if we nd the vector form of the solutions to the homogeneous system
LS(A;0), then the xed vectors (one per free variable) will have the desired property. Row-reduce A,
viewing it as the augmented matrix of a homogeneous system with an invisible columns of zeros as the last
column,2
410 4 5
012 3
0 0 0 03
5
Moving to the vector form of the solutions (Theorem VFSLS [118]), with free variables x3andx4, solutions
to the consistent system (it is homogeneous, Theorem HSC [71]) can be expressed as
2
664x1
x2
x3
x43
775=x32
664 4
2
1
03
775+x42
6645
3
0
13
775
Then with Sgiven by
S=8
>><
>>:2
664 4
2
1
03
775;2
6645
3
0
13
7759
>>=
>>;
Theorem SSNS [137] guarantees that
N(A) =hSi=*8
>><
>>:2
664 4
2
1
03
775;2
6645
3
0
13
7759
>>=
>>;+
M10 Contributed by Chris Black Statement [143]
1. The span of a single vector vis the set of all linear combinations of that vector. Thus, hvi=
fvj2Rg. This is the line through the origin and containing the (geometric) vector v. Thus, if
v=v1
v2
, then the span of vis the line through (0 ;0) and (v1;v2).
2. Two vectors will span the entire plane if they point in dierent directions, meaning that udoes not
lie on the line through vand vice-versa. That is, for vectors uandvinR2,hu;vi=R2ifuis not a
multiple of v.
M11 Contributed by Chris Black Statement [143]
1. The span of a single vector vis the set of all linear combinations of that vector. Thus, hvi=
fvj2Rg. This is the line through the origin and containing the (geometric) vector v. Thus, if
v=2
4v1
v2
v33
5, then the span of vis the line through (0 ;0;0) and (v1;v2;v3).
Version 2.30
152 Section SS Spanning Sets
2. If the two vectors point in the same direction, then their span is the line through them. Recall that
while two points determine a line, three points determine a plane. Two vectors will span a plane if
they point in dierent directions, meaning that udoes not lie on the line through vand vice-versa.
The plane spanned by u=2
4u1
u1
u13
5andv=2
4v1
v2
v33
5is determined by the origin and the points ( u1;u2;u3)
and (v1;v2;v3).
3. If all three vectors lie on the same line, then the span is that line. If one is a linear combination of
the other two, but they are not all on the same line, then they will lie in a plane. Otherwise, the span
of the set of three vectors will be all of 3-space.
M12 Contributed by Chris Black Statement [143]
1. If we can nd a vector w1that is a linear combination of uandv, thenhu;v;w1iwill be the
same set ashu;vi. Thus, w1can be any linear combination of uandv. One such example is
w1= 3u v=2
41
11
73
5.
2. Now we are looking for a vector w2that cannot be written as a linear combination of uandv. How
can we nd such a vector? Any vector that matches two components but not the third of any element
ofhu;viwill not be in the span (why?). One such example is w2=2
44
4
13
5(which is nearly 2 v, but
not quite).
M21 Contributed by Robert Beezer Statement [144]
If the columns of the coecient matrix from Archetype C [791] are named u1;u2;u3;u4then we can
discover the equation
( 2)u1+ ( 3)u2+u3+u4=0
by building a homogeneous system of equations and viewing a solution to the system as scalars in a linear
combination via Theorem SLSLC [112]. This particular vector equation can be rearranged to read
u4= (2)u1+ (3)u2+ ( 1)u3
This can be interpreted to mean that u4is unnecessary in hfu1;u2;u3;u4gi, so that
hfu1;u2;u3;u4gi=hfu1;u2;u3gi
If we try to repeat this process and nd a linear combination of u1;u2;u3that equals the zero vector,
we will fail. The required homogeneous system of equations (via Theorem SLSLC [112]) has only a trivial
solution, which will not provide the kind of equation we need to remove one of the three remaining vectors.
T10 Contributed by Robert Beezer Statement [144]
This is an equality of sets, so Denition SE [762] applies.
First show that X=hfv1;v2gihf v1;v2;5v1+ 3v2gi=Y.
Choose x2X. Then x=a1v1+a2v2for some scalars a1anda2. Then,
x=a1v1+a2v2=a1v1+a2v2+ 0(5v1+ 3v2)
which qualies xfor membership in Y, as it is a linear combination of v1;v2;5v1+ 3v2.
Version 2.30
Subsection SS.SOL Solutions 153
Now show the opposite inclusion, Y=hfv1;v2;5v1+ 3v2gihf v1;v2gi=X.
Choose y2Y. Then there are scalars a1; a2; a3such that
y=a1v1+a2v2+a3(5v1+ 3v2)
Rearranging, we obtain,
y=a1v1+a2v2+a3(5v1+ 3v2)
=a1v1+a2v2+ 5a3v1+ 3a3v2 Property DVAC [101]
=a1v1+ 5a3v1+a2v2+ 3a3v2 Property CC [100]
= (a1+ 5a3)v1+ (a2+ 3a3)v2 Property DSAC [101]
This is an expression for yas a linear combination of v1andv2, earning ymembership in X. SinceXis
a subset of Y, and vice versa, we see that X=Y, as desired.
T20 Contributed by Robert Beezer Statement [144]
No matter what the elements of the set Sare, we can choose the scalars in a linear combination to all be
zero. Suppose that S=fv1;v2;v3; :::; vpg. Then compute
0v1+ 0v2+ 0v3++ 0vp=0+0+0++0
=0
But what if we choose Sto be the empty set? The convention is that the empty sum in Denition SSCV
[131] evaluates to \zero," in this case this is the zero vector.
Version 2.30
154 Section SS Spanning Sets
Version 2.30
Section LI Linear Independence 155
Section LI
Linear Independence
Subsection LISV
Linearly Independent Sets of Vectors
Theorem SLSLC [112] tells us that a solution to a homogeneous system of equations is a linear combination
of the columns of the coecient matrix that equals the zero vector. We used just this situation to our
advantage (twice!) in Example SCAD [139] where we reduced the set of vectors used in a span construction
from four down to two, by declaring certain vectors as surplus. The next two denitions will allow us to
formalize this situation.
Denition RLDCV
Relation of Linear Dependence for Column Vectors
Given a set of vectors S=fu1;u2;u3; :::; ung, a true statement of the form
1u1+2u2+3u3++nun=0
is arelation of linear dependence onS. If this statement is formed in a trivial fashion, i.e. i= 0,
1in, then we say it is the trivial relation of linear dependence onS. 4
Denition LICV
Linear Independence of Column Vectors
The set of vectors S=fu1;u2;u3; :::; ungislinearly dependent if there is a relation of linear depen-
dence onSthat is not trivial. In the case where the only relation of linear dependence on Sis the trivial
one, thenSis alinearly independent set of vectors. 4
Notice that a relation of linear dependence is an equation . Though most of it is a linear combination, it
is not a linear combination (that would be a vector). Linear independence is a property of a setof vectors.
It is easy to take a set of vectors, and an equal number of scalars, all zero , and form a linear combination
that equals the zero vector. When the easy way is the only way, then we say the set is linearly independent.
Here's a couple of examples.
Example LDS
Linearly dependent set in C5
Consider the set of n= 4 vectors from C5,
S=8
>>>><
>>>>:2
666642
1
3
1
23
77775;2
666641
2
1
5
23
77775;2
666642
1
3
6
13
77775;2
66664 6
7
1
0
13
777759
>>>>=
>>>>;
To determine linear independence we rst form a relation of linear dependence,
12
666642
1
3
1
23
77775+22
666641
2
1
5
23
77775+32
666642
1
3
6
13
77775+42
66664 6
7
1
0
13
77775=0
Version 2.30
156 Section LI Linear Independence
We know that 1=2=3=4= 0 is a solution to this equation, but that is of no interest whatsoever.
That is always the case, no matter what four vectors we might have chosen. We are curious to know if there
are other, nontrivial, solutions. Theorem SLSLC [112] tells us that we can nd such solutions as solutions
to the homogeneous system LS(A;0) where the coecient matrix has these four vectors as columns,
A=2
666642 1 2 6
1 2 1 7
3 1 3 1
1 5 6 0
2 2 1 13
77775
Row-reducing this coecient matrix yields,
2
66666410 0 2
010 4
0 0 1 3
0 0 0 0
0 0 0 03
777775
We could solve this homogeneous system completely, but for this example all we need is one nontrivial
solution. Setting the lone free variable to any nonzero value, such as x4= 1, yields the nontrivial solution
x=2
6642
4
3
13
775
completing our application of Theorem SLSLC [112], we have
22
666642
1
3
1
23
77775+ ( 4)2
666641
2
1
5
23
77775+ 32
666642
1
3
6
13
77775+ 12
66664 6
7
1
0
13
77775=0
This is a relation of linear dependence on Sthat is not trivial, so we conclude that Sis linearly dependent.
Example LIS
Linearly independent set in C5
Consider the set of n= 4 vectors from C5,
T=8
>>>><
>>>>:2
666642
1
3
1
23
77775;2
666641
2
1
5
23
77775;2
666642
1
3
6
13
77775;2
66664 6
7
1
1
13
777759
>>>>=
>>>>;
To determine linear independence we rst form a relation of linear dependence,
12
666642
1
3
1
23
77775+22
666641
2
1
5
23
77775+32
666642
1
3
6
13
77775+42
66664 6
7
1
1
13
77775=0
Version 2.30
Subsection LI.LISV Linearly Independent Sets of Vectors 157
We know that 1=2=3=4= 0 is a solution to this equation, but that is of no interest whatsoever.
That is always the case, no matter what four vectors we might have chosen. We are curious to know
if there are other, nontrivial, solutions. Theorem SLSLC [112] tells us that we can nd such solutions
as solution to the homogeneous system LS(B;0) where the coecient matrix has these four vectors as
columns. Row-reducing this coecient matrix yields,
B=2
666642 1 2 6
1 2 1 7
3 1 3 1
1 5 6 1
2 2 1 13
77775RREF !2
66666410 0 0
010 0
0 0 10
0 0 0 1
0 0 0 03
777775
From the form of this matrix, we see that there are no free variables, so the solution is unique, and because
the system is homogeneous, this unique solution is the trivial solution. So we now know that there is but
one way to combine the four vectors of Tinto a relation of linear dependence, and that one way is the easy
and obvious way. In this situation we say that the set, T, is linearly independent.
Example LDS [153] and Example LIS [154] relied on solving a homogeneous system of equations to
determine linear independence. We can codify this process in a time-saving theorem.
Theorem LIVHS
Linearly Independent Vectors and Homogeneous Systems
Suppose that Ais anmnmatrix and S=fA1;A2;A3; :::; Angis the set of vectors in Cmthat are
the columns of A. ThenSis a linearly independent set if and only if the homogeneous system LS(A;0)
has a unique solution.
Proof (() Suppose thatLS(A;0) has a unique solution. Since it is a homogeneous system, this solution
must be the trivial solution x=0. By Theorem SLSLC [112], this means that the only relation of linear
dependence on Sis the trivial one. So Sis linearly independent.
()) We will prove the contrapositive. Suppose that LS(A;0) does not have a unique solution. Since it
is a homogeneous system, it is consistent (Theorem HSC [71]), and so must have innitely many solutions
(Theorem PSSLS [60]). One of these innitely many solutions must be nontrivial (in fact, almost all of
them are), so choose one. By Theorem SLSLC [112] this nontrivial solution will give a nontrivial relation
of linear dependence on S, so we can conclude that Sis a linearly dependent set.
Since Theorem LIVHS [155] is an equivalence, we can use it to determine the linear independence
or dependence of any set of column vectors, just by creating a corresponding matrix and analyzing the
row-reduced form. Let's illustrate this with two more examples.
Example LIHS
Linearly independent, homogeneous system
Is the set of vectors
S=8
>>>><
>>>>:2
666642
1
3
4
23
77775;2
666646
2
1
3
43
77775;2
666644
3
4
5
13
777759
>>>>=
>>>>;
linearly independent or linearly dependent?
Theorem LIVHS [155] suggests we study the matrix whose columns are the vectors in S,
A=2
666642 6 4
1 2 3
3 1 4
4 3 5
2 4 13
77775
Version 2.30
158 Section LI Linear Independence
Specically, we are interested in the size of the solution set for the homogeneous system LS(A;0). Row-
reducingA, we obtain2
66666410 0
010
0 0 1
0 0 0
0 0 03
777775
Now,r= 3, so there are n r= 3 3 = 0 free variables and we see that LS(A;0) has a unique solution
(Theorem HSC [71], Theorem FVCS [60]). By Theorem LIVHS [155], the set Sis linearly independent.
Example LDHS
Linearly dependent, homogeneous system
Is the set of vectors
S=8
>>>><
>>>>:2
666642
1
3
4
23
77775;2
666646
2
1
3
43
77775;2
666644
3
4
1
23
777759
>>>>=
>>>>;
linearly independent or linearly dependent?
Theorem LIVHS [155] suggests we study the matrix whose columns are the vectors in S,
A=2
666642 6 4
1 2 3
3 1 4
4 3 1
2 4 23
77775
Specically, we are interested in the size of the solution set for the homogeneous system LS(A;0). Row-
reducingA, we obtain2
6666410 1
01 1
0 0 0
0 0 0
0 0 03
77775
Now,r= 2, so there are n r= 3 2 = 1 free variables and we see that LS(A;0) has innitely
many solutions (Theorem HSC [71], Theorem FVCS [60]). By Theorem LIVHS [155], the set Sis linearly
dependent.
As an equivalence, Theorem LIVHS [155] gives us a straightforward way to determine if a set of vectors
is linearly independent or dependent.
Review Example LIHS [155] and Example LDHS [156]. They are very similar, diering only in the
last two slots of the third vector. This resulted in slightly dierent matrices when row-reduced, and
slightly dierent values of r, the number of nonzero rows. Notice, too, that we are less interested in the
actual solution set, and more interested in its form or size. These observations allow us to make a slight
improvement in Theorem LIVHS [155].
Theorem LIVRN
Linearly Independent Vectors, randn
Suppose that Ais anmnmatrix and S=fA1;A2;A3; :::; Angis the set of vectors in Cmthat are
Version 2.30
Subsection LI.LISV Linearly Independent Sets of Vectors 159
the columns of A. LetBbe a matrix in reduced row-echelon form that is row-equivalent to Aand letr
denote the number of non-zero rows in B. ThenSis linearly independent if and only if n=r.
Proof Theorem LIVHS [155] says the linear independence of Sis equivalent to the homogeneous linear
systemLS(A;0) having a unique solution. Since LS(A;0) is consistent (Theorem HSC [71]) we can apply
Theorem CSRN [59] to see that the solution is unique exactly when n=r.
So now here's an example of the most straightforward way to determine if a set of column vectors in
linearly independent or linearly dependent. While this method can be quick and easy, don't forget the
logical progression from the denition of linear independence through homogeneous system of equations
which makes it possible.
Example LDRN
Linearly dependent, r<n
Is the set of vectors
S=8
>>>>>><
>>>>>>:2
66666642
1
3
1
0
33
7777775;2
66666649
6
2
3
2
13
7777775;2
66666641
1
1
0
0
13
7777775;2
6666664 3
1
4
2
1
23
7777775;2
66666646
2
1
4
3
23
77777759
>>>>>>=
>>>>>>;
linearly independent or linearly dependent? Theorem LIVHS [155] suggests we place these vectors into a
matrix as columns and analyze the row-reduced version of the matrix,
2
66666642 9 1 3 6
1 6 1 1 2
3 2 1 4 1
1 3 0 2 4
0 2 0 1 3
3 1 1 2 23
7777775RREF !2
6666666410 0 0 1
010 0 1
0 0 10 2
0 0 0 1 1
0 0 0 0 0
0 0 0 0 03
77777775
Now we need only compute that r= 4<5 =nto recognize, via Theorem LIVHS [155] that Sis a linearly
dependent set. Boom!
Example LLDS
Large linearly dependent set in C4
Consider the set of n= 9 vectors from C4,
R=8
>><
>>:2
664 1
3
1
23
775;2
6647
1
3
63
775;2
6641
2
1
23
775;2
6640
4
2
93
775;2
6645
2
4
33
775;2
6642
1
6
43
775;2
6643
0
3
13
775;2
6641
1
5
33
775;2
664 6
1
1
13
7759
>>=
>>;:
To employ Theorem LIVHS [155], we form a 4 9 coecient matrix, C,
C=2
664 1 7 1 0 5 2 3 1 6
3 1 2 4 2 1 0 1 1
1 3 1 2 4 6 3 5 1
2 6 2 9 3 4 1 3 13
775:
To determine if the homogeneous system LS(C;0) has a unique solution or not, we would normally row-
reduce this matrix. But in this particular example, we can do better. Theorem HMVEI [73] tells us that
since the system is homogeneous with n= 9 variables in m= 4 equations, and n > m , there must be
Version 2.30
160 Section LI Linear Independence
innitely many solutions. Since there is not a unique solution, Theorem LIVHS [155] says the set is linearly
dependent.
The situation in Example LLDS [157] is slick enough to warrant formulating as a theorem.
Theorem MVSLD
More Vectors than Size implies Linear Dependence
Suppose that S=fu1;u2;u3; :::; ungis the set of vectors in Cm, and thatn>m . ThenSis a linearly
dependent set.
Proof Form themncoecient matrix Athat has the column vectors ui, 1inas its columns.
Consider the homogeneous system LS(A;0). By Theorem HMVEI [73] this system has innitely many
solutions. Since the system does not have a unique solution, Theorem LIVHS [155] says the columns of A
form a linearly dependent set, which is the desired conclusion.
Subsection LINM
Linear Independence and Nonsingular Matrices
We will now specialize to sets of nvectors from Cn. This will put Theorem MVSLD [158] o-limits, while
Theorem LIVHS [155] will involve square matrices. Let's begin by contrasting Archetype A [781] and
Archetype B [786].
Example LDCAA
Linearly dependent columns in Archetype A
Archetype A [781] is a system of linear equations with coecient matrix,
A=2
41 1 2
2 1 1
1 1 03
5
Do the columns of this matrix form a linearly independent or dependent set? By Example S [83] we
know that Ais singular. According to the denition of nonsingular matrices, Denition NM [83], the
homogeneous system LS(A;0) has innitely many solutions. So by Theorem LIVHS [155], the columns of
Aform a linearly dependent set.
Example LICAB
Linearly independent columns in Archetype B
Archetype B [786] is a system of linear equations with coecient matrix,
B=2
4 7 6 12
5 5 7
1 0 43
5
Do the columns of this matrix form a linearly independent or dependent set? By Example NM [84] we
know thatBis nonsingular. According to the denition of nonsingular matrices, Denition NM [83], the
homogeneous system LS(A;0) has a unique solution. So by Theorem LIVHS [155], the columns of Bform
a linearly independent set.
That Archetype A [781] and Archetype B [786] have opposite properties for the columns of their
coecient matrices is no accident. Here's the theorem, and then we will update our equivalences for
nonsingular matrices, Theorem NME1 [87].
Version 2.30
Subsection LI.NSSLI Null Spaces, Spans, Linear Independence 161
Theorem NMLIC
Nonsingular Matrices have Linearly Independent Columns
Suppose that Ais a square matrix. Then Ais nonsingular if and only if the columns of Aform a linearly
independent set.
Proof This is a proof where we can chain together equivalences, rather than proving the two halves
separately.
Anonsingular() LS (A;0) has a unique solution Denition NM [83]
() columns of Aare linearly independent Theorem LIVHS [155]
Here's an update to Theorem NME1 [87].
Theorem NME2
Nonsingular Matrix Equivalences, Round 2
Suppose that Ais a square matrix. The following are equivalent.
1.Ais nonsingular.
2.Arow-reduces to the identity matrix.
3. The null space of Acontains only the zero vector, N(A) =f0g.
4. The linear system LS(A;b) has a unique solution for every possible choice of b.
5. The columns of Aform a linearly independent set.
Proof Theorem NMLIC [159] is yet another equivalence for a nonsingular matrix, so we can add it to
the list in Theorem NME1 [87].
Subsection NSSLI
Null Spaces, Spans, Linear Independence
In Subsection SS.SSNS [136] we proved Theorem SSNS [137] which provided n rvectors that could be
used with the span construction to build the entire null space of a matrix. As we have hinted in Example
SCAD [139], and as we will see again going forward, linearly dependent sets carry redundant vectors
with them when used in building a set as a span. Our aim now is to show that the vectors provided by
Theorem SSNS [137] form a linearly independent set, so in one sense they are as ecient as possible a
way to describe the null space. Notice that the vectors zj, 1jn rrst appear in the vector form
of solutions to arbitrary linear systems (Theorem VFSLS [118]). The exact same vectors appear again
in the span construction in the conclusion of Theorem SSNS [137]. Since this second theorem specializes
to homogeneous systems the only real dierence is that the vector cin Theorem VFSLS [118] is the zero
vector for a homogeneous system. Finally, Theorem BNS [160] will now show that these same vectors are a
linearly independent set. We'll set the stage for the proof of this theorem with a moderately large example.
Study the example carefully, as it will make it easier to understand the proof.
Example LINSB
Linear independence of null space basis
Version 2.30
162 Section LI Linear Independence
Suppose that we are interested in the null space of the a 3 7 matrix,A, which row-reduces to
B=2
410 2 4 0 3 9
01 5 6 0 7 1
0 0 0 0 18 53
5
The setF=f3;4;6;7gis the set of indices for our four free variables that would be used in a description
of the solution set for the homogeneous system LS(A;0). Applying Theorem SSNS [137] we can begin to
construct a set of four vectors whose span is the null space of A, a set of vectors we will reference as T.
N(A) =hTi=hfz1;z2;z3;z4gi=*8
>>>>>>>><
>>>>>>>>:2
6666666641
0
0
03
777777775;2
6666666640
1
0
03
777777775;2
6666666640
0
1
03
777777775;2
6666666640
0
0
13
7777777759
>>>>>>>>=
>>>>>>>>;+
So far, we have constructed as much of these individual vectors as we can, based just on the knowledge of
the contents of the set F. This has allowed us to determine the entries in slots 3, 4, 6 and 7, while we have
left slots 1, 2 and 5 blank. Without doing any more, lets ask if Tis linearly independent? Begin with a
relation of linear dependence on T, and see what we can learn about the scalars,
0=1z1+2z2+3z3+4z42
6666666640
0
0
0
0
0
03
777777775=12
6666666641
0
0
03
777777775+22
6666666640
1
0
03
777777775+32
6666666640
0
1
03
777777775+42
6666666640
0
0
13
777777775
=2
6666666641
0
0
03
777777775+2
6666666640
2
0
03
777777775+2
6666666640
0
3
03
777777775+2
6666666640
0
0
43
777777775=2
6666666641
2
3
43
777777775
Applying Denition CVE [98] to the two ends of this chain of equalities, we see that 1=2=3=4= 0.
So the only relation of linear dependence on the set Tis a trivial one. By Denition LICV [153] the set T
is linearly independent. The important feature of this example is how the \pattern of zeros and ones" in
the four vectors led to the conclusion of linear independence.
The proof of Theorem BNS [160] is really quite straightforward, and relies on the \pattern of zeros
and ones" that arise in the vectors zi, 1in rin the entries that correspond to the free variables.
Play along with Example LINSB [159] as you study the proof. Also, take a look at Example VFSAD [114],
Example VFSAI [121] and Example VFSAL [122], especially at the conclusion of Step 2 (temporarily
ignore the construction of the constant vector, c). This proof is also a good rst example of how to prove
a conclusion that states a set is linearly independent.
Theorem BNS
Basis for Null Spaces
Suppose that Ais anmnmatrix, and Bis a row-equivalent matrix in reduced row-echelon form with r
Version 2.30
Subsection LI.NSSLI Null Spaces, Spans, Linear Independence 163
nonzero rows. Let D=fd1; d2; d3; :::; drgandF=ff1; f2; f3; :::; fn rgbe the sets of column indices
whereBdoes and does not (respectively) have leading 1's. Construct the n rvectors zj, 1jn r
of sizenas
[zj]i=8
><
>:1 if i2F,i=fj
0 if i2F,i6=fj
[B]k;fjifi2D,i=dk
Dene the set S=fz1;z2;z3; :::; zn rg. Then
1.N(A) =hSi.
2.Sis a linearly independent set.
Proof Notice rst that the vectors zj, 1jn rare exactly the same as the n rvectors dened in
Theorem SSNS [137]. Also, the hypotheses of Theorem SSNS [137] are the same as the hypotheses of the
theorem we are currently proving. So it is then simply the conclusion of Theorem SSNS [137] that tells us
thatN(A) =hSi. That was the easy half, but the second part is not much harder. What is new here is
the claim that Sis a linearly independent set.
To prove the linear independence of a set, we need to start with a relation of linear dependence and
somehow conclude that the scalars involved must all be zero , i.e. that the relation of linear dependence
only happens in the trivial fashion. So to establish the linear independence of S, we start with
1z1+2z2+3z3++n rzn r=0:
For eachj, 1jn r, consider the equality of the individual entries of the vectors on both sides of
this equality in position fj,
0 = [0]fj
= [1z1+2z2+3z3++n rzn r]fjDenition CVE [98]
= [1z1]fj+ [2z2]fj+ [3z3]fj++ [n rzn r]fjDenition CVA [98]
=1[z1]fj+2[z2]fj+3[z3]fj++
j 1[zj 1]fj+j[zj]fj+j+1[zj+1]fj++
n r[zn r]fjDenition CVSM [99]
=1(0) +2(0) +3(0) ++
j 1(0) +j(1) +j+1(0) ++n r(0) Denition of zj
=j
So for allj, 1jn r, we havej= 0, which is the conclusion that tells us that the only relation of
linear dependence on S=fz1;z2;z3; :::; zn rgis the trivial one. Hence, by Denition LICV [153] the
set is linearly independent, as desired.
Example NSLIL
Null space spanned by linearly independent set, Archetype L
In Example VFSAL [122] we previewed Theorem SSNS [137] by nding a set of two vectors such that their
span was the null space for the matrix in Archetype L [829]. Writing the matrix as L, we have
N(L) =*8
>>>><
>>>>:2
66664 1
2
2
1
03
77775;2
666642
2
1
0
13
777759
>>>>=
>>>>;+
Version 2.30
164 Section LI Linear Independence
Solving the homogeneous system LS(L;0) resulted in recognizing x4andx5as the free variables. So look
in entries 4 and 5 of the two vectors above and notice the pattern of zeros and ones that provides the linear
independence of the set.
Subsection READ
Reading Questions
1. LetSbe the set of three vectors below.
S=8
<
:2
41
2
13
5;2
43
4
23
5;2
44
2
13
59
=
;
IsSlinearly independent or linearly dependent? Explain why.
2. LetSbe the set of three vectors below.
S=8
<
:2
41
1
03
5;2
43
2
23
5;2
44
3
43
59
=
;
IsSlinearly independent or linearly dependent? Explain why.
3. Based on your answer to the previous question, is the matrix below singular or nonsingular? Explain.
2
41 3 4
1 2 3
0 2 43
5
Version 2.30
Subsection LI.EXC Exercises 165
Subsection EXC
Exercises
Determine if the sets of vectors in Exercises C20{C25 are linearly independent or linearly dependent. When
the set is linearly dependent, exhibit a nontrivial relation of linear dependence.
C208
<
:2
41
2
13
5;2
42
1
33
5;2
41
5
03
59
=
;
Contributed by Robert Beezer Solution [167]
C218
>><
>>:2
664 1
2
4
23
775;2
6643
3
1
33
775;2
6647
3
6
43
7759
>>=
>>;
Contributed by Robert Beezer Solution [167]
C228
<
:2
4 2
1
13
5;2
41
0
13
5;2
43
3
63
5;2
4 5
4
63
5;2
44
4
73
59
=
;
Contributed by Robert Beezer Solution [167]
C238
>>>><
>>>>:2
666641
2
2
5
33
77775;2
666643
3
1
2
43
77775;2
666642
1
2
1
13
77775;2
666641
0
1
2
23
777759
>>>>=
>>>>;
Contributed by Robert Beezer Solution [167]
C248
>>>><
>>>>:2
666641
2
1
0
13
77775;2
666643
2
1
2
23
77775;2
666644
4
2
2
33
77775;2
66664 1
2
1
2
03
777759
>>>>=
>>>>;
Contributed by Robert Beezer Solution [167]
C258
>>>><
>>>>:2
666642
1
3
1
23
77775;2
666644
2
1
3
23
77775;2
6666410
7
0
10
43
777759
>>>>=
>>>>;
Contributed by Robert Beezer Solution [168]
C30 For the matrix Bbelow, nd a set Sthat is linearly independent and spans the null space of B,
that is,N(B) =hSi.
B=2
4 3 1 2 7
1 2 1 4
1 1 2 13
5
Contributed by Robert Beezer Solution [168]
C31 For the matrix Abelow, nd a linearly independent set Sso that the null space of Ais spanned by
Version 2.30
166 Section LI Linear Independence
S, that is,N(A) =hSi.
A=2
664 1 2 2 1 5
1 2 1 1 5
3 6 1 2 7
2 4 0 1 23
775
Contributed by Robert Beezer Solution [168]
C32 Find a set of column vectors, T, such that (1) the span of Tis the null space of B,hTi=N(B)
and (2)Tis a linearly independent set.
B=2
42 1 1 1
4 3 1 7
1 1 1 33
5
Contributed by Robert Beezer Solution [169]
C33 Find a setSso thatSis linearly independent and N(A) =hSi, whereN(A) is the null space of the
matrixAbelow.
A=2
42 3 3 1 4
1 1 1 1 3
3 2 8 1 13
5
Contributed by Robert Beezer Solution [169]
C50 Consider each archetype that is a system of equations and consider the solutions listed for the
homogeneous version of the archetype. (If only the trivial solution is listed, then assume this is the only
solution to the system.) From the solution set, determine if the columns of the coecient matrix form
a linearly independent or linearly dependent set. In the case of a linearly dependent set, use one of the
sample solutions to provide a nontrivial relation of linear dependence on the set of columns of the coecient
matrix (Denition RLD [351]). Indicate when Theorem MVSLD [158] applies and connect this with the
number of variables and equations in the system of equations.
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]/Archetype E [799]
Archetype F [803]
Archetype G [808]/Archetype H [812]
Archetype I [816]
Archetype J [820]
Contributed by Robert Beezer
C51 For each archetype that is a system of equations consider the homogeneous version. Write elements
of the solution set in vector form (Theorem VFSLS [118]) and from this extract the vectors zjdescribed
in Theorem BNS [160]. These vectors are used in a span construction to describe the null space of the
coecient matrix for each archetype. What does it mean when we write a null space as hfgi ?
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]/Archetype E [799]
Archetype F [803]
Version 2.30
Subsection LI.EXC Exercises 167
Archetype G [808]/Archetype H [812]
Archetype I [816]
Archetype J [820]
Contributed by Robert Beezer
C52 For each archetype that is a system of equations consider the homogeneous version. Sample solutions
are given and a linearly independent spanning set is given for the null space of the coecient matrix. Write
each of the sample solutions individually as a linear combination of the vectors in the spanning set for the
null space of the coecient matrix.
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]/Archetype E [799]
Archetype F [803]
Archetype G [808]/Archetype H [812]
Archetype I [816]
Archetype J [820]
Contributed by Robert Beezer
C60 For the matrix Abelow, nd a set of vectors Sso that (1) Sis linearly independent, and (2) the
span ofSequals the null space of A,hSi=N(A). (See Exercise SS.C60 [143].)
A=2
41 1 6 8
1 2 0 1
2 1 6 73
5
Contributed by Robert Beezer Solution [170]
M20 Suppose that S=fv1;v2;v3gis a set of three vectors from C873. Prove that the set
T=f2v1+ 3v2+v3;v1 v2 2v3;2v1+v2 v3g
is linearly dependent.
Contributed by Robert Beezer Solution [170]
M21 Suppose that S=fv1;v2;v3gis a linearly independent set of three vectors from C873. Prove that
the set
T=f2v1+ 3v2+v3;v1 v2+ 2v3;2v1+v2 v3g
is linearly independent.
Contributed by Robert Beezer Solution [171]
M50 Consider the set of vectors from C3,W, given below. Find a set Tthat contains three vectors from
Wand such that W=hTi.
W=hfv1;v2;v3;v4;v5gi=*8
<
:2
42
1
13
5;2
4 1
1
13
5;2
41
2
33
5;2
43
1
33
5;2
40
1
33
59
=
;+
Contributed by Robert Beezer Solution [171]
Version 2.30
168 Section LI Linear Independence
M51 Consider the subspace W=hfv1;v2;v3;v4gi. Find a set Sso that (1) Sis a subset of W, (2)S
is linearly independent, and (3) W=hSi. Write each vector not included in Sas a linear combination of
the vectors that are in S.
v1=2
41
1
23
5 v2=2
44
4
83
5 v3=2
4 3
2
73
5 v4=2
42
1
73
5
Contributed by Manley Perkel Solution [172]
T10 Prove that if a set of vectors contains the zero vector, then the set is linearly dependent. (Ed. \The
zero vector is death to linearly independent sets.")
Contributed by Martin Jackson
T12 Suppose that Sis a linearly independent set of vectors, and Tis a subset of S,TS(Denition
SSET [761]). Prove that Tis linearly independent.
Contributed by Robert Beezer
T13 Suppose that Tis a linearly dependent set of vectors, and Tis a subset of S,TS(Denition
SSET [761]). Prove that Sis linearly dependent.
Contributed by Robert Beezer
T15 Suppose thatfv1;v2;v3; :::; vngis a set of vectors. Prove that
fv1 v2;v2 v3;v3 v4; :::; vn v1g
is a linearly dependent set.
Contributed by Robert Beezer Solution [172]
T20 Suppose thatfv1;v2;v3;v4gis a linearly independent set in C35. Prove that
fv1;v1+v2;v1+v2+v3;v1+v2+v3+v4g
is a linearly independent set.
Contributed by Robert Beezer Solution [172]
T50 Suppose that Ais anmnmatrix with linearly independent columns and the linear system
LS(A;b) is consistent. Show that this system has a unique solution. (Notice that we are not requiring A
to be square.)
Contributed by Robert Beezer Solution [173]
Version 2.30
Subsection LI.SOL Solutions 169
Subsection SOL
Solutions
C20 Contributed by Robert Beezer Statement [163]
With three vectors from C3, we can form a square matrix by making these three vectors the columns of a
matrix. We do so, and row-reduce to obtain,
2
410 0
010
0 0 13
5
the 33 identity matrix. So by Theorem NME2 [159] the original matrix is nonsingular and its columns
are therefore a linearly independent set.
C21 Contributed by Robert Beezer Statement [163]
Theorem LIVRN [156] says we can answer this question by putting theses vectors into a matrix as columns
and row-reducing. Doing this we obtain,2
66410 0
010
0 0 1
0 0 03
775
Withn= 3 (3 vectors, 3 columns) and r= 3 (3 leading 1's) we have n=rand the theorem says the
vectors are linearly independent.
C22 Contributed by Robert Beezer Statement [163]
Five vectors from C3. Theorem MVSLD [158] says the set is linearly dependent. Boom.
C23 Contributed by Robert Beezer Statement [163]
Theorem LIVRN [156] suggests we analyze a matrix whose columns are the vectors of S,
A=2
666641 3 2 1
2 3 1 0
2 1 2 1
5 2 1 2
3 4 1 23
77775
Row-reducing the matrix Ayields,2
66666410 0 0
010 0
0 0 10
0 0 0 1
0 0 0 03
777775
We see that r= 4 =n, whereris the number of nonzero rows and nis the number of columns. By
Theorem LIVRN [156], the set Sis linearly independent.
C24 Contributed by Robert Beezer Statement [163]
Theorem LIVRN [156] suggests we analyze a matrix whose columns are the vectors from the set,
A=2
666641 3 4 1
2 2 4 2
1 1 2 1
0 2 2 2
1 2 3 03
77775
Version 2.30
170 Section LI Linear Independence
Row-reducing the matrix Ayields,2
6666410 1 2
011 1
0 0 0 0
0 0 0 0
0 0 0 03
77775
We see that r= 26= 4 =n, whereris the number of nonzero rows and nis the number of columns. By
Theorem LIVRN [156], the set Sis linearly dependent.
C25 Contributed by Robert Beezer Statement [163]
Theorem LIVRN [156] suggests we analyze a matrix whose columns are the vectors from the set,
A=2
666642 4 10
1 2 7
3 1 0
1 3 10
2 2 43
77775
Row-reducing the matrix Ayields,2
6666410 1
01 3
0 0 0
0 0 0
0 0 03
77775
We see that r= 26= 3 =n, whereris the number of nonzero rows and nis the number of columns. By
Theorem LIVRN [156], the set Sis linearly dependent.
C30 Contributed by Robert Beezer Statement [163]
The requested set is described by Theorem BNS [160]. It is easiest to nd by using the procedure of Example
VFSAL [122]. Begin by row-reducing the matrix, viewing it as the coecient matrix of a homogeneous
system of equations. We obtain,2
410 1 2
011 1
0 0 0 03
5
Now build the vector form of the solutions to this homogeneous system (Theorem VFSLS [118]). The free
variables are x3andx4, corresponding to the columns without leading 1's,
2
664x1
x2
x3
x43
775=x32
664 1
1
1
03
775+x42
6642
1
0
13
775
The desired set Sis simply the constant vectors in this expression, and these are the vectors z1andz2
described by Theorem BNS [160].
S=8
>><
>>:2
664 1
1
1
03
775;2
6642
1
0
13
7759
>>=
>>;
C31 Contributed by Robert Beezer Statement [163]
Theorem BNS [160] provides formulas for n rvectors that will meet the requirements of this question.
Version 2.30
Subsection LI.SOL Solutions 171
These vectors are the same ones listed in Theorem VFSLS [118] when we solve the homogeneous system
LS(A;0), whose solution set is the null space (Denition NSM [73]).
To apply Theorem BNS [160] or Theorem VFSLS [118] we rst row-reduce the matrix, resulting in
B=2
66412 0 0 3
0 0 10 6
0 0 0 1 4
0 0 0 0 03
775
So we see that n r= 5 3 = 2 andF=f2;5g, so the vector form of a generic solution vector is
2
66664x1
x2
x3
x4
x53
77775=x22
66664 2
1
0
0
03
77775+x52
66664 3
0
6
4
13
77775
So we have
N(A) =*8
>>>><
>>>>:2
66664 2
1
0
0
03
77775;2
66664 3
0
6
4
13
777759
>>>>=
>>>>;+
C32 Contributed by Robert Beezer Statement [164]
The conclusion of Theorem BNS [160] gives us everything this question asks for. We need the reduced
row-echelon form of the matrix so we can determine the number of vectors in T, and their entries.
2
42 1 1 1
4 3 1 7
1 1 1 33
5RREF !2
410 2 2
01 3 5
0 0 0 03
5
We can build the set Tin immediately via Theorem BNS [160], but we will illustrate its construction in
two steps. Since F=f3;4g, we will have two vectors and can distribute strategically placed ones, and
many zeros. Then we distribute the negatives of the appropriate entries of the non-pivot columns of the
reduced row-echelon matrix.
T=8
>><
>>:2
6641
03
775;2
6640
13
7759
>>=
>>;T=8
>><
>>:2
664 2
3
1
03
775;2
6642
5
0
13
7759
>>=
>>;
C33 Contributed by Robert Beezer Statement [164]
A direct application of Theorem BNS [160] will provide the desired set. We require the reduced row-echelon
form ofA.
2
42 3 3 1 4
1 1 1 1 3
3 2 8 1 13
5RREF !2
410 6 0 3
01 5 0 2
0 0 0 1 43
5
The non-pivot columns have indices F=f3;5g. We build the desired set in two steps, rst placing the
requisite zeros and ones in locations based on F, then placing the negatives of the entries of columns 3 and
Version 2.30
172 Section LI Linear Independence
5 in the proper locations. This is all specied in Theorem BNS [160].
S=8
>>>><
>>>>:2
666641
03
77775;2
666640
13
777759
>>>>=
>>>>;=8
>>>><
>>>>:2
666646
5
1
0
03
77775;2
66664 3
2
0
4
13
777759
>>>>=
>>>>;
C60 Contributed by Robert Beezer Statement [165]
Theorem BNS [160] says that if we nd the vector form of the solutions to the homogeneous system
LS(A;0), then the xed vectors (one per free variable) will have the desired properties. Row-reduce A,
viewing it as the augmented matrix of a homogeneous system with an invisible columns of zeros as the last
column,2
410 4 5
012 3
0 0 0 03
5
Moving to the vector form of the solutions (Theorem VFSLS [118]), with free variables x3andx4, solutions
to the consistent system (it is homogeneous, Theorem HSC [71]) can be expressed as
2
664x1
x2
x3
x43
775=x32
664 4
2
1
03
775+x42
6645
3
0
13
775
Then with Sgiven by
S=8
>><
>>:2
664 4
2
1
03
775;2
6645
3
0
13
7759
>>=
>>;
Theorem BNS [160] guarantees the set has the desired properties.
M20 Contributed by Robert Beezer Statement [165]
By Denition LICV [153], we can complete this problem by nding scalars, 1; 2; 3, not all zero, such
that
1(2v1+ 3v2+v3) +2(v1 v2 2v3) +3(2v1+v2 v3) =0
Using various properties in Theorem VSPCV [100], we can rearrange this vector equation to
(21+2+ 23)v1+ (31 2+3)v2+ (1 22 3)v3=0
We can certainly make this vector equation true if we can determine values for the 's such that
21+2+ 23= 0
31 2+3= 0
1 22 3= 0
Aah, a homogeneous system of equations. And it has innitely many non-zero solutions. By the now
familiar techniques, one such solution is 1= 3,2= 4,3= 5, which you can check in the original
relation of linear dependence on Tabove.
Note that simply writing down the three scalars, and demonstrating that they provide a nontrivial
relation of linear dependence on T, could be considered an ironclad solution. But it wouldn't have been
Version 2.30
Subsection LI.SOL Solutions 173
very informative for you if we had only done just that here. Compare this solution very carefully with
Solution LI.M21 [171].
M21 Contributed by Robert Beezer Statement [165]
By Denition LICV [153] we can complete this problem by proving that if we assume that
1(2v1+ 3v2+v3) +2(v1 v2+ 2v3) +3(2v1+v2 v3) =0
then we must conclude that 1=2=3= 0. Using various properties in Theorem VSPCV [100], we can
rearrange this vector equation to
(21+2+ 23)v1+ (31 2+3)v2+ (1+ 22 3)v3=0
Because the set S=fv1;v2;v3gwas assumed to be linearly independent, by Denition LICV [153] we
must conclude that
21+2+ 23= 0
31 2+3= 0
1+ 22 3= 0
Aah, a homogeneous system of equations. And it has a unique solution, the trivial solution. So, 1=2=
3= 0, as desired. It is an inescapable conclusion from our assumption of a relation of linear dependence
above. Done.
Compare this solution very carefully with Solution LI.M20 [170], noting especially how this problem
required (and used) the hypothesis that the original set be linearly independent, and how this solution
feels more like a proof, while the previous problem could be solved with a fairly simple demonstration of
any nontrivial relation of linear dependence.
M50 Contributed by Robert Beezer Statement [165]
We want to rst nd some relations of linear dependence on fv1;v2;v3;v4;v5gthat will allow us to
\kick out" some vectors, in the spirit of Example SCAD [139]. To nd relations of linear dependence, we
formulate a matrix Awhose columns are v1;v2;v3;v4;v5. Then we consider the homogeneous system of
equationsLS(A;0) by row-reducing its coecient matrix (remember that if we formulated the augmented
matrix we would just add a column of zeros). After row-reducing, we obtain
2
410 0 2 1
010 1 2
0 0 10 03
5
From this we that solutions can be obtained employing the free variables x4andx5. With appropriate
choices we will be able to conclude that vectors v4andv5are unnecessary for creating Wvia a span. By
Theorem SLSLC [112] the choice of free variables below lead to solutions and linear combinations, which
are then rearranged.
x4= 1;x5= 0) ( 2)v1+ ( 1)v2+ (0)v3+ (1)v4+ (0)v5=0) v4= 2v1+v2
x4= 0;x5= 1) (1)v1+ (2)v2+ (0)v3+ (0)v4+ (1)v5=0) v5= v1 2v2
Since v4andv5can be expressed as linear combinations of v1andv2we can say that v4andv5are not
needed for the linear combinations used to build W(a claim that we could establish carefully with a pair
of set equality arguments). Thus
W=hfv1;v2;v3gi=*8
<
:2
42
1
13
5;2
4 1
1
13
5;2
41
2
33
59
=
;+
Version 2.30
174 Section LI Linear Independence
That thefv1;v2;v3gis linearly independent set can be established quickly with Theorem LIVRN [156].
There are other answers to this question, but notice that any nontrivial linear combination of v1;v2;v3;v4;v5
will have a zero coecient on v3, so this vector can never be eliminated from the set used to build the
span.
M51 Contributed by Robert Beezer Statement [166]
This problem can be solved using the approach in Solution LI.M50 [171]. We will provide a solution here
that is more ad-hoc, but note that we will have a more straight-forward procedure given by the upcoming
Theorem BS [180].
v1is a non-zero vector, so in a set all by itself we have a linearly independent set. As v2is a scalar
multiple of v1, the equation 4v1+v2=0is a relation of linear dependence on fv1;v2g, so we will pass on
v2. No such relation of linear dependence exists on fv1;v3g, though onfv1;v3;v4gwe have the relation
of linear dependence 7 v1+ 3v3+v4=0. So takeS=fv1;v3g, which is linearly independent.
Then
v2= 4v1+ 0v3 v4= 7v1 3v3
The two equations above are enough to justify the set equality
W=hfv1;v2;v3;v4gi=hfv1;v3gi=hSi
There are other solutions (for example, swap the roles of v1andv2, but by upcoming theorems we can
condently claim that any solution will be a set Swith exactly two vectors.
T15 Contributed by Robert Beezer Statement [166]
Consider the following linear combination
1 (v1 v2) +1 (v2 v3) + 1 ( v3 v4) ++ 1 (vn v1)
=v1 v2+v2 v3+v3 v4++vn v1
=v1+0+0++0 v1
=0
This is a nontrivial relation of linear dependence (Denition RLDCV [153]), so by Denition LICV [153]
the set is linearly dependent.
T20 Contributed by Robert Beezer Statement [166]
Our hypothesis and our conclusion use the term linear independence, so it will get a workout. To establish
linear independence, we begin with the denition (Denition LICV [153]) and write a relation of linear
dependence (Denition RLDCV [153]),
1(v1) +2(v1+v2) +3(v1+v2+v3) +4(v1+v2+v3+v4) =0
Using the distributive and commutative properties of vector addition and scalar multiplication (Theorem
VSPCV [100]) this equation can be rearranged as
(1+2+3+4)v1+ (2+3+4)v2+ (3+4)v3+ (4)v4=0
However, this is a relation of linear dependence (Denition RLDCV [153]) on a linearly independent set,
fv1;v2;v3;v4g(this was our lone hypothesis). By the denition of linear independence (Denition LICV
[153]) the scalars must all be zero. This is the homogeneous system of equations,
1+2+3+4= 0
2+3+4= 0
3+4= 0
Version 2.30
Subsection LI.SOL Solutions 175
4= 0
Row-reducing the coecient matrix of this system (or backsolving) gives the conclusion
1= 0 2= 0 3= 0 4= 0
This means, by Denition LICV [153], that the original set
fv1;v1+v2;v1+v2+v3;v1+v2+v3+v4g
is linearly independent.
T50 Contributed by Robert Beezer Statement [166]
LetA= [A1jA2jA3j:::jAn].LS(A;b) is consistent, so we know the system has at least one solution
(Denition CS [55]). We would like to show that there are no more than one solution to the system.
Employing Technique U [771], suppose that xandyare two solution vectors for LS(A;b). By Theorem
SLSLC [112] we know we can write,
b= [x]1A1+ [x]2A2+ [x]3A3++ [x]nAn
b= [y]1A1+ [y]2A2+ [y]3A3++ [y]nAn
Then
0=b b
= ([x]1A1+ [x]2A2++ [x]nAn) ([y]1A1+ [y]2A2++ [y]nAn)
= ([x]1 [y]1)A1+ ([x]2 [y]2)A2++ ([x]n [y]n)An
This is a relation of linear dependence (Denition RLDCV [153]) on a linearly independent set (the columns
ofA). So the scalars must all be zero,
[x]1 [y]1= 0 [ x]2 [y]2= 0 ::: [x]n [y]n= 0
Rearranging these equations yields the statement that [ x]i= [y]i, for 1in. However, this is exactly
how we dene vector equality (Denition CVE [98]), so x=yand the system has only one solution.
Version 2.30
176 Section LI Linear Independence
Version 2.30
Section LDS Linear Dependence and Spans 177
Section LDS
Linear Dependence and Spans
In any linearly dependent set there is always one vector that can be written as a linear combination of
the others. This is the substance of the upcoming Theorem DLDS [175]. Perhaps this will explain the use
of the word \dependent." In a linearly dependent set, at least one vector \depends" on the others (via a
linear combination).
Indeed, because Theorem DLDS [175] is an equivalence (Technique E [768]) some authors use this
condition as a denition (Technique D [765]) of linear dependence. Then linear independence is dened as
the logical opposite of linear dependence. Of course, we have chosen to take Denition LICV [153] as our
denition, and then follow with Theorem DLDS [175] as a theorem.
Subsection LDSS
Linearly Dependent Sets and Spans
If we use a linearly dependent set to construct a span, then we can always create the same innite set with
a starting set that is one vector smaller in size. We will illustrate this behavior in Example RSC5 [176].
However, this will not be possible if we build a span from a linearly independent set. So in a certain sense,
using a linearly independent set to formulate a span is the best possible way | there aren't any extra
vectors being used to build up all the necessary linear combinations. OK, here's the theorem, and then
the example.
Theorem DLDS
Dependency in Linearly Dependent Sets
Suppose that S=fu1;u2;u3; :::; ungis a set of vectors. Then Sis a linearly dependent set if and only if
there is an index t, 1tnsuch that utis a linear combination of the vectors u1;u2;u3; :::; ut 1;ut+1; :::; un.
Proof ()) Suppose that Sis linearly dependent, so there exists a nontrivial relation of linear dependence
by Denition LICV [153]. That is, there are scalars, i, 1in, which are not all zero, such that
1u1+2u2+3u3++nun=0:
Since theicannot all be zero, choose one, say t, that is nonzero. Then,
ut= 1
t( tut) Property MICN [759]
= 1
t(1u1++t 1ut 1+t+1ut+1++nun) Theorem VSPCV [100]
= 1
tu1++ t 1
tut 1+ t+1
tut+1++ n
tun Theorem VSPCV [100]
Since the values ofi
tare again scalars, we have expressed utas a linear combination of the other elements
ofS.
(() Assume that the vector utis a linear combination of the other vectors in S. Write this linear
combination, denoting the relevant scalars as 1,2, . . . ,t 1,t+1, . . .n, as
ut=1u1+2u2++t 1ut 1+t+1ut+1++nun
Then we have
1u1++t 1ut 1+ ( 1)ut+t+1ut+1++nun
Version 2.30
178 Section LDS Linear Dependence and Spans
=ut+ ( 1)ut Theorem VSPCV [100]
= (1 + ( 1))ut Property DSAC [101]
= 0ut Property AICN [759]
=0 Denition CVSM [99]
So the scalars 1; 2; 3; :::; t 1; t= 1;t+1; :::; nprovide a nontrivial linear combination of the
vectors inS, thus establishing that Sis a linearly dependent set (Denition LICV [153]).
This theorem can be used, sometimes repeatedly, to whittle down the size of a set of vectors used in a
span construction. We have seen some of this already in Example SCAD [139], but in the next example
we will detail some of the subtleties.
Example RSC5
Reducing a span in C5
Consider the set of n= 4 vectors from C5,
R=fv1;v2;v3;v4g=8
>>>><
>>>>:2
666641
2
1
3
23
77775;2
666642
1
3
1
23
77775;2
666640
7
6
11
23
77775;2
666644
1
2
1
63
777759
>>>>=
>>>>;
and deneV=hRi.
To employ Theorem LIVHS [155], we form a 5 4 coecient matrix, D,
D=2
666641 2 0 4
2 1 7 1
1 3 6 2
3 1 11 1
2 2 2 63
77775
and row-reduce to understand solutions to the homogeneous system LS(D;0),
2
66666410 0 4
010 0
0 0 11
0 0 0 0
0 0 0 03
777775
We can nd innitely many solutions to this system, most of them nontrivial, and we choose any one we
like to build a relation of linear dependence on R. Let's begin with x4= 1, to nd the solution
2
664 4
0
1
13
775
So we can write the relation of linear dependence,
( 4)v1+ 0v2+ ( 1)v3+ 1v4=0
Theorem DLDS [175] guarantees that we can solve this relation of linear dependence for some vector in
R, but the choice of which one is up to us. Notice however that v2has a zero coecient. In this case, we
cannot choose to solve for v2. Maybe some other relation of linear dependence would produce a nonzero
Version 2.30
Subsection LDS.COV Casting Out Vectors 179
coecient for v2if we just had to solve for this vector. Unfortunately, this example has been engineered
toalways produce a zero coecient here, as you can see from solving the homogeneous system. Every
solution has x2= 0!
OK, if we are convinced that we cannot solve for v2, let's instead solve for v3,
v3= ( 4)v1+ 0v2+ 1v4= ( 4)v1+ 1v4
We now claim that this particular equation will allow us to write
V=hRi=hfv1;v2;v3;v4gi=hfv1;v2;v4gi
in essence declaring v3as surplus for the task of building Vas a span. This claim is an equality of two
sets, so we will use Denition SE [762] to establish it carefully. Let R0=fv1;v2;v4gandV0=hR0i. We
want to show that V=V0.
First show that V0V. Since every vector of R0is inR, any vector we can construct in V0as a linear
combination of vectors from R0can also be constructed as a vector in Vby the same linear combination
of the same vectors in R. That was easy, now turn it around.
Next show that VV0. Choose any vfromV. Then there are scalars 1; 2; 3; 4so that
v=1v1+2v2+3v3+4v4
=1v1+2v2+3(( 4)v1+ 1v4) +4v4
=1v1+2v2+ (( 43)v1+3v4) +4v4
= (1 43)v1+2v2+ (3+4)v4:
This equation says that vcan then be written as a linear combination of the vectors in R0and hence
qualies for membership in V0. SoVV0and we have established that V=V0.
IfR0was also linearly dependent (it is not), we could reduce the set even further. Notice that we could
have chosen to eliminate any one of v1,v3orv4, but somehow v2is essential to the creation of Vsince it
cannot be replaced by any linear combination of v1,v3orv4.
Subsection COV
Casting Out Vectors
In Example RSC5 [176] we used four vectors to create a span. With a relation of linear dependence in
hand, we were able to \toss-out" one of these four vectors and create the same span from a subset of
just three vectors from the original set of four. We did have to take some care as to just which vector
we tossed-out. In the next example, we will be more methodical about just how we choose to eliminate
vectors from a linearly dependent set while preserving a span.
Example COV
Casting out vectors
We begin with a set Scontaining seven vectors from C4,
S=8
>><
>>:2
6641
2
0
13
775;2
6644
8
0
43
775;2
6640
1
2
23
775;2
664 1
3
3
43
775;2
6640
9
4
83
775;2
6647
13
12
313
775;2
664 9
7
8
373
7759
>>=
>>;
and deneW=hSi. The setSis obviously linearly dependent by Theorem MVSLD [158], since we have
n= 7 vectors from C4. So we can slim down Ssome, and still create Was the span of a smaller set of
Version 2.30
180 Section LDS Linear Dependence and Spans
vectors. As a device for identifying relations of linear dependence among the vectors of S, we place the
seven column vectors of Sinto a matrix as columns,
A= [A1jA2jA3j:::jA7] =2
6641 4 0 1 0 7 9
2 8 1 3 9 13 7
0 0 2 3 4 12 8
1 4 2 4 8 31 373
775
By Theorem SLSLC [112] a nontrivial solution to LS(A;0) will give us a nontrivial relation of linear
dependence (Denition RLDCV [153]) on the columns of A(which are the elements of the set S). The
row-reduced form for Ais the matrix
B=2
66414 0 0 2 1 3
0 0 10 1 3 5
0 0 0 12 6 6
0 0 0 0 0 0 03
775
so we can easily create solutions to the homogeneous system LS(A;0) using the free variables x2; x5; x6; x7.
Any such solution will correspond to a relation of linear dependence on the columns of B. These solutions
will allow us to solve for one column vector as a linear combination of some others, in the spirit of Theorem
DLDS [175], and remove that vector from the set. We'll set about forming these linear combinations
methodically. Set the free variable x2to one, and set the other free variables to zero. Then a solution to
LS(A;0) is
x=2
666666664 4
1
0
0
0
0
03
777777775
which can be used to create the linear combination
( 4)A1+ 1A2+ 0A3+ 0A4+ 0A5+ 0A6+ 0A7=0
This can then be arranged and solved for A2, resulting in A2expressed as a linear combination of
fA1;A3;A4g,
A2= 4A1+ 0A3+ 0A4
This means that A2is surplus, and we can create Wjust as well with a smaller set with this vector
removed,
W=hfA1;A3;A4;A5;A6;A7gi
Technically, this set equality for Wrequires a proof, in the spirit of Example RSC5 [176], but we will
bypass this requirement here, and in the next few paragraphs.
Now, set the free variable x5to one, and set the other free variables to zero. Then a solution to
LS(B;0) is
x=2
666666664 2
0
1
2
1
0
03
777777775
Version 2.30
Subsection LDS.COV Casting Out Vectors 181
which can be used to create the linear combination
( 2)A1+ 0A2+ ( 1)A3+ ( 2)A4+ 1A5+ 0A6+ 0A7=0
This can then be arranged and solved for A5, resulting in A5expressed as a linear combination of
fA1;A3;A4g,
A5= 2A1+ 1A3+ 2A4
This means that A5is surplus, and we can create Wjust as well with a smaller set with this vector
removed,
W=hfA1;A3;A4;A6;A7gi
Do it again, set the free variable x6to one, and set the other free variables to zero. Then a solution to
LS(B;0) is
x=2
666666664 1
0
3
6
0
1
03
777777775
which can be used to create the linear combination
( 1)A1+ 0A2+ 3A3+ 6A4+ 0A5+ 1A6+ 0A7=0
This can then be arranged and solved for A6, resulting in A6expressed as a linear combination of
fA1;A3;A4g,
A6= 1A1+ ( 3)A3+ ( 6)A4
This means that A6is surplus, and we can create Wjust as well with a smaller set with this vector
removed,
W=hfA1;A3;A4;A7gi
Set the free variable x7to one, and set the other free variables to zero. Then a solution to LS(B;0) is
x=2
6666666643
0
5
6
0
0
13
777777775
which can be used to create the linear combination
3A1+ 0A2+ ( 5)A3+ ( 6)A4+ 0A5+ 0A6+ 1A7=0
This can then be arranged and solved for A7, resulting in A7expressed as a linear combination of
fA1;A3;A4g,
A7= ( 3)A1+ 5A3+ 6A4
This means that A7is surplus, and we can create Wjust as well with a smaller set with this vector
removed,
W=hfA1;A3;A4gi
Version 2.30
182 Section LDS Linear Dependence and Spans
You might think we could keep this up, but we have run out of free variables. And not coincidentally,
the setfA1;A3;A4gis linearly independent (check this!). It should be clear how each free variable was
used to eliminate the corresponding column from the set used to span the column space, as this will be the
essence of the proof of the next theorem. The column vectors in Swere not chosen entirely at random, they
are the columns of Archetype I [816]. See if you can mimic this example using the columns of Archetype
J [820]. Go ahead, we'll go grab a cup of coee and be back before you nish up.
For extra credit, notice that the vector
b=2
6643
9
1
43
775
is the vector of constants in the denition of Archetype I [816]. Since the system LS(A;b) is consistent, we
know by Theorem SLSLC [112] that bis a linear combination of the columns of A, or stated equivalently,
b2W. This means that bmust also be a linear combination of just the three columns A1;A3;A4. Can
you nd such a linear combination? Did you notice that there is just a single (unique) answer? Hmmmm.
Example COV [177] deserves your careful attention, since this important example motivates the fol-
lowing very fundamental theorem.
Theorem BS
Basis of a Span
Suppose that S=fv1;v2;v3; :::; vngis a set of column vectors. Dene W=hSiand letAbe the
matrix whose columns are the vectors from S. LetBbe the reduced row-echelon form of A, withD=
fd1; d2; d3; :::; drgthe set of column indices corresponding to the pivot columns of B. Then
1.T=fvd1;vd2;vd3; :::vdrgis a linearly independent set.
2.W=hTi.
Proof To prove that Tis linearly independent, begin with a relation of linear dependence on T,
0=1vd1+2vd2+3vd3+:::+rvdr
and we will try to conclude that the only possibility for the scalars iis that they are all zero. Denote the
non-pivot columns of BbyF=ff1; f2; f3; :::; fn rg. Then we can preserve the equality by adding a big
fat zero to the linear combination,
0=1vd1+2vd2+3vd3+:::+rvdr+ 0vf1+ 0vf2+ 0vf3+:::+ 0vfn r
By Theorem SLSLC [112], the scalars in this linear combination (suitably reordered) are a solution to the
homogeneous system LS(A;0). But notice that this is the solution obtained by setting each free variable
to zero. If we consider the description of a solution vector in the conclusion of Theorem VFSLS [118], in
the case of a homogeneous system, then we see that if all the free variables are set to zero the resulting
solution vector is trivial (all zeros). So it must be that i= 0, 1ir. This implies by Denition LICV
[153] thatTis a linearly independent set.
The second conclusion of this theorem is an equality of sets (Denition SE [762]). Since Tis a subset of
S, any linear combination of elements of the set Tcan also be viewed as a linear combination of elements
of the setS. SohTihSi=W. It remains to prove that W=hSihTi.
For eachk, 1kn r, form a solution xtoLS(A;0) by setting the free variables as follows:
xf1= 0 xf2= 0 xf3= 0 ::: x fk= 1 ::: x fn r= 0
Version 2.30
Subsection LDS.COV Casting Out Vectors 183
By Theorem VFSLS [118], the remainder of this solution vector is given by,
xd1= [B]1;fkxd2= [B]2;fkxd3= [B]3;fk::: x dr= [B]r;fk
From this solution, we obtain a relation of linear dependence on the columns of A,
[B]1;fkvd1 [B]2;fkvd2 [B]3;fkvd3 ::: [B]r;fkvdr+ 1vfk=0
which can be arranged as the equality
vfk= [B]1;fkvd1+ [B]2;fkvd2+ [B]3;fkvd3+:::+ [B]r;fkvdr
Now, suppose we take an arbitrary element, w, ofW=hSiand write it as a linear combination of the
elements of S, but with the terms organized according to the indices in DandF,
w=1vd1+2vd2+3vd3+:::+rvdr+1vf1+2vf2+3vf3+:::+n rvfn r
From the above, we can replace each vfjby a linear combination of the vdi,
w=1vd1+2vd2+3vd3+:::+rvdr+
1
[B]1;f1vd1+ [B]2;f1vd2+ [B]3;f1vd3+:::+ [B]r;f1vdr
+
2
[B]1;f2vd1+ [B]2;f2vd2+ [B]3;f2vd3+:::+ [B]r;f2vdr
+
3
[B]1;f3vd1+ [B]2;f3vd2+ [B]3;f3vd3+:::+ [B]r;f3vdr
+
...
n r
[B]1;fn rvd1+ [B]2;fn rvd2+ [B]3;fn rvd3+:::+ [B]r;fn rvdr
With repeated applications of several of the properties of Theorem VSPCV [100] we can rearrange this
expression as,
=
1+1[B]1;f1+2[B]1;f2+3[B]1;f3+:::+n r[B]1;fn r
vd1+
2+1[B]2;f1+2[B]2;f2+3[B]2;f3+:::+n r[B]2;fn r
vd2+
3+1[B]3;f1+2[B]3;f2+3[B]3;f3+:::+n r[B]3;fn r
vd3+
...
r+1[B]r;f1+2[B]r;f2+3[B]r;f3+:::+n r[B]r;fn r
vdr
This mess expresses the vector was a linear combination of the vectors in
T=fvd1;vd2;vd3; :::vdrg
thus saying that w2hTi. Therefore, W=hSihTi.
In Example COV [177], we tossed-out vectors one at a time. But in each instance, we rewrote the
oending vector as a linear combination of those vectors that corresponded to the pivot columns of the
reduced row-echelon form of the matrix of columns. In the proof of Theorem BS [180], we accomplish this
reduction in one big step. In Example COV [177] we arrived at a linearly independent set at exactly the
same moment that we ran out of free variables to exploit. This was not a coincidence, it is the substance
of our conclusion of linear independence in Theorem BS [180].
Version 2.30
184 Section LDS Linear Dependence and Spans
Here's a straightforward application of Theorem BS [180].
Example RSC4
Reducing a span in C4
Begin with a set of ve vectors from C4,
S=8
>><
>>:2
6641
1
2
13
775;2
6642
2
4
23
775;2
6642
0
1
13
775;2
6647
1
1
43
775;2
6640
2
5
13
7759
>>=
>>;
and letW=hSi. To arrive at a (smaller) linearly independent set, follow the procedure described in
Theorem BS [180]. Place the vectors from Sinto a matrix as columns, and row-reduce,
2
6641 2 2 7 0
1 2 0 1 2
2 4 1 1 5
1 2 1 4 13
775RREF !2
66412 0 1 2
0 0 13 1
0 0 0 0 0
0 0 0 0 03
775
Columns 1 and 3 are the pivot columns ( D=f1;3g) so the set
T=8
>><
>>:2
6641
1
2
13
775;2
6642
0
1
13
7759
>>=
>>;
is linearly independent and hTi=hSi=W. Boom!
Since the reduced row-echelon form of a matrix is unique (Theorem RREFU [35]), the procedure of
Theorem BS [180] leads us to a unique set T. However, there is a wide variety of possibilities for sets T
that are linearly independent and which can be employed in a span to create W. Without proof, we list
two other possibilities:
T0=8
>><
>>:2
6642
2
4
23
775;2
6642
0
1
13
7759
>>=
>>;
T=8
>><
>>:2
6643
1
1
23
775;2
664 1
1
3
03
7759
>>=
>>;
Can you prove that T0andTare linearly independent sets and W=hSi=hT0i=hTi?
Example RES
Reworking elements of a span
Begin with a set of ve vectors from C4,
R=8
>><
>>:2
6642
1
3
23
775;2
664 1
1
0
13
775;2
664 8
1
9
43
775;2
6643
1
1
23
775;2
664 10
1
1
43
7759
>>=
>>;
Version 2.30
Subsection LDS.COV Casting Out Vectors 185
It is easy to create elements of X=hRi| we will create one at random,
y= 62
6642
1
3
23
775+ ( 7)2
664 1
1
0
13
775+ 12
664 8
1
9
43
775+ 62
6643
1
1
23
775+ 22
664 10
1
1
43
775=2
6649
2
1
33
775
We know we can replace Rby a smaller set (since it is obviously linearly dependent by Theorem MVSLD
[158]) that will create the same span. Here goes,
2
6642 1 8 3 10
1 1 1 1 1
3 0 9 1 1
2 1 4 2 43
775RREF !2
66410 3 0 1
01 2 0 2
0 0 0 1 2
0 0 0 0 03
775
So, if we collect the rst, second and fourth vectors from R,
P=8
>><
>>:2
6642
1
3
23
775;2
664 1
1
0
13
775;2
6643
1
1
23
7759
>>=
>>;
thenPis linearly independent and hPi=hRi=Xby Theorem BS [180]. Since we built yas an element
ofhRiit must also be an element of hPi. Can we write yas a linear combination of just the three vectors
inP? The answer is, of course, yes. But let's compute an explicit linear combination just for fun. By
Theorem SLSLC [112] we can get such a linear combination by solving a system of equations with the
column vectors of Ras the columns of a coecient matrix, and yas the vector of constants. Employing
an augmented matrix to solve this system,
2
6642 1 3 9
1 1 1 2
3 0 1 1
2 1 2 33
775RREF !2
66410 0 1
010 1
0 0 1 2
0 0 0 03
775
So we see, as expected, that
12
6642
1
3
23
775+ ( 1)2
664 1
1
0
13
775+ 22
6643
1
1
23
775=2
6649
2
1
33
775=y
A key feature of this example is that the linear combination that expresses yas a linear combination of the
vectors inPis unique. This is a consequence of the linear independence of P. The linearly independent
setPis smaller than R, but still just (barely) big enough to create elements of the set X=hRi. There
are many, many ways to write yas a linear combination of the ve vectors in R(the appropriate system
of equations to verify this claim has two free variables in the description of the solution set), yet there is
precisely one way to write yas a linear combination of the three vectors in P.
Version 2.30
186 Section LDS Linear Dependence and Spans
Subsection READ
Reading Questions
1. LetSbe the linearly dependent set of three vectors below.
S=8
>><
>>:2
6641
10
100
10003
775;2
6641
1
1
13
775;2
6645
23
203
20033
7759
>>=
>>;
Write one vector from Sas a linear combination of the other two and include this vector equality
in your response. (You should be able to do this on sight, rather than doing some computations.)
Convert this expression into a nontrivial relation of linear dependence on S.
2. Explain why the word \dependent" is used in the denition of linear dependence.
3. Suppose that Y=hPi=hQi, wherePis a linearly dependent set and Qis linearly independent.
Would you rather use PorQto describe Y? Why?
Version 2.30
Subsection LDS.EXC Exercises 187
Subsection EXC
Exercises
C20 LetTbe the set of columns of the matrix Bbelow. Dene W=hTi. Find a set Rso that (1) R
has 3 vectors, (2) Ris a subset of T, and (3)W=hRi.
B=2
4 3 1 2 7
1 2 1 4
1 1 2 13
5
Contributed by Robert Beezer Solution [187]
C40 Verify that the set R0=fv1;v2;v4gat the end of Example RSC5 [176] is linearly independent.
Contributed by Robert Beezer
C50 Consider the set of vectors from C3,W, given below. Find a linearly independent set Tthat contains
three vectors from Wand such thathWi=hTi.
W=fv1;v2;v3;v4;v5g=8
<
:2
42
1
13
5;2
4 1
1
13
5;2
41
2
33
5;2
43
1
33
5;2
40
1
33
59
=
;
Contributed by Robert Beezer Solution [187]
C51 Given the set Sbelow, nd a linearly independent set Tso thathTi=hSi.
S=8
<
:2
42
1
23
5;2
43
0
13
5;2
41
1
13
5;2
45
1
33
59
=
;
Contributed by Robert Beezer Solution [187]
C52 LetWbe the span of the set of vectors Sbelow,W=hSi. Find a set Tso that 1) the span of T
isW,hTi=W, (2)Tis a linearly independent set, and (3) Tis a subset of S.
S=8
<
:2
41
2
13
5;2
42
3
13
5;2
44
1
13
5;2
43
1
13
5;2
43
1
03
59
=
;
Contributed by Robert Beezer Solution [187]
C55 LetTbe the set of vectors T=8
<
:2
41
1
23
5;2
43
0
13
5;2
44
2
33
5;2
43
0
63
59
=
;. Find two dierent subsets of T, named
RandS, so thatRandSeach contain three vectors, and so that hRi=hTiandhSi=hTi. Prove that
bothRandSare linearly independent.
Contributed by Robert Beezer Solution [188]
C70 Reprise Example RES [182] by creating a new version of the vector y. In other words, form a new,
dierent linear combination of the vectors in Rto create a new vector y(but do not simplify the problem
too much by choosing any of the ve new scalars to be zero). Then express this new yas a combination
of the vectors in P.
Contributed by Robert Beezer
Version 2.30
188 Section LDS Linear Dependence and Spans
M10 At the conclusion of Example RSC4 [182] two alternative solutions, sets T0andT, are proposed.
Verify these claims by proving that hTi=hT0iandhTi=hTi.
Contributed by Robert Beezer
T40 Suppose that v1andv2are any two vectors from Cm. Prove the following set equality.
hfv1;v2gi=hfv1+v2;v1 v2gi
Contributed by Robert Beezer Solution [189]
Version 2.30
Subsection LDS.SOL Solutions 189
Subsection SOL
Solutions
C20 Contributed by Robert Beezer Statement [185]
LetT=fw1;w2;w3;w4g. The vector2
6642
1
0
13
775is a solution to the homogeneous system with the matrix
Bas the coecient matrix (check this!). By Theorem SLSLC [112] it provides the scalars for a linear
combination of the columns of B(the vectors in T) that equals the zero vector, a relation of linear
dependence on T,
2w1+ ( 1)w2+ (1)w4=0
We can rearrange this equation by solving for w4,
w4= ( 2)w1+w2
This equation tells us that the vector w4is super
uous in the span construction that creates W. So
W=hfw1;w2;w3gi. The requested set is R=fw1;w2;w3g.
C50 Contributed by Robert Beezer Statement [185]
To apply Theorem BS [180], we formulate a matrix Awhose columns are v1;v2;v3;v4;v5. Then we
row-reduce A. After row-reducing, we obtain
2
410 0 2 1
010 1 2
0 0 10 03
5
From this we that the pivot columns are D=f1;2;3g. Thus
T=fv1;v2;v3g=8
<
:2
42
1
13
5;2
4 1
1
13
5;2
41
2
33
59
=
;
is a linearly independent set and hTi=W. Compare this problem with Exercise LI.M50 [165].
C51 Contributed by Robert Beezer Statement [185]
Theorem BS [180] says we can make a matrix with these four vectors as columns, row-reduce, and just
keep the columns with indices in the set D. Here we go, forming the relevant matrix and row-reducing,
2
42 3 1 5
1 0 1 1
2 1 1 33
5RREF !2
410 1 1
01 1 1
0 0 0 03
5
Analyzing the row-reduced version of this matrix, we see that the rst two columns are pivot columns, so
D=f1;2g. Theorem BS [180] says we need only \keep" the rst two columns to create a set with the
requisite properties,
T=8
<
:2
42
1
23
5;2
43
0
13
59
=
;
C52 Contributed by Robert Beezer Statement [185]
Version 2.30
190 Section LDS Linear Dependence and Spans
This is a straight setup for the conclusion of Theorem BS [180]. The hypotheses of this theorem tell us
to pack the vectors of Winto the columns of a matrix and row-reduce,
2
41 2 4 3 3
2 3 1 1 1
1 1 1 1 03
5RREF !2
410 2 0 1
011 0 1
0 0 0 103
5
Pivot columns have indices D=f1;2;4g. Theorem BS [180] tells us to form Twith columns 1 ;2 and 4
ofS,
S=8
<
:2
41
2
13
5;2
42
3
13
5;2
43
1
13
59
=
;
C55 Contributed by Robert Beezer Statement [185]
LetAbe the matrix whose columns are the vectors in T. Then row-reduce A,
ARREF !B=2
410 0 2
010 1
0 0 1 13
5
From Theorem BS [180] we can form Rby choosing the columns of Athat correspond to the pivot columns
ofB. Theorem BS [180] also guarantees that Rwill be linearly independent.
R=8
<
:2
41
1
23
5;2
43
0
13
5;2
44
2
33
59
=
;
That was easy. To nd Swill require a bit more work. From Bwe can obtain a solution to LS(A;0),
which by Theorem SLSLC [112] will provide a nontrivial relation of linear dependence on the columns of
A, which are the vectors in T. To wit, choose the free variable x4to be 1, then x1= 2,x2= 1,x3= 1,
and so
( 2)2
41
1
23
5+ (1)2
43
0
13
5+ ( 1)2
44
2
33
5+ (1)2
43
0
63
5=2
40
0
03
5
this equation can be rewritten with the second vector staying put, and the other three moving to the other
side of the equality,2
43
0
13
5= (2)2
41
1
23
5+ (1)2
44
2
33
5+ ( 1)2
43
0
63
5
We could have chosen other vectors to stay put, but may have then needed to divide by a nonzero scalar.
This equation is enough to conclude that the second vector in Tis \surplus" and can be replaced (see the
careful argument in Example RSC5 [176]). So set
S=8
<
:2
41
1
23
5;2
44
2
33
5;2
43
0
63
59
=
;
and thenhSi=hTi.Tis also a linearly independent set, which we can show directly. Make a matrix
Cwhose columns are the vectors in S. Row-reduce Band you will obtain the identity matrix I3. By
Theorem LIVRN [156], the set Sis linearly independent.
Version 2.30
Subsection LDS.SOL Solutions 191
T40 Contributed by Robert Beezer Statement [186]
This is an equality of sets, so Denition SE [762] applies.
The \easy" half rst. Show that X=hfv1+v2;v1 v2gihf v1;v2gi=Y.
Choose x2X. Then x=a1(v1+v2) +a2(v1 v2) for some scalars a1anda2. Then,
x=a1(v1+v2) +a2(v1 v2)
=a1v1+a1v2+a2v1+ ( a2)v2
= (a1+a2)v1+ (a1 a2)v2
which qualies xfor membership in Y, as it is a linear combination of v1;v2.
Now show the opposite inclusion, Y=hfv1;v2gihf v1+v2;v1 v2gi=X.
Choose y2Y. Then there are scalars b1; b2such that y=b1v1+b2v2. Rearranging, we obtain,
y=b1v1+b2v2
=b1
2[(v1+v2) + (v1 v2)] +b2
2[(v1+v2) (v1 v2)]
=b1+b2
2(v1+v2) +b1 b2
2(v1 v2)
This is an expression for yas a linear combination of v1+v2andv1 v2, earning ymembership in X.
SinceXis a subset of Y, and vice versa, we see that X=Y, as desired.
Version 2.30
192 Section LDS Linear Dependence and Spans
Version 2.30
Section O Orthogonality 193
Section O
Orthogonality
In this section we dene a couple more operations with vectors, and prove a few theorems. At rst blush
these denitions and results will not appear central to what follows, but we will make use of them at key
points in the remainder of the course (such as Section MINM [259], Section OD [675]). Because we have
chosen to use Cas our set of scalars, this subsection is a bit more, uh, . . . complex than it would be for the
real numbers. We'll explain as we go along how things get easier for the real numbers R. If you haven't
already, now would be a good time to review some of the basic properties of arithmetic with complex
numbers described in Section CNO [757]. With that done, we can extend the basics of complex number
arithmetic to our study of vectors in Cm.
Subsection CAV
Complex Arithmetic and Vectors
We know how the addition and multiplication of complex numbers is employed in dening the operations
for vectors in Cm(Denition CVA [98] and Denition CVSM [99]). We can also extend the idea of the
conjugate to vectors.
Denition CCCV
Complex Conjugate of a Column Vector
Suppose that uis a vector from Cm. Then the conjugate of the vector, u, is dened by
[u]i=[u]i 1im
(This denition contains Notation CCCV.) 4
With this denition we can show that the conjugate of a column vector behaves as we would expect
with regard to vector addition and scalar multiplication.
Theorem CRVA
Conjugation Respects Vector Addition
Suppose xandyare two vectors from Cm. Then
x+y=x+y
Proof For each 1im,
[x+y]i=[x+y]i Denition CCCV [191]
=[x]i+ [y]i Denition CVA [98]
=[x]i+[y]i Theorem CCRA [759]
= [x]i+ [y]i Denition CCCV [191]
= [x+y]i Denition CVA [98]
Then by Denition CVE [98] we have x+y=x+y.
Theorem CRSM
Conjugation Respects Vector Scalar Multiplication
Version 2.30
194 Section O Orthogonality
Suppose xis a vector from Cm, and2Cis a scalar. Then
x=x
Proof For 1im,
[x]i=[x]i Denition CCCV [191]
=[x]i Denition CVSM [99]
=[x]i Theorem CCRM [760]
=[x]i Denition CCCV [191]
= [x]i Denition CVSM [99]
Then by Denition CVE [98] we have x=x.
These two theorems together tell us how we can \push" complex conjugation through linear combina-
tions.
Subsection IP
Inner products
Denition IP
Inner Product
Given the vectors u;v2Cmtheinner product ofuandvis the scalar quantity in C,
hu;vi= [u]1[v]1+ [u]2[v]2+ [u]3[v]3++ [u]m[v]m=mX
i=1[u]i[v]i
(This denition contains Notation IP.) 4
This operation is a bit dierent in that we begin with two vectors but produce a scalar. Computing
one is straightforward.
Example CSIP
Computing some inner products
The scalar product of
u=2
42 + 3i
5 + 2i
3 +i3
5 and v=2
41 + 2i
4 + 5i
0 + 5i3
5
is
hu;vi= (2 + 3i)(1 + 2i) + (5 + 2i)( 4 + 5i) + ( 3 +i)(0 + 5i)
= (2 + 3i)(1 2i) + (5 + 2i)( 4 5i) + ( 3 +i)(0 5i)
= (8 i) + ( 10 33i) + (5 + 15i)
= 3 19i
Version 2.30
Subsection O.IP Inner products 195
The scalar product of
w=2
666642
4
3
2
83
77775and x=2
666643
1
0
1
23
77775
is
hw;xi= 2(3) + 4( 1) + ( 3)(0) + 2( 1) + 8( 2) = 2(3) + 4(1) + ( 3)0 + 2( 1) + 8( 2) = 8:
In the case where the entries of our vectors are all real numbers (as in the second part of Example CSIP
[192]), the computation of the inner product may look familiar and be known to you as a dot product or
scalar product . So you can view the inner product as a generalization of the scalar product to vectors
fromCm(rather than Rm).
Also, note that we have chosen to conjugate the entries of the second vector listed in the inner product,
while many authors choose to conjugate entries from the rst component. It really makes no dierence
which choice is made, it just requires that subsequent denitions and theorems are consistent with the
choice. You can study the conclusion of Theorem IPAC [194] as an explanation of the magnitude of the
dierence that results from this choice. But be careful as you read other treatments of the inner product
or its use in applications, and be sure you know ahead of time which choice has been made.
There are several quick theorems we can now prove, and they will each be useful later.
Theorem IPVA
Inner Product and Vector Addition
Suppose u;v;w2Cm. Then
1.hu+v;wi=hu;wi+hv;wi
2.hu;v+wi=hu;vi+hu;wi
Proof The proofs of the two parts are very similar, with the second one requiring just a bit more eort
due to the conjugation that occurs. We will prove part 2 and you can prove part 1 (Exercise O.T10 [203]).
hu;v+wi=mX
i=1[u]i[v+w]i Denition IP [192]
=mX
i=1[u]i([v]i+ [w]i) Denition CVA [98]
=mX
i=1[u]i([v]i+[w]i) Theorem CCRA [759]
=mX
i=1[u]i[v]i+ [u]i[w]i Property DCN [759]
=mX
i=1[u]i[v]i+mX
i=1[u]i[w]i Property CACN [758]
=hu;vi+hu;wi Denition IP [192]
Version 2.30
196 Section O Orthogonality
Theorem IPSM
Inner Product and Scalar Multiplication
Suppose u;v2Cmand2C. Then
1.hu;vi=hu;vi
2.hu; vi=hu;vi
Proof The proofs of the two parts are very similar, with the second one requiring just a bit more eort
due to the conjugation that occurs. We will prove part 2 and you can prove part 1 (Exercise O.T11 [203]).
hu; vi=mX
i=1[u]i[v]i Denition IP [192]
=mX
i=1[u]i[v]i Denition CVSM [99]
=mX
i=1[u]i[v]i Theorem CCRM [760]
=mX
i=1[u]i[v]i Property CMCN [758]
=mX
i=1[u]i[v]i Property DCN [759]
=hu;vi Denition IP [192]
Theorem IPAC
Inner Product is Anti-Commutative
Suppose that uandvare vectors in Cm. Thenhu;vi=hv;ui.
Proof
hu;vi=mX
i=1[u]i[v]i Denition IP [192]
=mX
i=1[u]i[v]i Theorem CCT [760]
=mX
i=1[u]i[v]i Theorem CCRM [760]
= mX
i=1[u]i[v]i!
Theorem CCRA [759]
= mX
i=1[v]i[u]i!
Property CMCN [758]
=hv;ui Denition IP [192]
Version 2.30
Subsection O.N Norm 197
Subsection N
Norm
If treating linear algebra in a more geometric fashion, the length of a vector occurs naturally, and is what
you would expect from its name. With complex numbers, we will dene a similar function. Recall that if
cis a complex number, then jcjdenotes its modulus (Denition MCN [760]).
Denition NV
Norm of a Vector
Thenorm of the vector uis the scalar quantity in C
kuk=q
j[u]1j2+j[u]2j2+j[u]3j2++j[u]mj2=vuutmX
i=1j[u]ij2
(This denition contains Notation NV.) 4
Computing a norm is also easy to do.
Example CNSV
Computing the norm of some vectors
The norm of
u=2
6643 + 2i
1 6i
2 + 4i
2 +i3
775
is
kuk=q
j3 + 2ij2+j1 6ij2+j2 + 4ij2+j2 +ij2=p
13 + 37 + 20 + 5 =p
75 = 5p
3:
The norm of
v=2
666643
1
2
4
33
77775
is
kvk=q
j3j2+j 1j2+j2j2+j4j2+j 3j2=p
32+ 12+ 22+ 42+ 32=p
39:
Notice how the norm of a vector with real number entries is just the length of the vector. Inner products
and norms are related by the following theorem.
Theorem IPN
Inner Products and Norms
Suppose that uis a vector in Cm. Thenkuk2=hu;ui.
Proof
kuk2=0
@vuutmX
i=1j[u]ij21
A2
Denition NV [195]
=mX
i=1j[u]ij2
Version 2.30
198 Section O Orthogonality
=mX
i=1[u]i[u]i Denition MCN [760]
=hu;ui Denition IP [192]
When our vectors have entries only from the real numbers Theorem IPN [195] says that the dot product
of a vector with itself is equal to the length of the vector squared.
Theorem PIP
Positive Inner Products
Suppose that uis a vector in Cm. Thenhu;ui0 with equality if and only if u=0.
Proof From the proof of Theorem IPN [195] we see that
hu;ui=j[u]1j2+j[u]2j2+j[u]3j2++j[u]mj2
Since each modulus is squared, every term is positive, and the sum must also be positive. (Notice that in
general the inner product is a complex number and cannot be compared with zero, but in the special case
ofhu;uithe result is a real number.) The phrase, \with equality if and only if" means that we want to
show that the statement hu;ui= 0 (i.e. with equality) is equivalent (\if and only if") to the statement
u=0.
Ifu=0, then it is a straightforward computation to see that hu;ui= 0. In the other direction, assume
thathu;ui= 0. As before,hu;uiis a sum of moduli. So we have
0 =hu;ui=j[u]1j2+j[u]2j2+j[u]3j2++j[u]mj2
Now we have a sum of squares equaling zero, so each term must be zero. Then by similar logic, j[u]ij= 0
will imply that [ u]i= 0, since 0 + 0 iis the only complex number with zero modulus. Thus every entry of
uis zero and so u=0, as desired.
Notice that Theorem PIP [196] contains three implications:
u2Cm)hu;ui0
u=0)hu;ui= 0
hu;ui= 0)u=0
The results contained in Theorem PIP [196] are summarized by saying \the inner product is positive
denite ."
Subsection OV
Orthogonal Vectors
\Orthogonal" is a generalization of \perpendicular." You may have used mutually perpendicular vectors in
a physics class, or you may recall from a calculus class that perpendicular vectors have a zero dot product.
We will now extend these ideas into the realm of higher dimensions and complex scalars.
Denition OV
Orthogonal Vectors
A pair of vectors, uandv, from Cmareorthogonal if their inner product is zero, that is, hu;vi= 0.4
Example TOV
Two orthogonal vectors
Version 2.30
Subsection O.OV Orthogonal Vectors 199
The vectors
u=2
6642 + 3i
4 2i
1 +i
1 +i3
775v=2
6641 i
2 + 3i
4 6i
13
775
are orthogonal since
hu;vi= (2 + 3i)(1 +i) + (4 2i)(2 3i) + (1 +i)(4 + 6i) + (1 +i)(1)
= ( 1 + 5i) + (2 16i) + ( 2 + 10i) + (1 +i)
= 0 + 0i:
We extend this denition to whole sets by requiring vectors to be pairwise orthogonal. Despite using
the same word, careful thought about what objects you are using will eliminate any source of confusion.
Denition OSV
Orthogonal Set of Vectors
Suppose that S=fu1;u2;u3; :::; ungis a set of vectors from Cm. ThenSis anorthogonal set if every
pair of dierent vectors from Sis orthogonal, that is hui;uji= 0 whenever i6=j. 4
We now dene the prototypical orthogonal set, which we will reference repeatedly.
Denition SUV
Standard Unit Vectors
Letej2Cm, 1jmdenote the column vectors dened by
[ej]i=(
0 ifi6=j
1 ifi=j
Then the set
fe1;e2;e3; :::; emg=fejj1jmg
is the set of standard unit vectors inCm.
(This denition contains Notation SUV.) 4
Notice that ejis identical to column jof themmidentity matrix Im(Denition IM [84]). This
observation will often be useful. It is not hard to see that the set of standard unit vectors is an orthogonal
set. We will reserve the notation eifor these vectors.
Example SUVOS
Standard Unit Vectors are an Orthogonal Set
Compute the inner product of two distinct vectors from the set of standard unit vectors (Denition SUV
[197]), say ei,ej, wherei6=j,
hei;eji= 00 + 0 0 ++ 10 ++ 00 ++ 01 ++ 00 + 0 0
= 0(0) + 0(0) ++ 1(0) ++ 0(1) ++ 0(0) + 0(0)
= 0
So the setfe1;e2;e3; :::; emgis an orthogonal set.
Example AOS
An orthogonal set
Version 2.30
200 Section O Orthogonality
The set
fx1;x2;x3;x4g=8
>><
>>:2
6641 +i
1
1 i
i3
775;2
6641 + 5i
6 + 5i
7 i
1 6i3
775;2
664 7 + 34i
8 23i
10 + 22i
30 + 13i3
775;2
664 2 4i
6 +i
4 + 3i
6 i3
7759
>>=
>>;
is an orthogonal set. Since the inner product is anti-commutative (Theorem IPAC [194]) we can test pairs
of dierent vectors in any order. If the result is zero, then it will also be zero if the inner product is
computed in the opposite order. This means there are six pairs of dierent vectors to use in an inner
product computation. We'll do two and you can practice your inner products on the other four.
hx1;x3i= (1 +i)( 7 34i) + (1)( 8 + 23i) + (1 i)( 10 22i) + (i)(30 13i)
= (27 41i) + ( 8 + 23i) + ( 32 12i) + (13 + 30 i)
= 0 + 0i
and
hx2;x4i= (1 + 5i)( 2 + 4i) + (6 + 5i)(6 i) + ( 7 i)(4 3i) + (1 6i)(6 +i)
= ( 22 6i) + (41 + 24 i) + ( 31 + 17i) + (12 35i)
= 0 + 0i
So far, this section has seen lots of denitions, and lots of theorems establishing un-surprising conse-
quences of those denitions. But here is our rst theorem that suggests that inner products and orthogonal
vectors have some utility. It is also one of our rst illustrations of how to arrive at linear independence as
the conclusion of a theorem.
Theorem OSLI
Orthogonal Sets are Linearly Independent
Suppose that Sis an orthogonal set of nonzero vectors. Then Sis linearly independent.
Proof LetS=fu1;u2;u3; :::; ungbe an orthogonal set of nonzero vectors. To prove the linear
independence of S, we can appeal to the denition (Denition LICV [153]) and begin with an arbitrary
relation of linear dependence (Denition RLDCV [153]),
1u1+2u2+3u3++nun=0:
Then, for every 1 in, we have
i=1
hui;uii(ihui;uii) Theorem PIP [196]
=1
hui;uii(1(0) +2(0) ++ihui;uii++n(0)) Property ZCN [759]
=1
hui;uii(1hu1;uii++ihui;uii++nhun;uii) Denition OSV [197]
=1
hui;uii(h1u1;uii+h2u2;uii++hnun;uii) Theorem IPSM [194]
=1
hui;uiih1u1+2u2+3u3++nun;uii Theorem IPVA [193]
=1
hui;uiih0;uii Denition RLDCV [153]
=1
hui;uii0 Denition IP [192]
Version 2.30
Subsection O.GSP Gram-Schmidt Procedure 201
= 0 Property ZCN [759]
So we conclude that i= 0 for all 1inin any relation of linear dependence on S. But this says that
Sis a linearly independent set since the only way to form a relation of linear dependence is the trivial way
(Denition LICV [153]). Boom!
Subsection GSP
Gram-Schmidt Procedure
The Gram-Schmidt Procedure is really a theorem. It says that if we begin with a linearly independent set
ofpvectors,S, then we can do a number of calculations with these vectors and produce an orthogonal
set ofpvectors,T, so thathSi=hTi. Given the large number of computations involved, it is indeed a
procedure to do all the necessary computations, and it is best employed on a computer. However, it also
has value in proofs where we may on occasion wish to replace a linearly independent set by an orthogonal
set.
This is our rst occasion to use the technique of \mathematical induction" for a proof, a technique
we will see again several times, especially in Chapter D [423]. So study the simple example described in
Technique I [772] rst.
Theorem GSP
Gram-Schmidt Procedure
Suppose that S=fv1;v2;v3; :::; vpgis a linearly independent set of vectors in Cm. Dene the vectors
ui, 1ipby
ui=vi hvi;u1i
hu1;u1iu1 hvi;u2i
hu2;u2iu2 hvi;u3i
hu3;u3iu3 hvi;ui 1i
hui 1;ui 1iui 1
Then ifT=fu1;u2;u3; :::; upg, thenTis an orthogonal set of non-zero vectors, and hTi=hSi.
Proof We will prove the result by using induction on p(Technique I [772]). To begin, we prove that T
has the desired properties when p= 1. In this case u1=v1andT=fu1g=fv1g=S. BecauseSand
Tare equal,hSi=hTi. Equally trivial, Tis an orthogonal set. If u1=0, thenSwould be a linearly
dependent set, a contradiction.
Suppose that the theorem is true for any set of p 1 linearly independent vectors. Let S=fv1;v2;v3; :::; vpg
be a linearly independent set of pvectors. Then S0=fv1;v2;v3; :::; vp 1gis also linearly independent.
So we can apply the theorem to S0and construct the vectors T0=fu1;u2;u3; :::; up 1g.T0is therefore
an orthogonal set of nonzero vectors and hS0i=hT0i. Dene
up=vp hvp;u1i
hu1;u1iu1 hvp;u2i
hu2;u2iu2 hvp;u3i
hu3;u3iu3 hvp;up 1i
hup 1;up 1iup 1
and letT=T0[fupg. We need to now show that Thas several properties by building on what we know
aboutT0. But rst notice that the above equation has no problems with the denominators ( hui;uii) being
zero, since the uiare fromT0, which is composed of nonzero vectors.
We show thathTi=hSi, by rst establishing that hTihSi. Suppose x2hTi, so
x=a1u1+a2u2+a3u3++apup
The termapupis a linear combination of vectors from T0and the vector vp, while the remaining terms are
a linear combination of vectors from T0. SincehT0i=hS0i, any term that is a multiple of a vector from T0
can be rewritten as a linear combination of vectors from S0. The remaining term apvpis a multiple of a
vector inS. So we see that xcan be rewritten as a linear combination of vectors from S, i.e.x2hSi.
Version 2.30
202 Section O Orthogonality
To show thathSihTi, begin with y2hSi, so
y=a1v1+a2v2+a3v3++apvp
Rearrange our dening equation for upby solving for vp. Then the term apvpis a multiple of a linear
combination of elements of T. The remaining terms are a linear combination of v1;v2;v3; :::; vp 1,
hence an element of hS0i=hT0i. Thus these remaining terms can be written as a linear combination of
the vectors in T0. Soyis a linear combination of vectors from T, i.e.y2hTi.
The elements of T0are nonzero, but what about up? Suppose to the contrary that up=0,
0=up=vp hvp;u1i
hu1;u1iu1 hvp;u2i
hu2;u2iu2 hvp;u3i
hu3;u3iu3 hvp;up 1i
hup 1;up 1iup 1
vp=hvp;u1i
hu1;u1iu1+hvp;u2i
hu2;u2iu2+hvp;u3i
hu3;u3iu3++hvp;up 1i
hup 1;up 1iup 1
SincehS0i=hT0iwe can write the vectors u1;u2;u3; :::; up 1on the right side of this equation in terms
of the vectors v1;v2;v3; :::; vp 1and we then have the vector vpexpressed as a linear combination of
the otherp 1 vectors in S, implying that Sis a linearly dependent set (Theorem DLDS [175]), contrary
to our lone hypothesis about S.
Finally, it is a simple matter to establish that Tis an orthogonal set, though it will not appear so
simple looking. Think about your objects as you work through the following | what is a vector and what
is a scalar. Since T0is an orthogonal set by induction, most pairs of elements in Tare already known to
be orthogonal. We just need to test \new" inner products, between upandui, for 1ip 1. Here we
go, using summation notation,
hup;uii=*
vp p 1X
k=1hvp;uki
huk;ukiuk;ui+
=hvp;uii *p 1X
k=1hvp;uki
huk;ukiuk;ui+
Theorem IPVA [193]
=hvp;uii p 1X
k=1hvp;uki
huk;ukiuk;ui
Theorem IPVA [193]
=hvp;uii p 1X
k=1hvp;uki
huk;ukihuk;uii Theorem IPSM [194]
=hvp;uii hvp;uii
hui;uiihui;uii X
k6=ihvp;uki
huk;uki(0) Induction Hypothesis
=hvp;uii hvp;uii X
k6=i0
= 0
Example GSTV
Gram-Schmidt of three vectors
We will illustrate the Gram-Schmidt process with three vectors. Begin with the linearly independent (check
this!) set
S=fv1;v2;v3g=8
<
:2
41
1 +i
13
5;2
4 i
1
1 +i3
5;2
40
i
i3
59
=
;
Version 2.30
Subsection O.GSP Gram-Schmidt Procedure 203
Then
u1=v1=2
41
1 +i
13
5
u2=v2 hv2;u1i
hu1;u1iu1=1
42
4 2 3i
1 i
2 + 5i3
5
u3=v3 hv3;u1i
hu1;u1iu1 hv3;u2i
hu2;u2iu2=1
112
4 3 i
1 + 3i
1 i3
5
and
T=fu1;u2;u3g=8
<
:2
41
1 +i
13
5;1
42
4 2 3i
1 i
2 + 5i3
5;1
112
4 3 i
1 + 3i
1 i3
59
=
;
is an orthogonal set (which you can check) of nonzero vectors and hTi=hSi(all by Theorem GSP [199]).
Of course, as a by-product of orthogonality, the set Tis also linearly independent (Theorem OSLI [198]).
One nal denition related to orthogonal vectors.
Denition ONS
OrthoNormal Set
SupposeS=fu1;u2;u3; :::; ungis an orthogonal set of vectors such that kuik= 1 for all 1in.
ThenSis anorthonormal set of vectors. 4
Once you have an orthogonal set, it is easy to convert it to an orthonormal set | multiply each vector
by the reciprocal of its norm, and the resulting vector will have norm 1. This scaling of each vector will
not aect the orthogonality properties (apply Theorem IPSM [194]).
Example ONTV
Orthonormal set, three vectors
The set
T=fu1;u2;u3g=8
<
:2
41
1 +i
13
5;1
42
4 2 3i
1 i
2 + 5i3
5;1
112
4 3 i
1 + 3i
1 i3
59
=
;
from Example GSTV [200] is an orthogonal set. We compute the norm of each vector,
ku1k= 2 ku2k=1
2p
11 ku3k=p
2p
11
Converting each vector to a norm of 1, yields an orthonormal set,
w1=1
22
41
1 +i
13
5
w2=1
1
2p
111
42
4 2 3i
1 i
2 + 5i3
5=1
2p
112
4 2 3i
1 i
2 + 5i3
5
w3=1
p
2p
111
112
4 3 i
1 + 3i
1 i3
5=1p
222
4 3 i
1 + 3i
1 i3
5
Version 2.30
204 Section O Orthogonality
Example ONFV
Orthonormal set, four vectors
As an exercise convert the linearly independent set
S=8
>><
>>:2
6641 +i
1
1 i
i3
775;2
664i
1 +i
1
i3
775;2
664i
i
1 +i
13
775;2
664 1 i
i
1
13
7759
>>=
>>;
to an orthogonal set via the Gram-Schmidt Process (Theorem GSP [199]) and then scale the vectors to
norm 1 to create an orthonormal set. You should get the same set you would if you scaled the orthogonal
set of Example AOS [197] to become an orthonormal set.
It is crazy to do all but the simplest and smallest instances of the Gram-Schmidt procedure by hand.
Well, OK, maybe just once or twice to get a good understanding of Theorem GSP [199]. After that, let a
machine do the work for you. That's what they are for. See: Computation GSP.MMA [748]
We will see orthonormal sets again in Subsection MINM.UM [262]. They are intimately related to
unitary matrices (Denition UM [262]) through Theorem CUMOS [263]. Some of the utility of orthonormal
sets is captured by Theorem COB [378] in Subsection B.OBC [377]. Orthonormal sets appear once again
in Section OD [675] where they are key in orthonormal diagonalization.
Subsection READ
Reading Questions
1. Is the set 8
<
:2
41
1
23
5;2
45
3
13
5;2
48
4
23
59
=
;
an orthogonal set? Why?
2. What is the distinction between an orthogonal set and an orthonormal set?
3. What is nice about the output of the Gram-Schmidt process?
Version 2.30
Subsection O.EXC Exercises 205
Subsection EXC
Exercises
C20 Complete Example AOS [197] by verifying that the four remaining inner products are zero.
Contributed by Robert Beezer
C21 Verify that the set Tcreated in Example GSTV [200] by the Gram-Schmidt Procedure is an or-
thogonal set.
Contributed by Robert Beezer
M60 Suppose thatfu;v;wgCnis an orthonormal set. Prove that u+vis not orthogonal to v+w.
Contributed by Manley Perkel
T10 Prove part 1 of the conclusion of Theorem IPVA [193].
Contributed by Robert Beezer
T11 Prove part 1 of the conclusion of Theorem IPSM [194].
Contributed by Robert Beezer
T20 Suppose that u;v;w2Cn,; 2Canduis orthogonal to both vandw. Prove that uis
orthogonal to v+w.
Contributed by Robert Beezer Solution [204]
T30 Suppose that the set Sin the hypothesis of Theorem GSP [199] is not just linearly independent,
but is also orthogonal. Prove that the set Tcreated by the Gram-Schmidt procedure is equal to S. (Note
that we are getting a stronger conclusion than hTi=hSi| the conclusion is that T=S.) In other words,
it is pointless to apply the Gram-Schmidt procedure to a set that is already orthogonal.
Contributed by Steve Caneld
Version 2.30
206 Section O Orthogonality
Subsection SOL
Solutions
T20 Contributed by Robert Beezer Statement [203]
Vectors are orthogonal if their inner product is zero (Denition OV [196]), so we compute,
hv+w;ui=hv;ui+hw;ui Theorem IPVA [193]
=hv;ui+hw;ui Theorem IPSM [194]
=(0) +(0) Denition OV [196]
= 0
So by Denition OV [196], uandv+ware an orthogonal pair of vectors.
Version 2.30
Annotated Acronyms O.V Vectors 207
Annotated Acronyms V
Vectors
Theorem VSPCV [100]
These are the fundamental rules for working with the addition, and scalar multiplication, of column vectors.
We will see something very similar in the next chapter (Theorem VSPM [209]) and then this will be
generalized into what is arguably our most important denition, Denition VS [317].
Theorem SLSLC [112]
Vector addition and scalar multiplication are the two fundamental operations on vectors, and linear com-
binations roll them both into one. Theorem SLSLC [112] connects linear combinations with systems of
equations. This one we will see often enough that it is worth memorizing.
Theorem PSPHS [124]
This theorem is interesting in its own right, and sometimes the vaugeness surrounding the choice of zcan
seem mysterious. But we list it here because we will see an important theorem in Section ILT [541] which
will generalize this result (Theorem KPI [547]).
Theorem LIVRN [156]
If you have a set of column vectors, this is the fastest computational approach to determine if the set is
linearly independent. Make the vectors the columns of a matrix, row-reduce, compare randn. That's it
| and you always get an answer. Put this one in your toolkit.
Theorem BNS [160]
We will have several theorems (all listed in these \Annotated Acronyms" sections) whose conclusions will
provide a linearly independent set of vectors whose span equals some set of interest (the null space here).
While the notation in this theorem might appear gruesome, in practice it can become very routine to apply.
So practice this one | we'll be using it all through the book.
Theorem BS [180]
As promised, another theorem that provides a linearly independent set of vectors whose span equals some
set of interest (a span now). You can use this one to clean up anyspan.
Version 2.30
208 Section O Orthogonality
Version 2.30
Chapter M
Matrices
We have made frequent use of matrices for solving systems of equations, and we have begun to investigate
a few of their properties, such as the null space and nonsingularity. In this chapter, we will take a more
systematic approach to the study of matrices.
Section MO
Matrix Operations
In this section we will back up and start simple. First a denition of a totally general set of matrices.
Denition VSM
Vector Space of mnMatrices
The vector space Mmnis the set of all mnmatrices with entries from the set of complex numbers.
(This denition contains Notation VSM.) 4
Subsection MEASM
Matrix Equality, Addition, Scalar Multiplication
Just as we made, and used, a careful denition of equality for column vectors, so too, we have precise
denitions for matrices.
Denition ME
Matrix Equality
ThemnmatricesAandBareequal , writtenA=Bprovided [A]ij= [B]ijfor all 1im, 1jn.
(This denition contains Notation ME.) 4
So equality of matrices translates to the equality of complex numbers, on an entry-by-entry basis. Notice
that we now have yet another denition that uses the symbol \=" for shorthand. Whenever a theorem
has a conclusion saying two matrices are equal (think about your objects), we will consider appealing
to this denition as a way of formulating the top-level structure of the proof. We will now dene two
operations on the set Mmn. Again, we will overload a symbol (`+') and a convention (juxtaposition for
scalar multiplication).
Denition MA
Matrix Addition
Given themnmatricesAandB, dene the sum ofAandBas anmnmatrix, written A+B,
209
210 Section MO Matrix Operations
according to
[A+B]ij= [A]ij+ [B]ij 1im;1jn
(This denition contains Notation MA.) 4
So matrix addition takes two matrices of the same size and combines them (in a natural way!) to create
a new matrix of the same size. Perhaps this is the \obvious" thing to do, but it doesn't relieve us from
the obligation to state it carefully.
Example MA
Addition of two matrices in M23
If
A=2 3 4
1 0 7
B=6 2 4
3 5 2
then
A+B=2 3 4
1 0 7
+6 2 4
3 5 2
=2 + 6 3 + 2 4 + ( 4)
1 + 3 0 + 5 7 + 2
=8 1 0
4 5 5
Our second operation takes two objects of dierent types, specically a number and a matrix, and
combines them to create another matrix. As with vectors, in this context we call a number a scalar in
order to emphasize that it is not a matrix.
Denition MSM
Matrix Scalar Multiplication
Given themnmatrixAand the scalar 2C, thescalar multiple ofAis anmnmatrix, written
Aand dened according to
[A]ij=[A]ij 1im;1jn
(This denition contains Notation MSM.) 4
Notice again that we have yet another kind of multiplication, and it is again written putting two
symbols side-by-side. Computationally, scalar matrix multiplication is very easy.
Example MSM
Scalar multiplication in M32
If
A=2
42 8
3 5
0 13
5
and= 7, then
A= 72
42 8
3 5
0 13
5=2
47(2) 7(8)
7( 3) 7(5)
7(0) 7(1)3
5=2
414 56
21 35
0 73
5
Version 2.30
Subsection MO.VSP Vector Space Properties 211
Subsection VSP
Vector Space Properties
With denitions of matrix addition and scalar multiplication we can now state, and prove, several properties
of each operation, and some properties that involve their interplay. We now collect ten of them here for
later reference.
Theorem VSPM
Vector Space Properties of Matrices
Suppose that Mmnis the set of all mnmatrices (Denition VSM [207]) with addition and scalar
multiplication as dened in Denition MA [207] and Denition MSM [208]. Then
ACM Additive Closure, Matrices
IfA; B2Mmn, thenA+B2Mmn.
SCM Scalar Closure, Matrices
If2CandA2Mmn, thenA2Mmn.
CM Commutativity, Matrices
IfA; B2Mmn, thenA+B=B+A.
AAM Additive Associativity, Matrices
IfA; B; C2Mmn, thenA+ (B+C) = (A+B) +C.
ZM Zero Vector, Matrices
There is a matrix, O, called the zero matrix , such that A+O=Afor allA2Mmn.
AIM Additive Inverses, Matrices
IfA2Mmn, then there exists a matrix A2Mmnso thatA+ ( A) =O.
SMAM Scalar Multiplication Associativity, Matrices
If; 2CandA2Mmn, then(A) = ()A.
DMAM Distributivity across Matrix Addition, Matrices
If2CandA; B2Mmn, then(A+B) =A+B.
DSAM Distributivity across Scalar Addition, Matrices
If; 2CandA2Mmn, then (+)A=A+A.
OM One, Matrices
IfA2Mmn, then 1A=A.
Proof While some of these properties seem very obvious, they all require proof. However, the proofs are
not very interesting, and border on tedious. We'll prove one version of distributivity very carefully, and
you can test your proof-building skills on some of the others. We'll give our new notation for matrix entries
a workout here. Compare the style of the proofs here with those given for vectors in Theorem VSPCV
[100] | while the objects here are more complicated, our notation makes the proofs cleaner.
To prove Property DSAM [209], ( +)A=A+A, we need to establish the equality of two matrices
(see Technique GS [767]). Denition ME [207] says we need to establish the equality of their entries,
one-by-one. How do we do this, when we do not even know how many entries the two matrices might
have? This is where Notation ME [207] comes into play. Ready? Here we go.
Version 2.30
212 Section MO Matrix Operations
Foranyiandj, 1im, 1jn,
[(+)A]ij= (+) [A]ij Denition MSM [208]
=[A]ij+[A]ij Distributivity in C
= [A]ij+ [A]ij Denition MSM [208]
= [A+A]ij Denition MA [207]
There are several things to notice here. (1) Each equals sign is an equality of numbers. (2) The two ends
of the equation, being true for any iandj, allow us to conclude the equality of the matrices by Denition
ME [207]. (3) There are several plus signs, and several instances of juxtaposition. Identify each one, and
state exactly what operation is being represented by each.
For now, note the similarities between Theorem VSPM [209] about matrices and Theorem VSPCV
[100] about vectors.
The zero matrix described in this theorem, O, is what you would expect | a matrix full of zeros.
Denition ZM
Zero Matrix
Themnzero matrix is written asO=Omnand dened by [O]ij= 0, for all 1im, 1jn.
(This denition contains Notation ZM.) 4
Subsection TSM
Transposes and Symmetric Matrices
We describe one more common operation we can perform on matrices. Informally, to transpose a matrix
is to build a new matrix by swapping its rows and columns.
Denition TM
Transpose of a Matrix
Given anmnmatrixA, itstranspose is thenmmatrixAtgiven by
At
ij= [A]ji;1in;1jm:
(This denition contains Notation TM.) 4
Example TM
Transpose of a 34matrix
Suppose
D=2
43 7 2 3
1 4 2 8
0 3 2 53
5:
We could formulate the transpose, entry-by-entry, using the denition. But it is easier to just systematically
rewrite rows as columns (or vice-versa). The form of the denition given will be more useful in proofs. So
we have
Dt=2
6643 1 0
7 4 3
2 2 2
3 8 53
775
Version 2.30
Subsection MO.TSM Transposes and Symmetric Matrices 213
It will sometimes happen that a matrix is equal to its transpose. In this case, we will call a matrix
symmetric . These matrices occur naturally in certain situations, and also have some nice properties, so
it is worth stating the denition carefully. Informally a matrix is symmetric if we can \
ip" it about the
main diagonal (upper-left corner, running down to the lower-right corner) and have it look unchanged.
Denition SYM
Symmetric Matrix
The matrix Aissymmetric ifA=At. 4
Example SYM
A symmetric 55matrix
The matrix
E=2
666642 3 9 5 7
3 1 6 2 3
9 6 0 1 9
5 2 1 4 8
7 3 9 8 33
77775
is symmetric.
You might have noticed that Denition SYM [211] did not specify the size of the matrix A, as has been
our custom. That's because it wasn't necessary. An alternative would have been to state the denition
just for square matrices, but this is the substance of the next proof. Before reading the next proof, we
want to oer you some advice about how to become more procient at constructing proofs. Perhaps you
can apply this advice to the next theorem. Have a peek at Technique P [774] now.
Theorem SMS
Symmetric Matrices are Square
Suppose that Ais a symmetric matrix. Then Ais square.
Proof We start by specifying A's size, without assuming it is square, since we are trying to prove that,
so we can't also assume it. Suppose Ais anmnmatrix. Because Ais symmetric, we know by Denition
SM [428] that A=At. So, in particular, Denition ME [207] requires that AandAtmust have the same
size. The size of Atisnm. BecauseAhasmrows andAthasnrows, we conclude that m=n, and
henceAmust be square by Denition SQM [83].
We nish this section with three easy theorems, but they illustrate the interplay of our three new
operations, our new notation, and the techniques used to prove matrix equalities.
Theorem TMA
Transpose and Matrix Addition
Suppose that AandBaremnmatrices. Then ( A+B)t=At+Bt.
Proof The statement to be proved is an equality of matrices, so we work entry-by-entry and use Denition
ME [207]. Think carefully about the objects involved here, and the many uses of the plus sign. For
1im, 1jn,
(A+B)t
ij= [A+B]ji Denition TM [210]
= [A]ji+ [B]ji Denition MA [207]
=
At
ij+
Bt
ijDenition TM [210]
=
At+Bt
ijDenition MA [207]
Version 2.30
214 Section MO Matrix Operations
Since the matrices ( A+B)tandAt+Btagree at each entry, Denition ME [207] tells us the two matrices
are equal.
Theorem TMSM
Transpose and Matrix Scalar Multiplication
Suppose that 2CandAis anmnmatrix. Then ( A)t=At.
Proof The statement to be proved is an equality of matrices, so we work entry-by-entry and use Denition
ME [207]. Notice that the desired equality is of nmmatrices, and think carefully about the objects
involved here, plus the many uses of juxtaposition. For 1 im, 1jn,
(A)t
ji= [A]ij Denition TM [210]
=[A]ij Denition MSM [208]
=
At
jiDenition TM [210]
=
At
jiDenition MSM [208]
Since the matrices ( A)tandAtagree at each entry, Denition ME [207] tells us the two matrices are
equal.
Theorem TT
Transpose of a Transpose
Suppose that Ais anmnmatrix. Then
Att=A.
Proof We again want to prove an equality of matrices, so we work entry-by-entry and use Denition ME
[207]. For 1im, 1jn,
h
Atti
ij=
At
jiDenition TM [210]
= [A]ij Denition TM [210]
Its usually straightforward to coax the transpose of a matrix out of a computational device. See:
Computation TM.MMA [749] Computation TM.TI86 [751] Computation TM.SAGE [755]
Subsection MCC
Matrices and Complex Conjugation
As we did with vectors (Denition CCCV [191]), we can dene what it means to take the conjugate of a
matrix.
Denition CCM
Complex Conjugate of a Matrix
SupposeAis anmnmatrix. Then the conjugate ofA, writtenAis anmnmatrix dened by
A
ij=[A]ij
(This denition contains Notation CCM.) 4
Example CCM
Complex conjugate of a matrix
If
A=2 i 3 5 + 4i
3 + 6i2 3i 0
Version 2.30
Subsection MO.MCC Matrices and Complex Conjugation 215
then
A=2 +i 3 5 4i
3 6i2 + 3i 0
The interplay between the conjugate of a matrix and the two operations on matrices is what you might
expect.
Theorem CRMA
Conjugation Respects Matrix Addition
Suppose that AandBaremnmatrices. Then A+B=A+B.
Proof For 1im, 1jn,
A+B
ij=[A+B]ij Denition CCM [212]
=[A]ij+ [B]ij Denition MA [207]
=[A]ij+[B]ij Theorem CCRA [759]
=
A
ij+
B
ijDenition CCM [212]
=
A+B
ijDenition MA [207]
Since the matrices A+BandA+Bare equal in each entry, Denition ME [207] says that A+B=A+B.
Theorem CRMSM
Conjugation Respects Matrix Scalar Multiplication
Suppose that 2CandAis anmnmatrix. Then A=A.
Proof For 1im, 1jn,
A
ij=[A]ij Denition CCM [212]
=[A]ij Denition MSM [208]
=[A]ij Theorem CCRM [760]
=
A
ijDenition CCM [212]
=
A
ijDenition MSM [208]
Since the matrices AandAare equal in each entry, Denition ME [207] says that A=A.
Theorem CCM
Conjugate of the Conjugate of a Matrix
Suppose that Ais anmnmatrix. Then
A
=A.
Proof For 1im, 1jn,
h
Ai
ij=
A
ijDenition CCM [212]
=[A]ij Denition CCM [212]
= [A]ij Theorem CCT [760]
Since the matrices
A
andAare equal in each entry, Denition ME [207] says that
A
=A.
Finally, we will need the following result about matrix conjugation and transposes later.
Version 2.30
216 Section MO Matrix Operations
Theorem MCT
Matrix Conjugation and Transposes
Suppose that Ais anmnmatrix. Then (At) =
At.
Proof For 1im, 1jn,
h
(At)i
ji=[At]ji Denition CCM [212]
=[A]ij Denition TM [210]
=
A
ijDenition CCM [212]
=h
Ati
jiDenition TM [210]
Since the matrices (At) and
Atare equal in each entry, Denition ME [207] says that (At) =
At.
Subsection AM
Adjoint of a Matrix
The combination of transposing and conjugating a matrix will be important in subsequent sections, such
as Subsection MINM.UM [262] and Section OD [675]. We make a key denition here and prove some basic
results in the same spirit as those above.
Denition A
Adjoint
IfAis a matrix, then its adjoint isA=
At.
(This denition contains Notation A.) 4
You will see the adjoint written elsewhere variously as AH,AorAy. Notice that Theorem MCT [214]
says it does not really matter if we conjugate and then transpose, or transpose and then conjugate.
Theorem AMA
Adjoint and Matrix Addition
SupposeAandBare matrices of the same size. Then ( A+B)=A+B.
Proof
(A+B)=
A+BtDenition A [214]
=
A+BtTheorem CRMA [213]
=
At+
BtTheorem TMA [211]
=A+BDenition A [214]
Theorem AMSM
Adjoint and Matrix Scalar Multiplication
Suppose2Cis a scalar and Ais a matrix. Then ( A)=A.
Proof
(A)=
AtDenition A [214]
Version 2.30
Subsection MO.READ Reading Questions 217
=
AtTheorem CRMSM [213]
=
AtTheorem TMSM [212]
=ADenition A [214]
Theorem AA
Adjoint of an Adjoint
Suppose that Ais a matrix. Then ( A)=A
Proof
(A)=
(A)t
Denition A [214]
=
(A)t
Theorem MCT [214]
=
Att
Denition A [214]
=
A
Theorem TT [212]
=A Theorem CCM [213]
Take note of how the theorems in this section, while simple, build on earlier theorems and denitions
and never descend to the level of entry-by-entry proofs based on Denition ME [207]. In other words, the
equal signs that appear in the previous proofs are equalities of matrices, not scalars (which is the opposite
of a proof like that of Theorem TMA [211]).
Subsection READ
Reading Questions
1. Perform the following matrix computation.
(6)2
42 2 8 1
4 5 1 3
7 3 0 23
5+ ( 2)2
42 7 1 2
3 1 0 5
1 7 3 33
5
2. Theorem VSPM [209] reminds you of what previous theorem? How strong is the similarity?
3. Compute the transpose of the matrix below.
2
46 8 4
2 1 0
9 5 63
5
Version 2.30
218 Section MO Matrix Operations
Subsection EXC
Exercises
C10 LetA=1 4 3
6 3 0
,B=3 2 1
2 6 5
andC=2
42 4
4 0
2 23
5. Let= 4 and= 1=2. Perform the
following calculations:
1.A+B
2.A+C
3.Bt+C
4.A+Bt
5.C
6. 4A 3B
7.At+C
8.A+B Ct
9. 4A+ 2B 5Ct
Contributed by Chris Black Solution [219]
C11 Solve the given vector equation for x, or explain why no solution exists:
21 2 3
0 4 2
31 1 2
0 1x
= 1 1 0
0 5 2
Contributed by Chris Black Solution [219]
C12 Solve the given vector equation for , or explain why no solution exists:
1 3 4
2 1 1
+4 3 6
0 1 1
=7 12 6
6 4 2
Contributed by Chris Black Solution [219]
C13 Solve the given vector equation for , or explain why no solution exists:
2
43 1
2 0
1 43
5 2
44 1
3 2
0 13
5=2
42 1
1 2
2 63
5
Contributed by Chris Black Solution [220]
C14 Findandthat solve the following equation:
1 2
4 1
+2 1
3 1
= 1 4
6 1
Version 2.30
Subsection MO.EXC Exercises 219
Contributed by Chris Black Solution [220]
In Chapter V [97] we dened the operations of vector addition and vector scalar multiplication in
Denition CVA [98] and Denition CVSM [99]. These two operations formed the underpinnings of the
remainder of the chapter. We have now dened similar operations for matrices in Denition MA [207] and
Denition MSM [208]. You will have noticed the resulting similarities between Theorem VSPCV [100] and
Theorem VSPM [209].
In Exercises M20{M25, you will be asked to extend these similarities to other fundamental denitions
and concepts we rst saw in Chapter V [97]. This sequence of problems was suggested by Martin Jackson.
M20 SupposeS=fB1; B2; B3; :::; Bpgis a set of matrices from Mmn. Formulate appropriate def-
initions for the following terms and give an example of the use of each.
1. A linear combination of elements of S.
2. A relation of linear dependence on S, both trivial and non-trivial.
3.Sis a linearly independent set.
4.hSi.
Contributed by Robert Beezer
M21 Show that the set Sis linearly independent in M2;2.
S=1 0
0 0
;0 1
0 0
;0 0
1 0
;0 0
0 1
Contributed by Robert Beezer Solution [220]
M22 Determine if the set
S= 2 3 4
1 3 2
;4 2 2
0 1 1
; 1 2 2
2 2 2
; 1 1 0
1 0 2
; 1 2 2
0 1 2
is linearly independent in M2;3.
Contributed by Robert Beezer Solution [221]
M23 Determine if the matrix Ais in the span of S. In other words, is A2hSi? If so write Aas a linear
combination of the elements of S.
A= 13 24 2
8 2 20
S= 2 3 4
1 3 2
;4 2 2
0 1 1
; 1 2 2
2 2 2
; 1 1 0
1 0 2
; 1 2 2
0 1 2
Contributed by Robert Beezer Solution [221]
M24 SupposeYis the set of all 3 3 symmetric matrices (Denition SYM [211]). Find a set Tso that
Tis linearly independent and hTi=Y.
Contributed by Robert Beezer Solution [221]
Version 2.30
220 Section MO Matrix Operations
M25 Dene a subset of M3;3by
U33=n
A2M3;3j[A]ij= 0 whenever i>jo
Find a setRso thatRis linearly independent and hRi=U33.
Contributed by Robert Beezer
T13 Prove Property CM [209] of Theorem VSPM [209]. Write your proof in the style of the proof of
Property DSAM [209] given in this section.
Contributed by Robert Beezer Solution [221]
T14 Prove Property AAM [209] of Theorem VSPM [209]. Write your proof in the style of the proof of
Property DSAM [209] given in this section.
Contributed by Robert Beezer
T17 Prove Property SMAM [209] of Theorem VSPM [209]. Write your proof in the style of the proof of
Property DSAM [209] given in this section.
Contributed by Robert Beezer
T18 Prove Property DMAM [209] of Theorem VSPM [209]. Write your proof in the style of the proof of
Property DSAM [209] given in this section.
Contributed by Robert Beezer
A matrixAisskew-symmetric ifAt= AExercises T30{T37 employ this denition.
T30 Prove that a skew-symmetric matrix is square. (Hint: study the proof of Theorem SMS [211].)
Contributed by Robert Beezer
T31 Prove that a skew-symmetric matrix must have zeros for its diagonal elements. In other words, if A
is skew-symmetric of size n, then [A]ii= 0 for 1in. (Hint: carefully construct an example of a 3 3
skew-symmetric matrix before attempting a proof.)
Contributed by Manley Perkel
T32 Prove that a matrix Ais both skew-symmetric and symmetric if and only if Ais the zero matrix.
(Hint: one half of this proof is very easy, the other half takes a little more work.)
Contributed by Manley Perkel
T33 SupposeAandBare both skew-symmetric matrices of the same size and ; 2C. Prove that
A+Bis a skew-symmetric matrix.
Contributed by Manley Perkel
T34 SupposeAis a square matrix. Prove that A+Atis a symmetric matrix.
Contributed by Manley Perkel
T35 SupposeAis a square matrix. Prove that A Atis a skew-symmetric matrix.
Contributed by Manley Perkel
T36 SupposeAis a square matrix. Prove that there is a symmetric matrix Band a skew-symmetric
matrixCsuch thatA=B+C. In other words, any square matrix can be decomposed into a symmetric
matrix and a skew-symmetric matrix (Technique DC [772]). (Hint: consider building a proof on Exercise
MO.T34 [218] and Exercise MO.T35 [218].)
Contributed by Manley Perkel
T37 Prove that the decomposition in Exercise MO.T36 [218] is unique (see Technique U [771]). (Hint:
a proof can turn on Exercise MO.T31 [218].)
Contributed by Manley Perkel
Version 2.30
Subsection MO.SOL Solutions 221
Subsection SOL
Solutions
C10 Contributed by Chris Black Statement [216]
1.A+B=4 6 2
4 3 5
2.A+Cis undened; AandCare not the same size.
3.Bt+C=2
45 2
6 6
1 73
5
4.A+Btis undened; AandBtare not the same size.
5.C=2
41 2
2 0
1 13
5
6. 4A 3B= 5 10 15
30 30 15
7.At+C=2
49 22
20 3
11 83
5
8.A+B Ct=2 2 0
0 3 3
9. 4A+ 2B 5Ct=0 0 0
0 0 0
C11 Contributed by Chris Black Statement [216]
The given equation
1 1 0
0 5 2
= 21 2 3
0 4 2
31 1 2
0 1x
= 1 1 0
0 5 4 3x
is valid only if 4 3x= 2. Thus, the only solution is x= 2.
C12 Contributed by Chris Black Statement [216]
The given equation
7 12 6
6 4 2
=1 3 4
2 1 1
+4 3 6
0 1 1
=34
2
+4 3 6
0 1 1
=4 +3 + 3 6 + 4
2 1 + 1
leads to the 6 equations in :
4 += 7
Version 2.30
222 Section MO Matrix Operations
3 + 3= 12
6 + 4= 6
2= 6
1 += 4
1 = 2:
The only value that solves all 6 equations is = 3, which is the solution to the original matrix equation.
C13 Contributed by Chris Black Statement [216]
The given equation
2
42 1
1 2
2 63
5=2
43 1
2 0
1 43
5+2
44 1
3 2
0 13
5=2
43 4 1
2 3 2
4 13
5
gives a system of six equations in :
3 4 = 2
1 = 1
2 3 = 1
2 = 2
= 2
4 1 = 6:
Solving each of these equations, we see that the rst 3 and the fth all lead to the solution = 2, the
fourth equation is true no matter what the value of , but the last equation is only solved by = 7=4.
Thus, the system has no solution, and the original matrix equation also has no solution.
C14 Contributed by Chris Black Statement [216]
The given equation
1 4
6 1
=1 2
4 1
+2 1
3 1
=+ 22+
4+ 3 +
gives a system of four equations in two variables
+ 2= 1
2+= 4
4+ 3= 6
+= 1:
Solving this linear system by row-reducing the augmnented matrix shows that = 3,= 2 is the only
solution.
M21 Contributed by Chris Black Statement [217]
Suppose there exist constants ,,
, andso that
1 0
0 0
+0 1
0 0
+
0 0
1 0
+0 0
0 1
=0 0
0 0
:
Then,
0
0 0
+0
0 0
+0 0
0
+0 0
0
=0 0
0 0
Version 2.30
Subsection MO.SOL Solutions 223
so that
=0 0
0 0
. The only solution is then = 0,= 0,
= 0, and= 0, so that the set Sis a
linearly independent set of matrices.
M22 Contributed by Chris Black Statement [217]
Suppose that there exist constants a1,a2,a3,a4, anda5so that
a1 2 3 4
1 3 2
+a24 2 2
0 1 1
+a3 1 2 2
2 2 2
+a4 1 1 0
1 0 2
+a5 1 2 1
0 1 2
=0 0 0
0 0 0
:
Then, we have the matrix equality (Denition ME [207])
2a1+ 4a2 a3 a4 a53a1 2a2 2a3+a4+ 2a5 4a1+ 2a2 2a3 2a5
a1+ 2a3 a4 3a1 a2+ 2a3 a5 2a1+a2+ 2a3+ 2a4 2a5
=0 0 0
0 0 0
;
which yields the linear system of equations
2a1+ 4a2 a3 a4 a5= 0
3a1 2a2 2a3+a4+ 2a5= 0
4a1+ 2a2 2a3 2a5= 0
a1+ 2a3 a4= 0
3a1 a2+ 2a3 a5= 0
2a1+a2+ 2a3+ 2a4 2a5= 0:
By row-reducing the associated 6 5 homogeneous system, we see that the only solution is a1=a2=a3=
a4=a5= 0, so these matrices are a linearly independent subset of M2;3.
M23 Contributed by Chris Black Statement [217]
The matrix Ais in the span of S, since
13 24 2
8 2 20
= 2 2 3 4
1 3 2
24 2 2
0 1 1
3 1 2 2
2 2 2
+ 4 1 2 2
0 1 2
M24 Contributed by Chris Black Statement [217]
Since any symmetric matrix is of the form
2
4a b c
b d e
c e f3
5=2
4a0 0
0 0 0
0 0 03
5+2
40b0
b0 0
0 0 03
5+2
40 0c
0 0 0
c0 03
5+2
40 0 0
0d0
0 0 03
5+2
40 0 0
0 0e
0e03
5+2
40 0 0
0 0 0
0 0f3
5;
Any symmetric matrix is a linear combination of the linearly independent vectors in set Tbelow, so that
hTi=Y:
T=8
<
:2
41 0 0
0 0 0
0 0 03
5;2
40 1 0
1 0 0
0 0 03
5;2
40 0 1
0 0 0
1 0 03
5;2
40 0 0
0 1 0
0 0 03
5;2
40 0 0
0 0 1
0 1 03
5;2
40 0 0
0 0 0
0 0 13
59
=
;
(Something to think about: How do we know that these matrices are linearly independent?)
T13 Contributed by Robert Beezer Statement [218]
For allA; B2Mmnand for all 1im, 1in,
[A+B]ij= [A]ij+ [B]ij Denition MA [207]
= [B]ij+ [A]ij Commutativity in C
= [B+A]ij Denition MA [207]
With equality of each entry of the matrices A+BandB+Abeing equal Denition ME [207] tells us the
two matrices are equal.
Version 2.30
224 Section MO Matrix Operations
Version 2.30
Section MM Matrix Multiplication 225
Section MM
Matrix Multiplication
We know how to add vectors and how to multiply them by scalars. Together, these operations give us the
possibility of making linear combinations. Similarly, we know how to add matrices and how to multiply
matrices by scalars. In this section we mix all these ideas together and produce an operation known as
\matrix multiplication." This will lead to some results that are both surprising and central. We begin
with a denition of how to multiply a vector by a matrix.
Subsection MVP
Matrix-Vector Product
We have repeatedly seen the importance of forming linear combinations of the columns of a matrix. As
one example of this, the oft-used Theorem SLSLC [112], said that every solution to a system of linear
equations gives rise to a linear combination of the column vectors of the coecient matrix that equals the
vector of constants. This theorem, and others, motivate the following central denition.
Denition MVP
Matrix-Vector Product
SupposeAis anmnmatrix with columns A1;A2;A3; :::; Ananduis a vector of size n. Then the
matrix-vector product ofAwithuis the linear combination
Au= [u]1A1+ [u]2A2+ [u]3A3++ [u]nAn
(This denition contains Notation MVP.) 4
So, the matrix-vector product is yet another version of \multiplication," at least in the sense that we
have yet again overloaded juxtaposition of two symbols as our notation. Remember your objects, an mn
matrix times a vector of size nwill create a vector of size m. So ifAis rectangular, then the size of the
vector changes. With all the linear combinations we have performed so far, this computation should now
seem second nature.
Example MTV
A matrix times a vector
Consider
A=2
41 4 2 3 4
3 2 0 1 2
1 6 3 1 53
5 u=2
666642
1
2
3
13
77775
Then
Au= 22
41
3
13
5+ 12
44
2
63
5+ ( 2)2
42
0
33
5+ 32
43
1
13
5+ ( 1)2
44
2
53
5=2
47
1
63
5:
We can now represent systems of linear equations compactly with a matrix-vector product (Denition
MVP [223]) and column vector equality (Denition CVE [98]). This nally yields a very popular alternative
to our unconventional LS(A;b) notation.
Version 2.30
226 Section MM Matrix Multiplication
Theorem SLEMM
Systems of Linear Equations as Matrix Multiplication
The set of solutions to the linear system LS(A;b) equals the set of solutions for xin the vector equation
Ax=b.
Proof This theorem says that two sets (of solutions) are equal. So we need to show that one set of
solutions is a subset of the other, and vice versa (Denition SE [762]). Let A1;A2;A3; :::; Anbe the
columns of A. Both of these set inclusions then follow from the following chain of equivalences (Technique
E [768]),
xis a solution toLS(A;b)
() [x]1A1+ [x]2A2+ [x]3A3++ [x]nAn=b Theorem SLSLC [112]
()xis a solution to Ax=b Denition MVP [223]
Example MNSLE
Matrix notation for systems of linear equations
Consider the system of linear equations from Example NSLE [29].
2x1+ 4x2 3x3+ 5x4+x5= 9
3x1+x2+x4 3x5= 0
2x1+ 7x2 5x3+ 2x4+ 2x5= 3
has coecient matrix
A=2
42 4 3 5 1
3 1 0 1 3
2 7 5 2 23
5
and vector of constants
b=2
49
0
33
5
and so will be described compactly by the vector equation Ax=b.
The matrix-vector product is a very natural computation. We have motivated it by its connections
with systems of equations, but here is a another example.
Example MBC
Money's best cities
Every year Money magazine selects several cities in the United States as the \best" cities to live in, based
on a wide array of statistics about each city. This is an example of how the editors of Money might arrive
at a single number that consolidates the statistics about a city. We will analyze Los Angeles, Chicago and
New York City, based on four criteria: average high temperature in July (Farenheit), number of colleges
and universities in a 30-mile radius, number of toxic waste sites in the Superfund environmental clean-up
program and a personal crime index based on FBI statistics (average = 100, smaller is safer). It should be
apparent how to generalize the example to a greater number of cities and a greater number of statistics.
We begin by building a table of statistics. The rows will be labeled with the cities, and the columns
with statistical categories. These values are from Money 's website in early 2005.
City Temp Colleges Superfund Crime
Los Angeles 77 28 93 254
Chicago 84 38 85 363
New York 84 99 1 193
Version 2.30
Subsection MM.MVP Matrix-Vector Product 227
Conceivably these data might reside in a spreadsheet. Now we must combine the statistics for each city.
We could accomplish this by weighting each category, scaling the values and summing them. The sizes of
the weights would depend upon the numerical size of each statistic generally, but more importantly, they
would re
ect the editors opinions or beliefs about which statistics were most important to their readers. Is
the crime index more important than the number of colleges and universities? Of course, there is no right
answer to this question.
Suppose the editors nally decide on the following weights to employ: temperature, 0 :23; colleges, 0 :46;
Superfund, 0:05; crime, 0:20. Notice how negative weights are used for undesirable statistics. Then,
for example, the editors would compute for Los Angeles,
(0:23)(77) + (0 :46)(28) + ( 0:05)(93) + ( 0:20)(254) = 24:86
This computation might remind you of an inner product, but we will produce the computations for all of
the cities as a matrix-vector product. Write the table of raw statistics as a matrix
T=2
477 28 93 254
84 38 85 363
84 99 1 1933
5
and the weights as a vector
w=2
6640:23
0:46
0:05
0:203
775
then the matrix-vector product (Denition MVP [223]) yields
Tw= (0:23)2
477
84
843
5+ (0:46)2
428
38
993
5+ ( 0:05)2
493
85
13
5+ ( 0:20)2
4254
363
1933
5=2
4 24:86
40:05
26:213
5
This vector contains a single number for each of the cities being studied, so the editors would rank New
York best (26 :21), Los Angeles next ( 24:86), and Chicago third ( 40:05). Of course, the mayor's oces
in Chicago and Los Angeles are free to counter with a dierent set of weights that cause their city to be
ranked best. These alternative weights would be chosen to play to each cities' strengths, and minimize
their problem areas.
If a speadsheet were used to make these computations, a row of weights would be entered somewhere
near the table of data and the formulas in the spreadsheet would eect a matrix-vector product. This
example is meant to illustrate how \linear" computations (addition, multiplication) can be organized as a
matrix-vector product.
Another example would be the matrix of numerical scores on examinations and exercises for students in
a class. The rows would correspond to students and the columns to exams and assignments. The instructor
could then assign weights to the dierent exams and assignments, and via a matrix-vector product, compute
a single score for each student.
Later (much later) we will need the following theorem, which is really a technical lemma (see Technique
LC [774]). Since we are in a position to prove it now, we will. But you can safely skip it for the moment,
if you promise to come back later to study the proof when the theorem is employed. At that point you
will also be able to understand the comments in the paragraph following the proof.
Theorem EMMVP
Equal Matrices and Matrix-Vector Products
Suppose that AandBaremnmatrices such that Ax=Bxfor every x2Cn. ThenA=B.
Proof We are assuming Ax=Bxfor all x2Cn, so we can employ this equality for anychoice of
the vector x. However, we'll limit our use of this equality to the standard unit vectors, ej, 1jn
Version 2.30
228 Section MM Matrix Multiplication
(Denition SUV [197]). For all 1 jn, 1im,
[A]ij= 0 [A]i1++ 0 [A]i;j 1+ 1 [A]ij+ 0 [A]i;j+1++ 0 [A]in
= [A]i1[ej]1+ [A]i2[ej]2+ [A]i3[ej]3++ [A]in[ej]nDenition SUV [197]
= [Aej]iDenition MVP [223]
= [Bej]iDenition CVE [98]
= [B]i1[ej]1+ [B]i2[ej]2+ [B]i3[ej]3++ [B]in[ej]nDenition MVP [223]
= 0 [B]i1++ 0 [B]i;j 1+ 1 [B]ij+ 0 [B]i;j+1++ 0 [B]in Denition SUV [197]
= [B]ij
So by Denition ME [207] the matrices AandBare equal, as desired.
You might notice from studying the proof that the hypotheses of this theorem could be \weakened"
(i.e. made less restrictive). We need only suppose the equality of the matrix-vector products for just the
standard unit vectors (Denition SUV [197]) or any other spanning set (Denition TSVS [356]) of Cn
(Exercise LISS.T40 [363]). However, in practice, when we apply this theorem the stronger hypothesis will
be in eect so this version of the theorem will suce for our purposes. (If we changed the statement of the
theorem to have the less restrictive hypothesis, then we would call the theorem \stronger.")
Subsection MM
Matrix Multiplication
We now dene how to multiply two matrices together. Stop for a minute and think about how you might
dene this new operation.
Many books would present this denition much earlier in the course. However, we have taken great
care to delay it as long as possible and to present as many ideas as practical based mostly on the notion
of linear combinations. Towards the conclusion of the course, or when you perhaps take a second course
in linear algebra, you may be in a position to appreciate the reasons for this. For now, understand that
matrix multiplication is a central denition and perhaps you will appreciate its importance more by having
saved it for later.
Denition MM
Matrix Multiplication
SupposeAis anmnmatrix and Bis annpmatrix with columns B1;B2;B3; :::; Bp. Then the
matrix product ofAwithBis thempmatrix where column iis the matrix-vector product ABi.
Symbolically,
AB=A[B1jB2jB3j:::jBp] = [AB1jAB2jAB3j:::jABp]:
4
Example PTM
Product of two matrices
Set
A=2
41 2 1 4 6
0 4 1 2 3
5 1 2 3 43
5 B=2
666641 6 2 1
1 4 3 2
1 1 2 3
6 4 1 2
1 2 3 03
77775
Version 2.30
Subsection MM.MMEE Matrix Multiplication, Entry-by-Entry 229
Then
AB=2
66664A2
666641
1
1
6
13
77775A2
666646
4
1
4
23
77775A2
666642
3
2
1
33
77775A2
666641
2
3
2
03
777753
77775=2
428 17 20 10
20 13 3 1
18 44 12 33
5:
Is this the denition of matrix multiplication you expected? Perhaps our previous operations for
matrices caused you to think that we might multiply two matrices of the same size, entry-by-entry ? Notice
that our current denition uses matrices of dierent sizes (though the number of columns in the rst must
equal the number of rows in the second), and the result is of a third size. Notice too in the previous
example that we cannot even consider the product BA, since the sizes of the two matrices in this order
aren't right.
But it gets weirder than that. Many of your old ideas about \multiplication" won't apply to matrix
multiplication, but some still will. So make no assumptions, and don't do anything until you have a
theorem that says you can. Even if the sizes are right, matrix multiplication is not commutative | order
matters.
Example MMNC
Matrix multiplication is not commutative
Set
A=1 3
1 2
B=4 0
5 1
:
Then we have two square, 2 2 matrices, so Denition MM [226] allows us to multiply them in either
order. We nd
AB=19 3
6 2
BA=4 12
4 17
andAB6=BA. Not even close. It should not be hard for you to construct other pairs of matrices that do
not commute (try a couple of 3 3's). Can you nd a pair of non-identical matrices that docommute?
Matrix multiplication is fundamental, so it is a natural procedure for any computational device. See:
Computation MM.MMA [749]
Subsection MMEE
Matrix Multiplication, Entry-by-Entry
While certain \natural" properties of multiplication don't hold, many more do. In the next subsection,
we'll state and prove the relevant theorems. But rst, we need a theorem that provides an alternate means
of multiplying two matrices. In many texts, this would be given as the denition of matrix multiplication.
We prefer to turn it around and have the following formula as a consequence of our denition. It will prove
useful for proofs of matrix equality, where we need to examine products of matrices, entry-by-entry.
Theorem EMP
Entries of Matrix Products
SupposeAis anmnmatrix and Bis annpmatrix. Then for 1 im, 1jp, the individual
entries ofABare given by
[AB]ij= [A]i1[B]1j+ [A]i2[B]2j+ [A]i3[B]3j++ [A]in[B]nj
Version 2.30
230 Section MM Matrix Multiplication
=nX
k=1[A]ik[B]kj
Proof Denote the columns of Aas the vectors A1;A2;A3; :::; Anand the columns of Bas the vectors
B1;B2;B3; :::; Bp. Then for 1im, 1jp,
[AB]ij= [ABj]iDenition MM [226]
=
[Bj]1A1+ [Bj]2A2+ [Bj]3A3++ [Bj]nAn
iDenition MVP [223]
=
[Bj]1A1
i+
[Bj]2A2
i+
[Bj]3A3
i++
[Bj]nAn
iDenition CVA [98]
= [Bj]1[A1]i+ [Bj]2[A2]i+ [Bj]3[A3]i++ [Bj]n[An]i Denition CVSM [99]
= [B]1j[A]i1+ [B]2j[A]i2+ [B]3j[A]i3++ [B]nj[A]in Notation ME [207]
= [A]i1[B]1j+ [A]i2[B]2j+ [A]i3[B]3j++ [A]in[B]nj Property CMCN [758]
=nX
k=1[A]ik[B]kj
Example PTMEE
Product of two matrices, entry-by-entry
Consider again the two matrices from Example PTM [226]
A=2
41 2 1 4 6
0 4 1 2 3
5 1 2 3 43
5 B=2
666641 6 2 1
1 4 3 2
1 1 2 3
6 4 1 2
1 2 3 03
77775
Then suppose we just wanted the entry of ABin the second row, third column:
[AB]23= [A]21[B]13+ [A]22[B]23+ [A]23[B]33+ [A]24[B]43+ [A]25[B]53
=(0)(2) + ( 4)(3) + (1)(2) + (2)( 1) + (3)(3) = 3
Notice how there are 5 terms in the sum, since 5 is the common dimension of the two matrices (column
count forA, row count for B). In the conclusion of Theorem EMP [227], it would be the index kthat
would run from 1 to 5 in this computation. Here's a bit more practice.
The entry of third row, rst column:
[AB]31= [A]31[B]11+ [A]32[B]21+ [A]33[B]31+ [A]34[B]41+ [A]35[B]51
=( 5)(1) + (1)( 1) + (2)(1) + ( 3)(6) + (4)(1) = 18
To get some more practice on your own, complete the computation of the other 10 entries of this product.
Construct some other pairs of matrices (of compatible sizes) and compute their product two ways. First
use Denition MM [226]. Since linear combinations are straightforward for you now, this should be easy
to do and to do correctly. Then do it again, using Theorem EMP [227]. Since this process may take some
practice, use your rst computation to check your work.
Theorem EMP [227] is the way many people compute matrix products by hand. It will also be very
useful for the theorems we are going to prove shortly. However, the denition (Denition MM [226]) is
frequently the most useful for its connections with deeper ideas like the null space and the upcoming
column space.
Version 2.30
Subsection MM.PMM Properties of Matrix Multiplication 231
Subsection PMM
Properties of Matrix Multiplication
In this subsection, we collect properties of matrix multiplication and its interaction with the zero matrix
(Denition ZM [210]), the identity matrix (Denition IM [84]), matrix addition (Denition MA [207]),
scalar matrix multiplication (Denition MSM [208]), the inner product (Denition IP [192]), conjugation
(Theorem MMCC [232]), and the transpose (Denition TM [210]). Whew! Here we go. These are great
proofs to practice with, so try to concoct the proofs before reading them, they'll get progressively more
complicated as we go.
Theorem MMZM
Matrix Multiplication and the Zero Matrix
SupposeAis anmnmatrix. Then
1.AOnp=Omp
2.OpmA=Opn
Proof We'll prove (1) and leave (2) to you. Entry-by-entry, for 1 im, 1jp,
[AOnp]ij=nX
k=1[A]ik[Onp]kjTheorem EMP [227]
=nX
k=1[A]ik0 Denition ZM [210]
=nX
k=10
= 0 Property ZCN [759]
= [Omp]ijDenition ZM [210]
So by the denition of matrix equality (Denition ME [207]), the matrices AOnpandOmpare equal.
Theorem MMIM
Matrix Multiplication and Identity Matrix
SupposeAis anmnmatrix. Then
1.AIn=A
2.ImA=A
Proof Again, we'll prove (1) and leave (2) to you. Entry-by-entry, For 1 im, 1jn,
[AIn]ij=nX
k=1[A]ik[In]kj Theorem EMP [227]
= [A]ij[In]jj+nX
k=1
k6=j[A]ik[In]kj Property CACN [758]
= [A]ij(1) +nX
k=1;k6=j[A]ik(0) Denition IM [84]
= [A]ij+nX
k=1;k6=j0
= [A]ij
Version 2.30
232 Section MM Matrix Multiplication
So the matrices AandAInare equal, entry-by-entry, and by the denition of matrix equality (Denition
ME [207]) we can say they are equal matrices.
It is this theorem that gives the identity matrix its name. It is a matrix that behaves with matrix
multiplication like the scalar 1 does with scalar multiplication. To multiply by the identity matrix is to
have no eect on the other matrix.
Theorem MMDAA
Matrix Multiplication Distributes Across Addition
SupposeAis anmnmatrix and BandCarenpmatrices and Dis apsmatrix. Then
1.A(B+C) =AB+AC
2. (B+C)D=BD+CD
Proof We'll do (1), you do (2). Entry-by-entry, for 1 im, 1jp,
[A(B+C)]ij=nX
k=1[A]ik[B+C]kj Theorem EMP [227]
=nX
k=1[A]ik([B]kj+ [C]kj) Denition MA [207]
=nX
k=1[A]ik[B]kj+ [A]ik[C]kj Property DCN [759]
=nX
k=1[A]ik[B]kj+nX
k=1[A]ik[C]kj Property CACN [758]
= [AB]ij+ [AC]ij Theorem EMP [227]
= [AB+AC]ij Denition MA [207]
So the matrices A(B+C) andAB+ACare equal, entry-by-entry, and by the denition of matrix equality
(Denition ME [207]) we can say they are equal matrices.
Theorem MMSMM
Matrix Multiplication and Scalar Matrix Multiplication
SupposeAis anmnmatrix and Bis annpmatrix. Let be a scalar. Then (AB) = (A)B=A(B).
Proof These are equalities of matrices. We'll do the rst one, the second is similar and will be good
practice for you. For 1 im, 1jp,
[(AB)]ij=[AB]ij Denition MSM [208]
=nX
k=1[A]ik[B]kj Theorem EMP [227]
=nX
k=1[A]ik[B]kj Property DCN [759]
=nX
k=1[A]ik[B]kj Denition MSM [208]
= [(A)B]ij Theorem EMP [227]
So the matrices (AB) and (A)Bare equal, entry-by-entry, and by the denition of matrix equality
(Denition ME [207]) we can say they are equal matrices.
Theorem MMA
Version 2.30
Subsection MM.PMM Properties of Matrix Multiplication 233
Matrix Multiplication is Associative
SupposeAis anmnmatrix,Bis annpmatrix and Dis apsmatrix. Then A(BD) = (AB)D.
Proof A matrix equality, so we'll go entry-by-entry, no surprise there. For 1 im, 1js,
[A(BD)]ij=nX
k=1[A]ik[BD]kj Theorem EMP [227]
=nX
k=1[A]ik pX
`=1[B]k`[D]`j!
Theorem EMP [227]
=nX
k=1pX
`=1[A]ik[B]k`[D]`j Property DCN [759]
We can switch the order of the summation since these are nite sums,
=pX
`=1nX
k=1[A]ik[B]k`[D]`j Property CACN [758]
As [D]`jdoes not depend on the index k, we can use distributivity to move it outside of the inner sum,
=pX
`=1[D]`j nX
k=1[A]ik[B]k`!
Property DCN [759]
=pX
`=1[D]`j[AB]i` Theorem EMP [227]
=pX
`=1[AB]i`[D]`j Property CMCN [758]
= [(AB)D]ij Theorem EMP [227]
So the matrices ( AB)DandA(BD) are equal, entry-by-entry, and by the denition of matrix equality
(Denition ME [207]) we can say they are equal matrices.
The statement of our next theorem is technically inaccurate. If we upgrade the vectors u;vto matrices
with a single column, then the expression utvis a 11 matrix, though we will treat this small matrix as
if it was simply the scalar quantity in its lone entry. When we apply Theorem MMIP [231] there should
not be any confusion.
Theorem MMIP
Matrix Multiplication and Inner Products
If we consider the vectors u;v2Cmasm1 matrices then
hu;vi=utv
Proof
hu;vi=mX
k=1[u]k[v]k Denition IP [192]
=mX
k=1[u]k1[v]k1 Column vectors as matrices
Version 2.30
234 Section MM Matrix Multiplication
=mX
k=1
ut
1k[v]k1 Denition TM [210]
=mX
k=1
ut
1k[v]k1 Denition CCCV [191]
=
utv
11Theorem EMP [227]
To nish we just blur the distinction between a 1 1 matrix ( utv) and its lone entry.
Theorem MMCC
Matrix Multiplication and Complex Conjugation
SupposeAis anmnmatrix and Bis annpmatrix. Then AB=AB.
Proof To obtain this matrix equality, we will work entry-by-entry. For 1 im, 1jp,
AB
ij=[AB]ij Denition CCM [212]
=nX
k=1[A]ik[B]kj Theorem EMP [227]
=nX
k=1[A]ik[B]kj Theorem CCRA [759]
=nX
k=1[A]ik[B]kj Theorem CCRM [760]
=nX
k=1
A
ik
B
kjDenition CCM [212]
=
AB
ijTheorem EMP [227]
So the matrices ABandABare equal, entry-by-entry, and by the denition of matrix equality (Denition
ME [207]) we can say they are equal matrices.
Another theorem in this style, and its a good one. If you've been practicing with the previous proofs
you should be able to do this one yourself.
Theorem MMT
Matrix Multiplication and Transposes
SupposeAis anmnmatrix and Bis annpmatrix. Then ( AB)t=BtAt.
Proof This theorem may be surprising but if we check the sizes of the matrices involved, then maybe
it will not seem so far-fetched. First, ABhas sizemp, so its transpose has size pm. The product
ofBtwithAtis apnmatrix times an nmmatrix, also resulting in a pmmatrix. So at least
our objects are compatible for equality (and would not be, in general, if we didn't reverse the order of the
matrix multiplication).
Here we go again, entry-by-entry. For 1 im, 1jp,
(AB)t
ji= [AB]ij Denition TM [210]
=nX
k=1[A]ik[B]kj Theorem EMP [227]
=nX
k=1[B]kj[A]ik Property CMCN [758]
Version 2.30
Subsection MM.HM Hermitian Matrices 235
=nX
k=1
Bt
jk
At
kiDenition TM [210]
=
BtAt
jiTheorem EMP [227]
So the matrices ( AB)tandBtAtare equal, entry-by-entry, and by the denition of matrix equality (De-
nition ME [207]) we can say they are equal matrices.
This theorem seems odd at rst glance, since we have to switch the order of AandB. But if we simply
consider the sizes of the matrices involved, we can see that the switch is necessary for this reason alone.
That the individual entries of the products then come along to be equal is a bonus.
As the adjoint of a matrix is a composition of a conjugate and a transpose, its interaction with matrix
multiplication is similar to that of a transpose. Here's the last of our long list of basic properties of matrix
multiplication.
Theorem MMAD
Matrix Multiplication and Adjoints
SupposeAis anmnmatrix and Bis annpmatrix. Then ( AB)=BA.
Proof
(AB)=
ABtDenition A [214]
=
ABtTheorem MMCC [232]
=
Bt
AtTheorem MMT [232]
=BADenition A [214]
Notice how none of these proofs above relied on writing out huge general matrices with lots of ellipses
(\. . . ") and trying to formulate the equalities a whole matrix at a time. This messy business is a \proof
technique" to be avoided at all costs. Notice too how the proof of Theorem MMAD [233] does not use an
entry-by-entry approach, but simply builds on previous results about matrix multiplication's interaction
with conjugation and transposes.
These theorems, along with Theorem VSPM [209] and the other results in Section MO [207], give you
the \rules" for how matrices interact with the various operations we have dened on matrices (addition,
scalar multiplication, matrix multiplication, conjugation, transposes and adjoints). Use them and use them
often. But don't try to do anything with a matrix that you don't have a rule for. Together, we would
informally call all these operations, and the attendant theorems, \the algebra of matrices." Notice, too,
that every column vector is just a n1 matrix, so these theorems apply to column vectors also. Finally,
these results, taken as a whole, may make us feel that the denition of matrix multiplication is not so
unnatural.
Subsection HM
Hermitian Matrices
The adjoint of a matrix has a basic property when employed in a matrix-vector product as part of an inner
product. At this point, you could even use the following result as a motivation for the denition of an
adjoint.
Theorem AIP
Adjoint and Inner Product
Version 2.30
236 Section MM Matrix Multiplication
Suppose that Ais anmnmatrix and x2Cn,y2Cm. ThenhAx;yi=hx; Ayi.
Proof
hAx;yi= (Ax)ty Theorem MMIP [231]
=xtAty Theorem MMT [232]
=xt
At
y Theorem CCM [213]
=xt
At
y Theorem MCT [214]
=xt(A)y Denition A [214]
=xt(Ay) Theorem MMCC [232]
=hx; Ayi Theorem MMIP [231]
Sometimes a matrix is equal to its adjoint (Denition A [214]), and these matrices have interesting
properties. One of the most common situations where this occurs is when a matrix has only real number
entries. Then we are simply talking about symmetric matrices (Denition SYM [211]), so you can view
this as a generalization of a symmetric matrix.
Denition HM
Hermitian Matrix
The square matrix AisHermitian (orself-adjoint ) ifA=A. 4
Again, the set of real matrices that are Hermitian is exactly the set of symmetric matrices. In Section
PEE [479] we will uncover some amazing properties of Hermitian matrices, so when you get there, run
back here to remind yourself of this denition. Further properties will also appear in various sections of the
Topics (Part T [873]). Right now we prove a fundamental result about Hermitian matrices, matrix vector
products and inner products. As a characterization, this could be employed as a denition of a Hermitian
matrix and some authors take this approach.
Theorem HMIP
Hermitian Matrices and Inner Products
Suppose that Ais a square matrix of size n. ThenAis Hermitian if and only if hAx;yi=hx; Ayifor all
x;y2Cn.
Proof ()) This is the \easy half" of the proof, and makes the rationale for a denition of Hermitian
matrices most obvious. Assume Ais Hermitian,
hAx;yi=hx; Ayi Theorem AIP [233]
=hx; Ayi Denition HM [234]
(() This \half" will take a bit more work. Assume that hAx;yi=hx; Ayifor all x;y2Cn. Choose any
x2Cn. We want to show that A=Aby establishing that Ax=Ax. With only this much motivation,
consider the inner product,
hAx Ax; Ax Axi=hAx Ax; Axi hAx Ax; Axi Theorem IPVA [193]
=hAx Ax; Axi hA(Ax Ax);xi Theorem AIP [233]
=hAx Ax; Axi hAx Ax; Axi Hypothesis
= 0 Property AICN [759]
Version 2.30
Subsection MM.READ Reading Questions 237
Because this rst inner product equals zero, and has the same vector in each argument ( Ax Ax),
Theorem PIP [196] gives the conclusion that Ax Ax=0. WithAx=Axfor all x2Cn, Theorem
EMMVP [225] says A=A, which is the dening property of a Hermitian matrix (Denition HM [234]).
So, informally, Hermitian matrices are those that can be tossed around from one side of an inner
product to the other with reckless abandon. We'll see later what this buys us.
Subsection READ
Reading Questions
1. Form the matrix vector product of
2
42 3 1 0
1 2 7 3
1 5 3 23
5 with2
6642
3
0
53
775
2. Multiply together the two matrices below (in the order given).
2
42 3 1 0
1 2 7 3
1 5 3 23
52
6642 6
3 4
0 2
3 13
775
3. Rewrite the system of linear equations below as a vector equality and using a matrix-vector product.
(This question does not ask for a solution to the system. But it does ask you to express the system
of equations in a new form using tools from this section.)
2x1+ 3x2 x3= 0
x1+ 2x2+x3= 3
x1+ 3x2+ 3x3= 7
Version 2.30
238 Section MM Matrix Multiplication
Subsection EXC
Exercises
C20 Compute the product of the two matrices below, AB. Do this using the denitions of the matrix-
vector product (Denition MVP [223]) and the denition of matrix multiplication (Denition MM [226]).
A=2
42 5
1 3
2 23
5 B=1 5 3 4
2 0 2 3
Contributed by Robert Beezer Solution [239]
C21 Compute the product ABof the two matrices below using both the denition of the matrix-vector
product (Denition MVP [223]) and the denition of matrix multiplication (Denition MM [226]).
A=2
41 3 2
1 2 1
0 1 03
5 B=2
44 1 2
1 0 1
3 1 53
5
Contributed by Chris Black Solution [239]
C22 Compute the product ABof the two matrices below using both the denition of the matrix-vector
product (Denition MVP [223]) and the denition of matrix multiplication (Denition MM [226]).
A=1 0
2 1
B=2 3
4 6
Contributed by Chris Black Solution [239]
C23 Compute the product ABof the two matrices below using both the denition of the matrix-vector
product (Denition MVP [223]) and the denition of matrix multiplication (Denition MM [226]).
A=2
6643 1
2 4
6 5
1 23
775B= 3 1
4 2
Contributed by Chris Black Solution [239]
C24 Compute the product ABof the two matrices below.
A=2
41 2 3 2
0 1 2 1
1 1 3 13
5 B=2
6643
4
0
23
775
Contributed by Chris Black Solution [239]
C25 Compute the product ABof the two matrices below.
A=2
41 2 3 2
0 1 2 1
1 1 3 13
5 B=2
664 7
3
1
13
775
Version 2.30
Subsection MM.EXC Exercises 239
Contributed by Chris Black Solution [239]
C26 Compute the product ABof the two matrices below using both the denition of the matrix-vector
product (Denition MVP [223]) and the denition of matrix multiplication (Denition MM [226]).
A=2
41 3 1
0 1 0
1 1 23
5 B=2
42 5 1
0 1 0
1 2 13
5
Contributed by Chris Black Solution [239]
C30 For the matrix A=1 2
0 1
, ndA2,A3,A4. Find a general formula for Anfor any positive integer
n.
Contributed by Chris Black Solution [239]
C31 For the matrix A=1 1
0 1
, ndA2,A3,A4. Find a general formula for Anfor any positive integer
n.
Contributed by Chris Black Solution [240]
C32 For the matrix A=2
41 0 0
0 2 0
0 0 33
5, ndA2,A3,A4. Find a general formula for Anfor any positive
integern.
Contributed by Chris Black Solution [240]
C33 For the matrix A=2
40 1 2
0 0 1
0 0 03
5, ndA2,A3,A4. Find a general formula for Anfor any positive
integern.
Contributed by Chris Black Solution [240]
T10 Suppose that Ais a square matrix and there is a vector, b, such thatLS(A;b) has a unique solution.
Prove that Ais nonsingular. Give a direct proof (perhaps appealing to Theorem PSPHS [124]) rather than
just negating a sentence from the text discussing a similar situation.
Contributed by Robert Beezer Solution [240]
T20 Prove the second part of Theorem MMZM [229].
Contributed by Robert Beezer
T21 Prove the second part of Theorem MMIM [229].
Contributed by Robert Beezer
T22 Prove the second part of Theorem MMDAA [230].
Contributed by Robert Beezer
T23 Prove the second part of Theorem MMSMM [230].
Contributed by Robert Beezer Solution [240]
T31 Suppose that Ais anmnmatrix and x;y2N(A). Prove that x+y2N(A).
Contributed by Robert Beezer
T32 Suppose that Ais anmnmatrix,2C, and x2N(A). Prove that x2N(A).
Contributed by Robert Beezer
Version 2.30
240 Section MM Matrix Multiplication
T40 Suppose that Ais anmnmatrix and Bis annpmatrix. Prove that the null space of Bis a
subset of the null space of AB, that isN(B)N(AB). Provide an example where the opposite is false,
in other words give an example where N(AB)6N(B).
Contributed by Robert Beezer Solution [240]
T41 Suppose that Ais annnnonsingular matrix and Bis annpmatrix. Prove that the null space
ofBis equal to the null space of AB, that isN(B) =N(AB). (Compare with Exercise MM.T40 [238].)
Contributed by Robert Beezer Solution [241]
T50 Suppose uandvare any two solutions of the linear system LS(A;b). Prove that u vis an
element of the null space of A, that is, u v2N(A).
Contributed by Robert Beezer
T51 Give a new proof of Theorem PSPHS [124] replacing applications of Theorem SLSLC [112] with
matrix-vector products (Theorem SLEMM [224]).
Contributed by Robert Beezer Solution [241]
T52 Suppose that x;y2Cn,b2CmandAis anmnmatrix. If x,yandx+yare each a solution to
the linear system LS(A;b), what interesting can you say about b? Form an implication with the existence
of the three solutions as the hypothesis and an interesting statement about LS(A;b) as the conclusion,
and then give a proof.
Contributed by Robert Beezer Solution [241]
Version 2.30
Subsection MM.SOL Solutions 241
Subsection SOL
Solutions
C20 Contributed by Robert Beezer Statement [236]
By Denition MM [226],
AB=2
42
42 5
1 3
2 23
51
22
42 5
1 3
2 23
55
02
42 5
1 3
2 23
5 3
22
42 5
1 3
2 23
54
23
5
Repeated applications of Denition MVP [223] give
=2
412
42
1
23
5+ 22
45
3
23
552
42
1
23
5+ 02
45
3
23
5 32
42
1
23
5+ 22
45
3
23
542
42
1
23
5+ ( 3)2
45
3
23
53
5
=2
412 10 4 7
5 5 9 13
2 10 10 143
5
C21 Contributed by Chris Black Statement [236]
AB=2
413 3 15
1 0 5
1 0 13
5.
C22 Contributed by Chris Black Statement [236]
AB=2 3
0 0
.
C23 Contributed by Chris Black Statement [236]
AB=2
664 5 5
10 10
2 16
5 53
775.
C24 Contributed by Chris Black Statement [236]
AB=2
47
2
93
5.
C25 Contributed by Chris Black Statement [236]
AB=2
40
0
03
5.
C26 Contributed by Chris Black Statement [237]
AB=2
41 0 0
0 1 0
0 0 13
5.
C30 Contributed by Chris Black Statement [237]
A2=1 4
0 1
,A3=1 6
0 1
,A4=1 8
0 1
. From this pattern, we see that An=1 2n
0 1
.
Version 2.30
242 Section MM Matrix Multiplication
C31 Contributed by Chris Black Statement [237]
A2=1 2
0 1
,A3=1 3
0 1
,A4=1 4
0 1
. From this pattern, we see that An=1 n
0 1
.
C32 Contributed by Chris Black Statement [237]
A2=2
41 0 0
0 4 0
0 0 93
5,A3=2
41 0 0
0 8 0
0 0 273
5, andA4=2
41 0 0
0 16 0
0 0 813
5. The pattern emerges, and we see that
An=2
41 0 0
0 2n0
0 0 3n3
5.
C33 Contributed by Chris Black Statement [237]
We quickly compute A2=2
40 0 1
0 0 0
0 0 03
5, and we then see that A3and all subsequent powers of Aare the
33 zero matrix; that is, An=O3;3forn3.
T10 Contributed by Robert Beezer Statement [237]
SinceLS(A; b) has at least one solution, we can apply Theorem PSPHS [124]. Because the solution is
assumed to be unique, the null space of Amust be trivial. Then Theorem NMTNS [86] implies that Ais
nonsingular.
The converse of this statement is a trivial application of Theorem NMUS [86]. That said, we could
extend our NSMxx series of theorems with an added equivalence for nonsingularity, \Given a single vector
of constants, b, the systemLS(A;b) has a unique solution."
T23 Contributed by Robert Beezer Statement [237]
We'll run the proof entry-by-entry.
[(AB)]ij=[AB]ij Denition MSM [208]
=nX
k=1[A]ik[B]kj Theorem EMP [227]
=nX
k=1[A]ik[B]kj Distributivity in C
=nX
k=1[A]ik[B]kj Commutativity in C
=nX
k=1[A]ik[B]kj Denition MSM [208]
= [A(B)]ij Theorem EMP [227]
So the matrices (AB) andA(B) are equal, entry-by-entry, and by the denition of matrix equality
(Denition ME [207]) we can say they are equal matrices.
T40 Contributed by Robert Beezer Statement [238]
To prove that one set is a subset of another, we start with an element of the smaller set and see if we can
determine that it is a member of the larger set (Denition SSET [761]). Suppose x2N(B). Then we
know thatBx=0by Denition NSM [73]. Consider
(AB)x=A(Bx) Theorem MMA [231]
=A0 Hypothesis
=0 Theorem MMZM [229]
Version 2.30
Subsection MM.SOL Solutions 243
This establishes that x2N(AB), soN(B)N(AB).
To show that the inclusion does not hold in the opposite direction, choose Bto be any nonsingular
matrix of size n. ThenN(B) =f0gby Theorem NMTNS [86]. Let Abe the square zero matrix, O, of
the same size. Then AB=OB=Oby Theorem MMZM [229] and therefore N(AB) =Cn, and is nota
subset ofN(B) =f0g.
T41 Contributed by David Braithwaite Statement [238]
From the solution to Exercise MM.T40 [238] we know that N(B)N (AB). So to establish the set
equality (Denition SE [762]) we need to show that N(AB)N(B).
Suppose x2N(AB). Then we know that ABx=0by Denition NSM [73]. Consider
0= (AB)x Denition NSM [73]
=A(Bx) Theorem MMA [231]
So,Bx2N(A). Because Ais nonsingular, it has a trivial null space (Theorem NMTNS [86]) and we
conclude that Bx=0. This establishes that x2N (B), soN(AB)N (B) and combined with the
solution to Exercise MM.T40 [238] we have N(B) =N(AB) whenAis nonsingular.
T51 Contributed by Robert Beezer Statement [238]
We will work with the vector equality representations of the relevant systems of equations, as described by
Theorem SLEMM [224].
(() Suppose y=w+zandz2N(A). Then
Ay=A(w+z) Substitution
=Aw+Az Theorem MMDAA [230]
=b+0 z 2N(A)
=b Property ZC [100]
demonstrating that yis a solution.
()) Suppose yis a solution toLS(A; b). Then
A(y w) =Ay Aw Theorem MMDAA [230]
=b b y ;wsolutions to Ax=b
=0 Property AIC [100]
which says that y w2N(A). In other words, y w=zfor some vector z2N(A). Rewritten, this is
y=w+z, as desired.
T52 Contributed by Robert Beezer Statement [238]
LS(A;b) must be homogeneous. To see this consider that
b=Ax Theorem SLEMM [224]
=Ax+0 Property ZC [100]
=Ax+Ay Ay Property AIC [100]
=A(x+y) Ay Theorem MMDAA [230]
=b b Theorem SLEMM [224]
=0 Property AIC [100]
By Denition HS [71] we see that LS(A;b) is homogeneous.
Version 2.30
244 Section MM Matrix Multiplication
Version 2.30
Section MISLE Matrix Inverses and Systems of Linear Equations 245
Section MISLE
Matrix Inverses and Systems of Linear Equations
We begin with a familiar example, performed in a novel way.
Example SABMI
Solutions to Archetype B with a matrix inverse
Archetype B [786] is the system of m= 3 linear equations in n= 3 variables,
7x1 6x2 12x3= 33
5x1+ 5x2+ 7x3= 24
x1+ 4x3= 5
By Theorem SLEMM [224] we can represent this system of equations as
Ax=b
where
A=2
4 7 6 12
5 5 7
1 0 43
5 x=2
4x1
x2
x33
5 b=2
4 33
24
53
5
We'll pull a rabbit out of our hat and present the 3 3 matrixB,
B=2
4 10 12 9
13
2811
25
235
23
5
and note that
BA=2
4 10 12 9
13
2811
25
235
23
52
4 7 6 12
5 5 7
1 0 43
5=2
41 0 0
0 1 0
0 0 13
5
Now apply this computation to the problem of solving the system of equations,
x=I3x Theorem MMIM [229]
= (BA)x Substitution
=B(Ax) Theorem MMA [231]
=Bb Substitution
So we have
x=Bb=2
4 10 12 9
13
2811
25
235
23
52
4 33
24
53
5=2
4 3
5
23
5
So with the help and assistance of Bwe have been able to determine a solution to the system represented
byAx=bthrough judicious use of matrix multiplication. We know by Theorem NMUS [86] that since
the coecient matrix in this example is nonsingular, there would be a unique solution, no matter what the
choice of b. The derivation above amplies this result, since we were forced to conclude that x=Bband
Version 2.30
246 Section MISLE Matrix Inverses and Systems of Linear Equations
the solution couldn't be anything else. You should notice that this argument would hold for any particular
value of b.
The matrix Bof the previous example is called the inverse of A. WhenAandBare combined via
matrix multiplication, the result is the identity matrix, which can be inserted \in front" of xas the rst
step in nding the solution. This is entirely analogous to how we might solve a single linear equation like
3x= 12.
x= 1x=1
3(3)
x=1
3(3x) =1
3(12) = 4
Here we have obtained a solution by employing the \multiplicative inverse" of 3, 3 1=1
3. This works ne
for any scalar multiple of x, except for zero, since zero does not have a multiplicative inverse. Consider
seperately the two linear equations,
0x= 12 0 x= 0
The rst has no solutions, while the second has innitely many solutions. For matrices, it is all just a little
more complicated. Some matrices have inverses, some do not. And when a matrix does have an inverse,
just how would we compute it? In other words, just where did that matrix Bin the last example come
from? Are there other matrices that might have worked just as well?
Subsection IM
Inverse of a Matrix
Denition MI
Matrix Inverse
SupposeAandBare square matrices of size nsuch thatAB=InandBA=In. ThenAisinvertible
andBis the inverse ofA. In this situation, we write B=A 1.
(This denition contains Notation MI.) 4
Notice that if Bis the inverse of A, then we can just as easily say Ais the inverse of B, orAandB
are inverses of each other.
Not every square matrix has an inverse. In Example SABMI [243] the matrix Bis the inverse the
coecient matrix of Archetype B [786]. To see this it only remains to check that AB=I3. What about
Archetype A [781]? It is an example of a square matrix without an inverse.
Example MWIAA
A matrix without an inverse, Archetype A
Consider the coecient matrix from Archetype A [781],
A=2
41 1 2
2 1 1
1 1 03
5
Suppose that Ais invertible and does have an inverse, say B. Choose the vector of constants
b=2
41
3
23
5
and consider the system of equations LS(A;b). Just as in Example SABMI [243], this vector equation
would have the unique solution x=Bb.
Version 2.30
Subsection MISLE.CIM Computing the Inverse of a Matrix 247
However, the system LS(A;b) is inconsistent. Form the augmented matrix [ Ajb] and row-reduce to
2
410 1 0
01 1 0
0 0 0 13
5
which allows to recognize the inconsistency by Theorem RCLS [58].
So the assumption of A's inverse leads to a logical inconsistency (the system can't be both consistent
and inconsistent), so our assumption is false. Ais not invertible.
Its possible this example is less than satisfying. Just where did that particular choice of the vector b
come from anyway? Stay tuned for an application of the future Theorem CSCS [272] in Example CSAA
[276].
Let's look at one more matrix inverse before we embark on a more systematic study.
Example MI
Matrix inverse
Consider the matrices,
A=2
666641 2 1 2 1
2 3 0 5 1
1 1 0 2 1
2 3 1 3 2
1 3 1 3 13
77775B=2
66664 3 3 6 1 2
0 2 5 1 1
1 2 4 1 1
1 0 1 1 0
1 1 2 0 13
77775
Then
AB=2
666641 2 1 2 1
2 3 0 5 1
1 1 0 2 1
2 3 1 3 2
1 3 1 3 13
777752
66664 3 3 6 1 2
0 2 5 1 1
1 2 4 1 1
1 0 1 1 0
1 1 2 0 13
77775=2
666641 0 0 0 0
0 1 0 0 0
0 0 1 0 0
0 0 0 1 0
0 0 0 0 13
77775
and
BA=2
66664 3 3 6 1 2
0 2 5 1 1
1 2 4 1 1
1 0 1 1 0
1 1 2 0 13
777752
666641 2 1 2 1
2 3 0 5 1
1 1 0 2 1
2 3 1 3 2
1 3 1 3 13
77775=2
666641 0 0 0 0
0 1 0 0 0
0 0 1 0 0
0 0 0 1 0
0 0 0 0 13
77775
so by Denition MI [244], we can say that Ais invertible and write B=A 1.
We will now concern ourselves less with whether or not an inverse of a matrix exists, but instead with
how you can nd one when it does exist. In Section MINM [259] we will have some theorems that allow
us to more quickly and easily determine just when a matrix is invertible.
Subsection CIM
Computing the Inverse of a Matrix
We've seen that the matrices from Archetype B [786] and Archetype K [825] both have inverses, but these
inverse matrices have just dropped from the sky. How would we compute an inverse? And just when is
a matrix invertible, and when is it not? Writing a putative inverse with n2unknowns and solving the
Version 2.30
248 Section MISLE Matrix Inverses and Systems of Linear Equations
resultantn2equations is one approach. Applying this approach to 2 2 matrices can get us somewhere,
so just for fun, let's do it.
Theorem TTMI
Two-by-Two Matrix Inverse
Suppose
A=a b
c d
ThenAis invertible if and only if ad bc6= 0. When Ais invertible, then
A 1=1
ad bcd b
c a
Proof (() Assume that ad bc6= 0. We will use the denition of the inverse of a matrix to establish
thatAhas inverse (Denition MI [244]). Note that if ad bc6= 0 then the displayed formula for A 1is
legitimate since we are not dividing by zero). Using this proposed formula for the inverse of A, we compute
AA 1=a b
c d1
ad bcd b
c a
=1
ad bcad bc 0
0ad bc
=1 0
0 1
and
A 1A=1
ad bcd b
c aa b
c d
=1
ad bcad bc 0
0ad bc
=1 0
0 1
By Denition MI [244] this is sucient to establish that Ais invertible, and that the expression for A 1is
correct.
()) Assume that Ais invertible, and proceed with a proof by contradiction (Technique CD [770]),
by assuming also that ad bc= 0. This translates to ad=bc. Let
B=e f
g h
be a putative inverse of A. This means that
I2=AB=a b
c de f
g h
=ae+bg af +bh
ce+dg cf +dh
Working on the matrices on two ends of this equation, we will multiply the top row by cand the bottom
row bya.c0
0a
=ace+bcg acf +bch
ace+adg acf +adh
We are assuming that ad=bc, so we can replace two occurrences of adbybcin the bottom row of the
right matrix.c0
0a
=ace+bcg acf +bch
ace+bcg acf +bch
The matrix on the right now has two rows that are identical, and therefore the same must be true of the
matrix on the left. Identical rows for the matrix on the left implies that a= 0 andc= 0.
With this information, the product ABbecomes
1 0
0 1
=I2=AB=ae+bg af +bh
ce+dg cf +dh
=bg bh
dg dh
Version 2.30
Subsection MISLE.CIM Computing the Inverse of a Matrix 249
Sobg=dh= 1 and thus b;g;d;h are all nonzero. But then bhanddg(the \other corners") must also
be nonzero, so this is (nally) a contradiction. So our assumption was false and we see that ad bc6= 0
wheneverAhas an inverse.
There are several ways one could try to prove this theorem, but there is a continual temptation to divide
by one of the eight entries involved ( athroughf), but we can never be sure if these numbers are zero or
not. This could lead to an analysis by cases, which is messy, messy, messy. Note how the above proof
never divides, but always multiplies, and how zero/nonzero considerations are handled. Pay attention to
the expression ad bc, as we will see it again in a while (Chapter D [423]).
This theorem is cute, and it is nice to have a formula for the inverse, and a condition that tells us when
we can use it. However, this approach becomes impractical for larger matrices, even though it is possible
to demonstrate that, in theory, there is a general formula. (Think for a minute about extending this result
to just 33 matrices. For starters, we need 18 letters!) Instead, we will work column-by-column. Let's
rst work an example that will motivate the main theorem and remove some of the previous mystery.
Example CMI
Computing a matrix inverse
Consider the matrix dened in Example MI [245] as,
A=2
666641 2 1 2 1
2 3 0 5 1
1 1 0 2 1
2 3 1 3 2
1 3 1 3 13
77775
For its inverse, we desire a matrix Bso thatAB=I5. Emphasizing the structure of the columns and
employing the denition of matrix multiplication Denition MM [226],
AB=I5
A[B1jB2jB3jB4jB5] = [e1je2je3je4je5]
[AB1jAB2jAB3jAB4jAB5] = [e1je2je3je4je5]:
Equating the matrices column-by-column we have
AB1=e1AB2=e2AB3=e3AB4=e4AB5=e5:
Since the matrix Bis what we are trying to compute, we can view each column, Bi, as a column vector of
unknowns. Then we have ve systems of equations to solve, each with 5 equations in 5 variables. Notice
that all 5 of these systems have the same coecient matrix. We'll now solve each system in turn,
Row-reduce the augmented matrix of the linear system LS(A;e1),
2
666641 2 1 2 1 1
2 3 0 5 1 0
1 1 0 2 1 0
2 3 1 3 2 0
1 3 1 3 1 03
77775RREF !2
66666410 0 0 0 3
010 0 0 0
0 0 10 0 1
0 0 0 10 1
0 0 0 0 1 13
777775so B1=2
66664 3
0
1
1
13
77775
Row-reduce the augmented matrix of the linear system LS(A;e2),
2
666641 2 1 2 1 0
2 3 0 5 1 1
1 1 0 2 1 0
2 3 1 3 2 0
1 3 1 3 1 03
77775RREF !2
66666410 0 0 0 3
010 0 0 2
0 0 10 0 2
0 0 0 10 0
0 0 0 0 1 13
777775so B2=2
666643
2
2
0
13
77775
Version 2.30
250 Section MISLE Matrix Inverses and Systems of Linear Equations
Row-reduce the augmented matrix of the linear system LS(A;e3),
2
666641 2 1 2 1 0
2 3 0 5 1 0
1 1 0 2 1 1
2 3 1 3 2 0
1 3 1 3 1 03
77775RREF !2
66666410 0 0 0 6
010 0 0 5
0 0 10 0 4
0 0 0 10 1
0 0 0 0 1 23
777775so B3=2
666646
5
4
1
23
77775
Row-reduce the augmented matrix of the linear system LS(A;e4),
2
666641 2 1 2 1 0
2 3 0 5 1 0
1 1 0 2 1 0
2 3 1 3 2 1
1 3 1 3 1 03
77775RREF !2
66666410 0 0 0 1
010 0 0 1
0 0 10 0 1
0 0 0 10 1
0 0 0 0 1 03
777775so B4=2
66664 1
1
1
1
03
77775
Row-reduce the augmented matrix of the linear system LS(A;e5),
2
666641 2 1 2 1 0
2 3 0 5 1 0
1 1 0 2 1 0
2 3 1 3 2 0
1 3 1 3 1 13
77775RREF !2
66666410 0 0 0 2
010 0 0 1
0 0 10 0 1
0 0 0 10 0
0 0 0 0 1 13
777775so B5=2
66664 2
1
1
0
13
77775
We can now collect our 5 solution vectors into the matrix B,
B=[B1jB2jB3jB4jB5]
=2
666642
66664 3
0
1
1
13
777752
666643
2
2
0
13
777752
666646
5
4
1
23
777752
66664 1
1
1
1
03
777752
66664 2
1
1
0
13
777753
77775
=2
66664 3 3 6 1 2
0 2 5 1 1
1 2 4 1 1
1 0 1 1 0
1 1 2 0 13
77775
By this method, we know that AB=I5. Check that BA=I5, and then we will know that we have the
inverse ofA.
Notice how the ve systems of equations in the preceding example were all solved by exactly the same
sequence of row operations. Wouldn't it be nice to avoid this obvious duplication of eort? Our main
theorem for this section follows, and it mimics this previous example, while also avoiding all the overhead.
Theorem CINM
Computing the Inverse of a Nonsingular Matrix
SupposeAis a nonsingular square matrix of size n. Create the n2nmatrixMby placing the nn
identity matrix Into the right of the matrix A. LetNbe a matrix that is row-equivalent to Mand
Version 2.30
Subsection MISLE.CIM Computing the Inverse of a Matrix 251
in reduced row-echelon form. Finally, let Jbe the matrix formed from the nal ncolumns of N. Then
AJ=In.
ProofAis nonsingular, so by Theorem NMRRI [84] there is a sequence of row operations that will
convertAintoIn. It is this same sequence of row operations that will convert MintoN, since having
the identity matrix in the rst ncolumns of Nis sucient to guarantee that Nis in reduced row-echelon
form.
If we consider the systems of linear equations, LS(A;ei), 1in, we see that the aforementioned
sequence of row operations will also bring the augmented matrix of each of these systems into reduced row-
echelon form. Furthermore, the unique solution to LS(A;ei) appears in column n+ 1 of the row-reduced
augmented matrix of the system and is identical to column n+iofN. Let N1;N2;N3; :::; N2ndenote
the columns of N. So we nd,
AJ=A[Nn+1jNn+2jNn+3j:::jNn+n]
=[ANn+1jANn+2jANn+3j:::jANn+n] Denition MM [226]
=[e1je2je3j:::jen]
=In Denition IM [84]
as desired.
We have to be just a bit careful here about both what this theorem says and what it doesn't say. If A
is a nonsingular matrix, then we are guaranteed a matrix Bsuch thatAB=In, and the proof gives us a
process for constructing B. However, the denition of the inverse of a matrix (Denition MI [244]) requires
thatBA=Inalso. So at this juncture we must compute the matrix product in the \opposite" order before
we claimBas the inverse of A. However, we'll soon see that this is always the case, in Theorem OSIS
[260], so the title of this theorem is not inaccurate.
What ifAis singular? At this point we only know that Theorem CINM [248] cannot be applied.
The question of A's inverse is still open. (But see Theorem NI [261] in the next section.) We'll nish by
computing the inverse for the coecient matrix of Archetype B [786], the one we just pulled from a hat in
Example SABMI [243]. There are more examples in the Archetypes (Appendix A [777]) to practice with,
though notice that it is silly to ask for the inverse of a rectangular matrix (the sizes aren't right) and not
every square matrix has an inverse (remember Example MWIAA [244]?).
Example CMIAB
Computing a matrix inverse, Archetype B
Archetype B [786] has a coecient matrix given as
B=2
4 7 6 12
5 5 7
1 0 43
5
Exercising Theorem CINM [248] we set
M=2
4 7 6 12 1 0 0
5 5 7 0 1 0
1 0 4 0 0 13
5:
which row reduces to
N=2
41 0 0 10 12 9
0 1 013
2811
2
0 0 15
235
23
5:
Version 2.30
252 Section MISLE Matrix Inverses and Systems of Linear Equations
So
B 1=2
4 10 12 9
13
2811
25
235
23
5
once we check that B 1B=I3(the product in the opposite order is a consequence of the theorem).
While we can use a row-reducing procedure to compute any needed inverse, most computational devices
have a built-in procedure to compute the inverse of a matrix straightaway. See: Computation MI.MMA
[749] Computation MI.SAGE [755]
Subsection PMI
Properties of Matrix Inverses
The inverse of a matrix enjoys some nice properties. We collect a few here. First, a matrix can have but
one inverse.
Theorem MIU
Matrix Inverse is Unique
Suppose the square matrix Ahas an inverse. Then A 1is unique.
Proof As described in Technique U [771], we will assume that Ahas two inverses. The hypothesis tells
there is at least one. Suppose then that BandCare both inverses for A, so we know by Denition MI
[244] thatAB=BA=InandAC=CA=In. Then we have,
B=BIn Theorem MMIM [229]
=B(AC) Denition MI [244]
= (BA)C Theorem MMA [231]
=InC Denition MI [244]
=C Theorem MMIM [229]
So we conclude that BandCare the same, and cannot be dierent. So any matrix that acts like an
inverse, must be theinverse.
When most of us dress in the morning, we put on our socks rst, followed by our shoes. In the evening
we must then rst remove our shoes, followed by our socks. Try to connect the conclusion of the following
theorem with this everyday example.
Theorem SS
Socks and Shoes
SupposeAandBare invertible matrices of size n. ThenABis an invertible matrix and ( AB) 1=B 1A 1.
Proof At the risk of carrying our everyday analogies too far, the proof of this theorem is quite easy when
we compare it to the workings of a dating service. We have a statement about the inverse of the matrix
AB, which for all we know right now might not even exist. Suppose ABwas to sign up for a dating service
with two requirements for a compatible date. Upon multiplication on the left, and on the right, the result
should be the identity matrix. In other words, AB's ideal date would be its inverse.
Now along comes the matrix B 1A 1(which we know exists because our hypothesis says both Aand
Bare invertible and we can form the product of these two matrices), also looking for a date. Let's see if
Version 2.30
Subsection MISLE.PMI Properties of Matrix Inverses 253
B 1A 1is a good match for AB. First they meet at a non-committal neutral location, say a coee shop,
for quiet conversation:
(B 1A 1)(AB) =B 1(A 1A)B Theorem MMA [231]
=B 1InB Denition MI [244]
=B 1B Theorem MMIM [229]
=In Denition MI [244]
The rst date having gone smoothly, a second, more serious, date is arranged, say dinner and a show:
(AB)(B 1A 1) =A(BB 1)A 1Theorem MMA [231]
=AInA 1Denition MI [244]
=AA 1Theorem MMIM [229]
=In Denition MI [244]
So the matrix B 1A 1has met all of the requirements to be AB's inverse (date) and with the ensuing
marriage proposal we can announce that ( AB) 1=B 1A 1.
Theorem MIMI
Matrix Inverse of a Matrix Inverse
SupposeAis an invertible matrix. Then A 1is invertible and ( A 1) 1=A.
Proof As with the proof of Theorem SS [250], we examine if Ais a suitable inverse for A 1(by denition,
the opposite is true).
AA 1=In Denition MI [244]
and
A 1A=In Denition MI [244]
The matrix Ahas met all the requirements to be the inverse of A 1, and so is invertible and we can write
A= (A 1) 1.
Theorem MIT
Matrix Inverse of a Transpose
SupposeAis an invertible matrix. Then Atis invertible and ( At) 1= (A 1)t.
Proof As with the proof of Theorem SS [250], we see if ( A 1)tis a suitable inverse for At. Apply Theorem
MMT [232] to see that
(A 1)tAt= (AA 1)tTheorem MMT [232]
=It
n Denition MI [244]
=In Denition SYM [211]
and
At(A 1)t= (A 1A)tTheorem MMT [232]
=It
n Denition MI [244]
=In Denition SYM [211]
Version 2.30
254 Section MISLE Matrix Inverses and Systems of Linear Equations
The matrix ( A 1)thas met all the requirements to be the inverse of At, and so is invertible and we can
write (At) 1= (A 1)t.
Theorem MISM
Matrix Inverse of a Scalar Multiple
SupposeAis an invertible matrix and is a nonzero scalar. Then ( A) 1=1
A 1andAis invertible.
Proof As with the proof of Theorem SS [250], we see if1
A 1is a suitable inverse for A.
1
A 1
(A) =1
AA 1
Theorem MMSMM [230]
= 1In Scalar multiplicative inverses
=In Property OM [209]
and
(A)1
A 1
=
1
A 1A
Theorem MMSMM [230]
= 1In Scalar multiplicative inverses
=In Property OM [209]
The matrix1
A 1has met all the requirements to be the inverse of A, so we can write ( A) 1=1
A 1.
Notice that there are some likely theorems that are missing here. For example, it would be tempting
to think that ( A+B) 1=A 1+B 1, but this is false. Can you nd a counterexample? (See Exercise
MISLE.T10 [255].)
Subsection READ
Reading Questions
1. Compute the inverse of the matrix below.
4 10
2 6
2. Compute the inverse of the matrix below.
2
42 3 1
1 2 3
2 4 63
5
3. Explain why Theorem SS [250] has the title it does. (Do not just state the theorem, explain the
choice of the title making reference to the theorem itself.)
Version 2.30
Subsection MISLE.EXC Exercises 255
Subsection EXC
Exercises
C16 If it exists, nd the inverse of A=2
41 0 1
1 1 1
2 1 13
5, and check your answer.
Contributed by Chris Black Solution [256]
C17 If it exists, nd the inverse of A=2
42 1 1
1 2 1
3 1 23
5, and check your answer.
Contributed by Chris Black Solution [256]
C18 If it exists, nd the inverse of A=2
41 3 1
1 2 1
2 2 13
5, and check your answer.
Contributed by Chris Black Solution [256]
C19 If it exists, nd the inverse of A=2
41 3 1
0 2 1
2 2 13
5, and check your answer.
Contributed by Chris Black Solution [256]
C21 Verify that Bis the inverse of A.
A=2
6641 1 1 2
2 1 2 3
1 1 0 2
1 2 0 23
775B=2
6644 2 0 1
8 4 1 1
1 0 1 0
6 3 1 13
775
Contributed by Robert Beezer Solution [256]
C22 Recycle the matrices AandBfrom Exercise MISLE.C21 [253] and set
c=2
6642
1
3
23
775d=2
6641
1
1
13
775
Employ the matrix Bto solve the two linear systems LS(A;c) andLS(A;d).
Contributed by Robert Beezer Solution [256]
C23 If it exists, nd the inverse of the 2 2 matrix
A=7 3
5 2
and check your answer. (See Theorem TTMI [246].)
Contributed by Robert Beezer
C24 If it exists, nd the inverse of the 2 2 matrix
A=6 3
4 2
Version 2.30
256 Section MISLE Matrix Inverses and Systems of Linear Equations
and check your answer. (See Theorem TTMI [246].)
Contributed by Robert Beezer
C25 At the conclusion of Example CMI [247], verify that BA=I5by computing the matrix product.
Contributed by Robert Beezer
C26 Let
D=2
666641 1 3 2 1
2 3 5 3 0
1 1 4 2 2
1 4 1 0 4
1 0 5 2 53
77775
Compute the inverse of D,D 1, by forming the 5 10 matrix [ DjI5] and row-reducing (Theorem CINM
[248]). Then use a calculator to compute D 1directly.
Contributed by Robert Beezer Solution [256]
C27 Let
E=2
666641 1 3 2 1
2 3 5 3 1
1 1 4 2 2
1 4 1 0 2
1 0 5 2 43
77775
Compute the inverse of E,E 1, by forming the 5 10 matrix [ EjI5] and row-reducing (Theorem CINM
[248]). Then use a calculator to compute E 1directly.
Contributed by Robert Beezer Solution [256]
C28 Let
C=2
6641 1 3 1
2 1 4 1
1 4 10 2
2 0 4 53
775
Compute the inverse of C,C 1, by forming the 4 8 matrix [CjI4] and row-reducing (Theorem CINM
[248]). Then use a calculator to compute C 1directly.
Contributed by Robert Beezer Solution [257]
C40 Find all solutions to the system of equations below, making use of the matrix inverse found in
Exercise MISLE.C28 [254].
x1+x2+ 3x3+x4= 4
2x1 x2 4x3 x4= 4
x1+ 4x2+ 10x3+ 2x4= 20
2x1 4x3+ 5x4= 9
Contributed by Robert Beezer Solution [257]
C41 Use the inverse of a matrix to nd all the solutions to the following system of equations.
x1+ 2x2 x3= 3
2x1+ 5x2 x3= 4
x1 4x2= 2
Version 2.30
Subsection MISLE.EXC Exercises 257
Contributed by Robert Beezer Solution [257]
C42 Use a matrix inverse to solve the linear system of equations.
x1 x2+ 2x3= 5
x1 2x3= 8
2x1 x2 x3= 6
Contributed by Robert Beezer Solution [257]
T10 Construct an example to demonstrate that ( A+B) 1=A 1+B 1is not true for all square matrices
AandBof the same size.
Contributed by Robert Beezer Solution [258]
Version 2.30
258 Section MISLE Matrix Inverses and Systems of Linear Equations
Subsection SOL
Solutions
C16 Contributed by Chris Black Statement [253]
Answer:A 1=2
4 2 1 1
1 1 0
3 1 13
5.
C17 Contributed by Chris Black Statement [253]
The procedure we have for nding a matrix inverse fails for this matrix AsinceAdoes not row-reduce to
I3. We suspect in this case that Ais not invertible, although we do not yet know that concretely. (Stay
tuned for upcoming revelations in Section MINM [259]!)
C18 Contributed by Chris Black Statement [253]
Answer:A 1=2
40 1 1
1 1 0
2 4 13
5
C19 Contributed by Chris Black Statement [253]
Answer:A 1=2
40 1=2 1=2
1 1=2 1=2
2 2 13
5
C21 Contributed by Robert Beezer Statement [253]
Check that both matrix products (Denition MM [226]) ABandBAequal the 44 identity matrix I4
(Denition IM [84]).
C22 Contributed by Robert Beezer Statement [253]
Represent each of the two systems by a vector equality, Ax=candAy=d. Then in the spirit of Example
SABMI [243], solutions are given by
x=Bc=2
6648
21
5
163
775y=Bd=2
6645
10
0
73
775
Notice how we could solve many more systems having Aas the coecient matrix, and how each such system
has a unique solution. You might check your work by substituting the solutions back into the systems of
equations, or forming the linear combinations of the columns of Asuggested by Theorem SLSLC [112].
C26 Contributed by Robert Beezer Statement [254]
The inverse of Dis
D 1=2
66664 7 6 3 2 1
7 4 2 2 1
5 2 3 1 1
6 3 1 1 0
4 2 2 1 13
77775
C27 Contributed by Robert Beezer Statement [254]
The matrix Ehas no inverse, though we do not yet have a theorem that allows us to reach this conclusion.
However, when row-reducing the matrix [ EjI5], the rst 5 columns will not row-reduce to the 5 5 identity
matrix, so we are a t a loss on how we might compute the inverse. When requesting that your calculator
computeE 1, it should give some indication that Edoes not have an inverse.
Version 2.30
Subsection MISLE.SOL Solutions 259
C28 Contributed by Robert Beezer Statement [254]
Employ Theorem CINM [248],
2
6641 1 3 1 1 0 0 0
2 1 4 1 0 1 0 0
1 4 10 2 0 0 1 0
2 0 4 5 0 0 0 13
775RREF !2
666410 0 0 38 18 5 2
010 0 96 47 12 5
0 0 10 39 19 5 2
0 0 0 1 16 8 2 13
7775
And therefore we see that Cis nonsingular ( Crow-reduces to the identity matrix, Theorem NMRRI [84])
and by Theorem CINM [248],
C 1=2
66438 18 5 2
96 47 12 5
39 19 5 2
16 8 2 13
775
C40 Contributed by Robert Beezer Statement [254]
View this system as LS(C;b), whereCis the 44 matrix from Exercise MISLE.C28 [254] and b=2
664 4
4
20
93
775.
SinceCwas seen to be nonsingular in Exercise MISLE.C28 [254] Theorem SNCM [261] says the solution,
which is unique by Theorem NMUS [86], is given by
C 1b=2
66438 18 5 2
96 47 12 5
39 19 5 2
16 8 2 13
7752
664 4
4
20
93
775=2
6642
1
2
13
775
Notice that this solution can be easily checked in the original system of equations.
C41 Contributed by Robert Beezer Statement [254]
The coecient matrix of this system of equations is
A=2
41 2 1
2 5 1
1 4 03
5
and the vector of constants is b=2
4 3
4
23
5. So by Theorem SLEMM [224] we can convert the system to the
formAx=b. Row-reducing this matrix yields the identity matrix so by Theorem NMRRI [84] we know
Ais nonsingular. This allows us to apply Theorem SNCM [261] to nd the unique solution as
x=A 1b=2
4 4 4 3
1 1 1
3 2 13
52
4 3
4
23
5=2
42
1
33
5
Remember, you can check this solution easily by evaluating the matrix-vector product Ax(Denition MVP
[223]).
C42 Contributed by Robert Beezer Statement [255]
We can reformulate the linear system as a vector equality with a matrix-vector product via Theorem
SLEMM [224]. The system is then represented by Ax=bwhere
A=2
41 1 2
1 0 2
2 1 13
5 b=2
45
8
63
5
Version 2.30
260 Section MISLE Matrix Inverses and Systems of Linear Equations
According to Theorem SNCM [261], if Ais nonsingular then the (unique) solution will be given by A 1b.
We attempt the computation of A 1through Theorem CINM [248], or with our favorite computational
device and obtain,
A 1=2
42 3 2
3 5 4
1 1 13
5
So by Theorem NI [261], we know Ais nonsingular, and so the unique solution is
A 1b=2
42 3 2
3 5 4
1 1 13
52
45
8
63
5=2
4 2
1
33
5
T10 Contributed by Robert Beezer Statement [255]
For a large collection of small examples, let Dbe any 22 matrix that has an inverse (Theorem TTMI
[246] can help you construct such a matrix, I2is a simple choice). Set A=DandB= ( 1)D. WhileA 1
andB 1both exist, what is ( A+B) 1?
For a large collection of examples of any size, consider A=B=In. Can the proposed statement be
salvaged to become a theorem?
Version 2.30
Section MINM Matrix Inverses and Nonsingular Matrices 261
Section MINM
Matrix Inverses and Nonsingular Matrices
We saw in Theorem CINM [248] that if a square matrix Ais nonsingular, then there is a matrix Bso
thatAB=In. In other words, Bis halfway to being an inverse of A. We will see in this section that
Bautomatically fullls the second condition ( BA=In). Example MWIAA [244] showed us that the
coecient matrix from Archetype A [781] had no inverse. Not coincidentally, this coecient matrix is
singular. We'll make all these connections precise now. Not many examples or denitions in this section,
just theorems.
Subsection NMI
Nonsingular Matrices are Invertible
We need a couple of technical results for starters. Some books would call these minor, but essential, results
\lemmas." We'll just call 'em theorems. See Technique LC [774] for more on the distinction.
The rst of these technical results is interesting in that the hypothesis says something about a product
of two square matrices and the conclusion then says the same thing about each individual matrix in the
product. This result has an analogy in the algebra of complex numbers: suppose ; 2C, then6= 0
if and only if 6= 0 and6= 0. We can view this result as suggesting that the term \nonsingular" for
matrices is like the term \nonzero" for scalars.
Theorem NPNT
Nonsingular Product has Nonsingular Terms
Suppose that AandBare square matrices of size n. The product ABis nonsingular if and only if Aand
Bare both nonsingular.
Proof ()) We'll do this portion of the proof in two parts, each as a proof by contradiction (Technique
CD [770]). Assume that ABis nonsingular. Establishing that Bis nonsingular is the easier part, so we will
do it rst, but in reality, we will need to know that Bis nonsingular when we prove that Ais nonsingular.
You can also think of this proof as being a study of four possible conclusions in the table below. One
of the four rows must happen (the list is exhaustive). In the proof we learn that the rst three rows lead
to contradictions, and so are impossible. That leaves the fourth row as a certainty, which is our desired
conclusion.
A B Case
Singular Singular 1
Nonsingular Singular 1
Singular Nonsingular 2
Nonsingular Nonsingular
Part 1. Suppose Bis singular. Then there is a nonzero vector zthat is a solution to LS(B;0). So
(AB)z=A(Bz) Theorem MMA [231]
=A0 Theorem SLEMM [224]
=0 Theorem MMZM [229]
Because zis a nonzero solution to LS(AB;0), we conclude that ABis singular (Denition NM [83]). This
is a contradiction, so Bis nonsingular, as desired.
Version 2.30
262 Section MINM Matrix Inverses and Nonsingular Matrices
Part 2. Suppose Ais singular. Then there is a nonzero vector ythat is a solution to LS(A;0). Now
consider the linear system LS(B;y). Since we know Bis nonsingular from Case 1, the system has a unique
solution (Theorem NMUS [86]), which we will denote as w. We rst claim wis not the zero vector either.
Assuming the opposite, suppose that w=0(Technique CD [770]). Then
y=Bw Theorem SLEMM [224]
=B0 Hypothesis
=0 Theorem MMZM [229]
contrary to ybeing nonzero. So w6=0. The pieces are in place, so here we go,
(AB)w=A(Bw) Theorem MMA [231]
=Ay Theorem SLEMM [224]
=0 Theorem SLEMM [224]
Sowis a nonzero solution to LS(AB;0), and thus we can say that ABis singular (Denition NM [83]).
This is a contradiction, so Ais nonsingular, as desired.
(() Now assume that both AandBare nonsingular. Suppose that x2Cnis a solution toLS(AB;0).
Then
0= (AB)x Theorem SLEMM [224]
=A(Bx) Theorem MMA [231]
By Theorem SLEMM [224], Bxis a solution toLS(A;0), and by the denition of a nonsingular matrix
(Denition NM [83]), we conclude that Bx=0. Now, by an entirely similar argument, the nonsingularity
ofBforces us to conclude that x=0. So the only solution to LS(AB;0) is the zero vector and we
conclude that ABis nonsingular by Denition NM [83].
This is a powerful result in the \forward" direction, because it allows us to begin with a hypothesis
that something complicated (the matrix product AB) has the property of being nonsingular, and we can
then conclude that the simpler constituents ( AandBindividually) then also have the property of being
nonsingular. If we had thought that the matrix product was an articial construction, results like this
would make us begin to think twice.
The contrapositive of this result is equally interesting. It says that AorB(or both) is a singular matrix
if and only if the product ABis singular. Notice how the negation of the theorem's conclusion ( AandB
both nonsingular) becomes the statement \at least one of AandBis singular." (See Technique CP [769].)
Theorem OSIS
One-Sided Inverse is Sucient
SupposeAandBare square matrices of size nsuch thatAB=In. ThenBA=In.
Proof The matrix Inis nonsingular (since it row-reduces easily to In, Theorem NMRRI [84]). So A
andBare nonsingular by Theorem NPNT [259], so in particular Bis nonsingular. We can therefore
apply Theorem CINM [248] to assert the existence of a matrix Cso thatBC=In. This application of
Theorem CINM [248] could be a bit confusing, mostly because of the names of the matrices involved. B
is nonsingular, so there must be a \right-inverse" for B, and we're calling it C.
Now
BA= (BA)In Theorem MMIM [229]
= (BA)(BC) Theorem CINM [248]
=B(AB)C Theorem MMA [231]
Version 2.30
Subsection MINM.NMI Nonsingular Matrices are Invertible 263
=BInC Hypothesis
=BC Theorem MMIM [229]
=In Theorem CINM [248]
which is the desired conclusion.
So Theorem OSIS [260] tells us that if Ais nonsingular, then the matrix Bguaranteed by Theorem
CINM [248] will be both a \right-inverse" and a \left-inverse" for A, soAis invertible and A 1=B.
So if you have a nonsingular matrix, A, you can use the procedure described in Theorem CINM [248]
to nd an inverse for A. IfAis singular, then the procedure in Theorem CINM [248] will fail as the rst
ncolumns of Mwill not row-reduce to the identity matrix. However, we can say a bit more. When A
is singular, then Adoes not have an inverse (which is very dierent from saying that the procedure in
Theorem CINM [248] fails to nd an inverse). This may feel like we are splitting hairs, but its important
that we do not make unfounded assumptions. These observations motivate the next theorem.
Theorem NI
Nonsingularity is Invertibility
Suppose that Ais a square matrix. Then Ais nonsingular if and only if Ais invertible.
Proof (() SinceAis invertible, we can write In=AA 1(Denition MI [244]). Notice that Inis
nonsingular (Theorem NSRRI [ ??]) so Theorem NPNT [259] implies that A(andA 1) is nonsingular.
()) Suppose now that Ais nonsingular. By Theorem CINM [248] we nd Bso thatAB=In. Then
Theorem OSIS [260] tells us that BA=In. SoBisA's inverse, and by construction, Ais invertible.
So for a square matrix, the properties of having an inverse and of having a trivial null space are one
and the same. Can't have one without the other.
Theorem NME3
Nonsingular Matrix Equivalences, Round 3
Suppose that Ais a square matrix of size n. The following are equivalent.
1.Ais nonsingular.
2.Arow-reduces to the identity matrix.
3. The null space of Acontains only the zero vector, N(A) =f0g.
4. The linear system LS(A;b) has a unique solution for every possible choice of b.
5. The columns of Aare a linearly independent set.
6.Ais invertible.
Proof We can update our list of equivalences for nonsingular matrices (Theorem NME2 [159]) with the
equivalent condition from Theorem NI [261].
In the case that Ais a nonsingular coecient matrix of a system of equations, the inverse allows us to
very quickly compute the unique solution, for any vector of constants.
Theorem SNCM
Solution with Nonsingular Coecient Matrix
Suppose that Ais nonsingular. Then the unique solution to LS(A;b) isA 1b.
Proof By Theorem NMUS [86] we know already that LS(A;b) has a unique solution for every choice of
b. We need to show that the expression stated is indeed a solution ( thesolution). That's easy, just \plug
it in" to the corresponding vector equation representation (Theorem SLEMM [224]),
A
A 1b
=
AA 1
b Theorem MMA [231]
Version 2.30
264 Section MINM Matrix Inverses and Nonsingular Matrices
=Inb Denition MI [244]
=b Theorem MMIM [229]
SinceAx=bis true when we substitute A 1bforx,A 1bis a (the!) solution to LS(A;b).
Subsection UM
Unitary Matrices
Recall that the adjoint of a matrix is A=
At(Denition A [214]).
Denition UM
Unitary Matrices
Suppose that Uis a square matrix of size nsuch thatUU=In. Then we say Uisunitary .4
This condition may seem rather far-fetched at rst glance. Would there be anymatrix that behaved
this way? Well, yes, here's one.
Example UM3
Unitary matrix of size 3
U=2
641+ip
53+2ip
552+2ip
221 ip
52+2ip
55 3+ip
22ip
53 5ip
55 2p
223
75
The computations get a bit tiresome, but if you work your way through the computation of UU, you will
arrive at the 33 identity matrix I3.
Unitary matrices do not have to look quite so gruesome. Here's a larger one that is a bit more pleasing.
Example UPM
Unitary permutation matrix
The matrix
P=2
666640 1 0 0 0
0 0 0 1 0
1 0 0 0 0
0 0 0 0 1
0 0 1 0 03
77775
is unitary as can be easily checked. Notice that it is just a rearrangement of the columns of the 5 5
identity matrix, I5(Denition IM [84]).
An interesting exercise is to build another 5 5 unitary matrix, R, using a dierent rearrangement of
the columns of I5. Then form the product PR. This will be another unitary matrix (Exercise MINM.T10
[266]). If you were to build all 5! = 5 4321 = 120 matrices of this type you would have a set
that remains closed under matrix multiplication. It is an example of another algebraic structure known as
agroup since together the set and the one operation (matrix multiplication here) is closed, associative,
has an identity ( I5), and inverses (Theorem UMI [263]). Notice though that the operation in this group is
not commutative!
If a matrix Ahas only real number entries (we say it is a real matrix ) then the dening property of
being unitary simplies to AtA=In. In this case we, and everybody else, calls the matrix orthogonal ,
so you may often encounter this term in your other reading when the complex numbers are not under
consideration.
Unitary matrices have easily computed inverses. They also have columns that form orthonormal sets.
Here are the theorems that show us that unitary matrices are not as strange as they might initially appear.
Version 2.30
Subsection MINM.UM Unitary Matrices 265
Theorem UMI
Unitary Matrices are Invertible
Suppose that Uis a unitary matrix of size n. ThenUis nonsingular, and U 1=U.
Proof By Denition UM [262], we know that UU=In. The matrix Inis nonsingular (since it row-
reduces easily to In, Theorem NMRRI [84]). So by Theorem NPNT [259], UandUare both nonsingular
matrices.
The equation UU=Ingets us halfway to an inverse of U, and Theorem OSIS [260] tells us that then
UU=Inalso. SoUandUare inverses of each other (Denition MI [244]).
Theorem CUMOS
Columns of Unitary Matrices are Orthonormal Sets
Suppose that Ais a square matrix of size nwith columns S=fA1;A2;A3; :::; Ang. ThenAis a unitary
matrix if and only if Sis an orthonormal set.
Proof The proof revolves around recognizing that a typical entry of the product AAis an inner product
of columns of A. Here are the details to support this claim.
[AA]ij=nX
k=1[A]ik[A]kj Theorem EMP [227]
=nX
k=1h
Ati
ik[A]kj Theorem EMP [227]
=nX
k=1
A
ki[A]kj Denition TM [210]
=nX
k=1[A]ki[A]kj Denition CCM [212]
=nX
k=1[A]kj[A]ki Property CMCN [758]
=nX
k=1[Aj]k[Ai]k
=hAj;Aii Denition IP [192]
We now employ this equality in a chain of equivalences,
S=fA1;A2;A3; :::; Angis an orthonormal set
() h Aj;Aii=(
0 ifi6=j
1 ifi=jDenition ONS [201]
() [AA]ij=(
0 ifi6=j
1 ifi=j
() [AA]ij= [In]ij;1in;1jn Denition IM [84]
()AA=In Denition ME [207]
()Ais a unitary matrix Denition UM [262]
Example OSMC
Orthonormal set from matrix columns
Version 2.30
266 Section MINM Matrix Inverses and Nonsingular Matrices
The matrix
U=2
641+ip
53+2ip
552+2ip
221 ip
52+2ip
55 3+ip
22ip
53 5ip
55 2p
223
75
from Example UM3 [262] is a unitary matrix. By Theorem CUMOS [263], its columns
8
><
>:2
641+ip
51 ip
5ip
53
75;2
643+2ip
552+2ip
553 5ip
553
75;2
642+2ip
22 3+ip
22
2p
223
759
>=
>;
form an orthonormal set. You might nd checking the six inner products of pairs of these vectors easier
than doing the matrix product UU. Or, because the inner product is anti-commutative (Theorem IPAC
[194]) you only need check three inner products (see Exercise MINM.T12 [266]).
When using vectors and matrices that only have real number entries, orthogonal matrices are those
matrices with inverses that equal their transpose. Similarly, the inner product is the familiar dot product.
Keep this special case in mind as you read the next theorem.
Theorem UMPIP
Unitary Matrices Preserve Inner Products
Suppose that Uis a unitary matrix of size nanduandvare two vectors from Cn. Then
hUu; Uvi=hu;vi and kUvk=kvk
Proof
hUu; Uvi= (Uu)tUv Theorem MMIP [231]
=utUtUv Theorem MMT [232]
=utUtUv Theorem MMCC [232]
=ut
Ut
Uv Theorem CCT [760]
=ut
UtUv Theorem MCT [214]
=ut
UtUv Theorem MMCC [232]
=utUUv Denition A [214]
=utInv Denition UM [262]
=utInv Denition IM [84]
=utv Theorem MMIM [229]
=hu;vi Theorem MMIP [231]
The second conclusion is just a specialization of the rst conclusion.
kUvk=q
kUvk2
=p
hUv; Uvi Theorem IPN [195]
=p
hv;vi
=q
kvk2Theorem IPN [195]
Version 2.30
Subsection MINM.READ Reading Questions 267
=kvk
Aside from the inherent interest in this theorem, it makes a bigger statement about unitary matrices.
When we view vectors geometrically as directions or forces, then the norm equates to a notion of length. If
we transform a vector by multiplication with a unitary matrix, then the length (norm) of that vector stays
the same. If we consider column vectors with two or three slots containing only real numbers, then the inner
product of two such vectors is just the dot product, and this quantity can be used to compute the angle
between two vectors. When two vectors are multiplied (transformed) by the same unitary matrix, their
dot product is unchanged and their individual lengths are unchanged. The results in the angle between
the two vectors remaining unchanged.
A \unitary transformation" (matrix-vector products with unitary matrices) thus preserve geometrical
relationships among vectors representing directions, forces, or other physical quantities. In the case of a two-
slot vector with real entries, this is simply a rotation. These sorts of computations are exceedingly important
in computer graphics such as games and real-time simulations, especially when increased realism is achieved
by performing many such computations quickly. We will see unitary matrices again in subsequent sections
(especially Theorem OD [681]) and in each instance, consider the interpretation of the unitary matrix
as a sort of geometry-preserving transformation. Some authors use the term isometry to highlight this
behavior. We will speak loosely of a unitary matrix as being a sort of generalized rotation.
A nal reminder: the terms \dot product," \symmetric matrix" and \orthogonal matrix" used in refer-
ence to vectors or matrices with real number entries correspond to the terms \inner product," \Hermitian
matrix" and \unitary matrix" when we generalize to include complex number entries, so keep that in mind
as you read elsewhere.
Subsection READ
Reading Questions
1. Compute the inverse of the coecient matrix of the system of equations below and use the inverse to
solve the system.
4x1+ 10x2= 12
2x1+ 6x2= 4
2. In the reading questions for Section MISLE [243] you were asked to nd the inverse of the 3 3 matrix
below. 2
42 3 1
1 2 3
2 4 63
5
Because the matrix was not nonsingular, you had no theorems at that point that would allow you to
compute the inverse. Explain why you now know that the inverse does not exist (which is dierent
than not being able to compute it) by quoting the relevant theorem's acronym.
3. Is the matrix Aunitary? Why?
A="1p
22(4 + 2i)1p
374(5 + 3i)
1p
22( 1 i)1p
374(12 + 14i)#
Version 2.30
268 Section MINM Matrix Inverses and Nonsingular Matrices
Subsection EXC
Exercises
C20 LetA=2
41 2 1
0 1 1
1 0 23
5andB=2
4 1 1 0
1 2 1
0 1 13
5. Verify that ABis nonsingular.
Contributed by Chris Black
C40 Solve the system of equations below using the inverse of a matrix.
x1+x2+ 3x3+x4= 5
2x1 x2 4x3 x4= 7
x1+ 4x2+ 10x3+ 2x4= 9
2x1 4x3+ 5x4= 9
Contributed by Robert Beezer Solution [268]
M10 Find values of x,y zso that matrix A=2
41 2x
3 0y
1 1z3
5is invertible.
Contributed by Chris Black Solution [268]
M11 Find values of x,y zso that matrix A=2
41x1
1y4
0z53
5is singular.
Contributed by Chris Black Solution [268]
M15 IfAandBarennmatrices,Ais nonsingular, and Bis singular, show directly that ABis
singular, without using Theorem NPNT [259].
Contributed by Chris Black Solution [269]
M20 Construct an example of a 4 4 unitary matrix.
Contributed by Robert Beezer Solution [268]
M80 Matrix multiplication interacts nicely with many operations. But not always with transforming a
matrix to reduced row-echelon form. Suppose that Ais anmnmatrix and Bis annpmatrix. Let Pbe
a matrix that is row-equivalent to Aand in reduced row-echelon form, Qbe a matrix that is row-equivalent
toBand in reduced row-echelon form, and let Rbe a matrix that is row-equivalent to ABand in reduced
row-echelon form. Is PQ=R? (In other words, with nonstandard notation, is rref( A)rref(B) = rref(AB)?)
Construct a counterexample to show that, in general, this statement is false. Then nd a large class of
matrices where if AandBare in the class, then the statement is true.
Contributed by Mark Hamrick Solution [269]
T10 Suppose that QandPare unitary matrices of size n. Prove that QPis a unitary matrix.
Contributed by Robert Beezer
T11 Prove that Hermitian matrices (Denition HM [234]) have real entries on the diagonal. More
precisely, suppose that Ais a Hermitian matrix of size n. Then [A]ii2R, 1in.
Contributed by Robert Beezer
T12 Suppose that we are checking if a square matrix of size nis unitary. Show that a straightforward
application of Theorem CUMOS [263] requires the computation of n2inner products when the matrix is
Version 2.30
Subsection MINM.EXC Exercises 269
unitary, and fewer when the matrix is not orthogonal. Then show that this maximum number of inner
products can be reduced to1
2n(n+ 1) in light of Theorem IPAC [194].
Contributed by Robert Beezer
T25 The notation Akmeans a repeated matrix product between kcopies of the square matrix A.
(a) Assume Ais annnmatrix where A2=O(which does not imply that A=O.) Prove that In A
is invertible by showing that In+Ais an inverse of In A.
(b) Assume that Ais annnmatrix where A3=O. Prove that In Ais invertible.
(c) Form a general theorem based on your observations from parts (a) and (b) and provide a proof.
Contributed by Manley Perkel
Version 2.30
270 Section MINM Matrix Inverses and Nonsingular Matrices
Subsection SOL
Solutions
C40 Contributed by Robert Beezer Statement [266]
The coecient matrix and vector of constants for the system are
2
6641 1 3 1
2 1 4 1
1 4 10 2
2 0 4 53
775b=2
6645
7
9
93
775
A 1can be computed by using a calculator, or by the method of Theorem CINM [248]. Then Theorem
SNCM [261] says the unique solution is
A 1b=2
66438 18 5 2
96 47 12 5
39 19 5 2
16 8 2 13
7752
6645
7
9
93
775=2
6641
2
1
33
775
M20 Contributed by Robert Beezer Statement [266]
The 44 identity matrix, I4, would be one example (Denition IM [84]). Any of the 23 other rearrangements
of the columns of I4would be a simple, but less trivial, example. See Example UPM [262].
M10 Contributed by Chris Black Statement [266]
There are an innite number of possible answers. We want to nd a vector2
4x
y
z3
5so that the set
S=8
<
:2
41
3
13
5;2
42
0
13
5;2
4x
y
z3
59
=
;
is a linearly independent set. We need a vector not in the span of the rst two columns, which geometrically
means that we need it to not be in the same plane as the rst two columns of A. We can choose any values
we want for xandy, and then choose a value of zthat makes the three vectors independent.
I will (arbitrarily) choose x= 1,y= 1. Then, we have
A=2
41 2 1
3 0 1
1 1z3
5RREF !2
410 2z 1
011 z
0 0 4 6z3
5
which is invertible if and only if 4 6z6= 0. Thus, we can choose any value as long as z6=2
3, so we choose
z= 0, and we have found a matrix A=2
41 2 1
3 0 1
1 1 03
5that is invertible.
M11 Contributed by Chris Black Statement [266]
There are an innite number of possible answers. We need the set of vectors
S=8
<
:2
41
1
03
5;2
4x
y
z3
5;2
41
4
53
59
=
;
Version 2.30
Subsection MINM.SOL Solutions 271
to be linearly dependent. One way to do this by inspection is to have2
4x
y
z3
5=2
41
4
53
5. Thus, if we let x= 1,
y= 4,z= 5, then the matrix A=2
41 1 1
1 4 4
0 5 53
5is singular.
M15 Contributed by Chris Black Statement [266]
IfBis singular, then there exists a vector x6=0so that x2N(B). Thus,Bx=0, soA(Bx) = (AB)x=0,
sox2N(AB). Since the null space of ABis not trivial, ABis a nonsingular matrix.
M80 Contributed by Robert Beezer Statement [266]
Take
A=1 0
0 0
B=0 0
1 0
ThenAis already in reduced row-echelon form, and by swapping rows, Brow-reduces to A. So the product
of the row-echelon forms of AisAA=A6=O. However, the product ABis the 22 zero matrix, which
is in reduced-echelon form, and not equal to AA. When you get there, Theorem PEEF [298] or Theorem
EMDRO [425] might shed some light on why we would not expect this statement to be true in general.
IfAandBare nonsingular, then ABis nonsingular (Theorem NPNT [259]), and all three matrices
A,BandABrow-reduce to the identity matrix (Theorem NMRRI [84]). By Theorem MMIM [229], the
desired relationship is true.
Version 2.30
272 Section MINM Matrix Inverses and Nonsingular Matrices
Version 2.30
Section CRS Column and Row Spaces 273
Section CRS
Column and Row Spaces
Theorem SLSLC [112] showed us that there is a natural correspondence between solutions to linear sys-
tems and linear combinations of the columns of the coecient matrix. This idea motivates the following
important denition.
Denition CSM
Column Space of a Matrix
Suppose that Ais anmnmatrix with columns fA1;A2;A3; :::; Ang. Then the column space ofA,
writtenC(A), is the subset of Cmcontaining all linear combinations of the columns of A,
C(A) =hfA1;A2;A3; :::; Angi
(This denition contains Notation CSM.) 4
Some authors refer to the column space of a matrix as the range , but we will reserve this term for use
with linear transformations (Denition RLT [563]).
Subsection CSSE
Column Spaces and Systems of Equations
Upon encountering any new set, the rst question we ask is what objects are in the set, and which objects
are not? Here's an example of one way to answer this question, and it will motivate a theorem that will
then answer the question precisely.
Example CSMCS
Column space of a matrix and consistent systems
Archetype D [795] and Archetype E [799] are linear systems of equations, with an identical 3 4 coecient
matrix, which we call Ahere. However, Archetype D [795] is consistent, while Archetype E [799] is not.
We can explain this dierence by employing the column space of the matrix A.
The column vector of constants, b, in Archetype D [795] is
b=2
48
12
43
5
One solution toLS(A;b), as listed, is
x=2
6647
8
1
33
775
By Theorem SLSLC [112], we can summarize this solution as a linear combination of the columns of A
that equals b,
72
42
3
13
5+ 82
41
4
13
5+ 12
47
5
43
5+ 32
4 7
6
53
5=2
48
12
43
5=b:
This equation says that bis a linear combination of the columns of A, and then by Denition CSM [271],
we can say that b2C(A).
Version 2.30
274 Section CRS Column and Row Spaces
On the other hand, Archetype E [799] is the linear system LS(A;c), where the vector of constants is
c=2
42
3
23
5
and this system of equations is inconsistent. This means c62C(A), for if it were, then it would equal a
linear combination of the columns of Aand Theorem SLSLC [112] would lead us to a solution of the system
LS(A;c).
So if we x the coecient matrix, and vary the vector of constants, we can sometimes nd consistent
systems, and sometimes inconsistent systems. The vectors of constants that lead to consistent systems
are exactly the elements of the column space. This is the content of the next theorem, and since it is an
equivalence, it provides an alternate view of the column space.
Theorem CSCS
Column Spaces and Consistent Systems
SupposeAis anmnmatrix and bis a vector of size m. Then b2C(A) if and only ifLS(A;b) is
consistent.
Proof ()) Suppose b2C(A). Then we can write bas some linear combination of the columns of A. By
Theorem SLSLC [112] we can use the scalars from this linear combination to form a solution to LS(A;b),
so this system is consistent.
(() IfLS(A;b) is consistent, there is a solution that may be used with Theorem SLSLC [112] to write
bas a linear combination of the columns of A. This qualies bfor membership in C(A).
This theorem tells us that asking if the system LS(A;b) is consistent is exactly the same question as
asking if bis in the column space of A. Or equivalently, it tells us that the column space of the matrix A
is precisely those vectors of constants, b, that can be paired with Ato create a system of linear equations
LS(A;b) that is consistent.
Employing Theorem SLEMM [224] we can form the chain of equivalences
b2C(A)() LS (A;b) is consistent()Ax=bfor some x
Thus, an alternative (and popular) denition of the column space of an mnmatrixAis
C(A) =fy2Cmjy=Axfor some x2Cng=fAxjx2CngCm
We recognize this as saying create allthe matrix vector products possible with the matrix Aby letting x
range over all of the possibilities. By Denition MVP [223] we see that this means take all possible linear
combinations of the columns of A| precisely the denition of the column space (Denition CSM [271])
we have chosen.
Notice how this formulation of the column space looks very much like the denition of the null space of
a matrix (Denition NSM [73]), but for a rectangular matrix the column vectors of C(A) andN(A) have
dierent sizes, so the sets are very dierent.
Given a vector band a matrix Ait is now very mechanical to test if b2C(A). Form the linear system
LS(A;b), row-reduce the augmented matrix, [ Ajb], and test for consistency with Theorem RCLS [58].
Here's an example of this procedure.
Example MCSM
Membership in the column space of a matrix
Consider the column space of the 3 4 matrixA,
A=2
43 2 1 4
1 1 2 3
2 4 6 83
5
Version 2.30
Subsection CRS.CSSE Column Spaces and Systems of Equations 275
We rst show that v=2
418
6
123
5is in the column space of A,v2C(A). Theorem CSCS [272] says we need
only check the consistency of LS(A;v). Form the augmented matrix and row-reduce,
2
43 2 1 4 18
1 1 2 3 6
2 4 6 8 123
5RREF !2
410 1 2 6
01 1 1 0
0 0 0 0 03
5
Without a leading 1 in the nal column, Theorem RCLS [58] tells us the system is consistent and therefore
by Theorem CSCS [272], v2C(A).
If we wished to demonstrate explicitly that vis a linear combination of the columns of A, we can
nd a solution (any solution) of LS(A;v) and use Theorem SLSLC [112] to construct the desired linear
combination. For example, set the free variables to x3= 2 andx4= 1. Then a solution has x2= 1 and
x1= 6. Then by Theorem SLSLC [112],
v=2
418
6
123
5= 62
43
1
23
5+ 12
42
1
43
5+ 22
41
2
63
5+ 12
4 4
3
83
5
Now we show that w=2
42
1
33
5is not in the column space of A,w62C(A). Theorem CSCS [272] says we
need only check the consistency of LS(A;w). Form the augmented matrix and row-reduce,
2
43 2 1 4 2
1 1 2 3 1
2 4 6 8 33
5RREF !2
410 1 2 0
01 1 1 0
0 0 0 0 13
5
With a leading 1 in the nal column, Theorem RCLS [58] tells us the system is inconsistent and therefore
by Theorem CSCS [272], w62C(A).
Theorem CSCS [272] completes a collection of three theorems, and one denition, that deserve comment.
Many questions about spans, linear independence, null space, column spaces and similar objects can be
converted to questions about systems of equations (homogeneous or not), which we understand well from
our previous results, especially those in Chapter SLE [3]. These previous results include theorems like
Theorem RCLS [58] which allows us to quickly decide consistency of a system, and Theorem BNS [160]
which allows us to describe solution sets for homogeneous systems compactly as the span of a linearly
independent set of column vectors.
The table below lists these for denitions and theorems along with a brief reminder of the statement
and an example of how the statement is used.
Denition NSM [73]
Synopsis Null space is solution set of homogeneous system
Example General solution sets described by Theorem PSPHS [124]
Theorem SLSLC [112]
Synopsis Solutions for linear combinations with unknown scalars
Example Deciding membership in spans
Theorem SLEMM [224]
Synopsis System of equations represented by matrix-vector product
Example Solution toLS(A;b) isA 1bwhenAis nonsingular
Theorem CSCS [272]
Synopsis Column space vectors create consistent systems
Example Deciding membership in column spaces
Version 2.30
276 Section CRS Column and Row Spaces
Subsection CSSOC
Column Space Spanned by Original Columns
So we have a foolproof, automated procedure for determining membership in C(A). While this works just
ne a vector at a time, we would like to have a more useful description of the set C(A) as a whole. The
next example will preview the rst of two fundamental results about the column space of a matrix.
Example CSTW
Column space, two ways
Consider the 57 matrixA,2
666642 4 1 1 1 4 4
1 2 1 0 2 4 7
0 0 1 4 1 8 7
1 2 1 2 1 9 6
2 4 1 3 1 2 23
77775
According to the denition (Denition CSM [271]), the column space of Ais
C(A) =*8
>>>><
>>>>:2
666642
1
0
1
23
77775;2
666644
2
0
2
43
77775;2
666641
1
1
1
13
77775;2
66664 1
0
4
2
33
77775;2
666641
2
1
1
13
77775;2
666644
4
8
9
23
77775;2
666644
7
7
6
23
777759
>>>>=
>>>>;+
While this is a concise description of an innite set, we might be able to describe the span with fewer than
seven vectors. This is the substance of Theorem BS [180]. So we take these seven vectors and make them
the columns of matrix, which is simply the original matrix Aagain. Now we row-reduce,
2
666642 4 1 1 1 4 4
1 2 1 0 2 4 7
0 0 1 4 1 8 7
1 2 1 2 1 9 6
2 4 1 3 1 2 23
77775RREF !2
66666412 0 0 0 3 1
0 0 10 0 1 0
0 0 0 10 2 1
0 0 0 0 1 1 3
0 0 0 0 0 0 03
777775
The pivot columns are D=f1;3;4;5g, so we can create the set
T=8
>>>><
>>>>:2
666642
1
0
1
23
77775;2
666641
1
1
1
13
77775;2
66664 1
0
4
2
33
77775;2
666641
2
1
1
13
777759
>>>>=
>>>>;
and know thatC(A) =hTiandTis a linearly independent set of columns from the set of columns of A.
We will now formalize the previous example, which will make it trivial to determine a linearly inde-
pendent set of vectors that will span the column space of a matrix, and is constituted of just columns of
A.
Theorem BCS
Basis of the Column Space
Suppose that Ais anmnmatrix with columns A1;A2;A3; :::; An, andBis a row-equivalent matrix in
reduced row-echelon form with rnonzero rows. Let D=fd1; d2; d3; :::; drgbe the set of column indices
whereBhas leading 1's. Let T=fAd1;Ad2;Ad3; :::; Adrg. Then
Version 2.30
Subsection CRS.CSSOC Column Space Spanned by Original Columns 277
1.Tis a linearly independent set.
2.C(A) =hTi.
Proof Denition CSM [271] describes the column space as the span of the set of columns of A. Theorem
BS [180] tells us that we can reduce the set of vectors used in a span. If we apply Theorem BS [180] to
C(A), we would collect the columns of Ainto a matrix (which would just be Aagain) and bring the matrix
to reduced row-echelon form, which is the matrix Bin the statement of the theorem. In this case, the
conclusions of Theorem BS [180] applied to A,BandC(A) are exactly the conclusions we desire.
This is a nice result since it gives us a handful of vectors that describe the entire column space (through
the span), and we believe this set is as small as possible because we cannot create any more relations of
linear dependence to trim it down further. Furthermore, we dened the column space (Denition CSM
[271]) as all linear combinations of the columns of the matrix, and the elements of the set Sare still columns
of the matrix (we won't be so lucky in the next two constructions of the column space).
Procedurally this theorem is extremely easy to apply. Row-reduce the original matrix, identify r
columns with leading 1's in this reduced matrix, and grab the corresponding columns of the original
matrix. But it is still important to study the proof of Theorem BS [180] and its motivation in Example
COV [177] which lie at the root of this theorem. We'll trot through an example all the same.
Example CSOCD
Column space, original columns, Archetype D
Let's determine a compact expression for the entire column space of the coecient matrix of the system
of equations that is Archetype D [795]. Notice that in Example CSMCS [271] we were only determining if
individual vectors were in the column space or not, now we are describing the entire column space.
To start with the application of Theorem BCS [274], call the coecient matrix A
A=2
42 1 7 7
3 4 5 6
1 1 4 53
5:
and row-reduce it to reduced row-echelon form,
B=2
410 3 2
011 3
0 0 0 03
5:
There are leading 1's in columns 1 and 2, so D=f1;2g. To construct a set that spans C(A), just grab the
columns of Aindicated by the set D, so
C(A) =*8
<
:2
42
3
13
5;2
41
4
13
59
=
;+
:
That's it.
In Example CSMCS [271] we determined that the vector
c=2
42
3
23
5
was not in the column space of A. Try to write cas a linear combination of the rst two columns of A.
What happens?
Version 2.30
278 Section CRS Column and Row Spaces
Also in Example CSMCS [271] we determined that the vector
b=2
48
12
43
5
wasin the column space of A. Try to write bas a linear combination of the rst two columns of A. What
happens? Did you nd a unique solution to this question? Hmmmm.
Subsection CSNM
Column Space of a Nonsingular Matrix
Let's specialize to square matrices and contrast the column spaces of the coecient matrices in Archetype
A [781] and Archetype B [786].
Example CSAA
Column space of Archetype A
The coecient matrix in Archetype A [781] is
A=2
41 1 2
2 1 1
1 1 03
5
which row-reduces to2
410 1
01 1
0 0 03
5:
Columns 1 and 2 have leading 1's, so by Theorem BCS [274] we can write
C(A) =hfA1;A2gi=*8
<
:2
41
2
13
5;2
4 1
1
13
59
=
;+
:
We want to show in this example that C(A)6=C3. So take, for example, the vector b=2
41
3
23
5. Then there
is no solution to the system LS(A;b), or equivalently, it is not possible to write bas a linear combination
ofA1andA2. Try one of these two computations yourself. (Or try both!). Since b62C(A), the column
space ofAcannot be all of C3. So by varying the vector of constants, it is possible to create inconsistent
systems of equations with this coecient matrix (the vector bbeing one such example).
In Example MWIAA [244] we wished to show that the coecient matrix from Archetype A [781] was
not invertible as a rst example of a matrix without an inverse. Our device there was to nd an inconsistent
linear system with Aas the coecient matrix. The vector of constants in that example was b, deliberately
chosen outside the column space of A.
Example CSAB
Column space of Archetype B
The coecient matrix in Archetype B [786], call it Bhere, is known to be nonsingular (see Example NM
[84]). By Theorem NMUS [86], the linear system LS(B;b) has a (unique) solution for every choice of b.
Theorem CSCS [272] then says that b2C(B) for all b2C3. Stated dierently, there is no way to build
Version 2.30
Subsection CRS.CSNM Column Space of a Nonsingular Matrix 279
an inconsistent system with the coecient matrix B, but then we knew that already from Theorem NMUS
[86].
Example CSAA [276] and Example CSAB [276] together motivate the following equivalence, which says
that nonsingular matrices have column spaces that are as big as possible.
Theorem CSNM
Column Space of a Nonsingular Matrix
SupposeAis a square matrix of size n. ThenAis nonsingular if and only if C(A) =Cn.
Proof ()) SupposeAis nonsingular. We wish to establish the set equality C(A) =Cn. By Denition
CSM [271],C(A)Cn.
To show that CnC(A) choose b2Cn. By Theorem NMUS [86], we know the linear system LS(A;b)
has a (unique) solution and therefore is consistent. Theorem CSCS [272] then says that b2C(A). So by
Denition SE [762], C(A) =Cn.
(() Ifeiis columniof thennidentity matrix (Denition SUV [197]) and by hypothesis C(A) =Cn,
thenei2C(A) for 1in. By Theorem CSCS [272], the system LS(A;ei) is consistent for 1 in.
Letbidenote any one particular solution to LS(A;ei), 1in.
Dene thennmatrixB= [b1jb2jb3j:::jbn]. Then
AB=A[b1jb2jb3j:::jbn]
= [Ab1jAb2jAb3j:::jAbn] Denition MM [226]
= [e1je2je3j:::jen]
=In Denition SUV [197]
So the matrix Bis a \right-inverse" for A. By Theorem NMRRI [84], Inis a nonsingular matrix, so
by Theorem NPNT [259] both AandBare nonsingular. Thus, in particular, Ais nonsingular. (Travis
Osborne contributed to this proof.)
With this equivalence for nonsingular matrices we can update our list, Theorem NME3 [261].
Theorem NME4
Nonsingular Matrix Equivalences, Round 4
Suppose that Ais a square matrix of size n. The following are equivalent.
1.Ais nonsingular.
2.Arow-reduces to the identity matrix.
3. The null space of Acontains only the zero vector, N(A) =f0g.
4. The linear system LS(A;b) has a unique solution for every possible choice of b.
5. The columns of Aare a linearly independent set.
6.Ais invertible.
7. The column space of AisCn,C(A) =Cn.
Proof Since Theorem CSNM [277] is an equivalence, we can add it to the list in Theorem NME3 [261].
Version 2.30
280 Section CRS Column and Row Spaces
Subsection RSM
Row Space of a Matrix
The rows of a matrix can be viewed as vectors, since they are just lists of numbers, arranged horizontally.
So we will transpose a matrix, turning rows into columns, so we can then manipulate rows as column
vectors. As a result we will be able to make some new connections between row operations and solutions
to systems of equations. OK, here is the second primary denition of this section.
Denition RSM
Row Space of a Matrix
SupposeAis anmnmatrix. Then the row space ofA,R(A), is the column space of At, i.e.R(A) =
C
At
.
(This denition contains Notation RSM.) 4
Informally, the row space is the set of all linear combinations of the rows of A. However, we write
the rows as column vectors, thus the necessity of using the transpose to make the rows into columns.
Additionally, with the row space dened in terms of the column space, all of the previous results of this
section can be applied to row spaces.
Notice that if Ais a rectangular mnmatrix, thenC(A)Cm, whileR(A)Cnand the two sets
are not comparable since they do not even hold objects of the same type. However, when Ais square of
sizen, bothC(A) andR(A) are subsets of Cn, though usually the sets will not be equal (but see Exercise
CRS.M20 [287]).
Example RSAI
Row space of Archetype I
The coecient matrix in Archetype I [816] is
I=2
6641 4 0 1 0 7 9
2 8 1 3 9 13 7
0 0 2 3 4 12 8
1 4 2 4 8 31 373
775:
To build the row space, we transpose the matrix,
It=2
6666666641 2 0 1
4 8 0 4
0 1 2 2
1 3 3 4
0 9 4 8
7 13 12 31
9 7 8 373
777777775
Then the columns of this matrix are used in a span to build the row space,
R(I) =C
It
=*8
>>>>>>>><
>>>>>>>>:2
6666666641
4
0
1
0
7
93
777777775;2
6666666642
8
1
3
9
13
73
777777775;2
6666666640
0
2
3
4
12
83
777777775;2
666666664 1
4
2
4
8
31
373
7777777759
>>>>>>>>=
>>>>>>>>;+
:
Version 2.30
Subsection CRS.RSM Row Space of a Matrix 281
However, we can use Theorem BCS [274] to get a slightly better description. First, row-reduce It,
2
666666666410 0 31
7
01012
7
0 0 113
7
0 0 0 0
0 0 0 0
0 0 0 0
0 0 0 03
7777777775:
Since there are leading 1's in columns with indices D=f1;2;3g, the column space of Itcan be spanned
by just the rst three columns of It,
R(I) =C
It
=*8
>>>>>>>><
>>>>>>>>:2
6666666641
4
0
1
0
7
93
777777775;2
6666666642
8
1
3
9
13
73
777777775;2
6666666640
0
2
3
4
12
83
7777777759
>>>>>>>>=
>>>>>>>>;+
:
The row space would not be too interesting if it was simply the column space of the transpose. However,
when we do row operations on a matrix we have no eect on the many linear combinations that can be
formed with the rows of the matrix. This is stated more carefully in the following theorem.
Theorem REMRS
Row-Equivalent Matrices have equal Row Spaces
SupposeAandBare row-equivalent matrices. Then R(A) =R(B).
Proof Two matrices are row-equivalent (Denition REM [31]) if one can be obtained from another by
a sequence of (possibly many) row operations. We will prove the theorem for two matrices that dier
by a single row operation, and then this result can be applied repeatedly to get the full statement of the
theorem. The row spaces of AandBare spans of the columns of their transposes. For each row operation
we perform on a matrix, we can dene an analogous operation on the columns. Perhaps we should call
these column operations . Instead, we will still call them row operations, but we will apply them to the
columns of the transposes.
Refer to the columns of AtandBtasAiandBi, 1im. The row operation that switches rows will
just switch columns of the transposed matrices. This will have no eect on the possible linear combinations
formed by the columns.
Suppose that Btis formed from Atby multiplying column Atby6= 0. In other words, Bt=At,
andBi=Aifor alli6=t. We need to establish that two sets are equal, C
At
=C
Bt
. We will take a
generic element of one and show that it is contained in the other.
1B1+2B2+3B3++tBt++mBm
=1A1+2A2+3A3++t(At) ++mAm
=1A1+2A2+3A3++ (t)At++mAm
says thatC
Bt
C
At
. Similarly,
1A1+
2A2+
3A3++
tAt++
mAm
=
1A1+
2A2+
3A3++
t
At++
mAm
Version 2.30
282 Section CRS Column and Row Spaces
=
1A1+
2A2+
3A3++
t
(At) ++
mAm
=
1B1+
2B2+
3B3++
t
Bt++
mBm
says thatC
At
C
Bt
. SoR(A) =C
At
=C
Bt
=R(B) when a single row operation of the second
type is performed.
Suppose now that Btis formed from Atby replacing AtwithAs+Atfor some2Cands6=t. In
other words, Bt=As+At, and Bi=Aifori6=t.
1B1+2B2+3B3++sBs++tBt++mBm
=1A1+2A2+3A3++sAs++t(As+At) ++mAm
=1A1+2A2+3A3++sAs++ (t)As+tAt++mAm
=1A1+2A2+3A3++sAs+ (t)As++tAt++mAm
=1A1+2A2+3A3++ (s+t)As++tAt++mAm
says thatC
Bt
C
At
. Similarly,
1A1+
2A2+
3A3++
sAs++
tAt++
mAm
=
1A1+
2A2+
3A3++
sAs++ (
tAs+
tAs) +
tAt++
mAm
=
1A1+
2A2+
3A3++ (
tAs) +
sAs++ (
tAs+
tAt) ++
mAm
=
1A1+
2A2+
3A3++ (
t+
s)As++
t(As+At) ++
mAm
=
1B1+
2B2+
3B3++ (
t+
s)Bs++
tBt++
mBm
says thatC
At
C
Bt
. SoR(A) =C
At
=C
Bt
=R(B) when a single row operation of the third
type is performed.
So the row space of a matrix is preserved by each row operation, and hence row spaces of row-equivalent
matrices are equal sets.
Example RSREM
Row spaces of two row-equivalent matrices
In Example TREM [31] we saw that the matrices
A=2
42 1 3 4
5 2 2 3
1 1 0 63
5 B=2
41 1 0 6
3 0 2 9
2 1 3 43
5
are row-equivalent by demonstrating a sequence of two row operations that converted AintoB. Applying
Theorem REMRS [279] we can say
R(A) =*8
>><
>>:2
6642
1
3
43
775;2
6645
2
2
33
775;2
6641
1
0
63
7759
>>=
>>;+
=*8
>><
>>:2
6641
1
0
63
775;2
6643
0
2
93
775;2
6642
1
3
43
7759
>>=
>>;+
=R(B)
Theorem REMRS [279] is at its best when one of the row-equivalent matrices is in reduced row-echelon
form. The vectors that correspond to the zero rows can be ignored. (Who needs the zero vector when
building a span? See Exercise LI.T10 [166].) The echelon pattern insures that the nonzero rows yield
vectors that are linearly independent. Here's the theorem.
Theorem BRS
Basis for the Row Space
Suppose that Ais a matrix and Bis a row-equivalent matrix in reduced row-echelon form. Let Sbe the
set of nonzero columns of Bt. Then
Version 2.30
Subsection CRS.RSM Row Space of a Matrix 283
1.R(A) =hSi.
2.Sis a linearly independent set.
Proof From Theorem REMRS [279] we know that R(A) =R(B). IfBhas any zero rows, these
correspond to columns of Btthat are the zero vector. We can safely toss out the zero vector in the span
construction, since it can be recreated from the nonzero vectors by a linear combination where all the
scalars are zero. So R(A) =hSi.
SupposeBhasrnonzero rows and let D=fd1; d2; d3; :::; drgdenote the column indices of Bthat
have a leading one in them. Denote the rcolumn vectors of Bt, the vectors in S, asB1;B2;B3; :::; Br.
To show that Sis linearly independent, start with a relation of linear dependence
1B1+2B2+3B3++rBr=0
Now consider this vector equality in location di. SinceBis in reduced row-echelon form, the entries of
columndiofBare all zero, except for a (leading) 1 in row i. Thus, inBt, rowdiis all zeros, excepting a
1 in column i. So, for 1ir,
0 = [0]diDenition ZCV [28]
= [1B1+2B2+3B3++rBr]diDenition RLDCV [153]
= [1B1]di+ [2B2]di+ [3B3]di++ [rBr]di+ Denition MA [207]
=1[B1]di+2[B2]di+3[B3]di++r[Br]di+ Denition MSM [208]
=1(0) +2(0) +3(0) ++i(1) ++r(0) Denition RREF [33]
=i
So we conclude that i= 0 for all 1ir, establishing the linear independence of S(Denition LICV
[153]).
Example IAS
Improving a span
Suppose in the course of analyzing a matrix (its column space, its null space, its. . . ) we encounter the
following set of vectors, described by a span
X=*8
>>>><
>>>>:2
666641
2
1
6
63
77775;2
666643
1
2
1
63
77775;2
666641
1
0
1
23
77775;2
66664 3
2
3
6
103
777759
>>>>=
>>>>;+
LetAbe the matrix whose rows are the vectors in X, so by design X=R(A),
A=2
6641 2 1 6 6
3 1 2 1 6
1 1 0 1 2
3 2 3 6 103
775
Row-reduce Ato form a row-equivalent matrix in reduced row-echelon form,
B=2
66410 0 2 1
010 3 1
0 0 1 2 5
0 0 0 0 03
775
Version 2.30
284 Section CRS Column and Row Spaces
Then Theorem BRS [280] says we can grab the nonzero columns of Btand write
X=R(A) =R(B) =*8
>>>><
>>>>:2
666641
0
0
2
13
77775;2
666640
1
0
3
13
77775;2
666640
0
1
2
53
777759
>>>>=
>>>>;+
These three vectors provide a much-improved description of X. There are fewer vectors, and the pattern
of zeros and ones in the rst three entries makes it easier to determine membership in X. And all we
had to do was row-reduce the right matrix and toss out a zero row. Next to row operations themselves,
this is probably the most powerful computational technique at your disposal as it quickly provides a much
improved description of a span, any span.
Theorem BRS [280] and the techniques of Example IAS [281] will provide yet another description of
the column space of a matrix. First we state a triviality as a theorem, so we can reference it later.
Theorem CSRST
Column Space, Row Space, Transpose
SupposeAis a matrix. Then C(A) =R
At
.
Proof
C(A) =C
Att
Theorem TT [212]
=R
At
Denition RSM [278]
So to nd another expression for the column space of a matrix, build its transpose, row-reduce it, toss
out the zero rows, and convert the nonzero rows to column vectors to yield an improved set for the span
construction. We'll do Archetype I [816], then you do Archetype J [820].
Example CSROI
Column space from row operations, Archetype I
To nd the column space of the coecient matrix of Archetype I [816], we proceed as follows. The matrix
is
I=2
6641 4 0 1 0 7 9
2 8 1 3 9 13 7
0 0 2 3 4 12 8
1 4 2 4 8 31 373
775:
The transpose is2
6666666641 2 0 1
4 8 0 4
0 1 2 2
1 3 3 4
0 9 4 8
7 13 12 31
9 7 8 373
777777775:
Row-reduced this becomes,2
666666666410 0 31
7
01012
7
0 0 113
7
0 0 0 0
0 0 0 0
0 0 0 0
0 0 0 03
7777777775:
Version 2.30
Subsection CRS.READ Reading Questions 285
Now, using Theorem CSRST [282] and Theorem BRS [280]
C(I) =R
It
=*8
>><
>>:2
6641
0
0
31
73
775;2
6640
1
0
12
73
775;2
6640
0
1
13
73
7759
>>=
>>;+
:
This is a very nice description of the column space. Fewer vectors than the 7 involved in the denition,
and the pattern of the zeros and ones in the rst 3 slots can be used to advantage. For example, Archetype
I [816] is presented as a consistent system of equations with a vector of constants
b=2
6643
9
1
43
775:
SinceLS(I;b) is consistent, Theorem CSCS [272] tells us that b2C(I). But we could see this quickly
with the following computation, which really only involves any work in the 4th entry of the vectors as the
scalars in the linear combination are dictated by the rst three entries of b.
b=2
6643
9
1
43
775= 32
6641
0
0
31
73
775+ 92
6640
1
0
12
73
775+ 12
6640
0
1
13
73
775
Can you now rapidly construct several vectors, b, so thatLS(I;b) is consistent, and several more so that
the system is inconsistent?
Subsection READ
Reading Questions
1. Write the column space of the matrix below as the span of a set of three vectors and explain your
choice of method. 2
41 3 1 3
2 0 1 1
1 2 1 03
5
2. Suppose that Ais annnnonsingular matrix. What can you say about its column space?
3. Is the vector2
6640
5
2
33
775in the row space of the following matrix? Why or why not?
2
41 3 1 3
2 0 1 1
1 2 1 03
5
Version 2.30
286 Section CRS Column and Row Spaces
Subsection EXC
Exercises
C20 For parts (a), (b) and c, nd a set of linearly independent vectors Xso thatC(A) =hXi, and a set
of linearly independent vectors Yso thatR(A) =hYi.
a).A=2
6641 2 3 1
0 1 1 2
1 1 2 3
1 1 2 13
775
b).A=2
41 2 1 1 1
3 2 1 4 5
0 1 1 1 23
5
c).A=2
666642 1 0
3 0 3
1 2 3
1 1 1
1 1 13
77775
d). From your results in parts (a) - (c), can you formulate a conjecture about the sets XandY?
Contributed by Chris Black
C30 Example CSOCD [275] expresses the column space of the coecient matrix from Archetype D [795]
(call the matrix Ahere) as the span of the rst two columns of A. In Example CSMCS [271] we determined
that the vector
c=2
42
3
23
5
was not in the column space of Aand that the vector
b=2
48
12
43
5
wasin the column space of A. Attempt to write candbas linear combinations of the two vectors in the
span construction for the column space in Example CSOCD [275] and record your observations.
Contributed by Robert Beezer Solution [288]
C31 For the matrix Abelow nd a set of vectors Tmeeting the following requirements: (1) the span of
Tis the column space of A, that is,hTi=C(A), (2)Tis linearly independent, and (3) the elements of T
are columns of A.
A=2
6642 1 4 1 2
1 1 5 1 1
1 2 7 0 1
2 1 8 1 23
775
Contributed by Robert Beezer Solution [288]
C32 In Example CSAA [276], verify that the vector bis not in the column space of the coecient matrix.
Contributed by Robert Beezer
Version 2.30
Subsection CRS.EXC Exercises 287
C33 Find a linearly independent set Sso that the span of S,hSi, is row space of the matrix B, andS
is linearly independent.
B=2
42 3 1 1
1 1 0 1
1 2 3 43
5
Contributed by Robert Beezer Solution [288]
C34 For the 34 matrixAand the column vector y2C4given below, determine if yis in the row
space ofA. In other words, answer the question: y2R(A)?
A=2
4 2 6 7 1
7 3 0 3
8 0 7 63
5 y=2
6642
1
3
23
775
Contributed by Robert Beezer Solution [288]
C35 For the matrix Abelow, nd two dierent linearly independent sets whose spans equal the column
space ofA,C(A), such that
(a) the elements are each columns of A.
(b) the set is obtained by a procedure that is substantially dierent from the procedure you use in part
(a).
A=2
43 5 1 2
1 2 3 3
3 4 7 133
5
Contributed by Robert Beezer Solution [289]
C40 The following archetypes are systems of equations. For each system, write the vector of constants
as a linear combination of the vectors in the span construction for the column space provided by Theorem
BCS [274] (these vectors are listed for each of these archetypes).
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]
Archetype E [799]
Archetype F [803]
Archetype G [808]
Archetype H [812]
Archetype I [816]
Archetype J [820]
Contributed by Robert Beezer
C42 The following archetypes are either matrices or systems of equations with coecient matrices. For
each matrix, compute a set of column vectors such that (1) the vectors are columns of the matrix, (2) the
set is linearly independent, and (3) the span of the set is the column space of the matrix. See Theorem
BCS [274].
Archetype A [781]
Archetype B [786]
Archetype C [791]
Version 2.30
288 Section CRS Column and Row Spaces
Archetype D [795]/Archetype E [799]
Archetype F [803]
Archetype G [808]/Archetype H [812]
Archetype I [816]
Archetype J [820]
Archetype K [825]
Archetype L [829]
Contributed by Robert Beezer
C50 The following archetypes are either matrices or systems of equations with coecient matrices. For
each matrix, compute a set of column vectors such that (1) the set is linearly independent, and (2) the
span of the set is the row space of the matrix. See Theorem BRS [280].
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]/Archetype E [799]
Archetype F [803]
Archetype G [808]/Archetype H [812]
Archetype I [816]
Archetype J [820]
Archetype K [825]
Archetype L [829]
Contributed by Robert Beezer
C51 The following archetypes are either matrices or systems of equations with coecient matrices. For
each matrix, compute the column space as the span of a linearly independent set as follows: transpose the
matrix, row-reduce, toss out zero rows, convert rows into column vectors. See Example CSROI [282].
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]/Archetype E [799]
Archetype F [803]
Archetype G [808]/Archetype H [812]
Archetype I [816]
Archetype J [820]
Archetype K [825]
Archetype L [829]
Contributed by Robert Beezer
C52 The following archetypes are systems of equations. For each dierent coecient matrix build two
new vectors of constants. The rst should lead to a consistent system and the second should lead to an
inconsistent system. Descriptions of the column space as spans of linearly independent sets of vectors with
\nice patterns" of zeros and ones might be most useful and instructive in connection with this exercise.
(See the end of Example CSROI [282].)
Archetype A [781]
Version 2.30
Subsection CRS.EXC Exercises 289
Archetype B [786]
Archetype C [791]
Archetype D [795]/Archetype E [799]
Archetype F [803]
Archetype G [808]/Archetype H [812]
Archetype I [816]
Archetype J [820]
Contributed by Robert Beezer
M10 For the matrix Ebelow, nd vectors bandcso that the system LS(E;b) is consistent and
LS(E;c) is inconsistent.
E=2
4 2 1 1 0
3 1 0 2
4 1 1 63
5
Contributed by Robert Beezer Solution [289]
M20 Usually the column space and null space of a matrix contain vectors of dierent sizes. For a square
matrix, though, the vectors in these two sets are the same size. Usually the two sets will be dierent.
Construct an example of a square matrix where the column space and null space are equal.
Contributed by Robert Beezer Solution [290]
M21 We have a variety of theorems about how to create column spaces and row spaces and they frequently
involve row-reducing a matrix. Here is a procedure that some try to use to get a column space. Begin
with anmnmatrixAand row-reduce to a matrix Bwith columns B1;B2;B3; :::; Bn. Then form the
column space of Aas
C(A) =hfB1;B2;B3; :::; Bngi=C(B)
This is notnot a legitimate procedure, and therefore is nota theorem. Construct an example to show that
the procedure will not in general create the column space of A.
Contributed by Robert Beezer Solution [290]
T40 Suppose that Ais anmnmatrix and Bis annpmatrix. Prove that the column space of ABis
a subset of the column space of A, that isC(AB)C(A). Provide an example where the opposite is false,
in other words give an example where C(A)6C(AB). (Compare with Exercise MM.T40 [238].)
Contributed by Robert Beezer Solution [290]
T41 Suppose that Ais anmnmatrix and Bis annnnonsingular matrix. Prove that the column
space ofAis equal to the column space of AB, that isC(A) =C(AB). (Compare with Exercise MM.T41
[238] and Exercise CRS.T40 [287].)
Contributed by Robert Beezer Solution [290]
T45 Suppose that Ais anmnmatrix and Bis annmmatrix where ABis a nonsingular matrix.
Prove that
(1)N(B) =f0g
(2)C(B)\N(A) =f0g
Discuss the case when m=nin connection with Theorem NPNT [259].
Contributed by Robert Beezer Solution [290]
Version 2.30
290 Section CRS Column and Row Spaces
Subsection SOL
Solutions
C30 Contributed by Robert Beezer Statement [284]
In each case, begin with a vector equation where one side contains a linear combination of the two vectors
from the span construction that gives the column space of Awith unknowns for scalars, and then use
Theorem SLSLC [112] to set up a system of equations. For c, the corresponding system has no solution,
as we would expect.
Forbthere is a solution, as we would expect. What is interesting is that the solution is unique. This
is a consequence of the linear independence of the set of two vectors in the span construction. If we wrote
bas a linear combination of all four columns of A, then there would be innitely many ways to do this.
C31 Contributed by Robert Beezer Statement [284]
Theorem BCS [274] is the right tool for this problem. Row-reduce this matrix, identify the pivot columns
and then grab the corresponding columns of Afor the setT. The matrix Arow-reduces to
2
666410 3 0 0
01 2 0 0
0 0 0 10
0 0 0 0 13
7775
SoD=f1;2;4;5gand then
T=fA1;A2;A4;A5g=8
>><
>>:2
6642
1
1
23
775;2
6641
1
2
13
775;2
664 1
1
0
13
775;2
6642
1
1
23
7759
>>=
>>;
has the requested properties.
C33 Contributed by Robert Beezer Statement [285]
Theorem BRS [280] is the most direct route to a set with these properties. Row-reduce, toss zero rows,
keep the others. You could also transpose the matrix, then look for the column space by row-reducing the
transpose and applying Theorem BCS [274]. We'll do the former,
BRREF !2
410 1 2
01 1 1
0 0 0 03
5
So the setSis
S=8
>><
>>:2
6641
0
1
23
775;2
6640
1
1
13
7759
>>=
>>;
C34 Contributed by Robert Beezer Statement [285]
y2R(A)()y2C
At
Denition RSM [278]
() LS
At;y
is consistent Theorem CSCS [272]
Version 2.30
Subsection CRS.SOL Solutions 291
The augmented matrix
Aty
row reduces to
2
666410 0 0
010 0
0 0 10
0 0 0 13
7775
and with a leading 1 in the nal column Theorem RCLS [58] tells us the linear system is inconsistent and
soy62R(A).
C35 Contributed by Robert Beezer Statement [285]
(a) By Theorem BCS [274] we can row-reduce A, identify pivot columns with the set D, and \keep" those
columns of Aand we will have a set with the desired properties.
ARREF !2
410 13 19
01 8 11
0 0 0 03
5
So we have the set of pivot columns D=f1;2gand we \keep" the rst two columns of A,
8
<
:2
43
1
33
5;2
45
2
43
59
=
;
(b) We can view the column space as the row space of the transpose (Theorem CSRST [282]). We can
get a basis of the row space of a matrix quickly by bringing the matrix to reduced row-echelon form and
keeping the nonzero rows as column vectors (Theorem BRS [280]). Here goes.
AtRREF !2
66410 2
01 3
0 0 0
0 0 03
775
Taking the nonzero rows and tilting them up as columns gives us
8
<
:2
41
0
23
5;2
40
1
33
59
=
;
An approach based on the matrix Lfrom extended echelon form (Denition EEF [297]) and Theorem FS
[299] will work as well as an alternative approach.
M10 Contributed by Robert Beezer Statement [287]
Any vector from C3will lead to a consistent system, and therefore there is no vector that will lead to an
inconsistent system.
How do we convince ourselves of this? First, row-reduce E,
ERREF !2
410 0 1
010 1
0 0 113
5
If we augment Ewith any vector of constants, and row-reduce the augmented matrix, we will never nd
a leading 1 in the nal column, so by Theorem RCLS [58] the system will always be consistent.
Said another way, the column space of Eis all of C3,C(E) =C3. So by Theorem CSCS [272] any
vector of constants will create a consistent system (and none will create an inconsistent system).
Version 2.30
292 Section CRS Column and Row Spaces
M20 Contributed by Robert Beezer Statement [287]
The 22 matrix1 1
1 1
hasC(A) =N(A) =1
1
.
M21 Contributed by Robert Beezer Statement [287]
Begin with a matrix A(of any size) that does not have any zero rows, but which when row-reduced to B
yields at least one row of zeros. Such a matrix should be easy to construct (or nd, like say from Archetype
A [781]).
C(A) will contain some vectors whose nal slot (entry m) is non-zero, however, every column vector
from the matrix Bwill have a zero in slot mand so every vector in C(B) will also contain a zero in the
nal slot. This means that C(A)6=C(B), since we have vectors in C(A) that cannot be elements of C(B).
T40 Contributed by Robert Beezer Statement [287]
Choose x2C(AB). Then by Theorem CSCS [272] there is a vector wthat is a solution to LS(AB;x).
Dene the vector ybyy=Bw. We're set,
Ay=A(Bw) Denition of y
= (AB)w Theorem MMA [231]
=x w solution toLS(AB;x)
This says thatLS(A;x) is a consistent system, and by Theorem CSCS [272], we see that x2C(A) and
thereforeC(AB)C(A).
For an example where C(A)6C(AB) chooseAto be any nonzero matrix and choose Bto be a zero
matrix. ThenC(A)6=f0gandC(AB) =C(O) =f0g.
T41 Contributed by Robert Beezer Statement [287]
From the solution to Exercise CRS.T40 [287] we know that C(AB)C(A). So to establish the set equality
(Denition SE [762]) we need to show that C(A)C(AB).
Choose x2C(A). By Theorem CSCS [272] the linear system LS(A;x) is consistent, so let ybe one
such solution. Because Bis nonsingular, and linear system using Bas a coecient matrix will have a
solution (Theorem NMUS [86]). Let wbe the unique solution to the linear system LS(B;y). All set, here
we go,
(AB)w=A(Bw) Theorem MMA [231]
=Ay w solution toLS(B;y)
=x y solution toLS(A;x)
This says that the linear system LS(AB;x) is consistent, so by Theorem CSCS [272], x2C(AB). So
C(A)C(AB).
T45 Contributed by Robert Beezer Statement [287]
First, 02N(B) trivially. Now suppose that x2N(B). Then
ABx=A(Bx) Theorem MMA [231]
=A0 x 2N(B)
=0 Theorem MMZM [229]
Version 2.30
Subsection CRS.SOL Solutions 293
Since we have assumed ABis nonsingular, Denition NM [83] implies that x=0.
Second, 02C(B) and 02N(A) trivially, and so the zero vector is in the intersection as well (Denition
SI [763]). Now suppose that y2C(B)\N(A). Because y2C(B), Theorem CSCS [272] says the system
LS(B;y) is consistent. Let x2Cnbe one solution to this system. Then
ABx=A(Bx) Theorem MMA [231]
=Ay x solution toLS(B;y)
=0 y 2N(A)
Since we have assumed ABis nonsingular, Denition NM [83] implies that x=0. Then y=Bx=B0=0.
WhenABis nonsingular and m=nwe know that the rst condition, N(B) =f0g, means that
Bis nonsingular (Theorem NMTNS [86]). Because Bis nonsingular Theorem CSNM [277] implies that
C(B) =Cm. In order to have the second condition fullled, C(B)\N(A) =f0g, we must realize that
N(A) =f0g. However, a second application of Theorem NMTNS [86] shows that Amust be nonsingular.
This reproduces Theorem NPNT [259].
Version 2.30
294 Section CRS Column and Row Spaces
Version 2.30
Section FS Four Subsets 295
Section FS
Four Subsets
There are four natural subsets associated with a matrix. We have met three already: the null space,
the column space and the row space. In this section we will introduce a fourth, the left null space. The
objective of this section is to describe one procedure that will allow us to nd linearly independent sets
that span each of these four sets of column vectors. Along the way, we will make a connection with the
inverse of a matrix, so Theorem FS [299] will tie together most all of this chapter (and the entire course
so far).
Subsection LNS
Left Null Space
Denition LNS
Left Null Space
SupposeAis anmnmatrix. Then the left null space is dened asL(A) =N
At
Cm.
(This denition contains Notation LNS.) 4
The left null space will not feature prominently in the sequel, but we can explain its name and connect
it to row operations. Suppose y2L(A). Then by Denition LNS [293], Aty=0. We can then write
0t=
AtytDenition LNS [293]
=yt
AttTheorem MMT [232]
=ytA Theorem TT [212]
The product ytAcan be viewed as the components of yacting as the scalars in a linear combination of
therows ofA. And the result is a \row vector", 0tthat is totally zeros. When we apply a sequence
of row operations to a matrix, each row of the resulting matrix is some linear combination of the rows.
These observations tell us that the vectors in the left null space are scalars that record a sequence of row
operations that result in a row of zeros in the row-reduced version of the matrix. We will see this idea
more explicitly in the course of proving Theorem FS [299].
Example LNS
Left null space
We will nd the left null space of
A=2
6641 3 1
2 1 1
1 5 1
9 4 03
775
We transpose Aand row-reduce,
At=2
41 2 1 9
3 1 5 4
1 1 1 03
5RREF !2
410 0 2
010 3
0 0 1 13
5
Version 2.30
296 Section FS Four Subsets
Applying Denition LNS [293] and Theorem BNS [160] we have
L(A) =N
At
=*8
>><
>>:2
664 2
3
1
13
7759
>>=
>>;+
If you row-reduce Ayou will discover one zero row in the reduced row-echelon form. This zero row is
created by a sequence of row operations, which in total amounts to a linear combination, with scalars
a1= 2,a2= 3,a3= 1 anda4= 1, on the rows of Aand which results in the zero vector (check this!).
So the components of the vector describing the left null space of Aprovide a relation of linear dependence
on the rows of A.
Subsection CRS
Computing Column Spaces
We have three ways to build the column space of a matrix. First, we can use just the denition, Denition
CSM [271], and express the column space as a span of the columns of the matrix. A second approach gives
us the column space as the span of some of the columns of the matrix, but this set is linearly independent
(Theorem BCS [274]). Finally, we can transpose the matrix, row-reduce the transpose, kick out zero rows,
and transpose the remaining rows back into column vectors. Theorem CSRST [282] and Theorem BRS
[280] tell us that the resulting vectors are linearly independent and their span is the column space of the
original matrix.
We will now demonstrate a fourth method by way of a rather complicated example. Study this example
carefully, but realize that its main purpose is to motivate a theorem that simplies much of the apparent
complexity. So other than an instructive exercise or two, the procedure we are about to describe will not
be a usual approach to computing a column space.
Example CSANS
Column space as null space
Lets nd the column space of the matrix Abelow with a new approach.
A=2
666666410 0 3 8 7
16 1 4 10 13
6 1 3 6 6
0 2 2 3 2
3 0 1 2 3
1 1 1 1 03
7777775
By Theorem CSCS [272] we know that the column vector bis in the column space of Aif and only if the
linear systemLS(A;b) is consistent. So let's try to solve this system in full generality, using a vector of
variables for the vector of constants. In other words, which vectors blead to consistent systems? Begin
by forming the augmented matrix [ Ajb] with a general version of b,
[Ajb] =2
666666410 0 3 8 7 b1
16 1 4 10 13b2
6 1 3 6 6b3
0 2 2 3 2b4
3 0 1 2 3 b5
1 1 1 1 0 b63
7777775
Version 2.30
Subsection FS.CRS Computing Column Spaces 297
To identify solutions we will row-reduce this matrix and bring it to reduced row-echelon form. Despite the
presence of variables in the last column, there is nothing to stop us from doing this. Except our numerical
routines on calculators can't be used, and even some of the symbolic algebra routines do some unexpected
maneuvers with this computation. So do it by hand. Yes, it is a bit of work. But worth it. We'll still
be here when you get back. Notice along the way that the row operations are exactly the same ones you
would do if you were just row-reducing the coecient matrix alone, say in connection with a homogeneous
system of equations. The column with the biacts as a sort of bookkeeping device. There are many dierent
possibilities for the result, depending on what order you choose to perform the row operations, but shortly
we'll all be on the same page. Here's one possibility (you can nd this same result by doing additional row
operations with the fth and sixth rows to remove any occurrences of b1andb2from the rst four rows of
your result):
2
6666666410 0 0 2 b3 b4+ 2b5 b6
010 0 3 2b3+ 3b4 3b5+ 3b6
0 0 10 1 b3+b4+ 3b5+ 3b6
0 0 0 1 2 2b3+b4 4b5
0 0 0 0 0 b1+ 3b3 b4+ 3b5+b6
0 0 0 0 0 b2 2b3+b4+b5 b63
77777775
Our goal is to identify those vectors bwhich makeLS(A;b) consistent. By Theorem RCLS [58] we know
that the consistent systems are precisely those without a leading 1 in the last column. Are the expressions
in the last column of rows 5 and 6 equal to zero, or are they leading 1's? The answer is: maybe. It depends
onb. With a nonzero value for either of these expressions, we would scale the row and produce a leading
1. So we get a consistent system, and bis in the column space, if and only if these two expressions are
both simultaneously zero. In other words, members of the column space of Aare exactly those vectors b
that satisfy
b1+ 3b3 b4+ 3b5+b6= 0
b2 2b3+b4+b5 b6= 0
Hmmm. Looks suspiciously like a homogeneous system of two equations with six variables. If you've
been playing along (and we hope you have) then you may have a slightly dierent system, but you should
have just two equations. Form the coecient matrix and row-reduce (notice that the system above has a
coecient matrix that is already in reduced row-echelon form). We should all be together now with the
same matrix,
L=10 3 1 3 1
01 2 1 1 1
So,C(A) =N(L) and we can apply Theorem BNS [160] to obtain a linearly independent set to use in a
span construction,
C(A) =N(L) =*8
>>>>>><
>>>>>>:2
6666664 3
2
1
0
0
03
7777775;2
66666641
1
0
1
0
03
7777775;2
6666664 3
1
0
0
1
03
7777775;2
6666664 1
1
0
0
0
13
77777759
>>>>>>=
>>>>>>;+
Whew! As a postscript to this central example, you may wish to convince yourself that the four vectors
above really are elements of the column space? Do they create consistent systems with Aas coecient
matrix? Can you recognize the constant vector in your description of these solution sets?
OK, that was so much fun, let's do it again. But simpler this time. And we'll all get the same results
all the way through. Doing row operations by hand with variables can be a bit error prone, so let's see if
Version 2.30
298 Section FS Four Subsets
we can improve the process some. Rather than row-reduce a column vector bfull of variables, let's write
b=I6band we will row-reduce the matrix I6and when we nish row-reducing, then we will compute the
matrix-vector product. You should rst convince yourself that we can operate like this (this is the subject
of a future homework exercise). Rather than augmenting Awithb, we will instead augment it with I6
(does this feel familiar?),
M=2
666666410 0 3 8 7 1 0 0 0 0 0
16 1 4 10 13 0 1 0 0 0 0
6 1 3 6 6 0 0 1 0 0 0
0 2 2 3 2 0 0 0 1 0 0
3 0 1 2 3 0 0 0 0 1 0
1 1 1 1 0 0 0 0 0 0 13
7777775
We want to row-reduce the left-hand side of this matrix, but we will apply the same row operations to the
right-hand side as well. And once we get the left-hand side in reduced row-echelon form, we will continue on
to put leading 1's in the nal two rows, as well as clearing out the columns containing those two additional
leading 1's. It is these additional row operations that will ensure that we all get to the same place, since
the reduced row-echelon form is unique (Theorem RREFU [35]),
N=2
66666641 0 0 0 2 0 0 1 1 2 1
0 1 0 0 3 0 0 2 3 3 3
0 0 1 0 1 0 0 1 1 3 3
0 0 0 1 2 0 0 2 1 4 0
0 0 0 0 0 1 0 3 1 3 1
0 0 0 0 0 0 1 2 1 1 13
7777775
We are after the nal six columns of this matrix, which we will multiply by b
J=2
66666640 0 1 1 2 1
0 0 2 3 3 3
0 0 1 1 3 3
0 0 2 1 4 0
1 0 3 1 3 1
0 1 2 1 1 13
7777775
so
Jb=2
66666640 0 1 1 2 1
0 0 2 3 3 3
0 0 1 1 3 3
0 0 2 1 4 0
1 0 3 1 3 1
0 1 2 1 1 13
77777752
6666664b1
b2
b3
b4
b5
b63
7777775=2
6666664b3 b4+ 2b5 b6
2b3+ 3b4 3b5+ 3b6
b3+b4+ 3b5+ 3b6
2b3+b4 4b5
b1+ 3b3 b4+ 3b5+b6
b2 2b3+b4+b5 b63
7777775
So by applying the same row operations that row-reduce Ato the identity matrix (which we could do with
a calculator once I6is placed alongside of A), we can then arrive at the result of row-reducing a column
of symbols where the vector of constants usually resides. Since the row-reduced version of Ahas two zero
rows, for a consistent system we require that
b1+ 3b3 b4+ 3b5+b6= 0
b2 2b3+b4+b5 b6= 0
Now we are exactly back where we were on the rst go-round. Notice that we obtain the matrix Las
simply the last two rows and last six columns of N.
This example motivates the remainder of this section, so it is worth careful study. You might attempt
to mimic the second approach with the coecient matrices of Archetype I [816] and Archetype J [820]. We
will see shortly that the matrix Lcontains more information about Athan just the column space.
Version 2.30
Subsection FS.EEF Extended echelon form 299
Subsection EEF
Extended echelon form
The nal matrix that we row-reduced in Example CSANS [294] should look familiar in most respects to
the procedure we used to compute the inverse of a nonsingular matrix, Theorem CINM [248]. We will
now generalize that procedure to matrices that are not necessarily nonsingular, or even square. First a
denition.
Denition EEF
Extended Echelon Form
SupposeAis anmnmatrix. Extend Aon its right side with the addition of an mmidentity matrix
to form an m(n+m) matrix M. Use row operations to bring Mto reduced row-echelon form and call
the resultN.Nis the extended reduced row-echelon form ofA, and we will standardize on names
for ve submatrices ( B,C,J,K,L) ofN.
LetBdenote the mnmatrix formed from the rst ncolumns of Nand letJdenote the mm
matrix formed from the last mcolumns of N. Suppose that Bhasrnonzero rows. Further partition Nby
lettingCdenote the rnmatrix formed from all of the non-zero rows of B. LetKbe thermmatrix
formed from the rst rrows ofJ, whileLwill be the ( m r)mmatrix formed from the bottom m r
rows ofJ. Pictorially,
M= [AjIm]RREF !N= [BjJ] =CK
0L
4
Example SEEF
Submatrices of extended echelon form
We illustrate Denition EEF [297] with the matrix A,
A=2
6641 1 2 7 1 6
6 2 4 18 3 26
4 1 4 10 2 17
3 1 2 9 1 123
775
Augmenting with the 4 4 identity matrix, M=
2
6641 1 2 7 1 6 1 0 0 0
6 2 4 18 3 26 0 1 0 0
4 1 4 10 2 17 0 0 1 0
3 1 2 9 1 12 0 0 0 13
775
and row-reducing, we obtain
N=2
666410 2 1 0 3 0 1 1 1
014 6 0 1 0 2 3 0
0 0 0 0 1 2 0 1 0 2
0 0 0 0 0 0 1 2 2 13
7775
So we then obtain
B=2
66410 2 1 0 3
014 6 0 1
0 0 0 0 1 2
0 0 0 0 0 03
775
Version 2.30
300 Section FS Four Subsets
C=2
410 2 1 0 3
014 6 0 1
0 0 0 0 1 23
5
J=2
6640 1 1 1
0 2 3 0
0 1 0 2
1 2 2 13
775
K=2
40 1 1 1
0 2 3 0
0 1 0 23
5
L=
12 2 1
You can observe (or verify) the properties of the following theorem with this example.
Theorem PEEF
Properties of Extended Echelon Form
Suppose that Ais anmnmatrix and that Nis its extended echelon form. Then
1.Jis nonsingular.
2.B=JA.
3. Ifx2Cnandy2Cm, thenAx=yif and only if Bx=Jy.
4.Cis in reduced row-echelon form, has no zero rows and has rpivot columns.
5.Lis in reduced row-echelon form, has no zero rows and has m rpivot columns.
ProofJis the result of applying a sequence of row operations to Im, as suchJandImare row-equivalent.
LS(Im;0) has only the zero solution, since Imis nonsingular (Theorem NMRRI [84]). Thus, LS(J;0) also
has only the zero solution (Theorem REMES [31], Denition ESYS [14]) and Jis therefore nonsingular
(Denition NSM [73]).
To prove the second part of this conclusion, rst convince yourself that row operations and the matrix-
vector are commutative operations. By this we mean the following. Suppose that Fis anmnmatrix that
is row-equivalent to the matrix G. Apply to the column vector Fwthe same sequence of row operations
that converts FtoG. Then the result is Gw. So we can do row operations on the matrix, then do a
matrix-vector product, ordo a matrix-vector product and then do row operations on a column vector, and
the result will be the same either way. Since matrix multiplication is dened by a collection of matrix-
vector products (Denition MM [226]), if we apply to the matrix product FHthe same sequence of row
operations that converts FtoGthen the result will equal GH. Now apply these observations to A.
WriteAIn=ImAand apply the row operations that convert MtoN.Ais converted to B, whileIm
is converted to J, so we have BIn=JA. Simplifying the left side gives the desired conclusion.
For the third conclusion, we now establish the two equivalences
Ax=y() JAx=Jy() Bx=Jy
The forward direction of the rst equivalence is accomplished by multiplying both sides of the matrix
equality by J, while the backward direction is accomplished by multiplying by the inverse of J(which we
know exists by Theorem NI [261] since Jis nonsingular). The second equivalence is obtained simply by
the substitutions given by JA=B.
Version 2.30
Subsection FS.FS Four Subsets 301
The rstrrows ofNare in reduced row-echelon form, since any contiguous collection of rows taken
from a matrix in reduced row-echelon form will form a matrix that is again in reduced row-echelon form.
Since the matrix Cis formed by removing the last nentries of each these rows, the remainder is still in
reduced row-echelon form. By its construction, Chas no zero rows. Chasrrows and each contains a
leading 1, so there are rpivot columns in C.
The nalm rrows ofNare in reduced row-echelon form, since any contiguous collection of rows
taken from a matrix in reduced row-echelon form will form a matrix that is again in reduced row-echelon
form. Since the matrix Lis formed by removing the rst nentries of each these rows, and these entries
are all zero (they form the zero rows of B), the remainder is still in reduced row-echelon form. Lis the
nalm rrows of the nonsingular matrix J, so none of these rows can be totally zero, or Jwould not
row-reduce to the identity matrix. Lhasm rrows and each contains a leading 1, so there are m r
pivot columns in L.
Notice that in the case where Ais a nonsingular matrix we know that the reduced row-echelon form
ofAis the identity matrix (Theorem NMRRI [84]), so B=In. Then the second conclusion above says
JA=B=In, soJis the inverse of A. Thus this theorem generalizes Theorem CINM [248], though the
result is a \left-inverse" of Arather than a \right-inverse."
The third conclusion of Theorem PEEF [298] is the most telling. It says that xis a solution to the linear
systemLS(A;y) if and only if xis a solution to the linear system LS(B; Jy). Or said dierently, if we
row-reduce the augmented matrix [ Ajy] we will get the augmented matrix [ BjJy]. The matrix Jtracks
the cumulative eect of the row operations that converts Ato reduced row-echelon form, here eectively
applying them to the vector of constants in a system of equations having Aas a coecient matrix. When
Arow-reduces to a matrix with zero rows, then Jyshould also have zero entries in the same rows if the
system is to be consistent.
Subsection FS
Four Subsets
With all the preliminaries in place we can state our main result for this section. In essence this result will
allow us to say that we can nd linearly independent sets to use in span constructions for all four subsets
(null space, column space, row space, left null space) by analyzing only the extended echelon form of the
matrix, and specically, just the two submatrices CandL, which will be ripe for analysis since they are
already in reduced row-echelon form (Theorem PEEF [298]).
Theorem FS
Four Subsets
SupposeAis anmnmatrix with extended echelon form N. Suppose the reduced row-echelon form of
Ahasrnonzero rows. Then Cis the submatrix of Nformed from the rst rrows and the rst ncolumns
andLis the submatrix of Nformed from the last mcolumns and the last m rrows. Then
1. The null space of Ais the null space of C,N(A) =N(C).
2. The row space of Ais the row space of C,R(A) =R(C).
3. The column space of Ais the null space of L,C(A) =N(L).
4. The left null space of Ais the row space of L,L(A) =R(L).
Proof First,N(A) =N(B) sinceBis row-equivalent to A(Theorem REMES [31]). The zero rows of
Brepresent equations that are always true in the homogeneous system LS(B;0), so the removal of these
equations will not change the solution set. Thus, in turn, N(B) =N(C).
Version 2.30
302 Section FS Four Subsets
Second,R(A) =R(B) sinceBis row-equivalent to A(Theorem REMRS [279]). The zero rows of B
contribute nothing to the span that is the row space of B, so the removal of these rows will not change the
row space. Thus, in turn, R(B) =R(C).
Third, we prove the set equality C(A) =N(L) with Denition SE [762]. Begin by showing that
C(A)N(L). Choose y2C(A)Cm. Then there exists a vector x2Cnsuch thatAx=y(Theorem
CSCS [272]). Then for 1 km r,
[Ly]k= [Jy]r+k La submatrix of J
= [Bx]r+k Theorem PEEF [298]
= [Ox]k Zero matrix a submatrix of B
= [0]k Theorem MMZM [229]
So, for all 1km r, [Ly]k= [0]k. So by Denition CVE [98] we have Ly=0and thus y2N(L).
Now, show thatN(L)C(A). Choose y2N(L)Cm. Form the vector Ky2Cr. The linear system
LS(C; Ky) is consistent since Cis in reduced row-echelon form and has no zero rows (Theorem PEEF
[298]). Let x2Cndenote a solution to LS(C; Ky).
Then for 1jr,
[Bx]j= [Cx]j Ca submatrix of B
= [Ky]j xa solution toLS(C; Ky)
= [Jy]j Ka submatrix of J
And forr+ 1km,
[Bx]k= [Ox]k r Zero matrix a submatrix of B
= [0]k r Theorem MMZM [229]
= [Ly]k r yinN(L)
= [Jy]k La submatrix of J
So for all 1im, [Bx]i= [Jy]iand by Denition CVE [98] we have Bx=Jy. From Theorem PEEF
[298] we know then that Ax=y, and therefore y2C(A) (Theorem CSCS [272]). By Denition SE [762]
we now haveC(A) =N(L).
Fourth, we prove the set equality L(A) =R(L) with Denition SE [762]. Begin by showing that
R(L)L(A). Choose y2R(L)Cm. Then there exists a vector w2Cm rsuch that y=Ltw
(Denition RSM [278], Theorem CSCS [272]). Then for 1 in,
Aty
i=mX
k=1
At
ik[y]k Theorem EMP [227]
=mX
k=1
At
ik
Ltw
kDenition of w
=mX
k=1
At
ikm rX
`=1
Lt
k`[w]` Theorem EMP [227]
=mX
k=1m rX
`=1
At
ik
Lt
k`[w]` Property DCN [759]
Version 2.30
Subsection FS.FS Four Subsets 303
=m rX
`=1mX
k=1
At
ik
Lt
k`[w]` Property CACN [758]
=m rX
`=1 mX
k=1
At
ik
Lt
k`!
[w]` Property DCN [759]
=m rX
`=1 mX
k=1
At
ik
Jt
k;r+`!
[w]` La submatrix of J
=m rX
`=1
AtJt
i;r+`[w]` Theorem EMP [227]
=m rX
`=1
(JA)t
i;r+`[w]` Theorem MMT [232]
=m rX
`=1
Bt
i;r+`[w]` Theorem PEEF [298]
=m rX
`=10 [w]` Zero rows in B
= 0 Property ZCN [759]
= [0]i Denition ZCV [28]
Since
Aty
i= [0]ifor 1in, Denition CVE [98] implies that Aty=0. This means that y2N
At
.
Now, show thatL(A)R(L). Choose y2L(A)Cm. The matrix Jis nonsingular (Theorem PEEF
[298]), soJtis also nonsingular (Theorem MIT [251]) and therefore the linear system LS
Jt;y
has a
unique solution. Denote this solution as x2Cm. We will need to work with two \halves" of x, which we
will denote as zandwwith formal denitions given by
[z]j= [x]i 1jr; [w]k= [x]r+k 1km r
Now, for 1jr,
Ctz
j=rX
k=1
Ct
jk[z]k Theorem EMP [227]
=rX
k=1
Ct
jk[z]k+m rX
`=1[O]j`[w]` Denition ZM [210]
=rX
k=1
Bt
jk[z]k+m rX
`=1
Bt
j;r+`[w]` C,Osubmatrices of B
=rX
k=1
Bt
jk[x]k+m rX
`=1
Bt
j;r+`[x]r+` Denitions of zandw
=rX
k=1
Bt
jk[x]k+mX
k=r+1
Bt
jk[x]k Re-index second sum
=mX
k=1
Bt
jk[x]k Combine sums
=mX
k=1
(JA)t
jk[x]k Theorem PEEF [298]
Version 2.30
304 Section FS Four Subsets
=mX
k=1
AtJt
jk[x]k Theorem MMT [232]
=mX
k=1mX
`=1
At
j`
Jt
`k[x]k Theorem EMP [227]
=mX
`=1mX
k=1
At
j`
Jt
`k[x]k Property CACN [758]
=mX
`=1
At
j` mX
k=1
Jt
`k[x]k!
Property DCN [759]
=mX
`=1
At
j`
Jtx
`Theorem EMP [227]
=mX
`=1
At
j`[y]` Denition of x
=
Aty
jTheorem EMP [227]
= [0]j y2L(A)
So, by Denition CVE [98], Ctz=0and the vector zgives us a linear combination of the columns of Ct
that equals the zero vector. In other words, zgives a relation of linear dependence on the the rows of C.
However, the rows of Care a linearly independent set by Theorem BRS [280]. According to Denition
LICV [153] we must conclude that the entries of zare all zero, i.e. z=0.
Now, for 1im, we have
[y]i=
Jtx
iDenition of x
=mX
k=1
Jt
ik[x]k Theorem EMP [227]
=rX
k=1
Jt
ik[x]k+mX
k=r+1
Jt
ik[x]k Break apart sum
=rX
k=1
Jt
ik[z]k+mX
k=r+1
Jt
ik[w]k r Denition of zandw
=rX
k=1
Jt
ik0 +m rX
`=1
Jt
i;r+`[w]` z=0, re-index
= 0 +m rX
`=1
Lt
i;`[w]` La submatrix of J
=
Ltw
iTheorem EMP [227]
So by Denition CVE [98], y=Ltw. The existence of wimplies that y2R(L), and thereforeL(A)
R(L). So by Denition SE [762] we have L(A) =R(L).
The rst two conclusions of this theorem are nearly trivial. But they set up a pattern of results for C
that is re
ected in the latter two conclusions about L. In total, they tell us that we can compute all four
subsets just by nding null spaces and row spaces. This theorem does not tell us exactly how to compute
these subsets, but instead simply expresses them as null spaces and row spaces of matrices in reduced
row-echelon form without any zero rows ( CandL). A linearly independent set that spans the null space
Version 2.30
Subsection FS.FS Four Subsets 305
of a matrix in reduced row-echelon form can be found easily with Theorem BNS [160]. It is an even easier
matter to nd a linearly independent set that spans the row space of a matrix in reduced row-echelon form
with Theorem BRS [280], especially when there are no zero rows present. So an application of Theorem
FS [299] is typically followed by two applications each of Theorem BNS [160] and Theorem BRS [280].
The situation when r=mdeserves comment, since now the matrix Lhas no rows. What is C(A) when
we try to apply Theorem FS [299] and encounter N(L)? One interpretation of this situation is that Lis
the coecient matrix of a homogeneous system that has no equations. How hard is it to nd a solution
vector to this system? Some thought will convince you that anyproposed vector will qualify as a solution,
since it makes allof the equations true. So every possible vector is in the null space of Land therefore
C(A) =N(L) =Cm. OK, perhaps this sounds like some twisted argument from Alice in Wonderland . Let
us try another argument that might solidly convince you of this logic.
Ifr=m, when we row-reduce the augmented matrix of LS(A;b) the result will have no zero rows, and
all the leading 1's will occur in rst ncolumns, so by Theorem RCLS [58] the system will be consistent.
By Theorem CSCS [272], b2C(A). Since bwas arbitrary, every possible vector is in the column space of
A, so we again have C(A) =Cm. The situation when a matrix has r=mis known by the term full rank ,
and in the case of a square matrix coincides with nonsingularity (see Exercise FS.M50 [309]).
The properties of the matrix Ldescribed by this theorem can be explained informally as follows. A
column vector y2Cmis in the column space of Aif the linear system LS(A;y) is consistent (Theorem
CSCS [272]). By Theorem RCLS [58], the reduced row-echelon form of the augmented matrix [ Ajy] of a
consistent system will have zeros in the bottom m rlocations of the last column. By Theorem PEEF
[298] this nal column is the vector Jyand so should then have zeros in the nal m rlocations. But
sinceLcomprises the nal m rrows ofJ, this condition is expressed by saying y2N(L).
Additionally, the rows of Jare the scalars in linear combinations of the rows of Athat create the
rows ofB. That is, the rows of Jrecord the net eect of the sequence of row operations that takes Ato
its reduced row-echelon form, B. This can be seen in the equation JA=B(Theorem PEEF [298]). As
such, the rows of Lare scalars for linear combinations of the rows of Athat yield zero rows. But such
linear combinations are precisely the elements of the left null space. So any element of the row space of
Lis also an element of the left null space of A. We will now illustrate Theorem FS [299] with a few examples.
Example FS1
Four subsets, #1
In Example SEEF [297] we found the ve relevant submatrices of the matrix
A=2
6641 1 2 7 1 6
6 2 4 18 3 26
4 1 4 10 2 17
3 1 2 9 1 123
775
To apply Theorem FS [299] we only need CandL,
C=2
410 2 1 0 3
014 6 0 1
0 0 0 0 1 23
5 L=
12 2 1
Then we use Theorem FS [299] to obtain
N(A) =N(C) =*8
>>>>>><
>>>>>>:2
6666664 2
4
1
0
0
03
7777775;2
6666664 1
6
0
1
0
03
7777775;2
6666664 3
1
0
0
2
13
77777759
>>>>>>=
>>>>>>;+
Theorem BNS [160]
Version 2.30
306 Section FS Four Subsets
R(A) =R(C) =*8
>>>>>><
>>>>>>:2
66666641
0
2
1
0
33
7777775;2
66666640
1
4
6
0
13
7777775;2
66666640
0
0
0
1
23
77777759
>>>>>>=
>>>>>>;+
Theorem BRS [280]
C(A) =N(L) =*8
>><
>>:2
664 2
1
0
03
775;2
664 2
0
1
03
775;2
664 1
0
0
13
7759
>>=
>>;+
Theorem BNS [160]
L(A) =R(L) =*8
>><
>>:2
6641
2
2
13
7759
>>=
>>;+
Theorem BRS [280]
Boom!
Example FS2
Four subsets, #2
Now lets return to the matrix Athat we used to motivate this section in Example CSANS [294],
A=2
666666410 0 3 8 7
16 1 4 10 13
6 1 3 6 6
0 2 2 3 2
3 0 1 2 3
1 1 1 1 03
7777775
We form the matrix Mby adjoining the 6 6 identity matrix I6,
M=2
666666410 0 3 8 7 1 0 0 0 0 0
16 1 4 10 13 0 1 0 0 0 0
6 1 3 6 6 0 0 1 0 0 0
0 2 2 3 2 0 0 0 1 0 0
3 0 1 2 3 0 0 0 0 1 0
1 1 1 1 0 0 0 0 0 0 13
7777775
and row-reduce to obtain N
N=2
6666666410 0 0 2 0 0 1 1 2 1
010 0 3 0 0 2 3 3 3
0 0 10 1 0 0 1 1 3 3
0 0 0 1 2 0 0 2 1 4 0
0 0 0 0 0 10 3 1 3 1
0 0 0 0 0 0 1 2 1 1 13
77777775
To nd the four subsets for A, we only need identify the 4 5 matrixCand the 26 matrixL,
C=2
666410 0 0 2
010 0 3
0 0 10 1
0 0 0 1 23
7775L=10 3 1 3 1
01 2 1 1 1
Version 2.30
Subsection FS.FS Four Subsets 307
Then we apply Theorem FS [299],
N(A) =N(C) =*8
>>>><
>>>>:2
66664 2
3
1
2
13
777759
>>>>=
>>>>;+
Theorem BNS [160]
R(A) =R(C) =*8
>>>><
>>>>:2
666641
0
0
0
23
77775;2
666640
1
0
0
33
77775;2
666640
0
1
0
13
77775;2
666640
0
0
1
23
777759
>>>>=
>>>>;+
Theorem BRS [280]
C(A) =N(L) =*8
>>>>>><
>>>>>>:2
6666664 3
2
1
0
0
03
7777775;2
66666641
1
0
1
0
03
7777775;2
6666664 3
1
0
0
1
03
7777775;2
6666664 1
1
0
0
0
13
77777759
>>>>>>=
>>>>>>;+
Theorem BNS [160]
L(A) =R(L) =*8
>>>>>><
>>>>>>:2
66666641
0
3
1
3
13
7777775;2
66666640
1
2
1
1
13
77777759
>>>>>>=
>>>>>>;+
Theorem BRS [280]
The next example is just a bit dierent since the matrix has more rows than columns, and a trivial
null space.
Example FSAG
Four subsets, Archetype G
Archetype G [808] and Archetype H [812] are both systems of m= 5 equations in n= 2 variables. They
have identical coecient matrices, which we will denote here as the matrix G,
G=2
666642 3
1 4
3 10
3 1
6 93
77775
Adjoin the 55 identity matrix, I5, to form
M=2
666642 3 1 0 0 0 0
1 4 0 1 0 0 0
3 10 0 0 1 0 0
3 1 0 0 0 1 0
6 9 0 0 0 0 13
77775
This row-reduces to
N=2
66666410 0 0 03
111
33
010 0 0 2
111
11
0 0 10 0 0 1
3
0 0 0 10 1 1
3
0 0 0 0 1 1 13
777775
Version 2.30
308 Section FS Four Subsets
The rstn= 2 columns contain r= 2 leading 1's, so we obtain Cas the 22 identity matrix and extract
Lfrom the nal m r= 3 rows in the nal m= 5 columns.
C=10
01
L=2
410 0 0 1
3
010 1 1
3
0 0 11 13
5
Then we apply Theorem FS [299],
N(G) =N(C) =h;i=f0g Theorem BNS [160]
R(G) =R(C) =1
0
;0
1
=C2Theorem BRS [280]
C(G) =N(L) =*8
>>>><
>>>>:2
666640
1
1
1
03
77775;2
666641
31
3
1
0
13
777759
>>>>=
>>>>;+
Theorem BNS [160]
=*8
>>>><
>>>>:2
666640
1
1
1
03
77775;2
666641
1
3
0
33
777759
>>>>=
>>>>;+
L(G) =R(L) =*8
>>>><
>>>>:2
666641
0
0
0
1
33
77775;2
666640
1
0
1
1
33
77775;2
666640
0
1
1
13
777759
>>>>=
>>>>;+
Theorem BRS [280]
=*8
>>>><
>>>>:2
666643
0
0
0
13
77775;2
666640
3
0
3
13
77775;2
666640
0
1
1
13
777759
>>>>=
>>>>;+
As mentioned earlier, Archetype G [808] is consistent, while Archetype H [812] is inconsistent. See if you
can write the two dierent vectors of constants from these two archetypes as linear combinations of the two
vectors inC(G). How about the two columns of G, can you write each individually as a linear combination
of the two vectors in C(G)? They must be in the column space of Galso. Are your answers unique? Do
you notice anything about the scalars that appear in the linear combinations you are forming?
Example COV [177] and Example CSROI [282] each describes the column space of the coecient matrix
from Archetype I [816] as the span of a set of r= 3 linearly independent vectors. It is no accident that
these two dierent sets both have the same size. If we (you?) were to calculate the column space of this
matrix using the null space of the matrix Lfrom Theorem FS [299] then we would again nd a set of 3
linearly independent vectors that span the range. More on this later.
So we have three dierent methods to obtain a description of the column space of a matrix as the
span of a linearly independent set. Theorem BCS [274] is sometimes useful since the vectors it species
are equal to actual columns of the matrix. Theorem BRS [280] and Theorem CSRST [282] combine to
create vectors with lots of zeros, and strategically placed 1's near the top of the vector. Theorem FS [299]
and the matrix Lfrom the extended echelon form gives us a third method, which tends to create vectors
with lots of zeros, and strategically placed 1's near the bottom of the vector. If we don't care about linear
Version 2.30
Subsection FS.READ Reading Questions 309
independence we can also appeal to Denition CSM [271] and simply express the column space as the span
of all the columns of the matrix, giving us a fourth description.
With Theorem CSRST [282] and Denition RSM [278], we can compute column spaces with theorems
about row spaces, and we can compute row spaces with theorems about row spaces, but in each case
we must transpose the matrix rst. At this point you may be overwhelmed by all the possibilities for
computing column and row spaces. Diagram CSRST [307] is meant to help. For both the column space
and row space, it suggests four techniques. One is to appeal to the denition, another yields a span of a
linearly independent set, and a third uses Theorem FS [299]. A fourth suggests transposing the matrix
and the dashed line implies that then the companion set of techniques can be applied. This can lead to
a bit of silliness, since if you were to follow the dashed lines twice you would transpose the matrix twice,
and by Theorem TT [212] would accomplish nothing productive.
R(A)C(A)Definition CSM
Theorem BCS
Theorem FS,N(L)
Theorem CSRST, R(At)
Definition RSM, C(At)
Theorem FS,R(C)
Theorem BRS
Definition RSM
Diagram CSRST. Column Space and Row Space Techniques
Although we have many ways to describe a column space, notice that one tempting strategy will usually
fail. It is not possible to simply row-reduce a matrix directly and then use the columns of the row-reduced
matrix as a set whose span equals the column space. In other words, row operations do not preserve column
spaces (however row operations do preserve row spaces, Theorem REMRS [279]). See Exercise CRS.M21
[287].
Subsection READ
Reading Questions
1. Find a nontrivial element of the left null space of A.
A=2
42 1 3 4
1 1 2 1
0 1 1 23
5
2. Find the matrices CandLin the extended echelon form of A.
A=2
4 9 5 3
2 1 1
5 3 13
5
3. Why is Theorem FS [299] a great conclusion to Chapter M [207]?
Version 2.30
310 Section FS Four Subsets
Subsection EXC
Exercises
C20 Example FSAG [305] concludes with several questions. Perform the analysis suggested by these
questions.
Contributed by Robert Beezer
C25 Given the matrix Abelow, use the extended echelon form of Ato answer each part of this problem.
In each part, nd a linearly independent set of vectors, S, so that the span of S,hSi, equals the specied
set of vectors.
A=2
664 5 3 1
1 1 1
8 5 1
3 2 03
775
(a) The row space of A,R(A).
(b) The column space of A,C(A).
(c) The null space of A,N(A).
(d) The left null space of A,L(A).
Contributed by Robert Beezer Solution [310]
C26 For the matrix Dbelow use the extended echelon form to nd
(a) a linearly independent set whose span is the column space of D.
(b) a linearly independent set whose span is the left null space of D.
D=2
664 7 11 19 15
6 10 18 14
3 5 9 7
1 2 4 33
775
Contributed by Robert Beezer Solution [310]
C41 The following archetypes are systems of equations. For each system, write the vector of constants
as a linear combination of the vectors in the span construction for the column space provided by Theorem
FS [299] and Theorem BNS [160] (these vectors are listed for each of these archetypes).
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]
Archetype E [799]
Archetype F [803]
Archetype G [808]
Archetype H [812]
Archetype I [816]
Archetype J [820]
Contributed by Robert Beezer
C43 The following archetypes are either matrices or systems of equations with coecient matrices. For
each matrix, compute the extended echelon form Nand identify the matrices CandL. Using Theorem
Version 2.30
Subsection FS.EXC Exercises 311
FS [299], Theorem BNS [160] and Theorem BRS [280] express the null space, the row space, the column
space and left null space of each coecient matrix as a span of a linearly independent set.
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]/Archetype E [799]
Archetype F [803]
Archetype G [808]/Archetype H [812]
Archetype I [816]
Archetype J [820]
Archetype K [825]
Archetype L [829]
Contributed by Robert Beezer
C60 For the matrix Bbelow, nd sets of vectors whose span equals the column space of B(C(B)) and
which individually meet the following extra requirements.
(a) The set illustrates the denition of the column space.
(b) The set is linearly independent and the members of the set are columns of B.
(c) The set is linearly independent with a \nice pattern of zeros and ones" at the topof each vector.
(d) The set is linearly independent with a \nice pattern of zeros and ones" at the bottom of each vector.
B=2
42 3 1 1
1 1 0 1
1 2 3 43
5
Contributed by Robert Beezer Solution [311]
C61 LetAbe the matrix below, and nd the indicated sets with the requested properties.
A=2
42 1 5 3
5 3 12 7
1 1 4 33
5
(a) A linearly independent set Sso thatC(A) =hSiandSis composed of columns of A.
(b) A linearly independent set Sso thatC(A) =hSiand the vectors in Shave a nice pattern of zeros
and ones at the top of the vectors.
(c) A linearly independent set Sso thatC(A) =hSiand the vectors in Shave a nice pattern of zeros
and ones at the bottom of the vectors.
(d) A linearly independent set Sso thatR(A) =hSi.
Contributed by Robert Beezer Solution [312]
M50 Suppose that Ais a nonsingular matrix. Extend the four conclusions of Theorem FS [299] in this
special case and discuss connections with previous results (such as Theorem NME4 [277]).
Contributed by Robert Beezer
M51 Suppose that Ais a singular matrix. Extend the four conclusions of Theorem FS [299] in this
special case and discuss connections with previous results (such as Theorem NME4 [277]).
Contributed by Robert Beezer
Version 2.30
312 Section FS Four Subsets
Subsection SOL
Solutions
C25 Contributed by Robert Beezer Statement [308]
Add a 44 identity matrix to the right of Ato form the matrix Mand then row-reduce to the matrix N,
M=2
664 5 3 1 1 0 0 0
1 1 1 0 1 0 0
8 5 1 0 0 1 0
3 2 0 0 0 0 13
775RREF !2
666410 2 0 0 2 5
013 0 0 3 8
0 0 0 10 1 1
0 0 0 0 1 1 33
7775=N
To apply Theorem FS [299] in each of these four parts, we need the two matrices,
C=10 2
013
L=10 1 1
01 1 3
(a)
R(A) =R(C) Theorem FS [299]
=*2
41
0
23
5;2
40
1
33
5+
Theorem BRS [280]
(b)
C(A) =N(L) Theorem FS [299]
=*2
6641
1
1
03
775;2
6641
3
0
13
775+
Theorem BNS [160]
(c)
N(A) =N(C) Theorem FS [299]
=*2
4 2
3
13
5+
Theorem BNS [160]
(d)
L(A) =R(L) Theorem FS [299]
=*2
6641
0
1
13
775;2
6640
1
1
33
775+
Theorem BRS [280]
C26 Contributed by Robert Beezer Statement [308]
For both parts, we need the extended echelon form of the matrix.
2
664 7 11 19 15 1 0 0 0
6 10 18 14 0 1 0 0
3 5 9 7 0 0 1 0
1 2 4 3 0 0 0 13
775RREF !2
666410 2 1 0 0 2 5
01 3 2 0 0 1 3
0 0 0 0 10 3 2
0 0 0 0 0 1 2 03
7775
Version 2.30
Subsection FS.SOL Solutions 313
From this matrix we extract the last two rows, in the last four columns to form the matrix L,
L=10 3 2
01 2 0
(a) By Theorem FS [299] and Theorem BNS [160] we have
C(D) =N(L) =*8
>><
>>:2
664 3
2
1
03
775;2
664 2
0
0
13
7759
>>=
>>;+
(b) By Theorem FS [299] and Theorem BRS [280] we have
L(D) =R(L) =*8
>><
>>:2
6641
0
3
23
775;2
6640
1
2
03
7759
>>=
>>;+
C60 Contributed by Robert Beezer Statement [309]
(a) The denition of the column space is the span of the set of columns (Denition CSM [271]). So the
desired set is just the four columns of B,
S=8
<
:2
42
1
13
5;2
43
1
23
5;2
41
0
33
5;2
41
1
43
59
=
;
(b) Theorem BCS [274] suggests row-reducing the matrix and using the columns of Bthat correspond to
the pivot columns.
BRREF !2
410 1 2
01 1 1
0 0 0 03
5
So the pivot columns are numbered by elements of D=f1;2g, so the requested set is
S=8
<
:2
42
1
13
5;2
43
1
23
59
=
;
(c) We can nd this set by row-reducing the transpose of B, deleting the zero rows, and using the nonzero
rows as column vectors in the set. This is an application of Theorem CSRST [282] followed by Theorem
BRS [280].
BtRREF !2
66410 3
01 7
0 0 0
0 0 03
775
So the requested set is
S=8
<
:2
41
0
33
5;2
40
1
73
59
=
;
Version 2.30
314 Section FS Four Subsets
(d) With the column space expressed as a null space, the vectors obtained via Theorem BNS [160] will
be of the desired shape. So we rst proceed with Theorem FS [299] and create the extended echelon form,
[BjI3] =2
42 3 1 1 1 0 0
1 1 0 1 0 1 0
1 2 3 4 0 0 13
5RREF !2
410 1 2 02
3 1
3
01 1 1 01
31
3
0 0 0 0 1 7
3 1
33
5
So, employing Theorem FS [299], we have C(B) =N(L), where
L=
1 7
3 1
3
We can nd the desired set of vectors from Theorem BNS [160] as
S=8
<
:2
47
3
1
03
5;2
41
3
0
13
59
=
;
C61 Contributed by Robert Beezer Statement [309]
(a) First nd a matrix Bthat is row-equivalent to Aand in reduced row-echelon form
B=2
410 3 2
011 1
0 0 0 03
5
By Theorem BCS [274] we can choose the columns of Athat correspond to dependent variables ( D=f1;2g)
as the elements of Sand obtain the desired properties. So
S=8
<
:2
42
5
13
5;2
4 1
3
13
59
=
;
(b) We can write the column space of Aas the row space of the transpose (Theorem CSRST [282]). So
we row-reduce the transpose of Ato obtain the row-equivalent matrix Cin reduced row-echelon form
C=2
6641 0 8
0 1 3
0 0 0
0 0 03
775
The nonzero rows (written as columns) will be a linearly independent set that spans the row space of At,
by Theorem BRS [280], and the zeros and ones will be at the top of the vectors,
S=8
<
:2
41
0
83
5;2
40
1
33
59
=
;
(c) In preparation for Theorem FS [299], augment Awith the 33 identity matrix I3and row-reduce to
obtain the extended echelon form,
2
41 0 3 2 0 1
83
8
0 1 1 1 01
85
8
0 0 0 0 13
8 1
83
5
Then since the rst four columns of row 3 are all zeros, we extract
L=
13
8 1
8
Version 2.30
Subsection FS.SOL Solutions 315
Theorem FS [299] says that C(A) =N(L). We can then use Theorem BNS [160] to construct the desired
setS, based on the free variables with indices in F=f2;3gfor the homogeneous system LS(L;0), so
S=8
<
:2
4 3
8
1
03
5;2
41
8
0
13
59
=
;
Notice that the zeros and ones are at the bottom of the vectors.
(d) This is a straightforward application of Theorem BRS [280]. Use the row-reduced matrix Bfrom
part (a), grab the nonzero rows, and write them as column vectors,
S=8
>><
>>:2
6641
0
3
23
775;2
6640
1
1
13
7759
>>=
>>;
Version 2.30
316 Section FS Four Subsets
Version 2.30
Annotated Acronyms FS.M Matrices 317
Annotated Acronyms M
Matrices
Theorem VSPM [209]
These are the fundamental rules for working with the addition, and scalar multiplication, of matrices.
We saw something very similar in the previous chapter (Theorem VSPCV [100]). Together, these two
denitions will provide our denition for the key denition, Denition VS [317].
Theorem SLEMM [224]
Theorem SLSLC [112] connected linear combinations with systems of equations. Theorem SLEMM [224]
connects the matrix-vector product (Denition MVP [223]) and column vector equality (Denition CVE
[98]) with systems of equations. We'll see this one regularly.
Theorem EMP [227]
This theorem is a workhorse in Section MM [223] and will continue to make regular appearances. If you
want to get better at formulating proofs, the application of this theorem can be a key step in gaining that
broader understanding. While it might be hard to imagine Theorem EMP [227] as a denition of matrix
multiplication, we'll see in Exercise MR.T80 [637] that in theory it is actually a better denition of matrix
multiplication long-term.
Theorem CINM [248]
The inverse of a matrix is key. Here's how you can get one if you know how to row-reduce.
Theorem NPNT [259]
This theorem is a fundamental tool for proving subsequent important theorems, such as Theorem NI [261].
It may also be the best explantion for the term \nonsingular."
Theorem NI [261]
\Nonsingularity" or \invertibility"? Pick your favorite, or show your versatility by using one or the other
in the right context. They mean the same thing.
Theorem CSCS [272]
Given a coecient matrix, which vectors of constants create consistent systems. This theorem tells us that
the answer is exactly those column vectors in the column space. Conversely, we also use this teorem to test
for membership in the column space by checking the consistency of the appropriate system of equations.
Theorem BCS [274]
Another theorem that provides a linearly independent set of vectors whose span equals some set of interest
(a column space this time).
Theorem BRS [280]
Yet another theorem that provides a linearly independent set of vectors whose span equals some set of
interest (a row space).
Theorem CSRST [282]
Column spaces, row spaces, transposes, rows, columns. Many of the connections between these objects
are based on the simple observation captured in this theorem. This is not a deep result. We state it as a
theorem for convenience, so we can refer to it as needed.
Version 2.30
318 Section FS Four Subsets
Theorem FS [299]
This theorem is inherently interesting, if not computationally satisfying. Null space, row space, column
space, left null space | here they all are, simply by row reducing the extended matrix and applying
Theorem BNS [160] and Theorem BCS [274] twice (each). Nice.
Version 2.30
Chapter VS
Vector Spaces
We now have a computational toolkit in place and so we can begin our study of linear algebra in a more
theoretical style.
Linear algebra is the study of two fundamental objects, vector spaces and linear transformations (see
Chapter LT [515]). This chapter will focus on the former. The power of mathematics is often derived from
generalizing many dierent situations into one abstract formulation, and that is exactly what we will be
doing throughout this chapter.
Section VS
Vector Spaces
In this section we present a formal denition of a vector space, which will lead to an extra increment of
abstraction. Once dened, we study its most basic properties.
Subsection VS
Vector Spaces
Here is one of the two most important denitions in the entire course.
Denition VS
Vector Space
Suppose that Vis a set upon which we have dened two operations: (1) vector addition , which combines
two elements of Vand is denoted by \+", and (2) scalar multiplication , which combines a complex
number with an element of Vand is denoted by juxtaposition. Then V, along with the two operations, is
avector space overCif the following ten properties hold.
AC Additive Closure
Ifu;v2V, then u+v2V.
SC Scalar Closure
If2Candu2V, thenu2V.
C Commutativity
Ifu;v2V, then u+v=v+u.
AA Additive Associativity
Ifu;v;w2V, then u+ (v+w) = (u+v) +w.
319
320 Section VS Vector Spaces
Z Zero Vector
There is a vector, 0, called the zero vector , such that u+0=ufor all u2V.
AI Additive Inverses
Ifu2V, then there exists a vector u2Vso that u+ ( u) =0.
SMA Scalar Multiplication Associativity
If; 2Candu2V, then(u) = ()u.
DVA Distributivity across Vector Addition
If2Candu;v2V, then(u+v) =u+v.
DSA Distributivity across Scalar Addition
If; 2Candu2V, then (+)u=u+u.
O One
Ifu2V, then 1 u=u.
The objects in Vare called vectors , no matter what else they might really be, simply by virtue of being
elements of a vector space. 4
Now, there are several important observations to make. Many of these will be easier to understand
on a second or third reading, and especially after carefully studying the examples in Subsection VS.EVS
[319].
Anaxiom is often a \self-evident" truth. Something so fundamental that we all agree it is true and
accept it without proof. Typically, it would be the logical underpinning that we would begin to build
theorems upon. Some might refer to the ten properties of Denition VS [317] as axioms, implying that
a vector space is a very natural object and the ten properties are the essence of a vector space. We will
instead emphasize that we will begin with a denition of a vector space. After studying the remainder of
this chapter, you might return here and remind yourself how all our forthcoming theorems and denitions
rest on this foundation.
As we will see shortly, the objects in Vcan be anything , even though we will call them vectors. We
have been working with vectors frequently, but we should stress here that these have so far just been
column vectors | scalars arranged in a columnar list of xed length. In a similar vein, you have used the
symbol \+" for many years to represent the addition of numbers (scalars). We have extended its use to the
addition of column vectors and to the addition of matrices, and now we are going to recycle it even further
and let it denote vector addition in anypossible vector space. So when describing a new vector space, we
will have to dene exactly what \+" is. Similar comments apply to scalar multiplication. Conversely, we
candene our operations any way we like, so long as the ten properties are fullled (see Example CVS
[322]).
In Denition VS [317], the scalars do not have to be complex numbers. They can come from what are
called in more advanced mathematics, \elds" (see Section F [873] for more on these objects). Examples
of elds are the set of complex numbers, the set of real numbers, the set of rational numbers, and even
the nite set of \binary numbers", f0;1g. There are many, many others. In this case we would call Va
vector space over (the eld) F.
A vector space is composed of three objects, a set and two operations. Some would explicitly state
in the denition that Vmust be a non-empty set, be we can infer this from Property Z [318], since the
set cannot be empty and contain a vector that behaves as the zero vector. Also, we usually use the same
symbol for both the set and the vector space itself. Do not let this convenience fool you into thinking the
operations are secondary!
This discussion has either convinced you that we are really embarking on a new level of abstraction,
or they have seemed cryptic, mysterious or nonsensical. You might want to return to this section in a few
days and give it another read then. In any case, let's look at some concrete examples now.
Version 2.30
Subsection VS.EVS Examples of Vector Spaces 321
Subsection EVS
Examples of Vector Spaces
Our aim in this subsection is to give you a storehouse of examples to work with, to become comfortable
with the ten vector space properties and to convince you that the multitude of examples justies (at least
initially) making such a broad denition as Denition VS [317]. Some of our claims will be justied by
reference to previous theorems, we will prove some facts from scratch, and we will do one non-trivial
example completely. In other places, our usual thoroughness will be neglected, so grab paper and pencil
and play along.
Example VSCV
The vector space Cm
Set:Cm, all column vectors of size m, Denition VSCV [97].
Equality: Entry-wise, Denition CVE [98].
Vector Addition: The \usual" addition, given in Denition CVA [98].
Scalar Multiplication: The \usual" scalar multiplication, given in Denition CVSM [99].
Does this set with these operations fulll the ten properties? Yes. And by design all we need to do is
quote Theorem VSPCV [100]. That was easy.
Example VSM
The vector space of matrices, Mmn
Set:Mmn, the set of all matrices of size mnand entries from C, Example VSM [319].
Equality: Entry-wise, Denition ME [207].
Vector Addition: The \usual" addition, given in Denition MA [207].
Scalar Multiplication: The \usual" scalar multiplication, given in Denition MSM [208].
Does this set with these operations fulll the ten properties? Yes. And all we need to do is quote
Theorem VSPM [209]. Another easy one (by design).
So, the set of all matrices of a xed size forms a vector space. That entitles us to call a matrix a vector,
since a matrix is an element of a vector space. For example, if A; B2M3;4then we call AandB\vectors,"
and we even use our previous notation for column vectors to refer to AandB. So we could legitimately
write expressions like
u+v=A+B=B+A=v+u
This could lead to some confusion, but it is not too great a danger. But it is worth comment.
The previous two examples may be less than satisfying. We made all the relevant denitions long
ago. And the required verications were all handled by quoting old theorems. However, it is important to
consider these two examples rst. We have been studying vectors and matrices carefully (Chapter V [97],
Chapter M [207]), and both objects, along with their operations, have certain properties in common, as
you may have noticed in comparing Theorem VSPCV [100] with Theorem VSPM [209]. Indeed, it is these
two theorems that motivate us to formulate the abstract denition of a vector space, Denition VS [317].
Now, should we prove some general theorems about vector spaces (as we will shortly in Subsection VS.VSP
[323]), we can instantly apply the conclusions to bothCmandMmn. Notice too how we have taken six
denitions and two theorems and reduced them down to two examples . With greater generalization and
abstraction our old ideas get downgraded in stature.
Let us look at some more examples, now considering some new vector spaces.
Example VSP
The vector space of polynomials, Pn
Set:Pn, the set of all polynomials of degree nor less in the variable xwith coecients from C.
Version 2.30
322 Section VS Vector Spaces
Equality:
a0+a1x+a2x2++anxn=b0+b1x+b2x2++bnxnif and only if ai=bifor 0in
Vector Addition:
(a0+a1x+a2x2++anxn) + (b0+b1x+b2x2++bnxn) =
(a0+b0) + (a1+b1)x+ (a2+b2)x2++ (an+bn)xn
Scalar Multiplication:
(a0+a1x+a2x2++anxn) = (a0) + (a1)x+ (a2)x2++ (an)xn
This set, with these operations, will fulll the ten properties, though we will not work all the details
here. However, we will make a few comments and prove one of the properties. First, the zero vector
(Property Z [318]) is what you might expect, and you can check that it has the required property.
0= 0 + 0x+ 0x2++ 0xn
The additive inverse (Property AI [318]) is also no surprise, though consider how we have chosen to write
it.
a0+a1x+a2x2++anxn
= ( a0) + ( a1)x+ ( a2)x2++ ( an)xn
Now let's prove the associativity of vector addition (Property AA [317]). This is a bit tedious, though
necessary. Throughout, the plus sign (\+") does triple-duty. You might ask yourself what each plus sign
represents as you work through this proof.
u+(v+w)
= (a0+a1x++anxn) + ((b0+b1x++bnxn) + (c0+c1x++cnxn))
= (a0+a1x++anxn) + ((b0+c0) + (b1+c1)x++ (bn+cn)xn)
= (a0+ (b0+c0)) + (a1+ (b1+c1))x++ (an+ (bn+cn))xn
= ((a0+b0) +c0) + ((a1+b1) +c1)x++ ((an+bn) +cn)xn
= ((a0+b0) + (a1+b1)x++ (an+bn)xn) + (c0+c1x++cnxn)
= ((a0+a1x++anxn) + (b0+b1x++bnxn)) + (c0+c1x++cnxn)
= (u+v) +w
Notice how it is the application of the associativity of the (old) addition of complex numbers in the middle
of this chain of equalities that makes the whole proof happen. The remainder is successive applications of
our (new) denition of vector (polynomial) addition. Proving the remainder of the ten properties is similar
in style and tedium. You might try proving the commutativity of vector addition (Property C [317]), or
one of the distributivity properties (Property DVA [318], Property DSA [318]).
Example VSIS
The vector space of innite sequences
Set:C1=f(c0; c1; c2; c3; :::)jci2C; i2Ng.
Equality:
(c0; c1; c2; :::) = (d0; d1; d2; :::) if and only if ci=difor alli0
Vector Addition:
(c0; c1; c2; :::) + (d0; d1; d2; :::) = (c0+d0; c1+d1; c2+d2; :::)
Version 2.30
Subsection VS.EVS Examples of Vector Spaces 323
Scalar Multiplication:
(c0; c1; c2; c3; :::) = (c0; c 1; c 2; c 3; :::)
This should remind you of the vector space Cm, though now our lists of scalars are written horizontally
with commas as delimiters and they are allowed to be innite in length. What does the zero vector look
like (Property Z [318])? Additive inverses (Property AI [318])? Can you prove the associativity of vector
addition (Property AA [317])?
Example VSF
The vector space of functions
LetXbe any set. Set: F=ffjf:X!Cg.
Equality: f=gif and only if f(x) =g(x) for allx2X.
Vector Addition: f+gis the function with outputs dened by ( f+g)(x) =f(x) +g(x).
Scalar Multiplication: fis the function with outputs dened by ( f)(x) =f(x).
So this is the set of all functions of one variable that take elements of the set Xto a complex number.
You might have studied functions of one variable that take a real number to a real number, and that might
be a more natural set to use as X. But since we are allowing our scalars to be complex numbers, we need
to specify that the range of our functions is the complex numbers. Study carefully how the denitions of
the operation are made, and think about the dierent uses of \+" and juxtaposition. As an example of
what is required when verifying that this is a vector space, consider that the zero vector (Property Z [318])
is the function zwhose denition is z(x) = 0 for every input x2X.
Vector spaces of functions are very important in mathematics and physics, where the eld of scalars
may be the real numbers, so the ranges of the functions can in turn also be the set of real numbers.
Here's a unique example.
Example VSS
The singleton vector space
Set:Z=fzg.
Equality: Huh?
Vector Addition: z+z=z.
Scalar Multiplication: z=z.
This should look pretty wild. First, just what is z? Column vector, matrix, polynomial, sequence,
function? Mineral, plant, or animal? We aren't saying! zjust is. And we have denitions of vector
addition and scalar multiplication that are sucient for an occurrence of either that may come along.
Our only concern is if this set, along with the denitions of two operations, fullls the ten properties of
Denition VS [317]. Let's check associativity of vector addition (Property AA [317]). For all u;v;w2Z,
u+ (v+w) =z+ (z+z)
=z+z
= (z+z) +z
= (u+v) +w
What is the zero vector in this vector space (Property Z [318])? With only one element in the set, we do
not have much choice. Is z=0? It appears that zbehaves like the zero vector should, so it gets the title.
Maybe now the denition of this vector space does not seem so bizarre. It is a set whose only element is
the element that behaves like the zero vector, so that lone element isthe zero vector.
Perhaps some of the above denitions and verications seem obvious or like splitting hairs, but the
next example should convince you that they arenecessary. We will study this one carefully. Ready? Check
your preconceptions at the door.
Version 2.30
324 Section VS Vector Spaces
Example CVS
The crazy vector space
Set:C=f(x1; x2)jx1; x22Cg.
Vector Addition: ( x1; x2) + (y1; y2) = (x1+y1+ 1; x2+y2+ 1).
Scalar Multiplication: (x1; x2) = (x1+ 1; x 2+ 1).
Now, the rst thing I hear you say is \You can't do that!" And my response is, \Oh yes, I can!" I am
free to dene my set and my operations any way I please. They may not look natural, or even useful, but
we will now verify that they provide us with another example of a vector space. And that is enough. If
you are adventurous, you might try rst checking some of the properties yourself. What is the zero vector?
Additive inverses? Can you prove associativity? Ready, here we go.
Property AC [317], Property SC [317]: The result of each operation is a pair of complex numbers, so
these two closure properties are fullled.
Property C [317]:
u+v= (x1; x2) + (y1; y2) = (x1+y1+ 1; x2+y2+ 1)
= (y1+x1+ 1; y2+x2+ 1) = (y1; y2) + (x1; x2)
=v+u
Property AA [317]:
u+ (v+w) = (x1; x2) + ((y1; y2) + (z1; z2))
= (x1; x2) + (y1+z1+ 1; y2+z2+ 1)
= (x1+ (y1+z1+ 1) + 1; x2+ (y2+z2+ 1) + 1)
= (x1+y1+z1+ 2; x2+y2+z2+ 2)
= ((x1+y1+ 1) +z1+ 1;(x2+y2+ 1) +z2+ 1)
= (x1+y1+ 1; x2+y2+ 1) + (z1; z2)
= ((x1; x2) + (y1; y2)) + (z1; z2)
= (u+v) +w
Property Z [318]: The zero vector is . . . 0= ( 1; 1). Now I hear you say, \No, no, that can't be, it must
be (0;0)!" Indulge me for a moment and let us check my proposal.
u+0= (x1; x2) + ( 1; 1) = (x1+ ( 1) + 1; x2+ ( 1) + 1) = (x1; x2) =u
Feeling better? Or worse?
Property AI [318]: For each vector, u, we must locate an additive inverse, u. Here it is, (x1; x2) =
( x1 2; x2 2). As odd as it may look, I hope you are withholding judgment. Check:
u+ ( u) = (x1; x2) + ( x1 2; x2 2) = (x1+ ( x1 2) + 1; x2+ (x2 2) + 1) = ( 1; 1) =0
Property SMA [318]:
(u) =((x1; x2))
=(x1+ 1; x 2+ 1)
= ((x1+ 1) + 1; (x2+ 1) + 1)
= ((x 1+ ) + 1;(x 2+ ) + 1)
= (x 1+ 1; x 2+ 1)
= ()(x1; x2)
= ()u
Version 2.30
Subsection VS.VSP Vector Space Properties 325
Property DVA [318]: If you have hung on so far, here's where it gets even wilder. In the next two properties
we mix and mash the two operations.
(u+v) =((x1; x2) + (y1; y2))
=(x1+y1+ 1; x2+y2+ 1)
= ((x1+y1+ 1) + 1; (x2+y2+ 1) + 1)
= (x1+y1++ 1; x 2+y2++ 1)
= (x1+ 1 +y1+ 1 + 1; x 2+ 1 +y2+ 1 + 1)
= ((x1+ 1) + (y1+ 1) + 1;(x2+ 1) + (y2+ 1) + 1)
= (x1+ 1; x 2+ 1) + (y1+ 1; y 2+ 1)
=(x1; x2) +(y1; y2)
=u+v
Property DSA [318]:
(+)u= (+)(x1; x2)
= ((+)x1+ (+) 1;(+)x2+ (+) 1)
= (x1+x1++ 1; x 2+x2++ 1)
= (x1+ 1 +x1+ 1 + 1; x 2+ 1 +x2+ 1 + 1)
= ((x1+ 1) + (x1+ 1) + 1;(x2+ 1) + (x2+ 1) + 1)
= (x1+ 1; x 2+ 1) + (x1+ 1; x 2+ 1)
=(x1; x2) +(x1; x2)
=u+u
Property O [318]: After all that, this one is easy, but no less pleasing.
1u= 1(x1; x2) = (x1+ 1 1; x2+ 1 1) = (x1; x2) =u
That's it,Cis a vector space, as crazy as that may seem.
Notice that in the case of the zero vector and additive inverses, we only had to propose possibilities and
then verify that they were the correct choices. You might try to discover how you would arrive at these
choices, though you should understand why the process of discovering them is not a necessary component
of the proof itself.
Subsection VSP
Vector Space Properties
Subsection VS.EVS [319] has provided us with an abundance of examples of vector spaces, most of them
containing useful and interesting mathematical objects along with natural operations. In this subsection
we will prove some general properties of vector spaces. Some of these results will again seem obvious, but
it is important to understand why it is necessary to state and prove them. A typical hypothesis will be
\LetVbe a vector space." From this we may assume the ten properties of Denition VS [317], and nothing
more . Its like starting over, as we learn about what can happen in this new algebra we are learning. But
the power of this careful approach is that we can apply these theorems to any vector space we encounter
| those in the previous examples, or new ones we have not yet contemplated. Or perhaps new ones that
nobody has ever contemplated. We will illustrate some of these results with examples from the crazy vector
space (Example CVS [322]), but mostly we are stating theorems and doing proofs. These proofs do not
Version 2.30
326 Section VS Vector Spaces
get too involved, but are not trivial either, so these are good theorems to try proving yourself before you
study the proof given here. (See Technique P [774].)
First we show that there is just one zero vector. Notice that the properties only require there to be at
least one, and say nothing about there possibly being more. That is because we can use the ten properties
of a vector space (Denition VS [317]) to learn that there can never be more than one. To require that
this extra condition be stated as an eleventh property would make the denition of a vector space more
complicated than it needs to be.
Theorem ZVU
Zero Vector is Unique
Suppose that Vis a vector space. The zero vector, 0, is unique.
Proof To prove uniqueness, a standard technique is to suppose the existence of two objects (Technique
U [771]). So let 01and02be two zero vectors in V. Then
01=01+02 Property Z [318] for 02
=02+01 Property C [317]
=02 Property Z [318] for 01
This proves the uniqueness since the two zero vectors are really the same.
Theorem AIU
Additive Inverses are Unique
Suppose that Vis a vector space. For each u2V, the additive inverse, u, is unique.
Proof To prove uniqueness, a standard technique is to suppose the existence of two objects (Technique
U [771]). So let u1and u2be two additive inverses for u. Then
u1= u1+0 Property Z [318]
= u1+ (u+ u2) Property AI [318]
= ( u1+u) + u2 Property AA [317]
=0+ u2 Property AI [318]
= u2 Property Z [318]
So the two additive inverses are really the same.
As obvious as the next three theorems appear, nowhere have we guaranteed that the zero scalar, scalar
multiplication and the zero vector all interact this way. Until we have proved it, anyway.
Theorem ZSSM
Zero Scalar in Scalar Multiplication
Suppose that Vis a vector space and u2V. Then 0 u=0.
Proof Notice that 0 is a scalar, uis a vector, so Property SC [317] says 0 uis again a vector. As such,
0uhas an additive inverse, (0u) by Property AI [318].
0u=0+ 0u Property Z [318]
= ( (0u) + 0u) + 0u Property AI [318]
= (0u) + (0 u+ 0u) Property AA [317]
= (0u) + (0 + 0) u Property DSA [318]
= (0u) + 0u Property ZCN [759]
=0 Property AI [318]
Version 2.30
Subsection VS.VSP Vector Space Properties 327
Here's another theorem that looks like it should be obvious, but is still in need of a proof.
Theorem ZVSM
Zero Vector in Scalar Multiplication
Suppose that Vis a vector space and 2C. Then0=0.
Proof Notice that is a scalar, 0is a vector, so Property SC [317] means 0is again a vector. As such,
0has an additive inverse, (0) by Property AI [318].
0=0+0 Property Z [318]
= ( (0) +0) +0 Property AI [318]
= (0) + (0+0) Property AA [317]
= (0) +(0+0) Property DVA [318]
= (0) +0 Property Z [318]
=0 Property AI [318]
Here's another one that sure looks obvious. But understand that we have chosen to use certain notation
because it makes the theorem's conclusion look so nice. The theorem is not true because the notation looks
so good, it still needs a proof. If we had really wanted to make this point, we might have dened the additive
inverse of uasu]. Then we would have written the dening property, Property AI [318], as u+u]=0.
This theorem would become u]= ( 1)u. Not really quite as pretty, is it?
Theorem AISM
Additive Inverses from Scalar Multiplication
Suppose that Vis a vector space and u2V. Then u= ( 1)u.
Proof
u= u+0 Property Z [318]
= u+ 0u Theorem ZSSM [324]
= u+ (1 + ( 1))u
= u+ (1u+ ( 1)u) Property DSA [318]
= u+ (u+ ( 1)u) Property O [318]
= ( u+u) + ( 1)u Property AA [317]
=0+ ( 1)u Property AI [318]
= ( 1)u Property Z [318]
Because of this theorem, we can now write linear combinations like 6 u1+ ( 4)u2
as 6u1 4u2, even though we have not formally dened an operation called vector subtraction . Our
next theorem is a bit dierent from several of the others in the list. Rather than making a declaration
(\the zero vector is unique") it is an implication (\if. . . , then. . . ") and so can be used in proofs to convert
a vector equality into two possibilities, one a scalar equality and the other a vector equality. It should
remind you of the situation for complex numbers. If ; 2Cand= 0, then= 0 or= 0. This
critical property is the driving force behind using a factorization to solve a polynomial equation.
Version 2.30
328 Section VS Vector Spaces
Theorem SMEZV
Scalar Multiplication Equals the Zero Vector
Suppose that Vis a vector space and 2C. Ifu=0, then either = 0 or u=0.
Proof We prove this theorem by breaking up the analysis into two cases. The rst seems too trivial, and
it is, but the logic of the argument is still legitimate.
Case 1. Suppose = 0. In this case our conclusion is true (the rst part of the either/or is true) and
we are done. That was easy.
Case 2. Suppose 6= 0.
u= 1u Property O [318]
=1
u 6= 0
=1
(u) Property SMA [318]
=1
(0) Hypothesis
=0 Theorem ZVSM [325]
So in this case, the conclusion is true (the second part of the either/or is true) and we are done since the
conclusion was true in each of the two cases.
Example PCVS
Properties for the Crazy Vector Space
Several of the above theorems have interesting demonstrations when applied to the crazy vector space,
C(Example CVS [322]). We are not proving anything new here, or learning anything we did not know
already about C. It is just plain fun to see how these general theorems apply in a specic instance. For
most of our examples, the applications are obvious or trivial, but not with C.
Suppose u2C.
Then, as given by Theorem ZSSM [324],
0u= 0(x1; x2) = (0x1+ 0 1;0x2+ 0 1) = ( 1; 1) =0
And as given by Theorem ZVSM [325],
0=( 1; 1) = (( 1) + 1; ( 1) + 1)
= ( + 1; + 1) = ( 1; 1) =0
Finally, as given by Theorem AISM [325],
( 1)u= ( 1)(x1; x2) = (( 1)x1+ ( 1) 1;( 1)x2+ ( 1) 1)
= ( x1 2; x2 2) = u
Subsection RD
Recycling Denitions
When we say that Vis a vector space, we then know we have a set of objects (the \vectors"), but we also
know we have been provided with two operations (\vector addition" and \scalar multiplication") and these
operations behave with these objects according to the ten properties of Denition VS [317]. One combines
Version 2.30
Subsection VS.READ Reading Questions 329
two vectors and produces a vector, the other takes a scalar and a vector, producing a vector as the result.
So ifu1;u2;u32Vthen an expression like
5u1+ 7u2 13u3
would be unambiguous in anyof the vector spaces we have discussed in this section. And the resulting
object would be another vector in the vector space. If you were tempted to call the above expression a
linear combination, you would be right. Four of the denitions that were central to our discussions in
Chapter V [97] were stated in the context of vectors being column vectors , but were purposely kept broad
enough that they could be applied in the context of any vector space. They only rely on the presence of
scalars, vectors, vector addition and scalar multiplication to make sense. We will restate them shortly,
unchanged, except that their titles and acronyms no longer refer to column vectors, and the hypothesis of
being in a vector space has been added. Take the time now to look forward and review each one, and begin
to form some connections to what we have done earlier and what we will be doing in subsequent sections
and chapters. Specically, compare the following pairs of denitions:
Denition LCCV [109] and Denition LC [338]
Denition SSCV [131] and Denition SS [339]
Denition RLDCV [153] and Denition RLD [351]
Denition LICV [153] and Denition LI [351]
Subsection READ
Reading Questions
1. Comment on how the vector space Cmwent from a theorem (Theorem VSPCV [100]) to an example
(Example VSCV [319]).
2. In the crazy vector space, C, (Example CVS [322]) compute the linear combination
2(3;4) + ( 6)(1;2):
3. Suppose that is a scalar and 0is the zero vector. Why should we prove anything as obvious as
0=0such as we did in Theorem ZVSM [325]?
Version 2.30
330 Section VS Vector Spaces
Subsection EXC
Exercises
M10 Dene a possibly new vector space by beginning with the set and vector addition from C2(Example
VSCV [319]) but change the denition of scalar multiplication to
x=0=0
0
2C;x2C2
Prove that the rst nine properties required for a vector space hold, but Property O [318] does not hold.
This example shows us that we cannot expect to be able to derive Property O [318] as a consequence of
assuming the rst nine properties. In other words, we cannot slim down our list of properties by jettisoning
the last one, and still have the same collection of objects qualify as vector spaces.
Contributed by Robert Beezer
M11 LetVbe the set C2with the usual vector addition, but with scalar multiplication dened by
x
y
=y
x
Determine whether or not Vis a vector space with these operations.
Contributed by Chris Black Solution [330]
M12 LetVbe the set C2with the usual scalar multiplication, but with vector addition dened by
x
y
z
w
=y+w
x+z
Determine whether or not Vis a vector space with these operations.
Contributed by Chris Black Solution [330]
M13 LetVbe the setM2;2with the usual scalar multiplication, but with addition dened by AB=O2;2
for all 22 matricesAandB. Determine whether or not Vis a vector space with these operations.
Contributed by Chris Black Solution [330]
M14 LetVbe the setM2;2with the usual addition, but with scalar multiplication dened by A=O2;2
for all 22 matricesAand scalars . Determine whether or not Vis a vector space with these operations.
Contributed by Chris Black Solution [330]
M15 Consider the following sets of 3 3 matrices, where the symbol indicates the position of an arbitrary
complex number. Determine whether or not these sets form vector spaces with the usual operations of
addition and scalar multiplication for matrices.
1. All matrices of the form2
4 1
1
1 3
5
2. All matrices of the form2
40
00
03
5
3. All matrices of the form2
40 0
00
0 03
5(These are the diagonal matrices.)
Version 2.30
Subsection VS.EXC Exercises 331
4. All matrices of the form2
4
0
0 03
5(These are the upper triangular matrices.)
Contributed by Chris Black Solution [330]
M20 Explain why we need to dene the vector space Pnas the set of all polynomials with degree up to
and including ninstead of the more obvious set of all polynomials of degree exactlyn.
Contributed by Chris Black Solution [331]
M21 Does the set Z2=m
nm;n2Z
with the operations of standard addition and multiplication
of vectors form a vector space?
Contributed by Chris Black Solution [331]
T10 Prove each of the ten properties of Denition VS [317] for each of the following examples of a vector
space:
Example VSP [319]
Example VSIS [320]
Example VSF [321]
Example VSS [321]
Contributed by Robert Beezer
The next three problems suggest that under the right situations we can \cancel." In practice, these
techniques should be avoided in other proofs. Prove each of the following statements.
T21 Suppose that Vis a vector space, and u;v;w2V. Ifw+u=w+v, then u=v.
Contributed by Robert Beezer Solution [331]
T22 SupposeVis a vector space, u;v2Vandis a nonzero scalar from C. Ifu=v, then u=v.
Contributed by Robert Beezer Solution [331]
T23 SupposeVis a vector space, u6=0is a vector in Vand; 2C. Ifu=u, then=.
Contributed by Robert Beezer Solution [331]
T30 Suppose that Vis a vector space and 2Cis a scalar such that x=xfor every x2V. Prove
that= 1. In other words, Property O [318] is not duplicated for any other scalar but the \special" scalar,
1. (This question was suggested by James Gallagher.)
Contributed by Robert Beezer Solution [332]
Version 2.30
332 Section VS Vector Spaces
Subsection SOL
Solutions
M11 Contributed by Chris Black Statement [328]
The set C2with the proposed operations is not a vector space since Property O [318] is not valid. A
counterexample is 13
2
=2
3
6=3
2
, so in general, 1 u6=u.
M12 Contributed by Chris Black Statement [328]
Let's consider the existence of a zero vector, as required by Property Z [318] of a vector space. The
\regular" zero vector fails :x
y
0
0
=y
x
6=x
y
(remember that the property must hold for every
vector, not just for some). Is there another vector that lls the role of the zero vector? Suppose that
0=z1
z2
. Then for any vectorx
y
, we have
x
y
z1
z2
=y+z2
x+z1
=x
y
so thatx=y+z2andy=x+z1. This means that z1=y xandz2=x y. However, since xandy
can be any complex numbers, there are no xed complex numbers z1andz2that satisfy these equations.
Thus, there is no zero vector, Property Z [318] is not valid, and the set C2with the proposed operations
is not a vector space.
M13 Contributed by Chris Black Statement [328]
Since scalar multiplication remains unchanged, we only need to consider the axioms that involve vector
addition. Since every sum is the zero matrix, the rst 4 properties hold easily. However, there is no zero
vector in this set. Suppose that there was. Then there is a matrix Zso thatA+Z=Afor any 22
matrixA. However, A+Z=O2;2, which is in general not equal to A, so Property Z [318] fails and this
set is not a vector space.
M14 Contributed by Chris Black Statement [328]
Since addition is unchanged, we only need to check the axioms involving scalar multiplication. The proposed
scalar multiplication clearly fails Property O [318] : 1 A=O2;26=A. Thus, the proposed set is not a vector
space.
M15 Contributed by Chris Black Statement [328]
There is something to notice here that will make our job much easier: Since each of these sets are comprised
of 33 matrices with the standard operations of addition and scalar multiplication of matrices, the last 8
properties will automatically hold. That is, we really only need to verify Property AC [317] and Property
SC [317].
a). This set is not closed under either scalar multiplication or addition (fails Property AC [317] and
Property SC [317]). For example, 32
4 1
1
1 3
5=2
4 3
3
3 3
5is not a member of the proposed set.
b). This set is closed under both scalar multiplication and addition, so this set is a vector space with the
standard operation of addition and scalar multiplication.
c). This set is closed under both scalar multiplication and addition, so this set is a vector space with the
standard operation of addition and scalar multiplication.
d). This set is closed under both scalar multiplication and addition, so this set is a vector space with the
standard operation of addition and scalar multiplication.
Version 2.30
Subsection VS.SOL Solutions 333
M20 Contributed by Chris Black Statement [329]
Hint: The set of all polynomials of degree exactlynfails one of the closure properties of a vector space.
Which one, and why?
M21 Contributed by Robert Beezer Statement [329]
Additive closure will hold, but scalar closure will not. The best way to convince yourself of this is to
construct a counterexample. Such as,1
22Cand1
0
2Z2, however1
21
0
=1
2
0
62Z2, which violates
Property SC [317]. So Z2is not a vector space.
T21 Contributed by Robert Beezer Statement [329]
u=0+u Property Z [318]
= ( w+w) +u Property AI [318]
= w+ (w+u) Property AA [317]
= w+ (w+v) Hypothesis
= ( w+w) +v Property AA [317]
=0+v Property AI [318]
=v Property Z [318]
T22 Contributed by Robert Beezer Statement [329]
u= 1u Property O [318]
=1
u 6= 0
=1
(u) Property SMA [318]
=1
(v) Hypothesis
=1
v Property SMA [318]
= 1v
=v Property O [318]
T23 Contributed by Robert Beezer Statement [329]
0=u+ (u) Property AI [318]
=u+ (u) Hypothesis
=u+ ( 1) (u) Theorem AISM [325]
=u+ (( 1))u Property SMA [318]
=u+ ( )u
= ( )u Property DSA [318]
By hypothesis, u6=0, so Theorem SMEZV [326] implies
0 =
Version 2.30
334 Section VS Vector Spaces
=
T30 Contributed by Robert Beezer Statement [329]
We have,
0=x x Property AI [318]
=x x Hypothesis
=x 1x Property O [318]
= ( 1)x Property DSA [318]
So by Theorem SMEZV [326] we conclude that 1 = 0 or x=0. However, since our hypothesis was for
every x2V, we are left with the rst possibility and = 1.
There is one
aw in the proof above, and as stated, the problem is not correct either. Can you spot
the
aw and as a result correct the problem statement? (Hint: Example VSS [321]).
Version 2.30
Section S Subspaces 335
Section S
Subspaces
A subspace is a vector space that is contained within another vector space. So every subspace is a vector
space in its own right, but it is also dened relative to some other (larger) vector space. We will discover
shortly that we are already familiar with a wide variety of subspaces from previous sections. Here's the
denition.
Denition S
Subspace
Suppose that VandWare two vector spaces that have identical denitions of vector addition and scalar
multiplication, and that Wis a subset of V,WV. ThenWis asubspace ofV. 4
Lets look at an example of a vector space inside another vector space.
Example SC3
A subspace of C3
We know that C3is a vector space (Example VSCV [319]). Consider the subset,
W=8
<
:2
4x1
x2
x33
52x1 5x2+ 7x3= 09
=
;
It is clear that WC3, since the objects in Ware column vectors of size 3. But is Wa vector space? Does
it satisfy the ten properties of Denition VS [317] when we use the same operations? That is the main
question. Suppose x=2
4x1
x2
x33
5andy=2
4y1
y2
y33
5are vectors from W. Then we know that these vectors cannot
be totally arbitrary, they must have gained membership in Wby virtue of meeting the membership test.
For example, we know that xmust satisfy 2 x1 5x2+ 7x3= 0 while ymust satisfy 2 y1 5y2+ 7y3= 0.
Our rst property (Property AC [317]) asks the question, is x+y2W? When our set of vectors was C3,
this was an easy question to answer. Now it is not so obvious. Notice rst that
x+y=2
4x1
x2
x33
5+2
4y1
y2
y33
5=2
4x1+y1
x2+y2
x3+y33
5
and we can test this vector for membership in Was follows,
2(x1+y1) 5(x2+y2) + 7(x3+y3) = 2x1+ 2y1 5x2 5y2+ 7x3+ 7y3
= (2x1 5x2+ 7x3) + (2y1 5y2+ 7y3)
= 0 + 0 x2W;y2W
= 0
and by this computation we see that x+y2W. One property down, nine to go.
Ifis a scalar and x2W, is it always true that x2W? This is what we need to establish Property
SC [317]. Again, the answer is not as obvious as it was when our set of vectors was all of C3. Let's see.
x=2
4x1
x2
x33
5=2
4x1
x2
x33
5
Version 2.30
336 Section S Subspaces
and we can test this vector for membership in Wwith
2(x1) 5(x2) + 7(x3) =(2x1 5x2+ 7x3)
=0 x2W
= 0
and we see that indeed x2W. Always.
IfWhas a zero vector, it will be unique (Theorem ZVU [324]). The zero vector for C3should also
perform the required duties when added to elements of W. So the likely candidate for a zero vector in
Wis the same zero vector that we know C3has. You can check that 0=2
40
0
03
5is a zero vector in Wtoo
(Property Z [318]).
With a zero vector, we can now ask about additive inverses (Property AI [318]). As you might suspect,
the natural candidate for an additive inverse in Wis the same as the additive inverse from C3. However,
we must insure that these additive inverses actually are elements of W. Given x2W, is x2W?
x=2
4 x1
x2
x33
5
and we can test this vector for membership in Wwith
2( x1) 5( x2) + 7( x3) = (2x1 5x2+ 7x3)
= 0 x2W
= 0
and we now believe that x2W.
Is the vector addition in Wcommutative (Property C [317])? Is x+y=y+x? Of course! Nothing
about restricting the scope of our set of vectors will prevent the operation from still being commutative.
Indeed, the remaining ve properties are unaected by the transition to a smaller set of vectors, and so
remain true. That was convenient.
SoWsatises all ten properties, is therefore a vector space, and thus earns the title of being a subspace
ofC3.
Subsection TS
Testing Subspaces
In Example SC3 [333] we proceeded through all ten of the vector space properties before believing that
a subset was a subspace. But six of the properties were easy to prove, and we can lean on some of the
properties of the vector space (the superset) to make the other four easier. Here is a theorem that will
make it easier to test if a subset is a vector space. A shortcut if there ever was one.
Theorem TSS
Testing Subsets for Subspaces
Suppose that Vis a vector space and Wis a subset of V,WV. EndowWwith the same operations as
V. ThenWis a subspace if and only if three conditions are met
1.Wis non-empty, W6=;.
2. Ifx2Wandy2W, then x+y2W.
Version 2.30
Subsection S.TS Testing Subspaces 337
3. If2Candx2W, thenx2W.
Proof ()) We have the hypothesis that Wis a subspace, so by Property Z [318] we know that W
contains a zero vector. This is enough to show that W6=;. Also, since Wis a vector space it satises
the additive and scalar multiplication closure properties (Property AC [317], Property SC [317]), and so
exactly meets the second and third conditions. If that was easy, the the other direction might require a
bit more work.
(() We have three properties for our hypothesis, and from this we should conclude that Whas the
ten dening properties of a vector space. The second and third conditions of our hypothesis are exactly
Property AC [317] and Property SC [317]. Our hypothesis that Vis a vector space implies that Property
C [317], Property AA [317], Property SMA [318], Property DVA [318], Property DSA [318] and Property
O [318] all hold. They continue to be true for vectors from Wsince passing to a subset, and keeping the
operation the same, leaves their statements unchanged. Eight down, two to go.
Suppose x2W. Then by the third part of our hypothesis (scalar closure), we know that ( 1)x2W.
By Theorem AISM [325] ( 1)x= x, so together these statements show us that x2W. xis the
additive inverse of xinV, but will continue in this role when viewed as element of the subset W. So every
element of Whas an additive inverse that is an element of Wand Property AI [318] is completed. Just
one property left.
While we have implicitly discussed the zero vector in the previous paragraph, we need to be certain
that the zero vector (of V) really lives in W. SinceWis non-empty, we can choose some vector z2W.
Then by the argument in the previous paragraph, we know z2W. Now by Property AI [318] for Vand
then by the second part of our hypothesis (additive closure) we see that
0=z+ ( z)2W
SoWcontain the zero vector from V. Since this vector performs the required duties of a zero vector in V,
it will continue in that role as an element of W. This gives us, Property Z [318], the nal property of the
ten required. (Sarah Fellez contributed to this proof.)
So just three conditions, plus being a subset of a known vector space, gets us all ten properties.
Fabulous! This theorem can be paraphrased by saying that a subspace is \a non-empty subset (of a vector
space) that is closed under vector addition and scalar multiplication."
You might want to go back and rework Example SC3 [333] in light of this result, perhaps seeing where
we can now economize or where the work done in the example mirrored the proof and where it did not.
We will press on and apply this theorem in a slightly more abstract setting.
Example SP4
A subspace of P4
P4is the vector space of polynomials with degree at most 4 (Example VSP [319]). Dene a subset Was
W=fp(x)jp2P4; p(2) = 0g
soWis the collection of those polynomials (with degree 4 or less) whose graphs cross the x-axis atx= 2.
Whenever we encounter a new set it is a good idea to gain a better understanding of the set by nding a
few elements in the set, and a few outside it. For example x2 x 22W, whilex4+x3 762W.
IsWnonempty? Yes, x 22W.
Additive closure? Suppose p2Wandq2W. Isp+q2W?pandqare not totally arbitrary, we
know thatp(2) = 0 and q(2) = 0. Then we can check p+qfor membership in W,
(p+q)(2) =p(2) +q(2) Addition in P4
= 0 + 0 p2W; q2W
Version 2.30
338 Section S Subspaces
= 0
so we see that p+qqualies for membership in W.
Scalar multiplication closure? Suppose that 2Candp2W. Then we know that p(2) = 0. Testing
pfor membership,
(p)(2) =p(2) Scalar multiplication in P4
=0 p2W
= 0
sop2W.
We have shown that Wmeets the three conditions of Theorem TSS [334] and so qualies as a subspace
ofP4. Notice that by Denition S [333] we now know that Wis also a vector space. So all the properties
of a vector space (Denition VS [317]) and the theorems of Section VS [317] apply in full.
Much of the power of Theorem TSS [334] is that we can easily establish new vector spaces if we can
locate them as subsets of other vector spaces, such as the ones presented in Subsection VS.EVS [319].
It can be as instructive to consider some subsets that are notsubspaces. Since Theorem TSS [334] is an
equivalence (see Technique E [768]) we can be assured that a subset is not a subspace if it violates one of
the three conditions, and in any example of interest this will not be the \non-empty" condition. However,
since a subspace has to be a vector space in its own right, we can also search for a violation of any one of
the ten dening properties in Denition VS [317] or any inherent property of a vector space, such as those
given by the basic theorems of Subsection VS.VSP [323]. Notice also that a violation need only be for a
specic vector or pair of vectors.
Example NSC2Z
A non-subspace in C2, zero vector
Consider the subset Wbelow as a candidate for being a subspace of C2
W=x1
x23x1 5x2= 12
The zero vector of C2,0=0
0
will need to be the zero vector in Walso. However, 062Wsince
3(0) 5(0) = 06= 12. SoWhas no zero vector and fails Property Z [318] of Denition VS [317]. This
subspace also fails to be closed under addition and scalar multiplication. Can you nd examples of this?
Example NSC2A
A non-subspace in C2, additive closure
Consider the subset Xbelow as a candidate for being a subspace of C2
X=x1
x2x1x2= 0
You can check that 02X, so the approach of the last example will not get us anywhere. However, notice
thatx=1
0
2Xandy=0
1
2X. Yet
x+y=1
0
+0
1
=1
1
62X
Version 2.30
Subsection S.TS Testing Subspaces 339
SoXfails the additive closure requirement of either Property AC [317] or Theorem TSS [334], and is
therefore not a subspace.
Example NSC2S
A non-subspace in C2, scalar multiplication closure
Consider the subset Ybelow as a candidate for being a subspace of C2
Y=x1
x2x12Z; x22Z
Zis the set of integers, so we are only allowing \whole numbers" as the constituents of our vectors. Now,
02Y, and additive closure also holds (can you prove these claims?). So we will have to try something
dierent. Note that =1
22Cand2
3
2Y, but
x=1
22
3
=1
3
2
62Y
SoYfails the scalar multiplication closure requirement of either Property SC [317] or Theorem TSS [334],
and is therefore not a subspace.
There are two examples of subspaces that are trivial. Suppose that Vis any vector space. Then Vis
a subset of itself and is a vector space. By Denition S [333], Vqualies as a subspace of itself. The set
containing just the zero vector Z=f0gis also a subspace as can be seen by applying Theorem TSS [334]
or by simple modications of the techniques hinted at in Example VSS [321]. Since these subspaces are so
obvious (and therefore not too interesting) we will refer to them as being trivial.
Denition TS
Trivial Subspaces
Given the vector space V, the subspaces Vandf0gare each called a trivial subspace .4
We can also use Theorem TSS [334] to prove more general statements about subspaces, as illustrated
in the next theorem.
Theorem NSMS
Null Space of a Matrix is a Subspace
Suppose that Ais anmnmatrix. Then the null space of A,N(A), is a subspace of Cn.
Proof We will examine the three requirements of Theorem TSS [334]. Recall that N(A) =fx2CnjAx=0g.
First, 02N(A), which can be inferred as a consequence of Theorem HSC [71]. So N(A)6=;.
Second, check additive closure by supposing that x2N (A) and y2N (A). So we know a little
something about xandy:Ax=0andAy=0, and that is all we know. Question: Is x+y2N(A)?
Let's check.
A(x+y) =Ax+Ay Theorem MMDAA [230]
=0+0 x 2N(A);y2N(A)
=0 Theorem VSPCV [100]
So, yes, x+yqualies for membership in N(A).
Third, check scalar multiplication closure by supposing that 2Candx2N(A). So we know a little
something about x:Ax=0, and that is all we know. Question: Is x2N(A)? Let's check.
A(x) =(Ax) Theorem MMSMM [230]
=0 x 2N(A)
Version 2.30
340 Section S Subspaces
=0 Theorem ZVSM [325]
So, yes,xqualies for membership in N(A).
Having met the three conditions in Theorem TSS [334] we can now say that the null space of a matrix
is a subspace (and hence a vector space in its own right!).
Here is an example where we can exercise Theorem NSMS [337].
Example RSNS
Recasting a subspace as a null space
Consider the subset of C5dened as
W=8
>>>><
>>>>:2
66664x1
x2
x3
x4
x53
777753x1+x2 5x3+ 7x4+x5= 0;
4x1+ 6x2+ 3x3 6x4 5x5= 0;
2x1+ 4x2+ 7x4+x5= 09
>>>>=
>>>>;
It is possible to show that Wis a subspace of C5by checking the three conditions of Theorem TSS [334]
directly, but it will get tedious rather quickly. Instead, give Wa fresh look and notice that it is a set of
solutions to a homogeneous system of equations. Dene the matrix
A=2
43 1 5 7 1
4 6 3 6 5
2 4 0 7 13
5
and then recognize that W=N(A). By Theorem NSMS [337] we can immediately see that Wis a
subspace. Boom!
Subsection TSS
The Span of a Set
The span of a set of column vectors got a heavy workout in Chapter V [97] and Chapter M [207]. The
denition of the span depended only on being able to formulate linear combinations. In any of our more
general vector spaces we always have a denition of vector addition and of scalar multiplication. So we
can build linear combinations and manufacture spans. This subsection contains two denitions that are
just mild variants of denitions we have seen earlier for column vectors. If you haven't already, compare
them with Denition LCCV [109] and Denition SSCV [131].
Denition LC
Linear Combination
Suppose that Vis a vector space. Given nvectors u1;u2;u3; :::; unandnscalars1; 2; 3; :::; n,
theirlinear combination is the vector
1u1+2u2+3u3++nun:
4
Example LCM
A linear combination of matrices
In the vector space M23of 23 matrices, we have the vectors
x=1 3 2
2 0 7
y=3 1 2
5 5 1
z=4 2 4
1 1 1
Version 2.30
Subsection S.TSS The Span of a Set 341
and we can form linear combinations such as
2x+ 4y+ ( 1)z= 21 3 2
2 0 7
+ 43 1 2
5 5 1
+ ( 1)4 2 4
1 1 1
=2 6 4
4 0 14
+12 4 8
20 20 4
+ 4 2 4
1 1 1
=10 0 8
23 19 17
or,
4x 2y+ 3z= 41 3 2
2 0 7
23 1 2
5 5 1
+ 34 2 4
1 1 1
=4 12 8
8 0 28
+ 6 2 4
10 10 2
+12 6 12
3 3 3
=10 20 24
1 7 29
When we realize that we can form linear combinations in any vector space, then it is natural to revisit
our denition of the span of a set, since it is the set of allpossible linear combinations of a set of vectors.
Denition SS
Span of a Set
Suppose that Vis a vector space. Given a set of vectors S=fu1;u2;u3; :::; utg, their span ,hSi, is the
set of all possible linear combinations of u1;u2;u3; :::; ut. Symbolically,
hSi=f1u1+2u2+3u3++tutji2C;1itg
=(tX
i=1iuii2C;1it)
4
Theorem SSS
Span of a Set is a Subspace
SupposeVis a vector space. Given a set of vectors S=fu1;u2;u3; :::; utgV, their span,hSi, is a
subspace.
Proof We will verify the three conditions of Theorem TSS [334]. First,
0=0+0+0+:::+0 Property Z [318] for V
= 0u1+ 0u2+ 0u3++ 0ut Theorem ZSSM [324]
So we have written 0as a linear combination of the vectors in Sand by Denition SS [339] ;02hSiand
thereforeS6=;.
Second, suppose x2hSiandy2hSi. Can we conclude that x+y2hSi? What do we know about
xandyby virtue of their membership in hSi? There must be scalars from C,1; 2; 3; :::; tand
1; 2; 3; :::; tso that
x=1u1+2u2+3u3++tut
y=1u1+2u2+3u3++tut
Version 2.30
342 Section S Subspaces
Then
x+y=1u1+2u2+3u3++tut
+1u1+2u2+3u3++tut
=1u1+1u1+2u2+2u2
+3u3+3u3++tut+tut Property AA [317], Property C [317]
= (1+1)u1+ (2+2)u2
+ (3+3)u3++ (t+t)ut Property DSA [318]
Since eachi+iis again a scalar from Cwe have expressed the vector sum x+yas a linear combination
of the vectors from S, and therefore by Denition SS [339] we can say that x+y2hSi.
Third, suppose 2Candx2hSi. Can we conclude that x2hSi? What do we know about xby
virtue of its membership in hSi? There must be scalars from C,1; 2; 3; :::; tso that
x=1u1+2u2+3u3++tut
Then
x=(1u1+2u2+3u3++tut)
=(1u1) +(2u2) +(3u3) ++(tut) Property DVA [318]
= (1)u1+ (2)u2+ (3)u3++ (t)ut Property SMA [318]
Since eachiis again a scalar from Cwe have expressed the scalar multiple xas a linear combination
of the vectors from S, and therefore by Denition SS [339] we can say that x2hSi.
With the three conditions of Theorem TSS [334] met, we can say that hSiis a subspace (and so is
also vector space, Denition VS [317]). (See Exercise SS.T20 [144], Exercise SS.T21 [144], Exercise SS.T22
[144].)
Example SSP
Span of a set of polynomials
In Example SP4 [335] we proved that
W=fp(x)jp2P4; p(2) = 0g
is a subspace of P4, the vector space of polynomials of degree at most 4. Since Wis a vector space itself,
let's construct a span within W. First let
S=
x4 4x3+ 5x2 x 2;2x4 3x3 6x2+ 6x+ 4
and verify that Sis a subset of Wby checking that each of these two polynomials has x= 2 as a root.
Now, if we dene U=hSi, then Theorem SSS [339] tells us that Uis a subspace of W. So quite quickly
we have built a chain of subspaces, UinsideW, andWinsideP4.
Rather than dwell on how quickly we can build subspaces, let's try to gain a better understanding of
just how the span construction creates subspaces, in the context of this example. We can quickly build
representative elements of U,
3(x4 4x3+ 5x2 x 2) + 5(2x4 3x3 6x2+ 6x+ 4) = 13x4 27x3 15x2+ 27x+ 14
and
( 2)(x4 4x3+ 5x2 x 2) + 8(2x4 3x3 6x2+ 6x+ 4) = 14x4 16x3 58x2+ 50x+ 36
Version 2.30
Subsection S.TSS The Span of a Set 343
and each of these polynomials must be in Wsince it is closed under addition and scalar multiplication.
But you might check for yourself that both of these polynomials have x= 2 as a root.
I can tell you that y= 3x4 7x3 x2+ 7x 2 is not inU, but would you believe me? A rst check
shows that ydoes havex= 2 as a root, but that only shows that y2W. What does yhave to do to gain
membership in U=hSi? It must be a linear combination of the vectors in S,x4 4x3+ 5x2 x 2 and
2x4 3x3 6x2+ 6x+ 4. So let's suppose that yis such a linear combination,
y= 3x4 7x3 x2+ 7x 2
=1(x4 4x3+ 5x2 x 2) +2(2x4 3x3 6x2+ 6x+ 4)
= (1+ 22)x4+ ( 41 32)x3+ (51 62)x2+ ( 1+ 62)x ( 21+ 42)
Notice that operations above are done in accordance with the denition of the vector space of polynomials
(Example VSP [319]). Now, if we equate coecients, which is the denition of equality for polynomials,
then we obtain the system of ve linear equations in two variables
1+ 22= 3
41 32= 7
51 62= 1
1+ 62= 7
21+ 42= 2
Build an augmented matrix from the system and row-reduce,
2
666641 2 3
4 3 7
5 6 1
1 6 7
2 4 23
77775RREF !2
66666410 0
010
0 0 1
0 0 0
0 0 03
777775
With a leading 1 in the nal column of the row-reduced augmented matrix, Theorem RCLS [58] tells us
the system of equations is inconsistent. Therefore, there are no scalars, 1and2, to establish yas a
linear combination of the elements in U. Soy62U.
Let's again examine membership in a span.
Example SM32
A subspace of M32
The set of all 32 matrices forms a vector space when we use the operations of matrix addition (Denition
MA [207]) and scalar matrix multiplication (Denition MSM [208]), as was show in Example VSM [319].
Consider the subset
S=8
<
:2
43 1
4 2
5 53
5;2
41 1
2 1
14 13
5;2
43 1
1 2
19 113
5;2
44 2
1 2
14 23
5;2
43 1
4 0
17 73
59
=
;
and dene a new subset of vectors WinM32using the span (Denition SS [339]), W=hSi. So by
Theorem SSS [339] we know that Wis a subspace of M32. WhileWis an innite set, and this is a precise
description, it would still be worthwhile to investigate whether or not Wcontains certain elements.
First, is
y=2
49 3
7 3
10 113
5
Version 2.30
344 Section S Subspaces
inW? To answer this, we want to determine if ycan be written as a linear combination of the ve matrices
inS. Can we nd scalars, 1; 2; 3; 4; 5so that
2
49 3
7 3
10 113
5=12
43 1
4 2
5 53
5+22
41 1
2 1
14 13
5+32
43 1
1 2
19 113
5+42
44 2
1 2
14 23
5+52
43 1
4 0
17 73
5
=2
431+2+ 33+ 44+ 351+2 3+ 24+5
41+ 22 3+4 45 21 2+ 23 24
51+ 142 193+ 144 175 51 2 113 24+ 753
5
Using our denition of matrix equality (Denition ME [207]) we can translate this statement into six
equations in the ve unknowns,
31+2+ 33+ 44+ 35= 9
1+2 3+ 24+5= 3
41+ 22 3+4 45= 7
21 2+ 23 24= 3
51+ 142 193+ 144 175= 10
51 2 113 24+ 75= 11
This is a linear system of equations, which we can represent with an augmented matrix and row-reduce in
search of solutions. The matrix that is row-equivalent to the augmented matrix is
2
6666666410 0 05
82
010 0 19
4 1
0 0 10 7
80
0 0 0 117
81
0 0 0 0 0 0
0 0 0 0 0 03
77777775
So we recognize that the system is consistent since there is no leading 1 in the nal column (Theorem RCLS
[58]), and compute n r= 5 4 = 1 free variables (Theorem FVCS [60]). While there are innitely many
solutions, we are only in pursuit of a single solution, so let's choose the free variable 5= 0 for simplicity's
sake. Then we easily see that 1= 2,2= 1,3= 0,4= 1. So the scalars 1= 2,2= 1,3= 0,
4= 1,5= 0 will provide a linear combination of the elements of Sthat equals y, as we can verify by
checking,
2
49 3
7 3
10 113
5= 22
43 1
4 2
5 53
5+ ( 1)2
41 1
2 1
14 13
5+ (1)2
44 2
1 2
14 23
5
So with one particular linear combination in hand, we are convinced that ydeserves to be a member of
W=hSi. Second, is
x=2
42 1
3 1
4 23
5
inW? To answer this, we want to determine if xcan be written as a linear combination of the ve matrices
inS. Can we nd scalars, 1; 2; 3; 4; 5so that
2
42 1
3 1
4 23
5=12
43 1
4 2
5 53
5+22
41 1
2 1
14 13
5+32
43 1
1 2
19 113
5+42
44 2
1 2
14 23
5+52
43 1
4 0
17 73
5
Version 2.30
Subsection S.SC Subspace Constructions 345
=2
431+2+ 33+ 44+ 351+2 3+ 24+5
41+ 22 3+4 45 21 2+ 23 24
51+ 142 193+ 144 175 51 2 113 24+ 753
5
Using our denition of matrix equality (Denition ME [207]) we can translate this statement into six
equations in the ve unknowns,
31+2+ 33+ 44+ 35= 2
1+2 3+ 24+5= 1
41+ 22 3+4 45= 3
21 2+ 23 24= 1
51+ 142 193+ 144 175= 4
51 2 113 24+ 75= 2
This is a linear system of equations, which we can represent with an augmented matrix and row-reduce in
search of solutions. The matrix that is row-equivalent to the augmented matrix is
2
6666666410 0 05
80
010 0 38
80
0 0 10 7
80
0 0 0 1 17
80
0 0 0 0 0 1
0 0 0 0 0 03
77777775
With a leading 1 in the last column Theorem RCLS [58] tells us that the system is inconsistent. Therefore,
there are no values for the scalars that will place xinW, and so we conclude that x62W.
Notice how Example SSP [340] and Example SM32 [341] contained questions about membership in a
span, but these questions quickly became questions about solutions to a system of linear equations. This
will be a common theme going forward.
Subsection SC
Subspace Constructions
Several of the subsets of vectors spaces that we worked with in Chapter M [207] are also subspaces | they
are closed under vector addition and scalar multiplication in Cm.
Theorem CSMS
Column Space of a Matrix is a Subspace
Suppose that Ais anmnmatrix. ThenC(A) is a subspace of Cm.
Proof Denition CSM [271] shows us that C(A) is a subset of Cm, and that it is dened as the span of
a set of vectors from Cm(the columns of the matrix). Since C(A) is a span, Theorem SSS [339] says it is
a subspace.
That was easy! Notice that we could have used this same approach to prove that the null space is a
subspace, since Theorem SSNS [137] provided a description of the null space of a matrix as the span of a
set of vectors. However, I much prefer the current proof of Theorem NSMS [337]. Speaking of easy, here
is a very easy theorem that exposes another of our constructions as creating subspaces.
Version 2.30
346 Section S Subspaces
Theorem RSMS
Row Space of a Matrix is a Subspace
Suppose that Ais anmnmatrix. ThenR(A) is a subspace of Cn.
Proof Denition RSM [278] says R(A) =C
At
, so the row space of a matrix is a column space, and
every column space is a subspace by Theorem CSMS [343]. That's enough.
One more.
Theorem LNSMS
Left Null Space of a Matrix is a Subspace
Suppose that Ais anmnmatrix. ThenL(A) is a subspace of Cm.
Proof Denition LNS [293] says L(A) =N
At
, so the left null space is a null space, and every null
space is a subspace by Theorem NSMS [337]. Done.
So the span of a set of vectors, and the null space, column space, row space and left null space of a
matrix are all subspaces, and hence are all vector spaces, meaning they have all the properties detailed
in Denition VS [317] and in the basic theorems presented in Section VS [317]. We have worked with
these objects as just sets in Chapter V [97] and Chapter M [207], but now we understand that they have
much more structure. In particular, being closed under vector addition and scalar multiplication means a
subspace is also closed under linear combinations.
Subsection READ
Reading Questions
1. Summarize the three conditions that allow us to quickly test if a set is a subspace.
2. Consider the set of vectors
W=8
<
:2
4a
b
c3
53a 2b+c= 59
=
;
Is the setWa subspace of C3? Explain your answer.
3. Name ve general constructions of sets of column vectors (subsets of Cm) that we now know as
subspaces.
Version 2.30
Subsection S.EXC Exercises 347
Subsection EXC
Exercises
C15 Working within the vector space C3, determine if b=2
44
3
13
5is in the subspace W,
W=*8
<
:2
43
2
33
5;2
41
0
33
5;2
41
1
03
5;2
42
1
33
59
=
;+
Contributed by Chris Black Solution [347]
C16 Working within the vector space C4, determine if b=2
6641
1
0
13
775is in the subspace W,
W=*8
>><
>>:2
6641
2
1
13
775;2
6641
0
3
13
775;2
6642
1
1
23
7759
>>=
>>;+
Contributed by Chris Black Solution [347]
C17 Working within the vector space C4, determine if b=2
6642
1
2
13
775is in the subspace W,
W=*8
>><
>>:2
6641
2
0
23
775;2
6641
0
3
13
775;2
6640
1
0
23
775;2
6641
1
2
03
7759
>>=
>>;+
Contributed by Chris Black Solution [347]
C20 Working within the vector space P3of polynomials of degree 3 or less, determine if p(x) =x3+6x+4
is in the subspace Wbelow.
W=
x3+x2+x; x3+ 2x 6; x2 5
Contributed by Robert Beezer Solution [347]
C21 Consider the subspace
W=2 1
3 1
;4 0
2 3
; 3 1
2 1
of the vector space of 2 2 matrices, M22. IsC= 3 3
6 4
an element of W?
Contributed by Robert Beezer Solution [348]
Version 2.30
348 Section S Subspaces
C25 Show that the set W=x1
x23x1 5x2= 12
from Example NSC2Z [336] fails Property AC
[317] and Property SC [317].
Contributed by Robert Beezer
C26 Show that the set Y=x1
x2x12Z; x22Z
from Example NSC2S [337] has Property AC [317].
Contributed by Robert Beezer
M20 InC3, the vector space of column vectors of size 3, prove that the set Zis a subspace.
Z=8
<
:2
4x1
x2
x33
54x1 x2+ 5x3= 09
=
;
Contributed by Robert Beezer Solution [348]
T20 A square matrix Aof sizenis upper triangular if [ A]ij= 0 whenever i > j . LetUTnbe the set
of all upper triangular matrices of size n. Prove that UTnis a subspace of the vector space of all square
matrices of size n,Mnn.
Contributed by Robert Beezer Solution [349]
T30 LetPbe the set of all polynomials, of any degree. The set Pis a vector space. Let Ebe the
subset ofPconsisting of all polynomials with only terms of even degree. Prove or disprove: the set Eis a
subspace of P.
Contributed by Chris Black Solution [350]
T31 LetPbe the set of all polynomials, of any degree. The set Pis a vector space. Let Fbe the subset
ofPconsisting of all polynomials with only terms of odd degree. Prove or disprove: the set Fis a subspace
ofP.
Contributed by Chris Black Solution [350]
Version 2.30
Subsection S.SOL Solutions 349
Subsection SOL
Solutions
C15 Contributed by Chris Black Statement [345]
Forbto be an element of W=hSithere must be linear combination of the vectors in Sthat equals b
(Denition SSCV [131]). The existence of such scalars is equivalent to the linear system LS(A;b) being
consistent, where Ais the matrix whose columns are the vectors from S(Theorem SLSLC [112]).
2
43 1 1 2 4
2 0 1 1 3
3 3 0 3 13
5RREF !2
410 1=2 1=2 0
01 1=2 1=2 0
0 0 0 0 13
5
So by Theorem RCLS [58] the system is inconsistent, which indicates that bis not an element of the
subspaceW.
C16 Contributed by Chris Black Statement [345]
Forbto be an element of W=hSithere must be linear combination of the vectors in Sthat equals b
(Denition SSCV [131]). The existence of such scalars is equivalent to the linear system LS(A;b) being
consistent, where Ais the matrix whose columns are the vectors from S(Theorem SLSLC [112]).
2
6641 1 2 1
2 0 1 1
1 3 1 0
1 1 2 13
775RREF !2
66410 0 1=3
010 0
0 0 11=3
0 0 0 03
775
So by Theorem RCLS [58] the system is consistent, which indicates that bis in the subspace W.
C17 Contributed by Chris Black Statement [345]
Forbto be an element of W=hSithere must be linear combination of the vectors in Sthat equals b
(Denition SSCV [131]). The existence of such scalars is equivalent to the linear system LS(A;b) being
consistent, where Ais the matrix whose columns are the vectors from S(Theorem SLSLC [112]).
2
6641 1 0 1
2 0 1 1
0 3 0 2
2 1 2 03
775RREF !2
666410 0 0 3 =2
010 0 1
0 0 10 3=2
0 0 0 1 1=23
7775
So by Theorem RCLS [58] the system is consistent, which indicates that bis in the subspace W.
C20 Contributed by Robert Beezer Statement [345]
The question is if pcan be written as a linear combination of the vectors in W. To check this, we set p
equal to a linear combination and massage with the denitions of vector addition and scalar multiplication
that we get with P3(Example VSP [319])
p(x) =a1(x3+x2+x) +a2(x3+ 2x 6) +a3(x2 5)
x3+ 6x+ 4 = (a1+a2)x3+ (a1+a3)x2+ (a1+ 2a2)x+ ( 6a2 5a3)
Equating coecients of equal powers of x, we get the system of equations,
a1+a2= 1
a1+a3= 0
Version 2.30
350 Section S Subspaces
a1+ 2a2= 6
6a2 5a3= 4
The augmented matrix of this system of equations row-reduces to
2
666410 0 0
010 0
0 0 10
0 0 0 13
7775
There is a leading 1 in the last column, so Theorem RCLS [58] implies that the system is inconsistent. So
there is no way for pto gain membership in W, sop62W.
C21 Contributed by Robert Beezer Statement [345]
In order to belong to W, we must be able to express Cas a linear combination of the elements in the
spanning set of W. So we begin with such an expression, using the unknowns a; b; c for the scalars in the
linear combination.
C= 3 3
6 4
=a2 1
3 1
+b4 0
2 3
+c 3 1
2 1
Massaging the right-hand side, according to the denition of the vector space operations in M22(Example
VSM [319]), we nd the matrix equality,
3 3
6 4
=2a+ 4b 3c a +c
3a+ 2b+ 2c a+ 3b+c
Matrix equality allows us to form a system of four equations in three variables, whose augmented matrix
row-reduces as follows,2
6642 4 3 3
1 0 1 3
3 2 2 6
1 3 1 43
775RREF !2
66410 0 2
010 1
0 0 1 1
0 0 0 03
775
Since this system of equations is consistent (Theorem RCLS [58]), a solution will provide values for a; b
andcthat allow us to recognize Cas an element of W.
M20 Contributed by Robert Beezer Statement [346]
The membership criteria for Zis a single linear equation, which comprises a homogeneous system of
equations. As such, we can recognize Zas the solutions to this system, and therefore Zis a null space.
Specically, Z=N
4 1 5
. Every null space is a subspace by Theorem NSMS [337].
A less direct solution appeals to Theorem TSS [334].
First, we want to be certain Zis non-empty. The zero vector of C3,0=2
40
0
03
5, is a good candidate,
since if it fails to be in Z, we will know that Zisnota vector space. Check that
4(0) (0) + 5(0) = 0
so that 02Z.
Suppose x=2
4x1
x2
x33
5andy=2
4y1
y2
y33
5are vectors from Z. Then we know that these vectors cannot be
totally arbitrary, they must have gained membership in Zby virtue of meeting the membership test. For
Version 2.30
Subsection S.SOL Solutions 351
example, we know that xmust satisfy 4 x1 x2+ 5x3= 0 while ymust satisfy 4 y1 y2+ 5y3= 0. Our
second criteria asks the question, is x+y2Z? Notice rst that
x+y=2
4x1
x2
x33
5+2
4y1
y2
y33
5=2
4x1+y1
x2+y2
x3+y33
5
and we can test this vector for membership in Zas follows,
4(x1+y1) 1(x2+y2) + 5(x3+y3)
= 4x1+ 4y1 x2 y2+ 5x3+ 5y3
= (4x1 x2+ 5x3) + (4y1 y2+ 5y3)
= 0 + 0 x2Z;y2Z
= 0
and by this computation we see that x+y2Z.
Ifis a scalar and x2Z, is it always true that x2Z? To check our third criteria, we examine
x=2
4x1
x2
x33
5=2
4x1
x2
x33
5
and we can test this vector for membership in Zwith
4(x1) (x2) + 5(x3)
=(4x1 x2+ 5x3)
=0 x2Z
= 0
and we see that indeed x2Z. With the three conditions of Theorem TSS [334] fullled, we can conclude
thatZis a subspace of C3.
T20 Contributed by Robert Beezer Statement [346]
Apply Theorem TSS [334].
First, the zero vector of Mnnis the zero matrix, O, whose entries are all zero (Denition ZM [210]).
This matrix then meets the condition that [ O]ij= 0 fori>j and so is an element of UTn.
SupposeA;B2UTn. IsA+B2UTn? We examine the entries of A+B\below" the diagonal. That
is, in the following, assume that i>j .
[A+B]ij= [A]ij+ [B]ij Denition MA [207]
= 0 + 0 A;B2UTn
= 0
which qualies A+Bfor membership in UTn.
Suppose2CandA2UTn. IsA2UTn? We examine the entries of A\below" the diagonal.
That is, in the following, assume that i>j .
[A]ij=[A]ij Denition MSM [208]
=0 A2UTn
= 0
which qualies Afor membership in UTn.
Version 2.30
352 Section S Subspaces
Having fullled the three conditions of Theorem TSS [334] we see that UTnis a subspace of Mnn.
T30 Contributed by Chris Black Statement [346]
Proof: LetEbe the subset of Pcomprised of all polynomials with all terms of even degree. Clearly the
setEis non-empty, as z(x) = 0 is a polynomial of even degree. Let p(x) andq(x) be arbitrary elements
ofE. Then there exist nonnegative integers mandnso that
p(x) =a0+a2x2+a4x4++a2nx2n
q(x) =b0+b2x2+b4x4++b2mx2m
for some constants a0;a2;:::;a 2nandb0;b2;:::;b 2m. Without loss of generality, we can assume that mn.
Thus, we have
p(x) +q(x) = (a0+b0) + (a2+b2)x2++ (a2m+b2m)x2m+a2m+2x2m+2++a2nx2n
sop(x) +q(x) has all even terms, and thus p(x) +q(x)2E. Similarly, let be a scalar. Then
p(x) =(a0+a2x2+a4x4++a2nx2n)
=a0+ (a2)x2+ (a4)x4++ (a2n)x2n
so thatp(x) also has only terms of even degree, and p(x)2E. Thus,Eis a subspace of P.
T31 Contributed by Chris Black Statement [346]
This conjecture is false. We know that the zero vector in Pis the polynomial z(x) = 0, which does not
have odd degree. Thus, the set Fdoes not contain the zero vector, and cannot be a vector space.
Version 2.30
Section LISS Linear Independence and Spanning Sets 353
Section LISS
Linear Independence and Spanning Sets
A vector space is dened as a set with two operations, meeting ten properties (Denition VS [317]). Just
as the denition of span of a set of vectors only required knowing how to add vectors and how to multiply
vectors by scalars, so it is with linear independence. A denition of a linear independent set of vectors in
an arbitrary vector space only requires knowing how to form linear combinations and equating these with
the zero vector. Since every vector space must have a zero vector (Property Z [318]), we always have a
zero vector at our disposal.
In this section we will also put a twist on the notion of the span of a set of vectors. Rather than
beginning with a set of vectors and creating a subspace that is the span, we will instead begin with a
subspace and look for a set of vectors whose span equals the subspace.
The combination of linear independence and spanning will be very important going forward.
Subsection LI
Linear Independence
Our previous denition of linear independence (Denition LI [351]) employed a relation of linear dependence
that was a linear combination on one side of an equality and a zero vector on the other side. As a
linear combination in a vector space (Denition LC [338]) depends only on vector addition and scalar
multiplication, and every vector space must have a zero vector (Property Z [318]), we can extend our
denition of linear independence from the setting of Cmto the setting of a general vector space Vwith
almost no changes. Compare these next two denitions with Denition RLDCV [153] and Denition LICV
[153].
Denition RLD
Relation of Linear Dependence
Suppose that Vis a vector space. Given a set of vectors S=fu1;u2;u3; :::; ung, an equation of the
form
1u1+2u2+3u3++nun=0
is arelation of linear dependence onS. If this equation is formed in a trivial fashion, i.e. i= 0,
1in, then we say it is a trivial relation of linear dependence onS. 4
Denition LI
Linear Independence
Suppose that Vis a vector space. The set of vectors S=fu1;u2;u3; :::; ungfromVislinearly
dependent if there is a relation of linear dependence on Sthat is not trivial. In the case where the only
relation of linear dependence on Sis the trivial one, then Sis alinearly independent set of vectors.4
Notice the emphasis on the word \only." This might remind you of the denition of a nonsingular
matrix, where if the matrix is employed as the coecient matrix of a homogeneous system then the only
solution is the trivial one.
Example LIP4
Linear independence in P4
In the vector space of polynomials with degree 4 or less, P4(Example VSP [319]) consider the set
S=
2x4+ 3x3+ 2x2 x+ 10; x4 2x3+x2+ 5x 8;2x4+x3+ 10x2+ 17x 2
:
Version 2.30
354 Section LISS Linear Independence and Spanning Sets
Is this set of vectors linearly independent or dependent? Consider that
3
2x4+ 3x3+ 2x2 x+ 10
+ 4
x4 2x3+x2+ 5x 8
+ ( 1)
2x4+x3+ 10x2+ 17x 2
= 0x4+ 0x3+ 0x2+ 0x+ 0 = 0
This is a nontrivial relation of linear dependence (Denition RLD [351]) on the set Sand so convinces us
thatSis linearly dependent (Denition LI [351]).
Now, I hear you say, \Where did those scalars come from?" Do not worry about that right now, just
be sure you understand why the above explanation is sucient to prove that Sis linearly dependent. The
remainder of the example will demonstrate how we might nd these scalars if they had not been provided
so readily. Let's look at another set of vectors (polynomials) from P4. Let
T=
3x4 2x3+ 4x2+ 6x 1; 3x4+ 1x3+ 0x2+ 4x+ 2;
4x4+ 5x3 2x2+ 3x+ 1;2x4 7x3+ 4x2+ 2x+ 1
Suppose we have a relation of linear dependence on this set,
0= 0x4+ 0x3+ 0x2+ 0x+ 0
=1
3x4 2x3+ 4x2+ 6x 1
+2
3x4+ 1x3+ 0x2+ 4x+ 2
+3
4x4+ 5x3 2x2+ 3x+ 1
+4
2x4 7x3+ 4x2+ 2x+ 1
Using our denitions of vector addition and scalar multiplication in P4(Example VSP [319]), we arrive at,
0x4+ 0x3+ 0x2+ 0x+ 0 = (31 32+ 43+ 24)x4+ ( 21+2+ 53 74)x3
+ (41+ 23+ 44)x2+ (61+ 42+ 33+ 24)x
+ ( 1+ 22+3+4):
Equating coecients, we arrive at the homogeneous system of equations,
31 32+ 43+ 24= 0
21+2+ 53 74= 0
41+ 23+ 44= 0
61+ 42+ 33+ 24= 0
1+ 22+3+4= 0
We form the coecient matrix of this homogeneous system of equations and row-reduce to nd
2
66666410 0 0
010 0
0 0 10
0 0 0 1
0 0 0 03
777775
We expected the system to be consistent (Theorem HSC [71]) and so can compute n r= 4 4 = 0 and
Theorem CSRN [59] tells us that the solution is unique. Since this is a homogeneous system, this unique
solution is the trivial solution (Denition TSHSE [71]), 1= 0,2= 0,3= 0,4= 0. So by Denition
LI [351] the set Tis linearly independent.
A few observations. If we had discovered innitely many solutions, then we could have used one of
the non-trivial ones to provide a linear combination in the manner we used to show that Swas linearly
dependent. It is important to realize that it is not interesting that we can create a relation of linear
dependence with zero scalars | we can always do that | but that for T, this is the only way to create a
Version 2.30
Subsection LISS.LI Linear Independence 355
relation of linear dependence. It was no accident that we arrived at a homogeneous system of equations
in this example, it is related to our use of the zero vector in dening a relation of linear dependence. It is
easy to present a convincing statement that a set is linearly dependent (just exhibit a nontrivial relation of
linear dependence) but a convincing statement of linear independence requires demonstrating that there is
no relation of linear dependence other than the trivial one. Notice how we relied on theorems from Chapter
SLE [3] to provide this demonstration. Whew! There's a lot going on in this example. Spend some time
with it, we'll be waiting patiently right here when you get back.
Example LIM32
Linear independence in M32
Consider the two sets of vectors RandSfrom the vector space of all 3 2 matrices, M32(Example VSM
[319])
R=8
<
:2
43 1
1 4
6 63
5;2
4 2 3
1 3
2 63
5;2
46 6
1 0
7 93
5;2
47 9
4 5
2 53
59
=
;
S=8
<
:2
42 0
1 1
1 33
5;2
4 4 0
2 2
2 63
5;2
41 1
2 1
2 43
5;2
4 5 3
10 7
2 03
59
=
;
One set is linearly independent, the other is not. Which is which? Let's examine Rrst. Build a generic
relation of linear dependence (Denition RLD [351]),
12
43 1
1 4
6 63
5+22
4 2 3
1 3
2 63
5+32
46 6
1 0
7 93
5+42
47 9
4 5
2 53
5=0
Massaging the left-hand side with our denitions of vector addition and scalar multiplication in M32
(Example VSM [319]) we obtain,
2
431 22+ 63+ 74 11+ 32 63+ 94
11+ 12 3 44 41 32+ 54
61 22+ 73+ 24 61 62 93+ 543
5=2
40 0
0 0
0 03
5
Using our denition of matrix equality (Denition ME [207]) and equating corresponding entries we get
the homogeneous system of six equations in four variables,
31 22+ 63+ 74= 0
11+ 32 63+ 94= 0
11+ 12 3 44= 0
41 32+ 54= 0
61 22+ 73+ 24= 0
61 62 93+ 54= 0
Form the coecient matrix of this homogeneous system and row-reduce to obtain
2
6666666410 0 0
010 0
0 0 10
0 0 0 1
0 0 0 0
0 0 0 03
77777775
Version 2.30
356 Section LISS Linear Independence and Spanning Sets
Analyzing this matrix we are led to conclude that 1= 0,2= 0,3= 0,4= 0. This means there is
only a trivial relation of linear dependence on the vectors of Rand so we call Ra linearly independent set
(Denition LI [351]).
So it must be that Sis linearly dependent. Let's see if we can nd a non-trivial relation of linear
dependence on S. We will begin as with R, by constructing a relation of linear dependence (Denition
RLD [351]) with unknown scalars,
12
42 0
1 1
1 33
5+22
4 4 0
2 2
2 63
5+32
41 1
2 1
2 43
5+42
4 5 3
10 7
2 03
5=0
Massaging the left-hand side with our denitions of vector addition and scalar multiplication in M32
(Example VSM [319]) we obtain,
2
421 42+3 543+ 34
1 22 23 104 1+ 22+3+ 74
1 22+ 23+ 24 31 62+ 433
5=2
40 0
0 0
0 03
5
Using our denition of matrix equality (Denition ME [207]) and equating corresponding entries we get
the homogeneous system of six equations in four variables,
21 42+3 54= 0
+3+ 34= 0
1 22 23 104= 0
1+ 22+3+ 74= 0
1 22+ 23+ 24= 0
31 62+ 43= 0
Form the coecient matrix of this homogeneous system and row-reduce to obtain
2
66666641 2 0 4
0 0 1 3
0 0 0 0
0 0 0 0
0 0 0 0
0 0 0 03
7777775
Analyzing this we see that the system is consistent (we expected this since the system is homogeneous,
Theorem HSC [71]) and has n r= 4 2 = 2 free variables, namely 2and4. This means there are
innitely many solutions, and in particular, we can nd a non-trivial solution, so long as we do not pick all
of our free variables to be zero. The mere presence of a nontrivial solution for these scalars is enough to
conclude that Sis a linearly dependent set (Denition LI [351]). But let's go ahead and explicitly construct
a non-trivial relation of linear dependence.
Choose2= 1 and4= 1. There is nothing special about this choice, there are innitely many
possibilities, some \easier" than this one, just avoid picking both variables to be zero. Then we nd the
corresponding dependent variables to be 1= 2 and3= 3. So the relation of linear dependence,
( 2)2
42 0
1 1
1 33
5+ (1)2
4 4 0
2 2
2 63
5+ (3)2
41 1
2 1
2 43
5+ ( 1)2
4 5 3
10 7
2 03
5=2
40 0
0 0
0 03
5
Version 2.30
Subsection LISS.SS Spanning Sets 357
is an iron-clad demonstration that Sis linearly dependent. Can you construct another such demonstration?
Example LIC
Linearly independent set in the crazy vector space
Is the setR=f(1;0);(6;3)glinearly independent in the crazy vector space C(Example CVS [322])? We
begin with an arbitrary relation of linear independence on R
0=a1(1;0) +a2(6;3) Denition RLD [351]
and then massage it to a point where we can apply the denition of equality in C. Recall the denitions
of vector addition and scalar multiplication in Care not what you would expect.
( 1; 1) =0 Example CVS [322]
=a1(1;0) +a2(6;3) Denition RLD [351]
= (1a1+a1 1;0a1+a1 1) + (6a2+a2 1;3a2+a2 1) Example CVS [322]
= (2a1 1; a1 1) + (7a2 1;4a2 1)
= (2a1 1 + 7a2 1 + 1; a1 1 + 4a2 1 + 1) Example CVS [322]
= (2a1+ 7a2 1; a1+ 4a2 1)
Equality in C(Example CVS [322]) then yields the two equations,
2a1+ 7a2 1 = 1
a1+ 4a2 1 = 1
which becomes the homogeneous system
2a1+ 7a2= 0
a1+ 4a2= 0
Since the coecient matrix of this system is nonsingular (check this!) the system has only the trivial
solutiona1=a2= 0. By Denition LI [351] the set Ris linearly independent. Notice that even though the
zero vector of Cis not what we might rst suspected, a question about linear independence still concludes
with a question about a homogeneous system of equations. Hmmm.
Subsection SS
Spanning Sets
In a vector space V, suppose we are given a set of vectors SV. Then we can immediately construct a
subspace,hSi, using Denition SS [339] and then be assured by Theorem SSS [339] that the construction
does provide a subspace. We now turn the situation upside-down. Suppose we are rst given a subspace
WV. Can we nd a set Sso thathSi=W? Typically Wis innite and we are searching for a nite
set of vectors Sthat we can combine in linear combinations and \build" all of W.
I like to think of Sas the raw materials that are sucient for the construction of W. If you have
nails, lumber, wire, copper pipe, drywall, plywood, carpet, shingles, paint (and a few other things), then
you can combine them in many dierent ways to create a house (or innitely many dierent houses for
that matter). A fast-food restaurant may have beef, chicken, beans, cheese, tortillas, taco shells and hot
sauce and from this small list of ingredients build a wide variety of items for sale. Or maybe a better
Version 2.30
358 Section LISS Linear Independence and Spanning Sets
analogy comes from Ben Cordes | the additive primary colors (red, green and blue) can be combined to
create many dierent colors by varying the intensity of each. The intensity is like a scalar multiple, and
the combination of the three intensities is like vector addition. The three individual colors, red, green and
blue, are the elements of the spanning set.
Because we will use terms like \spanned by" and \spanning set," there is the potential for confusion
with \the span." Come back and reread the rst paragraph of this subsection whenever you are uncertain
about the dierence. Here's the working denition.
Denition TSVS
To Span a Vector Space
SupposeVis a vector space. A subset SofVis aspanning set forVifhSi=V. In this case, we also
saySspansV. 4
The denition of a spanning set requires that two sets (subspaces actually) be equal. If Sis a subset of
V, thenhSiV, always. Thus it is usually only necessary to prove that VhSi. Now would be a good
time to review Denition SE [762].
Example SSP4
Spanning set in P4
In Example SP4 [335] we showed that
W=fp(x)jp2P4; p(2) = 0g
is a subspace of P4, the vector space of polynomials with degree at most 4 (Example VSP [319]). In this
example, we will show that the set
S=
x 2; x2 4x+ 4; x3 6x2+ 12x 8; x4 8x3+ 24x2 32x+ 16
is a spanning set for W. To do this, we require that W=hSi. This is an equality of sets. We can check
that every polynomial in Shasx= 2 as a root and therefore SW. SinceWis closed under addition
and scalar multiplication, hSiWalso.
So it remains to show that WhSi(Denition SE [762]). To do this, begin by choosing an arbitrary
polynomial in W, sayr(x) =ax4+bx3+cx2+dx+e2W. This polynomial is not as arbitrary as it would
appear, since we also know it must have x= 2 as a root. This translates to
0 =a(2)4+b(2)3+c(2)2+d(2) +e= 16a+ 8b+ 4c+ 2d+e
as a condition on r.
We wish to show that ris a polynomial in hSi, that is, we want to show that rcan be written as a
linear combination of the vectors (polynomials) in S. So let's try.
r(x) =ax4+bx3+cx2+dx+e
=1(x 2) +2
x2 4x+ 4
+3
x3 6x2+ 12x 8
+4
x4 8x3+ 24x2 32x+ 16
=4x4+ (3 84)x3+ (2 63+ 244)x2
+ (1 42+ 123 324)x+ ( 21+ 42 83+ 164)
Equating coecients (vector equality in P4) gives the system of ve equations in four variables,
4=a
3 84=b
2 63+ 244=c
1 42+ 123 324=d
Version 2.30
Subsection LISS.SS Spanning Sets 359
21+ 42 83+ 164=e
Any solution to this system of equations will provide the linear combination we need to determine if r2hSi,
but we need to be convinced there is a solution for any values of a; b; c; d; e that qualify rto be a member
ofW. So the question is: is this system of equations consistent? We will form the augmented matrix, and
row-reduce. (We probably need to do this by hand, since the matrix is symbolic | reversing the order of
the rst four rows is the best way to start). We obtain a matrix in reduced row-echelon form
2
66666410 0 0 32 a+ 12b+ 4c+d
010 0 24 a+ 6b+c
0 0 10 8 a+b
0 0 0 1 a
0 0 0 0 16 a+ 8b+ 4c+ 2d+e3
777775=2
66666410 0 0 32 a+ 12b+ 4c+d
010 0 24 a+ 6b+c
0 0 10 8 a+b
0 0 0 1 a
0 0 0 0 03
777775
For your results to match our rst matrix, you may nd it necessary to multiply the nal row of your
row-reduced matrix by the appropriate scalar, and/or add multiples of this row to some of the other rows.
To obtain the second version of the matrix, the last entry of the last column has been simplied to zero
according to the one condition we were able to impose on an arbitrary polynomial from W. So with
no leading 1's in the last column, Theorem RCLS [58] tells us this system is consistent. Therefore, any
polynomial from Wcan be written as a linear combination of the polynomials in S, soWhSi. Therefore,
W=hSiandSis a spanning set for Wby Denition TSVS [356].
Notice that an alternative to row-reducing the augmented matrix by hand would be to appeal to
Theorem FS [299] by expressing the column space of the coecient matrix as a null space, and then
verifying that the condition on rguarantees that ris in the column space, thus implying that the system
is always consistent. Give it a try, we'll wait. This has been a complicated example, but worth studying
carefully.
Given a subspace and a set of vectors, as in Example SSP4 [356] it can take some work to determine
that the set actually is a spanning set. An even harder problem is to be confronted with a subspace and
required to construct a spanning set with no guidance. We will now work an example of this
avor, but
some of the steps will be unmotivated. Fortunately, we will have some better tools for this type of problem
later on.
Example SSM22
Spanning set in M22
In the space of all 2 2 matrices, M22consider the subspace
Z=a b
c da+ 3b c 5d= 0; 2a 6b+ 3c+ 14d= 0
and nd a spanning set for Z.
We need to construct a limited number of matrices in Zso that every matrix in Zcan be expressed as
a linear combination of this limited number of matrices. Suppose that B=a b
c d
is a matrix in Z. Then
we can form a column vector with the entries of Band write
2
664a
b
c
d3
7752N1 3 1 5
2 6 3 14
Version 2.30
360 Section LISS Linear Independence and Spanning Sets
Row-reducing this matrix and applying Theorem REMES [31] we obtain the equivalent statement,
2
664a
b
c
d3
7752N13 0 1
0 0 1 4
We can then express the subspace Zin the following equal forms,
Z=a b
c da+ 3b c 5d= 0; 2a 6b+ 3c+ 14d= 0
=a b
c da+ 3b d= 0; c+ 4d= 0
=a b
c da= 3b+d; c= 4d
= 3b+d b
4d db; d2C
= 3b b
0 0
+d0
4d db; d2C
=
b 3 1
0 0
+d1 0
4 1b; d2C
= 3 1
0 0
;1 0
4 1
So the set
Q= 3 1
0 0
;1 0
4 1
spansZby Denition TSVS [356].
Example SSC
Spanning set in the crazy vector space
In Example LIC [355] we determined that the set R=f(1;0);(6;3)gis linearly independent in the crazy
vector space C(Example CVS [322]). We now show that Ris a spanning set for C.
Given an arbitrary vector ( x; y)2Cwe desire to show that it can be written as a linear combination
of the elements of R. In other words, are there scalars a1anda2so that
(x; y) =a1(1;0) +a2(6;3)
We will act as if this equation is true and try to determine just what a1anda2would be (as functions of
xandy).
(x; y) =a1(1;0) +a2(6;3)
= (1a1+a1 1;0a1+a1 1) + (6a2+a2 1;3a2+a2 1) Scalar mult in C
= (2a1 1; a1 1) + (7a2 1;4a2 1)
= (2a1 1 + 7a2 1 + 1; a1 1 + 4a2 1 + 1) Addition in C
= (2a1+ 7a2 1; a1+ 4a2 1)
Equality in Cthen yields the two equations,
2a1+ 7a2 1 =x
Version 2.30
Subsection LISS.VR Vector Representation 361
a1+ 4a2 1 =y
which becomes the linear system with a matrix representation
2 7
1 4a1
a2
=x+ 1
y+ 1
The coecient matrix of this system is nonsingular, hence invertible (Theorem NI [261]), and we can
employ its inverse to nd a solution (Theorem TTMI [246], Theorem SNCM [261]),
a1
a2
=2 7
1 4 1x+ 1
y+ 1
=4 7
1 2x+ 1
y+ 1
=4x 7y 3
x+ 2y+ 1
We could chase through the above implications backwards and take the existence of these solutions as
sucient evidence for Rbeing a spanning set for C. Instead, let us view the above as simply scratchwork
and now get serious with a simple direct proof that Ris a spanning set. Ready? Suppose ( x; y) is any
vector from C, then compute the following linear combination using the denitions of the operations in C,
(4x 7y 3)(1;0) + ( x+ 2y+ 1)(6;3)
= (1(4x 7y 3) + (4x 7y 3) 1;0(4x 7y 3) + (4x 7y 3) 1) +
(6( x+ 2y+ 1) + ( x+ 2y+ 1) 1;3( x+ 2y+ 1) + ( x+ 2y+ 1) 1)
= (8x 14y 7;4x 7y 4) + ( 7x+ 14y+ 6; 4x+ 8y+ 3)
= ((8x 14y 7) + ( 7x+ 14y+ 6) + 1;(4x 7y 4) + ( 4x+ 8y+ 3) + 1)
= (x; y)
This nal sequence of computations in Cis sucient to demonstrate that any element of Ccanbe written
(or expressed) as a linear combination of the two vectors in R, soChRi. Since the reverse inclusion
hRiCis trivially true, C=hRiand we say RspansC(Denition TSVS [356]). Notice that this
demonstration is no more or less valid if we hide from the reader our scratchwork that suggested a1=
4x 7y 3 anda2= x+ 2y+ 1.
Subsection VR
Vector Representation
In Chapter R [603] we will take up the matter of representations fully, where Theorem VRRB [360] will
be critical for Denition VR [603]. We will now motivate and prove a critical theorem that tells us how
to \represent" a vector. This theorem could wait, but working with it now will provide some extra insight
into the nature of linearly independent spanning sets. First an example, then the theorem.
Example AVR
A vector representation
Consider the set
S=8
<
:2
4 7
5
13
5;2
4 6
5
03
5;2
4 12
7
43
59
=
;
from the vector space C3. LetAbe the matrix whose columns are the set S, and verify that Ais nonsingular.
By Theorem NMLIC [159] the elements of Sform a linearly independent set. Suppose that b2C3. Then
LS(A;b) has a (unique) solution (Theorem NMUS [86]) and hence is consistent. By Theorem SLSLC
[112], b2hSi. Since bis arbitrary, this is enough to show that hSi=C3, and therefore Sis a spanning set
Version 2.30
362 Section LISS Linear Independence and Spanning Sets
forC3(Denition TSVS [356]). (This set comes from the columns of the coecient matrix of Archetype
B [786].)
Now examine the situation for a particular choice of b, sayb=2
4 33
24
53
5. BecauseSis a spanning set
forC3, we know we can write bas a linear combination of the vectors in S,
2
4 33
24
53
5= ( 3)2
4 7
5
13
5+ (5)2
4 6
5
03
5+ (2)2
4 12
7
43
5:
The nonsingularity of the matrix Atells that the scalars in this linear combination are unique. More
precisely, it is the linear independence of Sthat provides the uniqueness. We will refer to the scalars
a1= 3,a2= 5,a3= 2 as a \representation of brelative toS." In other words, once we settle on Sas
a linearly independent set that spans C3, the vector bis recoverable just by knowing the scalars a1= 3,
a2= 5,a3= 2 (use these scalars in a linear combination of the vectors in S). This is all an illustration of
the following important theorem, which we prove in the setting of a general vector space.
Theorem VRRB
Vector Representation Relative to a Basis
Suppose that Vis a vector space and B=fv1;v2;v3; :::; vmgis a linearly independent set that spans
V. Let wbe any vector in V. Then there exist unique scalarsa1; a2; a3; :::; amsuch that
w=a1v1+a2v2+a3v3++amvm:
Proof That wcan be written as a linear combination of the vectors in Bfollows from the spanning
property of the set (Denition TSVS [356]). This is good, but not the meat of this theorem. We now know
that for any choice of the vector wthere exist some scalars that will create was a linear combination of
the basis vectors. The real question is: Is there more than one way to write was a linear combination of
fv1;v2;v3; :::; vmg? Are the scalars a1; a2; a3; :::; amunique? (Technique U [771])
Assume there are two ways to express was a linear combination of fv1;v2;v3; :::; vmg. In other
words there exist scalars a1; a2; a3; :::; amandb1; b2; b3; :::; bmso that
w=a1v1+a2v2+a3v3++amvm
w=b1v1+b2v2+b3v3++bmvm:
Then notice that
0=w+ ( w) Property AI [318]
=w+ ( 1)w Theorem AISM [325]
= (a1v1+a2v2+a3v3++amvm)+
( 1)(b1v1+b2v2+b3v3++bmvm)
= (a1v1+a2v2+a3v3++amvm)+
( b1v1 b2v2 b3v3 ::: bmvm) Property DVA [318]
= (a1 b1)v1+ (a2 b2)v2+ (a3 b3)v3+
+ (am bm)vm Property C [317], Property DSA [318]
But this is a relation of linear dependence on a linearly independent set of vectors (Denition RLD [351])!
Now we are using the other assumption about B, thatfv1;v2;v3; :::; vmgis a linearly independent set.
So by Denition LI [351] it must happen that the scalars are all zero. That is,
(a1 b1) = 0 ( a2 b2) = 0 ( a3 b3) = 0 ::: (am bm) = 0
Version 2.30
Subsection LISS.READ Reading Questions 363
a1=b1 a2=b2 a3=b3::: a m=bm:
And so we nd that the scalars are unique.
This is a very typical use of the hypothesis that a set is linearly independent | obtain a relation of
linear dependence and then conclude that the scalars must all be zero. The result of this theorem tells
us that we can write any vector in a vector space as a linear combination of the vectors in a linearly
independent spanning set, but only just. There is only enough raw material in the spanning set to write
each vector one way as a linear combination. So in this sense, we could call a linearly independent spanning
set a \minimal spanning set." These sets are so important that we will give them a simpler name (\basis")
and explore their properties further in the next section.
Subsection READ
Reading Questions
1. Is the set of matrices below linearly independent or linearly dependent in the vector space M22? Why
or why not?1 3
2 4
; 2 3
3 5
;0 9
1 3
2. Explain the dierence between the following two uses of the term \span":
(a)Sis a subset of the vector space Vand the span of Sis a subspace of V.
(b)Wis subspace of the vector space YandTspansW.
3. The set
S=8
<
:2
46
2
13
5;2
44
3
13
5;2
45
8
23
59
=
;
is linearly independent and spans C3. Write the vector x=2
4 6
2
23
5a linear combination of the elements
ofS. How many ways are there to answer this question, and which theorem allows you to say so?
Version 2.30
364 Section LISS Linear Independence and Spanning Sets
Subsection EXC
Exercises
C20 In the vector space of 2 2 matrices, M22, determine if the set Sbelow is linearly independent.
S=2 1
1 3
;0 4
1 2
;4 2
1 3
Contributed by Robert Beezer Solution [364]
C21 In the crazy vector space C(Example CVS [322]), is the set S=f(0;2);(2;8)glinearly indepen-
dent?
Contributed by Robert Beezer Solution [364]
C22 In the vector space of polynomials P3, determine if the set Sis linearly independent or linearly
dependent.
S=
2 +x 3x2 8x3;1 +x+x2+ 5x3;3 4x2 7x3
Contributed by Robert Beezer Solution [365]
C23 Determine if the set S=f(3;1);(7;3)gis linearly independent in the crazy vector space C(Example
CVS [322]).
Contributed by Robert Beezer Solution [365]
C24 In the vector space of real-valued functions F=ffjf:R!Rg, determine if the following set Sis
linearly independent.
S=
sin2x;cos2x;2
Contributed by Chris Black Solution [365]
C25 Let
S=1 2
2 1
;2 1
1 2
;0 1
1 2
1. Determine if SspansM2;2.
2. Determine if Sis linearly independent.
Contributed by Chris Black Solution [365]
C26 Let
S=1 2
2 1
;2 1
1 2
;0 1
1 2
;1 0
1 1
;1 4
0 3
1. Determine if SspansM2;2.
2. Determine if Sis linearly independent.
Contributed by Chris Black Solution [366]
C30 In Example LIM32 [353], nd another nontrivial relation of linear dependence on the linearly de-
pendent set of 32 matrices, S.
Contributed by Robert Beezer
Version 2.30
Subsection LISS.EXC Exercises 365
C40 Determine if the set T=
x2 x+ 5;4x3 x2+ 5x;3x+ 2
spans the vector space of polynomials
with degree 4 or less, P4.
Contributed by Robert Beezer Solution [367]
C41 The setWis a subspace of M22, the vector space of all 2 2 matrices. Prove that Sis a spanning
set forW.
W=a b
c d2a 3b+ 4c d= 0
S=1 0
0 2
;0 1
0 3
;0 0
1 4
Contributed by Robert Beezer Solution [367]
C42 Determine if the set S=f(3;1);(7;3)gspans the crazy vector space C(Example CVS [322]).
Contributed by Robert Beezer Solution [368]
M10 Halfway through Example SSP4 [356], we need to show that the system of equations
LS0
BBBB@2
666640 0 0 1
0 0 1 8
0 1 6 24
1 4 12 32
2 4 8 163
77775;2
66664a
b
c
d
e3
777751
CCCCA
is consistent for every choice of the vector of constants satisfying 16 a+ 8b+ 4c+ 2d+e= 0.
Express the column space of the coecient matrix of this system as a null space, using Theorem FS
[299]. From this use Theorem CSCS [272] to establish that the system is always consistent. Notice that
this approach removes from Example SSP4 [356] the need to row-reduce a symbolic matrix.
Contributed by Robert Beezer Solution [368]
T40 Prove the following variant of Theorem EMMVP [225] that has a weaker hypothesis: Suppose that
C=fu1;u2;u3; :::; upgis a linearly independent spanning set for Cn. Suppose also that AandBare
mnmatrices such that Aui=Buifor every 1in. ThenA=B.
Can you weaken the hypothesis even further while still preserving the conclusion?
Contributed by Robert Beezer
T50 Suppose that Vis a vector space and u;v2Vare two vectors in V. Use the denition of linear
independence to prove that S=fu;vgis a linearly dependent set if and only if one of the two vectors is
a scalar multiple of the other. Prove this directly in the context of an abstract vector space ( V), without
simply giving an upgraded version of Theorem DLDS [175] for the special case of just two vectors.
Contributed by Robert Beezer Solution [368]
Version 2.30
366 Section LISS Linear Independence and Spanning Sets
Subsection SOL
Solutions
C20 Contributed by Robert Beezer Statement [362]
Begin with a relation of linear dependence on the vectors in Sand massage it according to the denitions
of vector addition and scalar multiplication in M22,
O=a12 1
1 3
+a20 4
1 2
+a34 2
1 3
0 0
0 0
=2a1+ 4a3 a1+ 4a2+ 2a3
a1 a2+a33a1+ 2a2+ 3a3
By our denition of matrix equality (Denition ME [207]) we arrive at a homogeneous system of linear
equations,
2a1+ 4a3= 0
a1+ 4a2+ 2a3= 0
a1 a2+a3= 0
3a1+ 2a2+ 3a3= 0
The coecient matrix of this system row-reduces to the matrix,
2
66410 0
010
0 0 1
0 0 03
775
and from this we conclude that the only solution is a1=a2=a3= 0. Since the relation of linear
dependence (Denition RLD [351]) is trivial, the set Sis linearly independent (Denition LI [351]).
C21 Contributed by Robert Beezer Statement [362]
We begin with a relation of linear dependence using unknown scalars aandb. We wish to know if these
scalars must both be zero. Recall that the zero vector in Cis ( 1; 1) and that the denitions of vector
addition and scalar multiplication are not what we might expect.
0= ( 1; 1)
=a(0;2) +b(2;8) Denition RLD [351]
= (0a+a 1;2a+a 1) + (2b+b 1;8b+b 1) Scalar mult., Example CVS [322]
= (a 1;3a 1) + (3b 1;9b 1)
= (a 1 + 3b 1 + 1;3a 1 + 9b 1 + 1) Vector addition, Example CVS [322]
= (a+ 3b 1;3a+ 9b 1)
From this we obtain two equalities, which can be converted to a homogeneous system of equations,
1 =a+ 3b 1 a+ 3b= 0
1 = 3a+ 9b 1 3 a+ 9b= 0
This homogeneous system has a singular coecient matrix (Theorem SMZD [445]), and so has more than
just the trivial solution (Denition NM [83]). Any nontrivial solution will give us a nontrivial relation of
linear dependence on S. SoSis linearly dependent (Denition LI [351]).
Version 2.30
Subsection LISS.SOL Solutions 367
C22 Contributed by Robert Beezer Statement [362]
Begin with a relation of linear dependence (Denition RLD [351]),
a1
2 +x 3x2 8x3
+a2
1 +x+x2+ 5x3
+a3
3 4x2 7x3
=0
Massage according to the denitions of scalar multiplication and vector addition in the denition of P3
(Example VSP [319]) and use the zero vector dro this vector space,
(2a1+a2+ 3a3) + (a1+a2)x+ ( 3a1+a2 4a3)x2+ ( 8a1+ 5a2 7a3)x3= 0 + 0x+ 0x2+ 0x3
The denition of the equality of polynomials allows us to deduce the following four equations,
2a1+a2+ 3a3= 0
a1+a2= 0
3a1+a2 4a3= 0
8a1+ 5a2 7a3= 0
Row-reducing the coecient matrix of this homogeneous system leads to the unique solution a1=a2=
a3= 0. So the only relation of linear dependence on Sis the trivial one, and this is linear independence
forS(Denition LI [351]).
C23 Contributed by Robert Beezer Statement [362]
Notice, or discover, that the following gives a nontrivial relation of linear dependence on SinC, so by
Denition LI [351], the set Sis linearly dependent.
2(3;1) + ( 1)(7;3) = (7;3) + ( 9; 5) = ( 1; 1) =0
C24 Contributed by Chris Black Statement [362]
One of the fundamental identities of trigonometry is sin2(x) + cos2(x) = 1. Thus, we have a dependence
relation 2(sin2x) + 2(cos2x) + ( 1)(2) = 0, and the set is linearly dependent.
C25 Contributed by Chris Black Statement [362]
1. IfSspansM2;2, then for every 2 2 matrixB=x y
z w
, there exist constants ;;
so that
x y
z w
=1 2
2 1
+2 1
1 2
+
0 1
1 2
Applying Denition ME [207], this leads to the linear system
+ 2=x
2++
=y
2 +
=z
+ 2+ 2
=w:
We need to row-reduce the augmented matrix of this system by hand due to the symbols x,y,z, and
win the vector of constants.
2
6641 2 0 x
2 1 1 y
2 1 1z
1 2 2 w3
775RREF !2
66410 0 x y+z
0101
2(y z)
0 0 11
2(w x)
0 0 01
2(5y 3x 3z w)3
775
Version 2.30
368 Section LISS Linear Independence and Spanning Sets
With the apperance of a leading 1 possible in the last column, by Theorem RCLS [58] there will
exist some matrices B=x y
z w
so that the linear system above has no solution (namely, whenever
5y 3x 3z w6= 0), so the set Sdoes not span M2;2. (For example, you can verify that there is
no solution when B=3 3
3 2
.)
2. To check for linear independence, we need to see if there are nontrivial coecients ;;
that solve
0 0
0 0
=1 2
2 1
+2 1
1 2
+
0 1
1 2
This requires the same work that was done in part (a), with x=y=z=w= 0. In that case, the
coecient matrix row-reduces to have a leading 1 in each of the rst three columns and a row of zeros
on the bottom, so we know that the only solution to the matrix equation is ==
= 0. So the
setSis linearly independent.
C26 Contributed by Chris Black Statement [362]
1. The matrices in Swill spanM2;2if for anyx y
z w
, there are coecients a;b;c;d;e so that
a1 2
2 1
+b2 1
1 2
+c0 1
1 2
+d1 0
1 1
+e1 4
0 3
=x y
z w
Thus, we have
a+ 2b+d+e 2a+b+c+ 4e
2a b+c+d a + 2b+ 2c+d+ 3e
=x y
z w
so we have the matrix equation
2
6641 2 0 1 1
2 1 1 0 4
2 1 1 1 0
1 2 2 1 33
7752
66664a
b
c
d
e3
77775=2
664x
y
z
w3
775
This system will have a solution for every vector on the right side if the row-reduced coecient matrix
has a leading one in every row, since then it is never possible to have a leading 1 appear in the nal
column of a row-reduced augmented matrix.
2
6641 2 0 1 1
2 1 1 0 4
2 1 1 1 0
1 2 2 1 33
775RREF !2
666410 0 0 1
010 0 1
0 0 10 1
0 0 0 1 23
7775
Since there is a leading one in each row of the row-reduced coecient matrix, there is a solution for
every vector2
664x
y
z
w3
775, which means that there is a solution to the original equation for every matrix
x y
z w
. Thus, the original four matrices span M2;2.
Version 2.30
Subsection LISS.SOL Solutions 369
2. The matrices in Sare linearly independent if the only solution to
a1 2
2 1
+b2 1
1 2
+c0 1
1 2
+d1 0
1 1
+e1 4
0 3
=0 0
0 0
isa=b=c=d=e= 0.
We have
a+ 2b+d+e 2a+b+c+ 4e
2a b+c+d a + 2b+ 2c+d+ 3e
=2
6641 2 0 1 1
2 1 1 0 4
2 1 1 1 0
1 2 2 1 33
7752
66664a
b
c
d
e3
77775=0 0
0 0
so we need to nd the nullspace of the matrix
2
6641 2 0 1 1
2 1 1 0 4
2 1 1 1 0
1 2 2 1 33
775
We row-reduced this matrix in part (a), and found that there is a column without a leading 1, which
correspons to a free variable in a description of the solution set to the homogeneous system, so the
nullspace is nontrivial and there are an innite number of solutions to
a1 2
2 1
+b2 1
1 2
+c0 1
1 2
+d1 0
1 1
+e1 4
0 3
=0 0
0 0
Thus, this set of matrices is not linearly independent.
C40 Contributed by Robert Beezer Statement [363]
The polynomial x4is an element of P4. Can we write this element as a linear combination of the elements
ofT? To wit, are there scalars a1,a2,a3such that
x4=a1
x2 x+ 5
+a2
4x3 x2+ 5x
+a3(3x+ 2)
Massaging the right side of this equation, according to the denitions of Example VSP [319], and then
equating coecients, leads to an inconsistent system of equations (check this!). As such, Tis not a spanning
set forP4.
C41 Contributed by Robert Beezer Statement [363]
We want to show that W=hSi(Denition TSVS [356]), which is an equality of sets (Denition SE [762]).
First, show thathSiW. Begin by checking that each of the three matrices in Sis a member of the
setW. Then, since Wis a vector space, the closure properties (Property AC [317], Property SC [317])
guarantee that every linear combination of elements of Sremains in W.
Second, show that WhSi. We want to convince ourselves that an arbitrary element of Wis a linear
combination of elements of S. Choose
x=a b
c d
2W
The values of a; b; c; d are not totally arbitrary, since membership in Wrequires that 2 a 3b+ 4c d= 0.
Now, rewrite as follows,
x=a b
c d
Version 2.30
370 Section LISS Linear Independence and Spanning Sets
=a b
c2a 3b+ 4c
2a 3b+ 4c d= 0
=a0
0 2a
+0b
0 3b
+0 0
c4c
Denition MA [207]
=a1 0
0 2
+b0 1
0 3
+c0 0
1 4
Denition MSM [208]
2hSi Denition SS [339]
C42 Contributed by Robert Beezer Statement [363]
We will try to show that SspansC. Let (x; y) be an arbitrary element of Cand search for scalars a1and
a2such that
(x; y) =a1(3;1) +a2(7;3)
= (4a1 1;2a1 1) + (8a2 1;4a2 1)
= (4a1+ 8a2 1;2a1+ 4a2 1)
Equality in Cleads to the system
4a1+ 8a2=x+ 1
2a1+ 4a2=y+ 1
This system has a singular coecient matrix whose column space is simply2
1
. So any choice of x
andythat causes the column vectorx+ 1
y+ 1
to lie outside the column space will lead to an inconsistent
system, and hence create an element ( x; y) that is not in the span of S. SoSdoes not span C.
For example, choose x= 0 andy= 5, and then we can see that1
6
622
1
and we know that (0 ;5)
cannot be written as a linear combination of the vectors in S. A shorter solution might begin by asserting
that (0;5) is not inhSiand then establishing this claim alone.
M10 Contributed by Robert Beezer Statement [363]
Theorem FS [299] provides the matrix
L=
11
21
41
81
16
and so ifAdenotes the coecient matrix of the system, then C(A) =N(L). The single homogeneous
equation inLS(L;0) is equivalent to the condition on the vector of constants (use a; b; c; d; e as variables
and then multiply by 16).
T50 Contributed by Robert Beezer Statement [363]
()) IfSis linearly dependent, then there are scalars and, not both zero, such that u+v=0.
Suppose that 6= 0, the proof proceeds similarly if 6= 0. Now,
u= 1u Property O [318]
=1
u Property MICN [759]
=1
(u) Property SMA [318]
=1
(u+0) Property Z [318]
Version 2.30
Subsection LISS.SOL Solutions 371
=1
(u+v v) Property AI [318]
=1
(0 v) Denition LI [351]
=1
( v) Property Z [318]
=
v Property SMA [318]
which shows that uis a scalar multiple of v.
(() Suppose now that uis a scalar multiple of v. More precisely, suppose there is a scalar
such
thatu=
v. Then
( 1)u+
v= ( 1)u+u
= ( 1)u+ (1)u Property O [318]
= (( 1) + 1) u Property DSA [318]
= 0u Property AICN [759]
=0 Theorem ZSSM [324]
This is a relation of linear of linear dependence on S(Denition RLD [351]), which is nontrivial since one
of the scalars is 1. Therefore Sis linearly dependent by Denition LI [351].
Be careful using this theorem. It is only applicable to sets of two vectors. In particular, linear de-
pendence in a set of three or more vectors can be more complicated than just one vector being a scalar
multiple of another.
Version 2.30
372 Section LISS Linear Independence and Spanning Sets
Version 2.30
Section B Bases 373
Section B
Bases
A basis of a vector space is one of the most useful concepts in linear algebra. It often provides a concise,
nite description of an innite vector space.
Subsection B
Bases
We now have all the tools in place to dene a basis of a vector space.
Denition B
Basis
SupposeVis a vector space. Then a subset SVis abasis ofVif it is linearly independent and spans
V. 4
So, a basis is a linearly independent spanning set for a vector space. The requirement that the set
spansVinsures that Shas enough raw material to build V, while the linear independence requirement
insures that we do not have any more raw material than we need. As we shall see soon in Section D [391],
a basis is a minimal spanning set.
You may have noticed that we used the term basis for some of the titles of previous theorems (e.g.
Theorem BNS [160], Theorem BCS [274], Theorem BRS [280]) and if you review each of these theorems you
will see that their conclusions provide linearly independent spanning sets for sets that we now recognize
as subspaces of Cm. Examples associated with these theorems include Example NSLIL [161], Example
CSOCD [275] and Example IAS [281]. As we will see, these three theorems will continue to be powerful
tools, even in the setting of more general vector spaces.
Furthermore, the archetypes contain an abundance of bases. For each coecient matrix of a system
of equations, and for each archetype dened simply as a matrix, there is a basis for the null space, three
bases for the column space, and a basis for the row space. For this reason, our subsequent examples will
concentrate on bases for vector spaces other than Cm. Notice that Denition B [371] does not preclude
a vector space from having many bases, and this is the case, as hinted above by the statement that the
archetypes contain three bases for the column space of a matrix. More generally, we can grab any basis for
a vector space, multiply any one basis vector by a non-zero scalar and create a slightly dierent set that
is still a basis. For \important" vector spaces, it will be convenient to have a collection of \nice" bases.
When a vector space has a single particularly nice basis, it is sometimes called the standard basis though
there is nothing precise enough about this term to allow us to dene it formally | it is a question of style.
Here are some nice bases for important vector spaces.
Theorem SUVB
Standard Unit Vectors are a Basis
The set of standard unit vectors for Cm(Denition SUV [197]), B=fe1;e2;e3; :::; emg=feij1img
is a basis for the vector space Cm.
Proof We must show that the set Bis both linearly independent and a spanning set for Cm. First, the
vectors inBare, by Denition SUV [197], the columns of the identity matrix, which we know is nonsingular
(since it row-reduces to the identity matrix, Theorem NMRRI [84]). And the columns of a nonsingular
matrix are linearly independent by Theorem NMLIC [159].
Version 2.30
374 Section B Bases
Suppose we grab an arbitrary vector from Cm, say
v=2
666664v1
v2
v3
...
vm3
777775:
Can we write vas a linear combination of the vectors in B? Yes, and quite simply.
2
666664v1
v2
v3
...
vm3
777775=v12
6666641
0
0
...
03
777775+v22
6666640
1
0
...
03
777775+v32
6666640
0
1
...
03
777775++vm2
6666640
0
0
...
13
777775
v=v1e1+v2e2+v3e3++vmem
this shows that CmhBi, which is sucient to show that Bis a spanning set for Cm.
Example BP
Bases for Pn
The vector space of polynomials with degree at most n,Pn, has the basis
B=
1; x; x2; x3; :::; xn
:
Another nice basis for Pnis
C=
1;1 +x;1 +x+x2;1 +x+x2+x3; :::; 1 +x+x2+x3++xn
:
Checking that each of BandCis a linearly independent spanning set are good exercises.
Example BM
A basis for the vector space of matrices
In the vector space Mmnof matrices (Example VSM [319]) dene the matrices Bk`, 1km, 1`n
by
[Bk`]ij=(
1 ifk=i; `=j
0 otherwise
So these matrices have entries that are all zeros, with the exception of a lone entry that is one. The set of
allmnof them,
B=fBk`j1km;1`ng
forms a basis for Mmn. See Exercise B.M20 [383].
The bases described above will often be convenient ones to work with. However a basis doesn't have
to obviously look like a basis.
Example BSP4
A basis for a subspace of P4
In Example SSP4 [356] we showed that
S=
x 2; x2 4x+ 4; x3 6x2+ 12x 8; x4 8x3+ 24x2 32x+ 16
Version 2.30
Subsection B.B Bases 375
is a spanning set for W=fp(x)jp2P4; p(2) = 0g. We will now show that Sis also linearly independent
inW. Begin with a relation of linear dependence,
0 + 0x+ 0x2+ 0x3+ 0x4=1(x 2) +2
x2 4x+ 4
+3
x3 6x2+ 12x 8
+4
x4 8x3+ 24x2 32x+ 16
=4x4+ (3 84)x3+ (2 63+ 244)x2
+ (1 42+ 123 324)x+ ( 21+ 42 83+ 164)
Equating coecients (vector equality in P4) gives the homogeneous system of ve equations in four vari-
ables,
4= 0
3 84= 0
2 63+ 244= 0
1 42+ 123 324= 0
21+ 42 83+ 164= 0
We form the coecient matrix, and row-reduce to obtain a matrix in reduced row-echelon form
2
66666410 0 0
010 0
0 0 10
0 0 0 1
0 0 0 03
777775
With only the trivial solution to this homogeneous system, we conclude that only scalars that will form a
relation of linear dependence are the trivial ones, and therefore the set Sis linearly independent (Denition
LI [351]). Finally, Shas earned the right to be called a basis for W(Denition B [371]).
Example BSM22
A basis for a subspace of M22
In Example SSM22 [357] we discovered that
Q= 3 1
0 0
;1 0
4 1
is a spanning set for the subspace
Z=a b
c da+ 3b c 5d= 0; 2a 6b+ 3c+ 14d= 0
of the vector space of all 2 2 matrices, M22. If we can also determine that Qis linearly independent in
Z(or inM22), then it will qualify as a basis for Z. Let's begin with a relation of linear dependence.
0 0
0 0
=1 3 1
0 0
+21 0
4 1
= 31+21
422
Using our denition of matrix equality (Denition ME [207]) we equate corresponding entries and get a
homogeneous system of four equations in two variables,
31+2= 0
Version 2.30
376 Section B Bases
1= 0
42= 0
2= 0
We could row-reduce the coecient matrix of this homogeneous system, but it is not necessary. The second
and fourth equations tell us that 1= 0,2= 0 is the only solution to this homogeneous system. This
qualies the set Qas being linearly independent, since the only relation of linear dependence is trivial
(Denition LI [351]). Therefore Qis a basis for Z(Denition B [371]).
Example BC
Basis for the crazy vector space
In Example LIC [355] and Example SSC [358] we determined that the set R=f(1;0);(6;3)gfrom the
crazy vector space, C(Example CVS [322]), is linearly independent and is a spanning set for C. By
Denition B [371] we see that Ris a basis for C.
We have seen that several of the sets associated with a matrix are subspaces of vector spaces of column
vectors. Specically these are the null space (Theorem NSMS [337]), column space (Theorem CSMS [343]),
row space (Theorem RSMS [344]) and left null space (Theorem LNSMS [344]). As subspaces they are vector
spaces (Denition S [333]) and it is natural to ask about bases for these vector spaces. Theorem BNS [160],
Theorem BCS [274], Theorem BRS [280] each have conclusions that provide linearly independent spanning
sets for (respectively) the null space, column space, and row space. Notice that each of these theorems
contains the word \basis" in its title, even though we did not know the precise meaning of the word at
the time. To nd a basis for a left null space we can use the denition of this subspace as a null space
(Denition LNS [293]) and apply Theorem BNS [160]. Or Theorem FS [299] tells us that the left null space
can be expressed as a row space and we can then use Theorem BRS [280].
Theorem BS [180] is another early result that provides a linearly independent spanning set (i.e. a basis)
as its conclusion. If a vector space of column vectors can be expressed as a span of a set of column vectors,
then Theorem BS [180] can be employed in a straightforward manner to quickly yield a basis.
Subsection BSCV
Bases for Spans of Column Vectors
We have seen several examples of bases in dierent vector spaces. In this subsection, and the next (Sub-
section B.BNM [376]), we will consider building bases for Cmand its subspaces.
Suppose we have a subspace of Cmthat is expressed as the span of a set of vectors, S, andSis
not necessarily linearly independent, or perhaps not very attractive. Theorem REMRS [279] says that
row-equivalent matrices have identical row spaces, while Theorem BRS [280] says the nonzero rows of a
matrix in reduced row-echelon form are a basis for the row space. These theorems together give us a great
computational tool for quickly nding a basis for a subspace that is expressed originally as a span.
Example RSB
Row space basis
When we rst dened the span of a set of column vectors, in Example SCAD [139] we looked at the set
W=*8
<
:2
42
3
13
5;2
41
4
13
5;2
47
5
43
5;2
4 7
6
53
59
=
;+
with an eye towards realizing Was the span of a smaller set. By building relations of linear dependence
(though we did not know them by that name then) we were able to remove two vectors and write Was
Version 2.30
Subsection B.BSCV Bases for Spans of Column Vectors 377
the span of the other two vectors. These two remaining vectors formed a linearly independent set, even
though we did not know that at the time.
Now we know that Wis a subspace and must have a basis. Consider the matrix, C, whose rows are
the vectors in the spanning set for W,
C=2
6642 3 1
1 4 1
7 5 4
7 6 53
775
Then, by Denition RSM [278], the row space of Cwill beW,R(C) =W. Theorem BRS [280] tells us
that if we row-reduce C, the nonzero rows of the row-equivalent matrix in reduced row-echelon form will
be a basis forR(C), and hence a basis for W. Let's do it | Crow-reduces to
2
664107
11
011
11
0 0 0
0 0 03
775
If we convert the two nonzero rows to column vectors then we have a basis,
B=8
<
:2
41
0
7
113
5;2
40
1
1
113
59
=
;
and
W=*8
<
:2
41
0
7
113
5;2
40
1
1
113
59
=
;+
For aesthetic reasons, we might wish to multiply each vector in Bby 11, which will not change the spanning
or linear independence properties of Bas a basis. Then we can also write
W=*8
<
:2
411
0
73
5;2
40
11
13
59
=
;+
Example IAS [281] provides another example of this
avor, though now we can notice that Xis a
subspace, and that the resulting set of three vectors is a basis. This is such a powerful technique that we
should do one more example.
Example RS
Reducing a span
In Example RSC5 [176] we began with a set of n= 4 vectors from C5,
R=fv1;v2;v3;v4g=8
>>>><
>>>>:2
666641
2
1
3
23
77775;2
666642
1
3
1
23
77775;2
666640
7
6
11
23
77775;2
666644
1
2
1
63
777759
>>>>=
>>>>;
and dened V=hRi. Our goal in that problem was to nd a relation of linear dependence on the vectors
inR, solve the resulting equation for one of the vectors, and re-express Vas the span of a set of three
vectors.
Version 2.30
378 Section B Bases
Here is another way to accomplish something similar. The row space of the matrix
A=2
6641 2 1 3 2
2 1 3 1 2
0 7 6 11 2
4 1 2 1 63
775
is equal tohRi. By Theorem BRS [280] we can row-reduce this matrix, ignore any zero rows, and use
the non-zero rows as column vectors that are a basis for the row space of A. Row-reducing Acreates the
matrix 2
6641 0 0 1
1730
17
0 1 025
17 2
17
0 0 1 2
17 8
17
0 0 0 0 03
775
So 8
>>>><
>>>>:2
666641
0
0
1
1730
173
77775;2
666640
1
0
25
17
2
173
77775;2
666640
0
1
2
17
8
173
777759
>>>>=
>>>>;
is a basis for V. Our theorem tells us this is a basis, there is no need to verify that the subspace spanned
by three vectors (rather than four) is the identical subspace, and there is no need to verify that we have
reached the limit in reducing the set, since the set of three vectors is guaranteed to be linearly independent.
Subsection BNM
Bases and Nonsingular Matrices
A quick source of diverse bases for Cmis the set of columns of a nonsingular matrix.
Theorem CNMB
Columns of Nonsingular Matrix are a Basis
Suppose that Ais a square matrix of size m. Then the columns of Aare a basis of Cmif and only if Ais
nonsingular.
Proof ()) Suppose that the columns of Aare a basis for Cm. Then Denition B [371] says the set of
columns is linearly independent. Theorem NMLIC [159] then says that Ais nonsingular.
(() Suppose that Ais nonsingular. Then by Theorem NMLIC [159] this set of columns is linearly
independent. Theorem CSNM [277] says that for a nonsingular matrix, C(A) =Cm. This is equivalent
to saying that the columns of Aare a spanning set for the vector space Cm. As a linearly independent
spanning set, the columns of Aqualify as a basis for Cm(Denition B [371]).
Example CABAK
Columns as Basis, Archetype K
Archetype K [825] is the 5 5 matrix
K=2
6666410 18 24 24 12
12 2 6 0 18
30 21 23 30 39
27 30 36 37 30
18 24 30 30 203
77775
Version 2.30
Subsection B.OBC Orthonormal Bases and Coordinates 379
which is row-equivalent to the 5 5 identity matrix I5. So by Theorem NMRRI [84], Kis nonsingular.
Then Theorem CNMB [376] says the set
8
>>>><
>>>>:2
6666410
12
30
27
183
77775;2
6666418
2
21
30
243
77775;2
6666424
6
23
36
303
77775;2
6666424
0
30
37
303
77775;2
66664 12
18
39
30
203
777759
>>>>=
>>>>;
is a (novel) basis of C5.
Perhaps we should view the fact that the standard unit vectors are a basis (Theorem SUVB [371]) as
just a simple corollary of Theorem CNMB [376]? (See Technique LC [774].)
With a new equivalence for a nonsingular matrix, we can update our list of equivalences.
Theorem NME5
Nonsingular Matrix Equivalences, Round 5
Suppose that Ais a square matrix of size n. The following are equivalent.
1.Ais nonsingular.
2.Arow-reduces to the identity matrix.
3. The null space of Acontains only the zero vector, N(A) =f0g.
4. The linear system LS(A;b) has a unique solution for every possible choice of b.
5. The columns of Aare a linearly independent set.
6.Ais invertible.
7. The column space of AisCn,C(A) =Cn.
8. The columns of Aare a basis for Cn.
Proof With a new equivalence for a nonsingular matrix in Theorem CNMB [376] we can expand Theorem
NME4 [277].
Subsection OBC
Orthonormal Bases and Coordinates
We learned about orthogonal sets of vectors in Cmback in Section O [191], and we also learned that
orthogonal sets are automatically linearly independent (Theorem OSLI [198]). When an orthogonal set
also spans a subspace of Cm, then the set is a basis. And when the set is orthonormal, then the set is
an incredibly nice basis. We will back up this claim with a theorem, but rst consider how you might
manufacture such a set.
Suppose that Wis a subspace of Cmwith basisB. ThenBspansWand is a linearly independent
set of nonzero vectors. We can apply the Gram-Schmidt Procedure (Theorem GSP [199]) and obtain a
linearly independent set Tsuch thathTi=hBi=WandTis orthogonal. In other words, Tis a basis for
W, and is an orthogonal set. By scaling each vector of Tto norm 1, we can convert Tinto an orthonormal
set, without destroying the properties that make it a basis of W. In short, we can convert any basis into
an orthonormal basis. Example GSTV [200], followed by Example ONTV [201], illustrates this process.
Version 2.30
380 Section B Bases
Unitary matrices (Denition UM [262]) are another good source of orthonormal bases (and vice versa).
Suppose that Qis a unitary matrix of size n. Then the ncolumns of Qform an orthonormal set (Theorem
CUMOS [263]) that is therefore linearly independent (Theorem OSLI [198]). Since Qis invertible (Theorem
UMI [263]), we know Qis nonsingular (Theorem NI [261]), and then the columns of QspanCn(Theorem
CSNM [277]). So the columns of a unitary matrix of size nare an orthonormal basis for Cn.
Why all the fuss about orthonormal bases? Theorem VRRB [360] told us that any vector in a vector
space could be written, uniquely, as a linear combination of basis vectors. For an orthonormal basis,
nding the scalars for this linear combination is extremely easy, and this is the content of the next theorem.
Furthermore, with vectors written this way (as linear combinations of the elements of an orthonormal set)
certain computations and analysis become much easier. Here's the promised theorem.
Theorem COB
Coordinates and Orthonormal Bases
Suppose that B=fv1;v2;v3; :::; vpgis an orthonormal basis of the subspace WofCm. For any w2W,
w=hw;v1iv1+hw;v2iv2+hw;v3iv3++hw;vpivp
Proof BecauseBis a basis of W, Theorem VRRB [360] tells us that we can write wuniquely as a
linear combination of the vectors in B. So it is not this aspect of the conclusion that makes this theorem
interesting. What is interesting is that the particular scalars are so easy to compute. No need to solve big
systems of equations | just do an inner product of wwithvito arrive at the coecient of viin the linear
combination.
So begin the proof by writing was a linear combination of the vectors in B, using unknown scalars,
w=a1v1+a2v2+a3v3++apvp
and compute,
hw;vii=*pX
k=1akvk;vi+
Theorem VRRB [360]
=pX
k=1hakvk;vii Theorem IPVA [193]
=pX
k=1akhvk;vii Theorem IPSM [194]
=aihvi;vii+pX
i=1
k6=iakhvk;vii Property C [317]
=ai(1) +pX
i=1
k6=iak(0) Denition ONS [201]
=ai
So the (unique) scalars for the linear combination are indeed the inner products advertised in the conclusion
of the theorem's statement.
Example CROB4
Coordinatization relative to an orthonormal basis, C4
Version 2.30
Subsection B.OBC Orthonormal Bases and Coordinates 381
The set
fx1;x2;x3;x4g=8
>><
>>:2
6641 +i
1
1 i
i3
775;2
6641 + 5i
6 + 5i
7 i
1 6i3
775;2
664 7 + 34i
8 23i
10 + 22i
30 + 13i3
775;2
664 2 4i
6 +i
4 + 3i
6 i3
7759
>>=
>>;
was proposed, and partially veried, as an orthogonal set in Example AOS [197]. Let's scale each vector
to norm 1, so as to form an orthonormal set in C4. Then by Theorem OSLI [198] the set will be linearly
independent, and by Theorem NME5 [377] the set will be a basis for C4. So, once scaled to norm 1, the
adjusted set will be an orthonormal basis of C4. The norms are,
kx1k=p
6kx2k=p
174kx3k=p
3451kx4k=p
119
So an orthonormal basis is
B=fv1;v2;v3;v4g
=8
>><
>>:1p
62
6641 +i
1
1 i
i3
775;1p
1742
6641 + 5i
6 + 5i
7 i
1 6i3
775;1p
34512
664 7 + 34i
8 23i
10 + 22i
30 + 13i3
775;1p
1192
664 2 4i
6 +i
4 + 3i
6 i3
7759
>>=
>>;
Now, to illustrate Theorem COB [378], choose any vector from C4, sayw=2
6642
3
1
43
775, and compute
hw;v1i= 5ip
6;hw;v2i= 19 + 30ip
174;hw;v3i=120 211ip
3451;hw;v4i=6 + 12ip
119
Then Theorem COB [378] guarantees that
2
6642
3
1
43
775= 5ip
60
BB@1p
62
6641 +i
1
1 i
i3
7751
CCA+ 19 + 30ip
1740
BB@1p
1742
6641 + 5i
6 + 5i
7 i
1 6i3
7751
CCA
+120 211ip
34510
BB@1p
34512
664 7 + 34i
8 23i
10 + 22i
30 + 13i3
7751
CCA+6 + 12ip
1190
BB@1p
1192
664 2 4i
6 +i
4 + 3i
6 i3
7751
CCA
as you might want to check (if you have unlimited patience).
A slightly less intimidating example follows, in three dimensions and with just real numbers.
Example CROB3
Coordinatization relative to an orthonormal basis, C3
The set
fx1;x2;x3g=8
<
:2
41
2
13
5;2
4 1
0
13
5;2
42
1
13
59
=
;
is a linearly independent set, which the Gram-Schmidt Process (Theorem GSP [199]) converts to an
orthogonal set, and which can then be converted to the orthonormal set,
B=fv1;v2;v3g=8
<
:1p
62
41
2
13
5;1p
22
4 1
0
13
5;1p
32
41
1
13
59
=
;
Version 2.30
382 Section B Bases
which is therefore an orthonormal basis of C3. With three vectors in C3, all with real number entries,
the inner product (Denition IP [192]) reduces to the usual \dot product" (or scalar product) and the
orthogonal pairs of vectors can be interpreted as perpendicular pairs of directions. So the vectors in B
serve as replacements for our usual 3-D axes, or the usual 3-D unit vectors ~i;~jand~k. We would like
to decompose arbitrary vectors into \components" in the directions of each of these basis vectors. It is
Theorem COB [378] that tells us how to do this.
Suppose that we choose w=2
42
1
53
5. Compute
hw;v1i=5p
6hw;v2i=3p
2hw;v3i=8p
3
then Theorem COB [378] guarantees that
2
42
1
53
5=5p
60
@1p
62
41
2
13
51
A+3p
20
@1p
22
4 1
0
13
51
A+8p
30
@1p
32
41
1
13
51
A
which you should be able to check easily, even if you do not have much patience.
Not only do the columns of a unitary matrix form an orthonormal basis, but there is a deeper connection
between orthonormal bases and unitary matrices. Informally, the next theorem says that if we transform
each vector of an orthonormal basis by multiplying it by a unitary matrix, then the resulting set will be
another orthonormal basis. And more remarkably, any matrix with this property must be unitary! As an
equivalence (Technique E [768]) we could take this as our dening property of a unitary matrix, though it
might not have the same utility as Denition UM [262].
Theorem UMCOB
Unitary Matrices Convert Orthonormal Bases
LetAbe annnmatrix and B=fx1;x2;x3; :::; xngbe an orthonormal basis of Cn. Dene
C=fAx1; Ax2; Ax3; :::; A xng
ThenAis a unitary matrix if and only if Cis an orthonormal basis of Cn.
Proof ()) Assume Ais a unitary matrix and establish several facts about C. First we check that C
is an orthonormal set (Denition ONS [201]). By Theorem UMPIP [264], for i6=j,
hAxi; Axji=hxi;xji= 0
Similarly, Theorem UMPIP [264] also gives, for 1 in,
kAxik=kxik= 1
AsCis an orthogonal set (Denition OSV [197]), Theorem OSLI [198] yields the linear independence of C.
Having established that the column vectors on Cform a linearly independent set, a matrix whose columns
are the vectors of Cis nonsingular (Theorem NMLIC [159]), and hence these vectors form a basis of Cn
by Theorem CNMB [376].
(() Now assume that Cis an orthonormal set. Let ybe an arbitrary vector from Cn. SinceBspans
Cn, there are scalars, a1; a2; a3; :::; an, such that
y=a1x1+a2x2+a3x3++anxn
Version 2.30
Subsection B.OBC Orthonormal Bases and Coordinates 383
Now
AAy=nX
i=1hAAy;xiixi Theorem COB [378]
=nX
i=1*
AAnX
j=1ajxj;xi+
xi Denition TSVS [356]
=nX
i=1*nX
j=1AAajxj;xi+
xi Theorem MMDAA [230]
=nX
i=1*nX
j=1ajAAxj;xi+
xi Theorem MMSMM [230]
=nX
i=1nX
j=1hajAAxj;xiixi Theorem IPVA [193]
=nX
i=1nX
j=1ajhAAxj;xiixi Theorem IPSM [194]
=nX
i=1nX
j=1ajhAxj;(A)xiixi Theorem AIP [233]
=nX
i=1nX
j=1ajhAxj; Axiixi Theorem AA [215]
=nX
i=1nX
j=1
j6=iajhAxj; Axiixi+nX
`=1a`hAx`; Ax`ix` Property C [317]
=nX
i=1nX
j=1
j6=iaj(0)xi+nX
`=1a`(1)x` Denition ONS [201]
=nX
i=1nX
j=1
j6=i0+nX
`=1a`x` Theorem ZSSM [324]
=nX
`=1a`x` Property Z [318]
=y
=Iny Theorem MMIM [229]
Since the choice of ywas arbitrary, Theorem EMMVP [225] tells us that AA=In, soAis unitary
(Denition UM [262]).
Version 2.30
384 Section B Bases
Subsection READ
Reading Questions
1. The matrix below is nonsingular. What can you now say about its columns?
A=2
4 3 0 1
1 2 1
5 1 63
5
2. Write the vector w=2
46
6
153
5as a linear combination of the columns of the matrix Aabove. How many
ways are there to answer this question?
3. Why is an orthonormal basis desirable?
Version 2.30
Subsection B.EXC Exercises 385
Subsection EXC
Exercises
C10 Find a basis forhSi, where
S=8
>><
>>:2
6641
3
2
13
775;2
6641
2
1
13
775;2
6641
1
0
13
775;2
6641
2
2
13
775;2
6643
4
1
33
7759
>>=
>>;:
Contributed by Chris Black Solution [385]
C11 Find a basis for the subspace WofC4,
W=8
>><
>>:2
664a+b 2c
a+b 2c+d
2a+ 2b+ 4c d
b+d3
775a;b;c;d2C9
>>=
>>;
Contributed by Chris Black Solution [385]
C12 Find a basis for the vector space Tof lower triangular 3 3 matrices; that is, matrices of the form2
40 0
0
3
5where an asterisk represents any complex number.
Contributed by Chris Black Solution [386]
C13 Find a basis for the subspace QofP2, dened by Q=
p(x) =a+bx+cx2p(0) = 0
.
Contributed by Chris Black Solution [386]
C14 Find a basis for the subspace RofP2dened by R=
p(x) =a+bx+cx2p0(0) = 0
, wherep0
denotes the derivative.
Contributed by Chris Black Solution [386]
C40 From Example RSB [374], form an arbitrary (and nontrivial) linear combination of the four vectors
in the original spanning set for W. So the result of this computation is of course an element of W. As
such, this vector should be a linear combination of the basis vectors in B. Find the (unique) scalars that
provide this linear combination. Repeat with another linear combination of the original four vectors.
Contributed by Robert Beezer Solution [387]
C80 Prove thatf(1;2);(2;3)gis a basis for the crazy vector space C(Example CVS [322]).
Contributed by Robert Beezer
M20 In Example BM [372] provide the verications (linear independence and spanning) to show that B
is a basis of Mmn.
Contributed by Robert Beezer Solution [386]
T50 Theorem UMCOB [380] says that unitary matrices are characterized as those matrices that \carry"
orthonormal bases to orthonormal bases. This problem asks you to prove a similar result: nonsingular
matrices are characterized as those matrices that \carry" bases to bases.
More precisely, suppose that Ais a square matrix of size nandB=fx1;x2;x3; :::; xngis a basis of
Cn. Prove that Ais nonsingular if and only if C=fAx1; Ax2; Ax3; :::; A xngis a basis of Cn. (See also
Exercise PD.T33 [418], Exercise MR.T20 [637].)
Contributed by Robert Beezer Solution [387]
Version 2.30
386 Section B Bases
T51 Use the result of Exercise B.T50 [383] to build a very concise proof of Theorem CNMB [376]. (Hint:
make a judicious choice for the basis B.)
Contributed by Robert Beezer Solution [389]
Version 2.30
Subsection B.SOL Solutions 387
Subsection SOL
Solutions
C10 Contributed by Chris Black Statement [383]
Theorem BS [180] says that if we take these 5 vectors, put them into a matrix, and row-reduce to discover
the pivot columns, then the corresponding vectors in Swill be linearly independent and span S, and thus
will form a basis of S.
2
6641 1 1 1 3
3 2 1 2 4
2 1 0 2 1
1 1 1 1 33
775RREF !2
66410 1 0 2
01 2 0 5
0 0 0 1 0
0 0 0 0 03
775
Thus, the independent vectors that span Sare the rst, second and fourth of the set, so a basis of Sis
B=8
>><
>>:2
6641
3
2
13
775;2
6641
2
1
13
775;2
6641
2
2
13
7759
>>=
>>;
C11 Contributed by Chris Black Statement [383]
We can rewrite an arbitrary vector of Was
2
664a+b 2c
a+b 2c+d
2a+ 2b+ 4c d
b+d3
775=2
664a
a
2a
03
775+2
664b
b
2b
b3
775+2
664 2c
2c
4c
03
775+2
6640
d
d
d3
775
=a2
6641
1
2
03
775+b2
6641
1
2
13
775+c2
664 2
2
4
03
775+d2
6640
1
1
13
775
Thus, we can write Was
W=*8
>><
>>:2
6641
1
2
03
775;2
6641
1
2
13
775;2
664 2
2
4
03
775;2
6640
1
1
13
7759
>>=
>>;+
These four vectors span W, but we also need to determine if they are linearly independent (turns out they
are not). With an application of Theorem BS [180] we can see that the arrive at a basis employing three
of these vectors,
2
6641 1 2 0
1 1 2 1
2 2 4 1
0 1 0 13
775RREF !2
66410 2 0
01 0 0
0 0 0 1
0 0 0 03
775
Thus, we have the following basis of W,
B=8
>><
>>:2
6641
1
2
03
775;2
6641
1
2
13
775;2
6640
1
1
13
7759
>>=
>>;
Version 2.30
388 Section B Bases
C12 Contributed by Chris Black Statement [383]
LetAbe an arbitrary element of the specied vector space T. Then there exist a,b,c,d,eandfso that
A=2
4a0 0
b c 0
d e f3
5. Then
A=a2
41 0 0
0 0 0
0 0 03
5+b2
40 0 0
1 0 0
0 0 03
5+c2
40 0 0
0 1 0
0 0 03
5+d2
40 0 0
0 0 0
1 0 03
5+e2
40 0 0
0 0 0
0 1 03
5+f2
40 0 0
0 0 0
0 0 13
5
Consider the set
B=8
<
:2
41 0 0
0 0 0
0 0 03
5;2
40 0 0
1 0 0
0 0 03
5;2
40 0 0
0 1 0
0 0 03
5;2
40 0 0
0 0 0
1 0 03
5;2
40 0 0
0 0 0
0 1 03
5;2
40 0 0
0 0 0
0 0 13
59
=
;
The six vectors in Bspan the vector space T, and we can check rather simply that they are also linearly
independent. Thus, Bis a basis of T.
C13 Contributed by Chris Black Statement [383]
Ifp(0) = 0, then a+b(0) +c(02) = 0, soa= 0. Thus, we can write Q=
p(x) =bx+cx2b;c2C
. A
linearly independent set that spans QisB=
x;x2
, and this set forms a basis of Q.
C14 Contributed by Chris Black Statement [383]
The derivative of p(x) =a+bx+cx2isp0(x) =b+ 2cx. Thus, ifp2R, thenp0(0) =b+ 2c(0) = 0, so we
must haveb= 0. We see that we can rewrite RasR=
p(x) =a+cx2a;c2C
. A linearly independent
set that spans RisB=
1;x2
, andBis a basis of R.
M20 Contributed by Robert Beezer Statement [383]
We need to establish the linear independence and spanning properties of the set
B=fBk`j1km;1`ng
relative to the vector space Mmn.
This proof is more transparent if you write out individual matrices in the basis with lots of zeros and
dots and a lone one. But we don't have room for that here, so we will use summation notation. Think
carefully about each step, especially when the double summations seem to \disappear." Begin with a
relation of linear dependence, using double subscripts on the scalars to align with the basis elements.
O=mX
k=1nX
`=1k`Bk`
Now consider the entry in row iand column jfor these equal matrices,
0 = [O]ij Denition ZM [210]
="mX
k=1nX
`=1k`Bk`#
ijDenition ME [207]
=mX
k=1nX
`=1[k`Bk`]ij Denition MA [207]
=mX
k=1nX
`=1k`[Bk`]ij Denition MSM [208]
=ij[Bij]ij[Bk`]ij= 0 when ( k;`)6= (i;j)
Version 2.30
Subsection B.SOL Solutions 389
=ij(1) [ Bij]ij= 1
=ij
Sinceiandjwere arbitrary, we nd that each scalar is zero and so Bis linearly independent (Denition
LI [351]).
To establish the spanning property of Bwe need only show that an arbitrary matrix Acan be written
as a linear combination of the elements of B. So suppose that Ais an arbitrary mnmatrix and consider
the matrix Cdened as a linear combination of the elements of Bby
C=mX
k=1nX
`=1[A]k`Bk`
Then,
[C]ij="mX
k=1nX
`=1[A]k`Bk`#
ijDenition ME [207]
=mX
k=1nX
`=1[[A]k`Bk`]ijDenition MA [207]
=mX
k=1nX
`=1[A]k`[Bk`]ij Denition MSM [208]
= [A]ij[Bij]ij[Bk`]ij= 0 when ( k;`)6= (i;j)
= [A]ij(1) [ Bij]ij= 1
= [A]ij
So by Denition ME [207], A=C, and therefore A2hBi. By Denition B [371], the set Bis a basis of
the vector space Mmn.
C40 Contributed by Robert Beezer Statement [383]
An arbitrary linear combination is
y= 32
42
3
13
5+ ( 2)2
41
4
13
5+ 12
47
5
43
5+ ( 2)2
4 7
6
53
5=2
425
10
153
5
(You probably used a dierent collection of scalars.) We want to write yas a linear combination of
B=8
<
:2
41
0
7
113
5;2
40
1
1
113
59
=
;
We could set this up as vector equation with variables as scalars in a linear combination of the vectors
inB, but since the rst two slots of Bhave such a nice pattern of zeros and ones, we can determine the
necessary scalars easily and then double-check our answer with a computation in the third slot,
252
41
0
7
113
5+ ( 10)2
40
1
1
113
5=2
425
10
(25)7
11+ ( 10)1
113
5=2
425
10
153
5=y
Notice how the uniqueness of these scalars arises. They are forced to be 25 and 10.
T50 Contributed by Robert Beezer Statement [383]
Our rst proof relies mostly on denitions of linear independence and spanning, which is a good exercise.
Version 2.30
390 Section B Bases
The second proof is shorter and turns on a technical result from our work with matrix inverses, Theorem
NPNT [259].
()) Assume that Ais nonsingular and prove that Cis a basis of Cn. First show that Cis linearly
independent. Work on a relation of linear dependence on C,
0=a1Ax1+a2Ax2+a3Ax3++anAxn Denition RLD [351]
=Aa1x1+Aa2x2+Aa3x3++Aanxn Theorem MMSMM [230]
=A(a1x1+a2x2+a3x3++anxn) Theorem MMDAA [230]
SinceAis nonsingular, Denition NM [83] and Theorem SLEMM [224] allows us to conclude that
a1x1+a2x2++anxn=0
But this is a relation of linear dependence of the linearly independent set B, so the scalars are trivial,
a1=a2=a3==an= 0. By Denition LI [351], the set Cis linearly independent.
Now prove that Cspans Cn. Given an arbitrary vector y2Cn, can it be expressed as a linear
combination of the vectors in C? SinceAis a nonsingular matrix we can dene the vector wto be the
unique solution of the system LS(A;y) (Theorem NMUS [86]). Since w2Cnwe can write was a linear
combination of the vectors in the basis B. So there are scalars, b1; b2; b3; :::; bnsuch that
w=b1x1+b2x2+b3x3++bnxn
Then,
y=Aw Theorem SLEMM [224]
=A(b1x1+b2x2+b3x3++bnxn) Denition TSVS [356]
=Ab1x1+Ab2x2+Ab3x3++Abnxn Theorem MMDAA [230]
=b1Ax1+b2Ax2+b3Ax3++bnAxn Theorem MMSMM [230]
So we can write an arbitrary vector of Cnas a linear combination of the elements of C. In other words, C
spans Cn(Denition TSVS [356]). By Denition B [371], the set Cis a basis for Cn.
(() Assume that Cis a basis and prove that Ais nonsingular. Let xbe a solution to the homogeneous
systemLS(A;0). SinceBis a basis of Cnthere are scalars, a1; a2; a3; :::; an, such that
x=a1x1+a2x2+a3x3++anxn
Then
0=Ax Theorem SLEMM [224]
=A(a1x1+a2x2+a3x3++anxn) Denition TSVS [356]
=Aa1x1+Aa2x2+Aa3x3++Aanxn Theorem MMDAA [230]
=a1Ax1+a2Ax2+a3Ax3++anAxn Theorem MMSMM [230]
This is a relation of linear dependence on the linearly independent set C, so the scalars must all be zero,
a1=a2=a3==an= 0. Thus,
x=a1x1+a2x2+a3x3++anxn= 0x1+ 0x2+ 0x3++ 0xn=0:
By Denition NM [83] we see that Ais nonsingular.
Now for a second proof. Take the vectors for Band use them as the columns of a matrix, G=
[x1jx2jx3j:::jxn]. By Theorem CNMB [376], because we have the hypothesis that Bis a basis of Cn,Gis
Version 2.30
Subsection B.SOL Solutions 391
a nonsingular matrix. Notice that the columns of AGare exactly the vectors in the set C, by Denition
MM [226].
Anonsingular()AGnonsingular Theorem NPNT [259]
()Cbasis for CnTheorem CNMB [376]
That was easy!
T51 Contributed by Robert Beezer Statement [384]
ChooseBto be the set of standard unit vectors, a particularly nice basis of Cn(Theorem SUVB [371]).
For a vector ej(Denition SUV [197]) from this basis, what is Aej?
Version 2.30
392 Section B Bases
Version 2.30
Section D Dimension 393
Section D
Dimension
Almost every vector space we have encountered has been innite in size (an exception is Example VSS
[321]). But some are bigger and richer than others. Dimension, once suitably dened, will be a measure of
the size of a vector space, and a useful tool for studying its properties. You probably already have a rough
notion of what a mathematical denition of dimension might be | try to forget these imprecise ideas and
go with the new ones given here.
Subsection D
Dimension
Denition D
Dimension
Suppose that Vis a vector space and fv1;v2;v3; :::; vtgis a basis of V. Then the dimension ofVis
dened by dim ( V) =t. IfVhas no nite bases, we say Vhas innite dimension.
(This denition contains Notation D.) 4
This is a very simple denition, which belies its power. Grab a basis, any basis, and count up the
number of vectors it contains. That's the dimension. However, this simplicity causes a problem. Given a
vector space, you and I could each construct dierent bases | remember that a vector space might have
many bases. And what if your basis and my basis had dierent sizes? Applying Denition D [391] we
would arrive at dierent numbers! With our current knowledge about vector spaces, we would have to say
that dimension is not \well-dened." Fortunately, there is a theorem that will correct this problem.
In a strictly logical progression, the next two theorems would precede the denition of dimension. Many
subsequent theorems will trace their lineage back to the following fundamental result.
Theorem SSLD
Spanning Sets and Linear Dependence
Suppose that S=fv1;v2;v3; :::; vtgis a nite set of vectors which spans the vector space V. Then any
set oft+ 1 or more vectors from Vis linearly dependent.
Proof We want to prove that any set of t+ 1 or more vectors from Vis linearly dependent. So we will
begin with a totally arbitrary set of vectors from V,R=fu1;u2;u3; :::; umg, wherem>t . We will now
construct a nontrivial relation of linear dependence on R.
Each vector u1;u2;u3; :::; umcan be written as a linear combination of v1;v2;v3; :::; vtsinceSis
a spanning set of V. This means there exist scalars aij, 1it, 1jm, so that
u1=a11v1+a21v2+a31v3++at1vt
u2=a12v1+a22v2+a32v3++at2vt
u3=a13v1+a23v2+a33v3++at3vt
...
um=a1mv1+a2mv2+a3mv3++atmvt
Now we form, unmotivated, the homogeneous system of tequations in the mvariables,x1; x2; x3; :::; xm,
where the coecients are the just-discovered scalars aij,
a11x1+a12x2+a13x3++a1mxm= 0
Version 2.30
394 Section D Dimension
a21x1+a22x2+a23x3++a2mxm= 0
a31x1+a32x2+a33x3++a3mxm= 0
...
at1x1+at2x2+at3x3++atmxm= 0
This is a homogeneous system with more variables than equations (our hypothesis is expressed as m>t ),
so by Theorem HMVEI [73] there are innitely many solutions. Choose a nontrivial solution and denote
it byx1=c1; x2=c2; x3=c3; :::; xm=cm. As a solution to the homogeneous system, we then have
a11c1+a12c2+a13c3++a1mcm= 0
a21c1+a22c2+a23c3++a2mcm= 0
a31c1+a32c2+a33c3++a3mcm= 0
...
at1c1+at2c2+at3c3++atmcm= 0
As a collection of nontrivial scalars, c1; c2; c3; :::; cmwill provide the nontrivial relation of linear depen-
dence we desire,
c1u1+c2u2+c3u3++cmum
=c1(a11v1+a21v2+a31v3++at1vt) Denition TSVS [356]
+c2(a12v1+a22v2+a32v3++at2vt)
+c3(a13v1+a23v2+a33v3++at3vt)
...
+cm(a1mv1+a2mv2+a3mv3++atmvt)
=c1a11v1+c1a21v2+c1a31v3++c1at1vt Property DVA [318]
+c2a12v1+c2a22v2+c2a32v3++c2at2vt
+c3a13v1+c3a23v2+c3a33v3++c3at3vt
...
+cma1mv1+cma2mv2+cma3mv3++cmatmvt
= (c1a11+c2a12+c3a13++cma1m)v1 Property DSA [318]
+ (c1a21+c2a22+c3a23++cma2m)v2
+ (c1a31+c2a32+c3a33++cma3m)v3
...
+ (c1at1+c2at2+c3at3++cmatm)vt
= (a11c1+a12c2+a13c3++a1mcm)v1 Property CMCN [758]
+ (a21c1+a22c2+a23c3++a2mcm)v2
+ (a31c1+a32c2+a33c3++a3mcm)v3
...
+ (at1c1+at2c2+at3c3++atmcm)vt
= 0v1+ 0v2+ 0v3++ 0vt cjas solution
Version 2.30
Subsection D.D Dimension 395
=0+0+0++0 Theorem ZSSM [324]
=0 Property Z [318]
That does it. Rhas been undeniably shown to be a linearly dependent set.
The proof just given has some monstrous expressions in it, mostly owing to the double subscripts
present. Now is a great opportunity to show the value of a more compact notation. We will rewrite the key
steps of the previous proof using summation notation, resulting in a more economical presentation, and
even greater insight into the key aspects of the proof. So here is an alternate proof | study it carefully.
Proof (Alternate Proof of Theorem SSLD) We want to prove that any set of t+ 1 or more
vectors from Vis linearly dependent. So we will begin with a totally arbitrary set of vectors from V,
R=fujj1jmg, wherem > t . We will now construct a nontrivial relation of linear dependence on
R.
Each vector uj, 1jmcan be written as a linear combination of vi, 1itsinceSis a spanning
set ofV. This means there are scalars aij, 1it, 1jm, so that
uj=tX
i=1aijvi 1jm
Now we form, unmotivated, the homogeneous system of tequations in the mvariables,xj, 1jm,
where the coecients are the just-discovered scalars aij,
mX
j=1aijxj= 0 1 it
This is a homogeneous system with more variables than equations (our hypothesis is expressed as m>t ),
so by Theorem HMVEI [73] there are innitely many solutions. Choose one of these solutions that is not
trivial and denote it by xj=cj, 1jm. As a solution to the homogeneous system, we then havePm
j=1aijcj= 0 for 1it. As a collection of nontrivial scalars, cj, 1jm, will provide the nontrivial
relation of linear dependence we desire,
mX
j=1cjuj=mX
j=1cj tX
i=1aijvi!
Denition TSVS [356]
=mX
j=1tX
i=1cjaijvi Property DVA [318]
=tX
i=1mX
j=1cjaijvi Property CMCN [758]
=tX
i=1mX
j=1aijcjvi Commutativity in C
=tX
i=10
@mX
j=1aijcj1
Avi Property DSA [318]
=tX
i=10vi cjas solution
=tX
i=10 Theorem ZSSM [324]
Version 2.30
396 Section D Dimension
=0 Property Z [318]
That does it. Rhas been undeniably shown to be a linearly dependent set.
Notice how the swap of the two summations is so much easier in the third step above, as opposed to
all the rearranging and regrouping that takes place in the previous proof. In about half the space. And
there are no ellipses ( :::).
Theorem SSLD [391] can be viewed as a generalization of Theorem MVSLD [158]. We know that Cm
has a basis with mvectors in it (Theorem SUVB [371]), so it is a set of mvectors that spans Cm. By
Theorem SSLD [391], any set of more than mvectors from Cmwill be linearly dependent. But this is
exactly the conclusion we have in Theorem MVSLD [158]. Maybe this is not a total shock, as the proofs
of both theorems rely heavily on Theorem HMVEI [73]. The beauty of Theorem SSLD [391] is that it
applies in any vector space. We illustrate the generality of this theorem, and hint at its power, in the next
example.
Example LDP4
Linearly dependent set in P4
In Example SSP4 [356] we showed that
S=
x 2; x2 4x+ 4; x3 6x2+ 12x 8; x4 8x3+ 24x2 32x+ 16
is a spanning set for W=fp(x)jp2P4; p(2) = 0g. So we can apply Theorem SSLD [391] to Wwith
t= 4. Here is a set of ve vectors from W, as you may check by verifying that each is a polynomial of
degree 4 or less and has x= 2 as a root,
T=fp1; p2; p3; p4; p5gW
p1=x4 2x3+ 2x2 8x+ 8
p2= x3+ 6x2 5x 6
p3= 2x4 5x3+ 5x2 7x+ 2
p4= x4+ 4x3 7x2+ 6x
p5= 4x3 9x2+ 5x 6
By Theorem SSLD [391] we conclude that Tis linearly dependent, with no further computations.
Theorem SSLD [391] is indeed powerful, but our main purpose in proving it right now was to make
sure that our denition of dimension (Denition D [391]) is well-dened. Here's the theorem.
Theorem BIS
Bases have Identical Sizes
Suppose that Vis a vector space with a nite basis Band a second basis C. ThenBandChave the same
size.
Proof Suppose that Chas more vectors than B. (Allowing for the possibility that Cis innite, we can
replaceCby a subset that has more vectors than B.) As a basis, Bis a spanning set for V(Denition B
[371]), so Theorem SSLD [391] says that Cis linearly dependent. However, this contradicts the fact that
as a basisCis linearly independent (Denition B [371]). So Cmust also be a nite set, with size less than,
or equal to, that of B.
Suppose that Bhas more vectors than C. As a basis, Cis a spanning set for V(Denition B [371]), so
Theorem SSLD [391] says that Bis linearly dependent. However, this contradicts the fact that as a basis
Bis linearly independent (Denition B [371]). So Ccannot be strictly smaller than B.
The only possibility left for the sizes of BandCis for them to be equal.
Theorem BIS [394] tells us that if we nd one nite basis in a vector space, then they all have the same
size. This (nally) makes Denition D [391] unambiguous.
Version 2.30
Subsection D.DVS Dimension of Vector Spaces 397
Subsection DVS
Dimension of Vector Spaces
We can now collect the dimension of some common, and not so common, vector spaces.
Theorem DCM
Dimension of Cm
The dimension of Cm(Example VSCV [319]) is m.
Proof Theorem SUVB [371] provides a basis with mvectors.
Theorem DP
Dimension of Pn
The dimension of Pn(Example VSP [319]) is n+ 1.
Proof Example BP [372] provides twobases with n+ 1 vectors. Take your pick.
Theorem DM
Dimension of Mmn
The dimension of Mmn(Example VSM [319]) is mn.
Proof Example BM [372] provides a basis with mnvectors.
Example DSM22
Dimension of a subspace of M22
It should now be plausible that
Z=a b
c d2a+b+ 3c+ 4d= 0; a+ 3b 5c 2d= 0
is a subspace of the vector space M22(Example VSM [319]). (It is.) To nd the dimension of Zwe must
rst nd a basis, though any old basis will do.
First concentrate on the conditions relating a; b; c andd. They form a homogeneous system of two
equations in four variables with coecient matrix
2 1 3 4
1 3 5 2
We can row-reduce this matrix to obtain
10 2 2
01 1 0
Rewrite the two equations represented by each row of this matrix, expressing the dependent variables ( a
andb) in terms of the free variables ( candd), and we obtain,
a= 2c 2d
b=c
We can now write a typical entry of Zstrictly in terms of candd, and we can decompose the result,
a b
c d
= 2c 2d c
c d
= 2c c
c0
+ 2d0
0d
=c 2 1
1 0
+d 2 0
0 1
Version 2.30
398 Section D Dimension
this equation says that an arbitrary matrix in Zcan be written as a linear combination of the two vectors
in
S= 2 1
1 0
; 2 0
0 1
so we know that
Z=hSi= 2 1
1 0
; 2 0
0 1
Are these two matrices (vectors) also linearly independent? Begin with a relation of linear dependence on
S,
a1 2 1
1 0
+a2 2 0
0 1
=O
2a1 2a2a1
a1a2
=0 0
0 0
From the equality of the two entries in the last row, we conclude that a1= 0,a2= 0. Thus the only
possible relation of linear dependence is the trivial one, and therefore Sis linearly independent (Denition
LI [351]). So Sis a basis for V(Denition B [371]). Finally, we can conclude that dim ( Z) = 2 (Denition
D [391]) since Shas two elements.
Example DSP4
Dimension of a subspace of P4
In Example BSP4 [372] we showed that
S=
x 2; x2 4x+ 4; x3 6x2+ 12x 8; x4 8x3+ 24x2 32x+ 16
is a basis for W=fp(x)jp2P4; p(2) = 0g. Thus, the dimension of Wis four, dim ( W) = 4.
Note that dim ( P4) = 5 by Theorem DP [395], so Wis a subspace of dimension 4 within the vector
spaceP4of dimension 5, illustrating the upcoming Theorem PSSD [410].
Example DC
Dimension of the crazy vector space
In Example BC [374] we determined that the set R=f(1;0);(6;3)gfrom the crazy vector space, C
(Example CVS [322]), is a basis for C. By Denition D [391] we see that Chas dimension 2, dim ( C) = 2.
It is possible for a vector space to have no nite bases, in which case we say it has innite dimension.
Many of the best examples of this are vector spaces of functions, which lead to constructions like Hilbert
spaces. We will focus exclusively on nite-dimensional vector spaces. OK, one innite-dimensional example,
and then we will focus exclusively on nite-dimensional vector spaces.
Example VSPUD
Vector space of polynomials with unbounded degree
Dene the set Pby
P=fpjp(x) is a polynomial in xg
Our operations will be the same as those dened for Pn(Example VSP [319]).
With no restrictions on the possible degrees of our polynomials, any nite set that is a candidate for
spanningPwill come up short. We will give a proof by contradiction (Technique CD [770]). To this end,
suppose that the dimension of Pis nite, say dim ( P) =n.
The setT=
1; x; x2; :::; xn
is a linearly independent set (check this!) containing n+1 polynomials
fromP. However, a basis of Pwill be a spanning set of Pcontaining nvectors. This situation is a
contradiction of Theorem SSLD [391], so our assumption that Phas nite dimension is false. Thus, we
say dim (P) =1.
Version 2.30
Subsection D.RNM Rank and Nullity of a Matrix 399
Subsection RNM
Rank and Nullity of a Matrix
For any matrix, we have seen that we can associate several subspaces | the null space (Theorem NSMS
[337]), the column space (Theorem CSMS [343]), row space (Theorem RSMS [344]) and the left null space
(Theorem LNSMS [344]). As vector spaces, each of these has a dimension, and for the null space and
column space, they are important enough to warrant names.
Denition NOM
Nullity Of a Matrix
Suppose that Ais anmnmatrix. Then the nullity ofAis the dimension of the null space of A,
n(A) = dim (N(A)).
(This denition contains Notation NOM.) 4
Denition ROM
Rank Of a Matrix
Suppose that Ais anmnmatrix. Then the rank ofAis the dimension of the column space of A,
r(A) = dim (C(A)).
(This denition contains Notation ROM.) 4
Example RNM
Rank and nullity of a matrix
Let's compute the rank and nullity of
A=2
66666642 4 1 3 2 1 4
1 2 0 0 4 0 1
2 4 1 0 5 4 8
1 2 1 1 6 1 3
2 4 1 1 4 2 1
1 2 3 1 6 3 13
7777775
To do this, we will rst row-reduce the matrix since that will help us determine bases for the null space
and column space.2
666666641 2 0 0 4 0 1
0 0 10 3 0 2
0 0 0 1 1 0 3
0 0 0 0 0 1 1
0 0 0 0 0 0 0
0 0 0 0 0 0 03
77777775
From this row-equivalent matrix in reduced row-echelon form we record D=f1;3;4;6gandF=f2;5;7g.
For each index in D, Theorem BCS [274] creates a single basis vector. In total the basis will have 4
vectors, so the column space of Awill have dimension 4 and we write r(A) = 4.
For each index in F, Theorem BNS [160] creates a single basis vector. In total the basis will have 3
vectors, so the null space of Awill have dimension 3 and we write n(A) = 3.
There were no accidents or coincidences in the previous example | with the row-reduced version of a
matrix in hand, the rank and nullity are easy to compute.
Theorem CRN
Computing Rank and Nullity
Suppose that Ais anmnmatrix and Bis a row-equivalent matrix in reduced row-echelon form with r
Version 2.30
400 Section D Dimension
nonzero rows. Then r(A) =randn(A) =n r.
Proof Theorem BCS [274] provides a basis for the column space by choosing columns of Athat correspond
to the dependent variables in a description of the solutions to LS(A;0). In the analysis of B, there is
one dependent variable for each leading 1, one per nonzero row, or one per pivot column. So there are r
column vectors in a basis for C(A).
Theorem BNS [160] provide a basis for the null space by creating basis vectors of the null space of A
from entries of B, one for each independent variable, one per column with out a leading 1. So there are
n rcolumn vectors in a basis for n(A).
Every archetype (Appendix A [777]) that involves a matrix lists its rank and nullity. You may have
noticed as you studied the archetypes that the larger the column space is the smaller the null space is. A
simple corollary states this trade-o succinctly. (See Technique LC [774].)
Theorem RPNC
Rank Plus Nullity is Columns
Suppose that Ais anmnmatrix. Then r(A) +n(A) =n.
Proof Letrbe the number of nonzero rows in a row-equivalent matrix in reduced row-echelon form. By
Theorem CRN [397],
r(A) +n(A) =r+ (n r) =n
When we rst introduced ras our standard notation for the number of nonzero rows in a matrix in
reduced row-echelon form you might have thought rstood for \rows." Not really | it stands for \rank"!
Subsection RNNM
Rank and Nullity of a Nonsingular Matrix
Let's take a look at the rank and nullity of a square matrix.
Example RNSM
Rank and nullity of a square matrix
The matrix
E=2
6666666640 4 1 2 2 3 1
2 2 1 1 0 4 3
2 3 9 3 9 1 9
3 4 9 4 1 6 2
3 4 6 2 5 9 4
9 3 8 2 4 2 4
8 2 2 9 3 0 93
777777775
is row-equivalent to the matrix in reduced row-echelon form,
2
666666666410 0 0 0 0 0
010 0 0 0 0
0 0 10 0 0 0
0 0 0 10 0 0
0 0 0 0 10 0
0 0 0 0 0 10
0 0 0 0 0 0 13
7777777775
Version 2.30
Subsection D.RNNM Rank and Nullity of a Nonsingular Matrix 401
Withn= 7 columns and r= 7 nonzero rows Theorem CRN [397] tells us the rank is r(E) = 7 and the
nullity isn(E) = 7 7 = 0.
The value of either the nullity or the rank are enough to characterize a nonsingular matrix.
Theorem RNNM
Rank and Nullity of a Nonsingular Matrix
Suppose that Ais a square matrix of size n. The following are equivalent.
1. A is nonsingular.
2. The rank of Aisn,r(A) =n.
3. The nullity of Ais zero,n(A) = 0.
Proof (1)2) Theorem CSNM [277] says that if Ais nonsingular then C(A) =Cn. IfC(A) =Cn, then
the column space has dimension nby Theorem DCM [395], so the rank of Aisn.
(2)3) Suppose r(A) =n. Then Theorem RPNC [398] gives
n(A) =n r(A) Theorem RPNC [398]
=n n Hypothesis
= 0
(3)1) Suppose n(A) = 0, so a basis for the null space of Ais the empty set. This implies that N(A) =f0g
and Theorem NMTNS [86] says Ais nonsingular.
With a new equivalence for a nonsingular matrix, we can update our list of equivalences (Theorem
NME5 [377]) which now becomes a list requiring double digits to number.
Theorem NME6
Nonsingular Matrix Equivalences, Round 6
Suppose that Ais a square matrix of size n. The following are equivalent.
1.Ais nonsingular.
2.Arow-reduces to the identity matrix.
3. The null space of Acontains only the zero vector, N(A) =f0g.
4. The linear system LS(A;b) has a unique solution for every possible choice of b.
5. The columns of Aare a linearly independent set.
6.Ais invertible.
7. The column space of AisCn,C(A) =Cn.
8. The columns of Aare a basis for Cn.
9. The rank of Aisn,r(A) =n.
10. The nullity of Ais zero,n(A) = 0.
Proof Building on Theorem NME5 [377] we can add two of the statements from Theorem RNNM [399].
Version 2.30
402 Section D Dimension
Subsection READ
Reading Questions
1. What is the dimension of the vector space P6, the set of all polynomials of degree 6 or less?
2. How are the rank and nullity of a matrix related?
3. Explain why we might say that a nonsingular matrix has \full rank."
Version 2.30
Subsection D.EXC Exercises 403
Subsection EXC
Exercises
C20 The archetypes listed below are matrices, or systems of equations with coecient matrices. For
each, compute the nullity and rank of the matrix. This information is listed for each archetype (along with
the number of columns in the matrix, so as to illustrate Theorem RPNC [398]), and notice how it could
have been computed immediately after the determination of the sets DandFassociated with the reduced
row-echelon form of the matrix.
Archetype A [781]
Archetype B [786]
Archetype C [791]
Archetype D [795]/Archetype E [799]
Archetype F [803]
Archetype G [808]/Archetype H [812]
Archetype I [816]
Archetype J [820]
Archetype K [825]
Archetype L [829]
Contributed by Robert Beezer
C21 Find the dimension of the subspace W=8
>><
>>:2
664a+b
a+c
a+d
d3
775a;b;c;d2C9
>>=
>>;ofC4.
Contributed by Chris Black Solution [403]
C22 Find the dimension of the subspace W=
a+bx+cx2+dx3a+b+c+d= 0
ofP3.
Contributed by Chris Black Solution [403]
C23 Find the dimension of the subspace W=a b
c da+b=c;b+c=d;c+d=a
ofM2;2.
Contributed by Chris Black Solution [403]
C30 For the matrix Abelow, compute the dimension of the null space of A, dim (N(A)).
A=2
6642 1 3 11 9
1 2 1 7 3
3 1 3 6 8
2 1 2 5 33
775
Contributed by Robert Beezer Solution [404]
C31 The setWbelow is a subspace of C4. Find the dimension of W.
W=*8
>><
>>:2
6642
3
4
13
775;2
6643
0
1
23
775;2
664 4
3
2
53
7759
>>=
>>;+
Contributed by Robert Beezer Solution [404]
Version 2.30
404 Section D Dimension
C35 Find the rank and nullity of the matrix A=2
666641 0 1
1 2 2
2 1 1
1 0 1
1 1 23
77775.
Contributed by Chris Black Solution [404]
C36 Find the rank and nullity of the matrix A=2
41 2 1 1 1
1 3 2 0 4
1 2 1 1 13
5.
Contributed by Chris Black Solution [404]
C37 Find the rank and nullity of the matrix A=2
666643 2 1 1 1
2 3 0 1 1
1 1 2 1 0
1 1 0 1 1
0 1 1 2 13
77775.
Contributed by Chris Black Solution [405]
C40 In Example LDP4 [394] we determined that the set of ve polynomials, T, is linearly dependent by
a simple invocation of Theorem SSLD [391]. Prove that Tis linearly dependent from scratch, beginning
with Denition LI [351].
Contributed by Robert Beezer
M20M22is the vector space of 2 2 matrices. Let S22denote the set of all 2 2 symmetric matrices.
That is
S22=
A2M22jAt=A
(a) Show that S22is a subspace of M22.
(b) Exhibit a basis for S22and prove that it has the required properties.
(c) What is the dimension of S22?
Contributed by Robert Beezer Solution [405]
M21 A 22 matrixBis upper triangular if [ B]21= 0. LetUT2be the set of all 2 2 upper triangular
matrices. Then UT2is a subspace of the vector space of all 2 2 matrices, M22(you may assume this).
Determine the dimension of UT2providing allof the necessary justications for your answer.
Contributed by Robert Beezer Solution [405]
Version 2.30
Subsection D.SOL Solutions 405
Subsection SOL
Solutions
C21 Contributed by Chris Black Statement [401]
The subspace Wcan be written as
W=8
>><
>>:2
664a+b
a+c
a+d
d3
775a;b;c;d2C9
>>=
>>;
=8
>><
>>:a2
6641
1
1
03
775+b2
6641
0
0
03
775+c2
6640
1
0
03
775+d2
6640
0
1
13
775a;b;c;d2C9
>>=
>>;
=*8
>><
>>:2
6641
1
1
03
775;2
6641
0
0
03
775;2
6640
1
0
03
775;2
6640
0
1
13
7759
>>=
>>;+
Since the set of vectors8
>><
>>:2
6641
1
1
03
775;2
6641
0
0
03
775;2
6640
1
0
03
775;2
6640
0
1
13
7759
>>=
>>;is a linearly independent set (why?), it forms a basis of
W. Thus,Wis a subspace of C4with dimension 4 (and must therefore equal C4).
C22 Contributed by Chris Black Statement [401]
The subspace W=
a+bx+cx2+dx3a+b+c+d= 0
can be written as
W=
a+bx+cx2+ ( a b c)x3a;b;c2C
=
a(1 x3) +b(x x3) +c(x2 x3)a;b;c2C
=
1 x3;x x3;x2 x3
Since these vectors are linearly independent (why?), Wis a subspace of P3with dimension 3.
C23 Contributed by Chris Black Statement [401]
The equations specied are equivalent to the system
a+b c= 0
b+c d= 0
a c d= 0
The coecient matrix of this system row-reduces to
2
410 0 3
010 1
0 0 1 23
5
Thus, every solution can be decribed with a suitable choice of d, together with a= 3d,b= dandc= 2d.
Thus the subspace Wcan be described as
W=3d d
2d dd2C
=3 1
2 1
Version 2.30
406 Section D Dimension
So,Wis a subspace of M2;2with dimension 1.
C30 Contributed by Robert Beezer Statement [401]
Row reduce A,
ARREF !2
66410 0 1 1
010 3 1
0 0 1 2 2
0 0 0 0 03
775
Sor= 3 for this matrix. Then
dim (N(A)) =n(A) Denition NOM [397]
= (n(A) +r(A)) r(A)
= 5 r(A) Theorem RPNC [398]
= 5 3 Theorem CRN [397]
= 2
We could also use Theorem BNS [160] and create a basis for N(A) withn r= 5 3 = 2 vectors (because
the solutions are described with 2 free variables) and arrive at the dimension as the size of this basis.
C31 Contributed by Robert Beezer Statement [401]
We will appeal to Theorem BS [180] (or you could consider this an appeal to Theorem BCS [274]). Put
the three column vectors of this spanning set into a matrix as columns and row-reduce.
2
6642 3 4
3 0 3
4 1 2
1 2 53
775RREF !2
66410 1
01 2
0 0 0
0 0 03
775
The pivot columns are D=f1;2gso we can \keep" the vectors corresponding to the pivot columns and
set
T=8
>><
>>:2
6642
3
4
13
775;2
6643
0
1
23
7759
>>=
>>;
and conclude that W=hTiandTis linearly independent. In other words, Tis a basis with two vectors,
soWhas dimension 2.
C35 Contributed by Chris Black Statement [402]
The row reduced form of matrix Ais2
66666410 0
010
0 0 1
0 0 0
0 0 03
777775, so the rank of A(number of columns with leading
1's) is 3, and the nullity is 0.
C36 Contributed by Chris Black Statement [402]
The row reduced form of matrix Ais2
410 1 3 5
01 1 1 3
0 0 0 0 03
5, so the rank of A(number of columns with
leading 1's) is 2, and the nullity is 5 2 = 3.
Version 2.30
Subsection D.SOL Solutions 407
C37 Contributed by Chris Black Statement [402]
This matrix Arow reduces to the 5 5 identity matrix, so it has full rank. The rank of Ais 5, and the
nullity is 0.
M20 Contributed by Robert Beezer Statement [402]
(a) We will use the three criteria of Theorem TSS [334]. The zero vector of M22is the zero matrix, O
(Denition ZM [210]), which is a symmetric matrix. So S22is not empty, since O2S22.
Suppose that AandBare two matrices in S22. Then we know that At=AandBt=B. We want to
know ifA+B2S22, so testA+Bfor membership,
(A+B)t=At+BtTheorem TMA [211]
=A+B A; B 2S22
SoA+Bis symmetric and qualies for membership in S22.
Suppose that A2S22and2C. IsA2S22? We know that At=A. Now check that,
At=AtTheorem TMSM [212]
=A A 2S22
SoAis also symmetric and qualies for membership in S22.
With the three criteria of Theorem TSS [334] fullled, we see that S22is a subspace of M22.
(b) An arbitrary matrix from S22can be written asa b
b d
. We can express this matrix as
a b
b d
=a0
0 0
+0b
b0
+0 0
0d
=a1 0
0 0
+b0 1
1 0
+d0 0
0 1
this equation says that the set
T=1 0
0 0
;0 1
1 0
;0 0
0 1
spansS22. Is it also linearly independent?
Write a relation of linear dependence on S,
O=a11 0
0 0
+a20 1
1 0
+a30 0
0 1
0 0
0 0
=a1a2
a2a3
The equality of these two matrices (Denition ME [207]) tells us that a1=a2=a3= 0, and the only
relation of linear dependence on Tis trivial. So Tis linearly independent, and hence is a basis of S22.
(c) The basis Tfound in part (b) has size 3. So by Denition D [391], dim ( S22) = 3.
M21 Contributed by Robert Beezer Statement [402]
A typical matrix from UT2looks likea b
0c
wherea; b; c2Care arbitrary scalars. Observing this we can then write
a b
0c
=a1 0
0 0
+b0 1
0 0
+c0 0
0 1
Version 2.30
408 Section D Dimension
which says that
R=1 0
0 0
;0 1
0 0
;0 0
0 1
is a spanning set for UT2(Denition TSVS [356]). Is Ris linearly independent? If so, it is a basis for UT2.
So consider a relation of linear dependence on R,
11 0
0 0
+20 1
0 0
+30 0
0 1
=O=0 0
0 0
From this equation, one rapidly arrives at the conclusion that 1=2=3= 0. SoRis a linearly
independent set (Denition LI [351]), and hence is a basis (Denition B [371]) for UT2. Now, we simply
count up the size of the set Rto see that the dimension of UT2is dim (UT2) = 3.
Version 2.30
Section PD Properties of Dimension 409
Section PD
Properties of Dimension
Once the dimension of a vector space is known, then the determination of whether or not a set of vectors
is linearly independent, or if it spans the vector space, can often be much easier. In this section we will
state a workhorse theorem and then apply it to the column space and row space of a matrix. It will also
help us describe a super-basis for Cm.
Subsection GT
Goldilocks' Theorem
We begin with a useful theorem that we will need later, and in the proof of the main theorem in this
subsection. This theorem says that we can extend linearly independent sets, one vector at a time, by
adding vectors from outside the span of the linearly independent set, all the while preserving the linear
independence of the set.
Theorem ELIS
Extending Linearly Independent Sets
SupposeVis vector space and Sis a linearly independent set of vectors from V. Suppose wis a vector
such that w62hSi. Then the set S0=S[fwgis linearly independent.
Proof SupposeS=fv1;v2;v3; :::; vmgand begin with a relation of linear dependence on S0,
a1v1+a2v2+a3v3++amvm+am+1w=0:
There are two cases to consider. First suppose that am+1= 0. Then the relation of linear dependence on
S0becomes
a1v1+a2v2+a3v3++amvm=0:
and by the linear independence of the set S, we conclude that a1=a2=a3==am= 0. So all of the
scalars in the relation of linear dependence on S0are zero.
In the second case, suppose that am+16= 0. Then the relation of linear dependence on S0becomes
am+1w= a1v1 a2v2 a3v3 amvm
w= a1
am+1v1 a2
am+1v2 a3
am+1v3 am
am+1vm
This equation expresses was a linear combination of the vectors in S, contrary to the assumption that
w62hSi, so this case leads to a contradiction.
The rst case yielded only a trivial relation of linear dependence on S0and the second case led to a
contradiction. So S0is a linearly independent set since any relation of linear dependence is trivial.
In the story Goldilocks and the Three Bears , the young girl Goldilocks visits the empty house of the
three bears while out walking in the woods. One bowl of porridge is too hot, the other too cold, the third
is just right. One chair is too hard, one too soft, the third is just right. So it is with sets of vectors | some
are too big (linearly dependent), some are too small (they don't span), and some are just right (bases).
Here's Goldilocks' Theorem.
Theorem G
Goldilocks
Suppose that Vis a vector space of dimension t. LetS=fv1;v2;v3; :::; vmgbe a set of vectors from
V. Then
Version 2.30
410 Section PD Properties of Dimension
1. Ifm>t , thenSis linearly dependent.
2. Ifm<t , thenSdoes not span V.
3. Ifm=tandSis linearly independent, then SspansV.
4. Ifm=tandSspansV, thenSis linearly independent.
Proof LetBbe a basis of V. Since dim ( V) =t, Denition B [371] and Theorem BIS [394] imply that
Bis a linearly independent set of tvectors that spans V.
1. Suppose to the contrary that Sis linearly independent. Then Bis a smaller set of vectors that spans
V. This contradicts Theorem SSLD [391].
2. Suppose to the contrary that Sdoes spanV. ThenBis a larger set of vectors that is linearly
independent. This contradicts Theorem SSLD [391].
3. Suppose to the contrary that Sdoes not span V. Then we can choose a vector wsuch that w2V
andw62hSi. By Theorem ELIS [407], the set S0=S[fwgis again linearly independent. Then S0
is a set ofm+ 1 =t+ 1 vectors that are linearly independent, while Bis a set oftvectors that span
V. This contradicts Theorem SSLD [391].
4. Suppose to the contrary that Sis linearly dependent. Then by Theorem DLDS [175] (which can be
upgraded, with no changes in the proof, to the setting of a general vector space), there is a vector
inS, say vkthat is equal to a linear combination of the other vectors in S. LetS0=Snfvkg,
the set of \other" vectors in S. Then it is easy to show that V=hSi=hS0i. SoS0is a set of
m 1 =t 1 vectors that spans V, whileBis a set oftlinearly independent vectors in V. This
contradicts Theorem SSLD [391].
There is a tension in the construction of basis. Make a set too big and you will end up with relations
of linear dependence among the vectors. Make a set too small and you will not have enough raw material
to span the entire vector space. Make a set just the right size (the dimension) and you only need to have
linear independence or spanning, and you get the other property for free. These roughly-stated ideas are
made precise by Theorem G [407].
The structure and proof of this theorem also deserve comment. The hypotheses seem innocuous. We
presume we know the dimension of the vector space in hand, then we mostly just look at the size of the
setS. From this we get big conclusions about spanning and linear independence. Each of the four proofs
relies on ultimately contradicting Theorem SSLD [391], so in a way we could think of this entire theorem
as a corollary of Theorem SSLD [391]. (See Technique LC [774].) The proofs of the third and fourth parts
parallel each other in style (add w, toss vk) and then turn on Theorem ELIS [407] before contradicting
Theorem SSLD [391].
Theorem G [407] is useful in both concrete examples and as a tool in other proofs. We will use it often
to bypass verifying linear independence or spanning.
Example BPR
Bases for Pn, reprised
In Example BP [372] we claimed that
B=
1; x; x2; x3; :::; xn
C=
1;1 +x;1 +x+x2;1 +x+x2+x3; :::; 1 +x+x2+x3++xn
:
Version 2.30
Subsection PD.GT Goldilocks' Theorem 411
were both bases for Pn(Example VSP [319]). Suppose we had rst veried that Bwas a basis, so we
would then know that dim ( Pn) =n+ 1. The size of Cisn+ 1, the right size to be a basis. We could
then verify that Cis linearly independent. We would not have to make any special eorts to prove that C
spansPn, since Theorem G [407] would allow us to conclude this property of Cdirectly. Then we would
be able to say that Cis a basis of Pnalso.
Example BDM22
Basis by dimension in M22
In Example DSM22 [395] we showed that
B= 2 1
1 0
; 2 0
0 1
is a basis for the subspace ZofM22(Example VSM [319]) given by
Z=a b
c d2a+b+ 3c+ 4d= 0; a+ 3b 5c d= 0
This tells us that dim ( Z) = 2. In this example we will nd another basis. We can construct two new
matrices in Zby forming linear combinations of the matrices in B.
2 2 1
1 0
+ ( 3) 2 0
0 1
=2 2
2 3
3 2 1
1 0
+ 1 2 0
0 1
= 8 3
3 1
Then the set
C=2 2
2 3
; 8 3
3 1
has the right size to be a basis of Z. Let's see if it is a linearly independent set. The relation of linear
dependence
a12 2
2 3
+a2 8 3
3 1
=O
2a1 8a22a1+ 3a2
2a1+ 3a2 3a1+a2
=0 0
0 0
leads to the homogeneous system of equations whose coecient matrix
2
6642 8
2 3
2 3
3 13
775
row-reduces to2
66410
01
0 0
0 03
775
So witha1=a2= 0 as the only solution, the set is linearly independent. Now we can apply Theorem G
[407] to see that Calso spansZand therefore is a second basis for Z.
Example SVP4
Sets of vectors in P4
In Example BSP4 [372] we showed that
B=
x 2; x2 4x+ 4; x3 6x2+ 12x 8; x4 8x3+ 24x2 32x+ 16
Version 2.30
412 Section PD Properties of Dimension
is a basis for W=fp(x)jp2P4; p(2) = 0g. So dim (W) = 4.
The set
3x2 5x 2;2x2 7x+ 6; x3 2x2+x 2
is a subset of W(check this) and it happens to be linearly independent (check this, too). However, by
Theorem G [407] it cannot span W.
The set
3x2 5x 2;2x2 7x+ 6; x3 2x2+x 2; x4+ 2x3+ 5x2 10x; x4 16
is another subset of W(check this) and Theorem G [407] tells us that it must be linearly dependent.
The set
x 2; x2 2x; x3 2x2; x4 2x3
is a third subset of W(check this) and is linearly independent (check this). Since it has the right size to
be a basis, and is linearly independent, Theorem G [407] tells us that it also spans W, and therefore is a
basis ofW.
A simple consequence of Theorem G [407] is the observation that proper subspaces have strictly smaller
dimensions. Hopefully this may seem intuitively obvious, but it still requires proof, and we will cite this
result later.
Theorem PSSD
Proper Subspaces have Smaller Dimension
Suppose that UandVare subspaces of the vector space W, such that U(V. Then dim ( U)<dim (V).
Proof Suppose that dim ( U) =mand dim (V) =t. ThenUhas a basis Bof sizem. Ifm>t , then by
Theorem G [407], Bis linearly dependent, which is a contradiction. If m=t, then by Theorem G [407],
BspansV. ThenU=hBi=V, also a contradiction. All that remains is that m<t , which is the desired
conclusion.
The nal theorem of this subsection is an extremely powerful tool for establishing the equality of two
sets that are subspaces. Notice that the hypotheses include the equality of two integers (dimensions) while
the conclusion is the equality of two sets (subspaces). It is the extra \structure" of a vector space and its
dimension that makes possible this huge leap from an integer equality to a set equality.
Theorem EDYES
Equal Dimensions Yields Equal Subspaces
Suppose that UandVare subspaces of the vector space W, such that UVand dim (U) = dim (V).
ThenU=V.
Proof We give a proof by contradiction (Technique CD [770]). Suppose to the contrary that U6=V.
SinceUV, there must be a vector vsuch that v2Vandv62U. LetB=fu1;u2;u3; :::; utgbe a
basis forU. Then, by Theorem ELIS [407], the set C=B[fvg=fu1;u2;u3; :::; ut;vgis a linearly
independent set of t+ 1 vectors in V. However, by hypothesis, Vhas the same dimension as U(namelyt)
and therefore Theorem G [407] says that Cis too big to be linearly independent. This contradiction shows
thatU=V.
Subsection RT
Ranks and Transposes
We now prove one of the most surprising theorems about matrices. Notice the paucity of hypotheses
compared to the precision of the conclusion.
Version 2.30
Subsection PD.RT Ranks and Transposes 413
Theorem RMRT
Rank of a Matrix is the Rank of the Transpose
SupposeAis anmnmatrix. Then r(A) =r
At
.
Proof Suppose we row-reduce Ato the matrix Bin reduced row-echelon form, and Bhasrnon-zero
rows. The quantity rtells us three things about B: the number of leading 1's, the number of non-zero
rows and the number of pivot columns. For this proof we will be interested in the latter two.
Theorem BRS [280] and Theorem BCS [274] each has a conclusion that provides a basis, for the row
space and the column space, respectively. In each case, these bases contain rvectors. This observation
makes the following go.
r(A) = dim (C(A)) Denition ROM [397]
=r Theorem BCS [274]
= dim (R(A)) Theorem BRS [280]
= dim
C
At
Theorem CSRST [282]
=r
At
Denition ROM [397]
Jacob Linenthal helped with this proof.
This says that the row space and the column space of a matrix have the same dimension, which should
be very surprising. It does notsay that column space and the row space are identical. Indeed, if the matrix
is not square, then the sizes (number of slots) of the vectors in each space are dierent, so the sets are not
even comparable.
It is not hard to construct by yourself examples of matrices that illustrate Theorem RMRT [411], since
it applies equally well to anymatrix. Grab a matrix, row-reduce it, count the nonzero rows or the leading
1's. That's the rank. Transpose the matrix, row-reduce that, count the nonzero rows or the leading 1's.
That's the rank of the transpose. The theorem says the two will be equal. Here's an example anyway.
Example RRTI
Rank, rank of transpose, Archetype I
Archetype I [816] has a 4 7 coecient matrix which row-reduces to
2
66414 0 0 2 1 3
0 0 10 1 3 5
0 0 0 12 6 6
0 0 0 0 0 0 03
775
so the rank is 3. Row-reducing the transpose yields
2
666666666410 0 31
7
01012
7
0 0 113
7
0 0 0 0
0 0 0 0
0 0 0 0
0 0 0 03
7777777775:
demonstrating that the rank of the transpose is also 3.
Version 2.30
414 Section PD Properties of Dimension
Subsection DFS
Dimension of Four Subspaces
That the rank of a matrix equals the rank of its transpose is a fundamental and surprising result. However,
applying Theorem FS [299] we can easily determine the dimension of all four fundamental subspaces
associated with a matrix.
Theorem DFS
Dimensions of Four Subspaces
Suppose that Ais anmnmatrix, and Bis a row-equivalent matrix in reduced row-echelon form with r
nonzero rows. Then
1. dim (N(A)) =n r
2. dim (C(A)) =r
3. dim (R(A)) =r
4. dim (L(A)) =m r
Proof IfArow-reduces to a matrix in reduced row-echelon form with rnonzero rows, then the matrix C
of extended echelon form (Denition EEF [297]) will be an rnmatrix in reduced row-echelon form with
no zero rows and rpivot columns (Theorem PEEF [298]). Similarly, the matrix Lof extended echelon
form (Denition EEF [297]) will be an m rmmatrix in reduced row-echelon form with no zero rows
andm rpivot columns (Theorem PEEF [298]).
dim (N(A)) = dim (N(C)) Theorem FS [299]
=n r Theorem BNS [160]
dim (C(A)) = dim (N(L)) Theorem FS [299]
=m (m r) Theorem BNS [160]
=r
dim (R(A)) = dim (R(C)) Theorem FS [299]
=r Theorem BRS [280]
dim (L(A)) = dim (R(L)) Theorem FS [299]
=m r Theorem BRS [280]
There are many dierent ways to state and prove this result, and indeed, the equality of the dimensions
of the column space and row space is just a slight expansion of Theorem RMRT [411]. However, we
have restricted our techniques to applying Theorem FS [299] and then determining dimensions with bases
provided by Theorem BNS [160] and Theorem BRS [280]. This provides an appealing symmetry to the
results and the proof.
Version 2.30
Subsection PD.DS Direct Sums 415
Subsection DS
Direct Sums
Some of the more advanced ideas in linear algebra are closely related to decomposing (Technique DC [772])
vector spaces into direct sums of subspaces. With our previous results about bases and dimension, now
is the right time to state and collect a few results about direct sums, though we will only mention these
results in passing until we get to Section NLT [685], where they will get a heavy workout.
A direct sum is a short-hand way to describe the relationship between a vector space and two, or more,
of its subspaces. As we will use it, it is not a way to construct new vector spaces from others.
Denition DS
Direct Sum
Suppose that Vis a vector space with two subspaces UandWsuch that for every v2V,
1. There exists vectors u2U,w2Wsuch that v=u+w
2. Ifv=u1+w1andv=u2+w2where u1;u22U,w1;w22Wthenu1=u2andw1=w2.
ThenVis the direct sum ofUandWand we write V=UW.
(This denition contains Notation DS.) 4
Informally, when we say Vis the direct sum of the subspaces UandW, we are saying that each vector
ofVcan always be expressed as the sum of a vector from Uand a vector from W, and this expression
can only be accomplished in one way (i.e. uniquely). This statement should begin to feel something like
our denitions of nonsingular matrices (Denition NM [83]) and linear independence (Denition LI [351]).
It should not be hard to imagine the natural extension of this denition to the case of more than two
subspaces. Could you provide a careful denition of V=U1U2U3:::Um(Exercise PD.M50 [418])?
Example SDS
Simple direct sum
InC3, dene
v1=2
43
2
53
5 v2=2
4 1
2
13
5 v3=2
42
1
23
5
ThenC3=hfv1;v2gihf v3gi. This statement derives from the fact that B=fv1;v2;v3gis basis for
C3. The spanning property of Byields the decomposition of any vector into a sum of vectors from the two
subspaces, and the linear independence of Byields the uniqueness of the decomposition. We will illustrate
these claims with a numerical example.
Choose v=2
410
1
63
5. Then
v= 2v1+ ( 2)v2+ 1v3= (2v1+ ( 2)v2) + (1 v3)
where we have added parentheses for emphasis. Obviously 1 v32hfv3gi, while 2 v1+ ( 2)v22hfv1;v2gi.
Theorem VRRB [360] provides the uniqueness of the scalars in these linear combinations.
Example SDS [413] is easy to generalize into a theorem.
Theorem DSFB
Direct Sum From a Basis
Suppose that Vis a vector space with a basis B=fv1;v2;v3; :::; vngandmn. Dene
U=hfv1;v2;v3; :::; vmgi W=hfvm+1;vm+2;vm+3; :::; vngi
Version 2.30
416 Section PD Properties of Dimension
ThenV=UW.
Proof Choose any vector v2V. Then by Theorem VRRB [360] there are unique scalars, a1; a2; a3; :::; an
such that
v=a1v1+a2v2+a3v3++anvn
= (a1v1+a2v2+a3v3++amvm) +
(am+1vm+1+am+2vm+2+am+3vm+3++anvn)
=u+w
where we have implicitly dened uandwin the last line. It should be clear that u2U, and similarly,
w2W(and not simply by the choice of their names).
Suppose we had another decomposition of v, say v=u+w. Then we could write uas a linear
combination of v1through vm, say using scalars b1; b2; b3; :::; bm. And we could write was a linear
combination of vm+1through vn, say using scalars c1; c2; c3; :::; cn m. These two collections of scalars
would then together give a linear combination of v1through vnthat equals v. By the uniqueness of
a1; a2; a3; :::; an,ai=bifor 1imandam+i=cifor 1in m. From the equality of these
scalars we conclude that u=uandw=w. So with both conditions of Denition DS [413] fullled we
see thatV=UW.
Given one subspace of a vector space, we can always nd another subspace that will pair with the rst
to form a direct sum. The main idea of this theorem, and its proof, is the idea of extending a linearly
independent subset into a basis with repeated applications of Theorem ELIS [407].
Theorem DSFOS
Direct Sum From One Subspace
Suppose that Uis a subspace of the vector space V. Then there exists a subspace WofVsuch that
V=UW.
Proof IfU=V, then choose W=f0g. Otherwise, choose a basis B=fv1;v2;v3; :::; vmgforU.
Then since Bis a linearly independent set, Theorem ELIS [407] tells us there is a vector vm+1inV, but
not inU, such that B[fvm+1gis linearly independent. Dene the subspace U1=hB[fvm+1gi.
We can repeat this procedure, in the case were U16=V, creating a new vector vm+2inV, but
not inU1, and a new subspace U2=hB[fvm+1;vm+2gi. If we continue repeating this procedure,
eventually, Uk=Vfor somek, and we can no longer apply Theorem ELIS [407]. No matter, in this case
B[fvm+1;vm+2; :::; vm+kgis a linearly independent set that spans V, i.e. a basis for V.
DeneW=hfvm+1;vm+2; :::; vm+kgi. We now are exactly in position to apply Theorem DSFB [413]
and see that V=UW.
There are several dierent ways to dene a direct sum. Our next two theorems give equivalences
(Technique E [768]) for direct sums, and therefore could have been employed as denitions. The rst
should further cement the notion that a direct sum has some connection with linear independence.
Theorem DSZV
Direct Sums and Zero Vectors
SupposeUandWare subspaces of the vector space V. ThenV=UWif and only if
1. For every v2V, there exists vectors u2U,w2Wsuch that v=u+w.
2. Whenever 0=u+wwithu2U,w2Wthenu=w=0.
Proof The rst condition is identical in the denition and the theorem, so we only need to establish the
equivalence of the second conditions.
Version 2.30
Subsection PD.DS Direct Sums 417
()) Assume that V=UW, according to Denition DS [413]. By Property Z [318], 02Vand
0=0+0. If we also assume that 0=u+w, then the uniqueness of the decomposition gives u=0and
w=0.
(() Suppose that v2V,v=u1+w1andv=u2+w2where u1;u22U,w1;w22W. Then
0=v v Property AI [318]
= (u1+w1) (u2+w2)
= (u1 u2) + (w1 w2) Property AA [317]
By Property AC [317], u1 u22Uandw1 w22W. We can now apply our hypothesis, the second
statement of the theorem, to conclude that
u1 u2=0 w 1 w2=0
u1=u2 w1=w2
which establishes the uniqueness needed for the second condition of the denition.
Our second equivalence lends further credence to calling a direct sum a decomposition. The two
subspaces of a direct sum have no (nontrivial) elements in common.
Theorem DSZI
Direct Sums and Zero Intersection
SupposeUandWare subspaces of the vector space V. ThenV=UWif and only if
1. For every v2V, there exists vectors u2U,w2Wsuch that v=u+w.
2.U\W=f0g.
Proof The rst condition is identical in the denition and the theorem, so we only need to establish the
equivalence of the second conditions.
()) Assume that V=UW, according to Denition DS [413]. By Property Z [318] and Denition
SI [763],f0gU\W. To establish the opposite inclusion, suppose that x2U\W. Then, since xis an
element of both UandW, we can write two decompositions of xas a vector from Uplus a vector from W,
x=x+0 x =0+x
By the uniqueness of the decomposition, we see (twice) that x=0andU\Wf0g. Applying Denition
SE [762], we have U\W=f0g.
(() Assume that U\W=f0g. And assume further that v2Vis such that v=u1+w1and
v=u2+w2where u1;u22U,w1;w22W. Dene x=u1 u2. then by Property AC [317], x2U.
Also
x=u1 u2
= (v w1) (v w2)
= (v v) (w1 w2)
=w2 w1
Sox2Wby Property AC [317]. Thus, x2U\W=f0g(Denition SI [763]). So x=0and
u1 u2=0 w 2 w1=0
u1=u2 w2=w1
Version 2.30
418 Section PD Properties of Dimension
yielding the desired uniqueness of the second condition of the denition.
If the statement of Theorem DSZV [414] did not remind you of linear independence, the next theorem
should establish the connection.
Theorem DSLI
Direct Sums and Linear Independence
SupposeUandWare subspaces of the vector space VwithV=UW. Suppose that Ris a linearly
independent subset of UandSis a linearly independent subset of W. ThenR[Sis a linearly independent
subset ofV.
Proof LetR=fu1;u2;u3; :::; ukgandS=fw1;w2;w3; :::; w`g. Begin with a relation of linear
dependence (Denition RLD [351]) on the set R[Susing scalars a1; a2; a3; :::; akandb1; b2; b3; :::; b`.
Then,
0=a1u1+a2u2+a3u3++akuk+b1w1+b2w2+b3w3++b`w`
= (a1u1+a2u2+a3u3++akuk) + (b1w1+b2w2+b3w3++b`w`)
=u+w
where we have made an implicit denition of the vectors u2U,w2W. Applying Theorem DSZV [414]
we conclude that
u=a1u1+a2u2+a3u3++akuk=0
w=b1w1+b2w2+b3w3++b`w`=0
Now the linear independence of RandS(individually) yields
a1=a2=a3==ak= 0 b1=b2=b3==b`= 0
Forced to acknowledge that only a trivial linear combination yields the zero vector, Denition LI [351] says
the setR[Sis linearly independent in V.
Our last theorem in this collection will go some ways towards explaining the word \sum" in the moniker
\direct sum," while also partially explaining why these results appear in a section devoted to a discussion
of dimension.
Theorem DSD
Direct Sums and Dimension
SupposeUandWare subspaces of the vector space VwithV=UW. Then dim ( V) = dim (U)+dim (W).
Proof We will establish this equality of positive integers with two inequalities. We will need a basis of
U(call itB) and a basis of W(call itC).
First, note that BandChave sizes equal to the dimensions of the respective subspaces. The union
of these two linearly independent sets, B[Cwill be linearly independent in Vby Theorem DSLI [416].
Further, the two bases have no vectors in common by Theorem DSZI [415], since B\Cf0gand the zero
vector is never an element of a linearly independent set (Exercise LI.T10 [166]). So the size of the union is
exactly the sum of the dimensions of UandW. By Theorem G [407] the size of B[Ccannot exceed the
dimension of Vwithout being linearly dependent. These observations give us dim ( U)+dim (W)dim (V).
Grab any vector v2V. Then by Theorem DSZI [415] we can write v=u+wwithu2Uandw2W.
Individually, we can write uas a linear combination of the basis elements in B, and similarly, we can write
was a linear combination of the basis elements in C, since the bases are spanning sets for their respective
subspaces. These two sets of scalars will provide a linear combination of all of the vectors in B[Cwhich
Version 2.30
Subsection PD.READ Reading Questions 419
will equal v. The upshot of this is that B[Cis a spanning set for V. By Theorem G [407], the size of
B[Ccannot be smaller than the dimension of Vwithout failing to span V. These observations give us
dim (U) + dim (W)dim (V).
There is a certain appealling symmetry in the previous proof, where both linear independence and
spanning properties of the bases are used, both of the rst two conclusions of Theorem G [407] are employed,
and we have quoted both of the two conditions of Theorem DSZI [415].
One nal theorem tells us that we can successively decompose direct sums into sums of smaller and
smaller subspaces.
Theorem RDS
Repeated Direct Sums
SupposeVis a vector space with subspaces UandWwithV=UW. Suppose that XandYare
subspaces of WwithW=XY. ThenV=UXY.
Proof Suppose that v2V. Then due to V=UW, there exist vectors u2Uandw2Wsuch that
v=u+w. Due toW=XY, there exist vectors x2Xandy2Ysuch that w=x+y. All together,
v=u+w=u+x+y
which would be the rst condition of a denition of a 3-way direct product. Now consider the uniqueness.
Suppose that
v=u1+x1+y1 v=u2+x2+y2
Because x1+y12W,x2+y22W, andV=UW, we conclude that
u1=u2 x1+y1=x2+y2
From the second equality, an application of W=XYyields the conclusions x1=x2andy1=y2. This
establishes the uniqueness of the decomposition of vinto a sum of vectors from U,XandY.
Remember that when we write V=UWthere always needs to be a \superspace," in this case V. The
statementUWis meaningless. Writing V=UWis simply a shorthand for a somewhat complicated
relationship between V,UandW, as described in the two conditions of Denition DS [413], or Theorem
DSZV [414], or Theorem DSZI [415]. Theorem DSFB [413] and Theorem DSFOS [414] gives us sure-re
ways to build direct sums, while Theorem DSLI [416], Theorem DSD [416] and Theorem RDS [417] tell us
interesting properties of direct sums. This subsection has been long on theorems and short on examples.
If we were to use the term \lemma" we might have chosen to label some of these results as such, since
they will be important tools in other proofs, but may not have much interest on their own (see Technique
LC [774]). We will be referencing these results heavily in later sections, and will remind you then to come
back for a second look.
Subsection READ
Reading Questions
1. Why does Theorem G [407] have the title it does?
2. What is so surprising about Theorem RMRT [411]?
3. Row-reduce the matrix Ato reduced row-echelon form. Without any further computations, compute
the dimensions of the four subspaces, N(A),C(A),R(A) andL(A).
A=2
6641 1 2 8 5
1 1 1 4 1
0 2 3 8 6
2 0 1 8 43
775
Version 2.30
420 Section PD Properties of Dimension
Subsection EXC
Exercises
C10 Example SVP4 [409] leaves several details for the reader to check. Verify these ve claims.
Contributed by Robert Beezer
C40 Determine if the set T=
x2 x+ 5;4x3 x2+ 5x;3x+ 2
spans the vector space of polynomials
with degree 4 or less, P4. (Compare the solution to this exercise with Solution LISS.C40 [367].)
Contributed by Robert Beezer Solution [419]
M50 Mimic Denition DS [413] and construct a reasonable denition of V=U1U2U3:::Um.
Contributed by Robert Beezer
T05 Trivially, if UandVare two subspaces of W, then dim ( U) = dim (V). Combine this fact, Theorem
PSSD [410], and Theorem EDYES [410] all into one grand combined theorem. You might look to Theorem
PIP [196] stylistic inspiration. (Notice this problem does not ask you to prove anything. It just asks you
to roll up three theorems into one compact, logically equivalent statement.)
Contributed by Robert Beezer
T10 Prove the following theorem, which could be viewed as a reformulation of parts (3) and (4) of
Theorem G [407], or more appropriately as a corollary of Theorem G [407] (Technique LC [774]).
SupposeVis a vector space and Sis a subset of Vsuch that the number of vectors in Sequals the
dimension of V. ThenSis linearly independent if and only if SspansV.
Contributed by Robert Beezer
T15 Suppose that Ais anmnmatrix and let min( m;n) denote the minimum of mandn. Prove that
r(A)min(m;n).
Contributed by Robert Beezer
T20 Suppose that Ais anmnmatrix and b2Cm. Prove that the linear system LS(A;b) is consistent
if and only if r(A) =r([Ajb]).
Contributed by Robert Beezer Solution [419]
T25 Suppose that Vis a vector space with nite dimension. Let Wbe any subspace of V. Prove that
Whas nite dimension.
Contributed by Robert Beezer
T33 Part of Exercise B.T50 [383] is the half of the proof where we assume the matrix Ais nonsingular
and prove that a set is basis. In Solution B.T50 [387] we proved directly that the set was both linearly
independent and a spanning set. Shorten this part of the proof by applying Theorem G [407]. Be careful,
there is one subtlety.
Contributed by Robert Beezer Solution [419]
T60 Suppose that Wis a vector space with dimension 5, and UandVare subspaces of W, each of
dimension 3. Prove that U\Vcontains a non-zero vector. State a more general result.
Contributed by Joe Riegsecker Solution [419]
Version 2.30
Subsection PD.SOL Solutions 421
Subsection SOL
Solutions
C40 Contributed by Robert Beezer Statement [418]
The vector space P4has dimension 5 by Theorem DP [395]. Since Tcontains only 3 vectors, and 3 <5,
Theorem G [407] tells us that Tdoes not span P5.
T20 Contributed by Robert Beezer Statement [418]
()) Suppose rst that LS(A;b) is consistent. Then by Theorem CSCS [272], b2C(A). This means that
C(A) =C([Ajb]) and so it follows that r(A) =r([Ajb]).
(() Adding a column to a matrix will only increase the size of its column space, so in all cases,
C(A)C([Ajb]). However, if we assume that r(A) =r([Ajb]), then by Theorem EDYES [410] we
conclude thatC(A) =C([Ajb]). Then b2C([Ajb]) =C(A) so by Theorem CSCS [272], LS(A;b) is
consistent.
T33 Contributed by Robert Beezer Statement [418]
By Theorem DCM [395] we know that Cnhas dimension n. So by Theorem G [407] we need only establish
that the set Cis linearly independent or a spanning set. However, the hypotheses also require that C
be of sizen. We assumed that B=fx1;x2;x3; :::; xnghad sizen, but there is no guarantee that
C=fAx1; Ax2; Ax3; :::; A xngwill have size n. There could be some \collapsing" or \collisions."
Suppose we establish that Cis linearly independent. Then Cmust havendistinct elements or else we
could fashion a nontrivial relation of linear dependence involving duplicate elements.
If we instead to choose to prove that Cis a spanning set, then we could establish the uniqueness of the
elements of Cquite easily. Suppose that Axi=Axj. Then
A(xi xj) =Axi Axj=0
SinceAis nonsingular, we conclude that xi xj=0, orxi=xj, contrary to our description of B.
T60 Contributed by Robert Beezer Statement [418]
Letfu1;u2;u3gandfv1;v2;v3gbe bases for UandV(respectively). Then, the set fu1;u2;u3;v1;v2;v3g
is linearly dependent, since Theorem G [407] says we cannot have 6 linearly independent vectors in a vector
space of dimension 5. So we can assert that there is a non-trivial relation of linear dependence,
a1u1+a2u2+a3u3+b1v1+b2v2+b3v3=0
wherea1; a2; a3andb1; b2; b3are not all zero.
We can rearrange this equation as
a1u1+a2u2+a3u3= b1v1 b2v2 b3v3
This is an equality of two vectors, so we can give this common vector a name, say w,
w=a1u1+a2u2+a3u3= b1v1 b2v2 b3v3
This is the desired non-zero vector, as we will now show.
First, since w=a1u1+a2u2+a3u3, we can see that w2U. Similarly, w= b1v1 b2v2 b3v3, so
w2V. This establishes that w2U\V(Denition SI [763]).
Isw6=0? Suppose not, in other words, suppose w=0. Then
0=w=a1u1+a2u2+a3u3
Becausefu1;u2;u3gis a basis for U, it is a linearly independent set and the relation of linear dependence
above means we must conclude that a1=a2=a3= 0. By a similar process, we would conclude that
Version 2.30
422 Section PD Properties of Dimension
b1=b2=b3= 0. But this is a contradiction since a1; a2; a3; b1; b2; b3were chosen so that some were
nonzero. So w6=0.
How does this generalize? All we really needed was the original relation of linear dependence that
resulted because we had \too many" vectors in W. A more general statement would be: Suppose that W
is a vector space with dimension n,Uis a subspace of dimension pandVis a subspace of dimension q. If
p+q>n , thenU\Vcontains a non-zero vector.
Version 2.30
Annotated Acronyms PD.VS Vector Spaces 423
Annotated Acronyms VS
Vector Spaces
Denition VS [317]
The most fundamental object in linear algebra is a vector space. Or else the most fundamental object is
a vector, and a vector space is important because it is a collection of vectors. Either way, Denition VS
[317] is critical. All of our remaining theorems that assume we are working with a vector space can trace
their lineage back to this denition.
Theorem TSS [334]
Check all ten properties of a vector space (Denition VS [317]) can get tedious. But if you have a subset
of a known vector space, then Theorem TSS [334] considerably shortens the verication. Also, proofs of
closure (the last two conditions in Theorem TSS [334]) are a good way to practice a common style of proof.
Theorem VRRB [360]
The proof of uniqueness in this theorem is a very typical employment of the hypothesis of linear inde-
pendence. But that's not why we mention it here. This theorem is critical to our rst section about
representations, Section VR [603], via Denition VR [603].
Theorem CNMB [376]
Having just dened a basis (Denition B [371]) we discover that the columns of a nonsingular matrix form
a basis of Cm. Much of what we know about nonsingular matrices is either contained in this statement, or
much more evident because of it.
Theorem SSLD [391]
This theorem is a key juncture in our development of linear algebra. You have probably already realized
how useful Theorem G [407] is. All four parts of Theorem G [407] have proofs that nish with an application
of Theorem SSLD [391].
Theorem RPNC [398]
This simple relationship between the rank, nullity and number of columns of a matrix might be surprising.
But in simplicity comes power, as this theorem can be very useful. It will be generalized in the very last
theorem of Chapter LT [515], Theorem RPNDD [588].
Theorem G [407]
A whimsical title, but the intent is to make sure you don't miss this one. Much of the interaction between
bases, dimension, linear independence and spanning is captured in this theorem.
Theorem RMRT [411]
This one is a real surprise. Why should a matrix, and its transpose, both row-reduce to the same number
of non-zero rows?
Version 2.30
424 Section PD Properties of Dimension
Version 2.30
Chapter D
Determinants
The determinant is a function that takes a square matrix as an input and produces a scalar as an output.
So unlike a vector space, it is not an algebraic structure. However, it has many benecial properties for
studying vector spaces, matrices and systems of equations, so it is hard to ignore (though some have tried).
While the properties of a determinant can be very useful, they are also complicated to prove.
Section DM
Determinant of a Matrix
First, a slight detour, as we introduce elementary matrices, which will bring us back to the beginning of
the course and our old friend, row operations.
Subsection EM
Elementary Matrices
Elementary matrices are very simple, as you might have suspected from their name. Their purpose is
to eect row operations (Denition RO [31]) on a matrix through matrix multiplication (Denition MM
[226]). Their denitions look more complicated than they really are, so be sure to read ahead after you
read the denition for some explanations and an example.
Denition ELEM
Elementary Matrices
1. Fori6=j,Ei;jis the square matrix of size nwith
[Ei;j]k`=8
>>>>>>>>><
>>>>>>>>>:0k6=i;k6=j;`6=k
1k6=i;k6=j;`=k
0k=i;`6=j
1k=i;`=j
0k=j;`6=i
1k=j;`=i
425
426 Section DM Determinant of a Matrix
2. For6= 0,Ei() is the square matrix of size nwith
[Ei()]k`=8
><
>:0k6=i;`6=k
1k6=i;`=k
k =i;`=i
3. Fori6=j,Ei;j() is the square matrix of size nwith
[Ei;j()]k`=8
>>>>>><
>>>>>>:0k6=j;`6=k
1k6=j;`=k
0k=j;`6=i;`6=j
1k=j;`=j
k =j;`=i
(This denition contains Notation ELEM.)
4
Again, these matrices are not as complicated as they appear, since they are mostly perturbations of
thennidentity matrix (Denition IM [84]). Ei;jis the identity matrix with rows (or columns) iand
jtrading places, Ei() is the identity matrix where the diagonal entry in row iand column ihas been
replaced by , andEi;j() is the identity matrix where the entry in row jand column ihas been replaced
by. (Yes, those subscripts look backwards in the description of Ei;j()). Notice that our notation makes
no reference to the size of the elementary matrix, since this will always be apparent from the context, or
unimportant.
The raison d'^ etre for elementary matrices is to \do" row operations on matrices with matrix multi-
plication. So here is an example where we will both see some elementary matrices and see how they can
accomplish row operations.
Example EMRO
Elementary matrices and row operations
We will perform a sequence of row operations (Denition RO [31]) on the 3 4 matrixA, while also
multiplying the matrix on the left by the appropriate 3 3 elementary matrix.
A=2
42 1 3 1
1 3 2 4
5 0 3 13
5
R1$R3:2
45 0 3 1
1 3 2 4
2 1 3 13
5 E1;3:2
40 0 1
0 1 0
1 0 03
52
42 1 3 1
1 3 2 4
5 0 3 13
5=2
45 0 3 1
1 3 2 4
2 1 3 13
5
2R2:2
45 0 3 1
2 6 4 8
2 1 3 13
5 E2(2) :2
41 0 0
0 2 0
0 0 13
52
45 0 3 1
1 3 2 4
2 1 3 13
5=2
45 0 3 1
2 6 4 8
2 1 3 13
5
2R3+R1:2
49 2 9 3
2 6 4 8
2 1 3 13
5E3;1(2) :2
41 0 2
0 1 0
0 0 13
52
45 0 3 1
2 6 4 8
2 1 3 13
5=2
49 2 9 3
2 6 4 8
2 1 3 13
5
The next three theorems establish that each elementary matrix eects a row operation via matrix
multiplication.
Version 2.30
Subsection DM.EM Elementary Matrices 427
Theorem EMDRO
Elementary Matrices Do Row Operations
Suppose that Ais anmnmatrix, and Bis a matrix of the same size that is obtained from Aby a single
row operation (Denition RO [31]). Then there is an elementary matrix of size mthat will convert Ato
Bvia matrix multiplication on the left. More precisely,
1. If the row operation swaps rows iandj, thenB=Ei;jA.
2. If the row operation multiplies row iby, thenB=Ei()A.
3. If the row operation multiplies row ibyand adds the result to row j, thenB=Ei;j()A.
Proof In each of the three conclusions, performing the row operation on Awill create the matrix B
where only one or two rows will have changed. So we will establish the equality of the matrix entries row
by row, rst for the unchanged rows, then for the changed rows, showing in each case that the result of
the matrix product is the same as the result of the row operation. Here we go.
Rowkof the product Ei;jA, wherek6=i,k6=j, is unchanged from A,
[Ei;jA]k`=nX
p=1[Ei;j]kp[A]p` Theorem EMP [227]
= [Ei;j]kk[A]k`+nX
p=1
p6=k[Ei;j]kp[A]p` Property CACN [758]
= 1 [A]k`+nX
p=1
p6=k0 [A]p` Denition ELEM [423]
= [A]k`
Rowiof the product Ei;jAis rowjofA,
[Ei;jA]i`=nX
p=1[Ei;j]ip[A]p` Theorem EMP [227]
= [Ei;j]ij[A]j`+nX
p=1
p6=j[Ei;j]ip[A]p` Property CACN [758]
= 1 [A]j`+nX
p=1
p6=j0 [A]p` Denition ELEM [423]
= [A]j`
Rowjof the product Ei;jAis rowiofA,
[Ei;jA]j`=nX
p=1[Ei;j]jp[A]p` Theorem EMP [227]
= [Ei;j]ji[A]i`+nX
p=1
p6=i[Ei;j]jp[A]p` Property CACN [758]
Version 2.30
428 Section DM Determinant of a Matrix
= 1 [A]i`+nX
p=1
p6=i0 [A]p` Denition ELEM [423]
= [A]i`
So the matrix product Ei;jAis the same as the row operation that swaps rows iandj.
Rowkof the product Ei()A, wherek6=i, is unchanged from A,
[Ei()A]k`=nX
p=1[Ei()]kp[A]p` Theorem EMP [227]
= [Ei()]kk[A]k`+nX
p=1
p6=k[Ei()]kp[A]p` Property CACN [758]
= 1 [A]k`+nX
p=1
p6=k0 [A]p` Denition ELEM [423]
= [A]k`
Rowiof the product Ei()Aistimes rowiofA,
[Ei()A]i`=nX
p=1[Ei()]ip[A]p` Theorem EMP [227]
= [Ei()]ii[A]i`+nX
p=1
p6=i[Ei()]ip[A]p` Property CACN [758]
=[A]i`+nX
p=1
p6=i0 [A]p` Denition ELEM [423]
=[A]i`
So the matrix product Ei()Ais the same as the row operation that swaps multiplies row iby.
Rowkof the product Ei;j()A, wherek6=j, is unchanged from A,
[Ei;j()A]k`=nX
p=1[Ei;j()]kp[A]p` Theorem EMP [227]
= [Ei;j()]kk[A]k`+nX
p=1
p6=k[Ei;j()]kp[A]p` Property CACN [758]
= 1 [A]k`+nX
p=1
p6=k0 [A]p` Denition ELEM [423]
= [A]k`
Rowjof the product Ei;j()A, istimes rowiofAand then added to row jofA,
[Ei;j()A]j`=nX
p=1[Ei;j()]jp[A]p` Theorem EMP [227]
Version 2.30
Subsection DM.DD Denition of the Determinant 429
= [Ei;j()]jj[A]j`+
[Ei;j()]ji[A]i`+nX
p=1
p6=j;i[Ei;j()]jp[A]p` Property CACN [758]
= 1 [A]j`+[A]i`+nX
p=1
p6=j;i0 [A]p` Denition ELEM [423]
= [A]j`+[A]i`
So the matrix product Ei;j()Ais the same as the row operation that multiplies row ibyand adds the
result to row j.
Later in this section we will need two facts about elementary matrices.
Theorem EMN
Elementary Matrices are Nonsingular
IfEis an elementary matrix, then Eis nonsingular.
Proof We show that we can row-reduce each elementary matrix to the identity matrix. Given an
elementary matrix of the form Ei;j, perform the row operation that swaps row jwith rowi. Given an
elementary matrix of the form Ei(), with6= 0, perform the row operation that multiplies row iby 1=.
Given an elementary matrix of the form Ei;j(), with6= 0, perform the row operation that multiplies
rowiby and adds it to row j. In each case, the result of the single row operation is the identity
matrix. So each elementary matrix is row-equivalent to the identity matrix, and by Theorem NMRRI [84]
is nonsingular.
Notice that we have now made use of the nonzero restriction on in the denition of Ei(). One more
key property of elementary matrices.
Theorem NMPEM
Nonsingular Matrices are Products of Elementary Matrices
Suppose that Ais a nonsingular matrix. Then there exists elementary matrices E1; E2; E3; :::; Etso that
A=E1E2E3:::Et.
Proof SinceAis nonsingular, it is row-equivalent to the identity matrix by Theorem NMRRI [84], so
there is a sequence of trow operations that converts ItoA. For each of these row operations, form the as-
sociated elementary matrix from Theorem EMDRO [425] and denote these matrices by E1; E2; E3; :::; Et.
Applying the rst row operation to Iyields the matrix E1I. The second row operation yields E2(E1I),
and the third row operation creates E3E2E1I. The result of the full sequence of trow operations will yield
A, so
A=Et:::E 3E2E1I=Et:::E 3E2E1
Other than the cosmetic matter of re-indexing these elementary matrices in the opposite order, this is the
desired result.
Subsection DD
Denition of the Determinant
We'll now turn to the denition of a determinant and do some sample computations. The denition of the
determinant function is recursive , that is, the determinant of a large matrix is dened in terms of the
determinant of smaller matrices. To this end, we will make a few denitions.
Version 2.30
430 Section DM Determinant of a Matrix
Denition SM
SubMatrix
Suppose that Ais anmnmatrix. Then the submatrix A(ijj) is the (m 1)(n 1) matrix obtained
fromAby removing row iand column j.
(This denition contains Notation SM.) 4
Example SS
Some submatrices
For the matrix
A=2
41 2 3 9
4 2 0 1
3 5 2 13
5
we have the submatrices
A(2j3) =1 2 9
3 5 1
A(3j1) = 2 3 9
2 0 1
Denition DM
Determinant of a Matrix
SupposeAis a square matrix. Then its determinant , det (A) =jAj, is an element of Cdened recursively
by:
IfAis a 11 matrix, then det ( A) = [A]11.
IfAis a matrix of size nwithn2, then
det (A) = [A]11det (A(1j1)) [A]12det (A(1j2)) + [A]13det (A(1j3))
[A]14det (A(1j4)) ++ ( 1)n+1[A]1ndet (A(1jn))
(This denition contains Notation DM.) 4
So to compute the determinant of a 5 5 matrix we must build 5 submatrices, each of size 4. To
compute the determinants of each the 4 4 matrices we need to create 4 submatrices each, these now of
size 3 and so on. To compute the determinant of a 10 10 matrix would require computing the determinant
of 10! = 1098765432 = 3;628;800 11 matrices. Fortunately there are better ways.
However this does suggest an excellent computer programming exercise to write a recursive procedure to
compute a determinant.
Let's compute the determinant of a reasonable sized matrix by hand.
Example D33M
Determinant of a 33matrix
Suppose that we have the 3 3 matrix
A=2
43 2 1
4 1 6
3 1 23
5
Then
det (A) =jAj=3 2 1
4 1 6
3 1 2
Version 2.30
Subsection DM.CD Computing Determinants 431
= 31 6
1 2 24 6
3 2+ ( 1)4 1
3 1
= 3
12 6 1
2
42 6 3
4 1 1 3
= 3 (1(2) 6( 1)) 2 (4(2) 6( 3)) (4( 1) 1( 3))
= 24 52 + 1
= 27
In practice it is a bit silly to decompose a 2 2 matrix down into a couple of 1 1 matrices and then
compute the exceedingly easy determinant of these puny matrices. So here is a simple theorem.
Theorem DMST
Determinant of Matrices of Size Two
Suppose that A=a b
c d
. Then det ( A) =ad bc
Proof Applying Denition DM [428],
a b
c d=ad bc=ad bc
Do you recall seeing the expression ad bcbefore? (Hint: Theorem TTMI [246])
Subsection CD
Computing Determinants
There are a variety of ways to compute the determinant. We will establish rst that we can choose to
mimic our denition of the determinant, but by using matrix entries and submatrices based on a row other
than the rst one.
Theorem DER
Determinant Expansion about Rows
Suppose that Ais a square matrix of size n. Then
det (A) = ( 1)i+1[A]i1det (A(ij1)) + ( 1)i+2[A]i2det (A(ij2))
+ ( 1)i+3[A]i3det (A(ij3)) ++ ( 1)i+n[A]indet (A(ijn)) 1in
which is known as expansion about rowi.
Proof First, the statement of the theorem coincides with Denition DM [428] when i= 1, so throughout,
we need only consider i>1.
Given the recursive denition of the determinant, it should be no surprise that we will use induction
for this proof (Technique I [772]). When n= 1, there is nothing to prove since there is but one row. When
n= 2, we just examine expansion about the second row,
( 1)2+1[A]21det (A(2j1)) + ( 1)2+2[A]22det (A(2j2))
= [A]21[A]12+ [A]22[A]11 Denition DM [428]
= [A]11[A]22 [A]12[A]21
= det (A) Theorem DMST [429]
Version 2.30
432 Section DM Determinant of a Matrix
So the theorem is true for matrices of size n= 1 andn= 2. Now assume the result is true for all matrices
of sizen 1 as we derive an expression for expansion about row ifor a matrix of size n. We will abuse
our notation for a submatrix slightly, so A(i1;i2jj1;j2) will denote the matrix formed by removing rows
i1andi2, along with removing columns j1andj2. Also, as we take a determinant of a submatrix, we will
need to \jump up" the index of summation partway through as we \skip over" a missing column. To do
this smoothly we will set
`j=(
0`<j
1`>j
Now,
det (A) =nX
j=1( 1)1+j[A]1jdet (A(1jj)) Denition DM [428]
=nX
j=1( 1)1+j[A]1jX
1`n
`6=j( 1)i 1+` `j[A]i`det (A(1;ijj;`)) Induction Hypothesis
=nX
j=1X
1`n
`6=j( 1)j+i+` `j[A]1j[A]i`det (A(1;ijj;`)) Property DCN [759]
=nX
`=1X
1jn
j6=`( 1)j+i+` `j[A]1j[A]i`det (A(1;ijj;`)) Property CACN [758]
=nX
`=1( 1)i+`[A]i`X
1jn
j6=`( 1)j `j[A]1jdet (A(1;ijj;`)) Property DCN [759]
=nX
`=1( 1)i+`[A]i`X
1jn
j6=`( 1)`j+j[A]1jdet (A(i;1j`;j)) 2 `jis even
=nX
`=1( 1)i+`[A]i`det (A(ij`)) Denition DM [428]
We can also obtain a formula that computes a determinant by expansion about a column, but this will
be simpler if we rst prove a result about the interplay of determinants and transposes. Notice how the
following proof makes use of the ability to compute a determinant by expanding about anyrow.
Theorem DT
Determinant of the Transpose
Suppose that Ais a square matrix. Then det
At
= det (A).
Proof With our denition of the determinant (Denition DM [428]) and theorems like Theorem DER
[429], using induction (Technique I [772]) is a natural approach to proving properties of determinants. And
so it is here. Let nbe the size of the matrix A, and we will use induction on n.
Forn= 1, the transpose of a matrix is identical to the original matrix, so vacuously, the determinants
are equal.
Now assume the result is true for matrices of size n 1. Then,
det
At
=1
nnX
i=1det
At
Version 2.30
Subsection DM.CD Computing Determinants 433
=1
nnX
i=1nX
j=1( 1)i+j
At
ijdet
At(ijj)
Theorem DER [429]
=1
nnX
i=1nX
j=1( 1)i+j[A]jidet
At(ijj)
Denition TM [210]
=1
nnX
i=1nX
j=1( 1)i+j[A]jidet
(A(jji))t
Denition TM [210]
=1
nnX
i=1nX
j=1( 1)i+j[A]jidet (A(jji)) Induction Hypothesis
=1
nnX
j=1nX
i=1( 1)j+i[A]jidet (A(jji)) Property CACN [758]
=1
nnX
j=1det (A) Theorem DER [429]
= det (A)
Now we can easily get the result that a determinant can be computed by expansion about any column
as well.
Theorem DEC
Determinant Expansion about Columns
Suppose that Ais a square matrix of size n. Then
det (A) = ( 1)1+j[A]1jdet (A(1jj)) + ( 1)2+j[A]2jdet (A(2jj))
+ ( 1)3+j[A]3jdet (A(3jj)) ++ ( 1)n+j[A]njdet (A(njj)) 1jn
which is known as expansion about column j.
Proof
det (A) = det
At
Theorem DT [430]
=nX
i=1( 1)j+i
At
jidet
At(jji)
Theorem DER [429]
=nX
i=1( 1)j+i
At
jidet
(A(ijj))t
Denition TM [210]
=nX
i=1( 1)j+i
At
jidet (A(ijj)) Theorem DT [430]
=nX
i=1( 1)i+j[A]ijdet (A(ijj)) Denition TM [210]
That the determinant of an nnmatrix can be computed in 2 ndierent (albeit similar) ways is
nothing short of remarkable. For the doubters among us, we will do an example, computing a 4 4 matrix
in two dierent ways.
Version 2.30
434 Section DM Determinant of a Matrix
Example TCSD
Two computations, same determinant
Let
A=2
664 2 3 0 1
9 2 0 1
1 3 2 1
4 1 2 63
775
Then expanding about the fourth row (Theorem DER [429] with i= 4) yields,
jAj= (4)( 1)4+13 0 1
2 0 1
3 2 1+ (1)( 1)4+2 2 0 1
9 0 1
1 2 1
+ (2)( 1)4+3 2 3 1
9 2 1
1 3 1+ (6)( 1)4+4 2 3 0
9 2 0
1 3 2
= ( 4)(10) + (1)( 22) + ( 2)(61) + 6(46) = 92
while expanding about column 3 (Theorem DEC [431] with j= 3) gives
jAj= (0)( 1)1+39 2 1
1 3 1
4 1 6+ (0)( 1)2+3 2 3 1
1 3 1
4 1 6+
( 2)( 1)3+3 2 3 1
9 2 1
4 1 6+ (2)( 1)4+3 2 3 1
9 2 1
1 3 1
= 0 + 0 + ( 2)( 107) + ( 2)(61) = 92
Notice how much easier the second computation was. By choosing to expand about the third column, we
have two entries that are zero, so two 3 3 determinants need not be computed at all!
When a matrix has all zeros above (or below) the diagonal, exploiting the zeros by expanding about
the proper row or column makes computing a determinant insanely easy.
Example DUTM
Determinant of an upper triangular matrix
Suppose that
T=2
666642 3 1 3 3
0 1 5 2 1
0 0 3 9 2
0 0 0 1 3
0 0 0 0 53
77775
We will compute the determinant of this 5 5 matrix by consistently expanding about the rst column for
each submatrix that arises and does not have a zero entry multiplying it.
det (T) =2 3 1 3 3
0 1 5 2 1
0 0 3 9 2
0 0 0 1 3
0 0 0 0 5
= 2( 1)1+1 1 5 2 1
0 3 9 2
0 0 1 3
0 0 0 5
Version 2.30
Subsection DM.READ Reading Questions 435
= 2( 1)( 1)1+13 9 2
0 1 3
0 0 5
= 2( 1)(3)( 1)1+1 1 3
0 5
= 2( 1)(3)( 1)( 1)1+15
= 2( 1)(3)( 1)(5) = 30
If you consult other texts in your study of determinants, you may run into the terms \minor" and
\cofactor," especially in a discussion centered on expansion about rows and columns. We've chosen not to
make these denitions formally since we've been able to get along without them. However, informally, a
minor is a determinant of a submatrix, specically det ( A(ijj)) and is usually referenced as the minor of
[A]ij. Acofactor is a signed minor, specically the cofactor of [ A]ijis ( 1)i+jdet (A(ijj)).
Subsection READ
Reading Questions
1. Construct the elementary matrix that will eect the row operation 6R2+R3on a 47 matrix.
2. Compute the determinant of the matrix
2
42 3 1
3 8 2
4 1 33
5
3. Compute the determinant of the matrix
2
666643 9 2 4 2
0 1 4 2 7
0 0 2 5 2
0 0 0 1 6
0 0 0 0 43
77775
Version 2.30
436 Section DM Determinant of a Matrix
Subsection EXC
Exercises
C21 Doing the computations by hand, nd the determinant of the matrix below.
1 3
6 2
Contributed by Chris Black Solution [436]
C22 Doing the computations by hand, nd the determinant of the matrix below.
1 3
2 6
Contributed by Chris Black Solution [436]
C23 Doing the computations by hand, nd the determinant of the matrix below.
2
41 3 2
4 1 3
1 0 13
5
Contributed by Chris Black Solution [436]
C24 Doing the computations by hand, nd the determinant of the matrix below.
2
4 2 3 2
4 2 1
2 4 23
5
Contributed by Robert Beezer Solution [436]
C25 Doing the computations by hand, nd the determinant of the matrix below.
2
43 1 4
2 5 1
2 0 63
5
Contributed by Robert Beezer Solution [436]
C26 Doing the computations by hand, nd the determinant of the matrix A.
A=2
6642 0 3 2
5 1 2 4
3 0 1 2
5 3 2 13
775
Contributed by Robert Beezer Solution [436]
C27 Doing the computations by hand, nd the determinant of the matrix A.
A=2
6641 0 1 1
2 2 1 1
2 1 3 0
1 1 0 13
775
Version 2.30
Subsection DM.EXC Exercises 437
Contributed by Chris Black Solution [437]
C28 Doing the computations by hand, nd the determinant of the matrix A.
A=2
6641 0 1 1
2 1 1 1
2 5 3 0
1 1 0 13
775
Contributed by Chris Black Solution [437]
C29 Doing the computations by hand, nd the determinant of the matrix A.
A=2
666642 3 0 2 1
0 1 1 1 2
0 0 1 2 3
0 1 2 1 0
0 0 0 1 23
77775
Contributed by Chris Black Solution [437]
C30 Doing the computations by hand, nd the determinant of the matrix A.
A=2
666642 1 1 0 1
2 1 2 1 1
0 0 1 2 0
1 0 3 1 1
2 1 1 2 13
77775
Contributed by Chris Black Solution [437]
M10 Find a value of kso that the matrix A=2 4
3k
has det(A) = 0, or explain why it is not possible.
Contributed by Chris Black Solution [438]
M11 Find a value of kso that the matrix A=2
41 2 1
2 0 1
2 3k3
5has det(A) = 0, or explain why it is not
possible.
Contributed by Chris Black Solution [438]
M15 Given the matrix B=2 x 1
4 2 x
, nd all values of xthat are solutions of det( B) = 0.
Contributed by Chris Black Solution [438]
M16 Given the matrix B=2
44 x 4 4
2 2 x 4
3 3 4 x3
5, nd all values of xthat are solutions of det( B) =
0.
Contributed by Chris Black Solution [438]
Version 2.30
438 Section DM Determinant of a Matrix
Subsection SOL
Solutions
C21 Contributed by Chris Black Statement [434]
Using the formula in Theorem DMST [429] we have
1 3
6 2= 12 63 = 2 18 = 16
C22 Contributed by Chris Black Statement [434]
Using the formula in Theorem DMST [429] we have
1 3
2 6= 16 23 = 6 6 = 0
C23 Contributed by Chris Black Statement [434]
We can compute the determinant by expanding about any row or column; the most ecient ones to choose
are either the second column or the third row. In any case, the determinant will be 4.
C24 Contributed by Robert Beezer Statement [434]
We'll expand about the rst row since there are no zeros to exploit,
2 3 2
4 2 1
2 4 2= ( 2) 2 1
4 2+ ( 1)(3) 4 1
2 2+ ( 2) 4 2
2 4
= ( 2)(( 2)(2) 1(4)) + ( 3)(( 4)(2) 1(2)) + ( 2)(( 4)(4) ( 2)(2))
= ( 2)( 8) + ( 3)( 10) + ( 2)( 12) = 70
C25 Contributed by Robert Beezer Statement [434]
We can expand about any row or column, so the zero entry in the middle of the last row is attractive. Let's
expand about column 2. By Theorem DER [429] and Theorem DEC [431] you will get the same result by
expanding about a dierent row or column. We will use Theorem DMST [429] twice.
3 1 4
2 5 1
2 0 6= ( 1)( 1)1+22 1
2 6+ (5)( 1)2+23 4
2 6+ (0)( 1)3+23 4
2 1
= (1)(10) + (5)(10) + 0 = 60
C26 Contributed by Robert Beezer Statement [434]
With two zeros in column 2, we choose to expand about that column (Theorem DEC [431]),
det (A) =2 0 3 2
5 1 2 4
3 0 1 2
5 3 2 1
= 0( 1)5 2 4
3 1 2
5 2 1+ 1(1)2 3 2
3 1 2
5 2 1+ 0( 1)2 3 2
5 2 4
5 2 1+ 3(1)2 3 2
5 2 4
3 1 2
= (1) (2(1(1) 2(2)) 3(3(1) 5(2)) + 2(3(2) 5(1))) +
Version 2.30
Subsection DM.SOL Solutions 439
(3) (2(2(2) 4(1)) 3(5(2) 4(3)) + 2(5(1) 3(2)))
= ( 6 + 21 + 2) + (3)(0 + 6 2) = 29
C27 Contributed by Chris Black Statement [434]
Expanding on the rst row, we have
1 0 1 1
2 2 1 1
2 1 3 0
1 1 0 1=2 1 1
1 3 0
1 0 1 0 +2 2 1
2 1 0
1 1 1 2 2 1
2 1 3
1 1 0
= 4 + ( 1) ( 1) = 4
C28 Contributed by Chris Black Statement [435]
Expanding along the rst row, we have
1 0 1 1
2 1 1 1
2 5 3 0
1 1 0 1= 1 1 1
5 3 0
1 0 1 0 +2 1 1
2 5 0
1 1 1 2 1 1
2 5 3
1 1 0
= 5 0 + 5 10 = 0:
C29 Contributed by Chris Black Statement [435]
Expanding along the rst column, we have
2 3 0 2 1
0 1 1 1 2
0 0 1 2 3
0 1 2 1 0
0 0 0 1 2= 21 1 1 2
0 1 2 3
1 2 1 0
0 0 1 2+ 0 + 0 + 0 + 0
Now, expanding along the rst column again, we have
= 20
@1 2 3
2 1 0
0 1 2 0 +1 1 2
1 2 3
0 1 2 01
A
= 2([2 + 0 + 6 0 0 8] + [4 + 0 + 2 0 3 2])
= 2
C30 Contributed by Chris Black Statement [435]
In order to exploit the zeros, let's expand along row 3. We then have
2 3 0 2 1
0 1 1 1 2
0 0 1 2 0
1 0 3 1 1
2 1 1 2 1= ( 1)62 1 0 1
2 1 1 1
1 0 1 1
2 1 2 1+ ( 1)722 1 1 1
2 1 2 1
1 0 3 1
2 1 1 1
Notice that the second matrix here is singular since two rows are identical and thus it cannot row-reduce
to an identity matrix. We now have
=2 1 0 1
2 1 1 1
1 0 1 1
2 1 2 1+ 0
Version 2.30
440 Section DM Determinant of a Matrix
and now we expand on the rst row of the rst matrix:
= 21 1 1
0 1 1
1 2 1 2 1 1
1 1 1
2 2 1+ 0 2 1 1
1 0 1
2 1 2
= 2( 3) ( 3) ( 3) = 0
M10 Contributed by Chris Black Statement [435]
There is only one value of kthat will make this matrix have a zero determinant.
det (A) =2 4
3k= 2k 12
so det (A) = 0 only when k= 6.
M11 Contributed by Chris Black Statement [435]
det (A) =1 2 1
2 0 1
2 3k= 7 4k
Thus, det (A) = 0 only when k=7
4.
M15 Contributed by Chris Black Statement [435]
Using the formula for the determinant of a 2 2 matrix given in Theorem DMST [429], we have
det (B) =2 x 1
4 2 x= (2 x)(2 x) 4 =x2 4x=x(x 4)
and thus det ( B) = 0 only when x= 0 orx= 4.
M16 Contributed by Chris Black Statement [435]
det (B) = 8x 2x2 x3= x(x2+ 2x 8) = x(x 2)(x+ 4)
And thus, det ( B) = 0 when x= 0,x= 2, orx= 4.
Version 2.30
Section PDM Properties of Determinants of Matrices 441
Section PDM
Properties of Determinants of Matrices
We have seen how to compute the determinant of a matrix, and the incredible fact that we can perform
expansion about anyrow orcolumn to make this computation. In this largely theoretical section, we will
state and prove several more intriguing properties about determinants. Our main goal will be the two
results in Theorem SMZD [445] and Theorem DRMM [447], but more specically, we will see how the
value of a determinant will allow us to gain insight into the various properties of a square matrix.
Subsection DRO
Determinants and Row Operations
We start easy with a straightforward theorem whose proof presages the style of subsequent proofs in this
subsection.
Theorem DZRC
Determinant with Zero Row or Column
Suppose that Ais a square matrix with a row where every entry is zero, or a column where every entry is
zero. Then det ( A) = 0.
Proof Suppose that Ais a square matrix of size nand rowihas every entry equal to zero. We compute
det (A) via expansion about row i.
det (A) =nX
j=1( 1)i+j[A]ijdet (A(ijj)) Theorem DER [429]
=nX
j=1( 1)i+j0 det (A(ijj)) Row iis zeros
=nX
j=10 = 0
The proof for the case of a zero column is entirely similar, or could be derived from an application of
Theorem DT [430] employing the transpose of the matrix.
Theorem DRCS
Determinant for Row or Column Swap
Suppose that Ais a square matrix. Let Bbe the square matrix obtained from Aby interchanging the
location of two rows, or interchanging the location of two columns. Then det ( B) = det (A).
Proof Begin with the special case where Ais a square matrix of size nand we form Bby swapping
adjacent rowsiandi+ 1 for some 1in 1. Notice that the assumption about swapping adjacent
rows means that B(i+ 1jj) =A(ijj) for all 1jn, and [B]i+1;j= [A]ijfor all 1jn. We compute
det (B) via expansion about row i+ 1.
det (B) =nX
j=1( 1)(i+1)+j[B]i+1;jdet (B(i+ 1jj)) Theorem DER [429]
=nX
j=1( 1)(i+1)+j[A]ijdet (A(ijj)) Hypothesis
Version 2.30
442 Section PDM Properties of Determinants of Matrices
=nX
j=1( 1)1( 1)i+j[A]ijdet (A(ijj))
= ( 1)nX
j=1( 1)i+j[A]ijdet (A(ijj))
= det (A) Theorem DER [429]
So the result holds for the special case where we swap adjacent rows of the matrix. As any computer
scientist knows, we can accomplish anyrearrangement of an ordered list by swapping adjacent elements.
This principle can be demonstrated by na ve sorting algorithms such as \bubble sort." In any event, we
don't need to discuss every possible reordering, we just need to consider a swap of two rows, say rows s
andtwith 1s<tn.
Begin with row s, and repeatedly swap it with each row just below it, including row tand stopping
there. This will total t sswaps. Now swap the former row t, which currently lives in row t 1, with
each row above it, stopping when it becomes row s. This will total another t s 1 swaps. In this way,
we createBthrough a sequence of 2( t s) 1 swaps of adjacent rows, each of which adjusts det ( A) by a
multiplicative factor of 1. So
det (B) = ( 1)2(t s) 1det (A) =
( 1)2t s( 1) 1det (A) = det (A)
as desired.
The proof for the case of swapping two columns is entirely similar, or could be derived from an appli-
cation of Theorem DT [430] employing the transpose of the matrix.
So Theorem DRCS [439] tells us the eect of the rst row operation (Denition RO [31]) on the
determinant of a matrix. Here's the eect of the second row operation.
Theorem DRCM
Determinant for Row or Column Multiples
Suppose that Ais a square matrix. Let Bbe the square matrix obtained from Aby multiplying a single
row by the scalar , or by multiplying a single column by the scalar . Then det ( B) =det (A).
Proof Suppose that Ais a square matrix of size nand we form the square matrix Bby multiplying each
entry of row iofAby. Notice that the other rows of AandBare equal, so A(ijj) =B(ijj), for all
1jn. We compute det ( B) via expansion about row i.
det (B) =nX
j=1( 1)i+j[B]ijdet (B(ijj)) Theorem DER [429]
=nX
j=1( 1)i+j[B]ijdet (A(ijj)) Hypothesis
=nX
j=1( 1)i+j[A]ijdet (A(ijj)) Hypothesis
=nX
j=1( 1)i+j[A]ijdet (A(ijj))
=det (A) Theorem DER [429]
The proof for the case of a multiple of a column is entirely similar, or could be derived from an application
of Theorem DT [430] employing the transpose of the matrix.
Let's go for understanding the eect of all three row operations. But rst we need an intermediate
result, but it is an easy one.
Version 2.30
Subsection PDM.DRO Determinants and Row Operations 443
Theorem DERC
Determinant with Equal Rows or Columns
Suppose that Ais a square matrix with two equal rows, or two equal columns. Then det ( A) = 0.
Proof Suppose that Ais a square matrix of size nwhere the two rows sandtare equal. Form the matrix
Bby swapping rows sandt. Notice that as a consequence of our hypothesis, A=B. Then
det (A) =1
2(det (A) + det (A))
=1
2(det (A) det (B)) Theorem DRCS [439]
=1
2(det (A) det (A)) Hypothesis, A=B
=1
2(0) = 0
The proof for the case of two equal columns is entirely similar, or could be derived from an application of
Theorem DT [430] employing the transpose of the matrix.
Now explain the third row operation. Here we go.
Theorem DRCMA
Determinant for Row or Column Multiples and Addition
Suppose that Ais a square matrix. Let Bbe the square matrix obtained from Aby multiplying a row
by the scalar and then adding it to another row, or by multiplying a column by the scalar and then
adding it to another column. Then det ( B) = det (A).
Proof Suppose that Ais a square matrix of size n. Form the matrix Bby multiplying row sbyand
adding it to row t. LetCbe the auxiliary matrix where we replace row tofAby rowsofA. Notice that
A(tjj) =B(tjj) =C(tjj) for all 1jn. We compute the determinant of Bby expansion about row t.
det (B) =nX
j=1( 1)t+j[B]tjdet (B(tjj)) Theorem DER [429]
=nX
j=1( 1)t+j
[A]sj+ [A]tj
det (B(tjj)) Hypothesis
=nX
j=1( 1)t+j[A]sjdet (B(tjj))
+nX
j=1( 1)t+j[A]tjdet (B(tjj))
=nX
j=1( 1)t+j[A]sjdet (B(tjj))
+nX
j=1( 1)t+j[A]tjdet (B(tjj))
=nX
j=1( 1)t+j[C]tjdet (C(tjj))
+nX
j=1( 1)t+j[A]tjdet (A(tjj))
=det (C) + det (A) Theorem DER [429]
Version 2.30
444 Section PDM Properties of Determinants of Matrices
=0 + det (A) = det (A) Theorem DERC [441]
The proof for the case of adding a multiple of a column is entirely similar, or could be derived from an
application of Theorem DT [430] employing the transpose of the matrix.
Is this what you expected? We could argue that the third row operation is the most popular, and yet it
has no eect whatsoever on the determinant of a matrix! We can exploit this, along with our understanding
of the other two row operations, to provide another approach to computing a determinant. We'll explain
this in the context of an example.
Example DRO
Determinant by row operations
Suppose we desire the determinant of the 4 4 matrix
A=2
6642 0 2 3
1 3 1 1
1 1 1 2
3 5 4 03
775
We will perform a sequence of row operations on this matrix, shooting for an upper triangular matrix,
whose determinant will be simply the product of its diagonal entries. For each row operation, we will track
the eect on the determinant via Theorem DRCS [439], Theorem DRCM [440], Theorem DRCMA [441].
R1$R2 !A1=2
6641 3 1 1
2 0 2 3
1 1 1 2
3 5 4 03
775det (A) = det (A1) Theorem DRCS [439]
2R1+R2 !A2=2
6641 3 1 1
0 6 4 1
1 1 1 2
3 5 4 03
775= det (A2) Theorem DRCMA [441]
1R1+R3 !A3=2
6641 3 1 1
0 6 4 1
0 4 2 3
3 5 4 03
775= det (A3) Theorem DRCMA [441]
3R1+R4 !A4=2
6641 3 1 1
0 6 4 1
0 4 2 3
0 4 7 33
775= det (A4) Theorem DRCMA [441]
1R3+R2 !A5=2
6641 3 1 1
0 2 2 4
0 4 2 3
0 4 7 33
775= det (A5) Theorem DRCMA [441]
1
2R2 !A6=2
6641 3 1 1
0 1 1 2
0 4 2 3
0 4 7 33
775= 2 det (A6) Theorem DRCM [440]
4R2+R3 !A7=2
6641 3 1 1
0 1 1 2
0 0 2 11
0 4 7 33
775= 2 det (A7) Theorem DRCMA [441]
Version 2.30
Subsection PDM.DROEM Determinants, Row Operations, Elementary Matrices 445
4R2+R4 !A8=2
6641 3 1 1
0 1 1 2
0 0 2 11
0 0 3 113
775= 2 det (A8) Theorem DRCMA [441]
1R3+R4 !A9=2
6641 3 1 1
0 1 1 2
0 0 2 11
0 0 1 223
775= 2 det (A9) Theorem DRCMA [441]
2R4+R3 !A10=2
6641 3 1 1
0 1 1 2
0 0 0 55
0 0 1 223
775= 2 det (A10) Theorem DRCMA [441]
R3$R4 !A11=2
6641 3 1 1
0 1 1 2
0 0 1 22
0 0 0 553
775= 2 det (A11) Theorem DRCS [439]
1
55R4 !A12=2
6641 3 1 1
0 1 1 2
0 0 1 22
0 0 0 13
775= 110 det (A12) Theorem DRCM [440]
The matrix A12is upper triangular, so expansion about the rst column (repeatedly) will result in
det (A12) = (1)(1)(1)(1) = 1 (see Example DUTM [432]) and thus, det ( A) = 110(1) = 110.
Notice that our sequence of row operations was somewhat ad hoc , such as the transformation to A5.
We could have been even more methodical, and strictly followed the process that converts a matrix to
reduced row-echelon form (Theorem REMEF [34]), eventually achieving the same numerical result with
a nal matrix that equaled the 4 4 identity matrix. Notice too that we could have stopped with A8,
since at this point we could compute det ( A8) by two expansions about rst columns, followed by a simple
determinant of a 2 2 matrix (Theorem DMST [429]).
The beauty of this approach is that computationally we should already have written a procedure to
convert matrices to reduced-row echelon form, so all we need to do is track the multiplicative changes to
the determinant as the algorithm proceeds. Further, for a square matrix of size nthis approach requires on
the order of n3multiplications, while a recursive application of expansion about a row or column (Theorem
DER [429], Theorem DEC [431]) will require in the vicinity of ( n 1)(n!) multiplications. So even for very
small matrices, a computational approach utilizing row operations will have superior run-time. Tracking,
and controlling, the eects of round-o errors is another story, best saved for a numerical linear algebra
course.
Subsection DROEM
Determinants, Row Operations, Elementary Matrices
As a nal preparation for our two most important theorems about determinants, we prove a handful of
facts about the interplay of row operations and matrix multiplication with elementary matrices with regard
to the determinant. But rst, a simple, but crucial, fact about the identity matrix.
Theorem DIM
Determinant of the Identity Matrix
Version 2.30
446 Section PDM Properties of Determinants of Matrices
For everyn1, det (In) = 1.
Proof It may be overkill, but this is a good situation to run through a proof by induction on n(Technique
I [772]). Is the result true when n= 1? Yes,
det (I1) = [I1]11 Denition DM [428]
= 1 Denition IM [84]
Now assume the theorem is true for the identity matrix of size n 1 and investigate the determinant of
the identity matrix of size nwith expansion about row 1,
det (In) =nX
j=1( 1)1+j[In]1jdet (In(1jj)) Denition DM [428]
= ( 1)1+1[In]11det (In(1j1))
+nX
j=2( 1)1+j[In]1jdet (In(1jj))
= 1 det (In 1) +nX
j=2( 1)1+j0 det (In(1jj)) Denition IM [84]
= 1(1) +nX
j=20 = 1 Induction Hypothesis
Theorem DEM
Determinants of Elementary Matrices
For the three possible versions of an elementary matrix (Denition ELEM [423]) we have the determinants,
1. det (Ei;j) = 1
2. det (Ei()) =
3. det (Ei;j()) = 1
Proof Swapping rows iandjof the identity matrix will create Ei;j(Denition ELEM [423]), so
det (Ei;j) = det (In) Theorem DRCS [439]
= 1 Theorem DIM [443]
Multiplying row iof the identity matrix by will createEi() (Denition ELEM [423]), so
det (Ei()) =det (In) Theorem DRCM [440]
=(1) = Theorem DIM [443]
Version 2.30
Subsection PDM.DNMMM Determinants, Nonsingular Matrices, Matrix Multiplication 447
Multiplying row iof the identity matrix by and adding to row jwill create Ei;j() (Denition ELEM
[423]), so
det (Ei;j()) = det (In) Theorem DRCMA [441]
= 1 Theorem DIM [443]
Theorem DEMMM
Determinants, Elementary Matrices, Matrix Multiplication
Suppose that Ais a square matrix of size nandEis any elementary matrix of size n. Then
det (EA) = det (E) det (A)
Proof The proof procedes in three parts, one for each type of elementary matrix, with each part very
similar to the other two. First, let Bbe the matrix obtained from Aby swapping rows iandj,
det (Ei;jA) = det (B) Theorem EMDRO [425]
= det (A) Theorem DRCS [439]
= det (Ei;j) det (A) Theorem DEM [444]
Second, let Bbe the matrix obtained from Aby multiplying row iby,
det (Ei()A) = det (B) Theorem EMDRO [425]
=det (A) Theorem DRCM [440]
= det (Ei()) det (A) Theorem DEM [444]
Third, letBbe the matrix obtained from Aby multiplying row ibyand adding to row j,
det (Ei;j()A) = det (B) Theorem EMDRO [425]
= det (A) Theorem DRCMA [441]
= det (Ei;j()) det (A) Theorem DEM [444]
Since the desired result holds for each variety of elementary matrix individually, we are done.
Subsection DNMMM
Determinants, Nonsingular Matrices, Matrix Multiplication
If you asked someone with substantial experience working with matrices about the value of the determinant,
they'd be likely to quote the following theorem as the rst thing to come to mind.
Theorem SMZD
Singular Matrices have Zero Determinants
LetAbe a square matrix. Then Ais singular if and only if det ( A) = 0.
Proof Rather than jumping into the two halves of the equivalence, we rst establish a few items. Let
Bbe the unique square matrix that is row-equivalent to Aand in reduced row-echelon form (Theorem
REMEF [34], Theorem RREFU [35]). For each of the row operations that converts BintoA, there is an
Version 2.30
448 Section PDM Properties of Determinants of Matrices
elementary matrix Eiwhich eects the row operation by matrix multiplication (Theorem EMDRO [425]).
Repeated applications of Theorem EMDRO [425] allow us to write
A=EsEs 1:::E 2E1B
Then
det (A) = det (EsEs 1:::E 2E1B)
= det (Es) det (Es 1):::det (E2) det (E1) det (B) Theorem DEMMM [445]
From Theorem DEM [444] we can infer that the determinant of an elementary matrix is never zero (note
the ban on = 0 forEi() in Denition ELEM [423]). So the product on the right is composed of nonzero
scalars, with the possible exception of det ( B). More precisely, we can argue that det ( A) = 0 if and only
if det (B) = 0. With this established, we can take up the two halves of the equivalence.
()) IfAis singular, then by Theorem NMRRI [84], Bcannot be the identity matrix. Because (1) the
number of pivot columns is equal to the number of nonzero rows, (2) not every column is a pivot column,
and (3)Bis square, we see that Bmust have a zero row. By Theorem DZRC [439] the determinant of B
is zero, and by the above, we conclude that the determinant of Ais zero.
(() We will prove the contrapositive (Technique CP [769]). So assume Ais nonsingular, then by
Theorem NMRRI [84], Bis the identity matrix and Theorem DIM [443] tells us that det ( B) = 16= 0.
With the argument above, we conclude that the determinant of Ais nonzero as well.
For the case of 2 2 matrices you might compare the application of Theorem SMZD [445] with the
combination of the results stated in Theorem DMST [429] and Theorem TTMI [246].
Example ZNDAB
Zero and nonzero determinant, Archetypes A and B
The coecient matrix in Archetype A [781] has a zero determinant (check this!) while the coecient matrix
Archetype B [786] has a nonzero determinant (check this, too). These matrices are singular and nonsingular,
respectively. This is exactly what Theorem SMZD [445] says, and continues our list of contrasts between
these two archetypes.
Since Theorem SMZD [445] is an equivalence (Technique E [768]) we can expand on our growing list
of equivalences about nonsingular matrices. The addition of the condition det ( A)6= 0 is one of the best
motivations for learning about determinants.
Theorem NME7
Nonsingular Matrix Equivalences, Round 7
Suppose that Ais a square matrix of size n. The following are equivalent.
1.Ais nonsingular.
2.Arow-reduces to the identity matrix.
3. The null space of Acontains only the zero vector, N(A) =f0g.
4. The linear system LS(A;b) has a unique solution for every possible choice of b.
5. The columns of Aare a linearly independent set.
6.Ais invertible.
7. The column space of AisCn,C(A) =Cn.
8. The columns of Aare a basis for Cn.
9. The rank of Aisn,r(A) =n.
Version 2.30
Subsection PDM.DNMMM Determinants, Nonsingular Matrices, Matrix Multiplication 449
10. The nullity of Ais zero,n(A) = 0.
11. The determinant of Ais nonzero, det ( A)6= 0.
Proof Theorem SMZD [445] says Ais singular if and only if det ( A) = 0. If we negate each of these
statements, we arrive at two contrapositives that we can combine as the equivalence, Ais nonsingular if
and only if det ( A)6= 0. This allows us to add a new statement to the list found in Theorem NME6 [399].
Computationally, row-reducing a matrix is the most ecient way to determine if a matrix is nonsingular,
though the eect of using division in a computer can lead to round-o errors that confuse small quantities
with critical zero quantities. Conceptually, the determinant may seem the most ecient way to determine if
a matrix is nonsingular. The denition of a determinant uses just addition, subtraction and multiplication,
so division is never a problem. And the nal test is easy: is the determinant zero or not? However,
the number of operations involved in computing a determinant by the denition very quickly becomes so
excessive as to be impractical.
Now for the coup de gr^ ace . We will generalize Theorem DEMMM [445] to the case of anytwo square
matrices. You may recall thinking that matrix multiplication was dened in a needlessly complicated
manner. For sure, the denition of a determinant seems even stranger. (Though Theorem SMZD [445]
might be forcing you to reconsider.) Read the statement of the next theorem and contemplate how nicely
matrix multiplication and determinants play with each other.
Theorem DRMM
Determinant Respects Matrix Multiplication
Suppose that AandBare square matrices of the same size. Then det ( AB) = det (A) det (B).
Proof This proof is constructed in two cases. First, suppose that Ais singular. Then det ( A) = 0 by
Theorem SMZD [445]. By the contrapositive of Theorem NPNT [259], ABis singular as well. So by a
second application ofTheorem SMZD [445], det ( AB) = 0. Putting it all together
det (AB) = 0 = 0 det ( B) = det (A) det (B)
as desired.
For the second case, suppose that Ais nonsingular. By Theorem NMPEM [427] there are elementary
matricesE1; E2; E3; :::; Essuch thatA=E1E2E3:::Es. Then
det (AB) = det (E1E2E3:::EsB)
= det (E1) det (E2) det (E3):::det (Es) det (B) Theorem DEMMM [445]
= det (E1E2E3:::Es) det (B) Theorem DEMMM [445]
= det (A) det (B)
It is amazing that matrix multiplication and the determinant interact this way. Might it also be true
that det (A+B) = det (A) + det (B)? (See Exercise PDM.M30 [449].)
Version 2.30
450 Section PDM Properties of Determinants of Matrices
Subsection READ
Reading Questions
1. Consider the two matrices below, and suppose you already have computed det ( A) = 120. What is
det (B)? Why?
A=2
6640 8 3 4
1 2 2 5
2 8 4 3
0 4 2 33
775B=2
6640 8 3 4
0 4 2 3
2 8 4 3
1 2 2 53
775
2. State the theorem that allows us to make yet another extension to our NMEx series of theorems.
3. What is amazing about the interaction between matrix multiplication and the determinant?
Version 2.30
Subsection PDM.EXC Exercises 451
Subsection EXC
Exercises
C30 Each of the archetypes below is a system of equations with a square coecient matrix, or is a square
matrix itself. Compute the determinant of each matrix, noting how Theorem SMZD [445] indicates when
the matrix is singular or nonsingular.
Archetype A [781]
Archetype B [786]
Archetype F [803]
Archetype K [825]
Archetype L [829]
Contributed by Robert Beezer
M20 Construct a 33 nonsingular matrix and call it A. Then, for each entry of the matrix, compute
the corresponding cofactor, and create a new 3 3 matrix full of these cofactors by placing the cofactor of
an entry in the same location as the entry it was based on. Once complete, call this matrix C. Compute
ACt. Any observations? Repeat with a new matrix, or perhaps with a 4 4 matrix.
Contributed by Robert Beezer Solution [450]
M30 Construct an example to show that the following statement is not true for all square matrices A
andBof the same size: det ( A+B) = det (A) + det (B).
Contributed by Robert Beezer
T10 Theorem NPNT [259] says that if the product of square matrices ABis nonsingular, then the
individual matrices AandBare nonsingular also. Construct a new proof of this result making use of
theorems about determinants of matrices.
Contributed by Robert Beezer
T15 Use Theorem DRCM [440] to prove Theorem DZRC [439] as a corollary. (See Technique LC [774].)
Contributed by Robert Beezer
T20 Suppose that Ais a square matrix of size nand2Cis a scalar. Prove that det ( A) =ndet (A).
Contributed by Robert Beezer
T25 Employ Theorem DT [430] to construct the second half of the proof of Theorem DRCM [440] (the
portion about a multiple of a column).
Contributed by Robert Beezer
Version 2.30
452 Section PDM Properties of Determinants of Matrices
Subsection SOL
Solutions
M20 Contributed by Robert Beezer Statement [449]
The result of these computations should be a matrix with the value of det ( A) in the diagonal entries and
zeros elsewhere. The suggestion of using a nonsingular matrix was partially so that it was obvious that
the value of the determinant appears on the diagonal.
This result (which is true in general) provides a method for computing the inverse of a nonsingular
matrix. Since ACt= det (A)In, we can multiply by the reciprocal of the determinant (which is nonzero!)
and the inverse of A(it exists!) to arrive at an expression for the matrix inverse:
A 1=1
det (A)Ct
Version 2.30
Annotated Acronyms PDM.D Determinants 453
Annotated Acronyms D
Determinants
Theorem EMDRO [425]
The main purpose of elementary matrices is to provide a more formal foundation for row operations.
With this theorem we can convert the notion of \doing a row operation" into the slightly more precise,
and tractable, operation of matrix multiplication by an elementary matrix. The other big results in this
chapter are made possible by this connection and our previous understanding of the behavior of matrix
multiplication (such as results in Section MM [223]).
Theorem DER [429]
We dene the determinant by expansion about the rst row and then prove you can expand about any row
(and with Theorem DEC [431], about any column). Amazing. If the determinant seems contrived, these
results might begin to convince you that maybe something interesting is going on.
Theorem DRMM [447]
Theorem EMDRO [425] connects elementary matrices with matrix multiplication. Now we connect deter-
minants with matrix multiplication. If you thought the denition of matrix multiplication (as exemplied
by Theorem EMP [227]) was as outlandish as the denition of the determinant, then no more. They seem
to play together quite nicely.
Theorem SMZD [445]
This theorem provides a simple test for nonsingularity, even though it is stated and titled as a theorem about
singularity. It'll be helpful, especially in concert with Theorem DRMM [447], in establishing upcoming
results about nonsingular matrices or creating alternative proofs of earlier results. You might even use
this theorem as an indicator of how often a matrix is singular. Create a square matrix at random | what
are the odds it is singular? This theorem says the determinant has to be zero, which we might suspect is
a rare occurrence. Of course, we have to be a lot more careful about words like \random," \odds," and
\rare" if we want precise answers to this question.
Version 2.30
454 Section PDM Properties of Determinants of Matrices
Version 2.30
Chapter E
Eigenvalues
When we have a square matrix of size n,A, and we multiply it by a vector xfromCnto form the matrix-
vector product (Denition MVP [223]), the result is another vector in Cn. So we can adopt a functional
view of this computation | the act of multiplying by a square matrix is a function that converts one vector
(x) into another one ( Ax) of the same size. For some vectors, this seemingly complicated computation is
really no more complicated than scalar multiplication. The vectors vary according to the choice of A, so
the question is to determine, for an individual choice of A, if there are any such vectors, and if so, which
ones. It happens in a variety of situations that these vectors (and the scalars that go along with them) are
of special interest.
We will be solving polynomial equations in this chapter, which raises the specter of roots that are
complex numbers. This distinct possibility is our main reason for entertaining the complex numbers
throughout the course. You might be moved to revisit Section CNO [757] and Section O [191].
Section EE
Eigenvalues and Eigenvectors
We start with the principal denition for this chapter.
Subsection EEM
Eigenvalues and Eigenvectors of a Matrix
Denition EEM
Eigenvalues and Eigenvectors of a Matrix
Suppose that Ais a square matrix of size n,x6=0is a vector in Cn, andis a scalar in C. Then we say
xis aneigenvector ofAwitheigenvalue if
Ax=x
4
Before going any further, perhaps we should convince you that such things ever happen at all. Un-
derstand the next example, but do not concern yourself with where the pieces come from. We will have
methods soon enough to be able to discover these eigenvectors ourselves.
Example SEE
Some eigenvalues and eigenvectors
455
456 Section EE Eigenvalues and Eigenvectors
Consider the matrix
A=2
664204 98 26 10
280 134 36 14
716 348 90 36
472 232 60 283
775
and the vectors
x=2
6641
1
2
53
775y=2
664 3
4
10
43
775z=2
664 3
7
0
83
775w=2
6641
1
4
03
775
Then
Ax=2
664204 98 26 10
280 134 36 14
716 348 90 36
472 232 60 283
7752
6641
1
2
53
775=2
6644
4
8
203
775= 42
6641
1
2
53
775= 4x
soxis an eigenvector of Awith eigenvalue = 4. Also,
Ay=2
664204 98 26 10
280 134 36 14
716 348 90 36
472 232 60 283
7752
664 3
4
10
43
775=2
6640
0
0
03
775= 02
664 3
4
10
43
775= 0y
soyis an eigenvector of Awith eigenvalue = 0. Also,
Az=2
664204 98 26 10
280 134 36 14
716 348 90 36
472 232 60 283
7752
664 3
7
0
83
775=2
664 6
14
0
163
775= 22
664 3
7
0
83
775= 2z
sozis an eigenvector of Awith eigenvalue = 2. Also,
Aw=2
664204 98 26 10
280 134 36 14
716 348 90 36
472 232 60 283
7752
6641
1
4
03
775=2
6642
2
8
03
775= 22
6641
1
4
03
775= 2w
sowis an eigenvector of Awith eigenvalue = 2.
So we have demonstrated four eigenvectors of A. Are there more? Yes, any nonzero scalar multiple of
an eigenvector is again an eigenvector. In this example, set u= 30x. Then
Au=A(30x)
= 30Ax Theorem MMSMM [230]
= 30(4 x) xan eigenvector of A
= 4(30 x) Property SMAM [209]
= 4u
so that uis also an eigenvector of Afor the same eigenvalue, = 4.
The vectors zandware both eigenvectors of Afor the same eigenvalue = 2, yet this is not as simple
as the two vectors just being scalar multiples of each other (they aren't). Look what happens when we
add them together, to form v=z+w, and multiply by A,
Av=A(z+w)
Version 2.30
Subsection EE.PM Polynomials and Matrices 457
=Az+Aw Theorem MMDAA [230]
= 2z+ 2w z ,weigenvectors of A
= 2(z+w) Property DVAC [101]
= 2v
so that vis also an eigenvector of Afor the eigenvalue = 2. So it would appear that the set of eigenvectors
that are associated with a xed eigenvalue is closed under the vector space operations of Cn. Hmmm.
The vector yis an eigenvector of Afor the eigenvalue = 0, so we can use Theorem ZSSM [324] to
writeAy= 0y=0. But this also means that y2N(A). There would appear to be a connection here
also.
Example SEE [453] hints at a number of intriguing properties, and there are many more. We will
explore the general properties of eigenvalues and eigenvectors in Section PEE [479], but in this section we
will concern ourselves with the question of actually computing eigenvalues and eigenvectors. First we need
a bit of background material on polynomials and matrices.
Subsection PM
Polynomials and Matrices
A polynomial is a combination of powers, multiplication by scalar coecients, and addition (with subtrac-
tion just being the inverse of addition). We never have occasion to divide when computing the value of
a polynomial. So it is with matrices. We can add and subtract matrices, we can multiply matrices by
scalars, and we can form powers of square matrices by repeated applications of matrix multiplication. We
do not normally divide matrices (though sometimes we can multiply by an inverse). If a matrix is square,
all the operations constituting a polynomial will preserve the size of the matrix. So it is natural to consider
evaluating a polynomial with a matrix, eectively replacing the variable of the polynomial by a matrix.
We'll demonstrate with an example,
Example PM
Polynomial of a matrix
Let
p(x) = 14 + 19 x 3x2 7x3+x4D=2
4 1 3 2
1 0 2
3 1 13
5
and we will compute p(D). First, the necessary powers of D. Notice that D0is dened to be the multi-
plicative identity, I3, as will be the case in general.
D0=I3=2
41 0 0
0 1 0
0 0 13
5
D1=D=2
4 1 3 2
1 0 2
3 1 13
5
D2=DD1=2
4 1 3 2
1 0 2
3 1 13
52
4 1 3 2
1 0 2
3 1 13
5=2
4 2 1 6
5 1 0
1 8 73
5
D3=DD2=2
4 1 3 2
1 0 2
3 1 13
52
4 2 1 6
5 1 0
1 8 73
5=2
419 12 8
4 15 8
12 4 113
5
Version 2.30
458 Section EE Eigenvalues and Eigenvectors
D4=DD3=2
4 1 3 2
1 0 2
3 1 13
52
419 12 8
4 15 8
12 4 113
5=2
4 7 49 54
5 4 30
49 47 433
5
Then
p(D) = 14 + 19 D 3D2 7D3+D4
= 142
41 0 0
0 1 0
0 0 13
5+ 192
4 1 3 2
1 0 2
3 1 13
5 32
4 2 1 6
5 1 0
1 8 73
5
72
419 12 8
4 15 8
12 4 113
5+2
4 7 49 54
5 4 30
49 47 433
5
=2
4 139 193 166
27 98 124
193 118 203
5
Notice that p(x) factors as
p(x) = 14 + 19 x 3x2 7x3+x4= (x 2)(x 7)(x+ 1)2
BecauseDcommutes with itself ( DD =DD), we can use distributivity of matrix multiplication across
matrix addition (Theorem MMDAA [230]) without being careful with any of the matrix products, and just
as easily evaluate p(D) using the factored form of p(x),
p(D) = 14 + 19 D 3D2 7D3+D4= (D 2I3)(D 7I3)(D+I3)2
=2
4 3 3 2
1 2 2
3 1 13
52
4 8 3 2
1 7 2
3 1 63
52
40 3 2
1 1 2
3 1 23
52
=2
4 139 193 166
27 98 124
193 118 203
5
This example is not meant to be too profound. It ismeant to show you that it is natural to evaluate a
polynomial with a matrix, and that the factored form of the polynomial is as good as (or maybe better
than) the expanded form. And do not forget that constant terms in polynomials are really multiples of
the identity matrix when we are evaluating the polynomial with a matrix.
Subsection EEE
Existence of Eigenvalues and Eigenvectors
Before we embark on computing eigenvalues and eigenvectors, we will prove that every matrix has at least
one eigenvalue (and an eigenvector to go with it). Later, in Theorem MNEM [487], we will determine the
maximum number of eigenvalues a matrix may have.
The determinant (Denition D [391]) will be a powerful tool in Subsection EE.CEE [460] when it comes
time to compute eigenvalues. However, it is possible, with some more advanced machinery, to compute
eigenvalues without ever making use of the determinant. Sheldon Axler does just that in his book, Linear
Version 2.30
Subsection EE.EEE Existence of Eigenvalues and Eigenvectors 459
Algebra Done Right . Here and now, we give Axler's \determinant-free" proof that every matrix has an
eigenvalue. The result is not too startling, but the proof is most enjoyable.
Theorem EMHE
Every Matrix Has an Eigenvalue
SupposeAis a square matrix. Then Ahas at least one eigenvalue.
Proof Suppose that Ahas sizen, and choose xasanynonzero vector from Cn. (Notice how much
latitude we have in our choice of x. Only the zero vector is o-limits.) Consider the set
S=
x; Ax; A2x; A3x; :::; Anx
This is a set of n+ 1 vectors from Cn, so by Theorem MVSLD [158], Sis linearly dependent. Let
a0; a1; a2; :::; anbe a collection of n+ 1 scalars from C, not all zero, that provide a relation of linear
dependence on S. In other words,
a0x+a1Ax+a2A2x+a3A3x++anAnx=0
Some of the aiare nonzero. Suppose that just a06= 0, anda1=a2=a3==an= 0. Then a0x=0
and by Theorem SMEZV [326], either a0= 0 or x=0, which are both contradictions. So ai6= 0 for some
i1. Letmbe the largest integer such that am6= 0. From this discussion we know that m1. We can
also assume that am= 1, for if not, replace each aibyai=amto obtain scalars that serve equally well in
providing a relation of linear dependence on S.
Dene the polynomial
p(x) =a0+a1x+a2x2+a3x3++amxm
Because we have consistently used Cas our set of scalars (rather than R), we know that we can factor
p(x) into linear factors of the form ( x bi), wherebi2C. So there are scalars, b1; b2; b3; :::; bm, from C
so that,
p(x) = (x bm)(x bm 1)(x b3)(x b2)(x b1)
Put it all together and
0=a0x+a1Ax+a2A2x+a3A3x++anAnx
=a0x+a1Ax+a2A2x+a3A3x++amAmx ai= 0 fori>m
=
a0In+a1A+a2A2+a3A3++amAm
x Theorem MMDAA [230]
=p(A)x Denition of p(x)
= (A bmIn)(A bm 1In)(A b3In)(A b2In)(A b1In)x
Letkbe the smallest integer such that
(A bkIn)(A bk 1In)(A b3In)(A b2In)(A b1In)x=0:
From the preceding equation, we know that km. Dene the vector zby
z= (A bk 1In)(A b3In)(A b2In)(A b1In)x
Notice that by the denition of k, the vector zmust be nonzero. In the case where k= 1, we understand
thatzis dened by z=x, and zis still nonzero. Now
(A bkIn)z= (A bkIn)(A bk 1In)(A b3In)(A b2In)(A b1In)x=0
which allows us to write
Az= (A+O)z Property ZM [209]
Version 2.30
460 Section EE Eigenvalues and Eigenvectors
= (A bkIn+bkIn)z Property AIM [209]
= (A bkIn)z+bkInz Theorem MMDAA [230]
=0+bkInz Dening property of z
=bkInz Property ZM [209]
=bkz Theorem MMIM [229]
Since z6=0, this equation says that zis an eigenvector of Afor the eigenvalue =bk(Denition EEM
[453]), so we have shown that any square matrix Adoes have at least one eigenvalue.
The proof of Theorem EMHE [457] is constructive (it contains an unambiguous procedure that leads
to an eigenvalue), but it is not meant to be practical. We will illustrate the theorem with an example, the
purpose being to provide a companion for studying the proof and not to suggest this is the best procedure
for computing an eigenvalue.
Example CAEHW
Computing an eigenvalue the hard way
This example illustrates the proof of Theorem EMHE [457], so will employ the same notation as the proof
| look there for full explanations. It is notmeant to be an example of a reasonable computational approach
to nding eigenvalues and eigenvectors. OK, warnings in place, here we go.
Let
A=2
66664 7 1 11 0 4
4 1 0 2 0
10 1 14 0 4
8 2 15 1 5
10 1 16 0 63
77775
and choose
x=2
666643
0
3
5
43
77775
It is important to notice that the choice of xcould be anything , so long as it is notthe zero vector. We
have not chosen xtotally at random, but so as to make our illustration of the theorem as general as
possible. You could replicate this example with your own choice and the computations are guaranteed to
be reasonable, provided you have a computational tool that will factor a fth degree polynomial for you.
The set
S=
x; Ax; A2x; A3x; A4x; A5x
=8
>>>><
>>>>:2
666643
0
3
5
43
77775;2
66664 4
2
4
4
63
77775;2
666646
6
6
2
103
77775;2
66664 10
14
10
2
183
77775;2
6666418
30
18
10
343
77775;2
66664 34
62
34
26
663
777759
>>>>=
>>>>;
is guaranteed to be linearly dependent, as it has six vectors from C5(Theorem MVSLD [158]). We will
search for a non-trivial relation of linear dependence by solving a homogeneous system of equations whose
coecient matrix has the vectors of Sas columns through row operations,
2
666643 4 6 10 18 34
0 2 6 14 30 62
3 4 6 10 18 34
5 4 2 2 10 26
4 6 10 18 34 663
77775RREF !2
6666410 2 6 14 30
01 3 7 15 31
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 03
77775
Version 2.30
Subsection EE.EEE Existence of Eigenvalues and Eigenvectors 461
There are four free variables for describing solutions to this homogeneous system, so we have our pick of
solutions. The most expedient choice would be to set x3= 1 andx4=x5=x6= 0. However, we will again
opt to maximize the generality of our illustration of Theorem EMHE [457] and choose x3= 8,x4= 3,
x5= 1 andx6= 0. The leads to a solution with x1= 16 andx2= 12.
This relation of linear dependence then says that
0= 16x+ 12Ax 8A2x 3A3x+A4x+ 0A5x
0=
16 + 12A 8A2 3A3+A4
x
So we dene p(x) = 16 + 12x 8x2 3x3+x4, and as advertised in the proof of Theorem EMHE [457], we
have a polynomial of degree m= 4>1 such that p(A)x=0. Now we need to factor p(x) over C. If you
made your own choice of xat the start, this is where you might have a fth degree polynomial, and where
you might need to use a computational tool to nd roots and factors. We have
p(x) = 16 + 12 x 8x2 3x3+x4= (x 4)(x+ 2)(x 2)(x+ 1)
So we know that
0=p(A)x= (A 4I5)(A+ 2I5)(A 2I5)(A+ 1I5)x
We apply one factor at a time, until we get the zero vector, so as to determine the value of kdescribed in
the proof of Theorem EMHE [457],
(A+ 1I5)x=2
66664 6 1 11 0 4
4 2 0 2 0
10 1 15 0 4
8 2 15 0 5
10 1 16 0 53
777752
666643
0
3
5
43
77775=2
66664 1
2
1
1
23
77775
(A 2I5)(A+ 1I5)x=2
66664 9 1 11 0 4
4 1 0 2 0
10 1 12 0 4
8 2 15 3 5
10 1 16 0 83
777752
66664 1
2
1
1
23
77775=2
666644
8
4
4
83
77775
(A+ 2I5)(A 2I5)(A+ 1I5)x=2
66664 5 1 11 0 4
4 3 0 2 0
10 1 16 0 4
8 2 15 1 5
10 1 16 0 43
777752
666644
8
4
4
83
77775=2
666640
0
0
0
03
77775
Sok= 3 and
z= (A 2I5)(A+ 1I5)x=2
666644
8
4
4
83
77775
is an eigenvector of Afor the eigenvalue = 2, as you can check by doing the computation Az. If
you work through this example with your own choice of the vector x(strongly recommended) then the
eigenvalue you will nd may be dierent, but will be in the set f3;0;1; 1; 2g. See Exercise EE.M60
[472] for a suggested starting vector.
Version 2.30
462 Section EE Eigenvalues and Eigenvectors
Subsection CEE
Computing Eigenvalues and Eigenvectors
Fortunately, we need not rely on the procedure of Theorem EMHE [457] each time we need an eigenvalue.
It is the determinant, and specically Theorem SMZD [445], that provides the main tool for computing
eigenvalues. Here is an informal sequence of equivalences that is the key to determining the eigenvalues
and eigenvectors of a matrix,
Ax=x()Ax Inx=0() (A In)x=0
So, for an eigenvalue and associated eigenvector x6=0, the vector xwill be a nonzero element of the
null space of A In, while the matrix A Inwill be singular and therefore have zero determinant.
These ideas are made precise in Theorem EMRCP [461] and Theorem EMNS [462], but for now this brief
discussion should suce as motivation for the following denition and example.
Denition CP
Characteristic Polynomial
Suppose that Ais a square matrix of size n. Then the characteristic polynomial ofAis the polynomial
pA(x) dened by
pA(x) = det (A xIn)
4
Example CPMS3
Characteristic polynomial of a matrix, size 3
Consider
F=2
4 13 8 4
12 7 4
24 16 73
5
Then
pF(x) = det (F xI3)
= 13 x 8 4
12 7 x 4
24 16 7 xDenition CP [460]
= ( 13 x)7 x 4
16 7 x+ ( 8)( 1)12 4
24 7 xDenition DM [428]
+ ( 4)12 7 x
24 16
= ( 13 x)((7 x)(7 x) 4(16)) Theorem DMST [429]
+ ( 8)( 1)(12(7 x) 4(24))
+ ( 4)(12(16) (7 x)(24))
= 3 + 5x+x2 x3
= (x 3)(x+ 1)2
The characteristic polynomial is our main computational tool for nding eigenvalues, and will sometimes
be used to aid us in determining the properties of eigenvalues.
Version 2.30
Subsection EE.CEE Computing Eigenvalues and Eigenvectors 463
Theorem EMRCP
Eigenvalues of a Matrix are Roots of Characteristic Polynomials
SupposeAis a square matrix. Then is an eigenvalue of Aif and only if pA() = 0.
Proof SupposeAhas sizen.
is an eigenvalue of A
() there exists x6=0so thatAx=x Denition EEM [453]
() there exists x6=0so thatAx x=0
() there exists x6=0so thatAx Inx=0 Theorem MMIM [229]
() there exists x6=0so that (A In)x=0 Theorem MMDAA [230]
()A Inis singular Denition NM [83]
() det (A In) = 0 Theorem SMZD [445]
()pA() = 0 Denition CP [460]
Example EMS3
Eigenvalues of a matrix, size 3
In Example CPMS3 [460] we found the characteristic polynomial of
F=2
4 13 8 4
12 7 4
24 16 73
5
to bepF(x) = (x 3)(x+1)2. Factored, we can nd all of its roots easily, they are x= 3 andx= 1. By
Theorem EMRCP [461], = 3 and= 1 are both eigenvalues of F, and these are the only eigenvalues
ofF. We've found them all.
Let us now turn our attention to the computation of eigenvectors.
Denition EM
Eigenspace of a Matrix
Suppose that Ais a square matrix and is an eigenvalue of A. Then the eigenspace ofAfor,EA(),
is the set of all the eigenvectors of Afor, together with the inclusion of the zero vector. 4
Example SEE [453] hinted that the set of eigenvectors for a single eigenvalue might have some closure
properties, and with the addition of the non-eigenvector, 0, we indeed get a whole subspace.
Theorem EMS
Eigenspace for a Matrix is a Subspace
SupposeAis a square matrix of size nandis an eigenvalue of A. Then the eigenspace EA() is a subspace
of the vector space Cn.
Proof We will check the three conditions of Theorem TSS [334]. First, Denition EM [461] explicitly
includes the zero vector in EA(), so the set is non-empty.
Suppose that x;y2EA(), that is, xandyare two eigenvectors of Afor. Then
A(x+y) =Ax+Ay Theorem MMDAA [230]
=x+y x ;yeigenvectors of A
=(x+y) Property DVAC [101]
So either x+y=0, orx+yis an eigenvector of Afor(Denition EEM [453]). So, in either event,
x+y2EA(), and we have additive closure.
Version 2.30
464 Section EE Eigenvalues and Eigenvectors
Suppose that 2C, and that x2EA(), that is, xis an eigenvector of Afor. Then
A(x) =(Ax) Theorem MMSMM [230]
=x x an eigenvector of A
=(x) Property SMAC [100]
So eitherx=0, orxis an eigenvector of Afor(Denition EEM [453]). So, in either event, x2EA(),
and we have scalar closure.
With the three conditions of Theorem TSS [334] met, we know EA() is a subspace.
Theorem EMS [461] tells us that an eigenspace is a subspace (and hence a vector space in its own
right). Our next theorem tells us how to quickly construct this subspace.
Theorem EMNS
Eigenspace of a Matrix is a Null Space
SupposeAis a square matrix of size nandis an eigenvalue of A. Then
EA() =N(A In)
Proof The conclusion of this theorem is an equality of sets, so normally we would follow the advice of
Denition SE [762]. However, in this case we can construct a sequence of equivalences which will together
provide the two subset inclusions we need. First, notice that 02EA() by Denition EM [461] and
02N(A In) by Theorem HSC [71]. Now consider any nonzero vector x2Cn,
x2EA()()Ax=x Denition EM [461]
()Ax x=0
()Ax Inx=0 Theorem MMIM [229]
() (A In)x=0 Theorem MMDAA [230]
()x2N(A In) Denition NSM [73]
You might notice the close parallels (and dierences) between the proofs of Theorem EMRCP [461]
and Theorem EMNS [462]. Since Theorem EMNS [462] describes the set of all the eigenvectors of Aas a
null space we can use techniques such as Theorem BNS [160] to provide concise descriptions of eigenspaces.
Theorem EMNS [462] also provides a trivial proof for Theorem EMS [461].
Example ESMS3
Eigenspaces of a matrix, size 3
Example CPMS3 [460] and Example EMS3 [461] describe the characteristic polynomial and eigenvalues of
the 33 matrix
F=2
4 13 8 4
12 7 4
24 16 73
5
We will now take each eigenvalue in turn and compute its eigenspace. To do this, we row-reduce the matrix
F I3in order to determine solutions to the homogeneous system LS(F I3;0) and then express the
eigenspace as the null space of F I3(Theorem EMNS [462]). Theorem BNS [160] then tells us how to
write the null space as the span of a basis.
= 3 F 3I3=2
4 16 8 4
12 4 4
24 16 43
5RREF !2
4101
2
01 1
2
0 0 03
5
Version 2.30
Subsection EE.ECEE Examples of Computing Eigenvalues and Eigenvectors 465
EF(3) =N(F 3I3) =*8
<
:2
4 1
21
2
13
59
=
;+
=*8
<
:2
4 1
1
23
59
=
;+
= 1F+ 1I3=2
4 12 8 4
12 8 4
24 16 83
5RREF !2
412
31
3
0 0 0
0 0 03
5
EF( 1) =N(F+ 1I3) =*8
<
:2
4 2
3
1
03
5;2
4 1
3
0
13
59
=
;+
=*8
<
:2
4 2
3
03
5;2
4 1
0
33
59
=
;+
Eigenspaces in hand, we can easily compute eigenvectors by forming nontrivial linear combinations of
the basis vectors describing each eigenspace. In particular, notice that we can \pretty up" our basis
vectors by using scalar multiples to clear out fractions. More powerful scientic calculators, and most
every mathematical software package, will compute eigenvalues of a matrix along with basis vectors of the
eigenspaces. Be sure to understand how your device outputs complex numbers, since they are likely to
occur. Also, the basis vectors will not necessarily look like the results of an application of Theorem BNS
[160]. Duplicating the results of the next section (Subsection EE.ECEE [463]) with your device would be
very good practice. See: Computation E.SAGE [755]
Subsection ECEE
Examples of Computing Eigenvalues and Eigenvectors
No theorems in this section, just a selection of examples meant to illustrate the range of possibilities for the
eigenvalues and eigenvectors of a matrix. These examples can all be done by hand, though the computation
of the characteristic polynomial would be very time-consuming and error-prone. It can also be dicult
to factor an arbitrary polynomial, though if we were to suggest that most of our eigenvalues are going
to be integers, then it can be easier to hunt for roots. These examples are meant to look similar to a
concatenation of Example CPMS3 [460], Example EMS3 [461] and Example ESMS3 [462]. First, we will
sneak in a pair of denitions so we can illustrate them throughout this sequence of examples.
Denition AME
Algebraic Multiplicity of an Eigenvalue
Suppose that Ais a square matrix and is an eigenvalue of A. Then the algebraic multiplicity of,
A(), is the highest power of ( x ) that divides the characteristic polynomial, pA(x).
(This denition contains Notation AME.) 4
Since an eigenvalue is a root of the characteristic polynomial, there is always a factor of ( x ),
and the algebraic multiplicity is just the power of this factor in a factorization of pA(x). So in particular,
A()1. Compare the denition of algebraic multiplicity with the next denition.
Denition GME
Geometric Multiplicity of an Eigenvalue
Suppose that Ais a square matrix and is an eigenvalue of A. Then the geometric multiplicity of,
A(), is the dimension of the eigenspace EA().
(This denition contains Notation GME.) 4
Since every eigenvalue must have at least one eigenvector, the associated eigenspace cannot be trivial,
and so
A()1.
Example EMMS4
Eigenvalue multiplicities, matrix of size 4
Version 2.30
466 Section EE Eigenvalues and Eigenvectors
Consider the matrix
B=2
664 2 1 2 4
12 1 4 9
6 5 2 4
3 4 5 103
775
then
pB(x) = 8 20x+ 18x2 7x3+x4= (x 1)(x 2)3
So the eigenvalues are = 1;2 with algebraic multiplicities B(1) = 1 and B(2) = 3.
Computing eigenvectors,
= 1 B 1I4=2
664 3 1 2 4
12 0 4 9
6 5 3 4
3 4 5 93
775RREF !2
664101
30
01 1 0
0 0 0 1
0 0 0 03
775
EB(1) =N(B 1I4) =*8
>><
>>:2
664 1
3
1
1
03
7759
>>=
>>;+
=*8
>><
>>:2
664 1
3
3
03
7759
>>=
>>;+
= 2 B 2I4=2
664 4 1 2 4
12 1 4 9
6 5 4 4
3 4 5 83
775RREF !2
66410 0 1=2
010 1
0 0 11=2
0 0 0 03
775
EB(2) =N(B 2I4) =*8
>><
>>:2
664 1
2
1
1
2
13
7759
>>=
>>;+
=*8
>><
>>:2
664 1
2
1
23
7759
>>=
>>;+
So each eigenspace has dimension 1 and so
B(1) = 1 and
B(2) = 1. This example is of interest because
of the discrepancy between the two multiplicities for = 2. In many of our examples the algebraic and
geometric multiplicities will be equal for all of the eigenvalues (as it was for = 1 in this example), so keep
this example in mind. We will have some explanations for this phenomenon later (see Example NDMS4
[501]).
Example ESMS4
Eigenvalues, symmetric matrix of size 4
Consider the matrix
C=2
6641 0 1 1
0 1 1 1
1 1 1 0
1 1 0 13
775
then
pC(x) = 3 + 4x+ 2x2 4x3+x4= (x 3)(x 1)2(x+ 1)
So the eigenvalues are = 3;1; 1 with algebraic multiplicities C(3) = 1,C(1) = 2 and C( 1) = 1.
Computing eigenvectors,
= 3 C 3I4=2
664 2 0 1 1
0 2 1 1
1 1 2 0
1 1 0 23
775RREF !2
66410 0 1
010 1
0 0 1 1
0 0 0 03
775
Version 2.30
Subsection EE.ECEE Examples of Computing Eigenvalues and Eigenvectors 467
EC(3) =N(C 3I4) =*8
>><
>>:2
6641
1
1
13
7759
>>=
>>;+
= 1 C 1I4=2
6640 0 1 1
0 0 1 1
1 1 0 0
1 1 0 03
775RREF !2
66411 0 0
0 0 11
0 0 0 0
0 0 0 03
775
EC(1) =N(C 1I4) =*8
>><
>>:2
664 1
1
0
03
775;2
6640
0
1
13
7759
>>=
>>;+
= 1 C+ 1I4=2
6642 0 1 1
0 2 1 1
1 1 2 0
1 1 0 23
775RREF !2
66410 0 1
010 1
0 0 1 1
0 0 0 03
775
EC( 1) =N(C+ 1I4) =*8
>><
>>:2
664 1
1
1
13
7759
>>=
>>;+
So the eigenspace dimensions yield geometric multiplicities
C(3) = 1,
C(1) = 2 and
C( 1) = 1, the
same as for the algebraic multiplicities. This example is of interest because Ais a symmetric matrix, and
will be the subject of Theorem HMRE [487].
Example HMEM5
High multiplicity eigenvalues, matrix of size 5
Consider the matrix
E=2
6666429 14 2 6 9
47 22 1 11 13
19 10 5 4 8
19 10 3 2 8
7 4 3 1 33
77775
then
pE(x) = 16 + 16x+ 8x2 16x3+ 7x4 x5= (x 2)4(x+ 1)
So the eigenvalues are = 2; 1 with algebraic multiplicities E(2) = 4 and E( 1) = 1.
Computing eigenvectors,
= 2 E 2I5=2
6666427 14 2 6 9
47 24 1 11 13
19 10 3 4 8
19 10 3 4 8
7 4 3 1 53
77775RREF !2
66666410 0 1 0
010 3
2 1
2
0 0 1 0 1
0 0 0 0 0
0 0 0 0 03
777775
EE(2) =N(E 2I5) =*8
>>>><
>>>>:2
66664 1
3
2
0
1
03
77775;2
666640
1
2
1
0
13
777759
>>>>=
>>>>;+
=*8
>>>><
>>>>:2
66664 2
3
0
2
03
77775;2
666640
1
2
0
23
777759
>>>>=
>>>>;+
Version 2.30
468 Section EE Eigenvalues and Eigenvectors
= 1E+ 1I5=2
6666430 14 2 6 9
47 21 1 11 13
19 10 6 4 8
19 10 3 1 8
7 4 3 1 23
77775RREF !2
66666410 0 2 0
010 4 0
0 0 1 1 0
0 0 0 0 1
0 0 0 0 03
777775
EE( 1) =N(E+ 1I5) =*8
>>>><
>>>>:2
66664 2
4
1
1
03
777759
>>>>=
>>>>;+
So the eigenspace dimensions yield geometric multiplicities
E(2) = 2 and
E( 1) = 1. This example is
of interest because = 2 has such a large algebraic multiplicity, which is also not equal to its geometric
multiplicity.
Example CEMS6
Complex eigenvalues, matrix of size 6
Consider the matrix
F=2
6666664 59 34 41 12 25 30
1 7 46 36 11 29
233 119 58 35 75 54
157 81 43 21 51 39
91 48 32 5 32 26
209 107 55 28 69 503
7777775
then
pF(x) = 50 + 55x+ 13x2 50x3+ 32x4 9x5+x6
= (x 2)(x+ 1)(x2 4x+ 5)2
= (x 2)(x+ 1)((x (2 +i))(x (2 i)))2
= (x 2)(x+ 1)(x (2 +i))2(x (2 i))2
So the eigenvalues are = 2; 1;2 +i;2 iwith algebraic multiplicities F(2) = 1,F( 1) = 1,
F(2 +i) = 2 andF(2 i) = 2.
Computing eigenvectors,
= 2
F 2I6=2
6666664 61 34 41 12 25 30
1 5 46 36 11 29
233 119 56 35 75 54
157 81 43 19 51 39
91 48 32 5 30 26
209 107 55 28 69 523
7777775RREF !2
6666666410 0 0 01
5
010 0 0 0
0 0 10 03
5
0 0 0 10 1
5
0 0 0 0 14
5
0 0 0 0 0 03
77777775
EF(2) =N(F 2I6) =*8
>>>>>><
>>>>>>:2
6666664 1
5
0
3
51
5
4
5
13
77777759
>>>>>>=
>>>>>>;+
=*8
>>>>>><
>>>>>>:2
6666664 1
0
3
1
4
53
77777759
>>>>>>=
>>>>>>;+
Version 2.30
Subsection EE.ECEE Examples of Computing Eigenvalues and Eigenvectors 469
= 1
F+ 1I6=2
6666664 58 34 41 12 25 30
1 8 46 36 11 29
233 119 59 35 75 54
157 81 43 22 51 39
91 48 32 5 33 26
209 107 55 28 69 493
7777775RREF !2
6666666410 0 0 01
2
010 0 0 3
2
0 0 10 01
2
0 0 0 10 0
0 0 0 0 1 1
2
0 0 0 0 0 03
77777775
EF( 1) =N(F+I6) =*8
>>>>>><
>>>>>>:2
6666664 1
23
2
1
2
0
1
2
13
77777759
>>>>>>=
>>>>>>;+
=*8
>>>>>><
>>>>>>:2
6666664 1
3
1
0
1
23
77777759
>>>>>>=
>>>>>>;+
= 2 +i
F (2 +i)I6=2
6666664 61 i 34 41 12 25 30
1 5 i 46 36 11 29
233 119 56 i 35 75 54
157 81 43 19 i 51 39
91 48 32 5 30 i 26
209 107 55 28 69 52 i3
7777775
RREF !2
6666666410 0 0 01
5(7 +i)
010 0 01
5( 9 2i)
0 0 10 0 1
0 0 0 10 1
0 0 0 0 1 1
0 0 0 0 0 03
77777775
EF(2 +i) =N(F (2 +i)I6) =*8
>>>>>><
>>>>>>:2
6666664 1
5(7 +i)
1
5(9 + 2i)
1
1
1
13
77777759
>>>>>>=
>>>>>>;+
=*8
>>>>>><
>>>>>>:2
6666664 7 i
9 + 2i
5
5
5
53
77777759
>>>>>>=
>>>>>>;+
= 2 i
F (2 i)I6=2
6666664 61 +i 34 41 12 25 30
1 5 +i 46 36 11 29
233 119 56 +i 35 75 54
157 81 43 19 +i 51 39
91 48 32 5 30 +i 26
209 107 55 28 69 52 +i3
7777775
Version 2.30
470 Section EE Eigenvalues and Eigenvectors
RREF !2
6666666410 0 0 01
5(7 i)
010 0 01
5( 9 + 2i)
0 0 10 0 1
0 0 0 10 1
0 0 0 0 1 1
0 0 0 0 0 03
77777775
EF(2 i) =N(F (2 i)I6) =*8
>>>>>><
>>>>>>:2
66666641
5( 7 +i)
1
5(9 2i)
1
1
1
13
77777759
>>>>>>=
>>>>>>;+
=*8
>>>>>><
>>>>>>:2
6666664 7 +i
9 2i
5
5
5
53
77777759
>>>>>>=
>>>>>>;+
So the eigenspace dimensions yield geometric multiplicities
F(2) = 1,
F( 1) = 1,
F(2 +i) = 1
and
F(2 i) = 1. This example demonstrates some of the possibilities for the appearance of complex
eigenvalues, even when all the entries of the matrix are real. Notice how all the numbers in the analysis of
= 2 iare conjugates of the corresponding number in the analysis of = 2 +i. This is the content of
the upcoming Theorem ERMCP [483].
Example DEMS5
Distinct eigenvalues, matrix of size 5
Consider the matrix
H=2
6666415 18 8 6 5
5 3 1 1 3
0 4 5 4 2
43 46 17 14 15
26 30 12 8 103
77775
then
pH(x) = 6x+x2+ 7x3 x4 x5=x(x 2)(x 1)(x+ 1)(x+ 3)
So the eigenvalues are = 2;1;0; 1; 3 with algebraic multiplicities H(2) = 1,H(1) = 1,H(0) = 1,
H( 1) = 1 and H( 3) = 1.
Computing eigenvectors,
= 2 H 2I5=2
6666413 18 8 6 5
5 1 1 1 3
0 4 3 4 2
43 46 17 16 15
26 30 12 8 123
77775RREF !2
66666410 0 0 1
010 0 1
0 0 10 2
0 0 0 1 1
0 0 0 0 03
777775
EH(2) =N(H 2I5) =*8
>>>><
>>>>:2
666641
1
2
1
13
777759
>>>>=
>>>>;+
= 1 H 1I5=2
6666414 18 8 6 5
5 2 1 1 3
0 4 4 4 2
43 46 17 15 15
26 30 12 8 113
77775RREF !2
66666410 0 0 1
2
010 0 0
0 0 101
2
0 0 0 1 1
0 0 0 0 03
777775
Version 2.30
Subsection EE.ECEE Examples of Computing Eigenvalues and Eigenvectors 471
EH(1) =N(H 1I5) =*8
>>>><
>>>>:2
666641
2
0
1
2
1
13
777759
>>>>=
>>>>;+
=*8
>>>><
>>>>:2
666641
0
1
2
23
777759
>>>>=
>>>>;+
= 0 H 0I5=2
6666415 18 8 6 5
5 3 1 1 3
0 4 5 4 2
43 46 17 14 15
26 30 12 8 103
77775RREF !2
66666410 0 0 1
010 0 2
0 0 10 2
0 0 0 1 0
0 0 0 0 03
777775
EH(0) =N(H 0I5) =*8
>>>><
>>>>:2
66664 1
2
2
0
13
777759
>>>>=
>>>>;+
= 1H+ 1I5=2
6666416 18 8 6 5
5 4 1 1 3
0 4 6 4 2
43 46 17 13 15
26 30 12 8 93
77775RREF !2
66666410 0 0 1=2
010 0 0
0 0 10 0
0 0 0 1 1=2
0 0 0 0 03
777775
EH( 1) =N(H+ 1I5) =*8
>>>><
>>>>:2
666641
2
0
0
1
2
13
777759
>>>>=
>>>>;+
=*8
>>>><
>>>>:2
666641
0
0
1
23
777759
>>>>=
>>>>;+
= 3H+ 3I5=2
6666418 18 8 6 5
5 6 1 1 3
0 4 8 4 2
43 46 17 11 15
26 30 12 8 73
77775RREF !2
66666410 0 0 1
010 01
2
0 0 10 1
0 0 0 1 2
0 0 0 0 03
777775
EH( 3) =N(H+ 3I5) =*8
>>>><
>>>>:2
666641
1
2
1
2
13
777759
>>>>=
>>>>;+
=*8
>>>><
>>>>:2
66664 2
1
2
4
23
777759
>>>>=
>>>>;+
So the eigenspace dimensions yield geometric multiplicities
H(2) = 1,
H(1) = 1,
H(0) = 1,
H( 1) = 1
and
H( 3) = 1, identical to the algebraic multiplicities. This example is of interest for two reasons. First,
= 0 is an eigenvalue, illustrating the upcoming Theorem SMZE [480]. Second, all the eigenvalues are
distinct, yielding algebraic and geometric multiplicities of 1 for each eigenvalue, illustrating Theorem DED
[501].
Version 2.30
472 Section EE Eigenvalues and Eigenvectors
Subsection READ
Reading Questions
SupposeAis the 22 matrix
A= 5 8
4 7
1. Find the eigenvalues of A.
2. Find the eigenspaces of A.
3. For the polynomial p(x) = 3x2 x+ 2, compute p(A).
Version 2.30
Subsection EE.EXC Exercises 473
Subsection EXC
Exercises
C10 Find the characteristic polynomial of the matrix A=1 2
3 4
.
Contributed by Chris Black Solution [473]
C11 Find the characteristic polynomial of the matrix A=2
43 2 1
0 1 1
1 2 03
5.
Contributed by Chris Black Solution [473]
C12 Find the characteristic polynomial of the matrix A=2
6641 2 1 0
1 0 1 0
2 1 1 0
3 1 0 13
775.
Contributed by Chris Black Solution [473]
C19 Find the eigenvalues, eigenspaces, algebraic multiplicities and geometric multiplicities for the matrix
below. It is possible to do all these computations by hand, and it would be instructive to do so.
C= 1 2
6 6
Contributed by Robert Beezer Solution [473]
C20 Find the eigenvalues, eigenspaces, algebraic multiplicities and geometric multiplicities for the matrix
below. It is possible to do all these computations by hand, and it would be instructive to do so.
B= 12 30
5 13
Contributed by Robert Beezer Solution [473]
C21 The matrix Abelow has= 2 as an eigenvalue. Find the geometric multiplicity of = 2 using
your calculator only for row-reducing matrices.
A=2
66418 15 33 15
4 8 6 6
9 9 16 9
5 6 9 43
775
Contributed by Robert Beezer Solution [474]
C22 Without using a calculator, nd the eigenvalues of the matrix B.
B=2 1
1 1
Contributed by Robert Beezer Solution [474]
C23 Find the eigenvalues, eigenspaces, algebraic and geometric multiplicities for A=1 1
1 1
:
Contributed by Chris Black Solution [474]
Version 2.30
474 Section EE Eigenvalues and Eigenvectors
C24 Find the eigenvalues, eigenspaces, algebraic and geometric multiplicities for A=2
41 1 1
1 1 1
1 1 13
5.
Contributed by Chris Black Solution [475]
C25 Find the eigenvalues, eigenspaces, algebraic and geometric multiplicities for the 3 3 identity matrix
I3. Do your results make sense?
Contributed by Chris Black Solution [475]
C26 For matrix A=2
42 1 1
1 2 1
1 1 23
5, the characteristic polynomial of AispA() = (4 x)(1 x)2. Find the
eigenvalues and corresponding eigenspaces of A.
Contributed by Chris Black Solution [475]
C27 For matrix A=2
6640 4 1 1
2 6 1 1
2 8 1 1
2 8 3 13
775, the characteristic polynomial of Ais
pA() = (x+ 2)(x 2)2(x 4):
Find the eigenvalues and corresponding eigenspaces of A.
Contributed by Chris Black Solution [475]
M60 Repeat Example CAEHW [458] by choosing x=2
666640
8
2
1
23
77775and then arrive at an eigenvalue and eigen-
vector of the matrix A. The hard way.
Contributed by Robert Beezer Solution [476]
T10 A matrixAis idempotent if A2=A. Show that the only possible eigenvalues of an idempotent
matrix are = 0 and= 1. Then give an example of a matrix that is idempotent and has both of these
two values as eigenvalues.
Contributed by Robert Beezer Solution [476]
T15 The characteristic polynomial of the square matrix Ais usually dened as rA(x) = det (xIn A).
Find a specic relationship between our characteristic polynomial, pA(x), andrA(x), give a proof of
your relationship, and use this to explain why Theorem EMRCP [461] can remain essentially unchanged
with either denition. Explain the advantages of each denition over the other. (Computing with both
denitions, for a 2 2 and a 33 matrix, might be a good way to start.)
Contributed by Robert Beezer Solution [477]
T20 Suppose that andare two dierent eigenvalues of the square matrix A. Prove that the intersection
of the eigenspaces for these two eigenvalues is trivial. That is, EA()\EA() =f0g.
Contributed by Robert Beezer Solution [477]
Version 2.30
Subsection EE.SOL Solutions 475
Subsection SOL
Solutions
C10 Contributed by Chris Black Statement [471]
Answer:pA(x) = 2 5x+x2
C11 Contributed by Chris Black Statement [471]
Answer:pA(x) = 5 + 4x2 x3.
C12 Contributed by Chris Black Statement [471]
Answer:pA(x) = 2 + 2x 2x2 3x3+x4.
C19 Contributed by Robert Beezer Statement [471]
First compute the characteristic polynomial,
pC(x) = det (C xI2) Denition CP [460]
= 1 x 2
6 6 x
= ( 1 x)(6 x) (2)( 6)
=x2 5x+ 6
= (x 3)(x 2)
So the eigenvalues of Care the solutions to pC(x) = 0, namely, = 2 and= 3.
To obtain the eigenspaces, construct the appropriate singular matrices and nd expressions for the null
spaces of these matrices.
= 2
C (2)I2= 3 2
6 4
RREF !
1 2
3
0 0
EC(2) =N(C (2)I2) =2
3
1
=2
3
= 3
C (3)I2= 4 2
6 3
RREF !
1 1
2
0 0
EC(3) =N(C (3)I2) =1
2
1
=1
2
C20 Contributed by Robert Beezer Statement [471]
The characteristic polynomial of Bis
pB(x) = det (B xI2) Denition CP [460]
= 12 x 30
5 13 x
= ( 12 x)(13 x) (30)( 5) Theorem DMST [429]
=x2 x 6
= (x 3)(x+ 2)
Version 2.30
476 Section EE Eigenvalues and Eigenvectors
From this we nd eigenvalues = 3; 2 with algebraic multiplicities B(3) = 1 and B( 2) = 1.
For eigenvectors and geometric multiplicities, we study the null spaces of B I2(Theorem EMNS
[462]).
= 3 B 3I2= 15 30
5 10
RREF !
1 2
0 0
EB(3) =N(B 3I2) =2
1
= 2 B+ 2I2= 10 30
5 15
RREF !
1 3
0 0
EB( 2) =N(B+ 2I2) =3
1
Each eigenspace has dimension one, so we have geometric multiplicities
B(3) = 1 and
B( 2) = 1.
C21 Contributed by Robert Beezer Statement [471]
If= 2 is an eigenvalue of A, the matrix A 2I4will be singular, and its null space will be the eigenspace
ofA. So we form this matrix and row-reduce,
A 2I4=2
66416 15 33 15
4 6 6 6
9 9 18 9
5 6 9 63
775RREF !2
66410 3 0
011 1
0 0 0 0
0 0 0 03
775
With two free variables, we know a basis of the null space (Theorem BNS [160]) will contain two vectors.
Thus the null space of A 2I4has dimension two, and so the eigenspace of = 2 has dimension two also
(Theorem EMNS [462]),
A(2) = 2.
C22 Contributed by Robert Beezer Statement [471]
The characteristic polynomial (Denition CP [460]) is
pB(x) = det (B xI2)
=2 x 1
1 1 x
= (2 x)(1 x) (1)( 1) Theorem DMST [429]
=x2 3x+ 3
=
x 3 +p
3i
2!
x 3 p
3i
2!
where the factorization can be obtained by nding the roots of pB(x) = 0 with the quadratic equation.
By Theorem EMRCP [461] the eigenvalues of Bare the complex numbers 1=3+p
3i
2and2=3 p
3i
2.
C23 Contributed by Chris Black Statement [471]
Eigenvalues Eigenspaces Algebraic Multiplicity Geometric Multiplicity
= 0EA(0) = 1
1
A(0) = 1
A(0) = 1
= 2EA(2) =1
1
A(2) = 1
A(2) = 1
Version 2.30
Subsection EE.SOL Solutions 477
C24 Contributed by Chris Black Statement [472]
Eigenvalues Eigenspaces Algebraic Multiplicity Geometric Multiplicity
= 0EA(0) =*2
41
1
03
5;2
4 1
0
13
5+
A(0) = 2
A(0) = 2
= 3EA(3) =*2
41
1
13
5+
A(3) = 1
A(3) = 1
C25 Contributed by Chris Black Statement [472]
The characteristic polynomial for A=I3ispI3(x) = (1 x)3, which has eigenvalue = 1 with algebraic
multiplicity A(1) = 3. Looking for eigenvectors, we nd that A I=2
40 0 0
0 0 0
0 0 03
5. The nullspace of this
matrix is all of C3, so that the eigenspace is EI3(1) =*2
41
0
03
5;2
40
1
03
5;2
40
0
13
5+
, and the geometric multiplicity
is
A(1) = 3.
Does this make sense? Yes! Every vector xis a solution to I3x= 1x, so every nonzero vector is an
eigenvector with eigenvalue 1. Since every vector is unchanged when multiplied by I3, it makes sense that
= 1 is the only eigenvalue.
C26 Contributed by Chris Black Statement [472]
Since we are given that the characteristic polynomial of AispA(x) = (4 x)(1 x)2, we see that
the eigenvalues are = 4 with algebraic multiplicity A(4) = 1 and = 1 with algebraic multiplicity
A(1) = 2. The corresponding eigenspaces are
EA(4) =*2
41
1
13
5+
EA(1) =*2
41
1
03
5;2
41
0
13
5+
C27 Contributed by Chris Black Statement [472]
Since we are given that the characteristic polynomial of AispA(x) = (x+ 2)(x 2)2(x 4), we see that
the eigenvalues are = 2,= 2 and= 4. The eigenspaces are
EA( 2) =*2
6640
0
1
13
775+
EA(2) =*2
6641
1
2
03
775;2
6643
1
0
23
775+
EA(4) =*2
6641
1
1
13
775+
M60 Contributed by Robert Beezer Statement [472]
Version 2.30
478 Section EE Eigenvalues and Eigenvectors
Form the matrix Cwhose columns are x; Ax; A2x; A3x; A4x; A5xand row-reduce the matrix,
2
666640 6 32 102 320 966
8 10 24 58 168 490
2 12 50 156 482 1452
1 5 47 149 479 1445
2 12 50 156 482 14523
77775RREF !2
66666410 0 3 9 30
010 1 0 1
0 0 1 3 10 30
0 0 0 0 0 0
0 0 0 0 0 03
777775
The simplest possible relation of linear dependence on the columns of Ccomes from using scalars 4= 1
and5=6= 0 for the free variables in a solution to LS(C;0). The remainder of this solution is 1= 3,
2= 1,3= 3. This solution gives rise to the polynomial
p(x) = 3 x 3x2+x3= (x 3)(x 1)(x+ 1)
which then has the property that p(A)x=0.
No matter how you choose to order the factors of p(x), the value of k(in the language of Theorem
EMHE [457] and Example CAEHW [458]) is k= 2. For each of the three possibilities, we list the resulting
eigenvector and the associated eigenvalue:
(C 3I5)(C I5)z=2
666648
8
8
24
83
77775= 1
(C 3I5)(C+I5)z=2
6666420
20
20
40
203
77775= 1
(C+I5)(C I5)z=2
6666432
16
48
48
483
77775= 3
Note that each of these eigenvectors can be simplied by an appropriate scalar multiple, but we have shown
here the actual vector obtained by the product specied in the theorem.
T10 Contributed by Robert Beezer Statement [472]
Suppose that is an eigenvalue of A. Then there is an eigenvector x, such that Ax=x. We have,
x=Ax x eigenvector of A
=A2x Ais idempotent
=A(Ax)
=A(x) xeigenvector of A
=(Ax) Theorem MMSMM [230]
=(x) xeigenvector of A
=2x
From this we get
0=2x x
Version 2.30
Subsection EE.SOL Solutions 479
= (2 )x Property DSAC [101]
Since xis an eigenvector, it is nonzero, and Theorem SMEZV [326] leaves us with the conclusion that
2 = 0, and the solutions to this quadratic polynomial equation in are= 0 and= 1.
The matrix 1 0
0 0
is idempotent (check this!) and since it is a diagonal matrix, its eigenvalues are the diagonal entries, = 0
and= 1, so each of these possible values for an eigenvalue of an idempotent matrix actually occurs as an
eigenvalue of some idempotent matrix. So we cannot state any stronger conclusion about the eigenvalues
of an idempotent matrix, and we can say that this theorem is the \best possible."
T15 Contributed by Robert Beezer Statement [472]
Note in the following that the scalar multiple of a matrix is equivalent to multiplying each of the rows
by that scalar, so we actually apply Theorem DRCM [440] multiple times below (and are passing up an
opportunity to do a proof by induction in the process, which maybe you'd like to do yourself?).
pA(x) = det (A xIn) Denition CP [460]
= det (( 1)(xIn A)) Denition MSM [208]
= ( 1)ndet (xIn A) Theorem DRCM [440]
= ( 1)nrA(x)
Since the polynomials are scalar multiples of each other, their roots will be identical, so either polynomial
could be used in Theorem EMRCP [461].
Computing by hand, our denition of the characteristic polynomial is easier to use, as you only need
to subtract xdown the diagonal of the matrix before computing the determinant. However, the price to
be paid is that for odd values of n, the coecient of xnis 1, whilerA(x) always has the coecient 1 for
xn(we sayrA(x) is \monic.")
T20 Contributed by Robert Beezer Statement [472]
This problem asks you to prove that two sets are equal, so use Denition SE [762].
First show thatf0gEA()\EA(). Choose x2f0g. Then x=0. Eigenspaces are subspaces
(Theorem EMS [461]), so both EA() andEA() contain the zero vector, and therefore x2EA()\EA()
(Denition SI [763]).
To show thatEA()\EA()f0g, suppose that x2EA()\EA(). Then xis an eigenvector of A
for bothand(Denition SI [763]) and so
x= 1x Property O [318]
=1
( )x 6=; 6= 0
=1
(x x) Property DSAC [101]
=1
(Ax Ax) xeigenvector of Afor,
=1
(0)
=0 Theorem ZVSM [325]
Sox=0, and trivially, x2f0g.
Version 2.30
480 Section EE Eigenvalues and Eigenvectors
Version 2.30
Section PEE Properties of Eigenvalues and Eigenvectors 481
Section PEE
Properties of Eigenvalues and Eigenvectors
The previous section introduced eigenvalues and eigenvectors, and concentrated on their existence and
determination. This section will be more about theorems, and the various properties eigenvalues and
eigenvectors enjoy. Like a good 4 100 meter relay, we will lead-o with one of our better theorems and
save the very best for the anchor leg.
Theorem EDELI
Eigenvectors with Distinct Eigenvalues are Linearly Independent
Suppose that Ais annnsquare matrix and S=fx1;x2;x3; :::; xpgis a set of eigenvectors with
eigenvalues 1; 2; 3; :::; psuch thati6=jwheneveri6=j. ThenSis a linearly independent set.
Proof Ifp= 1, then the set S=fx1gis linearly independent since eigenvectors are nonzero (Denition
EEM [453]), so assume for the remainder that p2.
We will prove this result by contradiction (Technique CD [770]). Suppose to the contrary that Sis
a linearly dependent set. Dene Si=fx1;x2;x3; :::; xigand letkbe an integer such that Sk 1=
fx1;x2;x3; :::; xk 1gis linearly independent and Sk=fx1;x2;x3; :::; xkgis linearly dependent. We
have to ask if there is even such an integer k? First, since eigenvectors are nonzero, the set fx1gis
linearly independent. Since we are assuming that S=Spis linearly dependent, there must be an integer
k, 2kp, where the sets Sitransition from linear independence to linear dependence (and stay that
way). In other words, xkis the vector with the smallest index that is a linear combination of just vectors
with smaller indices.
Sincefx1;x2;x3; :::; xkgis linearly dependent there are scalars, a1; a2; a3; :::; ak, some non-zero
(Denition LI [351]), so that
0=a1x1+a2x2+a3x3++akxk
Then,
0= (A kIn)0 Theorem ZVSM [325]
= (A kIn) (a1x1+a2x2+a3x3++akxk) Denition RLD [351]
= (A kIn)a1x1+ (A kIn)a2x2++ (A kIn)akxk Theorem MMDAA [230]
=a1(A kIn)x1+a2(A kIn)x2++ak(A kIn)xk Theorem MMSMM [230]
=a1(Ax1 kInx1) +a2(Ax2 kInx2) ++ak(Axk kInxk) Theorem MMDAA [230]
=a1(Ax1 kx1) +a2(Ax2 kx2) ++ak(Axk kxk) Theorem MMIM [229]
=a1(1x1 kx1) +a2(2x2 kx2) ++ak(kxk kxk) Denition EEM [453]
=a1(1 k)x1+a2(2 k)x2++ak(k k)xk Theorem MMDAA [230]
=a1(1 k)x1+a2(2 k)x2++ak(0)xk Property AICN [759]
=a1(1 k)x1+a2(2 k)x2++ak 1(k 1 k)xk 1+0Theorem ZSSM [324]
=a1(1 k)x1+a2(2 k)x2++ak 1(k 1 k)xk 1 Property Z [318]
This is a relation of linear dependence on the linearly independent set fx1;x2;x3; :::; xk 1g, so the scalars
must all be zero. That is, ai(i k) = 0 for 1ik 1. However, we have the hypothesis that the
eigenvalues are distinct, so i6=kfor 1ik 1. Thusai= 0 for 1ik 1.
This reduces the original relation of linear dependence on fx1;x2;x3; :::; xkgto the simpler equation
akxk=0. By Theorem SMEZV [326] we conclude that ak= 0 or xk=0. Eigenvectors are never the zero
Version 2.30
482 Section PEE Properties of Eigenvalues and Eigenvectors
vector (Denition EEM [453]), so ak= 0. So all of the scalars ai, 1ikare zero, contradicting their in-
troduction as the scalars creating a nontrivial relation of linear dependence on the set fx1;x2;x3; :::; xkg.
With a contradiction in hand, we conclude that Smust be linearly independent.
There is a simple connection between the eigenvalues of a matrix and whether or not the matrix is
nonsingular.
Theorem SMZE
Singular Matrices have Zero Eigenvalues
SupposeAis a square matrix. Then Ais singular if and only if = 0 is an eigenvalue of A.
Proof We have the following equivalences:
Ais singular() there exists x6=0,Ax=0 Denition NSM [73]
() there exists x6=0,Ax= 0x Theorem ZSSM [324]
()= 0 is an eigenvalue of A Denition EEM [453]
With an equivalence about singular matrices we can update our list of equivalences about nonsingular
matrices.
Theorem NME8
Nonsingular Matrix Equivalences, Round 8
Suppose that Ais a square matrix of size n. The following are equivalent.
1.Ais nonsingular.
2.Arow-reduces to the identity matrix.
3. The null space of Acontains only the zero vector, N(A) =f0g.
4. The linear system LS(A;b) has a unique solution for every possible choice of b.
5. The columns of Aare a linearly independent set.
6.Ais invertible.
7. The column space of AisCn,C(A) =Cn.
8. The columns of Aare a basis for Cn.
9. The rank of Aisn,r(A) =n.
10. The nullity of Ais zero,n(A) = 0.
11. The determinant of Ais nonzero, det ( A)6= 0.
12.= 0 is not an eigenvalue of A.
Proof The equivalence of the rst and last statements is the contrapositive of Theorem SMZE [480], so
we are able to improve on Theorem NME7 [446].
Certain changes to a matrix change its eigenvalues in a predictable way.
Version 2.30
Section PEE Properties of Eigenvalues and Eigenvectors 483
Theorem ESMM
Eigenvalues of a Scalar Multiple of a Matrix
SupposeAis a square matrix and is an eigenvalue of A. Thenis an eigenvalue of A.
Proof Letx6=0be one eigenvector of Afor. Then
(A)x=(Ax) Theorem MMSMM [230]
=(x) xeigenvector of A
= ()x Property SMAC [100]
Sox6=0is an eigenvector of Afor the eigenvalue .
Unfortunately, there are not parallel theorems about the sum or product of arbitrary matrices. But we
can prove a similar result for powers of a matrix.
Theorem EOMP
Eigenvalues Of Matrix Powers
SupposeAis a square matrix, is an eigenvalue of A, ands0 is an integer. Then sis an eigenvalue
ofAs.
Proof Letx6=0be one eigenvector of Afor. SupposeAhas sizen. Then we proceed by induction on
s(Technique I [772]). First, for s= 0,
Asx=A0x
=Inx
=x Theorem MMIM [229]
= 1x Property OC [101]
=0x
=sx
sosis an eigenvalue of Asin this special case. If we assume the theorem is true for s, then we nd
As+1x=AsAx
=As(x) xeigenvector of Afor
=(Asx) Theorem MMSMM [230]
=(sx) Induction hypothesis
= (s)x Property SMAC [100]
=s+1x
Sox6=0is an eigenvector of As+1fors+1, and induction tells us the theorem is true for all s0.
While we cannot prove that the sum of two arbitrary matrices behaves in any reasonable way with
regard to eigenvalues, we can work with the sum of dissimilar powers of the same matrix. We have already
seen two connections between eigenvalues and polynomials, in the proof of Theorem EMHE [457] and the
characteristic polynomial (Denition CP [460]). Our next theorem strengthens this connection.
Theorem EPM
Eigenvalues of the Polynomial of a Matrix
SupposeAis a square matrix and is an eigenvalue of A. Letq(x) be a polynomial in the variable x.
Thenq() is an eigenvalue of the matrix q(A).
Proof Letx6=0be one eigenvector of Afor, and write q(x) =a0+a1x+a2x2++amxm. Then
q(A)x=
a0A0+a1A1+a2A2++amAm
x
Version 2.30
484 Section PEE Properties of Eigenvalues and Eigenvectors
= (a0A0)x+ (a1A1)x+ (a2A2)x++ (amAm)x Theorem MMDAA [230]
=a0(A0x) +a1(A1x) +a2(A2x) ++am(Amx) Theorem MMSMM [230]
=a0(0x) +a1(1x) +a2(2x) ++am(mx) Theorem EOMP [481]
= (a00)x+ (a11)x+ (a22)x++ (amm)x Property SMAC [100]
=
a00+a11+a22++amm
x Property DSAC [101]
=q()x
Sox6= 0 is an eigenvector of q(A) for the eigenvalue q().
Example BDE
Building desired eigenvalues
In Example ESMS4 [464] the 4 4 symmetric matrix
C=2
6641 0 1 1
0 1 1 1
1 1 1 0
1 1 0 13
775
is shown to have the three eigenvalues = 3;1; 1. Suppose we wanted a 4 4 matrix that has the three
eigenvalues = 4;0; 2. We can employ Theorem EPM [481] by nding a polynomial that converts 3 to
4, 1 to 0, and 1 to 2. Such a polynomial is called an interpolating polynomial , and in this example
we can use
r(x) =1
4x2+x 5
4
We will not discuss how to concoct this polynomial, but a text on numerical analysis should provide the
details or see Section CF [931]. For now, simply verify that r(3) = 4,r(1) = 0 and r( 1) = 2.
Now compute
r(C) =1
4C2+C 5
4I4
=1
42
6643 2 2 2
2 3 2 2
2 2 3 2
2 2 2 33
775+2
6641 0 1 1
0 1 1 1
1 1 1 0
1 1 0 13
775 5
42
6641 0 0 0
0 1 0 0
0 0 1 0
0 0 0 13
775
=1
22
6641 1 3 3
1 1 3 3
3 3 1 1
3 3 1 13
775
Theorem EPM [481] tells us that if r(x) transforms the eigenvalues in the desired manner, then r(C)
will have the desired eigenvalues. You can check this by computing the eigenvalues of r(C) directly.
Furthermore, notice that the multiplicities are the same, and the eigenspaces of Candr(C) are identical.
Inverses and transposes also behave predictably with regard to their eigenvalues.
Theorem EIM
Eigenvalues of the Inverse of a Matrix
SupposeAis a square nonsingular matrix and is an eigenvalue of A. Then1
is an eigenvalue of the
matrixA 1.
Proof Notice that since Ais assumed nonsingular, A 1exists by Theorem NI [261], but more importantly,
1
does not involve division by zero since Theorem SMZE [480] prohibits this possibility.
Version 2.30
Section PEE Properties of Eigenvalues and Eigenvectors 485
Letx6=0be one eigenvector of Afor. SupposeAhas sizen. Then
A 1x=A 1(1x) Property OC [101]
=A 1(1
x) Property MICN [759]
=1
A 1(x) Theorem MMSMM [230]
=1
A 1(Ax) Denition EEM [453]
=1
(A 1A)x Theorem MMA [231]
=1
Inx Denition MI [244]
=1
x Theorem MMIM [229]
Sox6= 0 is an eigenvector of A 1for the eigenvalue1
.
The theorems above have a similar style to them, a style you should consider using when confronted
with a need to prove a theorem about eigenvalues and eigenvectors. So far we have been able to reserve the
characteristic polynomial for strictly computational purposes. However, the next theorem, whose statement
resembles the preceding theorems, has an easier proof if we employ the characteristic polynomial and results
about determinants.
Theorem ETM
Eigenvalues of the Transpose of a Matrix
SupposeAis a square matrix and is an eigenvalue of A. Thenis an eigenvalue of the matrix At.
Proof SupposeAhas sizen. Then
pA(x) = det (A xIn) Denition CP [460]
= det
(A xIn)t
Theorem DT [430]
= det
At (xIn)t
Theorem TMA [211]
= det
At xIt
n
Theorem TMSM [212]
= det
At xIn
Denition IM [84]
=pAt(x) Denition CP [460]
SoAandAthave the same characteristic polynomial, and by Theorem EMRCP [461], their eigenvalues
are identical and have equal algebraic multiplicities. Notice that what we have proved here is a bit stronger
than the stated conclusion in the theorem.
If a matrix has only real entries, then the computation of the characteristic polynomial (Denition CP
[460]) will result in a polynomial with coecients that are real numbers. Complex numbers could result as
roots of this polynomial, but they are roots of quadratic factors with real coecients, and as such, come
in conjugate pairs. The next theorem proves this, and a bit more, without mentioning the characteristic
polynomial.
Theorem ERMCP
Eigenvalues of Real Matrices come in Conjugate Pairs
SupposeAis a square matrix with real entries and xis an eigenvector of Afor the eigenvalue . Then x
is an eigenvector of Afor the eigenvalue .
Proof
Ax=Ax Ahas real entries
Version 2.30
486 Section PEE Properties of Eigenvalues and Eigenvectors
=Ax Theorem MMCC [232]
=x x eigenvector of A
=x Theorem CRSM [191]
Soxis an eigenvector of Afor the eigenvalue .
This phenomenon is amply illustrated in Example CEMS6 [466], where the four complex eigenvalues
come in two pairs, and the two basis vectors of the eigenspaces are complex conjugates of each other.
Theorem ERMCP [483] can be a time-saver for computing eigenvalues and eigenvectors of real matrices
with complex eigenvalues, since the conjugate eigenvalue and eigenspace can be inferred from the theorem
rather than computed.
Subsection ME
Multiplicities of Eigenvalues
A polynomial of degree nwill have exactly nroots. From this fact about polynomial equations we can say
more about the algebraic multiplicities of eigenvalues.
Theorem DCP
Degree of the Characteristic Polynomial
Suppose that Ais a square matrix of size n. Then the characteristic polynomial of A,pA(x), has degree
n.
Proof We will prove a more general result by induction (Technique I [772]). Then the theorem will be
true as a special case. We will carefully state this result as a proposition indexed by m,m1.
P(m): Suppose that Ais anmmmatrix whose entries are complex numbers or linear polynomials
in the variable xof the form c x, wherecis a complex number. Suppose further that there are exactly
kentries that contain xand that no row or column contains more than one such entry. Then, when
k=m, det (A) is a polynomial in xof degreem, with leading coecient 1, and when k <m , det (A) is
a polynomial in xof degreekor less.
Base Case: Suppose Ais a 11 matrix. Then its determinant is equal to the lone entry (Denition
DM [428]). When k=m= 1, the entry is of the form c x, a polynomial in xof degreem= 1 with
leading coecient 1. Whenk<m , thenk= 0 and the entry is simply a complex number, a polynomial
of degree 0k. SoP(1) is true.
Induction Step: Assume P(m) is true, and that Ais an (m+ 1)(m+ 1) matrix with kentries of the
formc x. There are two cases to consider.
Supposek=m+ 1. Then every row and every column will contain an entry of the form c x. Suppose
that for the rst row, this entry is in column t. Compute the determinant of Aby an expansion about this
rst row (Denition DM [428]). The term associated with entry tof this row will be of the form
(c x)( 1)1+tdet (A(1jt))
The submatrix A(1jt) is anmmmatrix with k=mterms of the form c x, no more than one per row
or column. By the induction hypothesis, det ( A(1jt)) will be a polynomial in xof degreemwith coecient
1. So this entire term is then a polynomial of degree m+ 1 with leading coecient 1.
The remaining terms (which constitute the sum that is the determinant of A) are products of complex
numbers from the rst row with cofactors built from submatrices that lack the rst row of Aand lack some
column ofA, other than column t. As such, these submatrices are mmmatrices with k=m 1<m
entries of the form c x, no more than one per row or column. Applying the induction hypothesis, we
see that these terms are polynomials in xof degreem 1 or less. Adding the single term from the entry
Version 2.30
Subsection PEE.ME Multiplicities of Eigenvalues 487
in columntwith all these others, we see that det ( A) is a polynomial in xof degreem+ 1 and leading
coecient1.
The second case occurs when k < m + 1. Now there is a row of Athat does not contain an entry of
the formc x. We consider the determinant of Aby expanding about this row (Theorem DER [429]),
whose entries are all complex numbers. The cofactors employed are built from submatrices that are mm
matrices with either kork 1 entries of the form c x, no more than one per row or column. In either
case,km, and we can apply the induction hypothesis to see that the determinants computed for the
cofactors are all polynomials of degree kor less. Summing these contributions to the determinant of A
yields a polynomial in xof degreekor less, as desired.
Denition CP [460] tells us that the characteristic polynomial of an nnmatrix is the determinant of
a matrix having exactly nentries of the form c x, no more than one per row or column. As such we can
applyP(n) to see that the characteristic polynomial has degree n.
Theorem NEM
Number of Eigenvalues of a Matrix
Suppose that Ais a square matrix of size nwith distinct eigenvalues 1; 2; 3; :::; k. Then
kX
i=1A(i) =n
Proof By the denition of the algebraic multiplicity (Denition AME [463]), we can factor the charac-
teristic polynomial as
pA(x) =c(x 1)A(1)(x 2)A(2)(x 3)A(3)(x k)A(k)
wherecis a nonzero constant. (We could prove that c= ( 1)n, but we do not need that specicity right
now. See Exercise PEE.T30 [489]) The left-hand side is a polynomial of degree nby Theorem DCP [484]
and the right-hand side is a polynomial of degreePk
i=1A(i). So the equality of the polynomials' degrees
gives the equalityPk
i=1A(i) =n.
Theorem ME
Multiplicities of an Eigenvalue
Suppose that Ais a square matrix of size nandis an eigenvalue. Then
1
A()A()n
Proof Sinceis an eigenvalue of A, there is an eigenvector of Afor,x. Then x2EA(), so
A()1,
since we can extend fxginto a basis ofEA() (Theorem ELIS [407]).
To show that
A()A() is the most involved portion of this proof. To this end, let g=
A()
and let x1;x2;x3; :::; xgbe a basis for the eigenspace of ,EA(). Construct another n gvectors,
y1;y2;y3; :::; yn g, so that
fx1;x2;x3; :::; xg;y1;y2;y3; :::; yn gg
is a basis of Cn. This can be done by repeated applications of Theorem ELIS [407]. Finally, dene a matrix
Sby
S= [x1jx2jx3j:::jxgjy1jy2jy3j:::jyn g] = [x1jx2jx3j:::jxgjR]
Version 2.30
488 Section PEE Properties of Eigenvalues and Eigenvectors
whereRis ann(n g) matrix whose columns are y1;y2;y3; :::; yn g. The columns of Sare linearly
independent by design, so Sis nonsingular (Theorem NMLIC [159]) and therefore invertible (Theorem NI
[261]). Then,
[e1je2je3j:::jen] =In
=S 1S
=S 1[x1jx2jx3j:::jxgjR]
= [S 1x1jS 1x2jS 1x3j:::jS 1xgjS 1R]
So
S 1xi=ei1ig ()
Preparations in place, we compute the characteristic polynomial of A,
pA(x) = det (A xIn) Denition CP [460]
= 1 det (A xIn) Property OCN [759]
= det (In) det (A xIn) Denition DM [428]
= det
S 1S
det (A xIn) Denition MI [244]
= det
S 1
det (S) det (A xIn) Theorem DRMM [447]
= det
S 1
det (A xIn) det (S) Property CMCN [758]
= det
S 1(A xIn)S
Theorem DRMM [447]
= det
S 1AS S 1xInS
Theorem MMDAA [230]
= det
S 1AS xS 1InS
Theorem MMSMM [230]
= det
S 1AS xS 1S
Theorem MMIM [229]
= det
S 1AS xIn
Denition MI [244]
=pS 1AS(x) Denition CP [460]
What can we learn then about the matrix S 1AS?
S 1AS=S 1A[x1jx2jx3j:::jxgjR]
=S 1[Ax1jAx2jAx3j:::jAxgjAR] Denition MM [226]
=S 1[x1jx2jx3j:::jxgjAR] Denition EEM [453]
= [S 1x1jS 1x2jS 1x3j:::jS 1xgjS 1AR] Denition MM [226]
= [S 1x1jS 1x2jS 1x3j:::jS 1xgjS 1AR] Theorem MMSMM [230]
= [e1je2je3j:::jegjS 1AR] S 1S=In, (() above)
Now imagine computing the characteristic polynomial of Aby computing the characteristic polynomial
ofS 1ASusing the form just obtained. The rst gcolumns of S 1ASare all zero, save for a on the
diagonal. So if we compute the determinant by expanding about the rst column, successively, we will get
successive factors of ( x). More precisely, let Tbe the square matrix of size n gthat is formed from
the lastn grows and last n gcolumns of S 1AR. Then
pA(x) =pS 1AS(x) = ( x)gpT(x):
This says that ( x ) is a factor of the characteristic polynomial at leastgtimes, so the algebraic multiplicity
ofas an eigenvalue of Ais greater than or equal to g(Denition AME [463]). In other words,
A() =gA()
Version 2.30
Subsection PEE.EHM Eigenvalues of Hermitian Matrices 489
as desired.
Theorem NEM [485] says that the sum of the algebraic multiplicities for allthe eigenvalues of Ais equal
ton. Since the algebraic multiplicity is a positive quantity, no single algebraic multiplicity can exceed n
without the sum of all of the algebraic multiplicities doing the same.
Theorem MNEM
Maximum Number of Eigenvalues of a Matrix
Suppose that Ais a square matrix of size n. ThenAcannot have more than ndistinct eigenvalues.
Proof Suppose that Ahaskdistinct eigenvalues, 1; 2; 3; :::; k. Then
k=kX
i=11
kX
i=1A(i) Theorem ME [485]
=n Theorem NEM [485]
Subsection EHM
Eigenvalues of Hermitian Matrices
Recall that a matrix is Hermitian (or self-adjoint) if A=A(Denition HM [234]). In the case where A
is a matrix whose entries are all real numbers, being Hermitian is identical to being symmetric (Denition
SYM [211]). Keep this in mind as you read the next two theorems. Their hypotheses could be changed to
\supposeAis a real symmetric matrix."
Theorem HMRE
Hermitian Matrices have Real Eigenvalues
Suppose that Ais a Hermitian matrix and is an eigenvalue of A. Then2R.
Proof Letx6=0be one eigenvector of Afor the eigenvalue . Then by Theorem PIP [196] we know
hx;xi6= 0. So
=1
hx;xihx;xi Property MICN [759]
=1
hx;xihx;xi Theorem IPSM [194]
=1
hx;xihAx;xi Denition EEM [453]
=1
hx;xihx; Axi Theorem HMIP [234]
=1
hx;xihx; xi Denition EEM [453]
=1
hx;xihx;xi Theorem IPSM [194]
= Property MICN [759]
Version 2.30
490 Section PEE Properties of Eigenvalues and Eigenvectors
If a complex number is equal to its conjugate, then it has a complex part equal to zero, and therefore is a
real number.
Notice the appealing symmetry to the justications given for the steps of this proof. In the center is
the ability to pitch a Hermitian matrix from one side of the inner product to the other.
Look back and compare Example ESMS4 [464] and Example CEMS6 [466]. In Example CEMS6 [466]
the matrix has only real entries, yet the characteristic polynomial has roots that are complex numbers,
and so the matrix has complex eigenvalues. However, in Example ESMS4 [464], the matrix has only real
entries, but is also symmetric, and hence Hermitian. So by Theorem HMRE [487], we were guaranteed
eigenvalues that are real numbers.
In many physical problems, a matrix of interest will be real and symmetric, or Hermitian. Then if the
eigenvalues are to represent physical quantities of interest, Theorem HMRE [487] guarantees that these
values will not be complex numbers.
The eigenvectors of a Hermitian matrix also enjoy a pleasing property that we will exploit later.
Theorem HMOE
Hermitian Matrices have Orthogonal Eigenvectors
Suppose that Ais a Hermitian matrix and xandyare two eigenvectors of Afor dierent eigenvalues.
Then xandyare orthogonal vectors.
Proof Letxbe an eigenvector of Aforand let ybe an eigenvector of Afor a dierent eigenvalue .
So we have 6= 0. Then
hx;yi=1
( )hx;yi Property MICN [759]
=1
(hx;yi hx;yi) Property MICN [759]
=1
(hx;yi hx;yi) Theorem IPSM [194]
=1
(hx;yi hx; yi) Theorem HMRE [487]
=1
(hAx;yi hx; Ayi) Denition EEM [453]
=1
(hAx;yi hAx;yi) Theorem HMIP [234]
=1
(0) Property AICN [759]
= 0
This equality says that xandyare orthogonal vectors (Denition OV [196]).
Notice again how the key step in this proof is the fundamental property of a Hermitian matrix (Theorem
HMIP [234]) | the ability to swap Aacross the two arguments of the inner product. We'll build on these
results and continue to see some more interesting properties in Section OD [675].
Subsection READ
Reading Questions
1. How can you identify a nonsingular matrix just by looking at its eigenvalues?
2. How many dierent eigenvalues may a square matrix of size nhave?
3. What is amazing about the eigenvalues of a Hermitian matrix and why is it amazing?
Version 2.30
Subsection PEE.EXC Exercises 491
Subsection EXC
Exercises
T10 Suppose that Ais a square matrix. Prove that the constant term of the characteristic polynomial
ofAis equal to the determinant of A.
Contributed by Robert Beezer Solution [490]
T20 Suppose that Ais a square matrix. Prove that a single vector may not be an eigenvector of Afor
two dierent eigenvalues.
Contributed by Robert Beezer Solution [490]
T22 Suppose that Uis a unitary matrix with eigenvalue . Prove that had modulus 1, i.e. jj= 1.
This says that all of the eigenvalues of a unitary matrix lie on the unit circle of the complex plane.
Contributed by Robert Beezer
T30 Theorem DCP [484] tells us that the characteristic polynomial of a square matrix of size nhas
degreen. By suitably augmenting the proof of Theorem DCP [484] prove that the coecient of xnin the
characteristic polynomial is ( 1)n.
Contributed by Robert Beezer
T50 Theorem EIM [482] says that if is an eigenvalue of the nonsingular matrix A, then1
is an eigenvalue
ofA 1. Write an alternate proof of this theorem using the characteristic polynomial and without making
reference to an eigenvector of Afor.
Contributed by Robert Beezer Solution [490]
Version 2.30
492 Section PEE Properties of Eigenvalues and Eigenvectors
Subsection SOL
Solutions
T10 Contributed by Robert Beezer Statement [489]
Suppose that the characteristic polynomial of Ais
pA(x) =a0+a1x+a2x2++anxn
Then
a0=a0+a1(0) +a2(0)2++an(0)n
=pA(0)
= det (A 0In) Denition CP [460]
= det (A)
T20 Contributed by Robert Beezer Statement [489]
Suppose that the vector x6=0is an eigenvector of Afor the two eigenvalues and, where6=. Then
6= 0, and we also have
0=Ax Ax Property AIC [100]
=x x Denition EEM [453]
= ( )x Property DSAC [101]
By Theorem SMEZV [326], either = 0 or x=0, which are both contradictions.
T50 Contributed by Robert Beezer Statement [489]
Sinceis an eigenvalue of a nonsingular matrix, 6= 0 (Theorem SMZE [480]). Ais invertible (Theorem
NI [261]), and so Ais invertible (Theorem MISM [252]). Thus Ais nonsingular (Theorem NI [261])
and det ( A)6= 0 (Theorem SMZD [445]).
pA 11
= det
A 1 1
In
Denition CP [460]
= 1 det
A 1 1
In
Property OCN [759]
=1
det ( A)det ( A) det
A 1 1
In
Property MICN [759]
=1
det ( A)det
( A)
A 1 1
In
Theorem DRMM [447]
=1
det ( A)det
AA 1 ( A)1
In
Theorem MMDAA [230]
=1
det ( A)det
In ( A)1
In
Denition MI [244]
=1
det ( A)det
In+1
AIn
Theorem MMSMM [230]
=1
det ( A)det ( In+ 1AIn) Property MICN [759]
=1
det ( A)det ( In+AIn) Property OCN [759]
Version 2.30
Subsection PEE.SOL Solutions 493
=1
det ( A)det ( In+A) Theorem MMIM [229]
=1
det ( A)det (A In) Property ACM [209]
=1
det ( A)pA() Denition CP [460]
=1
det ( A)0 Theorem EMRCP [461]
= 0 Property ZCN [759]
So1
is a root of the characteristic polynomial of A 1and so is an eigenvalue of A 1. This proof is due to
Sara Bucht.
Version 2.30
494 Section PEE Properties of Eigenvalues and Eigenvectors
Version 2.30
Section SD Similarity and Diagonalization 495
Section SD
Similarity and Diagonalization
This section's topic will perhaps seem out of place at rst, but we will make the connection soon with
eigenvalues and eigenvectors. This is also our rst look at one of the central ideas of Chapter R [603].
Subsection SM
Similar Matrices
The notion of matrices being \similar" is a lot like saying two matrices are row-equivalent. Two similar
matrices are not equal, but they share many important properties. This section, and later sections in
Chapter R [603] will be devoted in part to discovering just what these common properties are.
First, the main denition for this section.
Denition SIM
Similar Matrices
SupposeAandBare two square matrices of size n. ThenAandBaresimilar if there exists a nonsingular
matrix of size n,S, such that A=S 1BS. 4
We will say \ Ais similar to BviaS" when we want to emphasize the role of Sin the relationship
betweenAandB. Also, it doesn't matter if we say Ais similar to B, orBis similar to A. If one statement
is true then so is the other, as can be seen by using S 1in place ofS(see Theorem SER [494] for the careful
proof). Finally, we will refer to S 1BSas asimilarity transformation when we want to emphasize the
waySchangesB. OK, enough about language, let's build a few examples.
Example SMS5
Similar matrices of size 5
If you wondered if there are examples of similar matrices, then it won't be hard to convince you they exist.
Dene
B=2
66664 4 1 3 2 2
1 2 1 3 2
4 1 3 2 2
3 4 2 1 3
3 1 1 1 43
77775S=2
666641 2 1 1 1
0 1 1 2 1
1 3 1 1 1
2 3 3 1 2
1 3 1 2 13
77775
Check that Sis nonsingular and then compute
A=S 1BS
=2
6666410 1 0 2 5
1 0 1 0 0
3 0 2 1 3
0 0 1 0 1
4 1 1 1 13
777752
66664 4 1 3 2 2
1 2 1 3 2
4 1 3 2 2
3 4 2 1 3
3 1 1 1 43
777752
666641 2 1 1 1
0 1 1 2 1
1 3 1 1 1
2 3 3 1 2
1 3 1 2 13
77775
=2
66664 10 27 29 80 25
2 6 6 10 2
3 11 9 14 9
1 13 0 10 1
11 35 6 49 193
77775
Version 2.30
496 Section SD Similarity and Diagonalization
So by this construction, we know that AandBare similar.
Let's do that again.
Example SMS3
Similar matrices of size 3
Dene
B=2
4 13 8 4
12 7 4
24 16 73
5 S=2
41 1 2
2 1 3
1 2 03
5
Check that Sis nonsingular and then compute
A=S 1BS
=2
4 6 4 1
3 2 1
5 3 13
52
4 13 8 4
12 7 4
24 16 73
52
41 1 2
2 1 3
1 2 03
5
=2
4 1 0 0
0 3 0
0 0 13
5
So by this construction, we know that AandBare similar. But before we move on, look at how pleasing the
form ofAis. Not convinced? Then consider that several computations related to Aare especially easy. For
example, in the spirit of Example DUTM [432], det ( A) = ( 1)(3)( 1) = 3. Similarly, the characteristic
polynomial is straightforward to compute by hand, pA(x) = ( 1 x)(3 x)( 1 x) = (x 3)(x+1)2and
since the result is already factored, the eigenvalues are transparently = 3; 1. Finally, the eigenvectors
ofAare just the standard unit vectors (Denition SUV [197]).
Subsection PSM
Properties of Similar Matrices
Similar matrices share many properties and it is these theorems that justify the choice of the word \similar."
First we will show that similarity is an equivalence relation . Equivalence relations are important in the
study of various algebras and can always be regarded as a kind of weak version of equality. Sort of alike, but
not quite equal. The notion of two matrices being row-equivalent is an example of an equivalence relation
we have been working with since the beginning of the course (see Exercise RREF.T11 [47]). Row-equivalent
matrices are not equal, but they are a lot alike. For example, row-equivalent matrices have the same rank.
Formally, an equivalence relation requires three conditions hold: re
exive, symmetric and transitive. We
will illustrate these as we prove that similarity is an equivalence relation.
Theorem SER
Similarity is an Equivalence Relation
SupposeA,BandCare square matrices of size n. Then
1.Ais similar to A. (Re
exive)
2. IfAis similar to B, thenBis similar to A. (Symmetric)
3. IfAis similar to BandBis similar to C, thenAis similar to C. (Transitive)
Version 2.30
Subsection SD.PSM Properties of Similar Matrices 497
Proof To see that Ais similar to A, we need only demonstrate a nonsingular matrix that eects a similarity
transformation of AtoA.Inis nonsingular (since it row-reduces to the identity matrix, Theorem NMRRI
[84]), and
I 1
nAIn=InAIn=A
If we assume that Ais similar to B, then we know there is a nonsingular matrix Sso thatA=S 1BS
by Denition SIM [493]. By Theorem MIMI [251], S 1is invertible, and by Theorem NI [261] is therefore
nonsingular. So
(S 1) 1A(S 1) =SAS 1Theorem MIMI [251]
=SS 1BSS 1Denition SIM [493]
=
SS 1
B
SS 1
Theorem MMA [231]
=InBIn Denition MI [244]
=B Theorem MMIM [229]
and we see that Bis similar to A.
Assume that Ais similar to B, andBis similar to C. This gives us the existence of two nonsingular
matrices,SandR, such that A=S 1BSandB=R 1CR, by Denition SIM [493]. (Notice how we have
to assumeS6=R, as will usually be the case.) Since SandRare invertible, so too RSis invertible by
Theorem SS [250] and then nonsingular by Theorem NI [261]. Now
(RS) 1C(RS) =S 1R 1CRS Theorem SS [250]
=S 1
R 1CR
S Theorem MMA [231]
=S 1BS Denition SIM [493]
=A
soAis similar to Cvia the nonsingular matrix RS.
Here's another theorem that tells us exactly what sorts of properties similar matrices share.
Theorem SMEE
Similar Matrices have Equal Eigenvalues
SupposeAandBare similar matrices. Then the characteristic polynomials of AandBare equal, that is,
pA(x) =pB(x).
Proof Letndenote the size of AandB. SinceAandBare similar, there exists a nonsingular matrix
S, such that A=S 1BS(Denition SIM [493]). Then
pA(x) = det (A xIn) Denition CP [460]
= det
S 1BS xIn
Denition SIM [493]
= det
S 1BS xS 1InS
Theorem MMIM [229]
= det
S 1BS S 1xInS
Theorem MMSMM [230]
= det
S 1(B xIn)S
Theorem MMDAA [230]
= det
S 1
det (B xIn) det (S) Theorem DRMM [447]
= det
S 1
det (S) det (B xIn) Property CMCN [758]
= det
S 1S
det (B xIn) Theorem DRMM [447]
= det (In) det (B xIn) Denition MI [244]
= 1 det (B xIn) Denition DM [428]
Version 2.30
498 Section SD Similarity and Diagonalization
=pB(x) Denition CP [460]
So similar matrices not only have the same setof eigenvalues, the algebraic multiplicities of these
eigenvalues will also be the same. However, be careful with this theorem. It is tempting to think the
converse is true, and argue that if two matrices have the same eigenvalues, then they are similar. Not so,
as the following example illustrates.
Example EENS
Equal eigenvalues, not similar
Dene
A=1 1
0 1
B=1 0
0 1
and check that
pA(x) =pB(x) = 1 2x+x2= (x 1)2
and soAandBhave equal characteristic polynomials. If the converse of Theorem SMEE [495] were true,
thenAandBwould be similar. Suppose this is the case. More precisely, suppose there is a nonsingular
matrixSso thatA=S 1BS. Then
A=S 1BS=S 1I2S=S 1S=I2
ClearlyA6=I2and this contradiction tells us that the converse of Theorem SMEE [495] is false.
Subsection D
Diagonalization
Good things happen when a matrix is similar to a diagonal matrix. For example, the eigenvalues of the
matrix are the entries on the diagonal of the diagonal matrix. And it can be a much simpler matter to
compute high powers of the matrix. Diagonalizable matrices are also of interest in more abstract settings.
Here are the relevant denitions, then our main theorem for this section.
Denition DIM
Diagonal Matrix
Suppose that Ais a square matrix. Then Ais adiagonal matrix if [A]ij= 0 whenever i6=j.4
Denition DZM
Diagonalizable Matrix
SupposeAis a square matrix. Then Aisdiagonalizable ifAis similar to a diagonal matrix. 4
Example DAB
Diagonalization of Archetype B
Archetype B [786] has a 3 3 coecient matrix
B=2
4 7 6 12
5 5 7
1 0 43
5
and is similar to a diagonal matrix, as can be seen by the following computation with the nonsingular
matrixS,
S 1BS=2
4 5 3 2
3 2 1
1 1 13
5 12
4 7 6 12
5 5 7
1 0 43
52
4 5 3 2
3 2 1
1 1 13
5
Version 2.30
Subsection SD.D Diagonalization 499
=2
4 1 1 1
2 3 1
1 2 13
52
4 7 6 12
5 5 7
1 0 43
52
4 5 3 2
3 2 1
1 1 13
5
=2
4 1 0 0
0 1 0
0 0 23
5
Example SMS3 [494] provides yet another example of a matrix that is subjected to a similarity trans-
formation and the result is a diagonal matrix. Alright, just how would we nd the magic matrix Sthat
can be used in a similarity transformation to produce a diagonal matrix? Before you read the statement
of the next theorem, you might study the eigenvalues and eigenvectors of Archetype B [786] and compute
the eigenvalues and eigenvectors of the matrix in Example SMS3 [494].
Theorem DC
Diagonalization Characterization
SupposeAis a square matrix of size n. ThenAis diagonalizable if and only if there exists a linearly
independent set Sthat contains neigenvectors of A.
Proof (() LetS=fx1;x2;x3; :::; xngbe a linearly independent set of eigenvectors of Afor the
eigenvalues 1; 2; 3; :::; n. Recall Denition SUV [197] and dene
R= [x1jx2jx3j:::jxn]
D=2
66666410 0 0
020 0
0 03 0
............
0 0 0n3
777775= [1e1j2e2j3e3j:::jnen]
The columns of Rare the vectors of the linearly independent set Sand so by Theorem NMLIC [159] the
matrixRis nonsingular. By Theorem NI [261] we know R 1exists.
R 1AR=R 1A[x1jx2jx3j:::jxn]
=R 1[Ax1jAx2jAx3j:::jAxn] Denition MM [226]
=R 1[1x1j2x2j3x3j:::jnxn] Denition EEM [453]
=R 1[1Re1j2Re2j3Re3j:::jnRen] Denition MVP [223]
=R 1[R(1e1)jR(2e2)jR(3e3)j:::jR(nen)] Theorem MMSMM [230]
=R 1R[1e1j2e2j3e3j:::jnen] Denition MM [226]
=InD Denition MI [244]
=D Theorem MMIM [229]
This says that Ais similar to the diagonal matrix Dvia the nonsingular matrix R. ThusAis diagonalizable
(Denition DZM [496]).
()) Suppose that Ais diagonalizable, so there is a nonsingular matrix of size n
T= [y1jy2jy3j:::jyn]
Version 2.30
500 Section SD Similarity and Diagonalization
and a diagonal matrix (recall Denition SUV [197])
E=2
666664d10 0 0
0d20 0
0 0d3 0
............
0 0 0dn3
777775= [d1e1jd2e2jd3e3j:::jdnen]
such thatT 1AT=E. Then consider,
[Ay1jAy2jAy3j:::jAyn] =A[y1jy2jy3j:::jyn] Denition MM [226]
=AT
=InAT Theorem MMIM [229]
=TT 1AT Denition MI [244]
=TE
=T[d1e1jd2e2jd3e3j:::jdnen]
= [T(d1e1)jT(d2e2)jT(d3e3)j:::jT(dnen)] Denition MM [226]
= [d1Te1jd2Te2jd3Te3j:::jdnTen] Denition MM [226]
= [d1y1jd2y2jd3y3j:::jdnyn] Denition MVP [223]
This equality of matrices (Denition ME [207]) allows us to conclude that the individual columns are equal
vectors (Denition CVE [98]). That is, Ayi=diyifor 1in. In other words, yiis an eigenvector of
Afor the eigenvalue di, 1in. (Why can't yi=0?). Because Tis nonsingular, the set containing T's
columns,S=fy1;y2;y3; :::; yng, is a linearly independent set (Theorem NMLIC [159]). So the set S
has all the required properties.
Notice that the proof of Theorem DC [497] is constructive. To diagonalize a matrix, we need only locate
nlinearly independent eigenvectors. Then we can construct a nonsingular matrix using the eigenvectors
as columns ( R) so thatR 1ARis a diagonal matrix ( D). The entries on the diagonal of Dwill be the
eigenvalues of the eigenvectors used to create R,in the same order as the eigenvectors appear in R. We
illustrate this by diagonalizing some matrices.
Example DMS3
Diagonalizing a matrix of size 3
Consider the matrix
F=2
4 13 8 4
12 7 4
24 16 73
5
of Example CPMS3 [460], Example EMS3 [461] and Example ESMS3 [462]. F's eigenvalues and eigenspaces
are
= 3 EF(3) =*8
<
:2
4 1
21
2
13
59
=
;+
= 1 EF( 1) =*8
<
:2
4 2
3
1
03
5;2
4 1
3
0
13
59
=
;+
Version 2.30
Subsection SD.D Diagonalization 501
Dene the matrix Sto be the 33 matrix whose columns are the three basis vectors in the eigenspaces
forF,
S=2
4 1
2 2
3 1
31
21 0
1 0 13
5
Check that Sis nonsingular (row-reduces to the identity matrix, Theorem NMRRI [84] or has a nonzero
determinant, Theorem SMZD [445]). Then the three columns of Sare a linearly independent set (Theorem
NMLIC [159]). By Theorem DC [497] we now know that Fis diagonalizable. Furthermore, the construction
in the proof of Theorem DC [497] tells us that if we apply the matrix StoFin a similarity transformation,
the result will be a diagonal matrix with the eigenvalues of Fon the diagonal. The eigenvalues appear on
the diagonal of the matrix in the same order as the eigenvectors appear in S. So,
S 1FS=2
4 1
2 2
3 1
31
21 0
1 0 13
5 12
4 13 8 4
12 7 4
24 16 73
52
4 1
2 2
3 1
31
21 0
1 0 13
5
=2
46 4 2
3 1 1
6 4 13
52
4 13 8 4
12 7 4
24 16 73
52
4 1
2 2
3 1
31
21 0
1 0 13
5
=2
43 0 0
0 1 0
0 0 13
5
Note that the above computations can be viewed two ways. The proof of Theorem DC [497] tells us that
the four matrices ( F,S,F 1and the diagonal matrix) willinteract the way we have written the equation.
Or as an example, we can actually perform the computations to verify what the theorem predicts.
The dimension of an eigenspace can be no larger than the algebraic multiplicity of the eigenvalue by
Theorem ME [485]. When every eigenvalue's eigenspace is this large, then we can diagonalize the matrix,
and only then. Three examples we have seen so far in this section, Example SMS5 [493], Example DAB
[496] and Example DMS3 [498], illustrate the diagonalization of a matrix, with varying degrees of detail
about just how the diagonalization is achieved. However, in each case, you can verify that the geometric
and algebraic multiplicities are equal for every eigenvalue. This is the substance of the next theorem.
Theorem DMFE
Diagonalizable Matrices have Full Eigenspaces
SupposeAis a square matrix. Then Ais diagonalizable if and only if
A() =A() for every eigenvalue
ofA.
Proof SupposeAhas sizenandkdistinct eigenvalues, 1; 2; 3; :::; k. LetSi=
xi1;xi2;xi3; :::; xi
A(i)
,
denote a basis for the eigenspace of i,EA(i), for 1ik. Then
S=S1[S2[S3[[Sk
is a set of eigenvectors for A. A vector cannot be an eigenvector for two dierent eigenvalues (see Exercise
EE.T20 [472]) so Si\Sj=;wheneveri6=j. In other words, Sis a disjoint union of Si, 1ik.
(() The size of Sis
jSj=kX
i=1
A(i) Sdisjoint union of Si
=kX
i=1A(i) Hypothesis
Version 2.30
502 Section SD Similarity and Diagonalization
=n Theorem NEM [485]
We next show that Sis a linearly independent set. So we will begin with a relation of linear dependence
onS, using doubly-subscripted scalars and eigenvectors,
0=
a11x11+a12x12++a1
A(1)x1
A(1)
+
a21x21+a22x22++a2
A(2)x2
A(2)
++
ak1xk1+ak2xk2++ak
A(k)xk
A(k)
Dene the vectors yi, 1ikby
y1=
a11x11+a12x12+a13x13++a
A(11)x1
A(1)
y2=
a21x21+a22x22+a23x23++a
A(22)x2
A(2)
y3=
a31x31+a32x32+a33x33++a
A(33)x3
A(3)
...
yk=
ak1xk1+ak2xk2+ak3xk3++a
A(kk)xk
A(k)
Then the relation of linear dependence becomes
0=y1+y2+y3++yk
Since the eigenspace EA(i) is closed under vector addition and scalar multiplication, yi2EA(i), 1
ik. Thus, for each i, the vector yiis an eigenvector of Afori, or is the zero vector. Recall that sets
of eigenvectors whose eigenvalues are distinct form a linearly independent set by Theorem EDELI [479].
Should any (or some) yibe nonzero, the previous equation would provide a nontrivial relation of linear
dependence on a set of eigenvectors with distinct eigenvalues, contradicting Theorem EDELI [479]. Thus
yi=0, 1ik.
Each of the kequations, yi=0is a relation of linear dependence on the corresponding set Si, a set
of basis vectors for the eigenspace EA(i), which is therefore linearly independent. From these relations
of linear dependence on linearly independent sets we conclude that the scalars are all zero, more precisely,
aij= 0, 1j
A(i) for 1ik. This establishes that our original relation of linear dependence on
Shas only the trivial relation of linear dependence, and hence Sis a linearly independent set.
We have determined that Sis a set ofnlinearly independent eigenvectors for A, and so by Theorem
DC [497] is diagonalizable.
()) Now we assume that Ais diagonalizable. Aiming for a contradiction (Technique CD [770]), suppose
that there is at least one eigenvalue, say t, such that
A(t)6=A(t). By Theorem ME [485] we must
have
A(t)<A(t), and
A(i)A(i) for 1ik,i6=t.
SinceAis diagonalizable, Theorem DC [497] guarantees a set of nlinearly independent vectors, all of
which are eigenvectors of A. Letnidenote the number of eigenvectors in Sthat are eigenvectors for i,
and recall that a vector cannot be an eigenvector for two dierent eigenvalues (Exercise EE.T20 [472]). S
is a linearly independent set, so the the subset Sicontaining the nieigenvectors for imust also be linearly
independent. Because the eigenspace EA(i) has dimension
A(i) andSiis a linearly independent subset
inEA(i), Theorem G [407] tells us that ni
A(i), for 1ik. Putting all these facts together gives,
n=n1+n2+n3++nt++nk Denition SU [763]
A(1) +
A(2) +
A(3) ++
A(t) ++
A(k) Theorem G [407]
<A(1) +A(2) +A(3) ++A(t) ++A(k) Theorem ME [485]
=n Theorem NEM [485]
This is a contradiction (we can't have n<n !) and so our assumption that some eigenspace had less than
full dimension was false.
Example SEE [453], Example CAEHW [458], Example ESMS3 [462], Example ESMS4 [464], Example
DEMS5 [468], Archetype B [786], Archetype F [803], Archetype K [825] and Archetype L [829] are all
Version 2.30
Subsection SD.D Diagonalization 503
examples of matrices that are diagonalizable and that illustrate Theorem DMFE [499]. While we have
provided many examples of matrices that are diagonalizable, especially among the archetypes, there are
many matrices that are not diagonalizable. Here's one now.
Example NDMS4
A non-diagonalizable matrix of size 4
In Example EMMS4 [463] the matrix
B=2
664 2 1 2 4
12 1 4 9
6 5 2 4
3 4 5 103
775
was determined to have characteristic polynomial
pB(x) = (x 1)(x 2)3
and an eigenspace for = 2 of
EB(2) =*8
>><
>>:2
664 1
2
1
1
2
13
7759
>>=
>>;+
So the geometric multiplicity of = 2 is
B(2) = 1, while the algebraic multiplicity is B(2) = 3. By
Theorem DMFE [499], the matrix Bis not diagonalizable.
Archetype A [781] is the lone archetype with a square matrix that is not diagonalizable, as the algebraic
and geometric multiplicities of the eigenvalue = 0 dier. Example HMEM5 [465] is another example of a
matrix that cannot be diagonalized due to the dierence between the geometric and algebraic multiplicities
of= 2, as is Example CEMS6 [466] which has two complex eigenvalues, each with diering multiplicities.
Likewise, Example EMMS4 [463] has an eigenvalue with dierent algebraic and geometric multiplicities
and so cannot be diagonalized.
Theorem DED
Distinct Eigenvalues implies Diagonalizable
SupposeAis a square matrix of size nwithndistinct eigenvalues. Then Ais diagonalizable.
Proof Let1; 2; 3; :::; ndenote the ndistinct eigenvalues of A. Then by Theorem NEM [485] we
haven=Pn
i=1A(i), which implies that A(i) = 1, 1in. From Theorem ME [485] it follows that
A(i) = 1, 1in. So
A(i) =A(i), 1inand Theorem DMFE [499] says Ais diagonalizable.
Example DEHD
Distinct eigenvalues, hence diagonalizable
In Example DEMS5 [468] the matrix
H=2
6666415 18 8 6 5
5 3 1 1 3
0 4 5 4 2
43 46 17 14 15
26 30 12 8 103
77775
has characteristic polynomial
pH(x) =x(x 2)(x 1)(x+ 1)(x+ 3)
Version 2.30
504 Section SD Similarity and Diagonalization
and so is a 55 matrix with 5 distinct eigenvalues. By Theorem DED [501] we know Hmust be diago-
nalizable. But just for practice, we exhibit the diagonalization itself. The matrix Scontains eigenvectors
ofHas columns, one from each eigenspace, guaranteeing linear independent columns and thus the non-
singularity of S. The diagonal matrix has the eigenvalues of Hin the same order that their respective
eigenvectors appear as the columns of S. Notice that we are using the versions of the eigenvectors from
Example DEMS5 [468] that have integer entries.
S 1HS
=2
666642 1 1 1 1
1 0 2 0 1
2 0 2 1 2
4 1 0 2 1
2 2 1 2 13
77775 12
6666415 18 8 6 5
5 3 1 1 3
0 4 5 4 2
43 46 17 14 15
26 30 12 8 103
777752
666642 1 1 1 1
1 0 2 0 1
2 0 2 1 2
4 1 0 2 1
2 2 1 2 13
77775
=2
66664 3 3 1 1 1
1 2 1 0 1
5 4 1 1 2
10 10 3 2 4
7 6 1 1 33
777752
6666415 18 8 6 5
5 3 1 1 3
0 4 5 4 2
43 46 17 14 15
26 30 12 8 103
777752
666642 1 1 1 1
1 0 2 0 1
2 0 2 1 2
4 1 0 2 1
2 2 1 2 13
77775
=2
66664 3 0 0 0 0
0 1 0 0 0
0 0 0 0 0
0 0 0 1 0
0 0 0 0 23
77775
Archetype B [786] is another example of a matrix that has as many distinct eigenvalues as its size, and
is hence diagonalizable by Theorem DED [501].
Powers of a diagonal matrix are easy to compute, and when a matrix is diagonalizable, it is almost as
easy. We could state a theorem here perhaps, but we will settle instead for an example that makes the
point just as well.
Example HPDM
High power of a diagonalizable matrix
Suppose that
A=2
66419 0 6 13
33 1 9 21
21 4 12 21
36 2 14 283
775
and we wish to compute A20. Normally this would require 19 matrix multiplications, but since Ais
diagonalizable, we can simplify the computations substantially. First, we diagonalize A. With
S=2
6641 1 2 1
2 3 3 3
1 1 3 3
2 1 4 03
775
we nd
D=S 1AS=2
664 6 1 3 6
0 2 2 3
3 0 1 2
1 1 1 13
7752
66419 0 6 13
33 1 9 21
21 4 12 21
36 2 14 283
7752
6641 1 2 1
2 3 3 3
1 1 3 3
2 1 4 03
775
Version 2.30
Subsection SD.FS Fibonacci Sequences 505
=2
664 1 0 0 0
0 0 0 0
0 0 2 0
0 0 0 13
775
Now we nd an alternate expression for A20,
A20=AAA:::A
=InAInAInAIn:::InAIn
=
SS 1
A
SS 1
A
SS 1
A
SS 1
:::
SS 1
A
SS 1
=S
S 1AS
S 1AS
S 1AS
:::
S 1AS
S 1
=SDDD:::DS 1
=SD20S 1
and sinceDis a diagonal matrix, powers are much easier to compute,
=S2
664 1 0 0 0
0 0 0 0
0 0 2 0
0 0 0 13
77520
S 1
=S2
664( 1)200 0 0
0 (0)200 0
0 0 (2)200
0 0 0 (1)203
775S 1
=2
6641 1 2 1
2 3 3 3
1 1 3 3
2 1 4 03
7752
6641 0 0 0
0 0 0 0
0 0 1048576 0
0 0 0 13
7752
664 6 1 3 6
0 2 2 3
3 0 1 2
1 1 1 13
775
=2
6646291451 2 2097148 4194297
9437175 5 3145719 6291441
9437175 2 3145728 6291453
12582900 2 4194298 83885963
775
Notice how we eectively replaced the twentieth power of Aby the twentieth power of D, and how a high
power of a diagonal matrix is just a collection of powers of scalars on the diagonal. The price we pay for
this simplication is the need to diagonalize the matrix (by computing eigenvalues and eigenvectors) and
nding the inverse of the matrix of eigenvectors. And we still need to do two matrix products. But the
higher the power, the greater the savings.
Subsection FS
Fibonacci Sequences
Example FSCF
Fibonacci sequence, closed form
TheFibonacci sequence is a sequence of integers dened recursively by
a0= 0 a1= 1 an+1=an+an 1; n1
Version 2.30
506 Section SD Similarity and Diagonalization
So the initial portion of the sequence is 0 ;1;1;2;3;5;8;13;21; :::. In this subsection we will illustrate
an application of eigenvalues and diagonalization through the determination of a closed-form expression
for an arbitrary term of this sequence.
To begin, verify that for any n1 the recursive statement above establishes the truth of the statement
an
an+1
=0 1
1 1an 1
an
LetAdenote this 22 matrix. Through repeated applications of the statement above we have
an
an+1
=Aan 1
an
=A2an 2
an 1
=A3an 3
an 2
==Ana0
a1
In preparation for working with this high power of A, not unlike in Example HPDM [502], we will diago-
nalizeA. The characteristic polynomial of AispA(x) =x2 x 1, with roots (the eigenvalues of Aby
Theorem EMRCP [461])
=1 +p
5
2=1 p
5
2
With two distinct eigenvalues, Theorem DED [501] implies that Ais diagonalizable. It will be easier to
compute with these eigenvalues once you conrm the following properties (all but the last can be derived
from the fact that andare roots of the characteristic polynomial, in a factored or unfactored form)
+= 1 = 1 1 + =21 +=2 =p
5
Then eigenvectors of A(forand, respectively) are
1
1
which can be easily conrmed, as we demonstrate for the eigenvector for ,
0 1
1 11
=
1 +
=
2
=1
From the proof of Theorem DC [497] we know Acan be diagonalized by a matrix Swith these eigenvectors
as columns, giving D=S 1AS. We listS,S 1and the diagonal matrix D,
S=1 1
S 1=1
1
1
D=0
0
OK, we have everything in place now. The main step in the following is to replace AbySDS 1. Here we
go,
an
an+1
=Ana0
a1
=
SDS 1na0
a1
=SDS 1SDS 1SDS 1SDS 1a0
a1
=SDDDDS 1a0
a1
Version 2.30
Subsection SD.FS Fibonacci Sequences 507
=SDnS 1a0
a1
=1 1
0
0n1
1
1a0
a1
=1
1 1
n0
0n 1
10
1
=1
1 1
n0
0n1
1
=1
1 1
n
n
=1
n n
n+1 n+1
Performing the scalar multiplication and equating the rst entries of the two vectors, we arrive at the
closed form expression
an=1
(n n)
=1p
5
1 +p
5
2!n
1 p
5
2!n!
=1
2np
5
1 +p
5n
1 p
5n
Notice that it does not matter whether we use the equality of the rst or second entries of the vectors, we
will arrive at the same formula, once in terms of nand again in terms of n+ 1. Also, our denition clearly
describes a sequence that will only contain integers, yet the presence of the irrational numberp
5 might
make us suspicious. But no, our expression for anwill always yield an integer!
The Fibonacci sequence, and generalizations of it, have been extensively studied (Fibonacci lived in the
12th and 13th centuries). There are many ways to derive the closed-form expression we just found, and
our approach may not be the most ecient route. But it is a nice demonstration of how diagonalization
can be used to solve a problem outside the eld of linear algebra.
We close this section with a comment about an important upcoming theorem that we prove in Chapter
R [603]. A consequence of Theorem OD [681] is that every Hermitian matrix (Denition HM [234]) is diag-
onalizable (Denition DZM [496]), and the similarity transformation that accomplishes the diagonalization
uses a unitary matrix (Denition UM [262]). This means that for every Hermitian matrix of size nthere
is a basis of Cnthat is composed entirely of eigenvectors for the matrix and also forms an orthonormal set
(Denition ONS [201]). Notice that for matrices with only real entries, we only need the hypothesis that
the matrix is symmetric (Denition SYM [211]) to reach this conclusion (Example ESMS4 [464]). Can you
imagine a prettier basis for use with a matrix? I can't.
These results in Section OD [675] explain much of our recurring interest in orthogonality, and make
the section a high point in your study of linear algebra. A precise statement of this diagonalization result
applies to a slightly broader class of matrices, known as \normal" matrices (Denition NRML [680]),
which are matrices that commute with their adjoints. With this expanded category of matrices, the result
becomes an equivalence (Technique E [768]). See Theorem OD [681] and Theorem OBNM [683] in Section
OD [675] for all the details.
Version 2.30
508 Section SD Similarity and Diagonalization
Subsection READ
Reading Questions
1. What is an equivalence relation?
2. State a condition that is equivalent to a matrix being diagonalizable, but is not the denition.
3. Find a diagonal matrix similar to
A= 5 8
4 7
Version 2.30
Subsection SD.EXC Exercises 509
Subsection EXC
Exercises
C20 Consider the matrix Abelow. First, show that Ais diagonalizable by computing the geometric
multiplicities of the eigenvalues and quoting the relevant theorem. Second, nd a diagonal matrix D
and a nonsingular matrix Sso thatS 1AS=D. (See Exercise EE.C20 [471] for some of the necessary
computations.)
A=2
66418 15 33 15
4 8 6 6
9 9 16 9
5 6 9 43
775
Contributed by Robert Beezer Solution [508]
C21 Determine if the matrix Abelow is diagonalizable. If the matrix is diagonalizable, then nd a
diagonal matrix Dthat is similar to A, and provide the invertible matrix Sthat performs the similarity
transformation. You should use your calculator to nd the eigenvalues of the matrix, but try only using
the row-reducing function of your calculator to assist with nding eigenvectors.
A=2
6641 9 9 24
3 27 29 68
1 11 13 26
1 7 7 183
775
Contributed by Robert Beezer Solution [508]
C22 Consider the matrix Abelow. Find the eigenvalues of Ausing a calculator and use these to construct
the characteristic polynomial of A,pA(x). State the algebraic multiplicity of each eigenvalue. Find all of
the eigenspaces for Aby computing expressions for null spaces, only using your calculator to row-reduce
matrices. State the geometric multiplicity of each eigenvalue. Is Adiagonalizable? If not, explain why. If
so, nd a diagonal matrix Dthat is similar to A.
A=2
66419 25 30 5
23 30 35 5
7 9 10 1
3 4 5 13
775
Contributed by Robert Beezer Solution [509]
T15 Suppose that AandBare similar matrices. Prove that A3andB3are similar matrices. Generalize.
Contributed by Robert Beezer Solution [510]
T16 Suppose that AandBare similar matrices, with Anonsingular. Prove that Bis nonsingular, and
thatA 1is similar to B 1.
Contributed by Robert Beezer Solution [510]
T17 Suppose that Bis a nonsingular matrix. Prove that ABis similar to BA.
Contributed by Robert Beezer Solution [510]
Version 2.30
510 Section SD Similarity and Diagonalization
Subsection SOL
Solutions
C20 Contributed by Robert Beezer Statement [507]
Using a calculator, we nd that Ahas three distinct eigenvalues, = 3;2; 1, with= 2 having algebraic
multiplicity two, A(2) = 2. The eigenvalues = 3; 1 have algebraic multiplicity one, and so by Theorem
ME [485] we can conclude that their geometric multiplicities are one as well. Together with the computation
of the geometric multiplicity of = 2 from Exercise EE.C20 [471], we know
A(3) =A(3) = 1
A(2) =A(2) = 2
A( 1) =A( 1) = 1
This satises the hypotheses of Theorem DMFE [499], and so we can conclude that Ais diagonalizable.
A calculator will give us four eigenvectors of A, the two for = 2 being linearly independent presumably.
Or, by hand, we could nd basis vectors for the three eigenspaces. For = 3; 1 the eigenspaces have
dimension one, and so any eigenvector for these eigenvalues will be multiples of the ones we use below. For
= 2 there are many dierent bases for the eigenspace, so your answer could vary. Our eigenvectors are
the basis vectors we would have obtained if we had actually constructed a basis in Exercise EE.C20 [471]
rather than just computing the dimension.
By the construction in the proof of Theorem DC [497], the required matrix Shas columns that are
four linearly independent eigenvectors of Aand the diagonal matrix has the eigenvalues on the diagonal
(in the same order as the eigenvectors in S). Here are the pieces, \doing" the diagonalization,
2
664 1 0 3 6
2 1 1 0
0 0 1 3
1 1 0 13
775 12
66418 15 33 15
4 8 6 6
9 9 16 9
5 6 9 43
7752
664 1 0 3 6
2 1 1 0
0 0 1 3
1 1 0 13
775=2
6643 0 0 0
0 2 0 0
0 0 2 0
0 0 0 13
775
C21 Contributed by Robert Beezer Statement [507]
A calculator will provide the eigenvalues = 2;2;1;0, so we can reconstruct the characteristic polynomial
as
pA(x) = (x 2)2(x 1)x
so the algebraic multiplicities of the eigenvalues are
A(2) = 2 A(1) = 1 A(0) = 1
Now compute eigenspaces by hand, obtaining null spaces for each of the three eigenvalues by constructing
the correct singular matrix (Theorem EMNS [462]),
A 2I4=2
664 1 9 9 24
3 29 29 68
1 11 11 26
1 7 7 163
775RREF !2
6641 0 0 3
2
0 1 15
2
0 0 0 0
0 0 0 03
775
EA(2) =N(A 2I4) =*8
>><
>>:2
6643
2
5
2
0
13
775;2
6640
1
1
03
7759
>>=
>>;+
=*8
>><
>>:2
6643
5
0
23
775;2
6640
1
1
03
7759
>>=
>>;+
A 1I4=2
6640 9 9 24
3 28 29 68
1 11 12 26
1 7 7 173
775RREF !2
6641 0 0 5
3
0 1 013
3
0 0 1 5
3
0 0 0 03
775
Version 2.30
Subsection SD.SOL Solutions 511
EA(1) =N(A I4) =*8
>><
>>:2
6645
3
13
35
3
13
7759
>>=
>>;+
=*8
>><
>>:2
6645
13
5
33
7759
>>=
>>;+
A 0I4=2
6641 9 9 24
3 27 29 68
1 11 13 26
1 7 7 183
775RREF !2
6641 0 0 3
0 1 0 5
0 0 1 2
0 0 0 03
775
EA(0) =N(A I4) =*8
>><
>>:2
6643
5
2
13
7759
>>=
>>;+
From this we can compute the dimensions of the eigenspaces to obtain the geometric multiplicities,
A(2) = 2
A(1) = 1
A(0) = 1
For each eigenvalue, the algebraic and geometric multiplicities are equal and so by Theorem DMFE [499]
we now know that Ais diagonalizable. The construction in Theorem DC [497] suggests we form a matrix
whose columns are eigenvectors of A
S=2
6643 0 5 3
5 1 13 5
0 1 5 2
2 0 3 13
775
Since det (S) = 16= 0, we know that Sis nonsingular (Theorem SMZD [445]), so the columns of Sare a
set of 4 linearly independent eigenvectors of A. By the proof of Theorem SMZD [445] we know
S 1AS=2
6642 0 0 0
0 2 0 0
0 0 1 0
0 0 0 03
775
a diagonal matrix with the eigenvalues of Aalong the diagonal, in the same order as the associated
eigenvectors appear as columns of S.
C22 Contributed by Robert Beezer Statement [507]
A calculator will report = 0 as an eigenvalue of algebraic multiplicity of 2, and = 1 as an eigenvalue
of algebraic multiplicity 2 as well. Since eigenvalues are roots of the characteristic polynomial (Theorem
EMRCP [461]) we have the factored version
pA(x) = (x 0)2(x ( 1))2=x2(x2+ 2x+ 1) =x4+ 2x3+x2
The eigenspaces are then
= 0
A (0)I4=2
66419 25 30 5
23 30 35 5
7 9 10 1
3 4 5 13
775RREF !2
66410 5 5
01 5 4
0 0 0 0
0 0 0 03
775
EA(0) =N(C (0)I4) =*8
>><
>>:2
6645
5
1
03
775;2
6645
4
0
13
7759
>>=
>>;+
Version 2.30
512 Section SD Similarity and Diagonalization
= 1
A ( 1)I4=2
66420 25 30 5
23 29 35 5
7 9 11 1
3 4 5 03
775RREF !2
66410 1 4
01 2 3
0 0 0 0
0 0 0 03
775
EA( 1) =N(C ( 1)I4) =*8
>><
>>:2
6641
2
1
03
775;2
664 4
3
0
13
7759
>>=
>>;+
Each eigenspace above is described by a spanning set obtained through an application of Theorem BNS [160]
and so is a basis for the eigenspace. In each case the dimension, and therefore the geometric multiplicity,
is 2.
For each of the two eigenvalues, the algebraic and geometric multiplicities are equal. Theorem DMFE
[499] says that in this situation the matrix is diagonalizable. We know from Theorem DC [497] that when
we diagonalize Athe diagonal matrix will have the eigenvalues of Aon the diagonal (in some order). So
we can claim that
D=2
6640 0 0 0
0 0 0 0
0 0 1 0
0 0 0 13
775
T15 Contributed by Robert Beezer Statement [507]
By Denition SIM [493] we know that there is a nonsingular matrix Sso thatA=S 1BS. Then
A3= (S 1BS)3
= (S 1BS)(S 1BS)(S 1BS)
=S 1B(SS 1)B(SS 1)BS Theorem MMA [231]
=S 1B(I3)B(I3)BS Denition MI [244]
=S 1BBBS Theorem MMIM [229]
=S 1B3S
This equation says that A3is similar to B3(via the matrix S).
More generally, if Ais similar to B, andmis a non-negative integer, then Amis similar to Bm. This
can be proved using induction (Technique I [772]).
T16 Contributed by Steve Caneld Statement [507]
Abeing similar to Bmeans that there exists an Ssuch thatA=S 1BS. So,B=SAS 1and because S,
A, andS 1are nonsingular, by Theorem NPNT [259], Bis nonsingular.
A 1=
S 1BS 1Denition SIM [493]
=S 1B 1
S 1 1Theorem SS [250] = S 1B 1S Theorem MIMI [251]
Then by Denition SIM [493], A 1is similar to B 1.
T17 Contributed by Robert Beezer Statement [507]
The nonsingular (invertible) matrix Bwill provide the desired similarity transformation,
B 1(BA)B=
B 1B
(AB) Theorem MMA [231]
=InAB Denition MI [244]
Version 2.30
Subsection SD.SOL Solutions 513
=AB Theorem MMIM [229]
Version 2.30
514 Section SD Similarity and Diagonalization
Version 2.30
Annotated Acronyms SD.E Eigenvalues 515
Annotated Acronyms E
Eigenvalues
Theorem EMRCP [461]
Much of what we know about eigenvalues can be traced to analysis of the characteristic polynomial. When
we rst dened eigenvalues, you might have wondered if they were scarce, or abundant. The characteristic
polynomial allows us to answer a question like this with a result like Theorem NEM [485] which tells us
there are always a few eigenvalues, but never too many.
Theorem EMNS [462]
If Theorem EMRCP [461] allows us to learn about eigenvalues through what we know about roots of
polynomials, then Theorem EMNS [462] allows us to learn about eigenvectors, and eigenspaces, from what
we already know about null spaces. These two theorems, along with Denition EEM [453], provide the
starting points for discerning the properties of eigenvalues and eigenvectors (to say nothing of actually
computing them).
Theorem HMRE [487]
As we have remarked before, we choose to include all of the complex numbers in our set of allowed scalars,
whereas many introductory texts restrict their attention to just the real numbers. Here is one of the
payos to this approach. Begin with a matrix, possibly containing complex entries, and require the matrix
to be Hermitian (Denition HM [234]). In the case of only real entries, this boils down to just requiring
the matrix to be symmetric (Denition SYM [211]). Generally, the roots of a characteristic polynomial,
even with all real coecients, can have complex numbers as roots. But for a Hermitian matrix, all of the
eigenvalues are real numbers! When somebody tells you mathematics can be beautiful, this is an example
of what they are talking about.
Theorem DC [497]
Diagonalizing a matrix, or the question of if a matrix is diagonalizable, could be viewed as one of a
handful of central questions in linear algebra. Here we have an unequivocal answer to the question of \if,"
along with a proof containing a construction for the diagonalization. So this theorem is of theoretical and
computational interest. This topic will be important again in Chapter R [603].
Theorem DMFE [499]
Another unequivocal answer to the question of if a matrix is diagonalizable, with perhaps a simpler condi-
tion to test. The proof also tells us how to construct the necessary set of nlinearly independent eigenvectors
| just round up bases for each eigenspace and join them together. No need to test the linear independence
of the combined set.
Version 2.30
516 Section SD Similarity and Diagonalization
Version 2.30
Chapter LT
Linear Transformations
In the next linear algebra course you take, the rst lecture might be a reminder about what a vector space
is (Denition VS [317]), their ten properties, basic theorems and then some examples. The second lecture
would likely be all about linear transformations. While it may seem we have waited a long time to present
what must be a central topic, in truth we have already been working with linear transformations for some
time.
Functions are important objects in the study of calculus, but have been absent from this course until
now (well, not really, it just seems that way). In your study of more advanced mathematics it is nearly
impossible to escape the use of functions | they are as fundamental as sets are.
Section LT
Linear Transformations
Early in Chapter VS [317] we prefaced the denition of a vector space with the comment that it was \one
of the two most important denitions in the entire course." Here comes the other. Any capsule summary
of linear algebra would have to describe the subject as the interplay of linear transformations and vector
spaces. Here we go.
Subsection LT
Linear Transformations
Denition LT
Linear Transformation
Alinear transformation ,T:U!V, is a function that carries elements of the vector space U(called
thedomain ) to the vector space V(called the codomain ), and which has two additional properties
1.T(u1+u2) =T(u1) +T(u2) for all u1;u22U
2.T(u) =T(u) for all u2Uand all2C
(This denition contains Notation LT.) 4
The two dening conditions in the denition of a linear transformation should \feel linear," whatever
that means. Conversely, these two conditions could be taken as exactly what it means to be linear. As
every vector space property derives from vector addition and scalar multiplication, so too, every property
517
518 Section LT Linear Transformations
of a linear transformation derives from these two dening properties. While these conditions may be
reminiscent of how we test subspaces, they really are quite dierent, so do not confuse the two.
Here are two diagrams that convey the essence of the two dening properties of a linear transformation.
In each case, begin in the upper left-hand corner, and follow the arrows around the rectangle to the lower-
right hand corner, taking two dierent routes and doing the indicated operations labeled on the arrows.
There are two results there. For a linear transformation these two expressions are always equal.
u1,u2
u1+u2T(u1),T(u2)
T(u1+u2)=T(u1)+T(u2)T
T+ +
Diagram DLTA. Denition of Linear Transformation, Additive
u
αuT(u)
T(αu)=αT(u)T
Tα α
Diagram DLTM. Denition of Linear Transformation, Multiplicative
A couple of words about notation. Tis the name of the linear transformation, and should be used when
we want to discuss the function as a whole. T(u) is how we talk about the output of the function, it is a
vector in the vector space V. When we write T(x+y) =T(x) +T(y), the plus sign on the left is the
operation of vector addition in the vector space U, since xandyare elements of U. The plus sign on the
right is the operation of vector addition in the vector space V, sinceT(x) andT(y) are elements of the
vector space V. These two instances of vector addition might be wildly dierent.
Let's examine several examples and begin to form a catalog of known linear transformations to work
with.
Example ALT
A linear transformation
DeneT:C3!C2by describing the output of the function for a generic input with the formula
T0
@2
4x1
x2
x33
51
A=2x1+x3
4x2
and check the two dening properties.
T(x+y) =T0
@2
4x1
x2
x33
5+2
4y1
y2
y33
51
A
=T0
@2
4x1+y1
x2+y2
x3+y33
51
A
Version 2.30
Subsection LT.LT Linear Transformations 519
=2(x1+y1) + (x3+y3)
4(x2+y2)
=(2x1+x3) + (2y1+y3)
4x2+ ( 4)y2
=2x1+x3
4x2
+2y1+y3
4y2
=T0
@2
4x1
x2
x33
51
A+T0
@2
4y1
y2
y33
51
A
=T(x) +T(y)
and
T(x) =T0
@2
4x1
x2
x33
51
A
=T0
@2
4x1
x2
x33
51
A
=2(x1) + (x3)
4(x2)
=(2x1+x3)
( 4x2)
=2x1+x3
4x2
=T0
@2
4x1
x2
x33
51
A
=T(x)
So by Denition LT [515], Tis a linear transformation.
It can be just as instructive to look at functions that are notlinear transformations. Since the dening
conditions must be true for allvectors and scalars, it is enough to nd just one situation where the
properties fail.
Example NLT
Not a linear transformation
DeneS:C3!C3by
S0
@2
4x1
x2
x33
51
A=2
44x1+ 2x2
0
x1+ 3x3 23
5
This function \looks" linear, but consider
3S0
@2
41
2
33
51
A= 32
48
0
83
5=2
424
0
243
5
Version 2.30
520 Section LT Linear Transformations
while
S0
@32
41
2
33
51
A=S0
@2
43
6
93
51
A=2
424
0
283
5
So the second required property fails for the choice of = 3 and x=2
41
2
33
5and by Denition LT [515], Sis
not a linear transformation. It is just about as easy to nd an example where the rst dening property
fails (try it!). Notice that it is the \-2" in the third component of the denition of Sthat prevents the
function from being a linear transformation.
Example LTPM
Linear transformation, polynomials to matrices
Dene a linear transformation T:P3!M22by
T
a+bx+cx2+dx3
=a+b a 2c
d b d
We verify the two dening conditions of a linear transformations.
T(x+y) =T
(a1+b1x+c1x2+d1x3) + (a2+b2x+c2x2+d2x3)
=T
(a1+a2) + (b1+b2)x+ (c1+c2)x2+ (d1+d2)x3
=(a1+a2) + (b1+b2) (a1+a2) 2(c1+c2)
d1+d2 (b1+b2) (d1+d2)
=(a1+b1) + (a2+b2) (a1 2c1) + (a2 2c2)
d1+d2 (b1 d1) + (b2 d2)
=a1+b1a1 2c1
d1b1 d1
+a2+b2a2 2c2
d2b2 d2
=T
a1+b1x+c1x2+d1x3
+T
a2+b2x+c2x2+d2x3
=T(x) +T(y)
and
T(x) =T
(a+bx+cx2+dx3)
=T
(a) + (b)x+ (c)x2+ (d)x3
=(a) + (b) (a) 2(c)
d (b) (d)
=(a+b)(a 2c)
d (b d)
=a+b a 2c
d b d
=T
a+bx+cx2+dx3
=T(x)
So by Denition LT [515], Tis a linear transformation.
Example LTPP
Linear transformation, polynomials to polynomials
Dene a function S:P4!P5by
S(p(x)) = (x 2)p(x)
Version 2.30
Subsection LT.LTC Linear Transformation Cartoons 521
Then
S(p(x) +q(x)) = (x 2)(p(x) +q(x)) = (x 2)p(x) + (x 2)q(x) =S(p(x)) +S(q(x))
S(p(x)) = (x 2)(p(x)) = (x 2)p(x) =(x 2)p(x) =S(p(x))
So by Denition LT [515], Sis a linear transformation.
Linear transformations have many amazing properties, which we will investigate through the next few
sections. However, as a taste of things to come, here is a theorem we can prove now and put to use
immediately.
Theorem LTTZZ
Linear Transformations Take Zero to Zero
SupposeT:U!Vis a linear transformation. Then T(0) =0.
Proof The two zero vectors in the conclusion of the theorem are dierent. The rst is from Uwhile the
second is from V. We will subscript the zero vectors in this proof to highlight the distinction. Think about
your objects. (This proof is contributed by Mark Shoemaker).
T(0U) =T(00U) Theorem ZSSM [324] in U
= 0T(0U) Denition LT [515]
=0V Theorem ZSSM [324] in V
Return to Example NLT [517] and compute S0
@2
40
0
03
51
A=2
40
0
23
5to quickly see again that Sis not
a linear transformation, while in Example LTPM [518] compute S
0 + 0x+ 0x2+ 0x3
=0 0
0 0
as an
example of Theorem LTTZZ [519] at work.
Subsection LTC
Linear Transformation Cartoons
Throughout this chapter, and Chapter R [603], we will include drawings of linear transformations. We will
call them \cartoons," not because they are humorous, but because they will only expose a portion of the
truth. A Bugs Bunny cartoon might give us some insights on human nature, but the rules of physics and
biology are routinely (and grossly) violated. So it will be with our linear transformation cartoons .
Here is our rst, followed by a guide to help you understand how these are meant to describe fundamental
truths about linear transformations, while simultaneously violating other truths.
Version 2.30
522 Section LT Linear Transformations
U VTu v
wv
0U 0V
xy
t
Diagram GLT. General Linear Transformation
Here we picture a linear transformation T:U!V, where this information will be consistently displayed
along the bottom edge. The ovals are meant to represent the vector spaces, in this case U, the domain,
on the left and V, the codomain, on the right. Of course, vector spaces are typically innite sets, so you'll
have to imagine that characteristic of these sets. A small dot inside of an oval will represent a vector
within that vector space, sometimes with a name, sometimes not (in this case every vector has a name).
The sizes of the ovals are meant to be proportional to the dimensions of the vector spaces. However, when
we make no assumptions about the dimensions, we will draw the ovals as the same size, as we have done
here (which is not meant to suggest that the dimensions have to be equal).
To convey that the linear transformation associates a certain input with a certain output, we will draw
an arrow from the input to the output. So, for example, in this cartoon we suggest that T(x) =y.
Nothing in the denition of a linear transformation prevents two dierent inputs being sent to the same
output and we see this in T(u) =v=T(w). Similarly, an output may not have any input being sent
its way, as illustrated by no arrow pointing at t. In this cartoon, we have captured the essence of our one
general theorem about linear transformations, Theorem LTTZZ [519], T(0U) =0V. On occasion we might
include this basic fact when it is relevant, at other times maybe not. Note that the denition of a linear
transformation requires that it be a function, so every element of the domain should be associated with
some element of the codomain. This will be re
ected by never having an element of the domain without
an arrow originating there.
These cartoons are of course no substitute for careful denitions and proofs, but they can be a handy
way to think about the various properties we will be studying.
Subsection MLT
Matrices and Linear Transformations
If you give me a matrix, then I can quickly build you a linear transformation. Always. First a motivating
example and then the theorem.
Example LTM
Linear transformation from a matrix
Let
A=2
43 1 8 1
2 0 5 2
1 1 3 73
5
Version 2.30
Subsection LT.MLT Matrices and Linear Transformations 523
and dene a function P:C4!C3by
P(x) =Ax
So we are using an old friend, the matrix-vector product (Denition MVP [223]) as a way to convert a
vector with 4 components into a vector with 3 components. Applying Denition MVP [223] allows us to
write the dening formula for Pin a slightly dierent form,
P(x) =Ax=2
43 1 8 1
2 0 5 2
1 1 3 73
52
664x1
x2
x3
x43
775=x12
43
2
13
5+x22
4 1
0
13
5+x32
48
5
33
5+x42
41
2
73
5
So we recognize the action of the function Pas using the components of the vector ( x1; x2; x3; x4) as
scalars to form the output of Pas a linear combination of the four columns of the matrix A, which are
all members of C3, so the result is a vector in C3. We can rearrange this expression further, using our
denitions of operations in C3(Section VO [97]).
P(x) =Ax Denition of P
=x12
43
2
13
5+x22
4 1
0
13
5+x32
48
5
33
5+x42
41
2
73
5 Denition MVP [223]
=2
43x1
2x1
x13
5+2
4 x2
0
x23
5+2
48x3
5x3
3x33
5+2
4x4
2x4
7x43
5 Denition CVSM [99]
=2
43x1 x2+ 8x3+x4
2x1+ 5x3 2x4
x1+x2+ 3x3 7x43
5 Denition CVA [98]
You might recognize this nal expression as being similar in style to some previous examples (Example ALT
[516]) and some linear transformations dened in the archetypes (Archetype M [833] through Archetype
R [848]). But the expression that says the output of this linear transformation is a linear combination of
the columns of Ais probably the most powerful way of thinking about examples of this type.
Almost forgot | we should verify that Pis indeed a linear transformation. This is easy with two
matrix properties from Section MM [223].
P(x+y) =A(x+y) Denition of P
=Ax+Ay Theorem MMDAA [230]
=P(x) +P(y) Denition of P
and
P(x) =A(x) Denition of P
=(Ax) Theorem MMSMM [230]
=P(x) Denition of P
So by Denition LT [515], Pis a linear transformation.
So the multiplication of a vector by a matrix \transforms" the input vector into an output vector,
possibly of a dierent size, by performing a linear combination. And this transformation happens in a
\linear" fashion. This \functional" view of the matrix-vector product is the most important shift you can
make right now in how you think about linear algebra. Here's the theorem, whose proof is very nearly an
exact copy of the verication in the last example.
Version 2.30
524 Section LT Linear Transformations
Theorem MBLT
Matrices Build Linear Transformations
Suppose that Ais anmnmatrix. Dene a function T:Cn!CmbyT(x) =Ax. ThenTis a linear
transformation.
Proof
T(x+y) =A(x+y) Denition of T
=Ax+Ay Theorem MMDAA [230]
=T(x) +T(y) Denition of T
and
T(x) =A(x) Denition of T
=(Ax) Theorem MMSMM [230]
=T(x) Denition of T
So by Denition LT [515], Tis a linear transformation.
So Theorem MBLT [522] gives us a rapid way to construct linear transformations. Grab an mn
matrixA, deneT(x) =Axand Theorem MBLT [522] tells us that Tis a linear transformation from Cn
toCm, without any further checking.
We can turn Theorem MBLT [522] around. You give me a linear transformation and I will give you a
matrix.
Example MFLT
Matrix from a linear transformation
Dene the function R:C3!C4by
R0
@2
4x1
x2
x33
51
A=2
6642x1 3x2+ 4x3
x1+x2+x3
x1+ 5x2 3x3
x2 4x33
775
You could verify that Ris a linear transformation by applying the denition, but we will instead massage the
expression dening a typical output until we recognize the form of a known class of linear transformations.
R0
@2
4x1
x2
x33
51
A=2
6642x1 3x2+ 4x3
x1+x2+x3
x1+ 5x2 3x3
x2 4x33
775
=2
6642x1
x1
x1
03
775+2
664 3x2
x2
5x2
x23
775+2
6644x3
x3
3x3
4x33
775Denition CVA [98]
=x12
6642
1
1
03
775+x22
664 3
1
5
13
775+x32
6644
1
3
43
775Denition CVSM [99]
=2
6642 3 4
1 1 1
1 5 3
0 1 43
7752
4x1
x2
x33
5 Denition MVP [223]
Version 2.30
Subsection LT.MLT Matrices and Linear Transformations 525
So if we dene the matrix
B=2
6642 3 4
1 1 1
1 5 3
0 1 43
775
thenR(x) =Bx. By Theorem MBLT [522], we can easily recognize Ras a linear transformation since it
has the form described in the hypothesis of the theorem.
Example MFLT [522] was not accident. Consider any one of the archetypes where both the domain
and codomain are sets of column vectors (Archetype M [833] through Archetype R [848]) and you should
be able to mimic the previous example. Here's the theorem, which is notable since it is our rst occasion
to use the full power of the dening properties of a linear transformation when our hypothesis includes a
linear transformation.
Theorem MLTCV
Matrix of a Linear Transformation, Column Vectors
Suppose that T:Cn!Cmis a linear transformation. Then there is an mnmatrixAsuch that
T(x) =Ax.
Proof The conclusion says a certain matrix exists. What better way to prove something exists than to
actually build it? So our proof will be constructive (Technique C [768]), and the procedure that we will
use abstractly in the proof can be used concretely in specic examples.
Lete1;e2;e3; :::; enbe the columns of the identity matrix of size n,In(Denition SUV [197]).
Evaluate the linear transformation Twith each of these standard unit vectors as an input, and record the
result. In other words, dene nvectors in Cm,Ai, 1inby
Ai=T(ei)
Then package up these vectors as the columns of a matrix
A= [A1jA2jA3j:::jAn]
DoesAhave the desired properties? First, Ais clearly an mnmatrix. Then
T(x) =T(Inx) Theorem MMIM [229]
=T([e1je2je3j:::jen]x) Denition SUV [197]
=T([x]1e1+ [x]2e2+ [x]3e3++ [x]nen) Denition MVP [223]
=T([x]1e1) +T([x]2e2) +T([x]3e3) ++T([x]nen) Denition LT [515]
= [x]1T(e1) + [x]2T(e2) + [x]3T(e3) ++ [x]nT(en) Denition LT [515]
= [x]1A1+ [x]2A2+ [x]3A3++ [x]nAn Denition of Ai
=Ax Denition MVP [223]
as desired.
So if we were to restrict our study of linear transformations to those where the domain and codomain are
both vector spaces of column vectors (Denition VSCV [97]), every matrix leads to a linear transformation
of this type (Theorem MBLT [522]), while every such linear transformation leads to a matrix (Theorem
MLTCV [523]). So matrices and linear transformations are fundamentally the same. We call the matrix
Aof Theorem MLTCV [523] the matrix representation ofT.
We have dened linear transformations for more general vector spaces than just Cm, can we extend
this correspondence between linear transformations and matrices to more general linear transformations
(more general domains and codomains)? Yes, and this is the main theme of Chapter R [603]. Stay tuned.
For now, let's illustrate Theorem MLTCV [523] with an example.
Version 2.30
526 Section LT Linear Transformations
Example MOLT
Matrix of a linear transformation
SupposeS:C3!C4is dened by
S0
@2
4x1
x2
x33
51
A=2
6643x1 2x2+ 5x3
x1+x2+x3
9x1 2x2+ 5x3
4x23
775
Then
C1=S(e1) =S0
@2
41
0
03
51
A=2
6643
1
9
03
775
C2=S(e2) =S0
@2
40
1
03
51
A=2
664 2
1
2
43
775
C3=S(e3) =S0
@2
40
0
13
51
A=2
6645
1
5
03
775
so dene
C= [C1jC2jC3] =2
6643 2 5
1 1 1
9 2 5
0 4 03
775
and Theorem MLTCV [523] guarantees that S(x) =Cx.
As an illuminating exercise, let z=2
42
3
33
5and compute S(z) two dierent ways. First, return to the
denition of Sand evaluate S(z) directly. Then do the matrix-vector product Cz. In both cases you
should obtain the vector S(z) =2
66427
2
39
123
775.
Subsection LTLC
Linear Transformations and Linear Combinations
It is the interaction between linear transformations and linear combinations that lies at the heart of many
of the important theorems of linear algebra. The next theorem distills the essence of this. The proof is
not deep, the result is hardly startling, but it will be referenced frequently. We have already passed by one
occasion to employ it, in the proof of Theorem MLTCV [523]. Paraphrasing, this theorem says that we
can \push" linear transformations \down into" linear combinations, or \pull" linear transformations \up
out" of linear combinations. We'll have opportunities to both push and pull.
Version 2.30
Subsection LT.LTLC Linear Transformations and Linear Combinations 527
Theorem LTLC
Linear Transformations and Linear Combinations
Suppose that T:U!Vis a linear transformation, u1;u2;u3; :::; utare vectors from Uanda1; a2; a3; :::; at
are scalars from C. Then
T(a1u1+a2u2+a3u3++atut) =a1T(u1) +a2T(u2) +a3T(u3) ++atT(ut)
Proof
T(a1u1+a2u2+a3u3++atut)
=T(a1u1) +T(a2u2) +T(a3u3) ++T(atut) Denition LT [515]
=a1T(u1) +a2T(u2) +a3T(u3) ++atT(ut) Denition LT [515]
Some authors, especially in more advanced texts, take the conclusion of Theorem LTLC [525] as the
dening condition of a linear transformation. This has the appeal of being a single condition, rather than
the two-part condition of Denition LT [515]. (See Exercise LT.T20 [536]).
Our next theorem says, informally, that it is enough to know how a linear transformation behaves for
inputs from any basis of the domain, and allthe other outputs are described by a linear combination of
these few values. Again, the statement of the theorem, and its proof, are not remarkable, but the insight
that goes along with it is very fundamental.
Theorem LTDB
Linear Transformation Dened on a Basis
SupposeB=fu1;u2;u3; :::; ungis a basis for the vector space Uandv1;v2;v3; :::; vnis a list of vectors
from the vector space V(which are not necessarily distinct). Then there is a unique linear transformation,
T:U!V, such that T(ui) =vi, 1in.
Proof To prove the existence of T, we construct a function and show that it is a linear transformation
(Technique C [768]). Suppose w2Uis an arbitrary element of the domain. Then by Theorem VRRB
[360] there are unique scalars a1; a2; a3; :::; ansuch that
w=a1u1+a2u2+a3u3++anun
Then dene
T(w) =a1v1+a2v2+a3v3++anvn
It should be clear that Tbehaves as required for ninputs from B. Since the scalars provided by Theorem
VRRB [360] are unique, there is no ambiguity in this denition, and Tqualies as a function with domain
Uand codomain V(i.e.Tis well-dened). But is Ta linear transformation as well?
Letx2Ube a second element of the domain, and suppose the scalars provided by Theorem VRRB
[360] (relative to B) areb1; b2; b3; :::; bn. Then
T(w+x) =T(a1u1+a2u2++anun+b1u1+b2u2++bnun)
=T((a1+b1)u1+ (a2+b2)u2++ (an+bn)un) Denition VS [317]
= (a1+b1)v1+ (a2+b2)v2++ (an+bn)vn Denition of T
=a1v1+a2v2++anvn+b1v1+b2v2++bnvn Denition VS [317]
=T(w) +T(x)
Version 2.30
528 Section LT Linear Transformations
Let2Cbe any scalar. Then
T(w) =T((a1u1+a2u2+a3u3++anun))
=T(a1u1+a2u2+a3u3++anun) Denition VS [317]
=a1v1+a2v2+a3v3++anvn Denition of T
=(a1v1+a2v2+a3v3++anvn) Denition VS [317]
=T(w)
So by Denition LT [515], Tis a linear transformation.
IsTunique (among all linear transformations that take the uito the vi)? Applying Technique U [771],
we posit the existence of a second linear transformation, S:U!Vsuch thatS(ui) =vi, 1in.
Again, let w2Urepresent an arbitrary element of Uand leta1; a2; a3; :::; anbe the scalars provided by
Theorem VRRB [360] (relative to B). We have,
T(w) =T(a1u1+a2u2+a3u3++anun) Theorem VRRB [360]
=a1T(u1) +a2T(u2) +a3T(u3) ++anT(un) Theorem LTLC [525]
=a1v1+a2v2+a3v3++anvn Denition of T
=a1S(u1) +a2S(u2) +a3S(u3) ++anS(un) Denition of S
=S(a1u1+a2u2+a3u3++anun) Theorem LTLC [525]
=S(w) Theorem VRRB [360]
So the output of TandSagree on every input, which means they are equal as functions, T=S. SoTis
unique.
You might recall facts from analytic geometry, such as \any two points determine a line" and \any
three non-collinear points determine a parabola." Theorem LTDB [525] has much of the same feel. By
specifying the noutputs for inputs from a basis, an entire linear transformation is determined. The analogy
is not perfect, but the style of these facts are not very dissimilar from Theorem LTDB [525].
Notice that the statement of Theorem LTDB [525] asserts the existence of a linear transformation with
certain properties, while the proof shows us exactly how to dene the desired linear transformation. The
next examples how to work with linear transformations that we nd this way.
Example LTDB1
Linear transformation dened on a basis
Consider the linear transformation T:C3!C2that is required to have the following three values,
T0
@2
41
0
03
51
A=2
1
T0
@2
40
1
03
51
A= 1
4
T0
@2
40
0
13
51
A=6
0
Because
B=8
<
:2
41
0
03
5;2
40
1
03
5;2
40
0
13
59
=
;
is a basis for C3(Theorem SUVB [371]), Theorem LTDB [525] says there is a unique linear transformation
Tthat behaves this way. How do we compute other values of T? Consider the input
w=2
42
3
13
5= (2)2
41
0
03
5+ ( 3)2
40
1
03
5+ (1)2
40
0
13
5
Version 2.30
Subsection LT.LTLC Linear Transformations and Linear Combinations 529
Then
T(w) = (2)2
1
+ ( 3) 1
4
+ (1)6
0
=13
10
Doing it again,
x=2
45
2
33
5= (5)2
41
0
03
5+ (2)2
40
1
03
5+ ( 3)2
40
0
13
5
so
T(x) = (5)2
1
+ (2) 1
4
+ ( 3)6
0
= 10
13
Any other value of Tcould be computed in a similar manner. So rather than being given a formula for
the outputs of T, the requirement thatTbehave in a certain way for the inputs chosen from a basis of
the domain, is as sucient as a formula for computing any value of the function. You might notice some
parallels between this example and Example MOLT [524] or Theorem MLTCV [523].
Example LTDB2
Linear transformation dened on a basis
Consider the linear transformation R:C3!C2with the three values,
R0
@2
41
2
13
51
A=5
1
R0
@2
4 1
5
13
51
A=0
4
R0
@2
43
1
43
51
A=2
3
You can check that
D=8
<
:2
41
2
13
5;2
4 1
5
13
5;2
43
1
43
59
=
;
is a basis for C3(make the vectors the columns of a square matrix and check that the matrix is nonsingular,
Theorem CNMB [376]). By Theorem LTDB [525] we know there is a unique linear transformation Rwith
the three specied outputs. However, we have to work just a bit harder to take an input vector and express
it as a linear combination of the vectors in D. For example, consider,
y=2
48
3
53
5
Then we must rst write yas a linear combination of the vectors in Dand solve for the unknown scalars,
to arrive at
y=2
48
3
53
5= (3)2
41
2
13
5+ ( 2)2
4 1
5
13
5+ (1)2
43
1
43
5
Then the proof of Theorem LTDB [525] gives us
R(y) = (3)5
1
+ ( 2)0
4
+ (1)2
3
=17
8
Any other value of Rcould be computed in a similar manner.
Here is a third example of a linear transformation dened by its action on a basis, only with more
abstract vector spaces involved.
Example LTDB3
Linear transformation dened on a basis
The setW=fp(x)2P3jp(1) = 0;p(3) = 0gP3is a subspace of the vector space of polynomials P3.
Version 2.30
530 Section LT Linear Transformations
This subspace has C=
3 4x+x2;12 13x+x3
as a basis (check this!). Suppose we consider the
linear transformation S:P3!M22with values
S
3 4x+x2
=1 3
2 0
S
12 13x+x3
=0 1
1 0
By Theorem LTDB [525] we know there is a unique linear transformation with these two values. To
illustrate a sample computation of S, considerq(x) = 9 6x 5x2+ 2x3. Verify that q(x) is an element of
W(does it have roots at x= 1 andx= 3?), then nd the scalars needed to write it as a linear combination
of the basis vectors in C. Because
q(x) = 9 6x 5x2+ 2x3= ( 5)(3 4x+x2) + (2)(12 13x+x3)
The proof of Theorem LTDB [525] gives us
S(q) = ( 5)1 3
2 0
+ (2)0 1
1 0
= 5 17
8 0
And all the other outputs of Scould be computed in the same manner. Every output of Swill have a zero
in the second row, second column. Can you see why this is so?
Informally, we can describe Theorem LTDB [525] by saying \it is enough to know what a linear
transformation does to a basis (of the domain)."
Subsection PI
Pre-Images
The denition of a function requires that for each input in the domain there is exactly one output in the
codomain. However, the correspondence does not have to behave the other way around. A member of the
codomain might have many inputs from the domain that create it, or it may have none at all. To formalize
our discussion of this aspect of linear transformations, we dene the pre-image.
Denition PI
Pre-Image
Suppose that T:U!Vis a linear transformation. For each v, dene the pre-image ofvto be the subset
ofUgiven by
T 1(v) =fu2UjT(u) =vg
4
In other words, T 1(v) is the set of all those vectors in the domain Uthat get \sent" to the vector v.
Example SPIAS
Sample pre-images, Archetype S
Archetype S [851] is the linear transformation dened by
T:C3!M22; T0
@2
4a
b
c3
51
A=a b 2a+ 2b+c
3a+b+c 2a 6b 2c
We could compute a pre-image for every element of the codomain M22. However, even in a free textbook,
we do not have the room to do that, so we will compute just two.
Choose
v=2 1
3 2
2M22
Version 2.30
Subsection LT.PI Pre-Images 531
for no particular reason. What is T 1(v)? Suppose u=2
4u1
u2
u33
52T 1(v). The condition that T(u) =v
becomes
2 1
3 2
=v=T(u) =T0
@2
4u1
u2
u33
51
A=u1 u2 2u1+ 2u2+u3
3u1+u2+u3 2u1 6u2 2u3
Using matrix equality (Denition ME [207]), we arrive at a system of four equations in the three unknowns
u1; u2; u3with an augmented matrix that we can row-reduce in the hunt for solutions,
2
6641 1 0 2
2 2 1 1
3 1 1 3
2 6 2 23
775RREF !2
664101
45
4
011
4 3
4
0 0 0 0
0 0 0 03
775
We recognize this system as having innitely many solutions described by the single free variable u3.
Eventually obtaining the vector form of the solutions (Theorem VFSLS [118]), we can describe the preimage
precisely as,
T 1(v) =
u2C3T(u) =v
=8
<
:2
4u1
u2
u33
5u1=5
4 1
4u3; u2= 3
4 1
4u39
=
;
=8
<
:2
45
4 1
4u3
3
4 1
4u3
u33
5u32C39
=
;
=8
<
:2
45
4
3
4
03
5+u32
4 1
4
1
4
13
5u32C39
=
;
=2
45
4
3
4
03
5+*8
<
:2
4 1
4
1
4
13
59
=
;+
This last line is merely a suggestive way of describing the set on the previous line. You might create three
or four vectors in the preimage, and evaluate Twith each. Was the result what you expected? For a hint
of things to come, you might try evaluating Twith just the lone vector in the spanning set above. What
was the result? Now take a look back at Theorem PSPHS [124]. Hmmmm.
OK, let's compute another preimage, but with a dierent outcome this time. Choose
v=1 1
2 4
2M22
What isT 1(v)? Suppose u=2
4u1
u2
u33
52T 1(v). ThatT(u) =vbecomes
1 1
2 4
=v=T(u) =T0
@2
4u1
u2
u33
51
A=u1 u2 2u1+ 2u2+u3
3u1+u2+u3 2u1 6u2 2u3
Version 2.30
532 Section LT Linear Transformations
Using matrix equality (Denition ME [207]), we arrive at a system of four equations in the three unknowns
u1; u2; u3with an augmented matrix that we can row-reduce in the hunt for solutions,
2
6641 1 0 1
2 2 1 1
3 1 1 2
2 6 2 43
775RREF !2
664101
40
011
40
0 0 0 1
0 0 0 03
775
By Theorem RCLS [58] we recognize this system as inconsistent. So no vector uis a member of T 1(v)
and so
T 1(v) =;
The preimage is just a set, it is almost never a subspace of U(you might think about just when T 1(v)
is a subspace, see Exercise ILT.T10 [553]). We will describe its properties going forward, and it will be
central to the main ideas of this chapter.
Subsection NLTFO
New Linear Transformations From Old
We can combine linear transformations in natural ways to create new linear transformations. So we will
dene these combinations and then prove that the results really are still linear transformations. First the
sum of two linear transformations.
Denition LTA
Linear Transformation Addition
Suppose that T:U!VandS:U!Vare two linear transformations with the same domain and
codomain. Then their sum is the function T+S:U!Vwhose outputs are dened by
(T+S) (u) =T(u) +S(u)
4
Notice that the rst plus sign in the denition is the operation being dened, while the second one is
the vector addition in V. (Vector addition in Uwill appear just now in the proof that T+Sis a linear
transformation.) Denition LTA [530] only provides a function. It would be nice to know that when the
constituents ( T,S) are linear transformations, then so too is T+S.
Theorem SLTLT
Sum of Linear Transformations is a Linear Transformation
Suppose that T:U!VandS:U!Vare two linear transformations with the same domain and
codomain. Then T+S:U!Vis a linear transformation.
Proof We simply check the dening properties of a linear transformation (Denition LT [515]). This is
a good place to consistently ask yourself which objects are being combined with which operations.
(T+S) (x+y) =T(x+y) +S(x+y) Denition LTA [530]
=T(x) +T(y) +S(x) +S(y) Denition LT [515]
=T(x) +S(x) +T(y) +S(y) Property C [317] in V
= (T+S) (x) + (T+S) (y) Denition LTA [530]
Version 2.30
Subsection LT.NLTFO New Linear Transformations From Old 533
and
(T+S) (x) =T(x) +S(x) Denition LTA [530]
=T(x) +S(x) Denition LT [515]
=(T(x) +S(x)) Property DVA [318] in V
=(T+S) (x) Denition LTA [530]
Example STLT
Sum of two linear transformations
Suppose that T:C2!C3andS:C2!C3are dened by
Tx1
x2
=2
4x1+ 2x2
3x1 4x2
5x1+ 2x23
5 Sx1
x2
=2
44x1 x2
x1+ 3x2
7x1+ 5x23
5
Then by Denition LTA [530], we have
(T+S)x1
x2
=Tx1
x2
+Sx1
x2
=2
4x1+ 2x2
3x1 4x2
5x1+ 2x23
5+2
44x1 x2
x1+ 3x2
7x1+ 5x23
5=2
45x1+x2
4x1 x2
2x1+ 7x23
5
and by Theorem SLTLT [530] we know T+Sis also a linear transformation from C2toC3.
Denition LTSM
Linear Transformation Scalar Multiplication
Suppose that T:U!Vis a linear transformation and 2C. Then the scalar multiple is the function
T:U!Vwhose outputs are dened by
(T) (u) =T(u)
4
Given that Tis a linear transformation, it would be nice to know that Tis also a linear transformation.
Theorem MLTLT
Multiple of a Linear Transformation is a Linear Transformation
Suppose that T:U!Vis a linear transformation and 2C. Then (T):U!Vis a linear transforma-
tion.
Proof We simply check the dening properties of a linear transformation (Denition LT [515]). This is
another good place to consistently ask yourself which objects are being combined with which operations.
(T) (x+y) =(T(x+y)) Denition LTSM [531]
=(T(x) +T(y)) Denition LT [515]
=T(x) +T(y) Property DVA [318] in V
= (T) (x) + (T) (y) Denition LTSM [531]
and
(T) (x) =T(x) Denition LTSM [531]
Version 2.30
534 Section LT Linear Transformations
=(T(x)) Denition LT [515]
= ()T(x) Property SMA [318] in V
= ()T(x) Commutativity in C
=(T(x)) Property SMA [318] in V
=((T) (x)) Denition LTSM [531]
Example SMLT
Scalar multiple of a linear transformation
Suppose that T:C4!C3is dened by
T0
BB@2
664x1
x2
x3
x43
7751
CCA=2
4x1+ 2x2 x3+ 2x4
x1+ 5x2 3x3+x4
2x1+ 3x2 4x3+ 2x43
5
For the sake of an example, choose = 2, so by Denition LTSM [531], we have
T0
BB@2
664x1
x2
x3
x43
7751
CCA= 2T0
BB@2
664x1
x2
x3
x43
7751
CCA= 22
4x1+ 2x2 x3+ 2x4
x1+ 5x2 3x3+x4
2x1+ 3x2 4x3+ 2x43
5=2
42x1+ 4x2 2x3+ 4x4
2x1+ 10x2 6x3+ 2x4
4x1+ 6x2 8x3+ 4x43
5
and by Theorem MLTLT [531] we know 2 Tis also a linear transformation from C4toC3.
Now, let's imagine we have two vector spaces, UandV, and we collect every possible linear transfor-
mation from UtoVinto one big set, and call it LT(U; V ). Denition LTA [530] and Denition LTSM
[531] tell us how we can \add" and \scalar multiply" two elements of LT(U; V ). Theorem SLTLT [530]
and Theorem MLTLT [531] tell us that if we do these operations, then the resulting functions are linear
transformations that are also in LT(U; V ). Hmmmm, sounds like a vector space to me! A set of objects,
an addition and a scalar multiplication. Why not?
Theorem VSLT
Vector Space of Linear Transformations
Suppose that UandVare vector spaces. Then the set of all linear transformations from UtoV,LT(U; V )
is a vector space when the operations are those given in Denition LTA [530] and Denition LTSM [531].
Proof Theorem SLTLT [530] and Theorem MLTLT [531] provide two of the ten properties in Denition
VS [317]. However, we still need to verify the remaining eight properties. By and large, the proofs are
straightforward and rely on concocting the obvious object, or by reducing the question to the same vector
space property in the vector space V.
The zero vector is of some interest, though. What linear transformation would we add to any other
linear transformation, so as to keep the second one unchanged? The answer is Z:U!Vdened by
Z(u) =0Vfor every u2U. Notice how we do not need to know any of the specics about UandVto
make this denition of Z.
Denition LTC
Linear Transformation Composition
Suppose that T:U!VandS:V!Ware linear transformations. Then the composition ofSandT
is the function ( ST):U!Wwhose outputs are dened by
(ST) (u) =S(T(u))
Version 2.30
Subsection LT.NLTFO New Linear Transformations From Old 535
4
Given that TandSare linear transformations, it would be nice to know that STis also a linear
transformation.
Theorem CLTLT
Composition of Linear Transformations is a Linear Transformation
Suppose that T:U!VandS:V!Ware linear transformations. Then ( ST):U!Wis a linear
transformation.
Proof We simply check the dening properties of a linear transformation (Denition LT [515]).
(ST) (x+y) =S(T(x+y)) Denition LTC [532]
=S(T(x) +T(y)) Denition LT [515] for T
=S(T(x)) +S(T(y)) Denition LT [515] for S
= (ST) (x) + (ST) (y) Denition LTC [532]
and
(ST) (x) =S(T(x)) Denition LTC [532]
=S(T(x)) Denition LT [515] for T
=S(T(x)) Denition LT [515] for S
=(ST) (x) Denition LTC [532]
Example CTLT
Composition of two linear transformations
Suppose that T:C2!C4andS:C4!C3are dened by
Tx1
x2
=2
664x1+ 2x2
3x1 4x2
5x1+ 2x2
6x1 3x23
775S0
BB@2
664x1
x2
x3
x43
7751
CCA=2
42x1 x2+x3 x4
5x1 3x2+ 8x3 2x4
4x1+ 3x2 4x3+ 5x43
5
Then by Denition LTC [532]
(ST)x1
x2
=S
Tx1
x2
=S0
BB@2
664x1+ 2x2
3x1 4x2
5x1+ 2x2
6x1 3x23
7751
CCA
=2
42(x1+ 2x2) (3x1 4x2) + (5x1+ 2x2) (6x1 3x2)
5(x1+ 2x2) 3(3x1 4x2) + 8(5x1+ 2x2) 2(6x1 3x2)
4(x1+ 2x2) + 3(3x1 4x2) 4(5x1+ 2x2) + 5(6x1 3x2)3
5
=2
4 2x1+ 13x2
24x1+ 44x2
15x1 43x23
5
and by Theorem CLTLT [533] STis a linear transformation from C2toC3.
Here is an interesting exercise that will presage an important result later. In Example STLT [531]
compute (via Theorem MLTCV [523]) the matrix of T,SandT+S. Do you see a relationship between
these three matrices?
Version 2.30
536 Section LT Linear Transformations
In Example SMLT [532] compute (via Theorem MLTCV [523]) the matrix of Tand 2T. Do you see a
relationship between these two matrices?
Here's the tough one. In Example CTLT [533] compute (via Theorem MLTCV [523]) the matrix of T,
SandST. Do you see a relationship between these three matrices???
Subsection READ
Reading Questions
1. Is the function below a linear transformation? Why or why not?
T:C3!C2; T0
@2
4x1
x2
x33
51
A=3x1 x2+x3
8x2 6
2. Determine the matrix representation of the linear transformation Sbelow.
S:C2!C3; Sx1
x2
=2
43x1+ 5x2
8x1 3x2
4x13
5
3. Theorem LTLC [525] has a fairly simple proof. Yet the result itself is very powerful. Comment on
why we might say this.
Version 2.30
Subsection LT.EXC Exercises 537
Subsection EXC
Exercises
C15 The archetypes below are all linear transformations whose domains and codomains are vector spaces
of column vectors (Denition VSCV [97]). For each one, compute the matrix representation described in
the proof of Theorem MLTCV [523].
Archetype M [833]
Archetype N [836]
Archetype O [839]
Archetype P [842]
Archetype Q [844]
Archetype R [848]
Contributed by Robert Beezer
C16 Find the matrix representation of T:C3!C4given byT0
@2
4x
y
z3
51
A=2
6643x+ 2y+z
x+y+z
x 3y
2x+ 3y+z3
775.
Contributed by Chris Black Solution [537]
C20 Letw=2
4 3
1
43
5. Referring to Example MOLT [524], compute S(w) two dierent ways. First use
the denition of S, then compute the matrix-vector product Cw(Denition MVP [223]).
Contributed by Robert Beezer Solution [537]
C25 Dene the linear transformation
T:C3!C2; T0
@2
4x1
x2
x33
51
A=2x1 x2+ 5x3
4x1+ 2x2 10x3
Verify that Tis a linear transformation.
Contributed by Robert Beezer Solution [537]
C26 Verify that the function below is a linear transformation.
T:P2!C2; T
a+bx+cx2
=2a b
b+c
Contributed by Robert Beezer Solution [537]
C30 Dene the linear transformation
T:C3!C2; T0
@2
4x1
x2
x33
51
A=2x1 x2+ 5x3
4x1+ 2x2 10x3
Compute the preimages, T 12
3
andT 14
8
.
Contributed by Robert Beezer Solution [538]
Version 2.30
538 Section LT Linear Transformations
C31 For the linear transformation Scompute the pre-images.
S:C3!C3; S0
@2
4a
b
c3
51
A=2
4a 2b c
3a b+ 2c
a+b+ 2c3
5
S 10
@2
4 2
5
33
51
A S 10
@2
4 5
5
73
51
A
Contributed by Robert Beezer Solution [538]
C40 IfT:C2!C2satisesT2
1
=3
4
andT1
1
= 1
2
, ndT4
3
.
Contributed by Chris Black Solution [539]
C41 IfT:C2!C3satisesT2
3
=2
42
2
13
5andT3
4
=2
4 1
0
23
5, nd the matrix representation of
T.
Contributed by Chris Black Solution [539]
C42 DeneT:M2;2!RbyTa b
c d
=a+b+c d. Find the pre-image T 1(3).
Contributed by Chris Black Solution [539]
C43 DeneT:P3!P2byT
a+bx+cx2+dx3
=b+ 2cx+ 3dx2. Find the pre-image of 0. Does this
linear transformation seem familiar?
Contributed by Chris Black Solution [539]
M10 Dene two linear transformations, T:C4!C3andS:C3!C2by
S0
@2
4x1
x2
x33
51
A=x1 2x2+ 3x3
5x1+ 4x2+ 2x3
T0
BB@2
664x1
x2
x3
x43
7751
CCA=2
4 x1+ 3x2+x3+ 9x4
2x1+x3+ 7x4
4x1+ 2x2+x3+ 2x43
5
Using the proof of Theorem MLTCV [523] compute the matrix representations of the three linear trans-
formations T,SandST. Discover and comment on the relationship between these three matrices.
Contributed by Robert Beezer Solution [540]
M60 SupposeUandVare vector spaces and dene a function Z:U!VbyT(u) =0Vfor every
u2U. Prove that Zis a (stupid) linear transformation. (See Exercise ILT.M60 [553], Exercise SLT.M60
[572], Exercise IVLT.M60 [594].)
Contributed by Robert Beezer
T20 Use the conclusion of Theorem LTLC [525] to motivate a new denition of a linear transformation.
Then prove that your new denition is equivalent to Denition LT [515]. (Technique D [765] and Technique
E [768] might be helpful if you are not sure what you are being asked to prove here.)
Contributed by Robert Beezer
Version 2.30
Subsection LT.SOL Solutions 539
Subsection SOL
Solutions
C16 Contributed by Chris Black Statement [535]
Answer:AT=2
6643 2 1
1 1 1
1 3 0
2 3 13
775.
C20 Contributed by Robert Beezer Statement [535]
In both cases the result will be S(w) =2
6649
2
9
43
775.
C25 Contributed by Robert Beezer Statement [535]
We can rewrite Tas follows:
T0
@2
4x1
x2
x33
51
A=2x1 x2+ 5x3
4x1+ 2x2 10x3
=x12
4
+x2 1
2
+x35
10
=2 1 5
4 2 102
4x1
x2
x33
5
and Theorem MBLT [522] tell us that any function of this form is a linear transformation.
C26 Contributed by Robert Beezer Statement [535]
Check the two conditions of Denition LT [515].
T(u+v) =T
a+bx+cx2
+
d+ex+fx2
=T
(a+d) + (b+e)x+ (c+f)x2
=2(a+d) (b+e)
(b+e) + (c+f)
=(2a b) + (2d e)
(b+c) + (e+f)
=2a b
b+c
+2d e
e+f
=T(u) +T(v)
and
T(u) =T
a+bx+cx2
=T
(a) + (b)x+ (c)x2
=2(a) (b)
(b) + (c)
=(2a b)
(b+c)
=2a b
b+c
=T(u)
SoTis indeed a linear transformation.
Version 2.30
540 Section LT Linear Transformations
C30 Contributed by Robert Beezer Statement [535]
For the rst pre-image, we want x2C3such thatT(x) =2
3
. This becomes,
2x1 x2+ 5x3
4x1+ 2x2 10x3
=2
3
Vector equality gives a system of two linear equations in three variables, represented by the augmented
matrix 2 1 5 2
4 2 10 3
RREF !1 1
25
20
0 0 0 1
so the system is inconsistent and the pre-image is the empty set. For the second pre-image the same
procedure leads to an augmented matrix with a dierent vector of constants
2 1 5 4
4 2 10 8
RREF !
1 1
25
22
0 0 0 0
This system is consistent and has innitely many solutions, as we can see from the presence of the two free
variables (x2andx3) both to zero. We apply Theorem VFSLS [118] to obtain
T 14
8
=8
<
:2
42
0
03
5+x22
41
2
1
03
5+x32
4 5
2
0
13
5x2; x32C9
=
;
C31 Contributed by Robert Beezer Statement [536]
We work from the denition of the pre-image, Denition PI [528]. Setting
S0
@2
4a
b
c3
51
A=2
4 2
5
33
5
we arrive at a system of three equations in three variables, with an augmented matrix that we row-reduce
in a search for solutions,2
41 2 1 2
3 1 2 5
1 1 2 33
5RREF !2
410 1 0
011 0
0 0 0 13
5
With a leading 1 in the last column, this system is inconsistent (Theorem RCLS [58]), and there are no
values ofa,bandcthat will create an element of the pre-image. So the preimage is the empty set.
We work from the denition of the pre-image, Denition PI [528]. Setting
S0
@2
4a
b
c3
51
A=2
4 5
5
73
5
we arrive at a system of three equations in three variables, with an augmented matrix that we row-reduce
in a search for solutions,2
41 2 1 5
3 1 2 5
1 1 2 73
5RREF !2
410 1 3
011 4
0 0 0 03
5
The solution set to this system, which is also the desired pre-image, can be expressed using the vector form
of the solutions (Theorem VFSLS [118])
S 10
@2
4 5
5
73
51
A=8
<
:2
43
4
03
5+c2
4 1
1
13
5c2C9
=
;=2
43
4
03
5+*8
<
:2
4 1
1
13
59
=
;+
Version 2.30
Subsection LT.SOL Solutions 541
Does the nal expression for this set remind you of Theorem KPI [547]?
C40 Contributed by Chris Black Statement [536]
Since4
3
=2
1
+ 21
1
, we have
T4
3
=T2
1
+ 21
1
=T2
1
+ 2T1
1
=3
4
+ 2 1
2
=1
8
:
C41 Contributed by Chris Black Statement [536]
First, we need to write the standard basis vectors e1ande2as linear combinations of2
3
and3
4
. Starting
withe1, we see that e1= 42
3
+ 33
4
, so we have
T(e1) =T
42
3
+ 33
4
= 4T2
3
+ 3T3
4
= 42
42
2
13
5+ 32
4 1
0
23
5=2
4 11
8
23
5:
Repeating the process for e2, we have e2= 32
3
23
4
, and we then see that
T(e2) =T
32
3
23
4
= 3T2
3
2T3
4
= 32
42
2
13
5 22
4 1
0
23
5=2
48
6
13
5:
Thus, the matrix representation of TisAT=2
4 11 8
8 6
2 13
5.
C42 Contributed by Chris Black Statement [536]
The preimage T 1(3) is the set of all matricesa b
c d
so thatTa b
c d
= 3. A matrixa b
c d
is in the
preimage if a+b+c d= 3, i.e.d=a+b+c 3. This is the set. (But the set is nota vector space.
Why not?)
T 1(3) =a b
c a +b+c 3a;b;c2C
C43 Contributed by Chris Black Statement [536]
The preimage T 1(0) is the set of all polynomials a+bx+cx2+dx3so thatT
a+bx+cx2+dx3
= 0.
Thus,b+ 2cx+ 3dx2= 0, where the 0 represents the zero polynomial. In order to satisfy this equation,
we must have b= 0,c= 0, andd= 0. Thus, T 1(0) is precisely the set of all constant polynomials {
polynomials of degree 0. Symbolically, this is T 1(0) =faja2Cg.
Does this seem familiar? What other operation sends constant functions to 0?
Version 2.30
542 Section LT Linear Transformations
M10 Contributed by Robert Beezer Statement [536]
1 2 3
5 4 22
4 1 3 1 9
2 0 1 7
4 2 1 23
5=7 9 2 1
11 19 11 77
Version 2.30
Section ILT Injective Linear Transformations 543
Section ILT
Injective Linear Transformations
Some linear transformations possess one, or both, of two key properties, which go by the names injective
and surjective. We will see that they are closely related to ideas like linear independence and spanning,
and subspaces like the null space and the column space. In this section we will dene an injective linear
transformation and analyze the resulting consequences. The next section will do the same for the surjective
property. In the nal section of this chapter we will see what happens when we have the two properties
simultaneously.
As usual, we lead with a denition.
Denition ILT
Injective Linear Transformation
SupposeT:U!Vis a linear transformation. Then Tisinjective if whenever T(x) =T(y), then x=y.
4
Given an arbitrary function, it is possible for two dierent inputs to yield the same output (think about
the function f(x) =x2and the inputs x= 3 andx= 3). For an injective function, this never happens. If
we have equal outputs ( T(x) =T(y)) then we must have achieved those equal outputs by employing equal
inputs ( x=y). Some authors prefer the term one-to-one where we use injective, and we will sometimes
refer to an injective linear transformation as an injection .
Subsection EILT
Examples of Injective Linear Transformations
It is perhaps most instructive to examine a linear transformation that is not injective rst.
Example NIAQ
Not injective, Archetype Q
Archetype Q [844] is the linear transformation
T:C5!C5; T0
BBBB@2
66664x1
x2
x3
x4
x53
777751
CCCCA=2
66664 2x1+ 3x2+ 3x3 6x4+ 3x5
16x1+ 9x2+ 12x3 28x4+ 28x5
19x1+ 7x2+ 14x3 32x4+ 37x5
21x1+ 9x2+ 15x3 35x4+ 39x5
9x1+ 5x2+ 7x3 16x4+ 16x53
77775
Notice that for
x=2
666641
3
1
2
43
77775y=2
666644
7
0
5
73
77775
we have
T0
BBBB@2
666641
3
1
2
43
777751
CCCCA=2
666644
55
72
77
313
77775T0
BBBB@2
666644
7
0
5
73
777751
CCCCA=2
666644
55
72
77
313
77775
Version 2.30
544 Section ILT Injective Linear Transformations
So we have two vectors from the domain, x6=y, yetT(x) =T(y), in violation of Denition ILT [541].
This is another example where you should not concern yourself with how xandywere selected, as this will
be explained shortly. However, do understand whythese two vectors provide enough evidence to conclude
thatTis not injective.
Here's a cartoon of a non-injective linear transformation. Notice that the central feature of this cartoon
is thatT(u) =v=T(w). Even though this happens again with some unnamed vectors, it only takes one
occurrence to destroy the possibility of injectivity. Note also that the two vectors displayed in the bottom
ofVhave no bearing, either way, on the injectivity of T.
U VTu
v
wv
Diagram NILT. Non-Injective Linear Transformation
To show that a linear transformation is not injective, it is enough to nd a single pair of inputs that get
sent to the identical output, as in Example NIAQ [541]. However, to show that a linear transformation is
injective we must establish that this coincidence of outputs never occurs. Here is an example that shows
how to establish this.
Example IAR
Injective, Archetype R
Archetype R [848] is the linear transformation
T:C5!C5; T0
BBBB@2
66664x1
x2
x3
x4
x53
777751
CCCCA=2
66664 65x1+ 128x2+ 10x3 262x4+ 40x5
36x1 73x2 x3+ 151x4 16x5
44x1+ 88x2+ 5x3 180x4+ 24x5
34x1 68x2 3x3+ 140x4 18x5
12x1 24x2 x3+ 49x4 5x53
77775
To establish that Ris injective we must begin with the assumption that T(x) =T(y) and somehow arrive
from this at the conclusion that x=y. Here we go,
T(x) =T(y)
T0
BBBB@2
66664x1
x2
x3
x4
x53
777751
CCCCA=T0
BBBB@2
66664y1
y2
y3
y4
y53
777751
CCCCA
Version 2.30
Subsection ILT.EILT Examples of Injective Linear Transformations 545
2
66664 65x1+ 128x2+ 10x3 262x4+ 40x5
36x1 73x2 x3+ 151x4 16x5
44x1+ 88x2+ 5x3 180x4+ 24x5
34x1 68x2 3x3+ 140x4 18x5
12x1 24x2 x3+ 49x4 5x53
77775=2
66664 65y1+ 128y2+ 10y3 262y4+ 40y5
36y1 73y2 y3+ 151y4 16y5
44y1+ 88y2+ 5y3 180y4+ 24y5
34y1 68y2 3y3+ 140y4 18y5
12y1 24y2 y3+ 49y4 5y53
77775
2
66664 65x1+ 128x2+ 10x3 262x4+ 40x5
36x1 73x2 x3+ 151x4 16x5
44x1+ 88x2+ 5x3 180x4+ 24x5
34x1 68x2 3x3+ 140x4 18x5
12x1 24x2 x3+ 49x4 5x53
77775 2
66664 65y1+ 128y2+ 10y3 262y4+ 40y5
36y1 73y2 y3+ 151y4 16y5
44y1+ 88y2+ 5y3 180y4+ 24y5
34y1 68y2 3y3+ 140y4 18y5
12y1 24y2 y3+ 49y4 5y53
77775=2
666640
0
0
0
03
77775
2
66664 65(x1 y1) + 128(x2 y2) + 10(x3 y3) 262(x4 y4) + 40(x5 y5)
36(x1 y1) 73(x2 y2) (x3 y3) + 151(x4 y4) 16(x5 y5)
44(x1 y1) + 88(x2 y2) + 5(x3 y3) 180(x4 y4) + 24(x5 y5)
34(x1 y1) 68(x2 y2) 3(x3 y3) + 140(x4 y4) 18(x5 y5)
12(x1 y1) 24(x2 y2) (x3 y3) + 49(x4 y4) 5(x5 y5)3
77775=2
666640
0
0
0
03
77775
2
66664 65 128 10 262 40
36 73 1 151 16
44 88 5 180 24
34 68 3 140 18
12 24 1 49 53
777752
66664x1 y1
x2 y2
x3 y3
x4 y4
x5 y53
77775=2
666640
0
0
0
03
77775
Now we recognize that we have a homogeneous system of 5 equations in 5 variables (the terms xi yiare
the variables), so we row-reduce the coecient matrix to
2
66666410 0 0 0
010 0 0
0 0 10 0
0 0 0 10
0 0 0 0 13
777775
So the only solution is the trivial solution
x1 y1= 0 x2 y2= 0 x3 y3= 0 x4 y4= 0 x5 y5= 0
and we conclude that indeed x=y. By Denition ILT [541], Tis injective.
Here's the cartoon for an injective linear transformation. It is meant to suggest that we never have two
inputs associated with a single output. Again, the two lonely vectors at the bottom of Vhave no bearing
either way on the injectivity of T.
Version 2.30
546 Section ILT Injective Linear Transformations
U VT
Diagram ILT. Injective Linear Transformation
Let's now examine an injective linear transformation between abstract vector spaces.
Example IAV
Injective, Archetype V
Archetype V [858] is dened by
T:P3!M22; T
a+bx+cx2+dx3
=a+b a 2c
d b d
To establish that the linear transformation is injective, begin by supposing that two polynomial inputs
yield the same output matrix,
T
a1+b1x+c1x2+d1x3
=T
a2+b2x+c2x2+d2x3
Then
O=0 0
0 0
=T
a1+b1x+c1x2+d1x3
T
a2+b2x+c2x2+d2x3
Hypothesis
=T
(a1+b1x+c1x2+d1x3) (a2+b2x+c2x2+d2x3)
Denition LT [515]
=T
(a1 a2) + (b1 b2)x+ (c1 c2)x2+ (d1 d2)x3
Operations in P3
=(a1 a2) + (b1 b2) (a1 a2) 2(c1 c2)
(d1 d2) (b1 b2) (d1 d2)
Denition of T
This single matrix equality translates to the homogeneous system of equations in the variables ai bi,
(a1 a2) + (b1 b2) = 0
(a1 a2) 2(c1 c2) = 0
(d1 d2) = 0
(b1 b2) (d1 d2) = 0
This system of equations can be rewritten as the matrix equation
2
6641 1 0 0
1 0 2 0
0 0 0 1
0 1 0 13
7752
664(a1 a2)
(b1 b2)
(c1 c2)
(d1 d2)3
775=2
6640
0
0
03
775
Version 2.30
Subsection ILT.KLT Kernel of a Linear Transformation 547
Since the coecient matrix is nonsingular (check this) the only solution is trivial, i.e.
a1 a2= 0 b1 b2= 0 c1 c2= 0 d1 d2= 0
so that
a1=a2 b1=b2 c1=c2 d1=d2
so the two inputs must be equal polynomials. By Denition ILT [541], Tis injective.
Subsection KLT
Kernel of a Linear Transformation
For a linear transformation T:U!V, the kernel is a subset of the domain U. Informally, it is the set
of all inputs that the transformation sends to the zero vector of the codomain. It will have some natural
connections with the null space of a matrix, so we will keep the same notation, and if you think about your
objects, then there should be little confusion. Here's the careful denition.
Denition KLT
Kernel of a Linear Transformation
SupposeT:U!Vis a linear transformation. Then the kernel ofTis the set
K(T) =fu2UjT(u) =0g
(This denition contains Notation KLT.) 4
Notice that the kernel of Tis just the preimage of 0,T 1(0) (Denition PI [528]). Here's an example.
Example NKAO
Nontrivial kernel, Archetype O
Archetype O [839] is the linear transformation
T:C3!C5; T0
@2
4x1
x2
x33
51
A=2
66664 x1+x2 3x3
x1+ 2x2 4x3
x1+x2+x3
2x1+ 3x2+x3
x1+ 2x33
77775
To determine the elements of C3inK(T), nd those vectors usuch thatT(u) =0, that is,
T(u) =0
2
66664 u1+u2 3u3
u1+ 2u2 4u3
u1+u2+u3
2u1+ 3u2+u3
u1+ 2u33
77775=2
666640
0
0
0
03
77775
Vector equality (Denition CVE [98]) leads us to a homogeneous system of 5 equations in the variables ui,
u1+u2 3u3= 0
u1+ 2u2 4u3= 0
Version 2.30
548 Section ILT Injective Linear Transformations
u1+u2+u3= 0
2u1+ 3u2+u3= 0
u1+ 2u3= 0
Row-reducing the coecient matrix gives
2
6666410 2
01 1
0 0 0
0 0 0
0 0 03
77775
The kernel of Tis the set of solutions to this homogeneous system of equations, which by Theorem BNS
[160] can be expressed as
K(T) =*8
<
:2
4 2
1
13
59
=
;+
We know that the span of a set of vectors is always a subspace (Theorem SSS [339]), so the kernel com-
puted in Example NKAO [545] is also a subspace. This is no accident, the kernel of a linear transformation
isalways a subspace.
Theorem KLTS
Kernel of a Linear Transformation is a Subspace
Suppose that T:U!Vis a linear transformation. Then the kernel of T,K(T), is a subspace of U.
Proof We can apply the three-part test of Theorem TSS [334]. First T(0U) =0Vby Theorem LTTZZ
[519], so 0U2K(T) and we know that the kernel is non-empty.
Suppose we assume that x;y2K(T). Isx+y2K(T)?
T(x+y) =T(x) +T(y) Denition LT [515]
=0+0 x ;y2K(T)
=0 Property Z [318]
This qualies x+yfor membership in K(T). So we have additive closure.
Suppose we assume that 2Candx2K(T). Isx2K(T)?
T(x) =T(x) Denition LT [515]
=0 x 2K(T)
=0 Theorem ZVSM [325]
This qualies xfor membership in K(T). So we have scalar closure and Theorem TSS [334] tells us that
K(T) is a subspace of U.
Let's compute another kernel, now that we know in advance that it will be a subspace.
Example TKAP
Trivial kernel, Archetype P
Archetype P [842] is the linear transformation
T:C3!C5; T0
@2
4x1
x2
x33
51
A=2
66664 x1+x2+x3
x1+ 2x2+ 2x3
x1+x2+ 3x3
2x1+ 3x2+x3
2x1+x2+ 3x33
77775
Version 2.30
Subsection ILT.KLT Kernel of a Linear Transformation 549
To determine the elements of C3inK(T), nd those vectors usuch thatT(u) =0, that is,
T(u) =0
2
66664 u1+u2+u3
u1+ 2u2+ 2u3
u1+u2+ 3u3
2u1+ 3u2+u3
2u1+u2+ 3u33
77775=2
666640
0
0
0
03
77775
Vector equality (Denition CVE [98]) leads us to a homogeneous system of 5 equations in the variables ui,
u1+u2+u3= 0
u1+ 2u2+ 2u3= 0
u1+u2+ 3u3= 0
2u1+ 3u2+u3= 0
2u1+u2+ 3u3= 0
Row-reducing the coecient matrix gives
2
66666410 0
010
0 0 1
0 0 0
0 0 03
777775
The kernel of Tis the set of solutions to this homogeneous system of equations, which is simply the trivial
solution u=0, so
K(T) =f0g=hfgi
Our next theorem says that if a preimage is a non-empty set then we can construct it by picking any
one element and adding on elements of the kernel.
Theorem KPI
Kernel and Pre-Image
SupposeT:U!Vis a linear transformation and v2V. If the preimage T 1(v) is non-empty, and
u2T 1(v) then
T 1(v) =fu+zjz2K(T)g=u+K(T)
Proof LetM=fu+zjz2K(T)g. First, we show that MT 1(v). Suppose that w2M, sowhas
the form w=u+z, where z2K(T). Then
T(w) =T(u+z)
=T(u) +T(z) Denition LT [515]
=v+0 u 2T 1(v);z2K(T)
=v Property Z [318]
which qualies wfor membership in the preimage of v,w2T 1(v).
For the opposite inclusion, suppose x2T 1(v). Then,
T(x u) =T(x) T(u) Denition LT [515]
Version 2.30
550 Section ILT Injective Linear Transformations
=v v x ;u2T 1(v)
=0
This qualies x ufor membership in the kernel of T,K(T). So there is a vector z2K(T) such that
x u=z. Rearranging this equation gives x=u+zand so x2M. SoT 1(v)Mand we see that
M=T 1(v), as desired.
This theorem, and its proof, should remind you very much of Theorem PSPHS [124]. Additionally, you
might go back and review Example SPIAS [528]. Can you tell now which is the only preimage to be a
subspace?
The next theorem is one we will cite frequently, as it characterizes injections by the size of the kernel.
Theorem KILT
Kernel of an Injective Linear Transformation
Suppose that T:U!Vis a linear transformation. Then Tis injective if and only if the kernel of Tis
trivial,K(T) =f0g.
Proof ()) We assume Tis injective and we need to establish that two sets are equal (Denition SE
[762]). Since the kernel is a subspace (Theorem KLTS [546]), f0gK (T). To establish the opposite
inclusion, suppose x2K(T).
T(x) =0 Denition KLT [545]
=T(0) Theorem LTTZZ [519]
We can apply Denition ILT [541] to conclude that x=0. ThereforeK(T)f0gand by Denition SE
[762],K(T) =f0g.
(() To establish that Tis injective, appeal to Denition ILT [541] and begin with the assumption that
T(x) =T(y). Then
T(x y) =T(x) T(y) Denition LT [515]
=0 Hypothesis
Sox y2K(T) by Denition KLT [545] and with the hypothesis that the kernel is trivial we conclude
thatx y=0. Then
y=y+0=y+ (x y) =x
thus establishing that Tis injective by Denition ILT [541].
Example NIAQR
Not injective, Archetype Q, revisited
We are now in a position to revisit our rst example in this section, Example NIAQ [541]. In that example,
we showed that Archetype Q [844] is not injective by constructing two vectors, which when used to evaluate
the linear transformation provided the same output, thus violating Denition ILT [541]. Just where did
those two vectors come from?
The key is the vector
z=2
666643
4
1
3
33
77775
which you can check is an element of K(T) for Archetype Q [844]. Choose a vector xat random, and then
compute y=x+z(verify this computation back in Example NIAQ [541]). Then
T(y) =T(x+z)
Version 2.30
Subsection ILT.ILTLI Injective Linear Transformations and Linear Independence 551
=T(x) +T(z) Denition LT [515]
=T(x) +0 z 2K(T)
=T(x) Property Z [318]
Whenever the kernel of a linear transformation is non-trivial, we can employ this device and conclude that
the linear transformation is not injective. This is another way of viewing Theorem KILT [548]. For an
injective linear transformation, the kernel is trivial and our only choice for zis the zero vector, which will
not help us create two dierent inputs forTthat yield identical outputs. For every one of the archetypes
that is not injective, there is an example presented of exactly this form.
Example NIAO
Not injective, Archetype O
In Example NKAO [545] the kernel of Archetype O [839] was determined to be
*8
<
:2
4 2
1
13
59
=
;+
a subspace of C3with dimension 1. Since the kernel is not trivial, Theorem KILT [548] tells us that Tis
not injective.
Example IAP
Injective, Archetype P
In Example TKAP [546] it was shown that the linear transformation in Archetype P [842] has a trivial
kernel. So by Theorem KILT [548], Tis injective.
Subsection ILTLI
Injective Linear Transformations and Linear Independence
There is a connection between injective linear transformations and linearly independent sets that we will
make precise in the next two theorems. However, more informally, we can get a feel for this connection
when we think about how each property is dened. A set of vectors is linearly independent if the only
relation of linear dependence is the trivial one. A linear transformation is injective if the only way two
input vectors can produce the same output is in the trivial way, when both input vectors are equal.
Theorem ILTLI
Injective Linear Transformations and Linear Independence
Suppose that T:U!Vis an injective linear transformation and S=fu1;u2;u3; :::; utgis a linearly
independent subset of U. ThenR=fT(u1); T(u2); T(u3); :::; T (ut)gis a linearly independent subset
ofV.
Proof Begin with a relation of linear dependence on R(Denition RLD [351], Denition LI [351]),
a1T(u1) +a2T(u2) +a3T(u3) +:::+atT(ut) =0
T(a1u1+a2u2+a3u3++atut) =0 Theorem LTLC [525]
a1u1+a2u2+a3u3++atut2K(T) Denition KLT [545]
a1u1+a2u2+a3u3++atut2f0g Theorem KILT [548]
a1u1+a2u2+a3u3++atut=0 Denition SET [761]
Version 2.30
552 Section ILT Injective Linear Transformations
Since this is a relation of linear dependence on the linearly independent set S, we can conclude that
a1= 0 a2= 0 a3= 0 ::: a t= 0
and this establishes that Ris a linearly independent set.
Theorem ILTB
Injective Linear Transformations and Bases
Suppose that T:U!Vis a linear transformation and B=fu1;u2;u3; :::; umgis a basis of U. ThenT
is injective if and only if C=fT(u1); T(u2); T(u3); :::; T (um)gis a linearly independent subset of V.
Proof ()) AssumeTis injective. Since Bis a basis, we know Bis linearly independent (Denition B
[371]). Then Theorem ILTLI [549] says that Cis a linearly independent subset of V.
(() Assume that Cis linearly independent. To establish that Tis injective, we will show that the
kernel ofTis trivial (Theorem KILT [548]). Suppose that u2K(T). As an element of U, we can write u
as a linear combination of the basis vectors in B(uniquely). So there are are scalars, a1; a2; a3; :::; am,
such that
u=a1u1+a2u2+a3u3++amum
Then,
0=T(u) Denition KLT [545]
=T(a1u1+a2u2+a3u3++amum) Denition TSVS [356]
=a1T(u1) +a2T(u2) +a3T(u3) ++amT(um) Theorem LTLC [525]
This is a relation of linear dependence (Denition RLD [351]) on the linearly independent set C, so the
scalars are all zero: a1=a2=a3==am= 0. Then
u=a1u1+a2u2+a3u3++amum
= 0u1+ 0u2+ 0u3++ 0um Theorem ZSSM [324]
=0+0+0++0 Theorem ZSSM [324]
=0 Property Z [318]
Since uwas chosen as an arbitrary vector from K(T), we haveK(T) =f0gand Theorem KILT [548] tells
us thatTis injective.
Subsection ILTD
Injective Linear Transformations and Dimension
Theorem ILTD
Injective Linear Transformations and Dimension
Suppose that T:U!Vis an injective linear transformation. Then dim ( U)dim (V).
Proof Suppose to the contrary that m= dim (U)>dim (V) =t. LetBbe a basis of U, which will then
containmvectors. Apply Tto each element of Bto form a set Cthat is a subset of V. By Theorem ILTB
[550],Cis linearly independent and therefore must contain mdistinct vectors. So we have found a set of
mlinearly independent vectors in V, a vector space of dimension t, withm>t . However, this contradicts
Theorem G [407], so our assumption is false and dim ( U)dim (V).
Example NIDAU
Not injective by dimension, Archetype U
Version 2.30
Subsection ILT.CILT Composition of Injective Linear Transformations 553
The linear transformation in Archetype U [856] is
T:M23!C4; Ta b c
d e f
=2
664a+ 2b+ 12c 3d+e+ 6f
2a b c+d 11f
a+b+ 7c+ 2d+e 3f
a+ 2b+ 12c+ 5e 5f3
775
Since dim (M23) = 6>4 = dim
C4
,Tcannot be injective for then Twould violate Theorem ILTD [550].
Notice that the previous example made no use of the actual formula dening the function. Merely
a comparison of the dimensions of the domain and codomain are enough to conclude that the linear
transformation is not injective. Archetype M [833] and Archetype N [836] are two more examples of linear
transformations that have \big" domains and \small" codomains, resulting in \collisions" of outputs and
thus are non-injective linear transformations.
Subsection CILT
Composition of Injective Linear Transformations
In Subsection LT.NLTFO [530] we saw how to combine linear transformations to build new linear trans-
formations, specically, how to build the composition of two linear transformations (Denition LTC [532]).
It will be useful later to know that the composition of injective linear transformations is again injective,
so we prove that here.
Theorem CILTI
Composition of Injective Linear Transformations is Injective
Suppose that T:U!VandS:V!Ware injective linear transformations. Then ( ST):U!Wis an
injective linear transformation.
Proof That the composition is a linear transformation was established in Theorem CLTLT [533], so we
need only establish that the composition is injective. Applying Denition ILT [541], choose x,yfromU.
Then if (ST) (x) = (ST) (y),
) S(T(x)) =S(T(y)) Denition LTC [532]
) T(x) =T(y) Denition ILT [541] for S
) x=y Denition ILT [541] for T
Subsection READ
Reading Questions
1. Suppose T:C8!C5is a linear transformation. Why can't Tbe injective?
2. Describe the kernel of an injective linear transformation.
3. Theorem KPI [547] should remind you of Theorem PSPHS [124]. Why do we say this?
Version 2.30
554 Section ILT Injective Linear Transformations
Subsection EXC
Exercises
C10 Each archetype below is a linear transformation. Compute the kernel for each.
Archetype M [833]
Archetype N [836]
Archetype O [839]
Archetype P [842]
Archetype Q [844]
Archetype R [848]
Archetype S [851]
Archetype T [854]
Archetype U [856]
Archetype V [858]
Archetype W [860]
Archetype X [862]
Contributed by Robert Beezer
C20 The linear transformation T:C4!C3is not injective. Find two inputs x;y2C4that yield the
same output (that is T(x) =T(y)).
T0
BB@2
664x1
x2
x3
x43
7751
CCA=2
42x1+x2+x3
x1+ 3x2+x3 x4
3x1+x2+ 2x3 2x43
5
Contributed by Robert Beezer Solution [555]
C25 Dene the linear transformation
T:C3!C2; T0
@2
4x1
x2
x33
51
A=2x1 x2+ 5x3
4x1+ 2x2 10x3
Find a basis for the kernel of T,K(T). IsTinjective?
Contributed by Robert Beezer Solution [555]
C26 LetA=2
6641 2 3 1 0
2 1 1 0 1
1 2 1 2 1
1 3 2 1 23
775and letT:C5!C4be given by T(x) =Ax. IsTinjective? (Hint:
No calculation is required.)
Contributed by Chris Black Solution [556]
C27 LetT:C3!C3be given by T0
@2
4x
y
z3
51
A=2
42x+y+z
x y+ 2z
x+ 2y z3
5. FindK(T). IsTinjective?
Contributed by Chris Black Solution [556]
Version 2.30
Subsection ILT.EXC Exercises 555
C28 LetA=2
6641 2 3 1
2 1 1 0
1 2 1 2
1 3 2 13
775and letT:C4!C4be given by T(x) =Ax. FindK(T). IsT
injective?
Contributed by Chris Black Solution [556]
C29 LetA=2
6641 2 1 1
2 1 1 0
1 2 1 2
1 2 1 13
775and letT:C4!C4be given by T(x) =Ax. FindK(T). IsTinjective?
Contributed by Chris Black Solution [556]
C30 LetT:M2;2!P2be given by Ta b
c d
= (a+b) + (a+c)x+ (a+d)x2. IsTinjective? Find
K(T).
Contributed by Chris Black Solution [556]
C31 Given that the linear transformation T:C3!C3,T0
@2
4x
y
z3
51
A=2
42x+y
2y+z
x+ 2z3
5is injective, show directly
thatfT(e1); T(e2); T(e3)gis a linearly independent set.
Contributed by Chris Black Solution [557]
C32 Given that the linear transformation T:C2!C3,Tx
y
=2
4x+y
2x+y
x+ 2y3
5is injective, show directly
thatfT(e1); T(e2)gis a linearly independent set.
Contributed by Chris Black Solution [557]
C33 Given that the linear transformation T:C3!C5,T0
@2
4x
y
z3
51
A=2
666641 3 2
0 1 1
1 2 1
1 0 1
3 1 23
777752
4x
y
z3
5is injective, show
directly thatfT(e1); T(e2); T(e3)gis a linearly independent set.
Contributed by Chris Black Solution [557]
C40 Show that the linear transformation Ris not injective by nding two dierent elements of the
domain, xandy, such that R(x) =R(y). (S22is the vector space of symmetric 2 2 matrices.)
R:S22!P1Ra b
b c
= (2a b+c) + (a+b+ 2c)x
Contributed by Robert Beezer Solution [557]
M60 SupposeUandVare vector spaces. Dene the function Z:U!VbyT(u) =0Vfor every
u2U. Then by Exercise LT.M60 [536], Zis a linear transformation. Formulate a condition on Uthat
is equivalent to Zbeing an injective linear transformation. In other words, ll in the blank to complete
the following statement (and then give a proof): Zis injective if and only if Uis . (See
Exercise SLT.M60 [572], Exercise IVLT.M60 [594].)
Contributed by Robert Beezer
T10 SupposeT:U!Vis a linear transformation. For which vectors v2VisT 1(v) a subspace of
Version 2.30
556 Section ILT Injective Linear Transformations
U?
Contributed by Robert Beezer
T15 Suppose that that T:U!VandS:V!Ware linear transformations. Prove the following
relationship between null spaces.
K(T)K(ST)
Contributed by Robert Beezer Solution [558]
T20 Suppose that Ais anmnmatrix. Dene the linear transformation Tby
T:Cn!Cm; T (x) =Ax
Prove that the kernel of Tequals the null space of A,K(T) =N(A).
Contributed by Andy Zimmer Solution [558]
Version 2.30
Subsection ILT.SOL Solutions 557
Subsection SOL
Solutions
C20 Contributed by Robert Beezer Statement [552]
A linear transformation that is not injective will have a non-trivial kernel (Theorem KILT [548]), and this
is the key to nding the desired inputs. We need one non-trivial element of the kernel, so suppose that
z2C4is an element of the kernel,
2
40
0
03
5=0=T(z) =2
42z1+z2+z3
z1+ 3z2+z3 z4
3z1+z2+ 2z3 2z43
5
Vector equality Denition CVE [98] leads to the homogeneous system of three equations in four variables,
2z1+z2+z3= 0
z1+ 3z2+z3 z4= 0
3z1+z2+ 2z3 2z4= 0
The coecient matrix of this system row-reduces as
2
42 1 1 0
1 3 1 1
3 1 2 23
5RREF !2
410 0 1
010 1
0 0 1 33
5
From this we can nd a solution (we only need one), that is an element of K(T),
z=2
664 1
1
3
13
775
Now, we choose a vector xat random and set y=x+z,
x=2
6642
3
4
23
775y=x+z=2
6642
3
4
23
775+2
664 1
1
3
13
775=2
6641
2
7
13
775
and you can check that
T(x) =2
411
13
213
5=T(y)
A quicker solution is to take two elements of the kernel (in this case, scalar multiples of z) which both get
sent to 0byT. Quicker yet, take 0andzasxandy, which also both get sent to 0byT.
C25 Contributed by Robert Beezer Statement [552]
To nd the kernel, we require all x2C3such thatT(x) =0. This condition is
2x1 x2+ 5x3
4x1+ 2x2 10x3
=0
0
This leads to a homogeneous system of two linear equations in three variables, whose coecient matrix
row-reduces to
1 1
25
2
0 0 0
Version 2.30
558 Section ILT Injective Linear Transformations
With two free variables Theorem BNS [160] yields the basis for the null space
8
<
:2
4 5
2
0
13
5;2
41
2
1
03
59
=
;
Withn(T)6= 0,K(T)6=f0g, so Theorem KILT [548] says Tis not injective.
C26 Contributed by Chris Black Statement [552]
By Theorem ILTD [550], if a linear transformation T:U!Vis injective, then dim( U)dim(V). In this
case,T:C5!C4, and 5 = dim
C5
>dim
C4
= 4. Thus, Tcannot possibly be injective.
C27 Contributed by Chris Black Statement [552]
IfT0
@2
4x
y
z3
51
A=0, then2
42x+y+z
x y+ 2z
x+ 2y z3
5=0. Thus, we have the system
2x+y+z= 0
x y+ 2z= 0
x+ 2y z= 0
. Thus, we are looking for the nullspace of the matrix AT=2
42 1 1
1 1 2
1 2 13
5. SinceATrow-reduces to
2
410 1
01 1
0 0 03
5, the kernel of Tis all vectors where x= zandy=z. Thus,K(T) =*8
<
:2
4 1
1
13
59
=
;+
.
C28 Contributed by Chris Black Statement [553]
SinceTis given by matrix multiplication, K(T) =N(A). We have
2
6641 2 3 1
2 1 1 0
1 2 1 2
1 3 2 13
775RREF !2
666410 0 0
010 0
0 0 10
0 0 0 13
7775
The nullspace of Aisf0g, so the kernel of Tis also trivial:K(T) =f0g.
C29 Contributed by Chris Black Statement [553]
SinceTis given by matrix multiplication, K(T) =N(A). We have
2
6641 2 1 1
2 1 1 0
1 2 1 2
1 2 1 13
775RREF !2
66410 1=3 0
011=3 0
0 0 0 1
0 0 0 03
775
Thus, a basis for the nullspace of Ais8
>><
>>:2
664 1
1
3
03
7759
>>=
>>;, and the kernel is K(T) =**2
664 1
1
3
03
775++
. Since the kernel
is nontrivial, this linear transformation is not injective.
C30 Contributed by Chris Black Statement [553]
We can see without computing that Tis not injective, since the degree of M2;2is larger than the degree
Version 2.30
Subsection ILT.SOL Solutions 559
ofP2. However, that doesn't address the question of the kernel of T. We need to nd all matricesa b
c d
so that (a+b) + (a+c)x+ (a+d)x2= 0. This means a+b= 0,a+c= 0, anda+d= 0,
or equivalently, b=d=c= a. Thus, the kernel is a one-dimensional subspace of M2;2spanned by1 1
1 1
. Symbolically, we have K(T) =1 1
1 1
.
C31 Contributed by Chris Black Statement [553]
We have
T(e1) =2
42
0
13
5 T(e2) =2
41
2
03
5 T(e3) =2
40
1
23
5
Let's put these vectors into a matrix and row reduce to test their linear independence.
2
42 1 0
0 2 1
1 0 23
5RREF !2
410 0
010
0 0 13
5
so the set of vectors fT(e1); T(e1); T(e1)gis linearly independent.
C32 Contributed by Chris Black Statement [553]
We haveT(e1) =2
41
2
13
5andT(e2) =2
41
1
23
5. Putting these into a matrix as columns and row-reducing, we
have
2
41 1
2 1
1 23
5RREF !2
410
01
0 03
5
Thus, the set of vectors fT(e1); T(e2)gis linearly independent.
C33 Contributed by Chris Black Statement [553]
We have
T(e1) =2
666641
0
1
1
33
77775T(e2) =2
666643
1
2
0
13
77775T(e3) =2
666642
1
1
1
23
77775
Let's row reduce the matrix of Tto test linear independence.
2
666641 3 2
0 1 1
1 2 1
1 0 1
3 1 23
77775RREF !2
66666410 0
010
0 0 1
0 0 0
0 0 03
777775
so the set of vectors fT(e1); T(e2); T(e3)gis linearly independent.
C40 Contributed by Robert Beezer Statement [553]
We choose xto be any vector we like. A particularly cocky choice would be to choose x=0, but we will
instead choose
x=2 1
1 4
Version 2.30
560 Section ILT Injective Linear Transformations
ThenR(x) = 9 + 9x. Now compute the kernel of R, which by Theorem KILT [548] we expect to be
nontrivial. Setting Ra b
b c
equal to the zero vector, 0= 0 + 0x, and equating coecients leads to
a homogeneous system of equations. Row-reducing the coecient matrix of this system will allow us to
determine the values of a,bandcthat create elements of the null space of R,
2 1 1
1 1 2
RREF !10 1
011
We only need a single element of the null space of this coecient matrix, so we will not compute a precise
description of the whole null space. Instead, choose the free variable c= 2. Then
z= 2 2
2 2
is the corresponding element of the kernel. We compute the desired yas
y=x+z=2 1
1 4
+ 2 2
2 2
=0 3
3 6
Then check that R(y) = 9 + 9x.
T15 Contributed by Robert Beezer Statement [554]
We are asked to prove that K(T) is a subset of K(ST). Employing Denition SSET [761], choose
x2K(T). Then we know that T(x) =0. So
(ST) (x) =S(T(x)) Denition LTC [532]
=S(0) x2K(T)
=0 Theorem LTTZZ [519]
This qualies xfor membership in K(ST).
T20 Contributed by Andy Zimmer Statement [554]
This is an equality of sets, so we want to establish two subset conditions (Denition SE [762]).
First, showN(A)K(T). Choose x2N(A). Check to see if x2K(T),
T(x) =Ax Denition of T
=0 x 2N(A)
So by Denition KLT [545], x2K(T) and thusN(A)K(T).
Now, showK(T)N(A). Choose x2K(T). Check to see if x2N(A),
Ax=T(x) Denition of T
=0 x 2K(T)
So by Denition NSM [73], x2N(A) and thusK(T)N(A).
Version 2.30
Section SLT Surjective Linear Transformations 561
Section SLT
Surjective Linear Transformations
The companion to an injection is a surjection. Surjective linear transformations are closely related to
spanning sets and ranges. So as you read this section re
ect back on Section ILT [541] and note the
parallels and the contrasts. In the next section, Section IVLT [579], we will combine the two properties.
As usual, we lead with a denition.
Denition SLT
Surjective Linear Transformation
SupposeT:U!Vis a linear transformation. Then Tissurjective if for every v2Vthere exists a
u2Uso thatT(u) =v. 4
Given an arbitrary function, it is possible for there to be an element of the codomain that is not an
output of the function (think about the function y=f(x) =x2and the codomain element y= 3). For
a surjective function, this never happens. If we choose any element of the codomain ( v2V) then there
must be an input from the domain ( u2U) which will create the output when used to evaluate the linear
transformation ( T(u) =v). Some authors prefer the term onto where we use surjective, and we will
sometimes refer to a surjective linear transformation as a surjection .
Subsection ESLT
Examples of Surjective Linear Transformations
It is perhaps most instructive to examine a linear transformation that is not surjective rst.
Example NSAQ
Not surjective, Archetype Q
Archetype Q [844] is the linear transformation
T:C5!C5; T0
BBBB@2
66664x1
x2
x3
x4
x53
777751
CCCCA=2
66664 2x1+ 3x2+ 3x3 6x4+ 3x5
16x1+ 9x2+ 12x3 28x4+ 28x5
19x1+ 7x2+ 14x3 32x4+ 37x5
21x1+ 9x2+ 15x3 35x4+ 39x5
9x1+ 5x2+ 7x3 16x4+ 16x53
77775
We will demonstrate that
v=2
66664 1
2
3
1
43
77775
is an unobtainable element of the codomain. Suppose to the contrary that uis an element of the domain
such thatT(u) =v. Then
2
66664 1
2
3
1
43
77775=v=T(u) =T0
BBBB@2
66664u1
u2
u3
u4
u53
777751
CCCCA
Version 2.30
562 Section SLT Surjective Linear Transformations
=2
66664 2u1+ 3u2+ 3u3 6u4+ 3u5
16u1+ 9u2+ 12u3 28u4+ 28u5
19u1+ 7u2+ 14u3 32u4+ 37u5
21u1+ 9u2+ 15u3 35u4+ 39u5
9u1+ 5u2+ 7u3 16u4+ 16u53
77775
=2
66664 2 3 3 6 3
16 9 12 28 28
19 7 14 32 37
21 9 15 35 39
9 5 7 16 163
777752
66664u1
u2
u3
u4
u53
77775
Now we recognize the appropriate input vector uas a solution to a linear system of equations. Form the
augmented matrix of the system, and row-reduce to
2
66666410 0 0 1 0
010 0 4
30
0 0 10 1
30
0 0 0 1 1 0
0 0 0 0 0 13
777775
With a leading 1 in the last column, Theorem RCLS [58] tells us the system is inconsistent. From the
absence of any solutions we conclude that no such vector uexists, and by Denition SLT [559], Tis not
surjective.
Again, do not concern yourself with how vwas selected, as this will be explained shortly. However, do
understand whythis vector provides enough evidence to conclude that Tis not surjective.
To show that a linear transformation is not surjective, it is enough to nd a single element of the
codomain that is never created by any input, as in Example NSAQ [559]. However, to show that a linear
transformation is surjective we must establish that every element of the codomain occurs as an output of
the linear transformation for some appropriate input.
Example SAR
Surjective, Archetype R
Archetype R [848] is the linear transformation
T:C5!C5; T0
BBBB@2
66664x1
x2
x3
x4
x53
777751
CCCCA=2
66664 65x1+ 128x2+ 10x3 262x4+ 40x5
36x1 73x2 x3+ 151x4 16x5
44x1+ 88x2+ 5x3 180x4+ 24x5
34x1 68x2 3x3+ 140x4 18x5
12x1 24x2 x3+ 49x4 5x53
77775
To establish that Ris surjective we must begin with a totally arbitrary element of the codomain, vand
somehow nd an input vector usuch thatT(u) =v. We desire,
T(u) =v
2
66664 65u1+ 128u2+ 10u3 262u4+ 40u5
36u1 73u2 u3+ 151u4 16u5
44u1+ 88u2+ 5u3 180u4+ 24u5
34u1 68u2 3u3+ 140u4 18u5
12u1 24u2 u3+ 49u4 5u53
77775=2
66664v1
v2
v3
v4
v53
77775
2
66664 65 128 10 262 40
36 73 1 151 16
44 88 5 180 24
34 68 3 140 18
12 24 1 49 53
777752
66664u1
u2
u3
u4
u53
77775=2
66664v1
v2
v3
v4
v53
77775
Version 2.30
Subsection SLT.ESLT Examples of Surjective Linear Transformations 563
We recognize this equation as a system of equations in the variables ui, but our vector of constants contains
symbols. In general, we would have to row-reduce the augmented matrix by hand, due to the symbolic
nal column. However, in this particular example, the 5 5 coecient matrix is nonsingular and so has
an inverse (Theorem NI [261], Denition MI [244]).
2
66664 65 128 10 262 40
36 73 1 151 16
44 88 5 180 24
34 68 3 140 18
12 24 1 49 53
77775 1
=2
66664 47 92 1 181 14
27 557
2221
211
32 64 1 126 12
25 503
2199
29
9 181
271
243
77775
so we nd that
2
66664u1
u2
u3
u4
u53
77775=2
66664 47 92 1 181 14
27 557
2221
211
32 64 1 126 12
25 503
2199
29
9 181
271
243
777752
66664v1
v2
v3
v4
v53
77775
=2
66664 47v1+ 92v2+v3 181v4 14v5
27v1 55v2+7
2v3+221
2v4+ 11v5
32v1+ 64v2 v3 126v4 12v5
25v1 50v2+3
2v3+199
2v4+ 9v5
9v1 18v2+1
2v3+71
2v4+ 4v53
77775
This establishes that if we are given anyoutput vector v, we can use its components in this nal expression
to formulate a vector usuch thatT(u) =v. So by Denition SLT [559] we now know that Tis surjective.
You might try to verify this condition in its full generality (i.e. evaluate Twith this nal expression and see
if you get vas the result), or test it more specically for some numerical vector v(see Exercise SLT.C20
[571]).
Let's now examine a surjective linear transformation between abstract vector spaces.
Example SAV
Surjective, Archetype V
Archetype V [858] is dened by
T:P3!M22; T
a+bx+cx2+dx3
=a+b a 2c
d b d
To establish that the linear transformation is surjective, begin by choosing an arbitrary output. In this
example, we need to choose an arbitrary 2 2 matrix, say
v=x y
z w
and we would like to nd an input polynomial
u=a+bx+cx2+dx3
so thatT(u) =v. So we have,
x y
z w
=v
=T(u)
Version 2.30
564 Section SLT Surjective Linear Transformations
=T
a+bx+cx2+dx3
=a+b a 2c
d b d
Matrix equality leads us to the system of four equations in the four unknowns, x;y;z;w ,
a+b=x
a 2c=y
d=z
b d=w
which can be rewritten as a matrix equation,
2
6641 1 0 0
1 0 2 0
0 0 0 1
0 1 0 13
7752
664a
b
c
d3
775=2
664x
y
z
w3
775
The coecient matrix is nonsingular, hence it has an inverse,
2
6641 1 0 0
1 0 2 0
0 0 0 1
0 1 0 13
775 1
=2
6641 0 1 1
0 0 1 1
1
2 1
2 1
2 1
2
0 0 1 03
775
so we have
2
664a
b
c
d3
775=2
6641 0 1 1
0 0 1 1
1
2 1
2 1
2 1
2
0 0 1 03
7752
664x
y
z
w3
775
=2
664x z w
z+w
1
2(x y z w)
z3
775
So the input polynomial u= (x z w) + (z+w)x+1
2(x y z w)x2+zx3will yield the output matrix
v, no matter what form vtakes. This means by Denition SLT [559] that Tis surjective. All the same,
let's do a concrete demonstration and evaluate Twithu,
T(u) =T
(x z w) + (z+w)x+1
2(x y z w)x2+zx3
=(x z w) + (z+w) (x z w) 2(1
2(x y z w))
z (z+w) z
=x y
z w
=v
Version 2.30
Subsection SLT.RLT Range of a Linear Transformation 565
Subsection RLT
Range of a Linear Transformation
For a linear transformation T:U!V, the range is a subset of the codomain V. Informally, it is the set
of all outputs that the transformation creates when fed every possible input from the domain. It will have
some natural connections with the column space of a matrix, so we will keep the same notation, and if you
think about your objects, then there should be little confusion. Here's the careful denition.
Denition RLT
Range of a Linear Transformation
SupposeT:U!Vis a linear transformation. Then the range ofTis the set
R(T) =fT(u)ju2Ug
(This denition contains Notation RLT.) 4
Example RAO
Range, Archetype O
Archetype O [839] is the linear transformation
T:C3!C5; T0
@2
4x1
x2
x33
51
A=2
66664 x1+x2 3x3
x1+ 2x2 4x3
x1+x2+x3
2x1+ 3x2+x3
x1+ 2x33
77775
To determine the elements of C5inR(T), nd those vectors vsuch thatT(u) =vfor some u2C3,
v=T(u)
=2
66664 u1+u2 3u3
u1+ 2u2 4u3
u1+u2+u3
2u1+ 3u2+u3
u1+ 2u33
77775
=2
66664 u1
u1
u1
2u1
u13
77775+2
66664u2
2u2
u2
3u2
03
77775+2
66664 3u3
4u3
u3
u3
2u33
77775
=u12
66664 1
1
1
2
13
77775+u22
666641
2
1
3
03
77775+u32
66664 3
4
1
1
23
77775
This says that every output of T(v) can be written as a linear combination of the three vectors
2
66664 1
1
1
2
13
777752
666641
2
1
3
03
777752
66664 3
4
1
1
23
77775
Version 2.30
566 Section SLT Surjective Linear Transformations
using the scalars u1; u2; u3. Furthermore, since ucan be any element of C3, every such linear combination
is an output. This means that
R(T) =*8
>>>><
>>>>:2
66664 1
1
1
2
13
77775;2
666641
2
1
3
03
77775;2
66664 3
4
1
1
23
777759
>>>>=
>>>>;+
The three vectors in this spanning set for R(T) form a linearly dependent set (check this!). So we can
nd a more economical presentation by any of the various methods from Section CRS [271] and Section
FS [293]. We will place the vectors into a matrix as rows, row-reduce, toss out zero rows and appeal to
Theorem BRS [280], so we can describe the range of Twith a basis,
R(T) =*8
>>>><
>>>>:2
666641
0
3
7
23
77775;2
666640
1
2
5
13
777759
>>>>=
>>>>;+
We know that the span of a set of vectors is always a subspace (Theorem SSS [339]), so the range
computed in Example RAO [563] is also a subspace. This is no accident, the range of a linear transformation
isalways a subspace.
Theorem RLTS
Range of a Linear Transformation is a Subspace
Suppose that T:U!Vis a linear transformation. Then the range of T,R(T), is a subspace of V.
Proof We can apply the three-part test of Theorem TSS [334]. First, 0U2UandT(0U) =0Vby
Theorem LTTZZ [519], so 0V2R(T) and we know that the range is non-empty.
Suppose we assume that x;y2R(T). Isx+y2R(T)? Ifx;y2R(T) then we know there are vectors
w;z2Usuch thatT(w) =xandT(z) =y. BecauseUis a vector space, additive closure (Property AC
[317]) implies that w+z2U. Then
T(w+z) =T(w) +T(z) Denition LT [515]
=x+y Denition of wandz
So we have found an input, w+z, which when fed into Tcreates x+yas an output. This qualies x+y
for membership in R(T). So we have additive closure.
Suppose we assume that 2Candx2R(T). Isx2R(T)? If x2R(T), then there is a vector
w2Usuch thatT(w) =x. BecauseUis a vector space, scalar closure implies that w2U. Then
T(w) =T(w) Denition LT [515]
=x Denition of w
So we have found an input ( w) which when fed into Tcreatesxas an output. This qualies xfor
membership inR(T). So we have scalar closure and Theorem TSS [334] tells us that R(T) is a subspace
ofV.
Let's compute another range, now that we know in advance that it will be a subspace.
Example FRAN
Full range, Archetype N
Version 2.30
Subsection SLT.RLT Range of a Linear Transformation 567
Archetype N [836] is the linear transformation
T:C5!C3; T0
BBBB@2
66664x1
x2
x3
x4
x53
777751
CCCCA=2
42x1+x2+ 3x3 4x4+ 5x5
x1 2x2+ 3x3 9x4+ 3x5
3x1+ 4x3 6x4+ 5x53
5
To determine the elements of C3inR(T), nd those vectors vsuch thatT(u) =vfor some u2C5,
v=T(u)
=2
42u1+u2+ 3u3 4u4+ 5u5
u1 2u2+ 3u3 9u4+ 3u5
3u1+ 4u3 6u4+ 5u53
5
=2
42u1
u1
3u13
5+2
4u2
2u2
03
5+2
43u3
3u3
4u33
5+2
4 4u4
9u4
6u43
5+2
45u5
3u5
5u53
5
=u12
42
1
33
5+u22
41
2
03
5+u32
43
3
43
5+u42
4 4
9
63
5+u52
45
3
53
5
This says that every output of T(v) can be written as a linear combination of the ve vectors
2
42
1
33
52
41
2
03
52
43
3
43
52
4 4
9
63
52
45
3
53
5
using the scalars u1; u2; u3; u4; u5. Furthermore, since ucan be any element of C5, every such linear
combination is an output. This means that
R(T) =*8
<
:2
42
1
33
5;2
41
2
03
5;2
43
3
43
5;2
4 4
9
63
5;2
45
3
53
59
=
;+
The ve vectors in this spanning set for R(T) form a linearly dependent set (Theorem MVSLD [158]). So
we can nd a more economical presentation by any of the various methods from Section CRS [271] and
Section FS [293]. We will place the vectors into a matrix as rows, row-reduce, toss out zero rows and
appeal to Theorem BRS [280], so we can describe the range of Twith a (nice) basis,
R(T) =*8
<
:2
41
0
03
5;2
40
1
03
5;2
40
0
13
59
=
;+
=C3
In contrast to injective linear transformations having small (trivial) kernels (Theorem KILT [548]),
surjective linear transformations have large ranges, as indicated in the next theorem.
Theorem RSLT
Range of a Surjective Linear Transformation
Suppose that T:U!Vis a linear transformation. Then Tis surjective if and only if the range of T
equals the codomain, R(T) =V.
Proof ()) By Denition RLT [563], we know that R(T)V. To establish the reverse inclusion, assume
v2V. Then since Tis surjective (Denition SLT [559]), there exists a vector u2Uso thatT(u) =v.
However, the existence of ugains vmembership inR(T), soVR(T). Thus,R(T) =V.
Version 2.30
568 Section SLT Surjective Linear Transformations
(() To establish that Tis surjective, choose v2V. Since we are assuming that R(T) =V,v2R(T).
This says there is a vector u2Uso thatT(u) =v, i.e.Tis surjective.
Example NSAQR
Not surjective, Archetype Q, revisited
We are now in a position to revisit our rst example in this section, Example NSAQ [559]. In that example,
we showed that Archetype Q [844] is not surjective by constructing a vector in the codomain where no
element of the domain could be used to evaluate the linear transformation to create the output, thus
violating Denition SLT [559]. Just where did this vector come from?
The short answer is that the vector
v=2
66664 1
2
3
1
43
77775
was constructed to lie outside of the range of T. How was this accomplished? First, the range of Tis given
by
R(T) =*8
>>>><
>>>>:2
666641
0
0
0
13
77775;2
666640
1
0
0
13
77775;2
666640
0
1
0
13
77775;2
666640
0
0
1
23
777759
>>>>=
>>>>;+
Suppose an element of the range vhas its rst 4 components equal to 1;2;3; 1, in that order. Then
to be an element of R(T), we would have
v= ( 1)2
666641
0
0
0
13
77775+ (2)2
666640
1
0
0
13
77775+ (3)2
666640
0
1
0
13
77775+ ( 1)2
666640
0
0
1
23
77775=2
66664 1
2
3
1
83
77775
So the only vector in the range with these rst four components specied, must have 8 in the fth
component. To set the fth component to any other value (say, 4) will result in a vector ( vin Example
NSAQ [559]) outside of the range. Any attempt to nd an input for Tthat will produce vas an output
will be doomed to failure.
Whenever the range of a linear transformation is not the whole codomain, we can employ this device
and conclude that the linear transformation is not surjective. This is another way of viewing Theorem
RSLT [565]. For a surjective linear transformation, the range is all of the codomain and there is no choice
for a vector vthat lies in V, yet not in the range. For every one of the archetypes that is not surjective,
there is an example presented of exactly this form.
Example NSAO
Not surjective, Archetype O
In Example RAO [563] the range of Archetype O [839] was determined to be
R(T) =*8
>>>><
>>>>:2
666641
0
3
7
23
77775;2
666640
1
2
5
13
777759
>>>>=
>>>>;+
Version 2.30
Subsection SLT.SSSLT Spanning Sets and Surjective Linear Transformations 569
a subspace of dimension 2 in C5. SinceR(T)6=C5, Theorem RSLT [565] says Tis not surjective.
Example SAN
Surjective, Archetype N
The range of Archetype N [836] was computed in Example FRAN [564] to be
R(T) =*8
<
:2
41
0
03
5;2
40
1
03
5;2
40
0
13
59
=
;+
Since the basis for this subspace is the set of standard unit vectors for C3(Theorem SUVB [371]), we have
R(T) =C3and by Theorem RSLT [565], Tis surjective.
Subsection SSSLT
Spanning Sets and Surjective Linear Transformations
Just as injective linear transformations are allied with linear independence (Theorem ILTLI [549], Theorem
ILTB [550]), surjective linear transformations are allied with spanning sets.
Theorem SSRLT
Spanning Set for Range of a Linear Transformation
Suppose that T:U!Vis a linear transformation and S=fu1;u2;u3; :::; utgspansU. Then
R=fT(u1); T(u2); T(u3); :::; T (ut)g
spansR(T).
Proof We need to establish that R(T) =hRi, a set equality. First we establish that R(T)hRi. To
this end, choose v2R(T). Then there exists a vector u2U, such that T(u) =v(Denition RLT [563]).
BecauseSspansUthere are scalars, a1; a2; a3; :::; at, such that
u=a1u1+a2u2+a3u3++atut
Then
v=T(u) Denition RLT [563]
=T(a1u1+a2u2+a3u3++atut) Denition TSVS [356]
=a1T(u1) +a2T(u2) +a3T(u3) +:::+atT(ut) Theorem LTLC [525]
which establishes that v2hRi(Denition SS [339]). So R(T)hRi.
To establish the opposite inclusion, choose an element of the span of R, say v2hRi. Then there are
scalarsb1; b2; b3; :::; btso that
v=b1T(u1) +b2T(u2) +b3T(u3) ++btT(ut) Denition SS [339]
=T(b1u1+b2u2+b3u3++btut) Theorem LTLC [525]
This demonstrates that vis an output of the linear transformation T, sov2R(T). ThereforehRiR (T),
so we have the set equality R(T) =hRi(Denition SE [762]). In other words, RspansR(T) (Denition
TSVS [356]).
Theorem SSRLT [567] provides an easy way to begin the construction of a basis for the range of a linear
transformation, since the construction of a spanning set requires simply evaluating the linear transformation
Version 2.30
570 Section SLT Surjective Linear Transformations
on a spanning set of the domain. In practice the best choice for a spanning set of the domain would be
as small as possible, in other words, a basis. The resulting spanning set for the codomain may not be
linearly independent, so to nd a basis for the range might require tossing out redundant vectors from the
spanning set. Here's an example.
Example BRLT
A basis for the range of a linear transformation
Dene the linear transformation T:M22!P2by
Ta b
c d
= (a+ 2b+ 8c+d) + ( 3a+ 2b+ 5d)x+ (a+b+ 5c)x2
A convenient spanning set for M22is the basis
S=1 0
0 0
;0 1
0 0
;0 0
1 0
;0 0
0 1
So by Theorem SSRLT [567], a spanning set for R(T) is
R=
T1 0
0 0
; T0 1
0 0
; T0 0
1 0
; T0 0
0 1
=
1 3x+x2;2 + 2x+x2;8 + 5x2;1 + 5x
The setRis not linearly independent, so if we desire a basis for R(T), we need to eliminate some redundant
vectors. Two particular relations of linear dependence on Rare
( 2)(1 3x+x2) + ( 3)(2 + 2x+x2) + (8 + 5x2) = 0 + 0x+ 0x2=0
(1 3x+x2) + ( 1)(2 + 2x+x2) + (1 + 5x) = 0 + 0x+ 0x2=0
These, individually, allow us to remove 8 + 5 x2and 1 + 5xfromRwith out destroying the property that
RspansR(T). The two remaining vectors are linearly independent (check this!), so we can write
R(T) =
1 3x+x2;2 + 2x+x2
and see that dim ( R(T)) = 2.
Elements of the range are precisely those elements of the codomain with non-empty preimages.
Theorem RPI
Range and Pre-Image
Suppose that T:U!Vis a linear transformation. Then
v2R(T) if and only if T 1(v)6=;
Proof ()) Ifv2R(T), then there is a vector u2Usuch thatT(u) =v. This qualies ufor membership
inT 1(v), and thus the preimage of vis not empty.
(() Suppose the preimage of vis not empty, so we can choose a vector u2Usuch thatT(u) =v.
Then v2R(T).
Theorem SLTB
Surjective Linear Transformations and Bases
Suppose that T:U!Vis a linear transformation and B=fu1;u2;u3; :::; umgis a basis of U. ThenT
is surjective if and only if C=fT(u1); T(u2); T(u3); :::; T (um)gis a spanning set for V.
Proof ()) AssumeTis surjective. Since Bis a basis, we know Bis a spanning set of U(Denition
B [371]). Then Theorem SSRLT [567] says that CspansR(T). But the hypothesis that Tis surjective
meansV=R(T) (Theorem RSLT [565]), so CspansV.
Version 2.30
Subsection SLT.SLTD Surjective Linear Transformations and Dimension 571
(() Assume that CspansV. To establish that Tis surjective, we will show that every element of V
is an output of Tfor some input (Denition SLT [559]). Suppose that v2V. As an element of V, we can
write vas a linear combination of the spanning set C. So there are are scalars, b1; b2; b3; :::; bm, such
that
v=b1T(u1) +b2T(u2) +b3T(u3) ++bmT(um)
Now dene the vector u2Uby
u=b1u1+b2u2+b3u3++bmum
Then
T(u) =T(b1u1+b2u2+b3u3++bmum)
=b1T(u1) +b2T(u2) +b3T(u3) ++bmT(um) Theorem LTLC [525]
=v
So, given any choice of a vector v2V, we can design an input u2Uto produce vas an output of T.
Thus, by Denition SLT [559], Tis surjective.
Subsection SLTD
Surjective Linear Transformations and Dimension
Theorem SLTD
Surjective Linear Transformations and Dimension
Suppose that T:U!Vis a surjective linear transformation. Then dim ( U)dim (V).
Proof Suppose to the contrary that m= dim (U)<dim (V) =t. LetBbe a basis of U, which will then
containmvectors. Apply Tto each element of Bto form a set Cthat is a subset of V. By Theorem SLTB
[568],Cis spanning set of Vwithmor fewer vectors. So we have a set of mor fewer vectors that span V,
a vector space of dimension t, withm<t . However, this contradicts Theorem G [407], so our assumption
is false and dim ( U)dim (V).
Example NSDAT
Not surjective by dimension, Archetype T
The linear transformation in Archetype T [854] is
T:P4!P5; T (p(x)) = (x 2)p(x)
Since dim (P4) = 5<6 = dim (P5),Tcannot be surjective for then it would violate Theorem SLTD [569].
Notice that the previous example made no use of the actual formula dening the function. Merely
a comparison of the dimensions of the domain and codomain are enough to conclude that the linear
transformation is not surjective. Archetype O [839] and Archetype P [842] are two more examples of linear
transformations that have \small" domains and \big" codomains, resulting in an inability to create all
possible outputs and thus they are non-surjective linear transformations.
Version 2.30
572 Section SLT Surjective Linear Transformations
Subsection CSLT
Composition of Surjective Linear Transformations
In Subsection LT.NLTFO [530] we saw how to combine linear transformations to build new linear trans-
formations, specically, how to build the composition of two linear transformations (Denition LTC [532]).
It will be useful later to know that the composition of surjective linear transformations is again surjective,
so we prove that here.
Theorem CSLTS
Composition of Surjective Linear Transformations is Surjective
Suppose that T:U!VandS:V!Ware surjective linear transformations. Then ( ST):U!Wis a
surjective linear transformation.
Proof That the composition is a linear transformation was established in Theorem CLTLT [533], so we
need only establish that the composition is surjective. Applying Denition SLT [559], choose w2W.
BecauseSis surjective, there must be a vector v2V, such that S(v) =w. With the existence of v
established, that Tis surjective guarantees a vector u2Usuch thatT(u) =v. Now,
(ST) (u) =S(T(u)) Denition LTC [532]
=S(v) Denition of u
=w Denition of v
This establishes that any element of the codomain ( w) can be created by evaluating STwith the right
input ( u). Thus, by Denition SLT [559], STis surjective.
Subsection READ
Reading Questions
1. Suppose T:C5!C8is a linear transformation. Why can't Tbe surjective?
2. What is the relationship between a surjective linear transformation and its range?
3. Compare and contrast injective and surjective linear transformations.
Version 2.30
Subsection SLT.EXC Exercises 573
Subsection EXC
Exercises
C10 Each archetype below is a linear transformation. Compute the range for each.
Archetype M [833]
Archetype N [836]
Archetype O [839]
Archetype P [842]
Archetype Q [844]
Archetype R [848]
Archetype S [851]
Archetype T [854]
Archetype U [856]
Archetype V [858]
Archetype W [860]
Archetype X [862]
Contributed by Robert Beezer
C20 Example SAR [560] concludes with an expression for a vector u2C5that we believe will create the
vector v2C5when used to evaluate T. That is,T(u) =v. Verify this assertion by actually evaluating T
withu. If you don't have the patience to push around all these symbols, try choosing a numerical instance
ofv, compute u, and then compute T(u), which should result in v.
Contributed by Robert Beezer
C22 The linear transformation S:C4!C3is not surjective. Find an output w2C3that has an empty
pre-image (that is S 1(w) =;.)
S0
BB@2
664x1
x2
x3
x43
7751
CCA=2
42x1+x2+ 3x3 4x4
x1+ 3x2+ 4x3+ 3x4
x1+ 2x2+x3+ 7x43
5
Contributed by Robert Beezer Solution [574]
C23 Determine whether or not the following linear transformation T:C5!P3is surjective:
T0
BBBB@2
66664a
b
c
d
e3
777751
CCCCA=a+ (b+c)x+ (c+d)x2+ (d+e)x3
Contributed by Chris Black Solution [574]
C24 Determine whether or not the linear transformation T:P3!C5below is surjective:
T
a+bx+cx2+dx3
=2
66664a+b
b+c
c+d
a+c
b+d3
77775:
Version 2.30
574 Section SLT Surjective Linear Transformations
Contributed by Chris Black Solution [575]
C25 Dene the linear transformation
T:C3!C2; T0
@2
4x1
x2
x33
51
A=2x1 x2+ 5x3
4x1+ 2x2 10x3
Find a basis for the range of T,R(T). IsTsurjective?
Contributed by Robert Beezer Solution [575]
C26 LetT:C3!C3be given by T0
@2
4a
b
c3
51
A=2
4a+b+ 2c
2c
a+b+c3
5. Find a basis of R(T). IsTsurjective?
Contributed by Chris Black Solution [575]
C27 LetT:C3!C4be given by T0
@2
4a
b
c3
51
A=2
664a+b c
a b+c
a+b+c
a+b+c3
775. Find a basis of R(T). IsTsurjective?
Contributed by Chris Black Solution [576]
C28 LetT:C4!M2;2be given by T0
BB@2
664a
b
c
d3
7751
CCA=a+b a +b+c
a+b+c a +d
. Find a basis of R(T). IsT
surjective?
Contributed by Chris Black Solution [576]
C29 LetT:P2!P4be given by T(p(x)) =x2p(x). Find a basis of R(T). IsTsurjective?
Contributed by Chris Black Solution [576]
C30 LetT:P4!P3be given by T(p(x)) =p0(x), wherep0(x) is the derivative. Find a basis of R(T).
IsTsurjective?
Contributed by Chris Black Solution [576]
C40 Show that the linear transformation Tis not surjective by nding an element of the codomain, v,
such that there is no vector uwithT(u) =v.
T:C3!C3; T0
@2
4a
b
c3
51
A=2
42a+ 3b c
2b 2c
a b+ 2c3
5
Contributed by Robert Beezer Solution [577]
M60 SupposeUandVare vector spaces. Dene the function Z:U!VbyT(u) =0Vfor every
u2U. Then by Exercise LT.M60 [536], Zis a linear transformation. Formulate a condition on Vthat
is equivalent to Zbeing an surjective linear transformation. In other words, ll in the blank to complete
the following statement (and then give a proof): Zis surjective if and only if Vis . (See
Exercise ILT.M60 [553], Exercise IVLT.M60 [594].)
Contributed by Robert Beezer
T15 Suppose that that T:U!VandS:V!Ware linear transformations. Prove the following
relationship between ranges.
R(ST)R(S)
Version 2.30
Subsection SLT.EXC Exercises 575
Contributed by Robert Beezer Solution [577]
T20 Suppose that Ais anmnmatrix. Dene the linear transformation Tby
T:Cn!Cm; T (x) =Ax
Prove that the range of Tequals the column space of A,R(T) =C(A).
Contributed by Andy Zimmer Solution [577]
Version 2.30
576 Section SLT Surjective Linear Transformations
Subsection SOL
Solutions
C22 Contributed by Robert Beezer Statement [571]
To nd an element of C3with an empty pre-image, we will compute the range of the linear transformation
R(S) and then nd an element outside of this set.
By Theorem SSRLT [567] we can evaluate Swith the elements of a spanning set of the domain and
create a spanning set for the range.
S0
BB@2
6641
0
0
03
7751
CCA=2
42
1
13
5S0
BB@2
6640
1
0
03
7751
CCA=2
41
3
23
5S0
BB@2
6640
0
1
03
7751
CCA=2
43
4
13
5S0
BB@2
6640
0
0
13
7751
CCA=2
4 4
3
73
5
So
R(S) =*8
<
:2
42
1
13
5;2
41
3
23
5;2
43
4
13
5;2
4 4
3
73
59
=
;+
This spanning set is obviously linearly dependent, so we can reduce it to a basis for R(S) using Theorem
BRS [280], where the elements of the spanning set are placed as the rows of a matrix. The result is that
R(S) =*8
<
:2
41
0
13
5;2
40
1
13
59
=
;+
Therefore, the unique vector in R(S) with a rst slot equal to 6 and a second slot equal to 15 will be the
linear combination
62
41
0
13
5+ 152
40
1
13
5=2
46
15
93
5
So, any vector with rst two components equal to 6 and 15, but with a third component dierent from 9,
such as
w=2
46
15
633
5
will not be an element of the range of Sand will therefore have an empty pre-image. Another strategy
on this problem is to guess . Almost any vector will lie outside the range of T, you have to be unlucky to
randomly choose an element of the range. This is because the codomain has dimension 3, while the range
is \much smaller" at a dimension of 2. You still need to check that your guess lies outside of the range,
which generally will involve solving a system of equations that turns out to be inconsistent.
C23 Contributed by Chris Black Statement [571]
The linear transformation Tis surjective if for any p(x) =+x+
x2+x3, there is a vector u=2
66664a
b
c
d
e3
77775
inC5so thatT(u) =p(x). We need to be able to solve the system
a=
b+c=
Version 2.30
Subsection SLT.SOL Solutions 577
c+d=
d+e=
This system has an innite number of solutions, one of which is a=,b=,c= 0,d=
ande=
,
so that
T0
BBBB@2
66664
0
3
777751
CCCCA=+ (+ 0)x+ (0 +
)x2+ (
+ (
))x3
=+x+
x2+x3
=p(x):
Thus,Tis surjective, since for every vector v2P3, there exists a vector u2C5so thatT(u) =v.
C24 Contributed by Chris Black Statement [571]
According to Theorem SLTD [569], if a linear transformation T:U!Vis surjective, then dim ( U)
dim (V). In this example, U=P3has dimension 4, and V=C5has dimension 5, so Tcannot be
surjective. (There is no way Tcan \expand" the domain P3to ll the codomain C5.)
C25 Contributed by Robert Beezer Statement [572]
To nd the range of T, applyTto the elements of a spanning set for C3as suggested in Theorem SSRLT
[567]. We will use the standard basis vectors (Theorem SUVB [371]).
R(T) =hfT(e1); T(e2); T(e3)gi=2
4
; 1
2
;5
10
Each of these vectors is a scalar multiple of the others, so we can toss two of them in reducing the spanning
set to a linearly independent set (or be more careful and apply Theorem BCS [274] on a matrix with these
three vectors as columns). The result is the basis of the range,
1
2
Withr(T)6= 2,R(T)6=C2, so Theorem RSLT [565] says Tis not surjective.
C26 Contributed by Chris Black Statement [572]
The range of Tis
R(T) =8
<
:2
4a+b+ 2c
2c
a+b+c3
5a;b;c2C9
=
;
=8
<
:a2
41
0
13
5+b2
41
0
13
5+c2
42
2
13
5a;b;c2C9
=
;
=*2
41
0
13
5;2
42
2
13
5+
Since the vectors2
41
0
13
5and2
42
2
13
5are linearly independent (why?), a basis of R(T) is8
<
:2
41
0
13
5;2
42
2
13
59
=
;. Since
the dimension of the range is 2 and the dimension of the codomain is 3, Tis not surjective.
Version 2.30
578 Section SLT Surjective Linear Transformations
C27 Contributed by Chris Black Statement [572]
The range of Tis
R(T) =8
>><
>>:2
664a+b c
a b+c
a+b+c
a+b+c3
775a;b;c2C9
>>=
>>;
=8
>><
>>:a2
6641
1
1
13
775+b2
6641
1
1
13
775+c2
664 1
1
1
13
775a;b;c2C9
>>=
>>;
=*2
6641
1
1
13
775;2
6641
1
1
13
775;2
664 1
1
1
13
775+
By row reduction (not shown), we can see that the set
8
>><
>>:2
6641
1
1
13
775;2
6641
1
1
13
775;2
664 1
1
1
13
7759
>>=
>>;
are linearly independent, so is a basis of R(T). Since the dimension of the range is 3 and the dimension
of the codomain is 4, Tis not surjective. (We should have anticipated that Twas not surjective since the
dimension of the domain is smaller than the dimension of the codomain.)
C28 Contributed by Chris Black Statement [572]
The range of Tis
R(T) =a+b a +b+c
a+b+c a +da;b;c;d2C
=
a1 1
1 1
+b1 1
1 0
+c0 1
1 0
+d0 0
0 1a;b;c;d2C
=1 1
1 1
;1 1
1 0
;0 1
1 0
;0 0
0 1
=1 1
1 0
;0 1
1 0
;0 0
0 1
:
Can you explain the last equality above?
These three matrices are linearly independent, so a basis of R(T) is1 1
1 0
;0 1
1 0
;0 0
0 1
. Thus,
Tis not surjective, since the range has dimension 3 which is shy of dim ( M2;2) = 4. (Notice that the range
is actually the subspace of symmetric 2 2 matrices in M2;2.)
C29 Contributed by Chris Black Statement [572]
If we transform the basis of P2, then Theorem SSRLT [567] guarantees we will have a spanning set of R(T).
A basis ofP2is
1;x;x2
. If we transform the elements of this set, we get the set
x2;x3;x4
which is a
spanning set forR(T). These three vectors are linearly independent, so
x2;x3;x4
is a basis ofR(T).
C30 Contributed by Chris Black Statement [572]
If we transform the basis of P4, then Theorem SSRLT [567] guarantees we will have a spanning set of R(T).
A basis ofP4is
1;x;x2;x3;x4
. If we transform the elements of this set, we get the set
0;1;2x;3x2;4x3
Version 2.30
Subsection SLT.SOL Solutions 579
which is a spanning set for R(T). Reducing this to a linearly independent set, we nd that f1;2x;3x2;4x3g
is a basis ofR(T). SinceR(T) andP3both have dimension 4, Tis surjective.
C40 Contributed by Robert Beezer Statement [572]
We wish to nd an output vector vthat has no associated input. This is the same as requiring that there
is no solution to the equality
v=T0
@2
4a
b
c3
51
A=2
42a+ 3b c
2b 2c
a b+ 2c3
5=a2
42
0
13
5+b2
43
2
13
5+c2
4 1
2
23
5
In other words, we would like to nd an element of C3not in the set
Y=*8
<
:2
42
0
13
5;2
43
2
13
5;2
4 1
2
23
59
=
;+
If we make these vectors the rows of a matrix, and row-reduce, Theorem BRS [280] provides an alternate
description of Y,
Y=*8
<
:2
42
0
13
5;2
40
4
53
59
=
;+
If we add these vectors together, and then change the third component of the result, we will create a vector
that lies outside of Y, sayv=2
42
4
93
5.
T15 Contributed by Robert Beezer Statement [572]
This question asks us to establish that one set ( R(ST)) is a subset of another ( R(S)). Choose an element
in the \smaller" set, say w2R(ST). Then we know that there is a vector u2Usuch that
w= (ST) (u) =S(T(u))
Now dene v=T(u), so that then
S(v) =S(T(u)) =w
This statement is sucient to show that w2R(S), sowis an element of the \larger" set, and R(ST)
R(S).
T20 Contributed by Andy Zimmer Statement [573]
This is an equality of sets, so we want to establish two subset conditions (Denition SE [762]).
First, showC(A)R(T). Choose y2C(A). Then by Denition CSM [271] and Denition MVP [223]
there is a vector x2Cnsuch thatAx=y. Then
T(x) =Ax Denition of T
=y
This statement qualies yas a member ofR(T) (Denition RLT [563]), so C(A)R(T).
Now, showR(T)C(A). Choose y2R(T). Then by Denition RLT [563], there is a vector xinCn
such thatT(x) =y. Then
Ax=T(x) Denition of T
=y
So by Denition CSM [271] and Denition MVP [223], yqualies for membership in C(A) and soR(T)
C(A).
Version 2.30
580 Section SLT Surjective Linear Transformations
Version 2.30
Section IVLT Invertible Linear Transformations 581
Section IVLT
Invertible Linear Transformations
In this section we will conclude our introduction to linear transformations by bringing together the twin
properties of injectivity and surjectivity and consider linear transformations with both of these proper-
ties.
Subsection IVLT
Invertible Linear Transformations
One preliminary denition, and then we will have our main denition for this section.
Denition IDLT
Identity Linear Transformation
Theidentity linear transformation on the vector space Wis dened as
IW:W!W; IW(w) =w
4
Informally, IWis the \do-nothing" function. You should check that IWis really a linear transformation,
as claimed, and then compute its kernel and range to see that it is both injective and surjective. All of
these facts should be straightforward to verify (Exercise IVLT.T05 [594]). With this in hand we can make
our main denition.
Denition IVLT
Invertible Linear Transformations
Suppose that T:U!Vis a linear transformation. If there is a function S:V!Usuch that
ST=IU TS=IV
thenTisinvertible . In this case, we call Stheinverse ofTand writeS=T 1. 4
Informally, a linear transformation Tis invertible if there is a companion linear transformation, S, which
\undoes" the action of T. When the two linear transformations are applied consecutively (composition),
in either order, the result is to have no real eect. It is entirely analogous to squaring a positive number
and then taking its (positive) square root.
Here is an example of a linear transformation that is invertible. As usual at the beginning of a section,
do not be concerned with where Scame from, just understand how it illustrates Denition IVLT [579].
Example AIVLT
An invertible linear transformation
Archetype V [858] is the linear transformation
T:P3!M22; T
a+bx+cx2+dx3
=a+b a 2c
d b d
Dene the function S:M22!P3dened by
Sa b
c d
= (a c d) + (c+d)x+1
2(a b c d)x2+cx3
Version 2.30
582 Section IVLT Invertible Linear Transformations
Then
(TS)a b
c d
=T
Sa b
c d
=T
(a c d) + (c+d)x+1
2(a b c d)x2+cx3
=(a c d) + (c+d) (a c d) 2(1
2(a b c d))
c (c+d) c
=a b
c d
=IM22a b
c d
And
(ST)
a+bx+cx2+dx3
=S
T
a+bx+cx2+dx3
=Sa+b a 2c
d b d
= ((a+b) d (b d)) + (d+ (b d))x
+1
2((a+b) (a 2c) d (b d))
x2+ (d)x3
=a+bx+cx2+dx3
=IP3
a+bx+cx2+dx3
For now, understand why these computations show that Tis invertible, and that S=T 1. Maybe even be
amazed by how Sworks so perfectly in concert with T! We will see later just how to arrive at the correct
form ofS(when it is possible).
It can be as instructive to study a linear transformation that is not invertible.
Example ANILT
A non-invertible linear transformation
Consider the linear transformation T:C3!M22dened by
T0
@2
4a
b
c3
51
A=a b 2a+ 2b+c
3a+b+c 2a 6b 2c
Suppose we were to search for an inverse function S:M22!C3.
First verify that the 2 2 matrixA=5 3
8 2
is not in the range of T. This will amount to nding an
input toT,2
4a
b
c3
5, such that
a b= 5
2a+ 2b+c= 3
3a+b+c= 8
2a 6b 2c= 2
Version 2.30
Subsection IVLT.IVLT Invertible Linear Transformations 583
As this system of equations is inconsistent, there is no input column vector, and A62R(T). How should
we deneS(A)? Note that
T(S(A)) = (TS) (A) =IM22(A) =A
So any denition we would provide for S(A) must then be a column vector that Tsends toAand we
would have A2R(T), contrary to the denition of T. This is enough to see that there is no function S
that will allow us to conclude that Tis invertible, since we cannot provide a consistent denition for S(A)
if we assume Tis invertible.
Even though we now know that Tis not invertible, let's not leave this example just yet. Check that
T0
@2
41
2
43
51
A=3 2
5 2
=B T0
@2
40
3
83
51
A=3 2
5 2
=B
How would we dene S(B)?
S(B) =S0
@T0
@2
41
2
43
51
A1
A= (ST)0
@2
41
2
43
51
A=IC30
@2
41
2
43
51
A=2
41
2
43
5
or
S(B) =S0
@T0
@2
40
3
83
51
A1
A= (ST)0
@2
40
3
83
51
A=IC30
@2
40
3
83
51
A=2
40
3
83
5
Which denition should we provide for S(B)? Both are necessary. But then Sis not a function. So we
have a second reason to know that there is no function Sthat will allow us to conclude that Tis invertible.
It happens that there are innitely many column vectors that Swould have to take to B. Construct the
kernel ofT,
K(T) =*8
<
:2
4 1
1
43
59
=
;+
Now choose either of the two inputs used above for Tand add to it a scalar multiple of the basis vector
for the kernel of T. For example,
x=2
41
2
43
5+ ( 2)2
4 1
1
43
5=2
43
0
43
5
then verify that T(x) =B. Practice creating a few more inputs for Tthat would be sent to B, and see
why it is hopeless to think that we could ever provide a reasonable denition for S(B)! There is a \whole
subspace's worth" of values that S(B) would have to take on.
In Example ANILT [580] you may have noticed that Tis not surjective, since the matrix Awas not in
the range of T. AndTis not injective since there are two dierent input column vectors that Tsends to
the matrix B. Linear transformations Tthat are not surjective lead to putative inverse functions Sthat
are undened on inputs outside of the range of T. Linear transformations Tthat are not injective lead
to putative inverse functions Sthat are multiply-dened on each of their inputs. We will formalize these
ideas in Theorem ILTIS [582].
But rst notice in Denition IVLT [579] that we only require the inverse (when it exists) to be a
function. When it does exist, it too is a linear transformation.
Version 2.30
584 Section IVLT Invertible Linear Transformations
Theorem ILTLT
Inverse of a Linear Transformation is a Linear Transformation
Suppose that T:U!Vis an invertible linear transformation. Then the function T 1:V!Uis a linear
transformation.
Proof We work through verifying Denition LT [515] for T 1, using the fact that Tis a linear trans-
formation to obtain the second equality in each half of the proof. To this end, suppose x;y2Vand
2C.
T 1(x+y) =T 1
T
T 1(x)
+T
T 1(y)
Denition IVLT [579]
=T 1
T
T 1(x) +T 1(y)
Denition LT [515]
=T 1(x) +T 1(y) Denition IVLT [579]
Now check the second dening property of a linear transformation for T 1,
T 1(x) =T 1
T
T 1(x)
Denition IVLT [579]
=T 1
T
T 1(x)
Denition LT [515]
=T 1(x) Denition IVLT [579]
SoT 1fullls the requirements of Denition LT [515] and is therefore a linear transformation. So when
Thas an inverse, T 1is also a linear transformation. Additionally, T 1is invertible and itsinverse is what
you might expect.
Theorem IILT
Inverse of an Invertible Linear Transformation
Suppose that T:U!Vis an invertible linear transformation. Then T 1is an invertible linear transfor-
mation and
T 1 1=T.
Proof BecauseTis invertible, Denition IVLT [579] tells us there is a function T 1:V!Usuch that
T 1T=IU TT 1=IV
Additionally, Theorem ILTLT [582] tells us that T 1is more than just a function, it is a linear trans-
formation. Now view these two statements as properties of the linear transformation T 1. In light of
Denition IVLT [579], they together say that T 1is invertible (let Tplay the role of Sin the statement
of the denition). Furthermore, the inverse of T 1is thenT, i.e.
T 1 1=T.
Subsection IV
Invertibility
We now know what an inverse linear transformation is, but just which linear transformations have inverses?
Here is a theorem we have been preparing for all chapter long.
Theorem ILTIS
Invertible Linear Transformations are Injective and Surjective
SupposeT:U!Vis a linear transformation. Then Tis invertible if and only if Tis injective and
surjective.
Proof ()) SinceTis presumed invertible, we can employ its inverse, T 1(Denition IVLT [579]). To
see thatTis injective, suppose x;y2Uand assume that T(x) =T(y),
x=IU(x) Denition IDLT [579]
Version 2.30
Subsection IVLT.IV Invertibility 585
=
T 1T
(x) Denition IVLT [579]
=T 1(T(x)) Denition LTC [532]
=T 1(T(y)) Denition ILT [541]
=
T 1T
(y) Denition LTC [532]
=IU(y) Denition IVLT [579]
=y Denition IDLT [579]
So by Denition ILT [541] Tis injective. To check that Tis surjective, suppose v2V. ThenT 1(v) is
a vector in U. Compute
T
T 1(v)
=
TT 1
(v) Denition LTC [532]
=IV(v) Denition IVLT [579]
=v Denition IDLT [579]
So there is an element from U, when used as an input to T(namelyT 1(v)) that produces the desired
output, v, and hence Tis surjective by Denition SLT [559].
(() Now assume that Tis both injective and surjective. We will build a function S:V!Uthat
will establish that Tis invertible. To this end, choose any v2V. SinceTis surjective, Theorem RSLT
[565] saysR(T) =V, so we have v2R(T). Theorem RPI [568] says that the pre-image of v,T 1(v),
is nonempty. So we can choose a vector from the pre-image of v, say u. In other words, there exists
u2T 1(v).
SinceT 1(v) is non-empty, Theorem KPI [547] then says that
T 1(v) =fu+zjz2K(T)g
However, because Tis injective, by Theorem KILT [548] the kernel is trivial, K(T) =f0g. So the pre-image
is a set with just one element, T 1(v) =fug. Now we can dene SbyS(v) =u. This is the key to
this half of this proof. Normally the preimage of a vector from the codomain might be an empty set, or
an innite set. But surjectivity requires that the preimage not be empty, and then injectivity limits the
preimage to a singleton. Since our choice of vwas arbitrary, we know that every pre-image for Tis a set
with a single element. This allows us to construct Sas a function . Now that it is dened, verifying that
it is the inverse of Twill be easy. Here we go.
Choose u2U. Dene v=T(u). ThenT 1(v) =fug, so thatS(v) =uand,
(ST) (u) =S(T(u)) =S(v) =u=IU(u)
and since our choice of uwas arbitrary we have function equality, ST=IU.
Now choose v2V. Dene uto be the single vector in the set T 1(v), in other words, u=S(v).
ThenT(u) =v, so
(TS) (v) =T(S(v)) =T(u) =v=IV(v)
and since our choice of vwas arbitrary we have function equality, TS=IV.
When a linear transformation is both injective and surjective, the pre-image of any element of the
codomain is a set of size one (a \singleton"). This fact allowed us to construct the inverse linear trans-
formation in one half of the proof of Theorem ILTIS [582] (see Technique C [768]). We can follow this
approach to construct the inverse of a specic linear transformation, as the next example shows.
Example CIVLT
Computing the Inverse of a Linear Transformations
Version 2.30
586 Section IVLT Invertible Linear Transformations
Consider the linear transformation T:S22!P2dened by
Ta b
b c
= (a+b+c) + ( a+ 2c)x+ (2a+ 3b+ 6c)x2
Tis invertible, which you are able to verify, perhaps by determining that the kernel of Tis empty and the
range ofTis all ofP2. This will be easier once we have Theorem RPNDD [588], which appears later in
this section.
By Theorem ILTIS [582] we know T 1exists, and it will be critical shortly to realize that T 1is
automatically known to be a linear transformation as well (Theorem ILTLT [582]). To determine the
complete behavior of T 1:P2!S22we can simply determine its action on a basis for the domain, P2.
This is the substance of Theorem LTDB [525], and an excellent example of its application. Choose any
basis ofP2, the simpler the better, such as B=
1; x; x2
. Values of T 1for these three basis elements
will be the single elements of their preimages. In turn, we have
T 1(1) :
Ta b
b c
= 1 + 0x+ 0x2
2
41 1 1 1
1 0 2 0
2 3 6 03
5RREF !2
41 0 0 6
0 1 0 10
0 0 1 33
5
(preimage) T 1(1) = 6 10
10 3
(function) T 1(1) = 6 10
10 3
T 1(x) :
Ta b
b c
= 0 + 1x+ 0x2
2
41 1 1 0
1 0 2 1
2 3 6 03
5RREF !2
41 0 0 3
0 1 0 4
0 0 1 13
5
(preimage) T 1(x) = 3 4
4 1
(function) T 1(x) = 3 4
4 1
T 1
x2
:
Ta b
b c
= 0 + 0x+ 1x2
2
41 1 1 0
1 0 2 0
2 3 6 13
5RREF !2
41 0 0 2
0 1 0 3
0 0 1 13
5
(preimage) T 1
x2
=2 3
3 1
(function) T 1
x2
=2 3
3 1
Theorem LTDB [525] says, informally, \it is enough to know what a linear transformation does to a basis."
Formally, we have the outputs of T 1for a basis, so by Theorem LTDB [525] there is a unique linear
Version 2.30
Subsection IVLT.IV Invertibility 587
transformation with these outputs. So we put this information to work. The key step here is that we can
convert any element of P2into a linear combination of the elements of the basis B(Theorem VRRB [360]).
We are after a \formula" for the value of T 1on a generic element of P2, sayp+qx+rx2.
T 1
p+qx+rx2
=T 1
p(1) +q(x) +r(x2)
Theorem VRRB [360]
=pT 1(1) +qT 1(x) +rT 1
x2
Theorem LTLC [525]
=p 6 10
10 3
+q 3 4
4 1
+r2 3
3 1
= 6p 3q+ 2r10p+ 4q 3r
10p+ 4q 3r 3p q+r
Notice how a linear combination in the domain of T 1has been translated into a linear combination in
the codomain of T 1since we know T 1is a linear transformation by Theorem ILTLT [582].
Also, notice how the augmented matrices used to determine the three pre-images could be combined into
one calculation of a matrix in extended echelon form, reminiscent of a procedure we know for computing
the inverse of a matrix (see Example CMI [247]). Hmmmm.
We will make frequent use of the characterization of invertible linear transformations provided by
Theorem ILTIS [582]. The next theorem is a good example of this, and we will use it often, too.
Theorem CIVLT
Composition of Invertible Linear Transformations
Suppose that T:U!VandS:V!Ware invertible linear transformations. Then the composition,
(ST) :U!Wis an invertible linear transformation.
Proof SinceSandTare both linear transformations, STis also a linear transformation by Theorem
CLTLT [533]. Since SandTare both invertible, Theorem ILTIS [582] says that SandTare both injective
and surjective. Then Theorem CILTI [551] says STis injective, and Theorem CSLTS [570] says STis
surjective. Now apply the \other half" of Theorem ILTIS [582] and conclude that STis invertible.
When a composition is invertible, the inverse is easy to construct.
Theorem ICLT
Inverse of a Composition of Linear Transformations
Suppose that T:U!VandS:V!Ware invertible linear transformations. Then STis invertible
and (ST) 1=T 1S 1.
Proof Compute, for all w2W
(ST)
T 1S 1
(w) =S
T
T 1
S 1(w)
=S
IV
S 1(w)
Denition IVLT [579]
=S
S 1(w)
Denition IDLT [579]
=w Denition IVLT [579]
=IW(w) Denition IDLT [579]
so (ST)
T 1S 1
=IWand also
T 1S 1
(ST)
(u) =T 1
S 1(S(T(u)))
=T 1(IV(T(u))) Denition IVLT [579]
=T 1(T(u)) Denition IDLT [579]
=u Denition IVLT [579]
=IU(u) Denition IDLT [579]
Version 2.30
588 Section IVLT Invertible Linear Transformations
so
T 1S 1
(ST) =IU. By Denition IVLT [579], STis invertible and ( ST) 1=T 1S 1.
Notice that this theorem not only establishes what the inverse of STis, it also duplicates the conclusion
of Theorem CIVLT [585] and also establishes the invertibility of ST. But somehow, the proof of Theorem
CIVLT [585] is nicer way to get this property.
Does Theorem ICLT [585] remind you of the
avor of any theorem we have seen about matrices? (Hint:
Think about getting dressed.) Hmmmm.
Subsection SI
Structure and Isomorphism
A vector space is dened (Denition VS [317]) as a set of objects (\vectors") endowed with a denition
of vector addition (+) and a denition of scalar multiplication (written with juxtaposition). Many of our
denitions about vector spaces involve linear combinations (Denition LC [338]), such as the span of a set
(Denition SS [339]) and linear independence (Denition LI [351]). Other denitions are built up from
these ideas, such as bases (Denition B [371]) and dimension (Denition D [391]). The dening properties
of a linear transformation require that a function \respect" the operations of the two vector spaces that
are the domain and the codomain (Denition LT [515]). Finally, an invertible linear transformation is one
that can be \undone" | it has a companion that reverses its eect. In this subsection we are going to
begin to roll all these ideas into one.
A vector space has \structure" derived from denitions of the two operations and the requirement
that these operations interact in ways that satisfy the ten properties of Denition VS [317]. When two
dierent vector spaces have an invertible linear transformation dened between them, then we can translate
questions about linear combinations (spans, linear independence, bases, dimension) from the rst vector
space to the second. The answers obtained in the second vector space can then be translated back, via
the inverse linear transformation, and interpreted in the setting of the rst vector space. We say that
these invertible linear transformations \preserve structure." And we say that the two vector spaces are
\structurally the same." The precise term is \isomorphic," from Greek meaning \of the same form." Let's
begin to try to understand this important concept.
Denition IVS
Isomorphic Vector Spaces
Two vector spaces UandVareisomorphic if there exists an invertible linear transformation Twith
domainUand codomain V,T:U!V. In this case, we write U=V, and the linear transformation Tis
known as an isomorphism betweenUandV. 4
A few comments on this denition. First, be careful with your language (Technique L [766]). Two
vector spaces are isomorphic, or not. It is a yes/no situation and the term only applies to a pair of vector
spaces. Any invertible linear transformation can be called an isomorphism, it is a term that applies to
functions. Second, a given pair of vector spaces there might be several dierent isomorphisms between the
two vector spaces. But it only takes the existence of one to call the pair isomorphic. Third, Uisomorphic
toV, orVisomorphic to U? Doesn't matter, since the inverse linear transformation will provide the
needed isomorphism in the \opposite" direction. Being \isomorphic to" is an equivalence relation on the
set of all vector spaces (see Theorem SER [494] for a reminder about equivalence relations).
Example IVSAV
Isomorphic vector spaces, Archetype V
Archetype V [858] is a linear transformation from P3toM22,
T:P3!M22; T
a+bx+cx2+dx3
=a+b a 2c
d b d
Version 2.30
Subsection IVLT.SI Structure and Isomorphism 589
Since it is injective and surjective, Theorem ILTIS [582] tells us that it is an invertible linear transformation.
By Denition IVS [586] we say P3andM22are isomorphic.
At a basic level, the term \isomorphic" is nothing more than a codeword for the presence of an invertible
linear transformation. However, it is also a description of a powerful idea, and this power only becomes
apparent in the course of studying examples and related theorems. In this example, we are led to believe
that there is nothing \structurally" dierent about P3andM22. In a certain sense they are the same. Not
equal, but the same. One is as good as the other. One is just as interesting as the other.
Here is an extremely basic application of this idea. Suppose we want to compute the following linear
combination of polynomials in P3,
5(2 + 3x 4x2+ 5x3) + ( 3)(3 5x+ 3x2+x3)
Rather than doing it straight-away (which is very easy), we will apply the transformation Tto convert
into a linear combination of matrices, and then compute in M22according to the denitions of the vector
space operations there (Example VSM [319]),
T
5(2 + 3x 4x2+ 5x3) + ( 3)(3 5x+ 3x2+x3)
= 5T
2 + 3x 4x2+ 5x3
+ ( 3)T
3 5x+ 3x2+x3
Theorem LTLC [525]
= 55 10
5 2
+ ( 3) 2 3
1 6
Denition of T
=31 59
22 8
Operations in M22
Now we will translate our answer back to P3by applying T 1, which we found in Example AIVLT [579],
T 1:M22!P3; T 1a b
c d
= (a c d) + (c+d)x+1
2(a b c d)x2+cx3
We compute,
T 131 59
22 8
= 1 + 30x 29x2+ 22x3
which is, as expected, exactly what we would have computed for the original linear combination had we
just used the denitions of the operations in P3(Example VSP [319]). Notice this is meant only as an
illustration and not a suggested route for doing this particular computation.
Checking the dimensions of two vector spaces can be a quick way to establish that they are not
isomorphic. Here's the theorem.
Theorem IVSED
Isomorphic Vector Spaces have Equal Dimension
SupposeUandVare isomorphic vector spaces. Then dim ( U) = dim (V).
Proof IfUandVare isomorphic, there is an invertible linear transformation T:U!V(Denition
IVS [586]). Tis injective by Theorem ILTIS [582] and so by Theorem ILTD [550], dim ( U)dim (V).
Similarly,Tis surjective by Theorem ILTIS [582] and so by Theorem SLTD [569], dim ( U)dim (V). The
net eect of these two inequalities is that dim ( U) = dim (V).
The contrapositive of Theorem IVSED [587] says that if UandVhave dierent dimensions, then they
are not isomorphic. Dimension is the simplest \structural" characteristic that will allow you to distinguish
non-isomorphic vector spaces. For example P6is not isomorphic to M34since their dimensions (7 and 12,
respectively) are not equal. With tools developed in Section VR [603] we will be able to establish that the
converse of Theorem IVSED [587] is true. Think about that one for a moment.
Version 2.30
590 Section IVLT Invertible Linear Transformations
Subsection RNLT
Rank and Nullity of a Linear Transformation
Just as a matrix has a rank and a nullity, so too do linear transformations. And just like the rank and
nullity of a matrix are related (they sum to the number of columns, Theorem RPNC [398]) the rank and
nullity of a linear transformation are related. Here are the denitions and theorems, see the Archetypes
(Appendix A [777]) for loads of examples.
Denition ROLT
Rank Of a Linear Transformation
Suppose that T:U!Vis a linear transformation. Then the rank ofT,r(T), is the dimension of the
range ofT,
r(T) = dim (R(T))
(This denition contains Notation ROLT.) 4
Denition NOLT
Nullity Of a Linear Transformation
Suppose that T:U!Vis a linear transformation. Then the nullity ofT,n(T), is the dimension of the
kernel ofT,
n(T) = dim (K(T))
(This denition contains Notation NOLT.) 4
Here are two quick theorems.
Theorem ROSLT
Rank Of a Surjective Linear Transformation
Suppose that T:U!Vis a linear transformation. Then the rank of Tis the dimension of V,r(T) =
dim (V), if and only if Tis surjective.
Proof By Theorem RSLT [565], Tis surjective if and only if R(T) =V. Applying Denition ROLT
[588],R(T) =Vif and only if r(T) = dim (R(T)) = dim (V).
Theorem NOILT
Nullity Of an Injective Linear Transformation
Suppose that T:U!Vis a linear transformation. Then the nullity of Tis zero,n(T) = 0, if and only if
Tis injective.
Proof By Theorem KILT [548], Tis injective if and only if K(T) =f0g. Applying Denition NOLT
[588],K(T) =f0gif and only if n(T) = 0.
Just as injectivity and surjectivity come together in invertible linear transformations, there is a clear
relationship between rank and nullity of a linear transformation. If one is big, the other is small.
Theorem RPNDD
Rank Plus Nullity is Domain Dimension
Suppose that T:U!Vis a linear transformation. Then
r(T) +n(T) = dim (U)
Proof Letr=r(T) ands=n(T). Suppose that R=fv1;v2;v3; :::; vrgVis a basis of the range
ofT,R(T), andS=fu1;u2;u3; :::; usgUis a basis of the kernel of T,K(T). Note that RandSare
Version 2.30
Subsection IVLT.RNLT Rank and Nullity of a Linear Transformation 591
possibly empty, which means that some of the sums in this proof are \empty" and are equal to the zero
vector.
Because the elements of Rare all in the range of T, each must have a non-empty pre-image by Theorem
RPI [568]. Choose vectors wi2U, 1irsuch that wi2T 1(vi). SoT(wi) =vi, 1ir. Consider
the set
B=fu1;u2;u3; :::; us;w1;w2;w3; :::; wrg
We claim that Bis a basis for U.
To establish linear independence for B, begin with a relation of linear dependence on B. So suppose
there are scalars a1; a2; a3; :::; asandb1; b2; b3; :::; br
0=a1u1+a2u2+a3u3++asus+b1w1+b2w2+b3w3++brwr
Then
0=T(0) Theorem LTTZZ [519]
=T(a1u1+a2u2+a3u3++asus+
b1w1+b2w2+b3w3++brwr) Denition LI [351]
=a1T(u1) +a2T(u2) +a3T(u3) ++asT(us) +
b1T(w1) +b2T(w2) +b3T(w3) ++brT(wr) Theorem LTLC [525]
=a10+a20+a30++as0+
b1T(w1) +b2T(w2) +b3T(w3) ++brT(wr) Denition KLT [545]
=0+0+0++0+
b1T(w1) +b2T(w2) +b3T(w3) ++brT(wr) Theorem ZVSM [325]
=b1T(w1) +b2T(w2) +b3T(w3) ++brT(wr) Property Z [318]
=b1v1+b2v2+b3v3++brvr Denition PI [528]
This is a relation of linear dependence on R(Denition RLD [351]), and since Ris a linearly independent
set (Denition LI [351]), we see that b1=b2=b3=:::=br= 0. Then the original relation of linear
dependence on Bbecomes
0=a1u1+a2u2+a3u3++asus+ 0w1+ 0w2+:::+ 0wr
=a1u1+a2u2+a3u3++asus+0+0+:::+0 Theorem ZSSM [324]
=a1u1+a2u2+a3u3++asus Property Z [318]
But this is again a relation of linear independence (Denition RLD [351]), now on the set S. SinceSis
linearly independent (Denition LI [351]), we have a1=a2=a3=:::=ar= 0. Since we now know
that all the scalars in the relation of linear dependence on Bmust be zero, we have established the linear
independence of Sthrough Denition LI [351].
To now establish that BspansU, choose an arbitrary vector u2U. ThenT(u)2R(T), so there are
scalarsc1; c2; c3; :::; crsuch that
T(u) =c1v1+c2v2+c3v3++crvr
Use the scalars c1; c2; c3; :::; crto dene a vector y2U,
y=c1w1+c2w2+c3w3++crwr
Then
T(u y) =T(u) T(y) Theorem LTLC [525]
Version 2.30
592 Section IVLT Invertible Linear Transformations
=T(u) T(c1w1+c2w2+c3w3++crwr) Substitution
=T(u) (c1T(w1) +c2T(w2) ++crT(wr)) Theorem LTLC [525]
=T(u) (c1v1+c2v2+c3v3++crvr) wi2T 1(vi)
=T(u) T(u) Substitution
=0 Property AI [318]
So the vector u yis sent to the zero vector by Tand hence is an element of the kernel of T. As such it
can be written as a linear combination of the basis vectors for K(T), the elements of the set S. So there
are scalars d1; d2; d3; :::; dssuch that
u y=d1u1+d2u2+d3u3++dsus
Then
u= (u y) +y
=d1u1+d2u2+d3u3++dsus+c1w1+c2w2+c3w3++crwr
This says that for any vector, u, fromU, there exist scalars ( d1; d2; d3; :::; ds; c1; c2; c3; :::; cr) that form
uas a linear combination of the vectors in the set B. In other words, BspansU(Denition SS [339]).
SoBis a basis (Denition B [371]) of Uwiths+rvectors, and thus
dim (U) =s+r=n(T) +r(T)
as desired.
Theorem RPNC [398] said that the rank and nullity of a matrix sum to the number of columns of the
matrix. This result is now an easy consequence of Theorem RPNDD [588] when we consider the linear
transformation T:Cn!Cmdened with the mnmatrixAbyT(x) =Ax. The range and kernel
ofTare identical to the column space and null space of the matrix A(Exercise ILT.T20 [554], Exercise
SLT.T20 [573]), so the rank and nullity of the matrix Aare identical to the rank and nullity of the linear
transformation T. The dimension of the domain of Tis the dimension of Cn, exactly the number of columns
for the matrix A.
This theorem can be especially useful in determining basic properties of linear transformations. For
example, suppose that T:C6!C6is a linear transformation and you are able to quickly establish that
the kernel is trivial. Then n(T) = 0. First this means that Tis injective by Theorem NOILT [588]. Also,
Theorem RPNDD [588] becomes
6 = dim
C6
=r(T) +n(T) =r(T) + 0 =r(T)
So the rank of Tis equal to the rank of the codomain, and by Theorem ROSLT [588] we know Tis
surjective. Finally, we know Tis invertible by Theorem ILTIS [582]. So from the determination that the
kernel is trivial, and consideration of various dimensions, the theorems of this section allow us to conclude
the existence of an inverse linear transformation for T.
Similarly, Theorem RPNDD [588] can be used to provide alternative proofs for Theorem ILTD [550],
Theorem SLTD [569] and Theorem IVSED [587]. It would be an interesting exercise to construct these
proofs.
It would be instructive to study the archetypes that are linear transformations and see how many of
their properties can be deduced just from considering only the dimensions of the domain and codomain.
Then add in just knowledge of either the nullity or rank, and so how much more you can learn about the
linear transformation. The table preceding all of the archetypes (Appendix A [777]) could be a good place
to start this analysis.
Version 2.30
Subsection IVLT.SLELT Systems of Linear Equations and Linear Transformations 593
Subsection SLELT
Systems of Linear Equations and Linear Transformations
This subsection does not really belong in this section, or any other section, for that matter. It is just the
right time to have a discussion about the connections between the central topic of linear algebra, linear
transformations, and our motivating topic from Chapter SLE [3], systems of linear equations. We will
discuss several theorems we have seen already, but we will also make some forward-looking statements that
will be justied in Chapter R [603].
Archetype D [795] and Archetype E [799] are ideal examples to illustrate connections with linear
transformations. Both have the same coecient matrix,
D=2
42 1 7 7
3 4 5 6
1 1 4 53
5
To apply the theory of linear transformations to these two archetypes, employ matrix multiplication (Def-
inition MM [226]) and dene the linear transformation,
T:C4!C3; T (x) =Dx=x12
42
3
13
5+x22
41
4
13
5+x32
47
5
43
5+x42
4 7
6
53
5
Theorem MBLT [522] tells us that Tis indeed a linear transformation. Archetype D [795] asks for solutions
toLS(D;b), where b=2
48
12
43
5. In the language of linear transformations this is equivalent to asking for
T 1(b). In the language of vectors and matrices it asks for a linear combination of the four columns of D
that will equal b. One solution listed is w=2
6647
8
1
33
775. With a non-empty preimage, Theorem KPI [547] tells
us that the complete solution set of the linear system is the preimage of b,
w+K(T) =fw+zjz2K(T)g
The kernel of the linear transformation Tis exactly the null space of the matrix D(see Exercise ILT.T20
[554]), so this approach to the solution set should be reminiscent of Theorem PSPHS [124]. The kernel
of the linear transformation is the preimage of the zero vector, exactly equal to the solution set of the
homogeneous system LS(D;0). SinceDhas a null space of dimension two, every preimage (and in
particular the preimage of b) is as \big" as a subspace of dimension two (but is not a subspace).
Archetype E [799] is identical to Archetype D [795] but with a dierent vector of constants, d=2
42
3
23
5.
We can use the same linear transformation Tto discuss this system of equations since the coecient matrix
is identical. Now the set of solutions to LS(D;d) is the pre-image of d,T 1(d). However, the vector d
is not in the range of the linear transformation (nor is it in the column space of the matrix, since these
two sets are equal by Exercise SLT.T20 [573]). So the empty pre-image is equivalent to the inconsistency
of the linear system.
These two archetypes each have three equations in four variables, so either the resulting linear systems
are inconsistent, or they are consistent and application of Theorem CMVEI [61] tells us that the system has
Version 2.30
594 Section IVLT Invertible Linear Transformations
innitely many solutions. Considering these same parameters for the linear transformation, the dimension
of the domain, C4, is four, while the codomain, C3, has dimension three. Then
n(T) = dim
C4
r(T) Theorem RPNDD [588]
= 4 dim (R(T)) Denition ROLT [588]
4 3 R(T) subspace of C3
= 1
So the kernel of Tis nontrivial simply by considering the dimensions of the domain (number of variables)
and the codomain (number of equations). Pre-images of elements of the codomain that are not in the range
ofTare empty (inconsistent systems). For elements of the codomain that are in the range of T(consistent
systems), Theorem KPI [547] tells us that the pre-images are built from the kernel, and with a non-trivial
kernel, these pre-images are innite (innitely many solutions).
When do systems of equations have unique solutions? Consider the system of linear equations LS(C;f)
and the linear transformation S(x) =Cx. IfShas a trivial kernel, then pre-images will either be empty or
be nite sets with single elements. Correspondingly, the coecient matrix Cwill have a trivial null space
and solution sets will either be empty (inconsistent) or contain a single solution (unique solution). Should
the matrix be square and have a trivial null space then we recognize the matrix as being nonsingular.
A square matrix means that the corresponding linear transformation, T, has equal-sized domain and
codomain. With a nullity of zero, Tis injective, and also Theorem RPNDD [588] tells us that rank of Tis
equal to the dimension of the domain, which in turn is equal to the dimension of the codomain. In other
words,Tis surjective. Injective and surjective, and Theorem ILTIS [582] tells us that Tis invertible. Just
as we can use the inverse of the coecient matrix to nd the unique solution of any linear system with a
nonsingular coecient matrix (Theorem SNCM [261]), we can use the inverse of the linear transformation
to construct the unique element of any pre-image (proof of Theorem ILTIS [582]).
The executive summary of this discussion is that to every coecient matrix of a system of linear equa-
tions we can associate a natural linear transformation. Solution sets for systems with this coecient matrix
are preimages of elements of the codomain of the linear transformation. For every theorem about systems
of linear equations there is an analogue about linear transformations. The theory of linear transformations
provides all the tools to recreate the theory of solutions to linear systems of equations.
We will continue this adventure in Chapter R [603].
Subsection READ
Reading Questions
1. What conditions allow us to easily determine if a linear transformation is invertible?
2. What does it mean to say two vector spaces are isomorphic? Both technically, and informally?
3. How do linear transformations relate to systems of linear equations?
Version 2.30
Subsection IVLT.EXC Exercises 595
Subsection EXC
Exercises
C10 The archetypes below are linear transformations of the form T:U!Vthat are invertible. For
each, the inverse linear transformation is given explicitly as part of the archetype's description. Verify for
each linear transformation that
T 1T=IU TT 1=IV
Archetype R [848],
Archetype V [858],
Archetype W [860]
Contributed by Robert Beezer
C20 Determine if the linear transformation T:P2!M22is (a) injective, (b) surjective, (c) invertible.
T
a+bx+cx2
=a+ 2b 2c 2a+ 2b
a+b 4c3a+ 2b+ 2c
Contributed by Robert Beezer Solution [596]
C21 Determine if the linear transformation S:P3!M22is (a) injective, (b) surjective, (c) invertible.
S
a+bx+cx2+dx3
= a+ 4b+c+ 2d4a b+ 6c d
a+ 5b 2c+ 2d a + 2c+ 5d
Contributed by Robert Beezer Solution [596]
C25 For each linear transformation below: (a) Find the matrix representation of T, (b) Calculate n(T),
(c) Calculate r(T), (d) Graph the image in either R2orR3as appropriate, (e) How many dimensions are
lost?, and (f) How many dimensions are preserved?
1.T:C3!C3given byT0
@2
4x
y
z3
51
A=2
4x
x
x3
5
2.T:C3!C3given byT0
@2
4x
y
z3
51
A=2
4x
y
03
5
3.T:C3!C2given byT0
@2
4x
y
z3
51
A=x
x
4.T:C3!C2given byT0
@2
4x
y
z3
51
A=x
y
5.T:C2!C3given byTx
y
=2
4x
y
03
5
Version 2.30
596 Section IVLT Invertible Linear Transformations
6.T:C2!C3given byTx
y
=2
4x
y
x+y3
5
Contributed by Chris Black
C50 Consider the linear transformation S:M12!P1from the set of 1 2 matrices to the set of
polynomials of degree at most 1, dened by
S
a b
= (3a+b) + (5a+ 2b)x
Prove that Sis invertible. Then show that the linear transformation
R:P1!M12; R (r+sx) =
(2r s) ( 5r+ 3s)
is the inverse of S, that isS 1=R.
Contributed by Robert Beezer Solution [597]
M30 The linear transformation Sbelow is invertible. Find a formula for the inverse linear transformation,
S 1.
S:P1!M1;2; S (a+bx) =
3a+b2a+b
Contributed by Robert Beezer Solution [597]
M31 The linear transformation R:M12!M21is invertible. Determine a formula for the inverse linear
transformation R 1:M21!M12.
R
a b
=a+ 3b
4a+ 11b
Contributed by Robert Beezer Solution [598]
M50 Rework Example CIVLT [583], only in place of the basis BforP2, choose instead to use the basis
C=
1;1 +x;1 +x+x2
. This will complicate writing a generic element of the domain of T 1as a linear
combination of the basis elements, and the algebra will be a bit messier, but in the end you should obtain
the same formula for T 1. The inverse linear transformation is what it is, and the choice of a particular
basis should not in
uence the outcome.
Contributed by Robert Beezer
M60 SupposeUandVare vector spaces. Dene the function Z:U!VbyT(u) =0Vfor every u2U.
Then by Exercise LT.M60 [536], Zis a linear transformation. Formulate a condition on UandVthat is
equivalent to Zbeing an invertible linear transformation. In other words, ll in the blank to complete the
following statement (and then give a proof): Zis invertible if and only if UandVare .
(See Exercise ILT.M60 [553], Exercise SLT.M60 [572], Exercise MR.M60 [637].)
Contributed by Robert Beezer
T05 Prove that the identity linear transformation (Denition IDLT [579]) is both injective and surjective,
and hence invertible.
Contributed by Robert Beezer
T15 Suppose that T:U!Vis a surjective linear transformation and dim ( U) = dim (V). Prove that T
is injective.
Contributed by Robert Beezer Solution [598]
Version 2.30
Subsection IVLT.EXC Exercises 597
T16 Suppose that T:U!Vis an injective linear transformation and dim ( U) = dim (V). Prove that T
is surjective.
Contributed by Robert Beezer
T30 Suppose that UandVare isomorphic vector spaces. Prove that there are innitely many isomor-
phisms between UandV.
Contributed by Robert Beezer Solution [599]
T40 SupposeT:U!VandS:V!Ware linear transformations and dim ( U) = dim (V) = dim (W).
Suppose that STis invertible. Prove that SandTare individually invertible (this could be construed
as a converse of Theorem CIVLT [585]).
Contributed by Robert Beezer Solution [599]
Version 2.30
598 Section IVLT Invertible Linear Transformations
Subsection SOL
Solutions
C20 Contributed by Robert Beezer Statement [593]
(a) We will compute the kernel of T. Suppose that a+bx+cx22K(T). Then
0 0
0 0
=T
a+bx+cx2
=a+ 2b 2c 2a+ 2b
a+b 4c3a+ 2b+ 2c
and matrix equality (Theorem ME [485]) yields the homogeneous system of four equations in three variables,
a+ 2b 2c= 0
2a+ 2b= 0
a+b 4c= 0
3a+ 2b+ 2c= 0
The coecient matrix of this system row-reduces as
2
6641 2 2
2 2 0
1 1 4
3 2 23
775RREF !2
66410 2
01 2
0 0 0
0 0 03
775
From the existence of non-trivial solutions to this system, we can infer non-zero polynomials in K(T). By
Theorem KILT [548] we then know that Tis not injective.
(b) Since 3 = dim ( P2)<dim (M22) = 4, by Theorem SLTD [569] Tis not surjective.
(c) SinceTis not surjective, it is not invertible by Theorem ILTIS [582].
C21 Contributed by Robert Beezer Statement [593]
(a) To check injectivity, we compute the kernel of S. To this end, suppose that a+bx+cx2+dx32K(S),
so 0 0
0 0
=S
a+bx+cx2+dx3
= a+ 4b+c+ 2d4a b+ 6c d
a+ 5b 2c+ 2d a + 2c+ 5d
this creates the homogeneous system of four equations in four variables,
a+ 4b+c+ 2d= 0
4a b+ 6c d= 0
a+ 5b 2c+ 2d= 0
a+ 2c+ 5d= 0
The coecient matrix of this system row-reduces as,
2
664 1 4 1 2
4 1 6 1
1 5 2 2
1 0 2 53
775RREF !2
666410 0 0
010 0
0 0 10
0 0 0 13
7775
We recognize the coecient matrix as being nonsingular, so the only solution to the system is a=b=c=
d= 0, and the kernel of Sis trivial,K(S) =
0 + 0x+ 0x2+ 0x3
. By Theorem KILT [548], we see that
Sis injective.
Version 2.30
Subsection IVLT.SOL Solutions 599
(b) We can establish that Sis surjective by considering the rank and nullity of S.
r(S) = dim (P3) n(S) Theorem RPNDD [588]
= 4 0
= dim (M22)
So,R(S) is a subspace of M22(Theorem RLTS [564]) whose dimension equals that of M22. By Theorem
EDYES [410], we gain the set equality R(S) =M22. Theorem RSLT [565] then implies that Sis surjective.
(c) SinceSis both injective and surjective, Theorem ILTIS [582] says Sis invertible.
C50 Contributed by Robert Beezer Statement [594]
Determine the kernel of Srst. The condition that S
a b
=0becomes (3a+b) + (5a+ 2b)x= 0 + 0x.
Equating coecients of these polynomials yields the system
3a+b= 0
5a+ 2b= 0
This homogeneous system has a nonsingular coecient matrix, so the only solution is a= 0,b= 0 and
thus
K(S) =
0 0
By Theorem KILT [548], we know Sis injective. With n(S) = 0 we employ Theorem RPNDD [588] to
nd
r(S) =r(S) + 0 =r(S) +n(S) = dim (M12) = 2 = dim ( P1)
SinceR(S)P1and dim (R(S)) = dim (P1), we can apply Theorem EDYES [410] to obtain the set
equalityR(S) =P1and therefore Sis surjective.
One of the two dening conditions of an invertible linear transformation is (Denition IVLT [579])
(SR) (a+bx) =S(R(a+bx))
=S
(2a b) ( 5a+ 3b)
= (3(2a b) + ( 5a+ 3b)) + (5(2a b) + 2( 5a+ 3b))x
= ((6a 3b) + ( 5a+ 3b)) + ((10a 5b) + ( 10a+ 6b))x
=a+bx
=IP1(a+bx)
That (RS)
a b
=IM12
a b
is similar.
M30 Contributed by Robert Beezer Statement [594]
(Another approach to this solution would follow Example CIVLT [583].)
Suppose that S 1:M1;2!P1has a form given by
S 1
z w
= (rz+sw) + (pz+qw)x
wherer; s; p; q are unknown scalars. Then
a+bx=S 1(S(a+bx))
=S 1
3a+b2a+b
= (r(3a+b) +s(2a+b)) + (p(3a+b) +q(2a+b))x
= ((3r+ 2s)a+ (r+s)b) + ((3p+ 2q)a+ (p+q)b)x
Version 2.30
600 Section IVLT Invertible Linear Transformations
Equating coecients of these two polynomials, and then equating coecients on aandb, gives rise to 4
equations in 4 variables,
3r+ 2s= 1
r+s= 0
3p+ 2q= 0
p+q= 1
This system has a unique solution: r= 1,s= 1,p= 2,q= 3. So the desired inverse linear
transformation is
S 1
z w
= (z w) + ( 2z+ 3w)x
Notice that the system of 4 equations in 4 variables could be split into two systems, each with two equations
in two variables (and identical coecient matrices). After making this split, the solution might feel like
computing the inverse of a matrix (Theorem CINM [248]). Hmmmm.
M31 Contributed by Robert Beezer Statement [594]
(Another approach to this solution would follow Example CIVLT [583].)
We are given that Ris invertible. The inverse linear transformation can be formulated by considering
the pre-image of a generic element of the codomain. With injectivity and surjectivity, we know that the
pre-image of any element will be a set of size one | it is this lone element that will be the output of the
inverse linear transformation.
Suppose that we set v=x
y
as a generic element of the codomain, M21. Then if
r s
=w2R 1(v),
x
y
=v=R(w)
=r+ 3s
4r+ 11s
So we obtain the system of two equations in the two variables rands,
r+ 3s=x
4r+ 11s=y
With a nonsingular coecient matrix, we can solve the system using the inverse of the coecient matrix,
r= 11x+ 3y
s= 4x y
So we dene,
R 1(v) =R 1x
y
=w=
r s
=
11x+ 3y4x y
T15 Contributed by Robert Beezer Statement [594]
IfTis surjective, then Theorem RSLT [565] says R(T) =V, sor(T) = dim (V). In turn, the hypothesis
givesr(T) = dim (U). Then, using Theorem RPNDD [588],
n(T) = (r(T) +n(T)) r(T) = dim (U) dim (U) = 0
With a null space of zero dimension, K(T) =f0g, and by Theorem KILT [548] we see that Tis injective.
Tis both injective and surjective so by Theorem ILTIS [582], Tis invertible.
Version 2.30
Subsection IVLT.SOL Solutions 601
T30 Contributed by Robert Beezer Statement [595]
SinceUandVare isomorphic, there is at least one isomorphism between them (Denition IVS [586]), say
T:U!V. As such,Tis an invertible linear transformation.
For2Cdene the linear transformation S:V!VbyS(v) =v. Convince yourself that when 6=
0,Sis an invertible linear transformation (Denition IVLT [579]). Then the composition, ST:U!V,
is an invertible linear transformation by Theorem CIVLT [585]. Once convinced that each non-zero value
ofgives rise to a dierent functions for ST, then we have constructed innitely many isomorphisms
fromUtoV.
T40 Contributed by Robert Beezer Statement [595]
SinceSTis invertible, by Theorem ILTIS [582] STis injective and therefore has a trivial kernel by
Theorem KILT [548]. Then
K(T)K(ST) Exercise ILT.T15 [554]
=f0g Theorem KILT [548]
SinceThas a trivial kernel, by Theorem KILT [548], Tis injective. Also,
r(T) = dim (U) n(T) Theorem RPNDD [588]
= dim (U) 0 Theorem NOILT [588]
= dim (V) Hypothesis
SinceR(T)V, Theorem EDYES [410] gives R(T) =V, so by Theorem RSLT [565], Tis surjective.
Finally, by Theorem ILTIS [582], Tis invertible.
SinceSTis invertible, by Theorem ILTIS [582] STis surjective and therefore has a full range by
Theorem RSLT [565]. Then
W=R(ST) Theorem RSLT [565]
R(S) Exercise SLT.T15 [572]
SinceR(S)Wwe haveR(S) =Wand by Theorem RSLT [565], Sis surjective. By an application
of Theorem RPNDD [588] similar to the rst part of this solution, we see that Shas a trivial kernel, is
therefore injective (Theorem KILT [548]), and thus invertible (Theorem ILTIS [582]).
Version 2.30
602 Section IVLT Invertible Linear Transformations
Version 2.30
Annotated Acronyms IVLT.LT Linear Transformations 603
Annotated Acronyms LT
Linear Transformations
Theorem MBLT [522]
You give me an mnmatrix and I'll give you a linear transformation T:Cn!Cm. This is our rst hint
that there is some relationship between linear transformations and matrices.
Theorem MLTCV [523]
You give me a linear transformation T:Cn!Cmand I'll give you an mnmatrix. This is our second hint
that there is some relationship between linear transformations and matrices. Generalizing this relationship
to arbitrary vector spaces (i.e. not just CnandCm) will be the most important idea of Chapter R [603].
Theorem LTLC [525]
A simple idea, and as described in Exercise LT.T20 [536], equivalent to the Denition LT [515]. The
statement is really just for convenience, as we'll quote this one often.
Theorem LTDB [525]
Another simple idea, but a powerful one. \It is enough to know what a linear transformation does to
a basis." At the outset of Chapter R [603], Theorem VRRB [360] will help us dene a very important
function, and then Theorem LTDB [525] will allow us to understand that this function is also a linear
transformation.
Theorem KPI [547]
The pre-image will be an important construction in this chapter, and this is one of the most important
descriptions of the pre-image. It should remind you of Theorem PSPHS [124], which is described in
Acronyms V [205]. See Theorem RPI [568], which is also described below.
Theorem KILT [548]
Kernels and injective linear transformations are intimately related. This result is the connection. Compare
with Theorem RSLT [565] below.
Theorem ILTB [550]
Injective linear transformations and linear independence are intimately related. This result is the connec-
tion. Compare with Theorem SLTB [568] below.
Theorem RSLT [565]
Ranges and surjective linear transformations are intimately related. This result is the connection. Compare
with Theorem KILT [548] above.
Theorem SSRLT [567]
This theorem provides the most direct way of forming the range of a linear transformation. The resulting
spanning set might well be linearly dependent, and beg for some clean-up, but that doesn't stop us from
having very quickly formed a reasonable description of the range. If you nd the determination of spanning
sets or ranges dicult, this is one worth remembering. You can view this as the analogue of forming a
column space by a direct application of Denition CSM [271].
Theorem SLTB [568]
Version 2.30
604 Section IVLT Invertible Linear Transformations
Surjective linear transformations and spanning sets are intimately related. This result is the connection.
Compare with Theorem ILTB [550] above.
Theorem RPI [568]
This is the analogue of Theorem KPI [547]. Membership in the range is equivalent to nonempty pre-images.
Theorem ILTIS [582]
Injectivity and surjectivity are independent concepts. You can have one without the other. But when you
have both, you get invertibility, a linear transformation that can be run \backwards." This result might
explain the entire structure of the four sections in this chapter.
Theorem RPNDD [588]
This is the promised generalization of Theorem RPNC [398] about matrices. So the number of columns of
a matrix is the analogue of the dimension of the domain. This will become even more precise in Chapter
R [603]. For now, this can be a powerful result for determining dimensions of kernels and ranges, and
consequently, the injectivity or surjectivity of linear transformations. Never underestimate a theorem that
counts something.
Version 2.30
Chapter R
Representations
Previous work with linear transformations may have convinced you that we can convert most questions
about linear transformations into questions about systems of equations or properties of subspaces of Cm.
In this section we begin to make these vague notions precise. We have used the word \representation"
prior, but it will get a heavy workout in this chapter. In many ways, everything we have studied so far
was in preparation for this chapter.
Section VR
Vector Representations
We begin by establishing an invertible linear transformation between any vector space Vof dimension m
andCm. This will allow us to \go back and forth" between the two vector spaces, no matter how abstract
the denition of Vmight be.
Denition VR
Vector Representation
Suppose that Vis a vector space with a basis B=fv1;v2;v3; :::; vng. Dene a function B:V!Cn
as follows. For w2Vdene the column vector B(w)2Cnby
w= [B(w)]1v1+ [B(w)]2v2+ [B(w)]3v3++ [B(w)]nvn
(This denition contains Notation VR.) 4
This denition looks more complicated that it really is, though the form above will be useful in proofs.
Simply stated, given w2V, we write was a linear combination of the basis elements of B. It is key
to realize that Theorem VRRB [360] guarantees that we can do this for every w, and furthermore this
expression as a linear combination is unique. The resulting scalars are just the entries of the vector B(w).
This discussion should convince you that Bis \well-dened" as a function. We can determine a precise
output for any input. Now we want to establish that Bis a function with additional properties - it is a
linear transformation.
Theorem VRLT
Vector Representation is a Linear Transformation
The function B(Denition VR [603]) is a linear transformation.
Proof We will take a novel approach in this proof. We will construct another function, which we will
easily determine is a linear transformation, and then show that this second function is really Bin disguise.
Here we go.
605
606 Section VR Vector Representations
SinceBis a basis, we can dene T:V!Cnto be the unique linear transformation such that T(vi) =ei,
1in, as guaranteed by Theorem LTDB [525], and where the eiare the standard unit vectors
(Denition SUV [197]). Then suppose for an arbitrary w2Vwe have,
[T(w)]i=2
4T0
@nX
j=1[B(w)]jvj1
A3
5
iDenition VR [603]
=2
4nX
j=1[B(w)]jT(vj)3
5
iTheorem LTLC [525]
=2
4nX
j=1[B(w)]jej3
5
i
=nX
j=1h
[B(w)]jeji
iDenition CVA [98]
=nX
j=1[B(w)]j[ej]iDenition CVSM [99]
= [B(w)]i[ei]i+nX
j=1
j6=i[B(w)]j[ej]iProperty CC [100]
= [B(w)]i(1) +nX
j=1
j6=i[B(w)]j(0) Denition SUV [197]
= [B(w)]i
As column vectors, Denition CVE [98] implies that T(w) =B(w). Since wwas an arbitrary element
ofV, as functions T=B. Now, since Tis known to be a linear transformation, it must follow that Bis
also a linear transformation.
The proof of Theorem VRLT [603] provides an alternate denition of vector representation relative to
a basisBthat we could state as a corollary (Technique LC [774]): Bis the unique linear transformation
that takesBto the standard unit basis.
Example VRC4
Vector representation in C4
Consider the vector y2C4
y=2
6646
14
6
73
775
We will nd several vector representations of yin this example. Notice that ynever changes, but the
representations ofydo change.
One basis for C4is
B=fu1;u2;u3;u4g=8
>><
>>:2
664 2
1
2
33
775;2
6643
6
2
43
775;2
6641
2
0
53
775;2
6644
3
1
63
7759
>>=
>>;
Version 2.30
Section VR Vector Representations 607
as can be seen by making these vectors the columns of a matrix, checking that the matrix is nonsingular
and applying Theorem CNMB [376]. To nd B(y), we need to nd scalars, a1; a2; a3; a4such that
y=a1u1+a2u2+a3u3+a4u4
By Theorem SLSLC [112] the desired scalars are a solution to the linear system of equations with a
coecient matrix whose columns are the vectors in Band with a vector of constants y. With a nonsingular
coecient matrix, the solution is unique, but this is no surprise as this is the content of Theorem VRRB
[360]. This unique solution is
a1= 2 a2= 1 a3= 3 a4= 4
Then by Denition VR [603], we have
B(y) =2
6642
1
3
43
775
Suppose now that we construct a representation of yrelative to another basis of C4,
C=8
>><
>>:2
664 15
9
4
23
775;2
66416
14
5
23
775;2
664 26
14
6
33
775;2
66414
13
4
63
7759
>>=
>>;
As withB, it is easy to check that Cis a basis. Writing yas a linear combination of the vectors in Cleads
to solving a system of four equations in the four unknown scalars with a nonsingular coecient matrix.
The unique solution can be expressed as
y=2
6646
14
6
73
775= ( 28)2
664 15
9
4
23
775+ ( 8)2
66416
14
5
23
775+ 112
664 26
14
6
33
775+ 02
66414
13
4
63
775
so that Denition VR [603] gives
C(y) =2
664 28
8
11
03
775
We often perform representations relative to standard bases, but for vectors in Cmits a little silly. Let's
nd the vector representation of yrelative to the standard basis (Theorem SUVB [371]),
D=fe1;e2;e3;e4g
Then, without any computation, we can check that
y=2
6646
14
6
73
775= 6e1+ 14e2+ 6e3+ 7e4
so by Denition VR [603],
D(y) =2
6646
14
6
73
775
Version 2.30
608 Section VR Vector Representations
which is not very exciting. Notice however that the order in which we place the vectors in the basis is
critical to the representation. Let's keep the standard unit vectors as our basis, but rearrange the order
we place them in the basis. So a fourth basis is
E=fe3;e4;e2;e1g
Then,
y=2
6646
14
6
73
775= 6e3+ 7e4+ 14e2+ 6e1
so by Denition VR [603],
E(y) =2
6646
7
14
63
775
So for every possible basis of C4we could construct a dierent representation of y.
Vector representations are most interesting for vector spaces that are not Cm.
Example VRP2
Vector representations in P2
Consider the vector u= 15 + 10x 6x22P2from the vector space of polynomials with degree at most 2
(Example VSP [319]). A nice basis for P2is
B=
1; x; x2
so that
u= 15 + 10x 6x2= 15(1) + 10( x) + ( 6)(x2)
so by Denition VR [603]
B(u) =2
415
10
63
5
Another nice basis for P2is
B=
1;1 +x;1 +x+x2
so that now it takes a bit of computation to determine the scalars for the representation. We want a1; a2; a3
so that
15 + 10x 6x2=a1(1) +a2(1 +x) +a3(1 +x+x2)
Performing the operations in P2on the right-hand side, and equating coecients, gives the three equations
in the three unknown scalars,
15 =a1+a2+a3
10 =a2+a3
6 =a3
The coecient matrix of this sytem is nonsingular, leading to a unique solution (no surprise there, see
Theorem VRRB [360]),
a1= 5 a2= 16 a3= 6
Version 2.30
Section VR Vector Representations 609
so by Denition VR [603]
C(u) =2
45
16
63
5
While we often form vector representations relative to \nice" bases, nothing prevents us from forming
representations relative to \nasty" bases. For example, the set
D=
2 x+ 3x2;1 2x2;5 + 4x+x2
can be veried as a basis of P2by checking linear independence with Denition LI [351] and then arguing
that 3 vectors from P2, a vector space of dimension 3 (Theorem DP [395]), must also be a spanning set
(Theorem G [407]). Now we desire scalars a1; a2; a3so that
15 + 10x 6x2=a1( 2 x+ 3x2) +a2(1 2x2) +a3(5 + 4x+x2)
Performing the operations in P2on the right-hand side, and equating coecients, gives the three equations
in the three unknown scalars,
15 = 2a1+a2+ 5a3
10 = a1+ 4a3
6 = 3a1 2a2+a3
The coecient matrix of this sytem is nonsingular, leading to a unique solution (no surprise there, see
Theorem VRRB [360]),
a1= 2 a2= 1 a3= 2
so by Denition VR [603]
D(u) =2
4 2
1
23
5
Theorem VRI
Vector Representation is Injective
The function B(Denition VR [603]) is an injective linear transformation.
Proof We will appeal to Theorem KILT [548]. Suppose Uis a vector space of dimension n, so vector
representation is of the form B:U!Cn. LetB=fu1;u2;u3; :::; ungbe the basis of Uused in the
denition of B. Suppose u2K(B). We write uas a linear combination of the vectors in the basis B
where the scalars are the components of the vector representation, B(u).
u= [B(u)]1u1+ [B(u)]2u2+ [B(u)]3u3++ [B(u)]nun Denition VR [603]
= [0]1u1+ [0]2u2+ [0]3u3++ [0]nun Denition KLT [545]
= 0u1+ 0u2+ 0u3++ 0un Denition ZCV [28]
=0+0+0++0 Theorem ZSSM [324]
=0 Property Z [318]
Version 2.30
610 Section VR Vector Representations
Thus an arbitrary vector, u, from the kernel , K(B), must equal the zero vector of U. SoK(B) =f0g
and by Theorem KILT [548], Bis injective.
Theorem VRS
Vector Representation is Surjective
The function B(Denition VR [603]) is a surjective linear transformation.
Proof We will appeal to Theorem RSLT [565]. Suppose Uis a vector space of dimension n, so vector
representation is of the form B:U!Cn. LetB=fu1;u2;u3; :::; ungbe the basis of Uused in the
denition of B. Suppose v2Cn. Dene the vector uby
u= [v]1u1+ [v]2u2+ [v]3u3++ [v]nun
Then for 1in
[B(u)]i= [B([v]1u1+ [v]2u2+ [v]3u3++ [v]nun)]i
= [v]i Denition VR [603]
so the entries of vectors B(u) and vare equal and Denition CVE [98] yields the vector equality B(u) =
v. This demonstrates that v2R(B), soCnR(B). SinceR(B)Cnby Denition RLT [563], we
haveR(B) =Cnand Theorem RSLT [565] says Bis surjective.
We will have many occasions later to employ the inverse of vector representation, so we will record the
fact that vector representation is an invertible linear transformation.
Theorem VRILT
Vector Representation is an Invertible Linear Transformation
The function B(Denition VR [603]) is an invertible linear transformation.
Proof The function B(Denition VR [603]) is a linear transformation (Theorem VRLT [603]) that is
injective (Theorem VRI [607]) and surjective (Theorem VRS [608]) with domain Vand codomain Cn. By
Theorem ILTIS [582] we then know that Bis an invertible linear transformation.
Informally, we will refer to the application of Bascoordinatizing a vector, while the application of
1
Bwill be referred to as un-coordinatizing a vector.
Subsection CVS
Characterization of Vector Spaces
Limiting our attention to vector spaces with nite dimension, we now describe every possible vector space.
All of them. Really.
Theorem CFDVS
Characterization of Finite Dimensional Vector Spaces
Suppose that Vis a vector space with dimension n. ThenVis isomorphic to Cn.
Proof SinceVhas dimension nwe can nd a basis of Vof sizen(Denition D [391]) which we will call
B. The linear transformation Bis an invertible linear transformation from VtoCn, so by Denition IVS
[586], we have that VandCnare isomorphic.
Theorem CFDVS [608] is the rst of several surprises in this chapter, though it might be a bit demor-
alizing too. It says that there really are not all that many dierent (nite dimensional) vector spaces, and
none are really any more complicated than Cn. Hmmm. The following examples should make this point.
Version 2.30
Subsection VR.CP Coordinatization Principle 611
Example TIVS
Two isomorphic vector spaces
The vector space of polynomials with degree 8 or less, P8, has dimension 9 (Theorem DP [395]). By
Theorem CFDVS [608], P8is isomorphic to C9.
Example CVSR
Crazy vector space revealed
The crazy vector space, Cof Example CVS [322], has dimension 2 by Example DC [396]. By Theorem
CFDVS [608], Cis isomorphic to C2. Hmmmm. Not really so crazy after all?
Example ASC
A subspace characterized
In Example DSP4 [396] we determined that a certain subspace WofP4has dimension 4. By Theorem
CFDVS [608], Wis isomorphic to C4.
Theorem IFDVS
Isomorphism of Finite Dimensional Vector Spaces
SupposeUandVare both nite-dimensional vector spaces. Then UandVare isomorphic if and only if
dim (U) = dim (V).
Proof ()) This is just the statement proved in Theorem IVSED [587].
(() This is the advertised converse of Theorem IVSED [587]. We will assume UandVhave equal
dimension and discover that they are isomorphic vector spaces. Let nbe the common dimension of Uand
V. Then by Theorem CFDVS [608] there are isomorphisms T:U!CnandS:V!Cn.
Tis therefore an invertible linear transformation by Denition IVS [586]. Similarly, Sis an invertible
linear transformation, and so S 1is an invertible linear transformation (Theorem IILT [582]). The com-
position of invertible linear transformations is again invertible (Theorem CIVLT [585]) so the composition
ofS 1withTis invertible. Then
S 1T
:U!Vis an invertible linear transformation from UtoV
and Denition IVS [586] says UandVare isomorphic.
Example MIVS
Multiple isomorphic vector spaces
C10,P9,M2;5andM5;2are all vector spaces and each has dimension 10. By Theorem IFDVS [609] each is
isomorphic to any other.
The subspace of M4;4that contains all the symmetric matrices (Denition SYM [211]) has dimension
10, so this subspace is also isomorphic to each of the four vector spaces above.
Subsection CP
Coordinatization Principle
WithBavailable as an invertible linear transformation, we can translate between vectors in a vector space
Uof dimension mandCm. Furthermore, as a linear transformation, Brespects the addition and scalar
multiplication in U, while 1
Brespects the addition and scalar multiplication in Cm. Since our denitions
of linear independence, spans, bases and dimension are all built up from linear combinations, we will nally
be able to translate fundamental properties between abstract vector spaces ( U) and concrete vector spaces
(Cm).
Theorem CLI
Coordinatization and Linear Independence
Suppose that Uis a vector space with a basis Bof sizen. ThenS=fu1;u2;u3; :::; ukgis a linearly inde-
Version 2.30
612 Section VR Vector Representations
pendent subset of Uif and only if R=fB(u1); B(u2); B(u3); :::; B(uk)gis a linearly independent
subset of Cn.
Proof The linear transformation Bis an isomorphism between UandCn(Theorem VRILT [608]). As
an invertible linear transformation, Bis an injective linear transformation (Theorem ILTIS [582]), and
1
Bis also an injective linear transformation (Theorem IILT [582], Theorem ILTIS [582]).
()) SinceBis an injective linear transformation and Sis linearly independent, Theorem ILTLI [549]
says thatRis linearly independent.
(() If we apply 1
Bto each element of R, we will create the set S. Since we are assuming Ris linearly
independent and 1
Bis injective, Theorem ILTLI [549] says that Sis linearly independent.
Theorem CSS
Coordinatization and Spanning Sets
Suppose that Uis a vector space with a basis Bof sizen. Then u2hfu1;u2;u3; :::; ukgiif and only if
B(u)2hfB(u1); B(u2); B(u3); :::; B(uk)gi.
Proof ()) Suppose u2hfu1;u2;u3; :::; ukgi. Then there are scalars, a1; a2; a3; :::; ak, such that
u=a1u1+a2u2+a3u3++akuk
Then,
B(u) =B(a1u1+a2u2+a3u3++akuk)
=a1B(u1) +a2B(u2) +a3B(u3) ++akB(uk) Theorem LTLC [525]
which says that B(u)2hfB(u1); B(u2); B(u3); :::; B(uk)gi.
(() Suppose that B(u)2hfB(u1); B(u2); B(u3); :::; B(uk)gi. Then there are scalars b1; b2; b3; :::; bk
such that
B(u) =b1B(u1) +b2B(u2) +b3B(u3) ++bkB(uk)
Recall that Bis invertible (Theorem VRILT [608]), so
u=IU(u) Denition IDLT [579]
=
1
BB
(u) Denition IVLT [579]
= 1
B(B(u)) Denition LTC [532]
= 1
B(b1B(u1) +b2B(u2) +b3B(u3) ++bkB(uk))
=b1 1
B(B(u1)) +b2 1
B(B(u2)) +b3 1
B(B(u3))
++bk 1
B(B(uk)) Theorem LTLC [525]
=b1IU(u1) +b2IU(u2) +b3IU(u3) ++bkIU(uk) Denition IVLT [579]
=b1u1+b2u2+b3u3++bkuk Denition IDLT [579]
which says that u2hfu1;u2;u3; :::; ukgi.
Here's a fairly simple example that illustrates a very, very important idea.
Example CP2
Coordinatizing in P2
In Example VRP2 [606] we needed to know that
D=
2 x+ 3x2;1 2x2;5 + 4x+x2
is a basis for P2. With Theorem CLI [609] and Theorem CSS [610] this task is much easier. First, choose
a known basis for P2, a basis that forms vector representations easily. We will choose
B=
1; x; x2
Version 2.30
Subsection VR.CP Coordinatization Principle 613
Now, form the subset of C3that is the result of applying Bto each element of D,
F=
B
2 x+ 3x2
; B
1 2x2
; B
5 + 4x+x2
=8
<
:2
4 2
1
33
5;2
41
0
23
5;2
45
4
13
59
=
;
and ask ifFis a linearly independent spanning set for C3. This is easily seen to be the case by forming a
matrixAwhose columns are the vectors of F, row-reducing Ato the identity matrix I3, and then using
the nonsingularity of Ato assert that Fis a basis for C3(Theorem CNMB [376]). Now, since Fis a basis
forC3, Theorem CLI [609] and Theorem CSS [610] tell us that Dis also a basis for P2.
Example CP2 [610] illustrates the broad notion that computations in abstract vector spaces can be
reduced to computations in Cm. You may have noticed this phenomenon as you worked through examples
in Chapter VS [317] or Chapter LT [515] employing vector spaces of matrices or polynomials. These
computations seemed to invariably result in systems of equations or the like from Chapter SLE [3], Chapter
V [97] and Chapter M [207]. It is vector representation, B, that allows us to make this connection formal
and precise.
Knowing that vector representation allows us to translate questions about linear combinations, linear
independence and spans from general vector spaces to Cmallows us to prove a great many theorems about
how to translate other properties. Rather than prove these theorems, each of the same style as the other,
we will oer some general guidance about how to best employ Theorem VRLT [603], Theorem CLI [609]
and Theorem CSS [610]. This comes in the form of a \principle": a basic truth, but most denitely not a
theorem (hence, no proof).
The Coordinatization Principle Suppose that Uis a vector space with a basis Bof sizen. Then any
question about U, or its elements, which ultimately depends on the vector addition or scalar multiplication
inU, or depends on linear independence or spanning, may be translated into the same question in Cn
by application of the linear transformation Bto the relevant vectors. Once the question is answered
inCn, the answer may be translated back to U(if necessary) through application of the inverse linear
transformation 1
B.
Example CM32
Coordinatization in M32
This is a simple example of the Coordinatization Principle [611], depending only on the fact that coordina-
tizing is an invertible linear transformation (Theorem VRILT [608]). Suppose we have a linear combination
to perform in M32, the vector space of 3 2 matrices, but we are adverse to doing the operations of M32
(Denition MA [207], Denition MSM [208]). More specically, suppose we are faced with the computation
62
43 7
2 4
0 33
5+ 22
4 1 3
4 8
2 53
5
We choose a nice basis for M32(or a nasty basis if we are so inclined),
B=8
<
:2
41 0
0 0
0 03
5;2
40 0
1 0
0 03
5;2
40 0
0 0
1 03
5;2
40 1
0 0
0 03
5;2
40 0
0 1
0 03
5;2
40 0
0 0
0 13
59
=
;
and applyBto each vector in the linear combination. This gives us a new computation, now in the vector
Version 2.30
614 Section VR Vector Representations
spaceC6,
62
66666643
2
0
7
4
33
7777775+ 22
6666664 1
4
2
3
8
53
7777775
which we can compute with the operations of C6(Denition CVA [98], Denition CVSM [99]), to arrive at
2
666666416
4
4
48
40
83
7777775
We are after the result of a computation in M32, so we now can apply 1
Bto obtain a 32 matrix,
162
41 0
0 0
0 03
5+ ( 4)2
40 0
1 0
0 03
5+ ( 4)2
40 0
0 0
1 03
5+ 482
40 1
0 0
0 03
5+ 402
40 0
0 1
0 03
5+ ( 8)2
40 0
0 0
0 13
5=2
416 48
4 40
4 83
5
which is exactly the matrix we would have computed had we just performed the matrix operations in the
rst place. So this was not meant to be an easier way to compute a linear combination of two matrices,
just a dierent way.
Subsection READ
Reading Questions
1. The vector space of 3 5 matrices, M3;5is isomorphic to what fundamental vector space?
2. A basis for C3is
B=8
<
:2
41
2
13
5;2
43
1
23
5;2
41
1
13
59
=
;
ComputeB0
@2
45
8
13
51
A.
3. What is the rst \surprise," and why is it surprising?
Version 2.30
Subsection VR.EXC Exercises 615
Subsection EXC
Exercises
C10 In the vector space C3, compute the vector representation B(v) for the basis Band vector vbelow.
B=8
<
:2
42
2
23
5;2
41
3
13
5;2
43
5
23
59
=
;v=2
411
5
83
5
Contributed by Robert Beezer Solution [614]
C20 Rework Example CM32 [611] replacing the basis Bby the basis
C=8
<
:2
4 14 9
10 10
6 23
5;2
4 7 4
5 5
3 13
5;2
4 3 1
0 2
1 13
5;2
4 7 4
3 2
1 03
5;2
44 2
3 3
2 13
5;2
40 0
1 2
1 13
59
=
;
Contributed by Robert Beezer Solution [614]
M10 Prove that the set Sbelow is a basis for the vector space of 2 2 matrices, M22. Do this choosing
a natural basis for M22and coordinatizing the elements of Swith respect to this basis. Examine the
resulting set of column vectors from C4and apply the Coordinatization Principle [611].
S=33 99
78 9
; 16 47
36 2
;10 27
17 3
; 2 7
6 4
Contributed by Andy Zimmer
Version 2.30
616 Section VR Vector Representations
Subsection SOL
Solutions
C10 Contributed by Robert Beezer Statement [613]
We need to express the vector vas a linear combination of the vectors in B. Theorem VRRB [360] tells
us we will be able to do this, and do it uniquely. The vector equation
a12
42
2
23
5+a22
41
3
13
5+a32
43
5
23
5=2
411
5
83
5
becomes (via Theorem SLSLC [112]) a system of linear equations with augmented matrix,
2
42 1 3 11
2 3 5 5
2 1 2 83
5
This system has the unique solution a1= 2,a2= 2,a3= 3. So by Denition VR [603],
B(v) =B0
@2
411
5
83
51
A=B0
@22
42
2
23
5+ ( 2)2
41
3
13
5+ 32
43
5
23
51
A=2
42
2
33
5
C20 Contributed by Robert Beezer Statement [613]
The following computations replicate the computations given in Example CM32 [611], only using the basis
C.
C0
@2
43 7
2 4
0 33
51
A=2
6666664 9
12
6
7
2
13
7777775C0
@2
4 1 3
4 8
2 53
51
A=2
6666664 11
34
4
1
16
53
7777775
62
6666664 9
12
6
7
2
13
7777775+ 22
6666664 11
34
4
1
16
53
7777775=2
6666664 76
140
44
40
20
43
7777775 1
C0
BBBBBB@2
6666664 76
140
44
40
20
43
77777751
CCCCCCA=2
416 48
4 30
4 83
5
Version 2.30
Section MR Matrix Representations 617
Section MR
Matrix Representations
We have seen that linear transformations whose domain and codomain are vector spaces of columns vec-
tors have a close relationship with matrices (Theorem MBLT [522], Theorem MLTCV [523]). In this
section, we will extend the relationship between matrices and linear transformations to the setting of linear
transformations between abstract vector spaces.
Denition MR
Matrix Representation
Suppose that T:U!Vis a linear transformation, B=fu1;u2;u3; :::; ungis a basis for Uof sizen,
andCis a basis for Vof sizem. Then the matrix representation ofTrelative toBandCis themn
matrix,
MT
B;C= [C(T(u1))jC(T(u2))jC(T(u3))j:::jC(T(un))]
(This denition contains Notation MR.) 4
Example OLTTR
One linear transformation, three representations
Consider the linear transformation
S:P3!M22; S
a+bx+cx2+dx3
=3a+ 7b 2c 5d8a+ 14b 2c 11d
4a 8b+ 2c+ 6d12a+ 22b 4c 17d
First, we build a representation relative to the bases,
B=
1 + 2x+x2 x3;1 + 3x+x2+x3; 1 2x+ 2x3;2 + 3x+ 2x2 5x3
C=1 1
1 2
;2 3
2 5
; 1 1
0 2
; 1 4
2 4
We evaluate Swith each element of the basis for the domain, B, and coordinatize the result relative to the
vectors in the basis for the codomain, C. Notice here how we take elements of vector spaces and decompose
them into linear combinations of basis elements as the key step in constructing coordinatizations of vectors.
There is a system of equations involved almost every time, but we will omit these details since this should
be a routine exercise at this stage.
C
S
1 + 2x+x2 x3
=C20 45
24 69
=C
( 90)1 1
1 2
+ 372 3
2 5
+ ( 40) 1 1
0 2
+ 4 1 4
2 4
=2
664 90
37
40
43
775
C
S
1 + 3x+x2+x3
=C17 37
20 57
=C
( 72)1 1
1 2
+ 292 3
2 5
+ ( 34) 1 1
0 2
+ 3 1 4
2 4
=2
664 72
29
34
33
775
C
S
1 2x+ 2x3
=C 27 58
32 90
Version 2.30
618 Section MR Matrix Representations
=C
1141 1
1 2
+ ( 46)2 3
2 5
+ 54 1 1
0 2
+ ( 5) 1 4
2 4
=2
664114
46
54
53
775
C
S
2 + 3x+ 2x2 5x3
=C48 109
58 167
=C
( 220)1 1
1 2
+ 912 3
2 5
+ 96 1 1
0 2
+ 10 1 4
2 4
=2
664 220
91
96
103
775
Thus, employing Denition MR [615]
MS
B;C=2
664 90 72 114 220
37 29 46 91
40 34 54 96
4 3 5 103
775
Often we use \nice" bases to build matrix representations and the work involved is much easier. Suppose
we take bases
D=
1; x; x2; x3
E=1 0
0 0
;0 1
0 0
;0 0
1 0
;0 0
0 1
The evaluation of Sat the elements of Dis easy and coordinatization relative to Ecan be done on sight,
E(S(1)) =E3 8
4 12
=E
31 0
0 0
+ 80 1
0 0
+ ( 4)0 0
1 0
+ 120 0
0 1
=2
6643
8
4
123
775
E(S(x)) =E7 14
8 22
=E
71 0
0 0
+ 140 1
0 0
+ ( 8)0 0
1 0
+ 220 0
0 1
=2
6647
14
8
223
775
E
S
x2
=E 2 2
2 4
=E
( 2)1 0
0 0
+ ( 2)0 1
0 0
+ 20 0
1 0
+ ( 4)0 0
0 1
=2
664 2
2
2
43
775
E
S
x3
=E 5 11
6 17
=E
( 5)1 0
0 0
+ ( 11)0 1
0 0
+ 60 0
1 0
+ ( 17)0 0
0 1
=2
664 5
11
6
173
775
Version 2.30
Section MR Matrix Representations 619
So the matrix representation of Srelative toDandEis
MS
D;E=2
6643 7 2 5
8 14 2 11
4 8 2 6
12 22 4 173
775
One more time, but now let's use bases
F=
1 +x x2+ 2x3; 1 + 2x+ 2x3;2 +x 2x2+ 3x3;1 +x+ 2x3
G=1 1
1 2
; 1 2
0 2
;2 1
2 3
;1 1
0 2
and evaluate Swith the elements of F, then coordinatize the results relative to G,
G
S
1 +x x2+ 2x3
=G2 2
2 4
=G
21 1
1 2
=2
6642
0
0
03
775
G
S
1 + 2x+ 2x3
=G1 2
0 2
=G
( 1) 1 2
0 2
=2
6640
1
0
03
775
G
S
2 +x 2x2+ 3x3
=G2 1
2 3
=G2 1
2 3
=2
6640
0
1
03
775
G
S
1 +x+ 2x3
=G0 0
0 0
=G
01 1
0 2
=2
6640
0
0
03
775
So we arrive at an especially economical matrix representation,
MS
F;G=2
6642 0 0 0
0 1 0 0
0 0 1 0
0 0 0 03
775
We may choose to use whatever terms we want when we make a denition. Some are arbitrary, while
others make sense, but only in light of subsequent theorems. Matrix representation is in the latter category.
We begin with a linear transformation and produce a matrix. So what? Here's the theorem that justies
the term \matrix representation."
Theorem FTMR
Fundamental Theorem of Matrix Representation
Suppose that T:U!Vis a linear transformation, Bis a basis for U,Cis a basis for VandMT
B;Cis the
matrix representation of Trelative toBandC. Then, for any u2U,
C(T(u)) =MT
B;C(B(u))
Version 2.30
620 Section MR Matrix Representations
or equivalently
T(u) = 1
C
MT
B;C(B(u))
Proof LetB=fu1;u2;u3; :::; ungbe the basis of U. Since u2U, there are scalars a1; a2; a3; :::; an
such that
u=a1u1+a2u2+a3u3++anun
Then,
MT
B;CB(u)
= [C(T(u1))jC(T(u2))jC(T(u3))j:::jC(T(un))]B(u) Denition MR [615]
= [C(T(u1))jC(T(u2))jC(T(u3))j:::jC(T(un))]2
666664a1
a2
a3
...
an3
777775Denition VR [603]
=a1C(T(u1)) +a2C(T(u2)) ++anC(T(un)) Denition MVP [223]
=C(a1T(u1) +a2T(u2) +a3T(u3) ++anT(un)) Theorem LTLC [525]
=C(T(a1u1+a2u2+a3u3++anun)) Theorem LTLC [525]
=C(T(u))
The alternative conclusion is obtained as
T(u) =IV(T(u)) Denition IDLT [579]
=
1
CC
(T(u)) Denition IVLT [579]
= 1
C(C(T(u))) Denition LTC [532]
= 1
C
MT
B;C(B(u))
This theorem says that we can apply Ttouand coordinatize the result relative to CinV, or we can
rst coordinatize urelative toBinU, then multiply by the matrix representation. Either way, the result is
the same. So the eect of a linear transformation can always be accomplished by a matrix-vector product
(Denition MVP [223]). That's important enough to say again. The eect of a linear transformation is a
matrix-vector product.
u
ρB(u)T(u)
MT
B,CρB(u)=ρC(T(u))T
MT
B,CρB ρC
Diagram FTMR. Fundamental Theorem of Matrix Representations
The alternative conclusion of this result might be even more striking. It says that to eect a linear trans-
formation ( T) of a vector ( u), coordinatize the input (with B), do a matrix-vector product (with MT
B;C),
and un-coordinatize the result (with 1
C). So, absent some bookkeeping about vector representations, a
Version 2.30
Section MR Matrix Representations 621
linear transformation isa matrix. To adjust the diagram, we \reverse" the arrow on the right, which
means inverting the vector representation ConV. Now we can go directly across the top of the diagram,
computing the linear transformation between the abstract vector spaces. Or, we can around the other
three sides, using vector representation, a matrix-vector product, followed by un-coordinatization.
u
ρB(u)T(u)=ρ−1
C/parenleftbig
MT
B,CρB(u)/parenrightbig
MT
B,CρB(u)T
MT
B,CρB ρ−1
C
Diagram FTMRA. Fundamental Theorem of Matrix Representations (Alternate)
Here's an example to illustrate how the \action" of a linear transformation can be eected by matrix
multiplication.
Example ALTMM
A linear transformation as matrix multiplication
In Example OLTTR [615] we found three representations of the linear transformation S. In this example,
we will compute a single output of Sin four dierent ways. First \normally," then three times over using
Theorem FTMR [617].
Choosep(x) = 3 x+ 2x2 5x3, for no particular reason. Then the straightforward application of S
top(x) yields
S(p(x)) =S
3 x+ 2x2 5x3
=3(3) + 7( 1) 2(2) 5( 5) 8(3) + 14( 1) 2(2) 11( 5)
4(3) 8( 1) + 2(2) + 6( 5) 12(3) + 22( 1) 4(2) 17( 5)
=23 61
30 91
Now use the representation of Srelative to the bases BandCand Theorem FTMR [617]. Note that we
will employ the following linear combination in moving from the second line to the third,
3 x+ 2x2 5x3= 48(1 + 2x+x2 x3) + ( 20)(1 + 3x+x2+x3)+
( 1)( 1 2x+ 2x3) + ( 13)(2 + 3x+ 2x2 5x3)
S(p(x)) = 1
C
MS
B;CB(p(x))
= 1
C
MS
B;CB
3 x+ 2x2 5x3
= 1
C0
BB@MS
B;C2
66448
20
1
133
7751
CCA
= 1
C0
BB@2
664 90 72 114 220
37 29 46 91
40 34 54 96
4 3 5 103
7752
66448
20
1
133
7751
CCA
= 1
C0
BB@2
664 134
59
46
73
7751
CCA
Version 2.30
622 Section MR Matrix Representations
= ( 134)1 1
1 2
+ 592 3
2 5
+ ( 46) 1 1
0 2
+ 7 1 4
2 4
=23 61
30 91
Again, but now with \nice" bases like DandE, and the computations are more transparent.
S(p(x)) = 1
E
MS
D;ED(p(x))
= 1
E
MS
D;ED
3 x+ 2x2 5x3
= 1
E
MS
D;ED
3(1) + ( 1)(x) + 2(x2) + ( 5)(x3)
= 1
E0
BB@MS
D;E2
6643
1
2
53
7751
CCA
= 1
E0
BB@2
6643 7 2 5
8 14 2 11
4 8 2 6
12 22 4 173
7752
6643
1
2
53
7751
CCA
= 1
E0
BB@2
66423
61
30
913
7751
CCA
= 231 0
0 0
+ 610 1
0 0
+ ( 30)0 0
1 0
+ 910 0
0 1
=23 61
30 91
OK, last time, now with the bases FandG. The coordinatizations will take some work this time, but the
matrix-vector product (Denition MVP [223]) (which is the actual action of the linear transformation) will
be especially easy, given the diagonal nature of the matrix representation, MS
F;G. Here we go,
S(p(x)) = 1
G
MS
F;GF(p(x))
= 1
G
MS
F;GF
3 x+ 2x2 5x3
= 1
G
MS
F;GF
32(1 +x x2+ 2x3) 7( 1 + 2x+ 2x3) 17(2 +x 2x2+ 3x3) 2(1 +x+ 2x3)
= 1
G0
BB@MS
F;G2
66432
7
17
23
7751
CCA
= 1
G0
BB@2
6642 0 0 0
0 1 0 0
0 0 1 0
0 0 0 03
7752
66432
7
17
23
7751
CCA
= 1
G0
BB@2
66464
7
17
03
7751
CCA
= 641 1
1 2
+ 7 1 2
0 2
+ ( 17)2 1
2 3
+ 01 1
0 2
Version 2.30
Subsection MR.NRFO New Representations from Old 623
=23 61
30 91
This example is not meant to necessarily illustrate that any one of these four computations is simpler than
the others. Instead, it is meant to illustrate the many dierent ways we can arrive at the same result, with
the last three all employing a matrix representation to eect the linear transformation.
We will use Theorem FTMR [617] frequently in the next few sections. A typical application will feel like
the linear transformation T\commutes" with a vector representation, C, and as it does the transformation
morphs into a matrix, MT
B;C, while the vector representation changes to a new basis, B. Or vice-versa.
Subsection NRFO
New Representations from Old
In Subsection LT.NLTFO [530] we built new linear transformations from other linear transformations.
Sums, scalar multiples and compositions. These new linear transformations will have matrix representations
as well. How do the new matrix representations relate to the old matrix representations? Here are the
three theorems.
Theorem MRSLT
Matrix Representation of a Sum of Linear Transformations
Suppose that T:U!VandS:U!Vare linear transformations, Bis a basis of UandCis a basis of
V. Then
MT+S
B;C=MT
B;C+MS
B;C
Proof Letxbe any vector in Cn. Dene u2Ubyu= 1
B(x), sox=B(u). Then,
MT+S
B;Cx=MT+S
B;CB(u) Substitution
=C((T+S) (u)) Theorem FTMR [617]
=C(T(u) +S(u)) Denition LTA [530]
=C(T(u)) +C(S(u)) Denition LT [515]
=MT
B;C(B(u)) +MS
B;C(B(u)) Theorem FTMR [617]
=
MT
B;C+MS
B;C
B(u) Theorem MMDAA [230]
=
MT
B;C+MS
B;C
x Substitution
Since the matrices MT+S
B;CandMT
B;C+MS
B;Chave equal matrix-vector products for every vector in Cn, by
Theorem EMMVP [225] they are equal matrices. (Now would be a good time to double-back and study
the proof of Theorem EMMVP [225]. You did promise to come back to this theorem sometime, didn't
you?)
Theorem MRMLT
Matrix Representation of a Multiple of a Linear Transformation
Suppose that T:U!Vis a linear transformation, 2C,Bis a basis of UandCis a basis of V. Then
MT
B;C=MT
B;C
Proof Letxbe any vector in Cn. Dene u2Ubyu= 1
B(x), sox=B(u). Then,
MT
B;Cx=MT
B;CB(u) Substitution
Version 2.30
624 Section MR Matrix Representations
=C((T) (u)) Theorem FTMR [617]
=C(T(u)) Denition LTSM [531]
=C(T(u)) Denition LT [515]
=
MT
B;CB(u)
Theorem FTMR [617]
=
MT
B;C
B(u) Theorem MMSMM [230]
=
MT
B;C
x Substitution
Since the matrices MT
B;CandMT
B;Chave equal matrix-vector products for every vector in Cn, by Theorem
EMMVP [225] they are equal matrices.
The vector space of all linear transformations from UtoVis now isomorphic to the vector space of all
mnmatrices.
Theorem MRCLT
Matrix Representation of a Composition of Linear Transformations
Suppose that T:U!VandS:V!Ware linear transformations, Bis a basis of U,Cis a basis of V,
andDis a basis of W. Then
MST
B;D=MS
C;DMT
B;C
Proof Letxbe any vector in Cn. Dene u2Ubyu= 1
B(x), sox=B(u). Then,
MST
B;Dx=MST
B;DB(u) Substitution
=D((ST) (u)) Theorem FTMR [617]
=D(S(T(u))) Denition LTC [532]
=MS
C;DC(T(u)) Theorem FTMR [617]
=MS
C;D
MT
B;CB(u)
Theorem FTMR [617]
=
MS
C;DMT
B;C
B(u) Theorem MMA [231]
=
MS
C;DMT
B;C
x Substitution
Since the matrices MST
B;DandMS
C;DMT
B;Chave equal matrix-vector products for every vector in Cn, by
Theorem EMMVP [225] they are equal matrices.
This is the second great surprise of introductory linear algebra. Matrices are linear transformations
(functions, really), and matrix multiplication is function composition! We can form the composition of
two linear transformations, then form the matrix representation of the result. Or we can form the matrix
representation of each linear transformation separately, then multiply the two representations together via
Denition MM [226]. In either case, we arrive at the same result.
Example MPMR
Matrix product of matrix representations
Consider the two linear transformations,
T:C2!P2Ta
b
= ( a+ 3b) + (2a+ 4b)x+ (a 2b)x2
S:P2!M22S
a+bx+cx2
=2a+b+ 2c a + 4b c
a+ 3c3a+b+ 2c
and bases for C2,P2andM22(respectively),
B=3
1
;2
1
Version 2.30
Subsection MR.NRFO New Representations from Old 625
C=
1 2x+x2; 1 + 3x;2x+ 3x2
D=1 2
1 1
;1 1
1 2
; 1 2
0 0
;2 3
2 2
Begin by computing the new linear transformation that is the composition of TandS(Denition LTC
[532], Theorem CLTLT [533]), ( ST) :C2!M22,
(ST)a
b
=S
Ta
b
=S
( a+ 3b) + (2a+ 4b)x+ (a 2b)x2
=2( a+ 3b) + (2a+ 4b) + 2(a 2b) ( a+ 3b) + 4(2a+ 4b) (a 2b)
( a+ 3b) + 3(a 2b) 3( a+ 3b) + (2a+ 4b) + 2(a 2b)
=2a+ 6b6a+ 21b
4a 9b a + 9b
Now compute the matrix representations (Denition MR [615]) for each of these three linear transformations
(T,S,ST), relative to the appropriate bases. First for T,
C
T3
1
=C
10x+x2
=C
28(1 2x+x2) + 28( 1 + 3x) + ( 9)(2x+ 3x2)
=2
428
28
93
5
C
T2
1
=C(1 + 8x)
=C
33(1 2x+x2) + 32( 1 + 3x) + ( 11)(2x+ 3x2)
=2
433
32
113
5
So we have the matrix representation of T,
MT
B;C=2
428 33
28 32
9 113
5
Now, a representation of S,
D
S
1 2x+x2
=D2 8
2 3
=D
( 11)1 2
1 1
+ ( 21)1 1
1 2
+ 0 1 2
0 0
+ (17)2 3
2 2
=2
664 11
21
0
173
775
D(S( 1 + 3x)) =D1 11
1 0
=D
261 2
1 1
+ 511 1
1 2
+ 0 1 2
0 0
+ ( 38)2 3
2 2
Version 2.30
626 Section MR Matrix Representations
=2
66426
51
0
383
775
D
S
2x+ 3x2
=D8 5
9 8
=D
341 2
1 1
+ 671 1
1 2
+ 1 1 2
0 0
+ ( 46)2 3
2 2
=2
66434
67
1
463
775
So we have the matrix representation of S,
MS
C;D=2
664 11 26 34
21 51 67
0 0 1
17 38 463
775
Finally, a representation of ST,
D
(ST)3
1
=D12 39
3 12
=D
1141 2
1 1
+ 2371 1
1 2
+ ( 9) 1 2
0 0
+ ( 174)2 3
2 2
=2
664114
237
9
1743
775
D
(ST)2
1
=D10 33
1 11
=D
951 2
1 1
+ 2021 1
1 2
+ ( 11) 1 2
0 0
+ ( 149)2 3
2 2
=2
66495
202
11
1493
775
So we have the matrix representation of ST,
MST
B;D=2
664114 95
237 202
9 11
174 1493
775
Version 2.30
Subsection MR.PMR Properties of Matrix Representations 627
Now, we are all set to verify the conclusion of Theorem MRCLT [622],
MS
C;DMT
B;C=2
664 11 26 34
21 51 67
0 0 1
17 38 463
7752
428 33
28 32
9 113
5
=2
664114 95
237 202
9 11
174 1493
775
=MST
B;D
We have intentionally used non-standard bases. If you were to choose \nice" bases for the three vector
spaces, then the result of the theorem might be rather transparent. But this would still be a worthwhile
exercise | give it a go.
A diagram, similar to ones we have seen earlier, might make the importance of this theorem clearer,
S,T
S◦TMS
C,D,MT
B,C
MS◦T
B,D=MS
C,DMT
B,CDefinition MR
Definition MRDefinition LTC Definition MM
Diagram MRCLT. Matrix Representation and Composition of Linear Transformations
One of our goals in the rst part of this book is to make the denition of matrix multiplication (Denition
MVP [223], Denition MM [226]) seem as natural as possible. However, many are brought up with an entry-
by-entry description of matrix multiplication (Theorem ME [485]) as the denition of matrix multiplication,
and then theorems about columns of matrices and linear combinations follow from that denition. With
this unmotivated denition, the realization that matrix multiplication is function composition is quite
remarkable. It is an interesting exercise to begin with the question, \What is the matrix representation
of the composition of two linear transformations?" and then, without using any theorems about matrix
multiplication, nally arrive at the entry-by-entry description of matrix multiplication. Try it yourself
(Exercise MR.T80 [637]).
Subsection PMR
Properties of Matrix Representations
It will not be a surprise to discover that the kernel and range of a linear transformation are closely related
to the null space and column space of the transformation's matrix representation. Perhaps this idea has
been bouncing around in your head already, even before seeing the denition of a matrix representation.
However, with a formal denition of a matrix representation (Denition MR [615]), and a fundamental
theorem to go with it (Theorem FTMR [617]) we can be formal about the relationship, using the idea of
isomorphic vector spaces (Denition IVS [586]). Here are the twin theorems.
Theorem KNSI
Kernel and Null Space Isomorphism
Suppose that T:U!Vis a linear transformation, Bis a basis for Uof sizen, andCis a basis for V.
Version 2.30
628 Section MR Matrix Representations
Then the kernel of Tis isomorphic to the null space of MT
B;C,
K(T)=N
MT
B;C
Proof To establish that two vector spaces are isomorphic, we must nd an isomorphism between them,
an invertible linear transformation (Denition IVS [586]). The kernel of the linear transformation T,K(T),
is a subspace of U, while the null space of the matrix representation, N
MT
B;C
is a subspace of Cn. The
functionBis dened as a function from UtoCn, but we can just as well employ the denition of Bas
a function fromK(T) toN
MT
B;C
.
We must rst insure that if we choose an input for BfromK(T) that then the output will be an
element ofN
MT
B;C
. So suppose that u2K(T). Then
MT
B;CB(u) =C(T(u)) Theorem FTMR [617]
=C(0) Denition KLT [545]
=0 Theorem LTTZZ [519]
This says that B(u)2N
MT
B;C
, as desired.
The restriction in the size of the domain and codomain Bwill not aect the fact that Bis a linear
transformation (Theorem VRLT [603]), nor will it aect the fact that Bis injective (Theorem VRI
[607]). Something must be done though to verify that Bis surjective. To this end, appeal to the
denition of surjective (Denition SLT [559]), and suppose that we have an element of the codomain,
x2N
MT
B;C
Cnand we wish to nd an element of the domain with xas its image. We now show
that the desired element of the domain is u= 1
B(x). First, verify that u2K(T),
T(u) =T
1
B(x)
= 1
C
MT
B;C
B
1
B(x)
Theorem FTMR [617]
= 1
C
MT
B;C(ICn(x))
Denition IVLT [579]
= 1
C
MT
B;Cx
Denition IDLT [579]
= 1
C(0Cn) Denition KLT [545]
=0V Theorem LTTZZ [519]
Second, verify that the proposed isomorphism, B, takes utox,
B(u) =B
1
B(x)
Substitution
=ICn(x) Denition IVLT [579]
=x Denition IDLT [579]
WithBdemonstrated to be an injective and surjective linear transformation from K(T) toN
MT
B;C
,
Theorem ILTIS [582] tells us Bis invertible, and so by Denition IVS [586], we say K(T) andN
MT
B;C
are isomorphic.
Example KVMR
Kernel via matrix representation
Consider the kernel of the linear transformation
T:M22!P2; Ta b
c d
= (2a b+c 5d) + (a+ 4b+ 5b+ 2d)x+ (3a 2b+c 8d)x2
Version 2.30
Subsection MR.PMR Properties of Matrix Representations 629
We will begin with a matrix representation of Trelative to the bases for M22andP2(respectively),
B=1 2
1 1
;1 3
1 4
;1 2
0 2
;2 5
2 4
C=
1 +x+x2;2 + 3x; 1 2x2
Then,
C
T1 2
1 1
=C
4 + 2x+ 6x2
=C
2(1 +x+x2) + 0(2 + 3x) + ( 2)( 1 2x2)
=2
42
0
23
5
C
T1 3
1 4
=C
18 + 28x2
=C
( 24)(1 +x+x2) + 8(2 + 3x) + ( 26)( 1 2x2)
=2
4 24
8
263
5
C
T1 2
0 2
=C
10 + 5x+ 15x2
=C
5(1 +x+x2) + 0(2 + 3x) + ( 5)( 1 2x2)
=2
45
0
53
5
C
T2 5
2 4
=C
17 + 4x+ 26x2
=C
( 8)(1 +x+x2) + (4)(2 + 3 x) + ( 17)( 1 2x2)
=2
4 8
4
173
5
So the matrix representation of T(relative to BandC) is
MT
B;C=2
42 24 5 8
0 8 0 4
2 26 5 173
5
We know from Theorem KNSI [625] that the kernel of the linear transformation Tis isomorphic to the
null space of the matrix representation MT
B;Cand by studying the proof of Theorem KNSI [625] we learn
thatBis an isomorphism between these null spaces. Rather than trying to compute the kernel of Tusing
denitions and techniques from Chapter LT [515] we will instead analyze the null space of MT
B;Cusing
techniques from way back in Chapter V [97]. First row-reduce MT
B;C,
2
42 24 5 8
0 8 0 4
2 26 5 173
5RREF !2
4105
22
0101
2
0 0 0 03
5
Version 2.30
630 Section MR Matrix Representations
So, by Theorem BNS [160], a basis for N
MT
B;C
is
*8
>><
>>:2
664 5
2
0
1
03
775;2
664 2
1
2
0
13
7759
>>=
>>;+
We can now convert this basis of N
MT
B;C
into a basis ofK(T) by applying 1
Bto each element of the
basis,
1
B0
BB@2
664 5
2
0
1
03
7751
CCA= ( 5
2)1 2
1 1
+ 01 3
1 4
+ 11 2
0 2
+ 02 5
2 4
= 3
2 3
5
21
2
1
B0
BB@2
664 2
1
2
0
13
7751
CCA= ( 2)1 2
1 1
+ ( 1
2)1 3
1 4
+ 01 2
0 2
+ 12 5
2 4
= 1
2 1
21
20
So the set 3
2 3
5
21
2
; 1
2 1
21
20
is a basis forK(T) Just for fun, you might evaluate Twith each of these two basis vectors and verify that
the output is the zero polynomial (Exercise MR.C10 [635]).
An entirely similar result applies to the range of a linear transformation and the column space of a
matrix representation of the linear transformation.
Theorem RCSI
Range and Column Space Isomorphism
Suppose that T:U!Vis a linear transformation, Bis a basis for Uof sizen, andCis a basis for Vof
sizem. Then the range of Tis isomorphic to the column space of MT
B;C,
R(T)=C
MT
B;C
Proof To establish that two vector spaces are isomorphic, we must nd an isomorphism between them,
an invertible linear transformation (Denition IVS [586]). The range of the linear transformation T,R(T),
is a subspace of V, while the column space of the matrix representation, C
MT
B;C
is a subspace of Cm.
The function Cis dened as a function from VtoCm, but we can just as well employ the denition of
Cas a function from R(T) toC
MT
B;C
.
We must rst insure that if we choose an input for CfromR(T) that then the output will be an
element ofC
MT
B;C
. So suppose that v2R(T). Then there is a vector u2U, such that T(u) =v.
Consider
MT
B;CB(u) =C(T(u)) Theorem FTMR [617]
Version 2.30
Subsection MR.PMR Properties of Matrix Representations 631
=C(v) Denition RLT [563]
This says that C(v)2C
MT
B;C
, as desired.
The restriction in the size of the domain and codomain will not aect the fact that Cis a linear
transformation (Theorem VRLT [603]), nor will it aect the fact that Cis injective (Theorem VRI [607]).
Something must be done though to verify that Cis surjective. This all gets a bit confusing, since the
domain of our isomorphism is the range of the linear transformation, so think about your objects as you go.
To establish that Cis surjective, appeal to the denition of a surjective linear transformation (Denition
SLT [559]), and suppose that we have an element of the codomain, y2C
MT
B;C
Cmand we wish to
nd an element of the domain with yas its image. Since y2C
MT
B;C
, there exists a vector, x2Cn
withMT
B;Cx=y. We now show that the desired element of the domain is v= 1
C(y). First, verify that
v2R(T) by applying Ttou= 1
B(x),
T(u) =T
1
B(x)
= 1
C
MT
B;C
B
1
B(x)
Theorem FTMR [617]
= 1
C
MT
B;C(ICn(x))
Denition IVLT [579]
= 1
C
MT
B;Cx
Denition IDLT [579]
= 1
C(y) Denition CSM [271]
=v Substitution
Second, verify that the proposed isomorphism, C, takes vtoy,
C(v) =C
1
C(y)
Substitution
=ICm(y) Denition IVLT [579]
=y Denition IDLT [579]
WithCdemonstrated to be an injective and surjective linear transformation from R(T) toC
MT
B;C
,
Theorem ILTIS [582] tells us Cis invertible, and so by Denition IVS [586], we say R(T) andC
MT
B;C
are isomorphic.
Example RVMR
Range via matrix representation
In this example, we will recycle the linear transformation Tand the bases BandCof Example KVMR
[626] but now we will compute the range of T,
T:M22!P2; Ta b
c d
= (2a b+c 5d) + (a+ 4b+ 5b+ 2d)x+ (3a 2b+c 8d)x2
With bases BandC,
B=1 2
1 1
;1 3
1 4
;1 2
0 2
;2 5
2 4
C=
1 +x+x2;2 + 3x; 1 2x2
we obtain the matrix representation
MT
B;C=2
42 24 5 8
0 8 0 4
2 26 5 173
5
Version 2.30
632 Section MR Matrix Representations
We know from Theorem RCSI [628] that the range of the linear transformation Tis isomorphic to the
column space of the matrix representation MT
B;Cand by studying the proof of Theorem RCSI [628] we
learn thatCis an isomorphism between these subspaces. Notice that since the range is a subspace of
the codomain, we will employ Cas the isomorphism, rather than B, which was the correct choice for an
isomorphism between the null spaces of Example KVMR [626].
Rather than trying to compute the range of Tusing denitions and techniques from Chapter LT [515]
we will instead analyze the column space of MT
B;Cusing techniques from way back in Chapter M [207].
First row-reduce
MT
B;Ct
,
2
6642 0 2
24 8 26
5 0 5
8 4 173
775RREF !2
66410 1
01 25
4
0 0 0
0 0 03
775
Now employ Theorem CSRST [282] and Theorem BRS [280] (there are other methods we could choose
here to compute the column space, such as Theorem BCS [274]) to obtain the basis for C
MT
B;C
,
8
<
:2
41
0
13
5;2
40
1
25
43
59
=
;
We can now convert this basis of C
MT
B;C
into a basis ofR(T) by applying 1
Cto each element of the
basis,
1
C0
@2
41
0
13
51
A= (1 +x+x2) ( 1 2x2) = 2 +x+ 3x2
1
C0
@2
40
1
25
43
51
A= (2 + 3x) 25
4( 1 2x2) =33
4+ 3x+31
2x2
So the set
2 + 3x+ 3x2;33
4+ 3x+31
2x2
is a basis forR(T).
Theorem KNSI [625] and Theorem RCSI [628] can be viewed as further formal evidence for the Coor-
dinatization Principle [611], though they are not direct consequences.
Subsection IVLT
Invertible Linear Transformations
We have seen, both in theorems and in examples, that questions about linear transformations are often
equivalent to questions about matrices. It is the matrix representation of a linear transformation that
makes this idea precise. Here's our nal theorem that solidies this connection.
Theorem IMR
Invertible Matrix Representations
Suppose that T:U!Vis a linear transformation, Bis a basis for UandCis a basis for V. ThenTis an
Version 2.30
Subsection MR.IVLT Invertible Linear Transformations 633
invertible linear transformation if and only if the matrix representation of Trelative toBandC,MT
B;Cis
an invertible matrix. When Tis invertible,
MT 1
C;B=
MT
B;C 1
Proof (() SupposeTis invertible, so the inverse linear transformation T 1:V!Uexists (Denition
IVLT [579]). Both linear transformations have matrix representations relative to the bases of UandV,
namelyMT
B;CandMT 1
C;B(Denition MR [615]). Then
MT 1
C;BMT
B;C=MT 1T
B;B Theorem MRCLT [622]
=MIU
B;BDenition IVLT [579]
= [B(IU(u1))jB(IU(u2))j:::jB(IU(un))] Denition MR [615]
= [B(u1)jB(u2)jB(u3)j:::jB(un)] Denition IDLT [579]
= [e1je2je3j:::jen] Denition VR [603]
=In Denition IM [84]
and
MT
B;CMT 1
C;B=MTT 1
C;C Theorem MRCLT [622]
=MIV
C;CDenition IVLT [579]
= [C(IV(v1))jC(IV(v2))j:::jC(IV(vn))] Denition MR [615]
= [C(v1)jC(v2)jC(v3)j:::jC(vn)] Denition IDLT [579]
= [e1je2je3j:::jen] Denition VR [603]
=In Denition IM [84]
These two equations show that MT
B;CandMT 1
C;Bare inverse matrices (Denition MI [244]) and establish
that whenTis invertible, then MT 1
C;B=
MT
B;C 1
.
(() Suppose now that MT
B;Cis an invertible matrix and hence nonsingular (Theorem NI [261]). We
compute the nullity of T,
n(T) = dim (K(T)) Denition KLT [545]
= dim
N
MT
B;C
Theorem KNSI [625]
=n
MT
B;C
Denition NOM [397]
= 0 Theorem RNNM [399]
So the kernel of Tis trivial, and by Theorem KILT [548], Tis injective.
We now compute the rank of T,
r(T) = dim (R(T)) Denition RLT [563]
= dim
C
MT
B;C
Theorem RCSI [628]
=r
MT
B;C
Denition ROM [397]
= dim (V) Theorem RNNM [399]
Since the dimension of the range of Tequals the dimension of the codomain V, by Theorem EDYES [410],
R(T) =V. Which says that Tis surjective by Theorem RSLT [565].
Version 2.30
634 Section MR Matrix Representations
BecauseTis both injective and surjective, by Theorem ILTIS [582], Tis invertible.
By now, the connections between matrices and linear transformations should be starting to become
more transparent, and you may have already recognized the invertibility of a matrix as being tantamount
to the invertibility of the associated matrix representation. The next example shows how to apply this
theorem to the problem of actually building a formula for the inverse of an invertible linear transformation.
Example ILTVR
Inverse of a linear transformation via a representation
Consider the linear transformation
R:P3!M22; R
a+bx+cx2+x3
=a+b c+ 2d2a+ 3b 2c+ 3d
a+b+ 2d a+b+ 2c 5d
If we wish to quickly nd a formula for the inverse of R(presuming it exists), then choosing \nice" bases
will work best. So build a matrix representation of Rrelative to the bases BandC,
B=
1; x; x2; x3
C=1 0
0 0
;0 1
0 0
;0 0
1 0
;0 0
0 1
Then,
C(R(1)) =C1 2
1 1
=2
6641
2
1
13
775
C(R(x)) =C1 3
1 1
=2
6641
3
1
13
775
C
R
x2
=C 1 2
0 2
=2
664 1
2
0
23
775
C
R
x3
=C2 3
2 5
=2
6642
3
2
53
775
So a representation of Ris
MR
B;C=2
6641 1 1 2
2 3 2 3
1 1 0 2
1 1 2 53
775
The matrix MR
B;Cis invertible (as you can check) so we know for sure that Ris invertible by Theorem
IMR [630]. Furthermore,
MR 1
C;B=
MR
B;C 1=2
6641 1 1 2
2 3 2 3
1 1 0 2
1 1 2 53
775 1
=2
66420 7 2 3
8 3 1 1
1 0 1 0
6 2 1 13
775
Version 2.30
Subsection MR.IVLT Invertible Linear Transformations 635
We can use this representation of the inverse linear transformation, in concert with Theorem FTMR [617],
to determine an explicit formula for the inverse itself,
R 1a b
c d
= 1
B
MR 1
C;BCa b
c d
Theorem FTMR [617]
= 1
B
MR
B;C 1Ca b
c d
Theorem IMR [630]
= 1
B0
BB@
MR
B;C 12
664a
b
c
d3
7751
CCADenition VR [603]
= 1
B0
BB@2
66420 7 2 3
8 3 1 1
1 0 1 0
6 2 1 13
7752
664a
b
c
d3
7751
CCADenition MI [244]
= 1
B0
BB@2
66420a 7b 2c+ 3d
8a+ 3b+c d
a+c
6a+ 2b+c d3
7751
CCADenition MVP [223]
= (20a 7b 2c+ 3d) + ( 8a+ 3b+c d)x
+ ( a+c)x2+ ( 6a+ 2b+c d)x3Denition VR [603]
You might look back at Example AIVLT [579], where we rst witnessed the inverse of a linear trans-
formation and recognize that the inverse ( S) was built from using the method of Example ILTVR [632]
with a matrix representation of T.
Theorem IMILT
Invertible Matrices, Invertible Linear Transformation
Suppose that Ais a square matrix of size nandT:Cn!Cnis the linear transformation dened by
T(x) =Ax. ThenAis invertible matrix if and only if Tis an invertible linear transformation.
Proof Choose bases B=C=fe1;e2;e3; :::; engconsisting of the standard unit vectors as a basis of
Cn(Theorem SUVB [371]) and build a matrix representation of Trelative toBandC. Then
C(T(ei)) =C(Aei)
=C(Ai)
=Ai
So then the matrix representation of T, relative to BandC, is simplyMT
B;C=A. with this observation,
the proof becomes a specialization of Theorem IMR [630],
Tis invertible()MT
B;Cis invertible()Ais invertible
This theorem may seem gratuitous. Why state such a special case of Theorem IMR [630]? Because
it adds another condition to our NMEx series of theorems, and in some ways it is the most fundamental
expression of what it means for a matrix to be nonsingular | the associated linear transformation is
invertible. This is our nal update.
Theorem NME9
Nonsingular Matrix Equivalences, Round 9
Suppose that Ais a square matrix of size n. The following are equivalent.
Version 2.30
636 Section MR Matrix Representations
1.Ais nonsingular.
2.Arow-reduces to the identity matrix.
3. The null space of Acontains only the zero vector, N(A) =f0g.
4. The linear system LS(A;b) has a unique solution for every possible choice of b.
5. The columns of Aare a linearly independent set.
6.Ais invertible.
7. The column space of AisCn,C(A) =Cn.
8. The columns of Aare a basis for Cn.
9. The rank of Aisn,r(A) =n.
10. The nullity of Ais zero,n(A) = 0.
11. The determinant of Ais nonzero, det ( A)6= 0.
12.= 0 is not an eigenvalue of A.
13. The linear transformation T:Cn!Cndened byT(x) =Axis invertible.
Proof By Theorem IMILT [633] the new addition to this list is equivalent to the statement that Ais
invertible so we can expand Theorem NME8 [480].
Subsection READ
Reading Questions
1. Why does Theorem FTMR [617] deserve the moniker \fundamental"?
2. Find the matrix representation, MT
B;Cof the linear transformation
T:C2!C2; Tx1
x2
=2x1 x2
3x1+ 2x2
relative to the bases
B=2
3
; 1
2
C=1
0
;1
1
3. What is the second \surprise," and why is it surprising?
Version 2.30
Subsection MR.EXC Exercises 637
Subsection EXC
Exercises
C10 Example KVMR [626] concludes with a basis for the kernel of the linear transformation T. Compute
the value of Tfor each of these two basis vectors. Did you get what you expected?
Contributed by Robert Beezer
C20 Compute the matrix representation of Trelative to the bases BandC.
T:P3!C3; T
a+bx+cx2+dx3
=2
42a 3b+ 4c 2d
a+b c+d
3a+ 2c 3d3
5
B=
1; x; x2; x3
C=8
<
:2
41
0
03
5;2
41
1
03
5;2
41
1
13
59
=
;
Contributed by Robert Beezer Solution [638]
C21 Find a matrix representation of the linear transformation Trelative to the bases BandC.
T:P2!C2; T (p(x)) =p(1)
p(3)
B=
2 5x+x2;1 +x x2; x2
C=3
4
;2
3
Contributed by Robert Beezer Solution [638]
C22 LetS22be the vector space of 2 2 symmetric matrices. Build the matrix representation of the
linear transformation T:P2!S22relative to the bases BandCand then use this matrix representation
to compute T
3 + 5x 2x2
.
B=
1;1 +x;1 +x+x2
C=1 0
0 0
;0 1
1 0
;0 0
0 1
T
a+bx+cx2
=2a b+c a + 3b c
a+ 3b c a c
Contributed by Robert Beezer Solution [638]
C25 Use a matrix representation to determine if the linear transformation T:P3!M22surjective.
T
a+bx+cx2+dx3
= a+ 4b+c+ 2d4a b+ 6c d
a+ 5b 2c+ 2d a + 2c+ 5d
Contributed by Robert Beezer Solution [639]
C30 Find bases for the kernel and range of the linear transformation Sbelow.
S:M22!P2; Sa b
c d
= (a+ 2b+ 5c 4d) + (3a b+ 8c+ 2d)x+ (a+b+ 4c 2d)x2
Version 2.30
638 Section MR Matrix Representations
Contributed by Robert Beezer Solution [640]
C40 LetS22be the set of 22 symmetric matrices. Verify that the linear transformation Ris invertible
and ndR 1.
R:S22!P2; Ra b
b c
= (a b) + (2a 3b 2c)x+ (a b+c)x2
Contributed by Robert Beezer Solution [640]
C41 Prove that the linear transformation Sis invertible. Then nd a formula for the inverse linear
transformation, S 1, by employing a matrix inverse.
S:P1!M1;2; S (a+bx) =
3a+b2a+b
Contributed by Robert Beezer Solution [641]
C42 The linear transformation R:M12!M21is invertible. Use a matrix representation to determine a
formula for the inverse linear transformation R 1:M21!M12.
R
a b
=a+ 3b
4a+ 11b
Contributed by Robert Beezer Solution [642]
C50 Use a matrix representation to nd a basis for the range of the linear transformation L.
L:M22!P2; Ta b
c d
= (a+ 2b+ 4c+d) + (3a+c 2d)x+ ( a+b+ 3c+ 3d)x2
Contributed by Robert Beezer Solution [642]
C51 Use a matrix representation to nd a basis for the kernel of the linear transformation L.
L:M22!P2; Ta b
c d
= (a+ 2b+ 4c+d) + (3a+c 2d)x+ ( a+b+ 3c+ 3d)x2
Contributed by Robert Beezer
C52 Find a basis for the kernel of the linear transformation T:P2!M22.
T
a+bx+cx2
=a+ 2b 2c 2a+ 2b
a+b 4c3a+ 2b+ 2c
Contributed by Robert Beezer Solution [643]
M20 The linear transformation Dperforms dierentiation on polynomials. Use a matrix representation
ofDto nd the rank and nullity of D.
D:Pn!Pn; D (p(x)) =p0(x)
Contributed by Robert Beezer Solution [644]
Version 2.30
Subsection MR.EXC Exercises 639
M60 SupposeUandVare vector spaces and dene a function Z:U!VbyT(u) =0Vfor every
u2U. Then Exercise IVLT.M60 [594] asks you to formulate the theorem: Zis invertible if and only if
U=f0UgandV=f0Vg. What would a matrix representation of Zlook like in this case? How does
Theorem IMR [630] read in this case?
Contributed by Robert Beezer
M80 In light of Theorem KNSI [625] and Theorem MRCLT [622], write a short comparison of Exercise
MM.T40 [238] with Exercise ILT.T15 [554].
Contributed by Robert Beezer
M81 In light of Theorem RCSI [628] and Theorem MRCLT [622], write a short comparison of Exercise
CRS.T40 [287] with Exercise SLT.T15 [572].
Contributed by Robert Beezer
M82 In light of Theorem MRCLT [622] and Theorem IMR [630], write a short comparison of Theorem
SS [250] and Theorem ICLT [585].
Contributed by Robert Beezer
M83 In light of Theorem MRCLT [622] and Theorem IMR [630], write a short comparison of Theorem
NPNT [259] and Exercise IVLT.T40 [595].
Contributed by Robert Beezer
T20 Construct a new solution to Exercise B.T50 [383] along the following outline. From the nnmatrix
A, construct the linear transformation T:Cn!Cn,T(x) =Ax. Use Theorem NI [261], Theorem IMILT
[633] and Theorem ILTIS [582] to translate between the nonsingularity of Aand the surjectivity/injectivity
ofT. Then apply Theorem ILTB [550] and Theorem SLTB [568] to connect these properties with bases.
Contributed by Robert Beezer Solution [644]
T60 Create an entirely dierent proof of Theorem IMILT [633] that relies on Denition IVLT [579] to
establish the invertibility of T, and that relies on Denition MI [244] to establish the invertibility of A.
Contributed by Robert Beezer
T80 Suppose that T:U!VandS:V!Ware linear transformations, and that B,CandDare
bases forU,V, andW. Using only Denition MR [615] dene matrix representations for TandS. Using
these two denitions, and Denition MR [615], derive a matrix representation for the composition STin
terms of the entries of the matrices MT
B;CandMS
C;D. Explain how you would use this result to motivate a
denition for matrix multiplication that is strikingly similar to Theorem EMP [227].
Contributed by Robert Beezer Solution [645]
Version 2.30
640 Section MR Matrix Representations
Subsection SOL
Solutions
C20 Contributed by Robert Beezer Statement [635]
Apply Denition MR [615],
C(T(1)) =C0
@2
42
1
33
51
A=C0
@12
41
0
03
5+ ( 2)2
41
1
03
5+ 32
41
1
13
51
A=2
41
2
33
5
C(T(x)) =C0
@2
4 3
1
03
51
A=C0
@( 4)2
41
0
03
5+ 12
41
1
03
5+ 02
41
1
13
51
A=2
4 4
1
03
5
C
T
x2
=C0
@2
44
1
23
51
A=C0
@52
41
0
03
5+ ( 3)2
41
1
03
5+ 22
41
1
13
51
A=2
45
3
23
5
C
T
x3
=C0
@2
4 2
1
33
51
A=C0
@( 3)2
41
0
03
5+ 42
41
1
03
5+ ( 3)2
41
1
13
51
A=2
4 3
4
33
5
These four vectors are the columns of the matrix representation,
MT
B;C=2
41 4 5 3
2 1 3 4
3 0 2 33
5
C21 Contributed by Robert Beezer Statement [635]
Applying Denition MR [615],
C
T
2 5x+x2
=C 2
4
=C
23
4
+ ( 4)2
3
=2
4
C
T
1 +x x2
=C1
5
=C
133
4
+ ( 19)2
3
=13
19
C
T
x2
=C1
9
=C
( 15)3
4
+ 232
3
= 15
23
So the resulting matrix representation is
MT
B;C=2 13 15
4 19 23
C22 Contributed by Robert Beezer Statement [635]
Input toTthe vectors of the basis Band coordinatize the outputs relative to C,
C(T(1)) =C2 1
1 1
=C
21 0
0 0
+ 10 1
1 0
+ 10 0
0 1
=2
42
1
13
5
C(T(1 +x)) =C1 4
4 1
=C
11 0
0 0
+ 40 1
1 0
+ 10 0
0 1
=2
41
4
13
5
Version 2.30
Subsection MR.SOL Solutions 641
C
T
1 +x+x2
=C2 3
3 0
=C
21 0
0 0
+ 30 1
1 0
+ 00 0
0 1
=2
42
3
03
5
Applying Denition MR [615] we have the matrix representation
MT
B;C=2
42 1 2
1 4 3
1 1 03
5
To compute T
3 + 5x 2x2
employ Theorem FTMR [617],
T
3 + 5x 2x2
= 1
C
MT
B;CB
3 + 5x 2x2
= 1
C
MT
B;CB
( 2)(1) + 7(1 + x) + ( 2)(1 +x+x2)
= 1
C0
@2
42 1 2
1 4 3
1 1 03
52
4 2
7
23
51
A
= 1
C0
@2
4 1
20
53
51
A
= ( 1)1 0
0 0
+ 200 1
1 0
+ 50 0
0 1
= 1 20
20 5
You can, of course, check your answer by evaluating T
3 + 5x 2x2
directly.
C25 Contributed by Robert Beezer Statement [635]
Choose bases BandCfor the matrix representation,
B=
1; x; x2; x3
C=1 0
0 0
;0 1
0 0
;0 0
1 0
;0 0
0 1
Input toTthe vectors of the basis Band coordinatize the outputs relative to C,
C(T(1)) =C 1 4
1 1
=C
( 1)1 0
0 0
+ 40 1
0 0
+ 10 0
1 0
+ 10 0
0 1
=2
664 1
4
1
13
775
C(T(x)) =C4 1
5 0
=C
41 0
0 0
+ ( 1)0 1
0 0
+ 50 0
1 0
+ 00 0
0 1
=2
6644
1
5
03
775
C
T
x2
=C1 6
2 2
=C
11 0
0 0
+ 60 1
0 0
+ ( 2)0 0
1 0
+ 20 0
0 1
=2
6641
6
2
23
775
C
T
x3
=C2 1
2 5
=C
21 0
0 0
+ ( 1)0 1
0 0
+ 20 0
1 0
+ 50 0
0 1
=2
6642
1
2
53
775
Version 2.30
642 Section MR Matrix Representations
Applying Denition MR [615] we have the matrix representation
MT
B;C=2
664 1 4 1 2
4 1 6 1
1 5 2 2
1 0 2 53
775
Properties of this matrix representation will translate to properties of the linear transformation The matrix
representation is nonsingular since it row-reduces to the identity matrix (Theorem NMRRI [84]) and
therefore has a column space equal to C4(Theorem CNMB [376]). The column space of the matrix
representation is isomorphic to the range of the linear transformation (Theorem RCSI [628]). So the range
ofThas dimension 4, equal to the dimension of the codomain M22. By Theorem ROSLT [588], Tis
surjective.
C30 Contributed by Robert Beezer Statement [635]
These subspaces will be easiest to construct by analyzing a matrix representation of S. Since we can
use any matrix representation, we might as well use natural bases that allow us to construct the matrix
representation quickly and easily,
B=1 0
0 0
;0 1
0 0
;0 0
1 0
;0 0
0 1
C=
1; x; x2
then we can practically build the matrix representation on sight,
MS
B;C=2
41 2 5 4
3 1 8 2
1 1 4 23
5
The rst step is to nd bases for the null space and column space of the matrix representation. Row-
reducing the matrix representation we nd,
2
410 3 0
011 2
0 0 0 03
5
So by Theorem BNS [160] and Theorem BCS [274], we have
N
MS
B;C
=*8
>><
>>:2
664 3
1
1
03
775;2
6640
2
0
13
7759
>>=
>>;+
C
MS
B;C
=*8
<
:2
41
3
13
5;2
42
1
13
59
=
;+
Now, the proofs of Theorem KNSI [625] and Theorem RCSI [628] tell us that we can apply 1
Band 1
C
(respectively) to \un-coordinatize" and get bases for the kernel and range of the linear transformation S
itself,
K(S) = 3 1
1 0
;0 2
0 1
R(S) =
1 + 3x+x2;2 x+x2
C40 Contributed by Robert Beezer Statement [636]
The analysis of Rwill be easiest if we analyze a matrix representation of R. Since we can use any matrix
representation, we might as well use natural bases that allow us to construct the matrix representation
quickly and easily,
B=1 0
0 0
;0 1
1 0
;0 0
0 1
C=
1; x; x2
Version 2.30
Subsection MR.SOL Solutions 643
then we can practically build the matrix representation on sight,
MR
B;C=2
41 1 0
2 3 2
1 1 13
5
This matrix representation is invertible (it has a nonzero determinant of 1, Theorem SMZD [445], The-
orem NI [261]) so Theorem IMR [630] tells us that the linear transformation Ris also invertible. To nd
a formula for R 1we compute,
R 1
a+bx+cx2
= 1
B
MR 1
C;BC
a+bx+cx2
Theorem FTMR [617]
= 1
B
MR
B;C 1C
a+bx+cx2
Theorem IMR [630]
= 1
B0
@
MR
B;C 12
4a
b
c3
51
A Denition VR [603]
= 1
B0
@2
45 1 2
4 1 2
1 0 13
52
4a
b
c3
51
A Denition MI [244]
= 1
B0
@2
45a b 2c
4a b 2c
a+c3
51
A Denition MVP [223]
=5a b 2c4a b 2c
4a b 2c a+c
Denition VR [603]
C41 Contributed by Robert Beezer Statement [636]
First, build a matrix representation of S(Denition MR [615]). We are free to choose whatever bases we
wish, so we should choose ones that are easy to work with, such as
B=f1; xg
C=
1 0
;
0 1
The resulting matrix representation is then
MT
B;C=3 1
2 1
this matrix is invertible, since it has a nonzero determinant, so by Theorem IMR [630] the linear transfor-
mationSis invertible. We can use the matrix inverse and Theorem IMR [630] to nd a formula for the
inverse linear transformation,
S 1
a b
= 1
B
MS 1
C;BC
a b
Theorem FTMR [617]
= 1
B
MS
B;C 1C
a b
Theorem IMR [630]
= 1
B
MS
B;C 1a
b
Denition VR [603]
= 1
B 3 1
2 1 1a
b!
= 1
B1 1
2 3a
b
Denition MI [244]
Version 2.30
644 Section MR Matrix Representations
= 1
Ba b
2a+ 3b
Denition MVP [223]
= (a b) + ( 2a+ 3b)x Denition VR [603]
C42 Contributed by Robert Beezer Statement [636]
Choose bases BandCforM12andM21(respectively),
B=
1 0
;
0 1
C=1
0
;0
1
The resulting matrix representation is
MR
B;C=1 3
4 11
This matrix is invertible (its determinant is nonzero, Theorem SMZD [445]), so by Theorem IMR [630],
we can compute the matrix representation of R 1with a matrix inverse (Theorem TTMI [246]),
MR 1
C;B=1 3
4 11 1
= 11 3
4 1
To obtain a general formula for R 1, use Theorem FTMR [617],
R 1x
y
= 1
B
MR 1
C;BCx
y
= 1
B 11 3
4 1x
y
= 1
B 11x+ 3y
4x y
=
11x+ 3y4x y
C50 Contributed by Robert Beezer Statement [636]
As usual, build any matrix representation of L, most likely using a \nice" bases, such as
B=1 0
0 0
;0 1
0 0
;0 0
1 0
;0 0
0 1
C=
1; x; x2
Then the matrix representation (Denition MR [615]) is,
ML
B;C=2
41 2 4 1
3 0 1 2
1 1 3 33
5
Theorem RCSI [628] tells us that we can compute the column space of the matrix representation, then use
the isomorphism 1
Cto convert the column space of the matrix representation into the range of the linear
transformation. So we rst analyze the matrix representation,
2
41 2 4 1
3 0 1 2
1 1 3 33
5RREF !2
410 0 1
010 1
0 0 1 13
5
With three nonzero rows in the reduced row-echelon form of the matrix, we know the column space has
dimension 3. Since P2has dimension 3 (Theorem DP [395]), the range must be all of P2. So anybasis of
P2would suce as a basis for the range. For instance, Citself would be a correct answer.
Version 2.30
Subsection MR.SOL Solutions 645
A more laborious approach would be to use Theorem BCS [274] and choose the rst three columns
of the matrix representation as a basis for the range of the matrix representation. These could then be
\un-coordinatized" with 1
Cto yield a (\not nice") basis for P2.
C52 Contributed by Robert Beezer Statement [636]
Choose bases BandCfor the matrix representation,
B=
1; x; x2
C=1 0
0 0
;0 1
0 0
;0 0
1 0
;0 0
0 1
Input toTthe vectors of the basis Band coordinatize the outputs relative to C,
C(T(1)) =C1 2
1 3
=C
11 0
0 0
+ 20 1
0 0
+ ( 1)0 0
1 0
+ 30 0
0 1
=2
6641
2
1
33
775
C(T(x)) =C2 2
1 2
=C
21 0
0 0
+ 20 1
0 0
+ 10 0
1 0
+ 20 0
0 1
=2
6642
2
1
23
775
C
T
x2
=C 2 0
4 2
=C
( 2)1 0
0 0
+ 00 1
0 0
+ ( 4)0 0
1 0
+ 20 0
0 1
=2
664 2
0
4
23
775
Applying Denition MR [615] we have the matrix representation
MT
B;C=2
6641 2 2
2 2 0
1 1 4
3 2 23
775
The null space of the matrix representation is isomorphic (via B) to the kernel of the linear transformation
(Theorem KNSI [625]). So we compute the null space of the matrix representation by rst row-reducing
the matrix to,2
66410 2
01 2
0 0 0
0 0 03
775
Employing Theorem BNS [160] we have
N
MT
B;C
=*8
<
:2
4 2
2
13
59
=
;+
We only need to uncoordinatize this one basis vector to get a basis for K(T),
K(T) =*8
<
: 1
B0
@2
4 2
2
13
51
A9
=
;+
=
2 + 2x+x2
Version 2.30
646 Section MR Matrix Representations
M20 Contributed by Robert Beezer Statement [636]
Build a matrix representation (Denition MR [615]) with the set
B=
1; x; x2; :::; xn
employed as a basis of both the domain and codomain. Then
B(D(1)) =B(0) =2
666666640
0
0
...
0
03
77777775B(D(x)) =B(1) =2
666666641
0
0
...
0
03
77777775
B
D
x2
=B(2x) =2
666666640
2
0
...
0
03
77777775B
D
x3
=B
3x2
=2
666666640
0
3
...
0
03
77777775
...
B(D(xn)) =B
nxn 1
=2
666666640
0
0
...
n
03
77777775
and the resulting matrix representation is
MD
B;B=2
666666640 1 0 0 ::: 0 0
0 0 2 0 ::: 0 0
0 0 0 3 ::: 0 0
.........
0 0 0 0 ::: 0n
0 0 0 0 ::: 0 03
77777775
This (n+ 1)(n+ 1) matrix is very close to being in reduced row-echelon form. Multiply row iby1
i, for
1in, to convert it to reduced row-echelon form. From this we can see that matrix representation
MD
B;Bhas ranknand nullity 1. Applying Theorem RCSI [628] and Theorem KNSI [625] tells us that the
linear transformation Dwill have the same values for the rank and nullity, as well.
T20 Contributed by Robert Beezer Statement [637]
Given the nonsingular nnmatrixA, create the linear transformation T:Cn!Cndened byT(x) =Ax.
Then
Anonsingular()Ainvertible Theorem NI [261]
()Tinvertible Theorem IMILT [633]
()Tinjective and surjective Theorem ILTIS [582]
Version 2.30
Subsection MR.SOL Solutions 647
()Clinearly independent, and Theorem ILTB [550]
Cspans CnTheorem SLTB [568]
()Cbasis for CnDenition B [371]
T80 Contributed by Robert Beezer Statement [637]
Suppose that B=fu1;u2;u3; :::; umg,C=fv1;v2;v3; :::; vngandD=fw1;w2;w3; :::; wpg. For
convenience, set M=MT
B;C,mij= [M]ij, 1in, 1jm, and similarly, set N=MS
C;D,nij= [N]ij,
1ip, 1jn. We want to learn about the matrix representation of ST:V!Wrelative toB
andD. We will examine a single (generic) entry of this representation.
MST
B;D
ij= [D((ST) (uj))]iDenition MR [615]
= [D(S(T(uj)))]iDenition LTC [532]
="
D
S nX
k=1mkjvk!!#
iDenition MR [615]
="
D nX
k=1mkjS(vk)!#
iTheorem LTLC [525]
="
D nX
k=1mkjpX
`=1n`kw`!#
iDenition MR [615]
="
D nX
k=1pX
`=1mkjn`kw`!#
iProperty DVA [318]
="
D pX
`=1nX
k=1mkjn`kw`!#
iProperty C [317]
="
D pX
`=1 nX
k=1mkjn`k!
w`!#
iProperty DSA [318]
=nX
k=1mkjnik Denition VR [603]
=nX
k=1nikmkj Property CMCN [758]
=nX
k=1
MS
C;D
ik
MT
B;C
kjProperty CMCN [758]
This formula for the entry of a matrix should remind you of Theorem EMP [227]. However, while the
theorem presumed we knew how to multiply matrices, the solution before us never uses any understanding
of matrix products. It uses the denitions of vector and matrix representations, properties of linear
transformations and vector spaces. So if we began a course by rst discussing vector space, and then
linear transformations between vector spaces, we could carry matrix representations into a motivation for
a denition of matrix multiplication that is grounded in function composition. That is worth saying again
| a denition of matrix representations of linear transformations results in a matrix product being the
representation of a composition of linear transformations.
This exercise is meant to explain why many authors take the formula in Theorem EMP [227] as their
denition of matrix multiplication, and why it is a natural choice when the proper motivation is in place.
If we rst dened matrix multiplication in the style of Theorem EMP [227], then the above argument,
Version 2.30
648 Section MR Matrix Representations
followed by a simple application of the denition of matrix equality (Denition ME [207]), would yield
Theorem MRCLT [622].
Version 2.30
Section CB Change of Basis 649
Section CB
Change of Basis
We have seen in Section MR [615] that a linear transformation can be represented by a matrix, once we
pick bases for the domain and codomain. How does the matrix representation change if we choose dierent
bases? Which bases lead to especially nice representations? From the innite possibilities, what is the
best possible representation? This section will begin to answer these questions. But rst we need to dene
eigenvalues for linear transformations and the change-of-basis matrix.
Subsection EELT
Eigenvalues and Eigenvectors of Linear Transformations
We now dene the notion of an eigenvalue and eigenvector of a linear transformation. It should not
be too surprising, especially if you remind yourself of the close relationship between matrices and linear
transformations.
Denition EELT
Eigenvalue and Eigenvector of a Linear Transformation
Suppose that T:V!Vis a linear transformation. Then a nonzero vector v2Vis aneigenvector ofT
for the eigenvalue ifT(v) =v. 4
We will see shortly the best method for computing the eigenvalues and eigenvectors of a linear trans-
formation, but for now, here are some examples to verify that such things really do exist.
Example ELTBM
Eigenvectors of linear transformation between matrices
Consider the linear transformation T:M22!M22dened by
Ta b
c d
= 17a+ 11b+ 8c 11d 57a+ 35b+ 24c 33d
14a+ 10b+ 6c 10d 41a+ 25b+ 16c 23d
and the vectors
x1=0 1
0 1
x2=1 1
1 0
x3=1 3
2 3
x4=2 6
1 4
Then compute
T(x1) =T0 1
0 1
=0 2
0 2
= 2x1
T(x2) =T1 1
1 0
=2 2
2 0
= 2x2
T(x3) =T1 3
2 3
= 1 3
2 3
= ( 1)x3
T(x4) =T2 6
1 4
= 4 12
2 8
= ( 2)x4
Sox1,x2,x3,x4are eigenvectors of Twith eigenvalues (respectively) 1= 2,2= 2,3= 1,4= 2.
Version 2.30
650 Section CB Change of Basis
Here's another.
Example ELTBP
Eigenvectors of linear transformation between polynomials
Consider the linear transformation R:P2!P2dened by
R
a+bx+cx2
= (15a+ 8b 4c) + ( 12a 6b+ 3c)x+ (24a+ 14b 7c)x2
and the vectors
w1= 1 x+x2w2=x+ 2x2w3= 1 + 4x2
Then compute
R(w1) =R
1 x+x2
= 3 3x+ 3x2= 3w1
R(w2) =R
x+ 2x2
= 0 + 0x+ 0x2= 0w2
R(w3) =R
1 + 4x2
= 1 4x2= ( 1)w3
Sow1,w2,w3are eigenvectors of Rwith eigenvalues (respectively) 1= 3,2= 0,3= 1. Notice how
the eigenvalue 2= 0 indicates that the eigenvector w2is a non-trivial element of the kernel of R, and
thereforeRis not injective (Exercise CB.T15 [669]).
Of course, these examples are meant only to illustrate the denition of eigenvectors and eigenvalues for
linear transformations, and therefore beg the question, \How would I ndeigenvectors?" We'll have an
answer before we nish this section. We need one more construction rst.
Subsection CBM
Change-of-Basis Matrix
Given a vector space, we know we can usually nd many dierent bases for the vector space, some nice,
some nasty. If we choose a single vector from this vector space, we can build many dierent representa-
tions of the vector by constructing the representations relative to dierent bases. How are these dierent
representations related to each other? A change-of-basis matrix answers this question.
Denition CBM
Change-of-Basis Matrix
Suppose that Vis a vector space, and IV:V!Vis the identity linear transformation on V. Let
B=fv1;v2;v3; :::; vngandCbe two bases of V. Then the change-of-basis matrix fromBtoCis
the matrix representation of IVrelative toBandC,
CB;C=MIV
B;C
= [C(IV(v1))jC(IV(v2))jC(IV(v3))j:::jC(IV(vn))]
= [C(v1)jC(v2)jC(v3)j:::jC(vn)]
4
Notice that this denition is primarily about a single vector space ( V) and two bases of V(B,C). The
linear transformation ( IV) is necessary but not critical. As you might expect, this matrix has something
to do with changing bases. Here is the theorem that gives the matrix its name (not the other way around).
Version 2.30
Subsection CB.CBM Change-of-Basis Matrix 651
Theorem CB
Change-of-Basis
Suppose that vis a vector in the vector space VandBandCare bases of V. Then
C(v) =CB;CB(v)
Proof
C(v) =C(IV(v)) Denition IDLT [579]
=MIV
B;CB(v) Theorem FTMR [617]
=CB;CB(v) Denition CBM [648]
So the change-of-basis matrix can be used with matrix multiplication to convert a vector representation
of a vector ( v) relative to one basis ( B(v)) to a representation of the same vector relative to a second
basis (C(v)).
Theorem ICBM
Inverse of Change-of-Basis Matrix
Suppose that Vis a vector space, and BandCare bases of V. Then the change-of-basis matrix CB;Cis
nonsingular and
C 1
B;C=CC;B
Proof The linear transformation IV:V!Vis invertible, and its inverse is itself, IV(check this!). So
by Theorem IMR [630], the matrix MIV
B;C=CB;Cis invertible. Theorem NI [261] says an invertible matrix
is nonsingular.
Then
C 1
B;C=
MIV
B;C 1
Denition CBM [648]
=MI 1
V
C;BTheorem IMR [630]
=MIV
C;BDenition IDLT [579]
=CC;B Denition CBM [648]
Example CBP
Change of basis with polynomials
The vector space P4(Example VSP [319]) has two nice bases (Example BP [372]),
B=
1;x;x2;x3;x4
C=
1;1 +x;1 +x+x2;1 +x+x2+x3;1 +x+x2+x3+x4
To build the change-of-basis matrix between BandC, we must rst build a vector representation of each
vector inBrelative toC,
C(1) =C((1) (1)) =2
666641
0
0
0
03
77775
Version 2.30
652 Section CB Change of Basis
C(x) =C(( 1) (1) + (1) (1 + x)) =2
66664 1
1
0
0
03
77775
C
x2
=C
( 1) (1 +x) + (1)
1 +x+x2
=2
666640
1
1
0
03
77775
C
x3
=C
( 1)
1 +x+x2
+ (1)
1 +x+x2+x3
=2
666640
0
1
1
03
77775
C
x4
=C
( 1)
1 +x+x2+x3
+ (1)
1 +x+x2+x3+x4
=2
666640
0
0
1
13
77775
Then we package up these vectors as the columns of a matrix,
CB;C=2
666641 1 0 0 0
0 1 1 0 0
0 0 1 1 0
0 0 0 1 1
0 0 0 0 13
77775
Now, to illustrate Theorem CB [649], consider the vector u= 5 3x+ 2x2+ 8x3 3x4. We can build the
representation of urelative toBeasily,
B(u) =B
5 3x+ 2x2+ 8x3 3x4
=2
666645
3
2
8
33
77775
Applying Theorem CB [649], we obtain a second representation of u, but now relative to C,
C(u) =CB;CB(u) Theorem CB [649]
=2
666641 1 0 0 0
0 1 1 0 0
0 0 1 1 0
0 0 0 1 1
0 0 0 0 13
777752
666645
3
2
8
33
77775
=2
666648
5
6
11
33
77775Denition MVP [223]
Version 2.30
Subsection CB.CBM Change-of-Basis Matrix 653
We can check our work by unraveling this second representation,
u= 1
C(C(u)) Denition IVLT [579]
= 1
C0
BBBB@2
666648
5
6
11
33
777751
CCCCA
= 8(1) + ( 5)(1 +x) + ( 6)(1 +x+x2)
+ (11)(1 + x+x2+x3) + ( 3)(1 +x+x2+x3+x4) Denition VR [603]
= 5 3x+ 2x2+ 8x3 3x4
The change-of-basis matrix from CtoBis actually easier to build. Grab each vector in the basis Cand
form its representation relative to B
B(1) =B((1)1) =2
666641
0
0
0
03
77775
B(1 +x) =B((1)1 + (1)x) =2
666641
1
0
0
03
77775
B
1 +x+x2
=B
(1)1 + (1)x+ (1)x2
=2
666641
1
1
0
03
77775
B
1 +x+x2+x3
=B
(1)1 + (1)x+ (1)x2+ (1)x3
=2
666641
1
1
1
03
77775
B
1 +x+x2+x3+x4
=B
(1)1 + (1)x+ (1)x2+ (1)x3+ (1)x4
=2
666641
1
1
1
13
77775
Then we package up these vectors as the columns of a matrix,
CC;B=2
666641 1 1 1 1
0 1 1 1 1
0 0 1 1 1
0 0 0 1 1
0 0 0 0 13
77775
Version 2.30
654 Section CB Change of Basis
We formed two representations of the vector uabove, so we can again provide a check on our computations
by converting from the representation of urelative toCto the representation of urelative toB,
B(u) =CC;BC(u) Theorem CB [649]
=2
666641 1 1 1 1
0 1 1 1 1
0 0 1 1 1
0 0 0 1 1
0 0 0 0 13
777752
666648
5
6
11
33
77775
=2
666645
3
2
8
33
77775Denition MVP [223]
One more computation that is either a check on our work, or an illustration of a theorem. The two change-
of-basis matrices, CB;CandCC;B, should be inverses of each other, according to Theorem ICBM [649].
Here we go,
CB;CCC;B=2
666641 1 0 0 0
0 1 1 0 0
0 0 1 1 0
0 0 0 1 1
0 0 0 0 13
777752
666641 1 1 1 1
0 1 1 1 1
0 0 1 1 1
0 0 0 1 1
0 0 0 0 13
77775=2
666641 0 0 0 0
0 1 0 0 0
0 0 1 0 0
0 0 0 1 0
0 0 0 0 13
77775
The computations of the previous example are not meant to present any labor-saving devices, but
instead are meant to illustrate the utility of the change-of-basis matrix. However, you might have noticed
thatCC;Bwas easier to compute than CB;C. If you needed CB;C, then you could rst compute CC;Band
then compute its inverse, which by Theorem ICBM [649], would equal CB;C.
Here's another illustrative example. We have been concentrating on working with abstract vector
spaces, but all of our theorems and techniques apply just as well to Cm, the vector space of column
vectors. We only need to use more complicated bases than the standard unit vectors (Theorem SUVB
[371]) to make things interesting.
Example CBCV
Change of basis with column vectors
For the vector space C4we have the two bases,
B=8
>><
>>:2
6641
2
1
23
775;2
664 1
3
1
13
775;2
6642
3
3
43
775;2
664 1
3
3
03
7759
>>=
>>;C=8
>><
>>:2
6641
6
4
13
775;2
664 4
8
5
83
775;2
664 5
13
2
93
775;2
6643
7
3
63
7759
>>=
>>;
The change-of-basis matrix from BtoCrequires writing each vector of Bas a linear combination the
vectors inC,
C0
BB@2
6641
2
1
23
7751
CCA=C0
BB@(1)2
6641
6
4
13
775+ ( 2)2
664 4
8
5
83
775+ (1)2
664 5
13
2
93
775+ ( 1)2
6643
7
3
63
7751
CCA=2
6641
2
1
13
775
Version 2.30
Subsection CB.CBM Change-of-Basis Matrix 655
C0
BB@2
664 1
3
1
13
7751
CCA=C0
BB@(2)2
6641
6
4
13
775+ ( 3)2
664 4
8
5
83
775+ (3)2
664 5
13
2
93
775+ (0)2
6643
7
3
63
7751
CCA=2
6642
3
3
03
775
C0
BB@2
6642
3
3
43
7751
CCA=C0
BB@(1)2
6641
6
4
13
775+ ( 3)2
664 4
8
5
83
775+ (1)2
664 5
13
2
93
775+ ( 2)2
6643
7
3
63
7751
CCA=2
6641
3
1
23
775
C0
BB@2
664 1
3
3
03
7751
CCA=C0
BB@(2)2
6641
6
4
13
775+ ( 2)2
664 4
8
5
83
775+ (4)2
664 5
13
2
93
775+ (3)2
6643
7
3
63
7751
CCA=2
6642
2
4
33
775
Then we package these vectors up as the change-of-basis matrix,
CB;C=2
6641 2 1 2
2 3 3 2
1 3 1 4
1 0 2 33
775
Now consider a single (arbitrary) vector y=2
6642
6
3
43
775. First, build the vector representation of yrelative to
B. This will require writing yas a linear combination of the vectors in B,
B(y) =B0
BB@2
6642
6
3
43
7751
CCA
=B0
BB@( 21)2
6641
2
1
23
775+ (6)2
664 1
3
1
13
775+ (11)2
6642
3
3
43
775+ ( 7)2
664 1
3
3
03
7751
CCA=2
664 21
6
11
73
775
Now, applying Theorem CB [649] we can convert the representation of yrelative toBinto a representation
relative toC,
C(y) =CB;CB(y) Theorem CB [649]
=2
6641 2 1 2
2 3 3 2
1 3 1 4
1 0 2 33
7752
664 21
6
11
73
775
=2
664 12
5
20
223
775Denition MVP [223]
We could continue further with this example, perhaps by computing the representation of yrelative to the
basisCdirectly as a check on our work (Exercise CB.C20 [669]). Or we could choose another vector to
Version 2.30
656 Section CB Change of Basis
play the role of yand compute two dierent representations of this vector relative to the two bases Band
C.
Subsection MRS
Matrix Representations and Similarity
Here is the main theorem of this section. It looks a bit involved at rst glance, but the proof should make
you realize it is not all that complicated. In any event, we are more interested in a special case.
Theorem MRCB
Matrix Representation and Change of Basis
Suppose that T:U!Vis a linear transformation, BandCare bases for U, andDandEare bases for
V. Then
MT
B;D=CE;DMT
C;ECB;C
Proof
CE;DMT
C;ECB;C=MIV
E;DMT
C;EMIU
B;CDenition CBM [648]
=MIV
E;DMTIU
B;ETheorem MRCLT [622]
=MIV
E;DMT
B;E Denition IDLT [579]
=MIVT
B;DTheorem MRCLT [622]
=MT
B;D Denition IDLT [579]
We will be most interested in a special case of this theorem (Theorem SCB [656]), but here's an example
that illustrates the full generality of Theorem MRCB [654].
Example MRCM
Matrix representations and change-of-basis matrices
Begin with two vector spaces, S2, the subspace of M22containing all 22 symmetric matrices, and P3
(Example VSP [319]), the vector space of all polynomials of degree 3 or less. Then dene the linear
transformation Q:S2!P3by
Qa b
b c
= (5a 2b+ 6c) + (3a b+ 2c)x+ (a+ 3b c)x2+ ( 4a+ 2b+c)x3
Here are two bases for each vector space, one nice, one nasty. First for S2,
B=5 3
3 2
;2 3
3 0
;1 2
2 4
C=1 0
0 0
;0 1
1 0
;0 0
0 1
and then for P3,
D=
2 +x 2x2+ 3x3; 1 2x2+ 3x3; 3 x+x3; x2+x3
E=
1; x; x2; x3
We'll begin with a matrix representation of Qrelative toCandE. We rst nd vector representations of
the elements of Crelative toE,
E
Q1 0
0 0
=E
5 + 3x+x2 4x3
=2
6645
3
1
43
775
Version 2.30
Subsection CB.MRS Matrix Representations and Similarity 657
E
Q0 1
1 0
=E
2 x+ 3x2+ 2x3
=2
664 2
1
3
23
775
E
Q0 0
0 1
=E
6 + 2x x2+x3
=2
6646
2
1
13
775
So
MQ
C;E=2
6645 2 6
3 1 2
1 3 1
4 2 13
775
Now we construct two change-of-basis matrices. First, CB;Crequires vector representations of the elements
ofB, relative to C. SinceCis a nice basis, this is straightforward,
C5 3
3 2
=C
(5)1 0
0 0
+ ( 3)0 1
1 0
+ ( 2)0 0
0 1
=2
45
3
23
5
C2 3
3 0
=C
(2)1 0
0 0
+ ( 3)0 1
1 0
+ (0)0 0
0 1
=2
42
3
03
5
C1 2
2 4
=C
(1)1 0
0 0
+ (2)0 1
1 0
+ (4)0 0
0 1
=2
41
2
43
5
So
CB;C=2
45 2 1
3 3 2
2 0 43
5
The other change-of-basis matrix we'll compute is CE;D. However, since Eis a nice basis (and Dis
not) we'll turn it around and instead compute CD;Eand apply Theorem ICBM [649] to use an inverse to
computeCE;D.
E
2 +x 2x2+ 3x3
=E
(2)1 + (1)x+ ( 2)x2+ (3)x3
=2
6642
1
2
33
775
E
1 2x2+ 3x3
=E
( 1)1 + (0)x+ ( 2)x2+ (3)x3
=2
664 1
0
2
33
775
E
3 x+x3
=E
( 3)1 + ( 1)x+ (0)x2+ (1)x3
=2
664 3
1
0
13
775
Version 2.30
658 Section CB Change of Basis
E
x2+x3
=E
(0)1 + (0)x+ ( 1)x2+ (1)x3
=2
6640
0
1
13
775
So, we can package these column vectors up as a matrix to obtain CD;Eand then,
CE;D= (CD;E) 1Theorem ICBM [649]
=2
6642 1 3 0
1 0 1 0
2 2 0 1
3 3 1 13
775 1
=2
6641 2 1 1
2 5 1 1
1 3 1 1
2 6 1 03
775
We are now in a position to apply Theorem MRCB [654]. The matrix representation of Qrelative toB
andDcan be obtained as follows,
MQ
B;D=CE;DMQ
C;ECB;C Theorem MRCB [654]
=2
6641 2 1 1
2 5 1 1
1 3 1 1
2 6 1 03
7752
6645 2 6
3 1 2
1 3 1
4 2 13
7752
45 2 1
3 3 2
2 0 43
5
=2
6641 2 1 1
2 5 1 1
1 3 1 1
2 6 1 03
7752
66419 16 25
14 9 9
2 7 3
28 14 43
775
=2
664 39 23 14
62 34 12
53 32 5
44 15 73
775
Now check our work by computing MQ
B;Ddirectly (Exercise CB.C21 [669]).
Here is a special case of the previous theorem, where we choose UandVto be the same vector space,
so the matrix representations and the change-of-basis matrices are all square of the same size.
Theorem SCB
Similarity and Change of Basis
Suppose that T:V!Vis a linear transformation and BandCare bases of V. Then
MT
B;B=C 1
B;CMT
C;CCB;C
Proof In the conclusion of Theorem MRCB [654], replace DbyB, and replace EbyC,
MT
B;B=CC;BMT
C;CCB;C Theorem MRCB [654]
=C 1
B;CMT
C;CCB;C Theorem ICBM [649]
Version 2.30
Subsection CB.MRS Matrix Representations and Similarity 659
This is the third surprise of this chapter. Theorem SCB [656] considers the special case where a
linear transformation has the same vector space for the domain and codomain ( V). We build a matrix
representation of Tusing the basis Bsimultaneously for both the domain and codomain ( MT
B;B), and then
we build a second matrix representation of T, now using the basis Cfor both the domain and codomain
(MT
C;C). Then these two representations are related via a similarity transformation (Denition SIM [493])
using a change-of-basis matrix ( CB;C)!
Example MRBE
Matrix representation with basis of eigenvectors
We return to the linear transformation T:M22!M22of Example ELTBM [647] dened by
Ta b
c d
= 17a+ 11b+ 8c 11d 57a+ 35b+ 24c 33d
14a+ 10b+ 6c 10d 41a+ 25b+ 16c 23d
In Example ELTBM [647] we showcased four eigenvectors of T. We will now put these four vectors in a
set,
B=fx1;x2;x3;x4g=0 1
0 1
;1 1
1 0
;1 3
2 3
;2 6
1 4
Check that Bis a basis of M22by rst establishing the linear independence of Band then employing
Theorem G [407] to get the spanning property easily. Here is a second set of 2 2 matrices, which also
forms a basis of M22(Example BM [372]),
C=fy1;y2;y3;y4g=1 0
0 0
;0 1
0 0
;0 0
1 0
;0 0
0 1
We can build two matrix representations of T, one relative to Band one relative to C. Each is easy, but
for wildly dierent reasons. In our computation of the matrix representation relative to Bwe borrow some
of our work in Example ELTBM [647]. Here are the representations, then the explanation.
B(T(x1)) =B(2x1) =B(2x1+ 0x2+ 0x3+ 0x4) =2
6642
0
0
03
775
B(T(x2)) =B(2x2) =B(0x1+ 2x2+ 0x3+ 0x4) =2
6640
2
0
03
775
B(T(x3)) =B(( 1)x3) =B(0x1+ 0x2+ ( 1)x3+ 0x4) =2
6640
0
1
03
775
B(T(x4)) =B(( 2)x4) =B(0x1+ 0x2+ 0x3+ ( 2)x4) =2
6640
0
0
23
775
So the resulting representation is
MT
B;B=2
6642 0 0 0
0 2 0 0
0 0 1 0
0 0 0 23
775
Version 2.30
660 Section CB Change of Basis
Very pretty. Now for the matrix representation relative to Crst compute,
C(T(y1)) =C 17 57
14 41
=C
( 17)1 0
0 0
+ ( 57)0 1
0 0
+ ( 14)0 0
1 0
+ ( 41)0 0
0 1
=2
664 17
57
14
413
775
C(T(y2)) =C11 35
10 25
=C
111 0
0 0
+ 350 1
0 0
+ 100 0
1 0
+ 250 0
0 1
=2
66411
35
10
253
775
C(T(y3)) =C8 24
6 16
=C
81 0
0 0
+ 240 1
0 0
+ 60 0
1 0
+ 160 0
0 1
=2
6648
24
6
163
775
C(T(y4)) =C 11 33
10 23
=C
( 11)1 0
0 0
+ ( 33)0 1
0 0
+ ( 10)0 0
1 0
+ ( 23)0 0
0 1
=2
664 11
33
10
233
775
So the resulting representation is
MT
C;C=2
664 17 11 8 11
57 35 24 33
14 10 6 10
41 25 16 233
775
Not quite as pretty. The purpose of this example is to illustrate Theorem SCB [656]. This theorem says
that the two matrix representations, MT
B;BandMT
C;C, of the one linear transformation, T, are related
by a similarity transformation using the change-of-basis matrix CB;C. Lets compute this change-of-basis
matrix. Notice that since Cis such a nice basis, this is fairly straightforward,
C(x1) =C0 1
0 1
=C
01 0
0 0
+ 10 1
0 0
+ 00 0
1 0
+ 10 0
0 1
=2
6640
1
0
13
775
C(x2) =C1 1
1 0
=C
11 0
0 0
+ 10 1
0 0
+ 10 0
1 0
+ 00 0
0 1
=2
6641
1
1
03
775
Version 2.30
Subsection CB.MRS Matrix Representations and Similarity 661
C(x3) =C1 3
2 3
=C
11 0
0 0
+ 30 1
0 0
+ 20 0
1 0
+ 30 0
0 1
=2
6641
3
2
33
775
C(x4) =C2 6
1 4
=C
21 0
0 0
+ 60 1
0 0
+ 10 0
1 0
+ 40 0
0 1
=2
6642
6
1
43
775
So we have,
CB;C=2
6640 1 1 2
1 1 3 6
0 1 2 1
1 0 3 43
775
Now, according to Theorem SCB [656] we can write,
MT
B;B=C 1
B;CMT
C;CCB;C
2
6642 0 0 0
0 2 0 0
0 0 1 0
0 0 0 23
775=2
6640 1 1 2
1 1 3 6
0 1 2 1
1 0 3 43
775 12
664 17 11 8 11
57 35 24 33
14 10 6 10
41 25 16 233
7752
6640 1 1 2
1 1 3 6
0 1 2 1
1 0 3 43
775
This should look and feel exactly like the process for diagonalizing a matrix, as was described in Section
SD [493]. And it is.
We can now return to the question of computing an eigenvalue or eigenvector of a linear transformation.
For a linear transformation of the form T:V!V, we know that representations relative to dierent
bases are similar matrices. We also know that similar matrices have equal characteristic polynomials by
Theorem SMEE [495]. We will now show that eigenvalues of a linear transformation Tare precisely the
eigenvalues of anymatrix representation of T. Since the choice of a dierent matrix representation leads to
a similar matrix, there will be no \new" eigenvalues obtained from this second representation. Similarly,
the change-of-basis matrix can be used to show that eigenvectors obtained from one matrix representation
will be precisely those obtained from any other representation. So we can determine the eigenvalues and
eigenvectors of a linear transformation by forming one matrix representation, using anybasis we please,
and analyzing the matrix in the manner of Chapter E [453].
Theorem EER
Eigenvalues, Eigenvectors, Representations
Suppose that T:V!Vis a linear transformation and Bis a basis of V. Then v2Vis an eigenvector of
Tfor the eigenvalue if and only if B(v) is an eigenvector of MT
B;Bfor the eigenvalue .
Proof ()) Assume that v2Vis an eigenvector of Tfor the eigenvalue . Then
MT
B;BB(v) =B(T(v)) Theorem FTMR [617]
=B(v) Denition EELT [647]
=B(v) Theorem VRLT [603]
which by Denition EEM [453] says that B(v) is an eigenvector of the matrix MT
B;Bfor the eigenvalue .
(() Assume that B(v) is an eigenvector of MT
B;Bfor the eigenvalue . Then
T(v) = 1
B(B(T(v))) Denition IVLT [579]
= 1
B
MT
B;BB(v)
Theorem FTMR [617]
Version 2.30
662 Section CB Change of Basis
= 1
B(B(v)) Denition EEM [453]
= 1
B(B(v)) Theorem ILTLT [582]
=v Denition IVLT [579]
which by Denition EELT [647] says vis an eigenvector of Tfor the eigenvalue .
Subsection CELT
Computing Eigenvectors of Linear Transformations
Knowing that the eigenvalues of a linear transformation are the eigenvalues of any representation, no matter
what the choice of the basis Bmight be, we could now unambiguously dene items such as the charac-
teristic polynomial of a linear transformation, rather than a matrix. We'll say that again | eigenvalues,
eigenvectors, and characteristic polynomials are intrinsic properties of a linear transformation, independent
of the choice of a basis used to construct a matrix representation.
As a practical matter, how does one compute the eigenvalues and eigenvectors of a linear transformation
of the form T:V!V? Choose a nice basis BforV, one where the vector representations of the values
of the linear transformations necessary for the matrix representation are easy to compute. Construct the
matrix representation relative to this basis, and nd the eigenvalues and eigenvectors of this matrix using
the techniques of Chapter E [453]. The resulting eigenvalues of the matrix are precisely the eigenvalues of
the linear transformation. The eigenvectors of the matrix are column vectors that need to be converted to
vectors inVthrough application of 1
B.
Now consider the case where the matrix representation of a linear transformation is diagonalizable. The
nlinearly independent eigenvectors that must exist for the matrix (Theorem DC [497]) can be converted (via
1
B) into eigenvectors of the linear transformation. A matrix representation of the linear transformation
relative to a basis of eigenvectors will be a diagonal matrix | an especially nice representation! Though we
did not know it at the time, the diagonalizations of Section SD [493] were really nding especially pleasing
matrix representations of linear transformations.
Here are some examples.
Example ELTT
Eigenvectors of a linear transformation, twice
Consider the linear transformation S:M22!M22dened by
Sa b
c d
= b c 3d 14a 15b 13c+d
18a+ 21b+ 19c+ 3d 6a 7b 7c 3d
To nd the eigenvalues and eigenvectors of Swe will build a matrix representation and analyze the matrix.
Since Theorem EER [659] places no restriction on the choice of the basis B, we may as well use a basis
that is easy to work with. So set
B=fx1;x2;x3;x4g=1 0
0 0
;0 1
0 0
;0 0
1 0
;0 0
0 1
Then to build the matrix representation of Srelative toBcompute,
B(S(x1)) =B0 14
18 6
=B(0x1+ ( 14)x2+ 18x3+ ( 6)x4) =2
6640
14
18
63
775
Version 2.30
Subsection CB.CELT Computing Eigenvectors of Linear Transformations 663
B(S(x2)) =B 1 15
21 7
=B(( 1)x1+ ( 15)x2+ 21x3+ ( 7)x4) =2
664 1
15
21
73
775
B(S(x3)) =B 1 13
19 7
=B(( 1)x1+ ( 13)x2+ 19x3+ ( 7)x4) =2
664 1
13
19
73
775
B(S(x4)) =B 3 1
3 3
=B(( 3)x1+ 1x2+ 3x3+ ( 3)x4) =2
664 3
1
3
33
775
So by Denition MR [615] we have
M=MS
B;B=2
6640 1 1 3
14 15 13 1
18 21 19 3
6 7 7 33
775
Now compute eigenvalues and eigenvectors of the matrix representation of Mwith the techniques of Section
EE [453]. First the characteristic polynomial,
pM(x) = det (M xI4) =x4 x3 10x2+ 4x+ 24 = (x 3)(x 2)(x+ 2)2
We could now make statements about the eigenvalues of M, but in light of Theorem EER [659] we can
refer to the eigenvalues of Sand mildly abuse (or extend) our notation for multiplicities to write
S(3) = 1 S(2) = 1 S( 2) = 2
Now compute the eigenvectors of M,
= 3 M 3I4=2
664 3 1 1 3
14 18 13 1
18 21 16 3
6 7 7 63
775RREF !2
66410 0 1
010 3
0 0 1 3
0 0 0 03
775
EM(3) =N(M 3I4) =*8
>><
>>:2
664 1
3
3
13
7759
>>=
>>;+
= 2 M 2I4=2
664 2 1 1 3
14 17 13 1
18 21 17 3
6 7 7 53
775RREF !2
66410 0 2
010 4
0 0 1 3
0 0 0 03
775
EM(2) =N(M 2I4) =*8
>><
>>:2
664 2
4
3
13
7759
>>=
>>;+
Version 2.30
664 Section CB Change of Basis
= 2 M ( 2)I4=2
6642 1 1 3
14 13 13 1
18 21 21 3
6 7 7 13
775RREF !2
66410 0 1
011 1
0 0 0 0
0 0 0 03
775
EM( 2) =N(M ( 2)I4) =*8
>><
>>:2
6640
1
1
03
775;2
6641
1
0
13
7759
>>=
>>;+
According to Theorem EER [659] the eigenvectors just listed as basis vectors for the eigenspaces of M
are vector representations (relative to B) of eigenvectors for S. So the application if the inverse function
1
Bwill convert these column vectors into elements of the vector space M22(22 matrices) that are
eigenvectors of S. SinceBis an isomorphism (Theorem VRILT [608]), so is 1
B. Applying the inverse
function will then preserve linear independence and spanning properties, so with a sweeping application
of the Coordinatization Principle [611] and some extensions of our previous notation for eigenspaces and
geometric multiplicities, we can write,
1
B0
BB@2
664 1
3
3
13
7751
CCA= ( 1)x1+ 3x2+ ( 3)x3+ 1x4= 1 3
3 1
1
B0
BB@2
664 2
4
3
13
7751
CCA= ( 2)x1+ 4x2+ ( 3)x3+ 1x4= 2 4
3 1
1
B0
BB@2
6640
1
1
03
7751
CCA= 0x1+ ( 1)x2+ 1x3+ 0x4=0 1
1 0
1
B0
BB@2
6641
1
0
13
7751
CCA= 1x1+ ( 1)x2+ 0x3+ 1x4=1 1
0 1
So
ES(3) = 1 3
3 1
ES(2) = 2 4
3 1
ES( 2) =0 1
1 0
;1 1
0 1
with geometric multiplicities given by
S(3) = 1
S(2) = 1
S( 2) = 2
Suppose we now decided to build another matrix representation of S, only now relative to a linearly
independent set of eigenvectors of S, such as
C= 1 3
3 1
; 2 4
3 1
;0 1
1 0
;1 1
0 1
Version 2.30
Subsection CB.CELT Computing Eigenvectors of Linear Transformations 665
At this point you should have computed enough matrix representations to predict that the result of
representing Srelative to Cwill be a diagonal matrix. Computing this representation is an example of
how Theorem SCB [656] generalizes the diagonalizations from Section SD [493]. For the record, here is the
diagonal representation,
MS
C;C=2
6643 0 0 0
0 2 0 0
0 0 2 0
0 0 0 23
775
Our interest in this example is not necessarily building nice representations, but instead we want to demon-
strate how eigenvalues and eigenvectors are an intrinsic property of a linear transformation, independent
of any particular representation. To this end, we will repeat the foregoing, but replace Bby another basis.
We will make this basis dierent, but not extremely so,
D=fy1;y2;y3;y4g=1 0
0 0
;1 1
0 0
;1 1
1 0
;1 1
1 1
Then to build the matrix representation of Srelative toDcompute,
D(S(y1)) =D0 14
18 6
=D(14y1+ ( 32)y2+ 24y3+ ( 6)y4) =2
66414
32
24
63
775
D(S(y2)) =D 1 29
39 13
=D(28y1+ ( 68)y2+ 52y3+ ( 13)y4) =2
66428
68
52
133
775
D(S(y3)) =D 2 42
58 20
=D(40y1+ ( 100)y2+ 78y3+ ( 20)y4) =2
66440
100
78
203
775
D(S(y4)) =D 5 41
61 23
=D(36y1+ ( 102)y2+ 84y3+ ( 23)y4) =2
66436
102
84
233
775
So by Denition MR [615] we have
N=MS
D;D=2
66414 28 40 36
32 68 100 102
24 52 78 84
6 13 20 233
775
Now compute eigenvalues and eigenvectors of the matrix representation of Nwith the techniques of Section
EE [453]. First the characteristic polynomial,
pN(x) = det (N xI4) =x4 x3 10x2+ 4x+ 24 = (x 3)(x 2)(x+ 2)2
Of course this is not news. We now know that M=MS
B;BandN=MS
D;Dare similar matrices (Theorem
SCB [656]). But Theorem SMEE [495] told us long ago that similar matrices have identical characteristic
Version 2.30
666 Section CB Change of Basis
polynomials. Now compute eigenvectors for the matrix representation, which will be dierent than what
we found for M,
= 3 N 3I4=2
66411 28 40 36
32 71 100 102
24 52 75 84
6 13 20 263
775RREF !2
6641 0 0 4
0 1 0 6
0 0 1 4
0 0 0 03
775
EN(3) =N(N 3I4) =*8
>><
>>:2
664 4
6
4
13
7759
>>=
>>;+
= 2 N 2I4=2
66412 28 40 36
32 70 100 102
24 52 76 84
6 13 20 253
775RREF !2
6641 0 0 6
0 1 0 7
0 0 1 4
0 0 0 03
775
EN(2) =N(N 2I4) =*8
>><
>>:2
664 6
7
4
13
7759
>>=
>>;+
= 2 N ( 2)I4=2
66416 28 40 36
32 66 100 102
24 52 80 84
6 13 20 213
775RREF !2
6641 0 1 3
0 1 2 3
0 0 0 0
0 0 0 03
775
EN( 2) =N(N ( 2)I4) =*8
>><
>>:2
6641
2
1
03
775;2
6643
3
0
13
7759
>>=
>>;+
Employing Theorem EER [659] we can apply 1
Dto each of the basis vectors of the eigenspaces of Nto
obtain eigenvectors for Sthat also form bases for eigenspaces of S,
1
D0
BB@2
664 4
6
4
13
7751
CCA= ( 4)y1+ 6y2+ ( 4)y3+ 1y4= 1 3
3 1
1
D0
BB@2
664 6
7
4
13
7751
CCA= ( 6)y1+ 7y2+ ( 4)y3+ 1y4= 2 4
3 1
1
D0
BB@2
6641
2
1
03
7751
CCA= 1y1+ ( 2)y2+ 1y3+ 0y4=0 1
1 0
1
D0
BB@2
6643
3
0
13
7751
CCA= 3y1+ ( 3)y2+ 0y3+ 1y4=1 2
1 1
Version 2.30
Subsection CB.CELT Computing Eigenvectors of Linear Transformations 667
The eigenspaces for the eigenvalues of algebraic multiplicity 1 are exactly as before,
ES(3) = 1 3
3 1
ES(2) = 2 4
3 1
However, the eigenspace for = 2 would at rst glance appear to be dierent. Here are the two
eigenspaces for = 2, rst the eigenspace obtained from M=MS
B;B, then followed by the eigenspace
obtained from M=MS
D;D.
ES( 2) =0 1
1 0
;1 1
0 1
ES( 2) =0 1
1 0
;1 2
1 1
Subspaces generally have many bases, and that is the situation here. With a careful proof of set equality,
you can show that these two eigenspaces are equal sets. The key observation to make such a proof go is
that 1 2
1 1
=0 1
1 0
+1 1
0 1
which will establish that the second set is a subset of the rst. With equal dimensions, Theorem EDYES
[410] will nish the task. So the eigenvalues of a linear transformation are independent of the matrix
representation employed to compute them!
Another example, this time a bit larger and with complex eigenvalues.
Example CELT
Complex eigenvectors of a linear transformation
Consider the linear transformation Q:P4!P4dened by
Q
a+bx+cx2+dx3+ex4
= ( 46a 22b+ 13c+ 5d+e) + (117a+ 57b 32c 15d 4e)x+
( 69a 29b+ 21c 7e)x2+ (159a+ 73b 44c 13d+ 2e)x3+
( 195a 87b+ 55c+ 10d 13e)x4
Choose a simple basis to compute with, say
B=
1; x; x2; x3; x4
Then it should be apparent that the matrix representation of Qrelative toBis
M=MQ
B;B=2
66664 46 22 13 5 1
117 57 32 15 4
69 29 21 0 7
159 73 44 13 2
195 87 55 10 133
77775
Compute the characteristic polynomial, eigenvalues and eigenvectors according to the techniques of Section
EE [453],
pQ(x) = x5+ 6x4 x3 88x2+ 252x 208
= (x 2)2(x+ 4)
x2 6x+ 13
Version 2.30
668 Section CB Change of Basis
= (x 2)2(x+ 4) (x (3 + 2i)) (x (3 2i))
Q(2) = 2 Q( 4) = 1 Q(3 + 2i) = 1 Q(3 2i) = 1
= 2
M (2)I5=2
66664 48 22 13 5 1
117 55 32 15 4
69 29 19 0 7
159 73 44 15 2
195 87 55 10 153
77775RREF !2
666641 0 01
2 1
2
0 1 0 5
2 5
2
0 0 1 2 6
0 0 0 0 0
0 0 0 0 03
77775
EM(2) =N(M (2)I5) =*8
>>>><
>>>>:2
66664 1
25
2
2
1
03
77775;2
666641
25
2
6
0
13
777759
>>>>=
>>>>;+
=*8
>>>><
>>>>:2
66664 1
5
4
2
03
77775;2
666641
5
12
0
23
777759
>>>>=
>>>>;+
= 4
M ( 4)I5=2
66664 42 22 13 5 1
117 61 32 15 4
69 29 25 0 7
159 73 44 9 2
195 87 55 10 93
77775RREF !2
666641 0 0 0 1
0 1 0 0 3
0 0 1 0 1
0 0 0 1 2
0 0 0 0 03
77775
EM( 4) =N(M ( 4)I5) =*8
>>>><
>>>>:2
66664 1
3
1
2
13
777759
>>>>=
>>>>;+
= 3 + 2i
M (3 + 2i)I5=2
66664 49 2i 22 13 5 1
117 54 2i 32 15 4
69 29 18 2i 0 7
159 73 44 16 2i 2
195 87 55 10 16 2i3
77775RREF !2
666641 0 0 0 3
4+i
4
0 1 0 07
4 i
4
0 0 1 0 1
2+i
2
0 0 0 17
4 i
4
0 0 0 0 03
77775
EM(3 + 2i) =N(M (3 + 2i)I5) =*8
>>>><
>>>>:2
666643
4 i
4
7
4+i
41
2 i
2
7
4+i
4
13
777759
>>>>=
>>>>;+
=*8
>>>><
>>>>:2
666643 i
7 +i
2 2i
7 +i
43
777759
>>>>=
>>>>;+
= 3 2i
M (3 2i)I5=2
66664 49 + 2i 22 13 5 1
117 54 + 2 i 32 15 4
69 29 18 + 2 i 0 7
159 73 44 16 + 2i 2
195 87 55 10 16 + 2i3
77775RREF !2
666641 0 0 0 3
4 i
4
0 1 0 07
4+i
4
0 0 1 0 1
2 i
2
0 0 0 17
4+i
4
0 0 0 0 03
77775
Version 2.30
Subsection CB.CELT Computing Eigenvectors of Linear Transformations 669
EM(3 2i) =N(M (3 2i)I5) =*8
>>>><
>>>>:2
666643
4+i
4
7
4 i
41
2+i
2
7
4 i
4
13
777759
>>>>=
>>>>;+
=*8
>>>><
>>>>:2
666643 +i
7 i
2 + 2i
7 i
43
777759
>>>>=
>>>>;+
It is straightforward to convert each of these basis vectors for eigenspaces of Mback to elements of P4by
applying the isomorphism 1
B,
1
B0
BBBB@2
66664 1
5
4
2
03
777751
CCCCA= 1 + 5x+ 4x2+ 2x3
1
B0
BBBB@2
666641
5
12
0
23
777751
CCCCA= 1 + 5x+ 12x2+ 2x4
1
B0
BBBB@2
66664 1
3
1
2
13
777751
CCCCA= 1 + 3x+x2+ 2x3+x4
1
B0
BBBB@2
666643 i
7 +i
2 2i
7 +i
43
777751
CCCCA= (3 i) + ( 7 +i)x+ (2 2i)x2+ ( 7 +i)x3+ 4x4
1
B0
BBBB@2
666643 +i
7 i
2 + 2i
7 i
43
777751
CCCCA= (3 +i) + ( 7 i)x+ (2 + 2i)x2+ ( 7 i)x3+ 4x4
So we apply Theorem EER [659] and the Coordinatization Principle [611] to get the eigenspaces for Q,
EQ(2) =
1 + 5x+ 4x2+ 2x3;1 + 5x+ 12x2+ 2x4
EQ( 4) =
1 + 3x+x2+ 2x3+x4
EQ(3 + 2i) =
(3 i) + ( 7 +i)x+ (2 2i)x2+ ( 7 +i)x3+ 4x4
EQ(3 2i) =
(3 +i) + ( 7 i)x+ (2 + 2i)x2+ ( 7 i)x3+ 4x4
with geometric multiplicities
Q(2) = 2
Q( 4) = 1
Q(3 + 2i) = 1
Q(3 2i) = 1
Version 2.30
670 Section CB Change of Basis
Subsection READ
Reading Questions
1. The change-of-basis matrix is a matrix representation of which linear transformation?
2. Find the change-of-basis matrix, CB;C, for the two bases of C2
B=2
3
; 1
2
C=1
0
;1
1
3. What is the third \surprise," and why is it surprising?
Version 2.30
Subsection CB.EXC Exercises 671
Subsection EXC
Exercises
C20 In Example CBCV [652] we computed the vector representation of yrelative to C,C(y), as an
example of Theorem CB [649]. Compute this same representation directly. In other words, apply Denition
VR [603] rather than Theorem CB [649].
Contributed by Robert Beezer
C21 Perform a check on Example MRCM [654] by computing MQ
B;Ddirectly. In other words, apply
Denition MR [615] rather than Theorem MRCB [654].
Contributed by Robert Beezer Solution [670]
C30 Find a basis for the vector space P3composed of eigenvectors of the linear transformation T. Then
nd a matrix representation of Trelative to this basis.
T:P3!P3; T
a+bx+cx2+dx3
= (a+c+d) + (b+c+d)x+ (a+b+c)x2+ (a+b+d)x3
Contributed by Robert Beezer Solution [670]
C40 LetS22be the vector space of 2 2 symmetric matrices. Find a basis BforS22that yields a diagonal
matrix representation of the linear transformation R.
R:S22!S22; Ra b
b c
= 5a+ 2b 3c 12a+ 5b 6c
12a+ 5b 6c6a 2b+ 4c
Contributed by Robert Beezer Solution [671]
C41 LetS22be the vector space of 2 2 symmetric matrices. Find a basis for S22composed of eigenvectors
of the linear transformation Q:S22!S22.
Qa b
b c
=25a+ 18b+ 30c 16a 11b 20c
16a 11b 20c 11a 9b 12c
Contributed by Robert Beezer Solution [672]
T10 Suppose that T:V!Vis an invertible linear transformation with a nonzero eigenvalue . Prove
that1
is an eigenvalue of T 1.
Contributed by Robert Beezer Solution [672]
T15 Suppose that Vis a vector space and T:V!Vis a linear transformation. Prove that Tis injective
if and only if = 0 is not an eigenvalue of T.
Contributed by Robert Beezer
Version 2.30
672 Section CB Change of Basis
Subsection SOL
Solutions
C21 Contributed by Robert Beezer Statement [669]
Apply Denition MR [615],
D
Q5 3
3 2
=D
19 + 14x 2x2 28x3
=D
( 39)(2 +x 2x2+ 3x3) + 62( 1 2x2+ 3x3) + ( 53)( 3 x+x3) + ( 44)( x2+x3)
=2
664 39
62
53
443
775
D
Q2 3
3 0
=D
16 + 9x 7x2 14x3
=D
( 23)(2 +x 2x2+ 3x3) + (34)( 1 2x2+ 3x3) + ( 32)( 3 x+x3) + ( 15)( x2+x3)
=2
664 23
34
32
153
775
D
Q1 2
2 4
=D
25 + 9x+ 3x2+ 4x3
=D
(14)(2 +x 2x2+ 3x3) + ( 12)( 1 2x2+ 3x3) + 5( 3 x+x3) + ( 7)( x2+x3)
=2
66414
12
5
73
775
These three vectors are the columns of the matrix representation,
MQ
B;D=2
664 39 23 14
62 34 12
53 32 5
44 15 73
775
which coincides with the result obtained in Example MRCM [654].
C30 Contributed by Robert Beezer Statement [669]
With the domain and codomain being identical, we will build a matrix representation using the same basis
for both the domain and codomain. The eigenvalues of the matrix representation will be the eigenvalues
of the linear transformation, and we can obtain the eigenvectors of the linear transformation by un-
coordinatizing (Theorem EER [659]). Since the method does not depend on which basis we choose, we can
choose a natural basis for ease of computation, say,
B=
1; x; x2;x3
Version 2.30
Subsection CB.SOL Solutions 673
The matrix representation is then,
MT
B;B=2
6641 0 1 1
0 1 1 1
1 1 1 0
1 1 0 13
775
The eigenvalues and eigenvectors of this matrix were computed in Example ESMS4 [464]. A basis for C4,
composed of eigenvectors of the matrix representation is,
C=8
>><
>>:2
6641
1
1
13
775;2
664 1
1
0
03
775;2
6640
0
1
13
775;2
664 1
1
1
13
7759
>>=
>>;
Applying 1
Bto each vector of this set, yields a basis of P3composed of eigenvectors of T,
D=
1 +x+x2+x3; 1 +x; x2+x3; 1 x+x2+x3
The matrix representation of Trelative to the basis Dwill be a diagonal matrix with the corresponding
eigenvalues along the diagonal, so in this case we get
MT
D;D=2
6643 0 0 0
0 1 0 0
0 0 1 0
0 0 0 13
775
C40 Contributed by Robert Beezer Statement [669]
Begin with a matrix representation of R, any matrix representation, but use the same basis for both
instances of S22. We'll choose a basis that makes it easy to compute vector representations in S22.
B=1 0
0 0
;0 1
1 0
;0 0
0 1
Then the resulting matrix representation of R(Denition MR [615]) is
MR
B;B=2
4 5 2 3
12 5 6
6 2 43
5
Now, compute the eigenvalues and eigenvectors of this matrix, with the goal of diagonalizing the matrix
(Theorem DC [497]),
= 2 EMR
B;B(2) =*8
<
:2
4 1
2
13
59
=
;+
= 1 EMR
B;B(1) =*8
<
:2
4 1
0
23
5;2
41
3
03
59
=
;+
The three vectors that occur as basis elements for these eigenspaces will together form a linearly inde-
pendent set (check this!). So these column vectors may be employed in a matrix that will diagonalize
the matrix representation. If we \un-coordinatize" these three column vectors relative to the basis B, we
Version 2.30
674 Section CB Change of Basis
will nd three linearly independent elements of S22that are eigenvectors of the linear transformation R
(Theorem EER [659]). A matrix representation relative to this basis of eigenvectors will be diagonal, with
the eigenvalues ( = 2;1) as the diagonal elements. Here we go,
1
B0
@2
4 1
2
13
51
A= ( 1)1 0
0 0
+ ( 2)0 1
1 0
+ 10 0
0 1
= 1 2
2 1
1
B0
@2
4 1
0
23
51
A= ( 1)1 0
0 0
+ 00 1
1 0
+ 20 0
0 1
= 1 0
0 2
1
B0
@2
41
3
03
51
A= 11 0
0 0
+ 30 1
1 0
+ 00 0
0 1
=1 3
3 0
So the requested basis of S22, yielding a diagonal matrix representation of R, is
1 2
2 1 1 0
0 2
;1 3
3 0
C41 Contributed by Robert Beezer Statement [669]
Use a single basis for both the domain and codomain, since they are equal.
B=1 0
0 0
;0 1
1 0
;0 0
0 1
The matrix representation of Qrelative toBis
M=MQ
B;B=2
425 18 30
16 11 20
11 9 123
5
We can analyze this matrix with the techniques of Section EE [453] and then apply Theorem EER [659].
The eigenvalues of this matrix are = 2;1;3 with eigenspaces
EM( 2) =*8
<
:2
4 6
4
33
59
=
;+
EM(1) =*8
<
:2
4 2
1
13
59
=
;+
EM(3) =*8
<
:2
4 3
2
13
59
=
;+
Because the three eigenvalues are distinct, the three basis vectors from the three eigenspaces for a linearly
independent set (Theorem EDELI [479]). Theorem EER [659] says we can uncoordinatize these eigenvectors
to obtain eigenvectors of Q. By Theorem ILTLI [549] the resulting set will remain linearly independent.
Set
C=8
<
: 1
B0
@2
4 6
4
33
51
A; 1
B0
@2
4 2
1
13
51
A; 1
B0
@2
4 3
2
13
51
A9
=
;= 6 4
4 3
; 2 1
1 1
; 3 2
2 1
ThenCis a linearly independent set of size 3 in the vector space M22, which has dimension 3 as well. By
Theorem G [407], Cis a basis of M22.
T10 Contributed by Robert Beezer Statement [669]
Letvbe an eigenvector of Tfor the eigenvalue . Then,
T 1(v) =1
T 1(v) 6= 0
Version 2.30
Subsection CB.SOL Solutions 675
=1
T 1(v) Theorem ILTLT [582]
=1
T 1(T(v)) veigenvector of T
=1
IV(v) Denition IVLT [579]
=1
v Denition IDLT [579]
which says that1
is an eigenvalue of T 1with eigenvector v. Note that it is possible to prove that any
eigenvalue of an invertible linear transformation is never zero. So the hypothesis that be nonzero is just
a convenience for this problem.
Version 2.30
676 Section CB Change of Basis
Version 2.30
Section OD Orthonormal Diagonalization 677
Section OD
Orthonormal Diagonalization
This section is in draft form
Theorems & definitions are complete, needs examples
We have seen in Section SD [493] that under the right conditions a square matrix is similar to a diagonal
matrix. We recognize now, via Theorem SCB [656], that a similarity transformation is a change of basis on
a matrix representation. So we can now discuss the choice of a basis used to build a matrix representation,
and decide if some bases are better than others for this purpose. This will be the tone of this section. We
will also see that every matrix has a reasonably useful matrix representation, and we will discover a new
class of diagonalizable linear transformations. First we need some basic facts about triangular matrices.
Subsection TM
Triangular Matrices
An upper, or lower, triangular matrix is exactly what it sounds like it should be, but here are the two
relevant denitions.
Denition UTM
Upper Triangular Matrix
Thennsquare matrix Aisupper triangular if [A]ij= 0 whenever i>j . 4
Denition LTM
Lower Triangular Matrix
Thennsquare matrix Aislower triangular if [A]ij= 0 whenever i<j . 4
Obviously, properties of a lower triangular matrices will have analogues for upper triangular matrices.
Rather than stating two very similar theorems, we will say that matrices are \triangular of the same type"
as a convenient shorthand to cover both possibilities and then give a proof for just one type.
Theorem PTMT
Product of Triangular Matrices is Triangular
Suppose that AandBare square matrices of size nthat are triangular of the same type. Then ABis also
triangular of that type.
Proof We prove this for lower triangular matrices and leave the proof for upper triangular matrices to
you. Suppose that AandBare both lower triangular. We need only establish that certain entries of the
productABare zero. Suppose that i<j , then
[AB]ij=nX
k=1[A]ik[B]kj Theorem EMP [227]
=j 1X
k=1[A]ik[B]kj+nX
k=j[A]ik[B]kj Property AACN [758]
=j 1X
k=1[A]ik0 +nX
k=j[A]ik[B]kj k<j , Denition LTM [675]
=j 1X
k=1[A]ik0 +nX
k=j0 [B]kj i<jk, Denition LTM [675]
Version 2.30
678 Section OD Orthonormal Diagonalization
=j 1X
k=10 +nX
k=j0
= 0
Since [AB]ij= 0 whenever i<j , by Denition LTM [675], ABis lower triangular.
The inverse of a triangular matrix is triangular, of the same type.
Theorem ITMT
Inverse of a Triangular Matrix is Triangular
Suppose that Ais a nonsingular matrix of size nthat is triangular. Then the inverse of A,A 1, is triangular
of the same type. Furthermore, the diagonal entries of A 1are the reciprocals of the corresponding diagonal
entries ofA. More precisely,
A 1
ii= [A] 1
ii.
Proof We give the proof for the case when Ais lower triangular, and leave the case when Ais upper
triangular for you. Consider the process for computing the inverse of a matrix that is outlined in the
proof of Theorem CINM [248]. We augment Awith the size nidentity matrix, In, and row-reduce the
n2nmatrix to reduced row-echelon form via the algorithm in Theorem REMEF [34]. The proof involves
tracking the peculiarities of this process in the case of a lower triangular matrix. Let M= [AjIn].
First, none of the diagonal elements of Aare zero. By repeated expansion about the rst row, the
determinant of a lower triangular matrix can be seen to be the product of the diagonal entries (Theorem
DER [429]). If just one of these diagonal elements was zero, then the determinant of Ais zero and Ais
singular by Theorem SMZD [445]. Slightly violating the exact algorithm for row reduction we can form a
matrix,M0, that is row-equivalent to M, by multiplying row iby the nonzero scalar [ A] 1
ii, for 1in.
This sets [M0]ii= 1 and [M0]i;n+1= [A] 1
ii, and leaves every zero entry of Munchanged.
LetMjdenote the matrix obtained form M0after converting column jto a pivot column. We can
convert column jofMj 1into a pivot column with a set of n j 1 row operations of the form Rj+Rk
withj+ 1kn. The key observation here is that we add multiples of row jonly to higher-numbered
rows. This means that none of the entries in rows 1 through j 1 is changed, and since row jhas zeros in
columnsj+ 1 through n, none of the entries in rows j+ 1 through nis changed in columns j+ 1 through
n. The rst ncolumns of M0form a lower triangular matrix with 1's on the diagonal. In its conversion
to the identity matrix through this sequence of row operations, it remains lower triangular with 1's on the
diagonal.
What happens in columns n+ 1 through 2 nofM0? These columns began in Mas the identity matrix,
and inM0each diagonal entry was scaled to a reciprocal of the corresponding diagonal entry of A. Notice
that trivially, these nal ncolumns of M0form a lower triangular matrix. Just as we argued for the rst
ncolumns, the row operations that convert Mj 1intoMjwill preserve the lower triangular form in the
nalncolumns and preserve the exact values of the diagonal entries. By Theorem CINM [248], the nal n
columns of Mnis the inverse of A, and this matrix has the necessary properties advertised in the conclusion
of this theorem.
Subsection UTMR
Upper Triangular Matrix Representation
Not every matrix is diagonalizable, but every linear transformation has a matrix representation that is an
upper triangular matrix, and the basis that achieves this representation is especially pleasing. Here's the
theorem.
Theorem UTMR
Upper Triangular Matrix Representation
Suppose that T:V!Vis a linear transformation. Then there is a basis BforVsuch that the matrix
Version 2.30
Subsection OD.UTMR Upper Triangular Matrix Representation 679
representation of Trelative toB,MT
B;B, is an upper triangular matrix. Each diagonal entry is an eigenvalue
ofT, and ifis an eigenvalue of T, thenoccursT() times on the diagonal.
Proof We begin with a proof by induction (Technique I [772]) of the rst statement in the conclusion of
the theorem. We use induction on the dimension of Vto show that if T:V!Vis a linear transformation,
then there is a basis BforVsuch that the matrix representation of Trelative toB,MT
B;B, is an upper
triangular matrix.
To start suppose that dim ( V) = 1. Choose any nonzero vector v2Vand realize that V=hfvgi.
Then we can determine Tuniquely by T(v) =vfor some2C(Theorem LTDB [525]). This description
ofTalso gives us a matrix representation relative to the basis B=fvgas the 11 matrix with lone entry
equal to. And this matrix representation is upper triangular (Denition UTM [675]).
For the induction step let dim ( V) =m, and assume the theorem is true for every linear transformation
dened on a vector space of dimension less than m. By Theorem EMHE [457] (suitably converted to the
setting of a linear transformation), Thas at least one eigenvalue, and we denote this eigenvalue as . (We
will remark later about how critical this step is.) We now consider properties of the linear transformation
T IV:V!V.
Letxbe an eigenvector of Tfor. By denition x6=0. Then
(T IV) (x) =T(x) IV(x) Theorem VSLT [532]
=T(x) x Denition IDLT [579]
=x x Denition EELT [647]
=0 Property AI [318]
SoT IVis not injective, as it has a nontrivial kernel (Theorem KILT [548]). With an application of
Theorem RPNDD [588] we bound the rank of T IV,
r(T IV) = dim (V) n(T IV)m 1
DeneWto be the subspace of Vthat is the range of T IV,W=R(T IV). We dene a new linear
transformation S, onW,
S:W!W S (w) =T(w)
This does not look we have accomplished much, since the action of Sis identical to the action of T. For
our purposes this will be a good thing. What is dierent is the domain and codomain. Sis dened on W,
a vector space with dimension less than m, and so is susceptible to our induction hypothesis. Verifying
thatSis really a linear transformation is almost entirely routine, with one exception. Employing Tin our
denition of Sraises the possibility that the outputs of Swill not be contained within W(but instead will
lie insideV, but outside W). To examine this possibility, suppose that w2W.
S(w) =T(w)
=T(w) +0 Property Z [318]
=T(w) + (IV(w) IV(w)) Property AI [318]
= (T(w) IV(w)) +IV(w) Property AA [317]
= (T(w) IV(w)) +w Denition IDLT [579]
= (T IV) (w) +w Theorem VSLT [532]
SinceWis the range of T IV, (T IV) (w)2W. And by Property SC [317], w2W. Finally,
applying Property AC [317] we see by closure that the sum is in Wand so we conclude that S(w)2W.
This argument convinces us that it is legitimate to dene Sas we did with Was the codomain.
Version 2.30
680 Section OD Orthonormal Diagonalization
Sis a linear transformation dened on a vector space with dimension less than m, so we can apply the
induction hypothesis and conclude that Whas a basis, C=fw1;w2;w3; :::; wkg, such that the matrix
representation of Srelative toCis an upper triangular matrix.
By Theorem DSFOS [414] there exists a second subspace of V, which we will call U, so thatVis a
direct sum of WandU,V=WU. Choose a basis D=fu1;u2;u3; :::; u`gforU. Som=k+`by
Theorem DSD [416], and B=C[Dis basis for Vby Theorem DSLI [416] and Theorem G [407]. Bis the
basis we desire. What does a matrix representation of Tlook like, relative to B?
Since the denition of TandSagree onW, the rstkcolumns of MT
B;Bwill have the upper triangular
matrix representation of Sin the rst krows. The remaining `=m krows of these rst kcolumns will
be all zeros since the outputs of TonCare all contained in W. The situation for TonDis not quite as
pretty, but it is close.
For 1i`, consider
B(T(ui)) =B(T(ui) +0) Property Z [318]
=B(T(ui) + (IV(ui) IV(ui))) Property AI [318]
=B((T(ui) IV(ui)) +IV(ui)) Property AA [317]
=B((T(ui) IV(ui)) +ui) Denition IDLT [579]
=B((T IV) (ui) +ui) Theorem VSLT [532]
=B(a1w1+a2w2+a3w3++akwk+ui) Denition RLT [563]
=2
66666666666666666664a1
a2
...
ak
0
...
0
0
...
03
77777777777777777775Denition VR [603]
In the penultimate step of this proof, we have rewritten an element of the range of T IVas a linear
combination of the basis vectors, C, for the range of T IV,W, using the scalars a1; a2; a3; :::; ak.
If we incorporate these `column vectors into the matrix representation MT
B;Bwe nd`occurrences of
on the diagonal, and any nonzero entries lying only in the rst krows. Together with the kkupper
triangular representation in the upper left-hand corner, the entire matrix representation is now clearly
upper triangular. This completes the induction step, so for any linear transformation there is a basis that
creates an upper triangular matrix representation.
We have one more statement in the conclusion of the theorem to verify. The eigenvalues of T, and their
multiplicities, can be computed with the techniques of Chapter E [453] relative to any matrix representation
(Theorem EER [659]). We take this approach with our upper triangular matrix representation MT
B;B. Let
dibe the diagonal entry of MT
B;Bin rowiand column i. Then the characteristic polynomial, computed as
a determinant (Denition CP [460]) with repeated expansions about the rst column, is
pMT
B;B(x) = (d1 x) (d2 x) (d3 x)(dm x)
The roots of the polynomial equation pMT
B;B(x) = 0 are the eigenvalues of the linear transformation
(Theorem EMRCP [461]). So each diagonal entry is an eigenvalue, and is repeated on the diagonal exactly
Version 2.30
Subsection OD.UTMR Upper Triangular Matrix Representation 681
T() times (Denition AME [463]).
A key step in this proof was the construction of the subspace Wwith dimension strictly less than that
ofV. This required an eigenvalue/eigenvector pair, which was guaranteed to us by Theorem EMHE [457].
Digging deeper, the proof of Theorem EMHE [457] requires that we can factor polynomials completely,
into linear factors. This will not always happen if our set of scalars is the reals, R. So this is our nal
explanation of our choice of the complex numbers, C, as our set of scalars. In Cpolynomials factor
completely, so every matrix has at least one eigenvalue, and an inductive argument will get us to upper
triangular matrix representations.
In the case of linear transformations dened on Cm, we can use the inner product (Denition IP [192])
protably to ne-tune the basis that yields an upper triangular matrix representation. Recall that the
adjoint of matrix A(Denition A [214]) is written as A.
Theorem OBUTR
Orthonormal Basis for Upper Triangular Representation
Suppose that Ais a square matrix. Then there is a unitary matrix U, and an upper triangular matrix T,
such that
UAU=T
andThas the eigenvalues of Aas the entries of the diagonal.
Proof This theorem is a statement about matrices and similarity. We can convert it to a statement
about linear transformations, matrix representations and bases (Theorem SCB [656]). Suppose that Ais
annnmatrix, and dene the linear transformation S:Cn!CnbyS(x) =Ax. Then Theorem UTMR
[676] gives us a basis B=fv1;v2;v3; :::; vngforCnsuch that a matrix representation of Srelative to
B,MS
B;B, is upper triangular.
Now convert the basis Binto an orthogonal basis, C, by an application of the Gram-Schmidt procedure
(Theorem GSP [199]). This is a messy business computationally, but here we have an excellent illustration
of the power of the Gram-Schmidt procedure. We need only be sure that Bis linearly independent and
spans Cn, and then we know that Cis linearly independent, spans Cnand is also an orthogonal set. We
will now consider the matrix representation of Srelative to C(rather than B). Write the new basis as
C=fy1;y2;y3; :::; yng. The application of the Gram-Schmidt procedure creates each vector of C, say
yj, as the dierence of vjand a linear combination of y1;y2;y3; :::; yj 1. We are not concerned here
with the actual values of the scalars in this linear combination, so we will write
yj=vj j 1X
k=1bjkyk
where thebjkare shorthand for the scalars. The equation above is in a form useful for creating the basis
CfromB. To better understand the relationship between BandCconvert it to read
vj=yj+j 1X
k=1bjkyk
In this form, we recognize that the change-of-basis matrix CB;C=MICn
B;C(Denition CBM [648]) is an
upper triangular matrix. By Theorem SCB [656] we have
MS
C;C=CB;CMS
B;BC 1
B;C
The inverse of an upper triangular matrix is upper triangular (Theorem ITMT [676]), and the product
of two upper triangular matrices is again upper triangular (Theorem PTMT [675]). So MS
C;Cis an upper
triangular matrix.
Version 2.30
682 Section OD Orthonormal Diagonalization
Now, multiply each vector of Cby a nonzero scalar, so that the result has norm 1. In this way we
create a new basis Dwhich is an orthonormal set (Denition ONS [201]). Note that the change-of-basis
matrixCC;Dis a diagonal matrix with nonzero entries equal to the norms of the vectors in C.
Now we can convert our results into the language of matrices. Let Ebe the basis of Cnformed with
the standard unit vectors (Denition SUV [197]). Then the matrix representation of Srelative to Eis
simplyA,A=MS
E;E. The change-of-basis matrix CD;Ehas columns that are simply the vectors in D,
the orthonormal basis. As such, Theorem CUMOS [263] tells us that CD;Eis a unitary matrix, and by
Denition UM [262] has an inverse equal to its adjoint. Write U=CD;E. We have
UAU=U 1AU Theorem UMI [263]
=C 1
D;EMS
E;ECD;E
=MS
D;D Theorem SCB [656]
=CC;DMS
C;CC 1
C;DTheorem SCB [656]
The inverse of a diagonal matrix is also a diagonal matrix, and so this nal expression is the product
of three upper triangular matrices, and so is again upper triangular (Theorem PTMT [675]). Thus the
desired upper triangular matrix, T, is the matrix representation of Srelative to the orthonormal basis D,
MS
D;D.
Subsection NM
Normal Matrices
Normal matrices comprise a broad class of interesting matrices, many of which we have met already. But
they are most interesting since they dene exactly which matrices we can diagonalize via a unitary matrix.
This is the upcoming Theorem OD [681]. Here's the denition.
Denition NRML
Normal Matrix
The square matrix Ais normal if AA=AA. 4
So a normal matrix commutes with its adjoint. Part of the beauty of this denition is that it includes
many other types of matrices. A diagonal matrix will commute with its adjoint, since the adjoint is again
diagonal and the entries are just conjugates of the entries of the original diagonal matrix. A Hermitian
(self-adjoint) matrix (Denition HM [234]) will trivially commute with its adjoint, since the two matrices
are the same. A real, symmetric matrix is Hermitian, so these matrices are also normal. A unitary matrix
(Denition UM [262]) has its adjoint as its inverse, and inverses commute (Theorem OSIS [260]), so unitary
matrices are normal. Another class of normal matrices is the skew-symmetric matrices. However, these
broad descriptions still do not capture all of the normal matrices, as the next example shows.
Example ANM
A normal matrix
Let
A=1 1
1 1
Then
1 1
1 11 1
1 1
=2 0
0 2
=1 1
1 11 1
1 1
so we see by Denition NRML [680] that Ais normal. However, Ais not symmetric (hence, as a real
matrix, not Hermitian), not unitary, and not skew-symmetric.
Version 2.30
Subsection OD.OD Orthonormal Diagonalization 683
Subsection OD
Orthonormal Diagonalization
A diagonal matrix is very easy to work with in matrix multiplication (Example HPDM [502]) and an
orthonormal basis also has many advantages (Theorem COB [378]). How about converting a matrix to a
diagonal matrix through a similarity transformation using a unitary matrix (i.e. build a diagonal matrix
representation with an orthonormal matrix)? That'd be fantastic! When can we do this? We can always
accomplish this feat when the matrix is normal, and normal matrices are the only ones that behave this
way. Here's the theorem.
Theorem OD
Orthonormal Diagonalization
Suppose that Ais a square matrix. Then there is a unitary matrix Uand a diagonal matrix D, with
diagonal entries equal to the eigenvalues of A, such that UAU=Dif and only if Ais a normal matrix.
Proof ()) Suppose there is a unitary matrix Uthat diagonalizes A, resulting in D, i.e.UAU=D.
We check the normality of A,
AA=InAInAIn Theorem MMIM [229]
=UUAUUAUUDenition UM [262]
=UUAUDU
=UUA(U)DUTheorem AA [215]
=U(UAU)DUAdjoint of a product
=UDDU
=U
DtDUDenition A [214]
=UDDUDiagonal matrix
=UDDUProperty CMCN [758]
=UD
DtUDiagonal matrix
=UDDUDenition A [214]
=UD(UAU)U
=UDUA(U)UAdjoint of a product
=UDUAUUTheorem AA [215]
=UUAUUAUU
=InAInAIn Denition UM [262]
=AATheorem MMIM [229]
So by Denition NRML [680], Ais a normal matrix.
(() For the converse, suppose that Ais a normal matrix. Whether or not Ais normal, Theorem
OBUTR [679] provides a unitary matrix Uand an upper triangular matrix T, whose diagonal entries are
the eigenvalues of A, and such that UAU=T. With the added condition that Ais normal, we will
determine that the entries of Tabove the diagonal must be all zero. Here we go. First we show that Tis
normal.
TT= (UAU)UAU
=UA(U)UAU Adjoint of a product
=UAUUAU Theorem AA [215]
=UAInAU Denition UM [262]
Version 2.30
684 Section OD Orthonormal Diagonalization
=UAAU Theorem MMIM [229]
=UAAU Denition NRML [680]
=UAInAU Theorem MMIM [229]
=UAUUAU Denition UM [262]
=UAUUA(U)Theorem AA [215]
=UAU(UAU)Adjoint of a product
=TT
So by Denition NRML [680], Tis a normal matrix.
We can translate the normality of Tinto the statement TT TT=O. We now establish an equality
we will use repeatedly. For 1 in,
0 = [O]ii Denition ZM [210]
= [TT TT]ii Denition NRML [680]
= [TT]ii [TT]ii Denition MA [207]
=nX
k=1[T]ik[T]ki nX
k=1[T]ik[T]ki Theorem EMP [227]
=nX
k=1[T]ik[T]ik nX
k=1[T]ki[T]ki Denition A [214]
=nX
k=i[T]ik[T]ik iX
k=1[T]ki[T]ki Denition UTM [675]
=nX
k=ij[T]ikj2 iX
k=1j[T]kij2Denition MCN [760]
To conclude, we use the above equality repeatedly, beginning with i= 1, and discover, row by row, that
the entries above the diagonal of Tare all zero. The key observation is that a sum of squares can only
equal zero when each term of the sum is zero. For i= 1 we have
0 =nX
k=1j[T]1kj2 1X
k=1j[T]k1j2=nX
k=2j[T]1kj2
which forces the conclusions
[T]12= 0 [ T]13= 0 [ T]14= 0 [T]1n= 0
Fori= 2 we use the same equality, but also incorporate the portion of the above conclusions that says
[T]12= 0,
0 =nX
k=2j[T]2kj2 2X
k=1j[T]k2j2=nX
k=2j[T]2kj2 2X
k=2j[T]k2j2=nX
k=3j[T]2kj2
which forces the conclusions
[T]23= 0 [ T]24= 0 [ T]25= 0 [T]2n= 0
We can repeat this process for the subsequent values of i= 3;4;5:::; n 1. Notice that it is critical we
do this in order, since we need to employ portions of each of the previous conclusions about rows having
Version 2.30
Subsection OD.OD Orthonormal Diagonalization 685
zero entries in order to successfully get the same conclusion for later rows. Eventually, we conclude that
all of the nondiagonal entries of Tare zero, so the extra assumption of normality forces Tto be diagonal.
We can rearrange the conclusion of this theorem to read A=UDU. Recall that a unitary matrix
can be viewed as a geometry-preserving transformation (isometry), or more loosely as a rotation of sorts.
Then a matrix-vector product, Ax, can be viewed instead as a sequence of three transformations. Uis
unitary, so is a rotation. Since Dis diagonal, it just multiplies each entry of a vector by a scalar. Diagonal
entries that are positive or negative, with absolute values bigger or smaller than 1 evoke descriptions like
re
ection, expansion and contraction. Generally we can say that D\stretches" a vector in each component.
Final multiplication by Uundoes (inverts) the rotation performed by U. So a normal matrix is a rotation-
stretch-rotation transformation.
The orthonormal basis formed from the columns of Ucan be viewed as a system of mutually perpendic-
ular axes. The rotation by Uallows the transformation by Ato be replaced by the simple transformation
Dalong these axes, and then Dbrings the result back to the original coordinate system. For this reason
Theorem OD [681] is known as the Principal Axis Theorem.
The columns of the unitary matrix in Theorem OD [681] create an especially nice basis for use with
the normal matrix. We record this observation as a theorem.
Theorem OBNM
Orthonormal Bases and Normal Matrices
Suppose that Ais a normal matrix of size n. Then there is an orthonormal basis of Cncomposed of
eigenvectors of A.
Proof LetUbe the unitary matrix promised by Theorem OD [681] and let Dbe the resulting diagonal
matrix. The desired set of vectors is formed by collecting the columns of Uinto a set. Theorem CUMOS
[263] says this set of columns is orthonormal. Since Uis nonsingular (Theorem UMI [263]), Theorem
CNMB [376] says the set is a basis.
SinceAis diagonalized by U, the diagonal entries of the matrix Dare the eigenvalues of A. An
argument exactly like the second half of the proof of Theorem DC [497] shows that each vector of the basis
is an eigenvector of A.
In a vague way Theorem OBNM [683] is an improvement on Theorem HMOE [488] which said that
eigenvectors of a Hermitian matrix for dierent eigenvalues are always orthogonal. Hermitian matrices
are normal and we see that we can nd at least one basis where every pair of eigenvectors is orthogonal.
Notice that this is not a generalization, since Theorem HMOE [488] states a weak result which applies to
many (but not all) pairs of eigenvectors, while Theorem OBNM [683] is a seemingly stronger result, but
only asserts that there is one collection of eigenvectors with the stronger property.
Version 2.30
686 Section OD Orthonormal Diagonalization
Version 2.30
Section NLT Nilpotent Linear Transformations 687
Section NLT
Nilpotent Linear Transformations
This section is in draft form
Nearly complete
We have seen that some matrices are diagonalizable and some are not. Some authors refer to a non-
diagonalizable matrix as defective , but we will study them carefully anyway. Examples of such matrices
include Example EMMS4 [463], Example HMEM5 [465], and Example CEMS6 [466]. Each of these matrices
has at least one eigenvalue with geometric multiplicity strictly less than its algebraic multiplicity, and
therefore Theorem DMFE [499] tells us these matrices are not diagonalizable.
Given a square matrix A, it is likely similar to many, many other matrices. Of all these possibilities,
which is the best? \Best" is a subjective term, but we might agree that a diagonal matrix is certainly a
very nice choice. Unfortunately, as we have seen, this will not always be possible. What form of a matrix is
\next-best"? Our goal, which will take us several sections to reach, is to show that every matrix is similar to
a matrix that is \nearly-diagonal" (Section JCF [721]). More precisely, every matrix is similar to a matrix
with elements on the diagonal, and zeros and ones on the diagonal just above the main diagonal (the
\super diagonal"), with zeros everywhere else. In the language of equivalence relations (see Theorem SER
[494]), we are determining a systematic representative for each equivalence class. Such a representative for
a set of similar matrices is called a canonical form .
We have just discussed the determination of a canonical form as a question about matrices. However,
we know that every square matrix creates a natural linear transformation (Theorem MBLT [522]) and
every linear transformation with identical domain and codomain has a square matrix representation for
each choice of a basis, with a change of basis creating a similarity transformation (Theorem SCB [656]). So
we will state, and prove, theorems using the language of linear transformations on abstract vector spaces,
while most of our examples will work with square matrices. You can, and should, mentally translate
between the two settings frequently and easily.
Subsection NLT
Nilpotent Linear Transformations
We will discover that nilpotent linear transformations are the essential obstacle in a non-diagonalizable
linear transformation. So we will study them carefully rst, both as an object of inherent mathematical
interest, but also as the object at the heart of the argument that leads to a pleasing canonical form for
any linear transformation. Once we understand these linear transformations thoroughly, we will be able
to easily analyze the structure of any linear transformation.
Denition NLT
Nilpotent Linear Transformation
Suppose that T:V!Vis a linear transformation such that there is an integer p>0 such that Tp(v) =0
for every v2V. The smallest pfor which this condition is met is called the index ofT.4
Of course, the linear transformation Tdened byT(v) =0will qualify as nilpotent of index 1. But
are there others?
Example NM64
Nilpotent matrix, size 6, index 4
Recall that our denitions and theorems are being stated for linear transformations on abstract vector
spaces, while our examples will work with square matrices (and use the same terms interchangeably). In
Version 2.30
688 Section NLT Nilpotent Linear Transformations
this case, to demonstrate the existence of nontrivial nilpotent linear transformations, we desire a matrix
such that some power of the matrix is the zero matrix. Consider
A=2
6666664 3 3 2 5 0 5
3 5 3 4 3 9
3 4 2 6 4 3
3 3 2 5 0 5
3 3 2 4 2 6
2 3 2 2 4 73
7777775
and compute powers of A,
A2=2
66666641 2 1 0 3 4
0 2 1 1 3 4
3 0 0 3 0 0
1 2 1 0 3 4
0 2 1 1 3 4
1 2 1 2 3 43
7777775
A3=2
66666641 0 0 1 0 0
1 0 0 1 0 0
0 0 0 0 0 0
1 0 0 1 0 0
1 0 0 1 0 0
1 0 0 1 0 03
7777775
A4=2
66666640 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 03
7777775
Thus we can say that Ais nilpotent of index 4.
Because it will presage some upcoming theorems, we will record some extra information about the
eigenvalues and eigenvectors of Ahere.Ahas just one eigenvalue, = 0, with algebraic multiplicity 6 and
geometric multiplicity 2. The eigenspace for this eigenvalue is
EA(0) =*2
66666642
2
5
2
1
03
7777775;2
6666664 1
1
5
1
0
13
7777775+
If there were degrees of singularity, we might say this matrix was very singular, since zero is an eigenvalue
with maximum algebraic multiplicity (Theorem SMZE [480], Theorem ME [485]). Notice too that Ais
\far" from being diagonalizable (Theorem DMFE [499]).
Another example.
Example NM62
Nilpotent matrix, size 6, index 2
Version 2.30
Subsection NLT.NLT Nilpotent Linear Transformations 689
Consider the matrix
B=2
6666664 1 1 1 4 3 1
1 1 1 2 3 1
9 10 5 9 5 15
1 1 1 4 3 1
1 1 0 2 4 2
4 3 1 1 5 53
7777775
and compute the second power of B,
B2=2
66666640 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 03
7777775
SoBis nilpotent of index 2. Again, the only eigenvalue of Bis zero, with algebraic multiplicity 6. The
geometric multiplicity of the eigenvalue is 3, as seen in the eigenspace,
EB(0) =*2
66666641
3
6
1
0
03
7777775;2
66666640
4
7
0
1
03
7777775;2
66666640
2
1
0
0
13
7777775+
Again, Theorem DMFE [499] tells us that Bis far from being diagonalizable.
On a rst encounter with the denition of a nilpotent matrix, you might wonder if such a thing was
possible at all. That a high power of a nonzero object could be zero is so very dierent from our experience
with scalars that it seems very unnatural. Hopefully the two previous examples were somewhat surprising.
But we have seen that matrix algebra does not always behave the way we expect (Example MMNC [227]),
and we also now recognize matrix products not just as arithmetic, but as function composition (Theorem
MRCLT [622]). We will now turn to some examples of nilpotent matrices which might be more transparent.
Denition JB
Jordan Block
Given the scalar 2C, the Jordan block Jn() is thennmatrix dened by
[Jn()]ij=8
><
>: i =j
1j=i+ 1
0 otherwise
(This denition contains Notation JB.) 4
Example JB4
Jordan block, size 4
A simple example of a Jordan block,
J4(5) =2
6645 1 0 0
0 5 1 0
0 0 5 1
0 0 0 53
775
Version 2.30
690 Section NLT Nilpotent Linear Transformations
We will return to general Jordan blocks later, but in this section we are just interested in Jordan blocks
where= 0. Here's an example of why we are specializing in these matrices now.
Example NJB5
Nilpotent Jordan block, size 5
Consider
J5(0) =2
666640 1 0 0 0
0 0 1 0 0
0 0 0 1 0
0 0 0 0 1
0 0 0 0 03
77775
and compute powers,
(J5(0))2=2
666640 0 1 0 0
0 0 0 1 0
0 0 0 0 1
0 0 0 0 0
0 0 0 0 03
77775
(J5(0))3=2
666640 0 0 1 0
0 0 0 0 1
0 0 0 0 0
0 0 0 0 0
0 0 0 0 03
77775
(J5(0))4=2
666640 0 0 0 1
0 0 0 0 0
0 0 0 0 0
0 0 0 0 0
0 0 0 0 03
77775
(J5(0))5=2
666640 0 0 0 0
0 0 0 0 0
0 0 0 0 0
0 0 0 0 0
0 0 0 0 03
77775
SoJ5(0) is nilpotent of index 5. As before, we record some information about the eigenvalues and eigen-
vectors of this matrix. The only eigenvalue is zero, with algebraic multiplicity 5, the maximum possible
(Theorem ME [485]). The geometric multiplicity of this eigenvalue is just 1, the minimum possible (The-
orem ME [485]), as seen in the eigenspace,
EJ5(0)(0) =*2
666641
0
0
0
03
77775+
There should not be any real surprises in this example. We can watch the ones in the powers of J5(0)
slowly march o to the upper-right hand corner of the powers. In some vague way, the eigenvalues and
Version 2.30
Subsection NLT.NLT Nilpotent Linear Transformations 691
eigenvectors of this matrix are equally extreme.
We can form combinations of Jordan blocks to build a variety of nilpotent matrices. Simply place
Jordan blocks on the diagonal of a matrix with zeros everywhere else, to create a block diagonal matrix.
Example NM83
Nilpotent matrix, size 8, index 3
Consider the matrix
C=2
4J3(0)O O
OJ3(0)O
O O J2(0)3
5=2
666666666640 1 0 0 0 0 0 0
0 0 1 0 0 0 0 0
0 0 0 0 0 0 0 0
0 0 0 0 1 0 0 0
0 0 0 0 0 1 0 0
0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 1
0 0 0 0 0 0 0 03
77777777775
and compute powers,
C2=2
666666666640 0 1 0 0 0 0 0
0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0
0 0 0 0 0 1 0 0
0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 03
77777777775
C3=2
666666666640 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 03
77777777775
SoCis nilpotent of index 3. You should notice how block diagonal matrices behave in products (much like
diagonal matrices) and that it was the largest Jordan block that determined the index of this combination.
All eight eigenvalues are zero, and each of the three Jordan blocks contributes one eigenvector to a basis
for the eigenspace, resulting in zero having a geometric multiplicity of 3.
It would appear that nilpotent matrices only have zero as an eigenvalue, so the algebraic multiplicity
will be the maximum possible. However, by creating block diagonal matrices with Jordan blocks on the
diagonal you should be able to attain any desired geometric multiplicity for this lone eigenvalue. Likewise,
the size of the largest Jordan block employed will determine the index of the matrix. So nilpotent matrices
with various combinations of index and geometric multiplicities are easy to manufacture. The predictable
properties of block diagonal matrices in matrix products and eigenvector computations, along with the
next theorem, make this possible. You might nd Example NJB5 [688] a useful companion to this proof.
Theorem NJB
Nilpotent Jordan Blocks
The Jordan block Jn(0) is nilpotent of index n.
Proof While not phrased as an if-then statement, the statement in the theorem is understood to mean
that if we have a specic matrix ( Jn(0)) then we need to establish it is nilpotent of a specied index. The
Version 2.30
692 Section NLT Nilpotent Linear Transformations
rst column of Jn(0) is the zero vector, and the remaining n 1 columns are the standard unit vectors ei,
1in 1 (Denition SUV [197]), which are also the rst n 1 columns of the size nidentity matrix
In. As shorthand, write J=Jn(0).
J= [0je1je2je3j:::jen 1]
We will use the denition of matrix multiplication (Denition MM [226]), together with a proof by induction
(Technique I [772]), to study the powers of J. Our claim is that
Jk= [0j0j:::j0je1je2j:::jen k]
for 1kn. For the base case, k= 1, and the denition of J1=Jn(0) establishes the claim. For the
induction step, rst note that Je1=0andJei=ei 1for 2in. Then, assuming the claim is true for
k, we examine the k+ 1 case,
Jk+1=JJk
=J[0j0j:::j0je1je2j:::jen k] Induction Hypothesis
= [J0jJ0j:::jJ0jJe1jJe2j:::jJen k] Denition MM [226]
= [0j0j:::j0j0je1je2j:::jen k 1] Denition MVP [223]
=
0j0j:::j0je1je2j:::en (k+1)
This concludes the induction. So Jkhas a nonzero entry (a one) in row n kand column n, for 1kn 1,
and is therefore a nonzero matrix. However, Jn= [0j0j:::j0] =O. By Denition NLT [685], Jis nilpotent
of indexn.
Subsection PNLT
Properties of Nilpotent Linear Transformations
In this subsection we collect some basic properties of nilpotent linear transformations. After studying the
examples in the previous section, some of these will be no surprise.
Theorem ENLT
Eigenvalues of Nilpotent Linear Transformations
Suppose that T:V!Vis a nilpotent linear transformation and is an eigenvalue of T. Then= 0.
Proof Letxbe an eigenvector of Tfor the eigenvalue , and suppose that Tis nilpotent with index p.
Then
0=Tp(x) Denition NLT [685]
=px Theorem EOMP [481]
Because xis an eigenvector, it is nonzero, and therefore Theorem SMEZV [326] tells us that p= 0 and
so= 0.
Paraphrasing, all of the eigenvalues of a nilpotent linear transformation are zero. So in particular,
the characteristic polynomial of a nilpotent linear transformation, T, on a vector space of dimension n, is
simplypT(x) =xn.
The next theorem is not critical for what follows, but it will explain our interest in nilpotent linear
transformations. More specically, it is the rst step in backing up the assertion that nilpotent linear trans-
formations are the essential obstacle in a non-diagonalizable linear transformation. While it is not obvious
from the statement of the theorem, it says that a nilpotent linear transformation is not diagonalizable,
unless it is trivially so.
Version 2.30
Subsection NLT.PNLT Properties of Nilpotent Linear Transformations 693
Theorem DNLT
Diagonalizable Nilpotent Linear Transformations
Suppose the linear transformation T:V!Vis nilpotent. Then Tis diagonalizable if and only Tis the
zero linear transformation.
Proof We start with the easy direction. Let n= dim (V).
(() The linear transformation Z:V!Vdened by Z(v) =0for all v2Vis nilpotent of index
p= 1 and a matrix representation relative to any basis of Vis thennzero matrix,O. Quite obviously,
the zero matrix is a diagonal matrix (Denition DIM [496]) and hence Zis diagonalizable (Denition DZM
[496]).
()) Assume now that Tis diagonalizable, so
T() =T() for every eigenvalue (Theorem DMFE
[499]). By Theorem ENLT [690], Thas only one eigenvalue (zero), which therefore must have algebraic
multiplicity n(Theorem NEM [485]). So the geometric multiplicity of zero will be nas well,
T(0) =n.
LetBbe a basis for the eigenspace ET(0). ThenBis a linearly independent subset of Vof sizen, and
by Theorem G [407] will be a basis for V. For any x2Bwe have
T(x) = 0x Denition EM [461]
=0 Theorem ZSSM [324]
SoTis identically zero on a basis for B, and since the action of a linear transformation on a basis determines
all of the values of the linear transformation (Theorem LTDB [525]), it is easy to see that T(v) =0for
every v2V.
So, other than one trivial case (the zero matrix), every nilpotent linear transformation is not diag-
onalizable. It remains to see what is so \essential" about this broad class of non-diagonalizable linear
transformations. For this we now turn to a discussion of kernels of powers of nilpotent linear transforma-
tions, beginning with a result about general linear transformations that may not necessarily be nilpotent.
Theorem KPLT
Kernels of Powers of Linear Transformations
SupposeT:V!Vis a linear transformation, where dim ( V) =n. Then there is an integer m, 0mn,
such that
f0g=K
T0
(K
T1
(K
T2
((K(Tm) =K
Tm+1
=K
Tm+2
=
Proof There are several items to verify in the conclusion as stated. First, we show that K
Tk
K
Tk+1
for anyk. Choose z2K
Tk
. Then
Tk+1(z) =T
Tk(z)
Denition LTC [532]
=T(0) Denition KLT [545]
=0 Theorem LTTZZ [519]
So by Denition KLT [545], z2K
Tk+1
and by Denition SSET [761] we have K
Tk
K
Tk+1
.
Second, we demonstrate the existence of a power mwhere consecutive powers result in equal kernels.
A by-product will be the condition that mcan be chosen so that mn. To the contrary, suppose that
f0g=K
T0
(K
T1
(K
T2
((K
Tn 1
(K(Tn)(K
Tn+1
(
SinceK
Tk
(K
Tk+1
, Theorem PSSD [410] implies that dim
K
Tk+1
dim
K
Tk
+ 1. Repeated
application of this observation yields
dim
K
Tn+1
dim (K(Tn)) + 1
Version 2.30
694 Section NLT Nilpotent Linear Transformations
dim
K
Tn 1
+ 2
...
dim
K
T0
+ (n+ 1)
= dim (f0g) +n+ 1
=n+ 1
Thus,K
Tn+1
has a basis of size at least n+ 1, which is a linearly independent set of size greater than n
in the vector space Vof dimension n. This contradicts Theorem G [407].
This contradiction yields the existence of an integer ksuch thatK
Tk
=K
Tk+1
, so we can dene
mto be smallest such integer with this property. From the argument above about dimensions resulting
from a strictly increasing chain of subspaces, it should be clear that mn.
It remains to show that once two consecutive kernels are equal, then all of the remaining kernels are
equal. More formally, if K(Tm) =K
Tm+1
, thenK(Tm) =K
Tm+j
for allj1. We will give a proof
by induction on j(Technique I [772]). The base case ( j= 1) is precisely our dening property for m.
In the induction step, we assume that K(Tm) =K
Tm+j
and endeavor to show that K(Tm) =
K
Tm+j+1
. At the outset of this proof we established that K(Tm)K
Tm+j+1
. So Denition SE
[762] requires only that we establish the subset inclusion in the opposite direction. To wit, choose z2
K
Tm+j+1
. Then
0=Tm+j+1(z) Denition KLT [545]
=Tm+j(T(z)) Denition LTC [532]
=Tm(T(z)) Induction Hypothesis
=Tm+1(z) Denition LTC [532]
=Tm(z) Base Case
So by Denition KLT [545], z2K(Tm) as desired.
We now specialize Theorem KPLT [691] to the case of nilpotent linear transformations, which buys us
just a bit more precision in the conclusion.
Theorem KPNLT
Kernels of Powers of Nilpotent Linear Transformations
SupposeT:V!Vis a nilpotent linear transformation with index pand dim (V) =n. Then 0pn
and
f0g=K
T0
(K
T1
(K
T2
((K(Tp) =K
Tp+1
==V
Proof SinceTp= 0 it follows that Tp+j= 0 for allj0 and thusK
Tp+j
=Vforj0. So the value
ofmguaranteed by Theorem KPLT [691] is at most p. The only remaining aspect of our conclusion that
does not follow from Theorem KPLT [691] is that m=p. To see this we must show that K
Tk
(K
Tk+1
for 0kp 1. IfK
Tk
=K
Tk+1
for somek < p , thenK
Tk
=K(Tp) =V. This implies that
Tk= 0, violating the fact that Thas indexp. So the smallest value of mis indeedp, and we learn that
p<n .
The structure of the kernels of powers of nilpotent linear transformations will be crucial to what follows.
But immediately we can see a practical benet. Suppose we are confronted with the question of whether
or not annnmatrix,A, is nilpotent or not. If we don't quickly nd a low power that equals the zero
matrix, when do we stop trying higher and higher powers? Theorem KPNLT [692] gives us the answer: if
we don't see a zero matrix by the time we nish computing An, then it is not going to ever happen. We'll
now take a look at one example of Theorem KPNLT [692] in action.
Version 2.30
Subsection NLT.PNLT Properties of Nilpotent Linear Transformations 695
Example KPNLT
Kernels of powers of a nilpotent linear transformation
We will recycle the nilpotent matrix Aof index 4 from Example NM64 [685]. We now know that would
have only needed to look at the rst 6 powers of Aif the matrix had not been nilpotent. We list bases for
the null spaces of the powers of A. (Notice how we are using null spaces for matrices interchangeably with
kernels of linear transformations, see Theorem KNSI [625] for justication.)
N(A) =N0
BBBBBB@2
6666664 3 3 2 5 0 5
3 5 3 4 3 9
3 4 2 6 4 3
3 3 2 5 0 5
3 3 2 4 2 6
2 3 2 2 4 73
77777751
CCCCCCA=*8
>>>>>><
>>>>>>:2
66666642
2
5
2
1
03
7777775;2
6666664 1
1
5
1
0
13
77777759
>>>>>>=
>>>>>>;+
N
A2
=N0
BBBBBB@2
66666641 2 1 0 3 4
0 2 1 1 3 4
3 0 0 3 0 0
1 2 1 0 3 4
0 2 1 1 3 4
1 2 1 2 3 43
77777751
CCCCCCA=*8
>>>>>><
>>>>>>:2
66666640
1
2
0
0
03
7777775;2
66666642
1
0
2
0
03
7777775;2
66666640
3
0
0
2
03
7777775;2
66666640
2
0
0
0
13
77777759
>>>>>>=
>>>>>>;+
N
A3
=N0
BBBBBB@2
66666641 0 0 1 0 0
1 0 0 1 0 0
0 0 0 0 0 0
1 0 0 1 0 0
1 0 0 1 0 0
1 0 0 1 0 03
77777751
CCCCCCA=*8
>>>>>><
>>>>>>:2
66666640
1
0
0
0
03
7777775;2
66666640
0
1
0
0
03
7777775;2
66666641
0
0
1
0
03
7777775;2
66666640
0
0
0
1
03
7777775;2
66666640
0
0
0
0
13
77777759
>>>>>>=
>>>>>>;+
N
A4
=N0
BBBBBB@2
66666640 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 03
77777751
CCCCCCA=*8
>>>>>><
>>>>>>:2
66666641
0
0
0
0
03
7777775;2
66666640
1
0
0
0
03
7777775;2
66666640
0
1
0
0
03
7777775;2
66666640
0
0
1
0
03
7777775;2
66666640
0
0
0
1
03
7777775;2
66666640
0
0
0
0
13
77777759
>>>>>>=
>>>>>>;+
With the exception of some convenience scaling of the basis vectors in N
A2
these are exactly the basis
vectors described in Theorem BNS [160]. We can see that the dimension of N(A) equals the geometric
multiplicity of the zero eigenvalue. Why is this not an accident? We can see the dimensions of the kernels
consistently increasing, and we can see that N
A4
=C6. But Theorem KPNLT [692] says a little more.
Each successive kernel should be a superset of the previous one. We ought to be able to begin with a basis
ofN(A) and extend it to a basis of N
A2
. Then we should be able to extend a basis of N
A2
into a
basis ofN
A3
, all with repeated applications of Theorem ELIS [407]. Verify the following,
N(A) =*8
>>>>>><
>>>>>>:2
66666642
2
5
2
1
03
7777775;2
6666664 1
1
5
1
0
13
77777759
>>>>>>=
>>>>>>;+
Version 2.30
696 Section NLT Nilpotent Linear Transformations
N
A2
=*8
>>>>>><
>>>>>>:2
66666642
2
5
2
1
03
7777775;2
6666664 1
1
5
1
0
13
7777775;2
66666640
3
0
0
2
03
7777775;2
66666640
2
0
0
0
13
77777759
>>>>>>=
>>>>>>;+
N
A3
=*8
>>>>>><
>>>>>>:2
66666642
2
5
2
1
03
7777775;2
6666664 1
1
5
1
0
13
7777775;2
66666640
3
0
0
2
03
7777775;2
66666640
2
0
0
0
13
7777775;2
66666640
0
0
0
0
13
77777759
>>>>>>=
>>>>>>;+
N
A4
=*8
>>>>>><
>>>>>>:2
66666642
2
5
2
1
03
7777775;2
6666664 1
1
5
1
0
13
7777775;2
66666640
3
0
0
2
03
7777775;2
66666640
2
0
0
0
13
7777775;2
66666640
0
0
0
0
13
7777775;2
66666640
0
0
1
0
03
77777759
>>>>>>=
>>>>>>;+
Do not be concerned at the moment about how these bases were constructed since we are not describing
the applications of Theorem ELIS [407] here. Do verify carefully for each alleged basis that, (1) it is a
superset of the basis for the previous kernel, (2) the basis vectors really are members of the kernel of the
right power of A, (3) the basis is a linearly independent set, (4) the size of the basis is equal to the size of
the basis found previously for each kernel. With these verications, Theorem G [407] will tell us that we
have successfully demonstrated what Theorem KPNLT [692] guarantees.
Subsection CFNLT
Canonical Form for Nilpotent Linear Transformations
Our main purpose in this section is to nd a basis so that a nilpotent linear transformation will have a
pleasing, nearly-diagonal matrix representation. Of course, we will not have a denition for \pleasing," nor
for \nearly-diagonal." But the short answer is that our preferred matrix representation will be built up
from Jordan blocks, Jn(0). Here's the theorem. You will nd Example CFNLT [698] helpful as you study
this proof, since it uses the same notation, and is large enough to (barely) illustrate the full generality of
the theorem (see ).
Theorem CFNLT
Canonical Form for Nilpotent Linear Transformations
Suppose that T:V!Vis a nilpotent linear transformation of index p. Then there is a basis for Vso
that the matrix representation, MT
B;B, is block diagonal with each block being a Jordan block, Jn(0). The
size of the largest block is the index p, and the total number of blocks is the nullity of T,n(T).
Proof We will explicitly construct the desired basis, so the proof is constructive (Technique C [768]),
and can be used in practice. As we begin, the basis vectors will not be in the proper order, but we will
rearrange them at the end of the proof. For convenience, dene ni=n
Ti
, so for example, n0= 0,
n1=n(T) andnp=n(Tp) = dim (V). Denesi=ni ni 1, for 1ip, so we can think of sias
\how much bigger" K
Ti
is thanK
Ti 1
. In particular, Theorem KPNLT [692] implies that si>0 for
1ip.
Version 2.30
Subsection NLT.CFNLT Canonical Form for Nilpotent Linear Transformations 697
We are going to build a set of vectors zi;j, 1ip, 1jsi. Each zi;jwill be an element of
K
Ti
and not an element of K
Ti 1
. In total, we will obtain a linearly independent set ofPp
i=1si=Pp
i=1ni ni 1=np n0= dim (V) vectors that form a basis of V. We construct this set in pieces, starting
at the \wrong" end. Our procedure will build a series of subspaces, Zi, each lying in between K
Ti 1
and
K
Ti
, having bases zi;j, 1jsi, and which together equal Vas a direct sum. Now would be a good
time to review the results on direct sums collected in Subsection PD.DS [413]. OK, here we go.
We build the subspace Zprst (this is what we meant by \starting at the wrong end"). K
Tp 1
is
a proper subspace of K(Tp) =V(Theorem KPNLT [692]). Theorem DSFOS [414] says that there is a
subspace of Vthat will pair with the subspace K
Tp 1
to form a direct sum of V. Call this subspace
Zp, and choose vectors zp;j, 1jspas a basis of Zp, which we will denote as Bp. Note that we have a
fair amount of freedom in how to choose these rst basis vectors. Several observations will be useful in the
next step. First V=K
Tp 1
Zp. The basis Bp=
zp;1;zp;2;zp;3; :::; zp;sp
is linearly independent.
For 1jsp,zp;j2K(Tp) =V. Since the two subspaces of a direct sum have no nonzero vectors in
common (Theorem DSZI [415]), for 1 jsp,zp;j62K
Tp 1
. That was comparably easy.
If obtaining Zpwas easy, getting Zp 1will be harder. We will repeat the next step p 1 times, and
so will do it carefully the rst time. Eventually, Zp 1will have dimension sp 1. However, the rst sp
vectors of a basis are straightforward. Dene zp 1;j=T(zp;j), 1jsp. Notice that we have no choice
in creating these vectors, they are a consequence of our choices for zp;j. In retrospect (i.e. on a second
reading of this proof), you will recognize this as the key step in realizing a matrix representation of a
nilpotent linear transformation with Jordan blocks. We need to know that this set of vectors in linearly
independent, so start with a relation of linear dependence (Denition RLD [351]), and massage it,
0=a1zp 1;1+a2zp 1;2+a3zp 1;3++aspzp 1;sp
=a1T(zp;1) +a2T(zp;2) +a3T(zp;3) ++aspT
zp;sp
=T
a1zp;1+a2zp;2+a3zp;3++aspzp;sp
Theorem LTLC [525]
Dene x=a1zp;1+a2zp;2+a3zp;3++aspzp;sp. The statement just above means that x2K(T)K
Tp 1
(Denition KLT [545], Theorem KPNLT [692]). As dened, xis a linear combination of the basis vectors
Bp, and therefore x2Zp. Thus x2K
Tp 1
\Zp(Denition SI [763]). Because V=K
Tp 1
Zp,
Theorem DSZI [415] tells us that x=0. Now we recognize the denition of xas a relation of linear
dependence on the linearly independent set Bp, and therefore a1=a2==asp= 0 (Denition LI
[351]). This establishes the linear independence of zp 1;j, 1jsp(Denition LI [351]).
We also need to know where the vectors zp 1;j, 1jsplive. First we demonstrate that they are
members ofK
Tp 1
.
Tp 1(zp 1;j) =Tp 1(T(zp;j))
=Tp(zp;j)
=0
Sozp 1;j2K
Tp 1
, 1jsp. However, we now show that these vectors are not elements of K
Tp 2
.
Suppose to the contrary (Technique CD [770]) that zp 1;j2K
Tp 2
. Then
0=Tp 2(zp 1;j)
=Tp 2(T(zp;j))
=Tp 1(zp;j)
which contradicts the earlier statement that zp;j62K
Tp 1
. Sozp 1;j62K
Tp 2
, 1jsp.
Version 2.30
698 Section NLT Nilpotent Linear Transformations
Now choose a basis Cp 2=
u1;u2;u3; :::; unp 2
forK
Tp 2
. We want to extend this basis by
adding in the zp 1;jto span a subspace of K
Tp 1
. But rst we want to know that this set is linearly
independent. Let ak, 1knp 2andbj, 1jspbe the scalars in a relation of linear dependence,
0=a1u1+a2u2++anp 2unp 2+b1zp 1;1+b2zp 1;2++bspzp 1;sp
Then,
0=Tp 2(0)
=Tp 2
a1u1+a2u2++anp 2unp 2+b1zp 1;1+b2zp 1;2++bspzp 1;sp
=a1Tp 2(u1) +a2Tp 2(u2) ++anp 2Tp 2
unp 2
+
b1Tp 2(zp 1;1) +b2Tp 2(zp 1;2) ++bspTp 2
zp 1;sp
=a10+a20++anp 20+b1Tp 2(zp 1;1) +b2Tp 2(zp 1;2) ++bspTp 2
zp 1;sp
=b1Tp 2(zp 1;1) +b2Tp 2(zp 1;2) ++bspTp 2
zp 1;sp
=b1Tp 2(T(zp;1)) +b2Tp 2(T(zp;2)) ++bspTp 2
T
zp;sp
=b1Tp 1(zp;1) +b2Tp 1(zp;2) ++bspTp 1
zp;sp
=Tp 1
b1zp;1+b2zp;2++bspzp;sp
Dene y=b1zp;1+b2zp;2++bspzp;sp. The statement just above means that y2K
Tp 1
(Denition
KLT [545]). As dened, yis a linear combination of the basis vectors Bp, and therefore y2Zp. Thus
y2K
Tp 1
\Zp. BecauseV=K
Tp 1
Zp, Theorem DSZI [415] tells us that y=0. Now we recognize
the denition of yas a relation of linear dependence on the linearly independent set Bp, and therefore
b1=b2==bsp= 0 (Denition LI [351]). Return to the full relation of linear dependence with both sets
of scalars (the aiandbj). Now that we know that bj= 0 for 1jsp, this relation of linear dependence
simplies to a relation of linear dependence on just the basis Cp 1. Therefore, ai= 0, 1ainp 1and
we have the desired linear independence.
Dene a new subspace of K
Tp 1
as
Qp 1=
u1;u2;u3; :::; unp 1;zp 1;1;zp 1;2;zp 1;3; :::; zp 1;sp
By Theorem DSFOS [414] there exists a subspace of K
Tp 1
which will pair with Qp 1to form a direct
sum. Call this subspace Rp 1, so by denition, K
Tp 1
=Qp 1Rp 1. We are interested in the dimension
ofRp 1. Note rst, that since the spanning set of Qp 1is linearly independent, dim ( Qp 1) =np 2+sp.
Then
dim (Rp 1) = dim
K
Tp 1
dim (Qp 1) Theorem DSD [416]
=np 1 (np 2+sp)
= (np 1 np 2) sp
=sp 1 sp
Notice that if sp 1=sp, thenRp 1is trivial. Now choose a basis of Rp 1, and denote these sp 1 sp
vectors as zp 1;sp+1,zp 1;sp+2,zp 1;sp+3, . . . , zp 1;sp 1. This is another occassion to notice that we have
some freedom in this choice.
We now haveK
Tp 1
=Qp 1Rp 1, and we have bases for each of the two subspaces. The union of
these two bases will therefore be a linearly independent set in K
Tp 1
with size
(np 2+sp) + (sp 1 sp) =np 2+sp 1
=np 2+np 1 np 2
=np 1= dim
K
Tp 1
Version 2.30
Subsection NLT.CFNLT Canonical Form for Nilpotent Linear Transformations 699
So, by Theorem G [407], the following set is a basis of K
Tp 1
,
u1;u2;u3; :::; unp 2;zp 1;1;zp 1;2; :::; zp 1;sp;zp 1;sp+1;zp 1;sp+2; :::; zp 1;sp 1
We built up this basis in three parts, we will now split it in half. Dene the subspace Zp 1by
Zp 1=hBp 1i=
zp 1;1;zp 1;2; :::; zp 1;sp 1
where we have implicitly denoted the basis as Bp 1. Then Theorem DSFB [413] allows us to split up the
basis forK
Tp 1
asCp 1[Bp 1and write
K
Tp 1
=K
Tp 2
Zp 1
Whew! This is a good place to recap what we have achieved. The vectors zi;jform bases for the subspaces
Ziand right now
V=K
Tp 1
Zp=K
Tp 2
Zp 1Zp
The key feature of this decomposition of Vis that the rst spvectors in the basis for Zp 1are outputs of
the linear transformation Tusing the basis vectors of Zpas inputs.
Now we want to further decompose K
Tp 2
(intoK
Tp 3
andZp 2). The procedure is the same as
above, so we will only sketch the key steps. Checking the details proceeds in the same manner as above.
Technically, we could have set up the preceding as the induction step in a proof by induction (Technique
I [772]), but this probably would make the proof harder to understand.
Hit each element of Bp 1withT, to create vectors zp 2;j, 1jsp 1. These vectors form a linearly
independent set, and each is an element of K
Tp 2
, but not an element of K
Tp 3
. Grab a basis Cp 3
ofK
Tp 3
and tack on the newly-created vectors zp 2;j, 1jsp 1. This expanded set is linearly
independent, and we can dene a subspace Qp 2using it as a basis. Theorem DSFOS [414] gives us a
subspaceRp 2such thatK
Tp 2
=Qp 2Rp 2. Vectors zp 2;j,sp 1+ 1jsp 2are chosen as a
basis forRp 2once the relevant dimensions have been veried. The union of Cp 3andzp 2;j, 1jsp 2
then form a basis of K
Tp 2
, which can be split into two parts to yield the decomposition
K
Tp 2
=K
Tp 3
Zp 2
HereZp 2is the subspace of K
Tp 2
with basisBp 2=fzp 2;jj1jsp 2g. Finally,
V=K
Tp 1
Zp=K
Tp 2
Zp 1Zp=K
Tp 3
Zp 2Zp 1Zp
Again, the key feature of this decomposition is that the rst vectors in the basis of Zp 2are outputs of T
using vectors from the basis Zp 1as inputs (and in turn, some of these inputs are outputs of Tderived
from inputs in Zp).
Now assume we repeat this procedure until we decompose K
T2
into subspacesK(T) andZ2. Finally,
decomposeK(T) into subspaces K
T0
=K(In) =f0gandZ1, so that we recognize the vectors z1;j,
1js1=n1as elements ofK(T). The set
B=B1[B2[B3[[Bp=fzi;jj1ip;1jsig
is linearly independent by Theorem DSLI [416] and has size
pX
i=1si=pX
i=1ni ni 1=np n0= dim (V)
So by Theorem G [407], Bis a basis of V. We desire a matrix representation of Trelative toB(Denition
MR [615]), but rst we will reorder the elements of B. The following display lists the elements of Bin
Version 2.30
700 Section NLT Nilpotent Linear Transformations
the desired order, when read across the rows left-to-right in the usual way. Notice that we established the
existence of these vectors column-by-column, and beginning on the right.
z1;1 z2;1 z3;1 zp;1
z1;2 z2;2 z3;2 zp;2
......
z1;sp z2;sp z3;sp zp;sp
z1;sp+1 z2;sp+1 z3;sp+1
......
z1;s3 z2;s3 z3;s3
...
z1;s2 z2;s2
...
z1;s1
It is dicult to layout this table with the notation we have been using, but it would not be especially
useful to invent some notation to overcome the diculty. (One approach would be to dene something like
the inverse of the nonincreasing function, i!si.) Do notice that there are s1=n1rows andpcolumns.
Columniis the basis Bi. The vectors in the rst column are elements of K(T). Each row is the same
length, or shorter, than the one above it. If we apply Tto any vector in the table, other than those in the
rst column, the output is the preceding vector in the row.
Now contemplate the matrix representation of Trelative toBas we read across the rows of the table
above. In the rst row, T(z1;1) =0, so the rst column of the representation is the zero column. Next,
T(z2;1) =z1;1, so the second column of the representation is a vector with a single one in the rst entry,
and zeros elsewhere. Next, T(z3;1) =z2;1, so column 3 of the representation is a zero, then a one, then
all zeros. Continuing in this vein, we obtain the rst pcolumns of the representation, which is the Jordan
blockJp(0) followed by rows of zeros.
When we apply Tto the basis vectors of the second row, what happens? Applying Tto the rst vector,
the result is the zero vector, so the representation gets a zero column. Applying Tto the second vector in
the row, the output is simply the rst vector in that row, making the next column of the representation
all zeros plus a lone one, sitting just above the diagonal. Continuing, we create a Jordan block, sitting on
the diagonal of the matrix representation. It is not possible in general to state the size of this block, but
since the second row is no longer than the rst, it cannot have size larger than p.
Since there are as many rows as the dimension of K(T), the representation contains as many Jordan
blocks as the nullity of T,n(T). Each successive block is smaller than the preceding one, with the rst,
and largest, having size p. The blocks are Jordan blocks since the basis vectors zi;jwere often dened as
the result of applying Tto other elements of the basis already determined, and then we rearranged the
basis into an order that placed outputs of Tjust before their inputs, excepting the start of each row, which
was an element of K(T).
The proof of Theorem CFNLT [694] is constructive (Technique C [768]), so we can use it to create bases
of nilpotent linear transformations with pleasing matrix representations. Recall that Theorem DNLT [691]
told us that nilpotent linear transformations are almost never diagonalizable, so this is progress. As we
have hinted before, with a nice representation of nilpotent matrices, it will not be dicult to build up
representations of other non-diagonalizable matrices. Here is the promised example which illustrates the
previous theorem. It is a useful companion to your study of the proof of Theorem CFNLT [694].
Version 2.30
Subsection NLT.CFNLT Canonical Form for Nilpotent Linear Transformations 701
Example CFNLT
Canonical form for a nilpotent linear transformation
The 66 matrix,A, of Example NM64 [685] is nilpotent of index p= 4. If we dene the linear trans-
formationT:C6!C6byT(x) =Ax, thenTis nilpotent of index 4 and we can seek a basis of C6that
yields a matrix representation with Jordan blocks on the diagonal. The nullity of Tis 2, so from Theorem
CFNLT [694] we can expect the largest Jordan block to be J4(0), and there will be just two blocks. This
only leaves enough room for the second block to have size 2.
We will recycle the bases for the null spaces of the powers of Afrom Example KPNLT [693] rather than
recomputing them here. We will also use the same notation used in the proof of Theorem CFNLT [694].
To begin,s4=n4 n3= 6 5 = 1, so we need one vector of K
T4
=C6, that is not inK
T3
, to
be a basis for Z4. We have a lot of latitude in this choice, and we have not described any sure-re method
for constructing a vector outside of a subspace. Looking at the basis for K
T3
we see that if a vector is
in this subspace, and has a nonzero value in the rst entry, then it must also have a nonzero value in the
fourth entry. So the vector
z4;1=2
66666641
0
0
0
0
03
7777775
will not be an element of K
T3
(notice that many other choices could be made here, so our basis will not
be unique). This completes the determination of Zp=Z4.
Next,s3=n3 n2= 5 4 = 1, so we again need just a single basis vector for Z3. We start by
evaluatingTwith each basis vector of Z4,
z3;1=T(z4;1) =Az4;1=2
6666664 3
3
3
3
3
23
7777775
Sinces3=s4, the subspace R3is trivial, and there is nothing left to do, z3;1is the lone basis vector of Z3.
Nows2=n2 n1= 4 2 = 2, so the construction of Z2will not be as simple as the construction of
Z3. We rst apply Tto the basis vector of Z2,
z2;1=T(z3;1) =Az3;1=2
66666641
0
3
1
0
13
7777775
The two basis vectors of K
T1
, together with z2;1, form a basis for Q2. Because dim
K
T2
dim (Q2) =
4 3 = 1 we need only nd a single basis vector for R2. This vector must be an element of K
T2
, but
not an element of Q2. Again, there is a variety of vectors that t this description, and we have no precise
algorithm for nding them. Since they are plentiful, they are not too hard to nd. We add up the four basis
vectors ofK
T2
, ensuring an element of K
T2
. Then we check to see if the vector is a linear combination
Version 2.30
702 Section NLT Nilpotent Linear Transformations
of three vectors: the two basis vectors of K
T1
andz2;1. Having passed the tests, we have chosen
z2;2=2
66666642
1
2
2
2
13
7777775
Thus,Z2=hfz2;1;z2;2gi.
Lastly,s1=n1 n0= 2 0 = 2. Since s2=s1, we again have a trivial R1and need only complete our
basis by evaluating the basis vectors of Z2withT,
z1;1=T(z2;1) =Az2;1=2
66666641
1
0
1
1
13
7777775
z1;2=T(z2;2) =Az2;2=2
6666664 2
2
5
2
1
03
7777775
Now we reorder these vectors as the desired basis,
B=fz1;1;z2;1;z3;1;z4;1;z1;2;z2;2g
We now apply Denition MR [615] to build a matrix representation of Trelative toB,
B(T(z1;1)) =B(Az1;1) =B(0) =2
66666640
0
0
0
0
03
7777775
B(T(z2;1)) =B(Az2;1) =B(z1;1) =2
66666641
0
0
0
0
03
7777775
B(T(z3;1)) =B(Az3;1) =B(z2;1) =2
66666640
1
0
0
0
03
7777775
Version 2.30
Subsection NLT.CFNLT Canonical Form for Nilpotent Linear Transformations 703
B(T(z4;1)) =B(Az4;1) =B(z3;1) =2
66666640
0
1
0
0
03
7777775
B(T(z1;2)) =B(Az1;2) =B(0) =2
66666640
0
0
0
0
03
7777775
B(T(z2;2)) =B(Az2;2) =B(z1;2) =2
66666640
0
0
0
1
03
7777775
Installing these vectors as the columns of the matrix representation we have
MT
B;B=2
66666640 1 0 0 0 0
0 0 1 0 0 0
0 0 0 1 0 0
0 0 0 0 0 0
0 0 0 0 0 1
0 0 0 0 0 03
7777775
which is a block diagonal matrix with Jordan blocks J4(0) andJ2(0). If we constructed the matrix S
having the vectors of Bas columns, then Theorem SCB [656] tells us that a similarity transformation
withSrelates the original matrix representation of Twith the matrix representation consisting of Jordan
blocks., i.e. S 1AS=MT
B;B.
Notice that constructing interesting examples of matrix representations requires domains with dimen-
sions bigger than just two or three. Going forward we will see several more big examples.
Version 2.30
704 Section NLT Nilpotent Linear Transformations
Version 2.30
Section IS Invariant Subspaces 705
Section IS
Invariant Subspaces
This section is in draft form
Nearly complete
We have seen in Section NLT [685] that nilpotent linear transformations are almost never diagonalizable
(Theorem DNLT [691]), yet have matrix representations that are very nearly diagonal (Theorem CFNLT
[694]). Our goal in this section, and the next (Section JCF [721]), is to obtain a matrix representation of any
linear transformation that is very nearly diagonal. A key step in reaching this goal is an understanding of
invariant subspaces, and a particular type of invariant subspace that contains vectors known as \generalized
eigenvectors."
Subsection IS
Invariant Subspaces
As is often the case, we start with a denition.
Denition IS
Invariant Subspace
Suppose that T:V!Vis a linear transformation and Wis a subspace of V. Suppose further that
T(w)2Wfor every w2W. ThenWis aninvariant subspace ofVrelative toT. 4
We do not have any special notation for an invariant subspace, so it is important to recognize that an
invariant subspace is always relative to both a superspace ( V) and a linear transformation ( T), which will
sometimes not be mentioned, yet will be clear from the context. Note also that the linear transformation
involved must have an equal domain and codomain | the denition would not make much sense if our
outputs were not of the same type as our inputs.
As usual, we begin with an example that demonstrates the existence of invariant subspaces. We will
return later to understand how this example was constructed, but for now, just understand how we check
the existence of the invariant subspaces.
Example TIS
Two invariant subspaces
Consider the linear transformation T:C4!C4dened byT(x) =AxwhereAis given by
A=2
664 8 6 15 9
8 14 10 18
1 1 3 0
3 8 2 113
775
Dene (with zero motivation),
w1=2
664 7
2
3
03
775w2=2
664 1
2
0
13
775
and setW=hfw1;w2gi. We verify that Wis an invariant subspace of C4with respect to T. By the
denition of W, any vector chosen from Wcan be written as a linear combination of w1andw2. Suppose
Version 2.30
706 Section IS Invariant Subspaces
thatw2W, and then check the details of the following verication,
T(w) =T(a1w1+a2w2) Denition SS [339]
=a1T(w1) +a2T(w2) Theorem LTLC [525]
=a12
664 1
2
0
13
775+a22
6645
2
3
23
775
=a1w2+a2(( 1)w1+ 2w2)
= ( a2)w1+ (a1+ 2a2)w2
2W Denition SS [339]
So, by Denition IS [703], Wis an invariant subspace of C4relative toT. In an entirely similar manner
we construct another invariant subspace of T.
With zero motivation, dene
x1=2
664 3
1
1
03
775x2=2
6640
1
0
13
775
and setX=hfx1;x2gi. We verify that Xis an invariant subspace of C4with respect to T. By the
denition of X, any vector chosen from Xcan be written as a linear combination of x1andx2. Suppose
thatx2X, and then check the details of the following verication,
T(x) =T(b1x1+b2x2) Denition SS [339]
=b1T(x1) +b2T(x2) Theorem LTLC [525]
=b12
6643
0
1
13
775+b22
6643
4
1
33
775
=b1(( 1)x1+x2) +b2(( 1)x1+ ( 3)x2)
= ( b1 b2)x1+ (b1 3b2)x2
2X Denition SS [339]
So, by Denition IS [703], Xis an invariant subspace of C4relative toT.
There is a bit of magic in each of these verications where the two outputs of Thappen to equal linear
combinations of the two inputs. But this is the essential nature of an invariant subspace. We'll have a
peek under the hood later, and it won't look so magical after all.
As a hint of things to come, verify that B=fw1;w2;x1;x2gis a basis of C4. Splitting this basis in
half, Theorem DSFB [413], tells us that C4=WX. To see why a decomposition of a vector space into
a direct sum of invariant subspaces might be interesting, construct the matrix representation of Trelative
toB,MT
B;B. Hmmmmmm.
Example TIS [703] is a bit mysterious at this stage. Do we know any other examples of invariant
subspaces? Yes, as it turns out, we have already seen quite a few. We'll give some examples now,
and in more general situations, describe broad classes of invariant subspaces with theorems. First up is
eigenspaces.
Version 2.30
Subsection IS.IS Invariant Subspaces 707
Theorem EIS
Eigenspaces are Invariant Subspaces
Suppose that T:V!Vis a linear transformation with eigenvalue and associated eigenspace ET().
LetWbe any subspace of ET(). ThenWis an invariant subspace of Vrelative toT.
Proof Choose w2W. Then
T(w) =w Denition EELT [647]
2W Property SC [317]
So by Denition IS [703], Wis an invariant subspace of Vrelative toT.
Theorem EIS [705] is general enough to determine that an entire eigenspace is an invariant subspace,
or that simply the span of a single eigenvector is an invariant subspace. It is not always the case that any
subspace of an invariant subspace is again an invariant subspace, but eigenspaces do have this property.
Here is an example of the theorem, which also allows us to very quickly build several several invariant (4x4,
2 evs, 1 2x2 jordan, 1 2x2 diag)
Example EIS
Eigenspaces as invariant subspaces
Dene the linear transformation S:M22!M22by
Sa b
c d
= 2a+ 19b 33c+ 21d 3a+ 16b 24c+ 15d
2a+ 9b 13c+ 9d a+ 4b 6c+ 5d
Build a matrix representation of Srelative to the standard basis (Denition MR [615], Example BM [372])
and compute eigenvalues and eigenspaces of Swith the computational techniques of Chapter E [453] in
concert with Theorem EER [659]. Then
ES(1) =4 3
2 1
ES(2) =6 3
1 0
; 9 3
0 1
So by Theorem EIS [705], both ES(1) andES(2) are invariant subspaces of M22relative toS. However,
Theorem EIS [705] provides even more invariant subspaces. Since ES(1) has dimension 1, it has no
interesting subspaces, however ES(2) has dimension 2 and has a plethora of subspaces. For example, set
u= 26 3
1 0
+ 3 9 3
0 1
= 6 3
2 3
and deneU=hfugi. Then since Uis a subspace ofES(2), Theorem EIS [705] says that Uis an invariant
subspace of M22(or we could check this claim directly based simply on the fact that uis an eigenvector of
S).
For every linear transformation there are some obvious, trivial invariant subspaces. Suppose that
T:V!Vis a linear transformation. Then simply because Tis a function (Denition LT [515]), the
subspaceVis an invariant subspace of T. In only a minor twist on this theme, the range of T,R(T), is an
invariant subspace of Tby Denition RLT [563]. Finally, Theorem LTTZZ [519] provides the justication
for claiming that f0gis an invariant subspace of T.
That the trivial subspace is always an invariant subspace is a special case of the next theorem. As an
easy exercise before reading the next theorem, prove that the kernel of a linear transformation (Denition
KLT [545]),K(T), is an invariant subspace. We'll wait.
Theorem KPIS
Kernels of Powers are Invariant Subspaces
Suppose that T:V!Vis a linear transformation. Then K
Tk
is an invariant subspace of V.
Proof Suppose that z2K
Tk
. Then
Tk(T(z)) =Tk+1(z) Denition LTC [532]
Version 2.30
708 Section IS Invariant Subspaces
=T
Tk(z)
Denition LTC [532]
=T(0) Denition KLT [545]
=0 Theorem LTTZZ [519]
So by Denition KLT [545], we see that T(z)2K
Tk
. ThusK
Tk
is an invariant subspace of Vrelative
toT(Denition IS [703]).
Two interesting special cases of Theorem KPIS [705] occur when choose k= 0 andk= 1. Rather than
give an example of this theorem, we will refer you back to Example KPNLT [693] where we work with null
spaces of the rst four powers of a nilpotent matrix. By Theorem KPIS [705] each of these null spaces is
an invariant subspace of the associated linear transformation.
Here's one more example of invariant subspaces we have encountered previously.
Example ISJB
Invariant subspaces and Jordan blocks
Refer back to Example CFNLT [698]. We decomposed the vector space C6into a direct sum of the
subspacesZ1; Z2; Z3; Z4. The union of the basis vectors for these subspaces is a basis of C6, which we
reordered prior to building a matrix representation of the linear transformation T. A principal reason for
this reordering was to create invariant subspaces (though it was not obvious then).
Dene
X1=hfz1;1;z2;1;z3;1;z4;1gi=*8
>>>>>><
>>>>>>:2
66666641
1
0
1
1
13
7777775;2
66666641
0
3
1
0
13
7777775;2
6666664 3
3
3
3
3
23
7777775;2
66666641
0
0
0
0
03
77777759
>>>>>>=
>>>>>>;+
X2=hfz1;2;z2;2gi=*8
>>>>>><
>>>>>>:2
6666664 2
2
5
2
1
03
7777775;2
66666642
1
2
2
2
13
77777759
>>>>>>=
>>>>>>;+
Recall from the proof of Theorem CFNLT [694] or the computations in Example CFNLT [698] that rst
elements of X1andX2are in the kernel of T,K(T), and each element of X1andX2is the output of T
when evaluated with the subsequent element of the set. This was by design, and it is this feature of these
basis vectors that leads to the nearly diagonal matrix representation with Jordan blocks. However, we also
recognize now that this property of these basis vectors allow us to conclude easily that X1andX2are
invariant subspaces of C6relative toT.
Furthermore, C6=X1X2(Theorem DSFB [413]). So the domain of Tis the direct sum of invariant
subspaces and the resulting matrix representation has a block diagonal form. Hmmmmm.
Subsection GEE
Generalized Eigenvectors and Eigenspaces
We now dene a new type of invariant subspace and explore its key properties. This generalization of
eigenvalues and eigenspaces will allow us to move from diagonal matrix representations of diagonalizable
matrices to nearly diagonal matrix representations of arbitrary matrices. Here are the denitions.
Version 2.30
Subsection IS.GEE Generalized Eigenvectors and Eigenspaces 709
Denition GEV
Generalized Eigenvector
Suppose that T:V!Vis a linear transformation. Suppose further that for x6=0, (T IV)k(x) =0
for somek>0. Then xis ageneralized eigenvector ofTwith eigenvalue . 4
Denition GES
Generalized Eigenspace
Suppose that T:V!Vis a linear transformation. Dene the generalized eigenspace ofTforas
GT() =n
xj(T IV)k(x) =0for somek0o
(This denition contains Notation GES.) 4
So the generalized eigenspace is composed of generalized eigenvectors, plus the zero vector. As the
name implies, the generalized eigenspace is a subspace of V. But more topically, it is an invariant subspace
ofVrelative toT.
Theorem GESIS
Generalized Eigenspace is an Invariant Subspace
Suppose that T:V!Vis a linear transformation. Then the generalized eigenspace GT() is an invariant
subspace of Vrelative toT.
Proof First we establish that GT() is a subspace of V. First (T IV)1(0) =0by Theorem LTTZZ
[519], so 02GT().
Suppose that x;y2GT(). Then there are integers k; `such that (T IV)k(x) =0and (T IV)`(y) =
0. Setm=k+`,
(T IV)m(x+y) = (T IV)m(x) + (T IV)m(y) Denition LT [515]
= (T IV)k+`(x) + (T IV)k+`(y)
= (T IV)`
(T IV)k(x)
+
(T IV)k
(T IV)`(y)
Denition LTC [532]
= (T IV)`(0) + (T IV)k(0) Denition GES [707]
=0+0 Theorem LTTZZ [519]
=0 Property Z [318]
Sox+y2GT().
Suppose that x2GT() and2C. Then there is an integer ksuch that (T IV)k(x) =0.
(T IV)k(x) =(T IV)k(x) Denition LT [515]
=0 Denition GES [707]
=0 Theorem ZVSM [325]
Sox2GT(). By Theorem TSS [334], GT() is a subspace of V.
Now we show that GT() is invariant relative to T. Suppose that x2GT(). Then by Denition GES
[707] there is an integer ksuch that (T IV)k(x) =0. The following argument is due to Zoltan Toth.
(T IV)k(T(x)) = (T IV)k(T(x)) 0 Property Z [318]
= (T IV)k(T(x)) 0 Theorem ZVSM [325]
= (T IV)k(T(x)) (T IV)k(x) Denition GES [707]
Version 2.30
710 Section IS Invariant Subspaces
= (T IV)k(T(x)) (T IV)k(x) Denition LT [515]
= (T IV)k(T(x) x) Denition LT [515]
= (T IV)k((T IV) (x)) Denition LTA [530]
= (T IV)k+1(x) Denition LTC [532]
= (T IV)
(T IV)k(x)
Denition LTC [532]
= (T IV) (0) Denition GES [707]
=0 Theorem LTTZZ [519]
This qualies T(x) for membership in GT(), so by Denition GES [707], GT() is invariant relative to
T.
Before we compute some generalized eigenspaces, we state and prove one theorem that will make it
much easier to create a generalized eigenspace, since it will allow us to use tools we already know well, and
will remove some the ambiguity of the clause \for some k" in the denition.
Theorem GEK
Generalized Eigenspace as a Kernel
Suppose that T:V!Vis a linear transformation, dim ( V) =n, andis an eigenvalue of T. Then
GT() =K((T IV)n).
Proof The conclusion of this theorem is a set equality, so we will apply Denition SE [762] by establishing
two set inclusions. First, suppose that x2GT(). Then there is an integer ksuch that (T IV)k(x) =0.
This is equivalent to the statement that x2K
(T IV)k
. No matter what the value of kis, Theorem
KPLT [691] gives
x2K
(T IV)k
K((T IV)n)
So,GT()K((T IV)n). For the opposite inclusion, suppose y2K((T IV)n). Then (T IV)n(y) =
0, soy2GT() and thusK((T IV)n)GT(). By Denition SE [762] we have the desired equality of
sets.
Theorem GEK [708] allows us to compute generalized eigenspaces as a single kernel (or null space of a
matrix representation) with tools like Theorem KNSI [625] and Theorem BNS [160]. Also, we do not need
to consider all possible powers kand can simply consider the case where k=n. It is worth noting that
the \regular" eigenspace is a subspace of the generalized eigenspace since
ET() =K
(T IV)1
K((T IV)n) =GT()
where the subset inclusion is a consequence of Theorem KPLT [691]. Also, there is no such thing as a
\generalized eigenvalue." If is not an eigenvalue of T, then the kernel of T IVis trivial and therefore
subsequent powers of T IValso have trivial kernels (Theorem KPLT [691]). So the generalized eigenspace
of a scalar that is not already an eigenvalue would be trivial. Alright, we know enough now to compute
some generalized eigenspaces. We will record some information about algebraic and geometric multiplicities
of eigenvalues (Denition AME [463], Denition GME [463]) as we go, since these observations will be of
interest in light of some future theorems.
Example GE4
Generalized eigenspaces, dimension 4 domain
In Example TIS [703] we presented two invariant subspaces of C4. There was some mystery about just
how these were constructed, but we can now reveal that they are generalized eigenspaces. Example TIS
Version 2.30
Subsection IS.GEE Generalized Eigenvectors and Eigenspaces 711
[703] featured T:C4!C4dened byT(x) =AxwithAgiven by
A=2
664 8 6 15 9
8 14 10 18
1 1 3 0
3 8 2 113
775
A matrix representation of Trelative to the standard basis (Denition SUV [197]) will equal A. So we
can analyze Awith the techniques of Chapter E [453]. Doing so, we nd two eigenvalues, = 1; 2, with
multiplicities,
T(1) = 2
T(1) = 1
T( 2) = 2
T( 2) = 1
To apply Theorem GEK [708] we subtract each eigenvalue from the diagonal entries of A, raise the result
to the power dim
C4
= 4, and compute a basis for the null space.
= 2 (A ( 2)I4)4=2
664648 1215 729 1215
324 486 486 486
405 729 486 729
297 486 405 4863
775RREF !2
6641 0 3 0
0 1 1 1
0 0 0 0
0 0 0 03
775
GT( 2) =*8
>><
>>:2
664 3
1
1
03
775;2
6640
1
0
13
7759
>>=
>>;+
= 1 ( A (1)I4)4=2
66481 405 81 729
108 189 378 486
27 135 27 243
135 54 351 2433
775RREF !2
6641 07
31
0 12
32
0 0 0 0
0 0 0 03
775
GT(1) =*8
>><
>>:2
664 7
2
3
03
775;2
664 1
2
0
13
7759
>>=
>>;+
In Example TIS [703] we concluded that these two invariant subspaces formed a direct sum of C4, only at
that time, they were called XandW. Now we can write
C4=GT(1)GT( 2)
This is no accident. Notice that the dimension of each of these invariant subspaces is equal to the algebraic
multiplicity of the associated eigenvalue. Not an accident either. (See the upcoming Theorem GESD [721].)
Example GE6
Generalized eigenspaces, dimension 6 domain
Dene the linear transformation S:C6!C6byS(x) =Bxwhere
2
66666642 4 25 54 90 37
2 3 4 16 26 8
2 3 4 15 24 7
10 18 6 36 51 2
8 14 0 21 28 4
5 7 6 7 8 73
7777775
Version 2.30
712 Section IS Invariant Subspaces
ThenBwill be the matrix representation of Srelative to the standard basis (Denition SUV [197]) and
we can use the techniques of Chapter E [453] applied to Bin order to nd the eigenvalues of S.
S(3) = 2
S(3) = 1
S( 1) = 4
S( 1) = 2
To nd the generalized eigenspaces of Swe need to subtract an eigenvalue from the diagonal elements of
B, raise the result to the power dim
C6
= 6 and compute the null space. Here are the results for the two
eigenvalues of S,
= 3 ( B 3I6)6=2
666666464000 152576 59904 26112 95744 133632
15872 39936 11776 8704 29184 36352
12032 30208 9984 6400 20736 26368
1536 11264 23040 17920 17920 1536
9728 27648 6656 9728 1536 17920
7936 17920 5888 1792 4352 140803
7777775
RREF !2
66666641 0 0 0 4 5
0 1 0 0 1 1
0 0 1 0 1 1
0 0 0 1 2 1
0 0 0 0 0 0
0 0 0 0 0 03
7777775
GS(3) =*8
>>>>>><
>>>>>>:2
66666644
1
1
2
1
03
7777775;2
6666664 5
1
1
1
0
13
77777759
>>>>>>=
>>>>>>;+
= 1 (B ( 1)I6)6=2
66666646144 16384 18432 36864 57344 18432
4096 8192 4096 16384 24576 4096
4096 8192 4096 16384 24576 4096
18432 32768 6144 61440 90112 6144
14336 24576 2048 45056 65536 2048
10240 16384 2048 28672 40960 20483
7777775
RREF !2
66666641 0 5 2 4 5
0 1 3 3 5 3
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 03
7777775
GS( 1) =*8
>>>>>><
>>>>>>:2
66666645
3
1
0
0
03
7777775;2
6666664 2
3
0
1
0
03
7777775;2
66666644
5
0
0
1
03
7777775;2
6666664 5
3
0
0
0
13
77777759
>>>>>>=
>>>>>>;+
If we take the union of the two bases for these two invariant subspaces we obtain the set
C=fv1;v2;v3;v4;v5;v6g
Version 2.30
Subsection IS.RLT Restrictions of Linear Transformations 713
=8
>>>>>><
>>>>>>:2
66666644
1
1
2
1
03
7777775;2
6666664 5
1
1
1
0
13
7777775;2
66666645
3
1
0
0
03
7777775;2
6666664 2
3
0
1
0
03
7777775;2
66666644
5
0
0
1
03
7777775;2
6666664 5
3
0
0
0
13
77777759
>>>>>>=
>>>>>>;
You can check that this set is linearly independent (right now we have no guarantee this will happen).
Once this is veried, we have a linearly independent set of size 6 inside a vector space of dimension 6, so by
Theorem G [407], the set Cis a basis for C6. This is enough to apply Theorem DSFB [413] and conclude
that
C6=GS(3)GS( 1)
This is no accident. Notice that the dimension of each of these invariant subspaces is equal to the algebraic
multiplicity of the associated eigenvalue. Not an accident either. (See the upcoming Theorem GESD [721].)
Subsection RLT
Restrictions of Linear Transformations
Generalized eigenspaces will prove to be an important type of invariant subspace. A second reason for our
interest in invariant subspaces is they provide us with another method for creating new linear transforma-
tions from old ones.
Denition LTR
Linear Transformation Restriction
Suppose that T:V!Vis a linear transformation, and Uis an invariant subspace of Vrelative to T.
Dene the restriction ofTtoUby
TjU:U!U T jU(u) =T(u)
(This denition contains Notation LTR.) 4
It might appear that this denition has not accomplished anything, as TjUwould appear to take on
exactly the same values as T. And this is true. However, TjUdiers from Tin the choice of domain and
codomain. We tend to give little attention to the domain and codomain of functions, while their dening
rules get the spotlight. But the restriction of a linear transformation is all about the choice of domain and
codomain. We are restricting the rule of the function to a smaller subspace. Notice the importance of only
using this construction with an invariant subspace, since otherwise we cannot be assured that the outputs
of the function are even contained in the codomain. Maybe this observation should be the key step in the
proof of a theorem saying that TjUis also a linear transformation, but we won't bother.
Example LTRGE
Linear transformation restriction on generalized eigenspace
In order to gain some experience with restrictions of linear transformations, we construct one and then also
construct a matrix representation for the restriction. Furthermore, we will use a generalized eigenspace as
the invariant subspace for the construction of the restriction.
Version 2.30
714 Section IS Invariant Subspaces
Consider the linear transformation T:C5!C5dened byT(x) =Ax, where
A=2
66664 22 24 24 24 46
3 2 6 0 11
12 16 6 14 17
6 8 4 10 8
11 14 8 13 183
77775
One of the eigenvalues of Ais= 2, with geometric multiplicity
T(2) = 1, and algebraic multiplicity
T(2) = 3. We get the generalized eigenspace in the usual manner,
W=GT(2) =K
(T 2IC5)5
=*8
>>>><
>>>>:2
66664 2
1
1
0
03
77775;2
666640
1
0
1
03
77775;2
66664 4
2
0
0
13
777759
>>>>=
>>>>;+
=hfw1;w2;w3gi
By Theorem GESIS [707], we know Wis invariant relative to T, so we can employ Denition LTR [711]
to form the restriction, TjW:W!W.
To better understand exactly what a restriction is (and isn't), we'll form a matrix representation of TjW.
This will also be a skill we will use in subsequent examples. For a basis of Wwe will useC=fw1;w2;w3g.
Notice that dim ( W) = 3, so our matrix representation will be a square matrix of size 3. Applying Denition
MR [615], we compute
C(T(w1)) =C(Aw1) =C0
BBBB@2
66664 4
2
2
0
03
777751
CCCCA=C0
BBBB@22
66664 2
1
1
0
03
77775+ 02
666640
1
0
1
03
77775+ 02
66664 4
2
0
0
13
777751
CCCCA=2
42
0
03
5
C(T(w2)) =C(Aw2) =C0
BBBB@2
666640
2
2
2
13
777751
CCCCA=C0
BBBB@22
66664 2
1
1
0
03
77775+ 22
666640
1
0
1
03
77775+ ( 1)2
66664 4
2
0
0
13
777751
CCCCA=2
42
2
13
5
C(T(w3)) =C(Aw3) =C0
BBBB@2
66664 6
3
1
0
23
777751
CCCCA=C0
BBBB@( 1)2
66664 2
1
1
0
03
77775+ 02
666640
1
0
1
03
77775+ 22
66664 4
2
0
0
13
777751
CCCCA=2
4 1
0
23
5
So the matrix representation of TjWrelative toCis
MTjW
C;C=2
42 2 1
0 2 0
0 1 23
5
The question arises: how do we use a 3 3 matrix to compute with vectors from C5? To answer this
question, consider the randomly chosen vector
w=2
66664 4
4
4
2
13
77775
Version 2.30
Subsection IS.RLT Restrictions of Linear Transformations 715
First check that w2GT(2). There are two ways to do this, rst verify that
(T 2IC5)5(w) = (A 2I5)5w=0
meeting Denition GES [707] (with k= 5). Or, express was a linear combination of the basis CforW,
to wit, w= 4w1 2w2 w3. Now compute TjW(w) directly using Denition LTR [711],
TjW(w) =T(w) =Aw=2
66664 10
9
5
4
03
77775
It was necessary to verify that w2GT(2), and if we trust our work so far, then this output will also be
an element of W, but it would be wise to check this anyway (using either of the methods we used for w).
We'll wait.
Now we will repeat this sample computation, but instead using the matrix representation of TjW
relative toC.
TjW(w) = 1
C
MTjW
C;CC(w)
Theorem FTMR [617]
= 1
C
MTjW
C;CC(4w1 2w2 w3)
= 1
C0
@2
42 2 1
0 2 0
0 1 23
52
44
2
13
51
A Denition VR [603]
= 1
C0
@2
45
4
03
51
A Denition MVP [223]
= 5w1 4w2+ 0w3 Denition VR [603]
= 52
66664 2
1
1
0
03
77775+ ( 4)2
666640
1
0
1
03
77775+ 02
66664 4
2
0
0
13
77775
=2
66664 10
9
5
4
03
77775
which matches the previous computation. Notice how the \action" of TjWis accomplished by a 3 3 matrix
multiplying a column vector of size 3. If you would like more practice with these sorts of computations,
mimic the above using the other eigenvalue of T, which is = 2. The generalized eigenspace has
dimension 2, so the matrix representation of the restriction to the generalized eigenspace will be a 2 2
matrix.
Suppose that T:V!Vis a linear transformation and we can nd a decomposition of Vas a direct
sum, sayV=U1U2U3Umwhere each Uiis an invariant subspace of Vrelative toT. Then,
for any v2Vthere is a unique decomposition v=u1+u2+u3++umwithui2Ui, 1imand
furthermore
T(v) =T(u1+u2+u3++um) Denition DS [413]
Version 2.30
716 Section IS Invariant Subspaces
=T(u1) +T(u2) +T(u3) ++T(um) Theorem LTLC [525]
=TjU1(u1) +TjU2(u2) +TjU3(u3) ++TjUm(um)
So in a very real sense, we obtain a decomposition of the linear transformation Tinto the restrictions TjUi,
1im. If we wanted to be more careful, we could extend each restriction to a linear transformation
dened on Vby setting the output of TjUito be the zero vector for inputs outside of Ui. ThenTwould
be exactly equal to the sum (Denition LTA [530]) of these extended restrictions. However, the irony of
extending our restrictions is more than we could handle right now.
Our real interest is in the matrix representation of a linear transformation when the domain decomposes
as a direct sum of invariant subspaces. Consider forming a basis BofVas the union of bases Bifrom the
individualUi, i.e.B=[m
i=1Bi. Now form the matrix representation of Trelative toB. The result will be
block diagonal, where each block is the matrix representation of a restriction TjUirelative to a basis Bi,
MTjUi
Bi;Bi. Though we did not have the denitions to describe it then, this is exactly what was going on in
the latter portion of the proof of Theorem CFNLT [694]. Two examples should help to clarify these ideas.
Example ISMR4
Invariant subspaces, matrix representation, dimension 4 domain
Example TIS [703] and Example GE4 [708] describe a basis of C4which is derived from bases for two
invariant subspaces (both generalized eigenspaces). In this example we will construct a matrix representa-
tion of the linear transformation Trelative to this basis. Recycling the notation from Example TIS [703],
we work with the basis,
B=fw1;w2;x1;x2g=8
>><
>>:2
664 7
2
3
03
775;2
664 1
2
0
13
775;2
664 3
1
1
03
775;2
6640
1
0
13
7759
>>=
>>;
Now we compute the matrix representation of Trelative toB, borrowing some computations from Example
TIS [703],
B(T(w1)) =B0
BB@2
664 1
2
0
13
7751
CCA=B((0)w1+ (1)w2) =2
6640
1
0
03
775
B(T(w2)) =B0
BB@2
6645
2
3
23
7751
CCA=B(( 1)w1+ (2)w2) =2
664 1
2
0
03
775
B(T(x1)) =B0
BB@2
6643
0
1
13
7751
CCA=B(( 1)x1+ (1)x2) =2
6640
0
1
13
775
B(T(x2)) =B0
BB@2
6643
4
1
33
7751
CCA=B(( 1)x1+ ( 3)x2) =2
6640
0
1
33
775
Applying Denition MR [615], we have
MT
B;B=2
6640 1 0 0
1 2 0 0
0 0 1 1
0 0 1 33
775
Version 2.30
Subsection IS.RLT Restrictions of Linear Transformations 717
The interesting feature of this representation is the two 2 2 blocks on the diagonal that arise from the
decomposition of C4into a direct sum (of generalized eigenspaces). Or maybe the interesting feature of
this matrix is the two 2 2 submatrices in the \other" corners that are all zero. You decide.
Example ISMR6
Invariant subspaces, matrix representation, dimension 6 domain
In Example GE6 [709] we computed the generalized eigenspaces of the linear transformation S:C6!C6
byS(x) =Bxwhere
2
66666642 4 25 54 90 37
2 3 4 16 26 8
2 3 4 15 24 7
10 18 6 36 51 2
8 14 0 21 28 4
5 7 6 7 8 73
7777775
From this we found the basis
C=fv1;v2;v3;v4;v5;v6g
=8
>>>>>><
>>>>>>:2
66666644
1
1
2
1
03
7777775;2
6666664 5
1
1
1
0
13
7777775;2
66666645
3
1
0
0
03
7777775;2
6666664 2
3
0
1
0
03
7777775;2
66666644
5
0
0
1
03
7777775;2
6666664 5
3
0
0
0
13
77777759
>>>>>>=
>>>>>>;
ofC6wherefv1;v2gis a basis ofGS(3) andfv3;v4;v5;v6gis a basis ofGS( 1). We can employ Cin
the construction of a matrix representation of S(Denition MR [615]). Here are the computations,
C(S(v1)) =C0
BBBBBB@2
666666411
3
3
7
4
13
77777751
CCCCCCA=C(4v1+ 1v2) =2
66666644
1
0
0
0
03
7777775
C(S(v2)) =C0
BBBBBB@2
6666664 14
3
3
4
1
23
77777751
CCCCCCA=C(( 1)v1+ 2v2) =2
6666664 1
2
0
0
0
03
7777775
C(S(v3)) =C0
BBBBBB@2
666666423
5
5
2
2
23
77777751
CCCCCCA=C(5v3+ 2v4+ ( 2)v5+ ( 2)v6) =2
66666640
0
5
2
2
23
7777775
C(S(v4)) =C0
BBBBBB@2
6666664 46
11
10
2
5
43
77777751
CCCCCCA=C(( 10)v3+ ( 2)v4+ 5v5+ 4v6) =2
66666640
0
10
2
5
43
7777775
Version 2.30
718 Section IS Invariant Subspaces
C(S(v5)) =C0
BBBBBB@2
666666478
19
17
1
10
73
77777751
CCCCCCA=C(17v3+ 1v4+ ( 10)v5+ ( 7)v6) =2
66666640
0
17
1
10
73
7777775
C(S(v6)) =C0
BBBBBB@2
6666664 35
9
8
2
6
33
77777751
CCCCCCA=C(( 8)v3+ 2v4+ 6v5+ 3v6) =2
66666640
0
8
2
6
33
7777775
These column vectors are the columns of the matrix representation, so we obtain
MS
C;C=2
66666644 1 0 0 0 0
1 2 0 0 0 0
0 0 5 10 17 8
0 0 2 2 1 2
0 0 2 5 10 6
0 0 2 4 7 33
7777775
As before, the key feature of this representation is the 2 2 and 44 blocks on the diagonal. We will
discover in the nal theorem of this section (Theorem RGEN [716]) that we already understand these
blocks fairly well. For now, we recognize them as arising from generalized eigenspaces and suspect that
their sizes are equal to the algebraic multiplicities of the eigenvalues.
The paragraph prior to these last two examples is worth repeating. A basis derived from a direct sum
decomposition into invariant subspaces will provide a matrix representation of a linear transformation with
a block diagonal form.
Diagonalizing a linear transformation is the most extreme example of decomposing a vector space
into invariant subspaces. When a linear transformation is diagonalizable, then there is a basis composed
of eigenvectors (Theorem DC [497]). Each of these basis vectors can be used individually as the lone
element of a spanning set for an invariant subspace (Theorem EIS [705]). So the domain decomposes into
a direct sum of one-dimensional invariant subspaces (Theorem DSFB [413]). The corresponding matrix
representation is then block diagonal with all the blocks of size 1, i.e. the matrix is diagonal. Section NLT
[685], Section IS [703] and Section JCF [721] are all devoted to generalizing this extreme situation when
there are not enough eigenvectors available to make such a complete decomposition and arrive at such an
elegant matrix representation.
One last theorem will roll up much of this section and Section NLT [685] into one nice, neat package.
Theorem RGEN
Restriction to Generalized Eigenspace is Nilpotent
SupposeT:V!Vis a linear transformation with eigenvalue . Then the linear transformation TjGT()
IGT()is nilpotent.
Proof Notice rst that every subspace of Vis invariant with respect to IV, soIGT()=IVjGT(). Let
n= dim (V) and choose v2GT(). Then
TjGT() IGT()n(v) = (T IV)n(v) Denition LTR [711]
=0 Theorem GEK [708]
So by Denition NLT [685], TjGT() IGT()is nilpotent.
The proof of Theorem RGEN [716] indicates that the index of the nilpotent linear transformation is
less than or equal to the dimension of V. In practice, it will be less than or equal to the dimension of the
Version 2.30
Subsection IS.RLT Restrictions of Linear Transformations 719
domain of the linear transformation, GT(). In any event, the exact value of this index will be of some
interest, so we dene it now. Notice that this is a property of the eigenvalue , similar to the algebraic
and geometric multiplicities (Denition AME [463], Denition GME [463]).
Denition IE
Index of an Eigenvalue
SupposeT:V!Vis a linear transformation with eigenvalue . Then the index of,T(), is the index
of the nilpotent linear transformation TjGT() IGT().
(This denition contains Notation IE.) 4
Example GENR6
Generalized eigenspaces and nilpotent restrictions, dimension 6 domain
In Example GE6 [709] we computed the generalized eigenspaces of the linear transformation S:C6!C6
dened byS(x) =Bxwhere
2
66666642 4 25 54 90 37
2 3 4 16 26 8
2 3 4 15 24 7
10 18 6 36 51 2
8 14 0 21 28 4
5 7 6 7 8 73
7777775
The generalized eigenspace, GS(3), has dimension 2, while GS( 1), has dimension 4. We'll investigate each
thoroughly in turn, with the intent being to illustrate Theorem RGEN [716]. Much of our computations
will be repeats of those done in Example ISMR6 [715].
ForU=GS(3) we compute a matrix representation of SjUusing the basis found in Example GE6 [709],
B=fu1;u2g=8
>>>>>><
>>>>>>:2
66666644
1
1
2
1
03
7777775;2
6666664 5
1
1
1
0
13
77777759
>>>>>>=
>>>>>>;
SinceBhas size 2, we obtain a 2 2 matrix representation (Denition MR [615]) from
B(SjU(u1)) =B0
BBBBBB@2
666666411
3
3
7
4
13
77777751
CCCCCCA=B(4u1+u2) =4
1
B(SjU(u2)) =B0
BBBBBB@2
6666664 14
3
3
4
1
23
77777751
CCCCCCA=B(( 1)u1+ 2u2) = 1
2
Thus
M=MSjU
U;U=4 1
1 2
Version 2.30
720 Section IS Invariant Subspaces
Now we can illustrate Theorem RGEN [716] with powers of the matrix representation (rather than the
restriction itself),
M 3I2=1 1
1 1
(M 3I2)2=0 0
0 0
SoM 3I2is a nilpotent matrix of index 2 (meaning that SjU 3IUis a nilpotent linear transformation
of index 2) and according to Denition IE [717] we say S(3) = 2.
ForW=GS( 1) we compute a matrix representation of SjWusing the basis found in Example GE6
[709],
C=fw1;w2;w3;w4g=8
>>>>>><
>>>>>>:2
66666645
3
1
0
0
03
7777775;2
6666664 2
3
0
1
0
03
7777775;2
66666644
5
0
0
1
03
7777775;2
6666664 5
3
0
0
0
13
77777759
>>>>>>=
>>>>>>;
SinceChas size 4, we obtain a 4 4 matrix representation (Denition MR [615]) from
C(SjW(w1)) =C0
BBBBBB@2
666666423
5
5
2
2
23
77777751
CCCCCCA=C(5w1+ 2w2+ ( 2)w3+ ( 2)w4) =2
6645
2
2
23
775
C(SjW(w2)) =C0
BBBBBB@2
6666664 46
11
10
2
5
43
77777751
CCCCCCA=C(( 10)w1+ ( 2)w2+ 5w3+ 4w4) =2
664 10
2
5
43
775
C(SjW(w3)) =C0
BBBBBB@2
666666478
19
17
1
10
73
77777751
CCCCCCA=C(17w1+w2+ ( 10)w3+ ( 7)w4) =2
66417
1
10
73
775
C(SjW(w4)) =C0
BBBBBB@2
6666664 35
9
8
2
6
33
77777751
CCCCCCA=C(( 8)w1+ 2w2+ 6w3+ 3w4) =2
664 8
2
6
33
775
Thus
N=MSjW
W;W=2
6645 10 17 8
2 2 1 2
2 5 10 6
2 4 7 33
775
Version 2.30
Subsection IS.RLT Restrictions of Linear Transformations 721
Now we can illustrate Theorem RGEN [716] with powers of the matrix representation (rather than the
restriction itself),
N ( 1)I4=2
6646 10 17 8
2 1 1 2
2 5 9 6
2 4 7 43
775
(N ( 1)I4)2=2
664 2 3 5 2
4 6 10 4
4 6 10 4
2 3 5 23
775
(N ( 1)I4)3=2
6640 0 0 0
0 0 0 0
0 0 0 0
0 0 0 03
775
SoN ( 1)I4is a nilpotent matrix of index 3 (meaning that SjW ( 1)IWis a nilpotent linear transfor-
mation of index 3) and according to Denition IE [717] we say S( 1) = 3.
Notice that if we were to take the union of the two bases of the generalized eigenspaces, we would have
a basis for C6. Then a matrix representation of Srelative to this basis would be the same block diagonal
matrix we found in Example ISMR6 [715], only we now understand each of these blocks as being very close
to being a nilpotent matrix.
Invariant subspaces, and restrictions of linear transformations, are topics you will see again and again
if you continue with further study of linear algebra. Our reasons for discussing them now is to arrive at a
nice matrix representation of the restriction of a linear transformation to one of its generalized eigenspaces.
Here's the theorem.
Theorem MRRGE
Matrix Representation of a Restriction to a Generalized Eigenspace
Suppose that T:V!Vis a linear transformation with eigenvalue . Then there is a basis of the the
generalized eigenspace GT() such that the restriction TjGT():GT()!GT() has a matrix representation
that is block diagonal where each block is a Jordan block of the form Jn().
Proof Theorem RGEN [716] tells us that TjGT() IGT()is a nilpotent linear transformation. Theorem
CFNLT [694] tells us that a nilpotent linear transformation has a basis for its domain that yields a matrix
representation that is block diagonal where the blocks are Jordan blocks of the form Jn(0). LetBbe a
basis ofGT() that yields such a matrix representation for TjGT() IGT().
By Denition LTA [530], we can write
TjGT()=
TjGT() IGT()
+IGT()
The matrix representation of IGT()relative to the basis Bis then simply the diagonal matrix Im, where
m= dim (GT()). By Theorem MRSLT [621] we have the rather unwieldy expression,
MTjGT()
B;B=M(TjGT() IGT())+IGT()
B;B
=MTjGT() IGT()
B;B+MIGT()
B;B
The rst of these matrix representations has Jordan blocks with zero in every diagonal entry, while the
second matrix representation has in every diagonal entry. The result of adding the two representations
is to convert the Jordan blocks from the form Jn(0) to the form Jn().
Of course, Theorem CFNLT [694] provides some extra information on the sizes of the Jordan blocks in
a representation and we could carry over this information to Theorem MRRGE [719], but will save that
for a subsequent application of this result.
Version 2.30
722 Section IS Invariant Subspaces
Subsection EXC
Exercises
T10 Suppose that T:V!Vis linear transformation, and p(x) is a polynomial. Then dene the new
linear transformation p(T):V!Vby interpreting the coecients of the terms of the polynomial as
scalar mutliples of linear transformations (Denition LTSM [531]), addition of terms as the sum of linear
transformations (Denition LTA [530]), and powers as repeated composition of linear transformations
(Denition LTC [532]). Prove that Tp(T) =p(T)T.
Use this observation to give a shorter argument for the proof of the invariance of the generalized
eigenspace in Theorem GESIS [707].
Contributed by Robert Beezer
Version 2.30
Section JCF Jordan Canonical Form 723
Section JCF
Jordan Canonical Form
This section is in draft form
Needs examples near beginning
We have seen in Section IS [703] that generalized eigenspaces are invariant subspaces that in every
instance have led to a direct sum decomposition of the domain of the associated linear transformation.
This allows us to create a block diagonal matrix representation (Example ISMR4 [714], Example ISMR6
[715]). We also know from Theorem RGEN [716] that the restriction of a linear transformation to a
generalized eigenspace is almost a nilpotent linear transformation. Of course, we understand nilpotent
linear transformations very well from Section NLT [685] and we have carefully determined a nice matrix
representation for them.
So here is the game plan for the nal push. Prove that the domain of a linear transformation always
decomposes into a direct sum of generalized eigenspaces. We have unravelled Theorem RGEN [716] at
Theorem MRRGE [719] so that we can formulate the matrix representations of the restrictions on the
generalized eigenspaces using our storehouse of results about nilpotent linear transformations. Arrive at a
matrix representation of anylinear transformation that is block diagonal with each block being a Jordan
block.
Subsection GESD
Generalized Eigenspace Decomposition
In Theorem UTMR [676] we were able to show that any linear transformation from VtoVhas an upper
triangular matrix representation (Denition UTM [675]). We will now show that we can improve on the
basis yielding this representation by massaging the basis so that the matrix representation is also block
diagonal. The subspaces associated with each block will be generalized eigenspaces, so the most general
result will be a decomposition of the domain of a linear transformation into a direct sum of generalized
eigenspaces.
Theorem GESD
Generalized Eigenspace Decomposition
Suppose that T:V!Vis a linear transformation with distinct eigenvalues 1; 2; 3; :::; m. Then
V=GT(1)GT(2)GT(3)G T(m)
Proof Suppose that dim ( V) =nand then(not necessarily distinct) eigenvalues of Tare1; 2; 3; :::; n.
We begin with a basis of Vthat yields an upper triangular matrix representation, as guaranteed by Theo-
rem UTMR [676], B=fx1;x2;x3; :::; xng. Since the matrix representation is upper triangular, and the
eigenvalues of the linear transformation are the diagonal elements we can choose this basis so that there
are then scalars aij, 1jn, 1ij 1 such that
T(xj) =j 1X
i=1aijxi+jxj
We now dene a new basis for Vwhich is just a slight variation in the basis B. Choose any kand`such that
1k<`nandk6=`. Dene the scalar =akl=(` k). The new basis is C=fy1;y2;y3; :::; yng
Version 2.30
724 Section JCF Jordan Canonical Form
where
yj=xj; j6=`;1jn y`=x`+xk
We now compute the values of the linear transformation Twith inputs from C, noting carefully the
changed scalars in the linear combinations of Cdescribing the outputs. These changes will translate to
minor changes in the matrix representation built using the basis C. There are three cases to consider,
depending on which column of the matrix representation we are examining. First, assume j <` . Then
T(yj) =T(xj)
=j 1X
i=1aijxi+jxj
=j 1X
i=1aijyi+jyj
That seems a bit pointless. The rst ` 1 columns of the matrix representations of Trelative toBandC
are identical. OK, if that was too easy, here's the main act. Assume j=`. Then
T(y`) =T(x`+xk)
=T(x`) +T(xk)
= ` 1X
i=1ai`xi+`x`!
+ k 1X
i=1aikxi+kxk!
=` 1X
i=1ai`xi+`x`+k 1X
i=1aikxi+kxk
=` 1X
i=1ai`xi+k 1X
i=1aikxi+kxk+`x`
=` 1X
i=1
i6=kai`xi+k 1X
i=1aikxi+aklxk+kxk+`x`
=` 1X
i=1
i6=kai`xi+k 1X
i=1aikxi+aklxk+kxk `xk+`xk+`x`
=` 1X
i=1
i6=kai`xi+k 1X
i=1aikxi+ (akl+k `)xk+`(xk+x`)
=` 1X
i=1
i6=kai`xi+k 1X
i=1aikxi+ (akl+(k `))xk+`(x`+xk)
=` 1X
i=1
i6=kai`yi+k 1X
i=1aikyi+ (akl+(k `))yk+`y`
So how dierent are the matrix representations relative to BandCin column`? Fori>k , the coecient
ofyiisaij, as in the representation relative to B. It is a dierent story for ik, where the coecients of
yimay be very dierent. We are especially interested in the coecient of yk. In fact, this whole rst part
Version 2.30
Subsection JCF.GESD Generalized Eigenspace Decomposition 725
of this proof is about this particular entry of the matrix representation. The coecient of ykis
akl+(k `) =akl+akl
` k(k `)
=akl+ ( 1)akl
= 0
If the denition of was a mystery, then no more. In the matrix representation of Trelative toC, the
entry in column `, rowkis a zero. Nice. The only price we pay is that other entries in column `, specically
rows 1 through k 1, may also change in a way we can't control.
One more case to consider. Assume j >` . Then
T(yj) =T(xj)
=j 1X
i=1aijxi+jxj
=j 1X
i=1
i6=`;kaijxi+a`jx`+akjxk+jxj
=j 1X
i=1
i6=`;kaijxi+a`jx`+a`jxk a`jxk+akjxk+jxj
=j 1X
i=1
i6=`;kaijxi+a`j(x`+xk) + (akj a`j)xk+jxj
=j 1X
i=1
i6=`;kaijyi+a`jy`+ (akj a`j)yk+jyj
As before, we ask: how dierent are the matrix representations relative to BandCin columnj? Only
ykhas a coecient dierent from the corresponding coecient when the basis is B. So in the matrix
representations, the only entries to change are in row k, for columns `+ 1 through n.
What have we accomplished? With a change of basis, we can place a zero in a desired entry (row
k, column`) of the matrix representation, leaving most of the entries untouched. The only entries to
possibly change are above the new zero entry, or to the right of the new zero entry. Suppose we repeat
this procedure, starting by \zeroing out" the entry above the diagonal in the second column and rst row.
Then we move right to the third column, and zero out the element just above the diagonal in the second
row. Next we zero out the element in the third column and rst row. Then tackle the fourth column, work
upwards from the diagonal, zeroing out elements as we go. Entries above, and to the right will repeatedly
change, but newly created zeros will never get wrecked, since they are below, or just to the left of the
entry we are working on. Similarly the values on the diagonal do not change either. This entire argument
can be retooled in the language of change-of-basis matrices and similarity transformations, and this is the
approach taken by Noble in his Applied Linear Algebra . It is interesting to concoct the change-of-basis
matrix between the matrices BandCand compute the inverse.
Perhaps you have noticed that we have to be just a bit more careful than the previous paragraph
suggests. The denition of has a denominator that cannot be zero, which restricts our maneuvers to
zeroing out entries in row kand column `only whenk6=`. So we do not necessarily arrive at a diagonal
matrix. More carefully we can write
T(yj) =j 1X
i=1
i:i=jbijyi+jyj
Version 2.30
726 Section JCF Jordan Canonical Form
where thebijare our new coecients after repeated changes, the yjare the new basis vectors, and the
condition \i:i=j" means that we only have terms in the sum involving vectors whose nal coecients
are identical diagonal values (the eigenvalues). Now reorder the basis vectors carefully. Group together
vectors that have equal diagonal entries in the matrix representation, but within each group preserve
the order of the precursor basis. This grouping will create a block diagonal structure for the matrix
representation, while otherwise preserving the order of the basis will retain the upper triangular form of
the representation. So we can arrive at a basis that yields a matrix representation that is upper triangular
and block diagonal, with the diagonal entries of each block all equal to a common eigenvalue of the linear
transformation.
More carefully, employing the distinct eigenvalues of T,i, 1im, we can assert there is a set of
basis vectors for V,uij, 1im, 1jT(i), such that
T(uij) =j 1X
k=1bijkuik+iuij
So the subspace Ui=hfuijj1jT(i)gi, 1imis an invariant subspace of Vrelative toTand the
restrictionTjUihas an upper triangular matrix representation relative to the basis fuijj1jT(i)g
where the diagonal entries are all equal to i. Notice too that with this denition,
V=U1U2U3Um
Whew. This is a good place to take a break, grab a cup of coee, use the toilet, or go for a short stroll,
before we show that Uiis a subspace of the generalized eigenspace GT(i). This will follow if we can prove
that each of the basis vectors for Uiis a generalized eigenvector of Tfori(Denition GEV [707]). We
need some power of T iIVthat takes uijto the zero vector. We prove by induction on j(Technique I
[772]) the claim that ( T iIV)j(uij) =0. Forj= 1 we have,
(T iIV) (ui1) =T(ui1) iIV(ui1)
=T(ui1) iui1
=iui1 iui1
=0
For the induction step, assume that if k<j , then (T iIV)ktakes uikto the zero vector. Then
(T iIV)j(uij) = (T iIV)j 1((T iIV) (uij))
= (T iIV)j 1(T(uij) iIV(uij))
= (T iIV)j 1(T(uij) iuij)
= (T iIV)j 1 j 1X
k=1bijkuik+iuij iuij!
= (T iIV)j 1 j 1X
k=1bijkuik!
=j 1X
k=1bijk(T iIV)j 1(uik)
=j 1X
k=1bijk(T iIV)j 1 k
(T iIV)k(uik)
=j 1X
k=1bijk(T iIV)j 1 k(0)
Version 2.30
Subsection JCF.GESD Generalized Eigenspace Decomposition 727
=j 1X
k=1bijk0
=0
This completes the induction step. Since every vector of the spanning set for Uiis an element of the
subspaceGT(i), Property AC [317] and Property SC [317] allow us to conclude that UiGT(i). Then
by Denition S [333], Uiis a subspace ofGT(i). Notice that this inductive proof could be interpreted to
say that every element of Uiis a generalized eigenvector of Tfori, and the algebraic multiplicity of iis
a suciently high power to demonstrate this via the denition for each vector.
We are now prepared for our nal argument in this long proof. We wish to establish that the dimension
of the subspaceGT(i) is the algebraic multiplicity of i. This will be enough to show that UiandGT(i)
are equal, and will nally provide the desired direct sum decomposition.
We will prove by induction (Technique I [772]) the following claim. Suppose that T:V!Vis a linear
transformation and Bis a basis for Vthat provides an upper triangular matrix representation of T. The
number of times any eigenvalue occurs on the diagonal of the representation is greater than or equal to
the dimension of the generalized eigenspace GT().
We will use the symbol mfor the dimension of Vso as to avoid confusion with our notation for the
nullity. So dim V=mand our proof will proceed by induction on m. Use the notation # T() to count
the number of times occurs on the diagonal of a matrix representation of T. We want to show that
#T()dim (GT())
= dim (K((T )m)) Theorem GEK [708]
=n((T )m) Denition NOLT [588]
For the base case, dim V= 1. Every matrix representation of Tis an upper triangular matrix with the
lone eigenvalue of T,, as the single diagonal entry. So # T() = 1. The generalized eigenspace of is
not trivial (since by Theorem GEK [708] it equals the regular eigenspace), so it cannot be a subspace of
dimension zero, and thus dim ( GT()) = 1.
Now for the induction step, assume the claim is true for any linear transformation dened on a vector
space with dimension m 1 or less. Suppose that B=fv1;v2;v3; :::; vmgis a basis for Vthat yields
an upper triangular matrix representation for Twith diagonal entries 1; 2; 3; :::; m. ThenU=
hfv1;v2;v3; :::; vm 1giis a subspace of Vthat is invariant relative to T. The restriction TjU:U!U
is then a linear transformation dened on U, a vector space of dimension m 1. A matrix representation
ofTjUrelative to the basis C=fv1;v2;v3; :::; vm 1gwill be an upper triangular matrix with diagonal
entries1; 2; 3; :::; m 1. We can therefore apply the induction hypothesis to TjUand its representation
relative toC.
Suppose that is any eigenvalue of T. Then suppose that v2K((T IV)m). As an element of V,
we can write vas a linear combination of the basis elements of B, or more compactly, there is a vector
u2Uand a scalar such that v=u+vm. Then,
(m )mvm
=(T IV)m(vm) Theorem EOMP [481]
=0+(T IV)m(vm) Property Z [318]
= (T IV)m(u) + (T IV)m(u) +(T IV)m(vm) Property AI [318]
= (T IV)m(u) + (T IV)m(u+vm) Theorem LTLC [525]
= (T IV)m(u) + (T IV)m(v) Theorem LTLC [525]
= (T IV)m(u) +0 Denition KLT [545]
= (T IV)m(u) Property Z [318]
Version 2.30
728 Section JCF Jordan Canonical Form
The nal expression in this string of equalities is an element of UsinceUis invariant relative to both
TandIV. The expression at the beginning is a scalar multiple of vm, and as such cannot be a nonzero
element ofUwithout violating the linear independence of B. So
(m )mvm=0
The vector vmis nonzero since Bis linearly independent, so Theorem SMEZV [326] tells us that (m )m=
0. From the properties of scalar multiplication, we are confronted with two possibilities.
Our rst case is that 6=m. Notice then that occurs the same number of times along the diagonal
in the representations of TjUandT. Now= 0 and v=u+ 0vm=u. Since vwas chosen as an arbitrary
element ofK((T IV)m), Denition SSET [761] says that K((T IV)m)U. It is always the case that
K((TjU IU)m)K((T IV)m). However, we can also see that in this case, the opposite set inclusion
is true as well. By Denition SE [762] we have K((TjU IU)m) =K((T IV)m). Then
#T() = #TjU()
dim
GTjU()
Induction Hypothesis
= dim
K
(TjU IU)m 1
Theorem GEK [708]
= dim (K((TjU IU)m)) Theorem KPLT [691]
= dim (K((T IV)m))
= dim (GT()) Theorem GEK [708]
The second case is that =m. Notice then that occurs one more time along the diagonal in the
representation of Tcompared to the representation of TjU. Then
(TjU IU)m(u) = (T IV)m(u)
= (T IV)m(u) +0 Property Z [318]
= (T IV)m(u) +(m )mvm Theorem ZSSM [324]
= (T IV)m(u) +(T IV)m(vm) Theorem EOMP [481]
= (T IV)m(u+vm) Theorem LTLC [525]
= (T IV)m(v)
=0 Denition KLT [545]
Sou2K((TjU IU)m). The vector vwas chosen as an arbitrary member of K((T IV)m). From the
expression v=u+vmwe can now see valso as an element of K((TjU IU)m) plus a scalar multiple
ofvm. This observation yields
dim (K((T IV)m))dim (K((TjU IU)m)) + 1
Now count eigenvalues on the diagonal,
#T() = #TjU() + 1
dim
GTjU()
+ 1 Induction Hypothesis
= dim
K
(TjU IU)m 1
+ 1 Theorem GEK [708]
= dim (K((TjU IU)m)) + 1 Theorem KPLT [691]
dim (K((T IV)m))
Version 2.30
Subsection JCF.JCF Jordan Canonical Form 729
= dim (GT()) Theorem GEK [708]
In Theorem UTMR [676] we constructed an upper triangular matrix representation of Twhere each
eigenvalue occurred T() times on the diagonal. So
T(i) = #T(i) Theorem UTMR [676]
dim (GT(i))
dim (Ui) Theorem PSSD [410]
=T(i) Theorem PSSD [410]
Thus, dim (GT(i)) =T(i) and by Theorem EDYES [410], Ui=GT(i) and we can write
V=U1U2U3Um
=GT(1)GT(2)GT(3)G T(m)
Besides a nice decomposition into invariant subspaces, this proof has a bonus for us.
Theorem DGES
Dimension of Generalized Eigenspaces
SupposeT:V!Vis a linear transformation with eigenvalue . Then the dimension of the generalized
eigenspace for is the algebraic multiplicity of , dim (GT(i)) =T(i).
Proof At the very end of the proof of Theorem GESD [721] we obtain the inequalities
T(i)dim (GT(i))T(i)
which establishes the desired equality.
Subsection JCF
Jordan Canonical Form
Now we are in a position to dene what we (and others) regard as an especially nice matrix representation.
The word \canonical" has at its root, the word \canon," which has various meanings. One is the set of laws
established by a church council. Another is a set of writings that are authentic, important or representative.
Here we take it to mean the accepted, or best, representative among a variety of choices. Every linear
transformation admits a variety of representations, and we will declare one as the best. Hopefully you will
agree.
Denition JCF
Jordan Canonical Form
A square matrix is in Jordan canonical form if it meets the following requirements:
1. The matrix is block diagonal.
2. Each block is a Jordan block.
3. If< then the block Jk() occupies rows with indices greater than the indices of the rows occupied
byJ`().
Version 2.30
730 Section JCF Jordan Canonical Form
4. If=and` < k , then the block J`() occupies rows with indices greater than the indices of the
rows occupied by Jk().
4
Theorem JCFLT
Jordan Canonical Form for a Linear Transformation
SupposeT:V!Vis a linear transformation. Then there is a basis BforVsuch that the matrix
representation of Twith the following properties:
1. The matrix representation is in Jordan canonical form.
2. IfJk() is one of the Jordan blocks, then is an eigenvalue of T.
3. For a xed value of , the largest block of the form Jk() has size equal to the index of ,T().
4. For a xed value of , the number of blocks of the form Jk() is the geometric multiplicity of ,
T().
5. For a xed value of , the number of rows occupied by blocks of the form Jk() is the algebraic
multiplicity of ,T().
Proof This theorem is really just the consequence of applying to T, consecutively Theorem GESD [721],
Theorem MRRGE [719] and Theorem CFNLT [694].
Theorem GESD [721] gives us a decomposition of Vinto generalized eigenspaces, one for each distinct
eigenvalue. Since these generalized eigenspaces ar invariant relative to T, this provides a block diagonal
matrix representation where each block is the matrix representation of the restriction of Tto the generalized
eigenspace.
Restricting Tto a generalized eigenspace results in a \nearly nilpotent" linear transformation, as stated
more precisely in Theorem RGEN [716]. We unravel Theorem RGEN [716] in the proof of Theorem MRRGE
[719] so that we can apply Theorem CFNLT [694] about representations of nilpotent linear transformations.
We know the dimension of a generalized eigenspace is the algebraic multiplicity of the eigenvalue
(Theorem DGES [727]), so the blocks associated with the generalized eigenspaces are square with a size
equal to the algebraic multiplicity. In rening the basis for this block, and producing Jordan blocks the
results of Theorem CFNLT [694] apply. The total number of blocks will be the nullity of TjGT() IGT(),
which is the geometric multiplicity of as an eigenvalue of T(Denition GME [463]). The largest of the
Jordan blocks will have size equal to the index of the nilpotent linear transformation TjGT() IGT(),
which is exactly the denition of the index of the eigenvalue (Denition IE [717]).
Before we do some examples of this result, notice how close Jordan canonical form is to a diagonal
matrix. Or, equivalently, notice how close we have come to diagonalizing a matrix (Denition DZM
[496]). We have a matrix representation which has diagonal entries that are the eigenvalues of a matrix.
Each occurs on the diagonal as many times as the algebraic multiplicity. However, when the geometric
multiplicity is strictly less than the algebraic multiplicity, we have some entries in the representation just
above the diagonal (the \superdiagonal"). Furthermore, we have some idea how often this happens if we
know the geometric multiplicity and the index of the eigenvalue.
We now recognize just how simple a diagonalizable linear transformation really is. For each eigenvalue,
the generalized eigenspace is just the regular eigenspace, and it decomposes into a direct sum of one-
dimensional subspaces, each spanned by a dierent eigenvector chosen from a basis of eigenvectors for the
eigenspace.
Some authors create matrix representations of nilpotent linear transformations where the Jordan block
has the ones just below the diagonal (the \subdiagonal"). No matter, it is really the same, just dierent.
Version 2.30
Subsection JCF.JCF Jordan Canonical Form 731
We have also dened Jordan canonical form to place blocks for the larger eigenvalues earlier, and for blocks
with the same eigenvalue, we place the bigger ones earlier. This is fairly standard, but there is no reason we
couldn't order the blocks dierently. It'd be the same, just dierent. The reason for choosing some ordering
is to be assured that there is just onecanonical matrix representation for each linear transformation.
Example JCF10
Jordan canonical form, size 10
Suppose that T:C10!C10is the linear transformation dened by T(x) =Axwhere
A=2
666666666666664 6 9 7 5 5 12 22 14 8 21
3 5 3 1 2 7 12 9 1 12
8 9 8 6 0 14 25 13 4 26
7 9 7 5 0 13 23 13 2 24
0 1 0 1 3 2 3 4 2 3
3 2 1 2 9 1 1 5 5 5
1 3 3 2 4 3 6 4 4 3
3 4 3 2 1 5 9 5 1 9
0 2 0 0 2 2 4 4 2 4
4 4 5 4 1 6 11 4 1 103
777777777777775
We'll nd a basis for C10that will yield a matrix representation of Tin Jordan canonical form. First we
nd the eigenvalues, and their multiplicities, with the techniques of Chapter E [453].
= 2 T(2) = 2
T(2) = 2
= 0 T(0) = 3
T( 1) = 2
= 1 T( 1) = 5
T( 1) = 2
For each eigenvalue, we can compute a generalized eigenspace. By Theorem GESD [721] we know that
C10will decompose into a direct sum of these eigenspaces, and we can restrict Tto each part of this
decomposition. At this stage we know that the Jordan canonical form will be block diagonal with blocks of
size 2, 3 and 5, since the dimensions of the generalized eigenspaces are equal to the algebraic multiplicities
of the eigenvalues (Theorem DGES [727]). The geometric multiplicities tell us how many Jordan blocks
occupy each of the three larger blocks, but we will discuss this as we analyze each eigenvalue. We do not
yet know the index of each eigenvalue (though we can easily infer it for = 2) and even if we did have this
information, it only determines the size of the largest Jordan block (per eigenvalue). We will press ahead,
considering each eigenvalue one at a time.
The eigenvalue = 2 has \full" geometric multiplicity, and is not an impediment to diagonalizing T.
We will treat it in full generality anyway. First we compute the generalized eigenspace. Since Theorem
GEK [708] says that GT(2) =K
(T 2IC10)10
we can compute this generalized eigenspace as a null space
derived from the matrix A,
(A 2I10)10RREF !2
6666666666666666410 0 0 0 0 0 0 2 1
010 0 0 0 0 0 1 1
0 0 10 0 0 0 0 1 2
0 0 0 10 0 0 0 1 2
0 0 0 0 10 0 0 1 0
0 0 0 0 0 10 0 2 1
0 0 0 0 0 0 10 1 0
0 0 0 0 0 0 0 1 0 1
0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 03
77777777777777775
Version 2.30
732 Section JCF Jordan Canonical Form
GT(2) =K
(A 2I10)10
=*8
>>>>>>>>>>>>>><
>>>>>>>>>>>>>>:2
6666666666666642
1
1
1
1
2
1
0
1
03
777777777777775;2
6666666666666641
1
2
2
0
1
0
1
0
13
7777777777777759
>>>>>>>>>>>>>>=
>>>>>>>>>>>>>>;+
The restriction of TtoGT(2) relative to the two basis vectors above has a matrix representation that is a
22 diagonal matrix with the eigenvalue = 2 as the diagonal entries. So these two vectors will be the
rst two vectors in our basis for C10,
v1=2
6666666666666642
1
1
1
1
2
1
0
1
03
777777777777775v2=2
6666666666666641
1
2
2
0
1
0
1
0
13
777777777777775
Notice that it was not strictly necessary to compute the 10-th power of A 2I10. WithT(2) =
T(2)
the null space of the matrix A 2I10contains allof the generalized eigenvectors of Tfor the eigenvalue
= 2. But there was no harm in computing the 10-th power either. This discussion is equivalent to the
observation that the linear transformation TjGT(2):GT(2)!GT(2) is nilpotent of index 1. In other words,
T(2) = 1.
The eigenvalue = 0 will not be quite as simple, since the geometric multiplicity is strictly less than
the geometric multiplicity. As before, we rst compute the generalized eigenspace. Since Theorem GEK
[708] says thatGT(0) =K
(T 0IC10)10
we can compute this generalized eigenspace as a null space
derived from the matrix A,
(A 0I10)10RREF !2
666666666666666410 0 0 0 0 0 0 1 1
010 0 0 0 1 0 1 0
0 0 10 0 0 0 0 1 2
0 0 0 10 0 0 0 2 1
0 0 0 0 10 0 0 1 0
0 0 0 0 0 1 1 0 1 2
0 0 0 0 0 0 0 1 1 0
0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 03
7777777777777775
Version 2.30
Subsection JCF.JCF Jordan Canonical Form 733
GT(0) =K
(A 0I10)10
=*8
>>>>>>>>>>>>>><
>>>>>>>>>>>>>>:2
6666666666666640
1
0
0
0
1
1
0
0
03
777777777777775;2
6666666666666641
1
1
2
1
1
0
1
1
03
777777777777775;2
6666666666666641
0
2
1
0
2
0
0
0
13
7777777777777759
>>>>>>>>>>>>>>=
>>>>>>>>>>>>>>;+
=hFi
So dim (GT(0)) = 3 = T(0), as expected. We will use these three basis vectors for the generalized
eigenspace to construct a matrix representation of TjGT(0), whereFis being dened implicitly as the basis
ofGT(0). We construct this representation as usual, applying Denition MR [615],
F0
BBBBBBBBBBBBBB@TjGT(0)0
BBBBBBBBBBBBBB@2
6666666666666640
1
0
0
0
1
1
0
0
03
7777777777777751
CCCCCCCCCCCCCCA1
CCCCCCCCCCCCCCA=F0
BBBBBBBBBBBBBB@2
666666666666664 1
0
2
1
0
2
0
0
0
13
7777777777777751
CCCCCCCCCCCCCCA=F0
BBBBBBBBBBBBBB@( 1)2
6666666666666641
0
2
1
0
2
0
0
0
13
7777777777777751
CCCCCCCCCCCCCCA=2
40
0
13
5
F0
BBBBBBBBBBBBBB@TjGT(0)0
BBBBBBBBBBBBBB@2
6666666666666641
1
1
2
1
1
0
1
1
03
7777777777777751
CCCCCCCCCCCCCCA1
CCCCCCCCCCCCCCA=F0
BBBBBBBBBBBBBB@2
6666666666666641
0
2
1
0
2
0
0
0
13
7777777777777751
CCCCCCCCCCCCCCA=F0
BBBBBBBBBBBBBB@(1)2
6666666666666641
0
2
1
0
2
0
0
0
13
7777777777777751
CCCCCCCCCCCCCCA=2
40
0
13
5
F0
BBBBBBBBBBBBBB@TjGT(0)0
BBBBBBBBBBBBBB@2
6666666666666641
0
2
1
0
2
0
0
0
13
7777777777777751
CCCCCCCCCCCCCCA1
CCCCCCCCCCCCCCA=F0
BBBBBBBBBBBBBB@2
6666666666666640
0
0
0
0
0
0
0
0
03
7777777777777751
CCCCCCCCCCCCCCA=2
40
0
03
5
So we have the matrix representation
M=MTjGT(0)
F;F=2
40 0 0
0 0 0
1 1 03
5
Version 2.30
734 Section JCF Jordan Canonical Form
By Theorem RGEN [716] we can obtain a nilpotent matrix from this matrix representation by subtracting
the eigenvalue from the diagonal elements, and then we can apply Theorem CFNLT [694] to M (0)I3.
First check that ( M (0)I3)2=O, so we know that the index of M (0)I3as a nilpotent matrix, and that
therefore= 0 is an eigenvalue of Twith index 2, T(0) = 2. To determine a basis of C3that converts
M (0)I3to canonical form, we need the null spaces of the powers of M (0)I3. For convenience, set
N=M (0)I3.
N
N1
=*8
<
:2
41
1
03
5;2
40
0
13
59
=
;+
N
N2
=*8
<
:2
41
0
03
5;2
40
1
03
5;2
40
0
13
59
=
;+
=C3
Then we choose a vector from N
N2
that is not an element of N
N1
. Any vector with unequal rst
two entries will t the bill, say
z2;1=2
41
0
03
5
where we are employing the notation in Theorem CFNLT [694]. The next step is to multiply this vector
byNto get part of the basis for N
N1
,
z1;1=Nz2;1=2
40 0 0
0 0 0
1 1 03
52
41
0
03
5=2
40
0
13
5
We need a vector to pair with z1;1that will make a basis for the two-dimensional subspace N
N1
.
Examining the basis for N
N1
we see that a vector with its rst two entries equal will do the job.
z1;2=2
41
1
03
5
Reordering, we nd the basis,
C=fz1;1;z2;1;z1;2g=8
<
:2
40
0
13
5;2
41
0
03
5;2
41
1
03
59
=
;
From this basis, we can get a matrix representation of N(when viewed as a linear transformation) relative
to the basis CforC3,
2
40 1 0
0 0 0
0 0 03
5=J2(0)O
OJ1(0)
Now we add back the eigenvalue = 0 to the representation of Nto obtain a representation for M. Of
course, with an eigenvalue of zero, the change is not apparent, so we won't display the same matrix again.
This is the second block of the Jordan canonical form for T. However, the three vectors in Cwill not
suce as basis vectors for the domain of T| they have the wrong size! The vectors in Care vectors in
the domain of a linear transformation dened by the matrix M. ButMwas a matrix representation of
Version 2.30
Subsection JCF.JCF Jordan Canonical Form 735
TjGT(0) 0IGT(0)relative to the basis FforGT(0). We need to \uncoordinatize" each of the basis vectors
inCto produce a linear combination of vectors in Fthat will be an element of the generalized eigenspace
GT(0). These will be the next three vectors of our nal answer, a basis for C10that has a pleasing matrix
representation.
v3= 1
F0
@2
40
0
13
51
A= 02
6666666666666640
1
0
0
0
1
1
0
0
03
777777777777775+ 02
6666666666666641
1
1
2
1
1
0
1
1
03
777777777777775+ ( 1)2
6666666666666641
0
2
1
0
2
0
0
0
13
777777777777775=2
666666666666664 1
0
2
1
0
2
0
0
0
13
777777777777775
v4= 1
F0
@2
41
0
03
51
A= 12
6666666666666640
1
0
0
0
1
1
0
0
03
777777777777775+ 02
6666666666666641
1
1
2
1
1
0
1
1
03
777777777777775+ 02
6666666666666641
0
2
1
0
2
0
0
0
13
777777777777775=2
6666666666666640
1
0
0
0
1
1
0
0
03
777777777777775
v5= 1
F0
@2
41
1
03
51
A= 12
6666666666666640
1
0
0
0
1
1
0
0
03
777777777777775+ 12
6666666666666641
1
1
2
1
1
0
1
1
03
777777777777775+ 02
6666666666666641
0
2
1
0
2
0
0
0
13
777777777777775=2
6666666666666641
2
1
2
1
2
1
1
1
03
777777777777775
Five down, ve to go. Basis vectors, that is. = 1 is the smallest eigenvalue, but it will require the
most computation. First we compute the generalized eigenspace. Since Theorem GEK [708] says that
GT( 1) =K
(T ( 1)IC10)10
we can compute this generalized eigenspace as a null space derived from
the matrix A,
(A ( 1)I10)10RREF !2
666666666666666410 1 0 1 0 1 1 0 1
010 0 1 0 0 1 0 0
0 0 0 11 0 1 0 0 2
0 0 0 0 0 1 2 1 0 2
0 0 0 0 0 0 0 0 1 0
0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 03
7777777777777775
Version 2.30
736 Section JCF Jordan Canonical Form
GT( 1) =K
(A ( 1)I10)10
=*8
>>>>>>>>>>>>>><
>>>>>>>>>>>>>>:2
666666666666664 1
0
1
0
0
0
0
0
0
03
777777777777775;2
666666666666664 1
1
0
1
1
0
0
0
0
03
777777777777775;2
6666666666666641
0
0
1
0
2
1
0
0
03
777777777777775;2
666666666666664 1
1
0
0
0
1
0
1
0
03
777777777777775;2
666666666666664 1
0
0
2
0
2
0
0
0
13
7777777777777759
>>>>>>>>>>>>>>=
>>>>>>>>>>>>>>;+
=hFi
So dim (GT( 1)) = 5 = T( 1), as expected. We will use these ve basis vectors for the generalized
eigenspace to construct a matrix representation of TjGT( 1), whereFis being recycled and dened now
implicitly as the basis of GT( 1). We construct this representation as usual, applying Denition MR [615],
F0
BBBBBBBBBBBBBB@TjGT( 1)0
BBBBBBBBBBBBBB@2
666666666666664 1
0
1
0
0
0
0
0
0
03
7777777777777751
CCCCCCCCCCCCCCA1
CCCCCCCCCCCCCCA=F0
BBBBBBBBBBBBBB@2
666666666666664 1
0
0
0
0
2
2
0
0
13
7777777777777751
CCCCCCCCCCCCCCA
=F0
BBBBBBBBBBBBBB@02
666666666666664 1
0
1
0
0
0
0
0
0
03
777777777777775+ 02
666666666666664 1
1
0
1
1
0
0
0
0
03
777777777777775+ ( 2)2
6666666666666641
0
0
1
0
2
1
0
0
03
777777777777775+ 02
666666666666664 1
1
0
0
0
1
0
1
0
03
777777777777775+ ( 1)2
666666666666664 1
0
0
2
0
2
0
0
0
13
7777777777777751
CCCCCCCCCCCCCCA=2
666640
0
2
0
13
77775
F0
BBBBBBBBBBBBBB@TjGT( 1)0
BBBBBBBBBBBBBB@2
666666666666664 1
1
0
1
1
0
0
0
0
03
7777777777777751
CCCCCCCCCCCCCCA1
CCCCCCCCCCCCCCA=F0
BBBBBBBBBBBBBB@2
6666666666666647
1
5
3
1
2
4
0
0
33
7777777777777751
CCCCCCCCCCCCCCA
Version 2.30
Subsection JCF.JCF Jordan Canonical Form 737
=F0
BBBBBBBBBBBBBB@( 5)2
666666666666664 1
0
1
0
0
0
0
0
0
03
777777777777775+ ( 1)2
666666666666664 1
1
0
1
1
0
0
0
0
03
777777777777775+ 42
6666666666666641
0
0
1
0
2
1
0
0
03
777777777777775+ 02
666666666666664 1
1
0
0
0
1
0
1
0
03
777777777777775+ 32
666666666666664 1
0
0
2
0
2
0
0
0
13
7777777777777751
CCCCCCCCCCCCCCA=2
66664 5
1
4
0
33
77775
F0
BBBBBBBBBBBBBB@TjGT( 1)0
BBBBBBBBBBBBBB@2
6666666666666641
0
0
1
0
2
1
0
0
03
7777777777777751
CCCCCCCCCCCCCCA1
CCCCCCCCCCCCCCA=F0
BBBBBBBBBBBBBB@2
6666666666666641
0
1
1
0
0
1
0
0
13
7777777777777751
CCCCCCCCCCCCCCA
=F0
BBBBBBBBBBBBBB@( 1)2
666666666666664 1
0
1
0
0
0
0
0
0
03
777777777777775+ 02
666666666666664 1
1
0
1
1
0
0
0
0
03
777777777777775+ 12
6666666666666641
0
0
1
0
2
1
0
0
03
777777777777775+ 02
666666666666664 1
1
0
0
0
1
0
1
0
03
777777777777775+ 12
666666666666664 1
0
0
2
0
2
0
0
0
13
7777777777777751
CCCCCCCCCCCCCCA=2
66664 1
0
1
0
13
77775
F0
BBBBBBBBBBBBBB@TjGT( 1)0
BBBBBBBBBBBBBB@2
666666666666664 1
1
0
0
0
1
0
1
0
03
7777777777777751
CCCCCCCCCCCCCCA1
CCCCCCCCCCCCCCA=F0
BBBBBBBBBBBBBB@2
666666666666664 1
0
2
2
1
1
1
1
0
23
7777777777777751
CCCCCCCCCCCCCCA
Version 2.30
738 Section JCF Jordan Canonical Form
=F0
BBBBBBBBBBBBBB@22
666666666666664 1
0
1
0
0
0
0
0
0
03
777777777777775+ ( 1)2
666666666666664 1
1
0
1
1
0
0
0
0
03
777777777777775+ ( 1)2
6666666666666641
0
0
1
0
2
1
0
0
03
777777777777775+ 12
666666666666664 1
1
0
0
0
1
0
1
0
03
777777777777775+ ( 2)2
666666666666664 1
0
0
2
0
2
0
0
0
13
7777777777777751
CCCCCCCCCCCCCCA=2
666642
1
1
1
23
77775
F0
BBBBBBBBBBBBBB@TjGT( 1)0
BBBBBBBBBBBBBB@2
666666666666664 1
0
0
2
0
2
0
0
0
13
7777777777777751
CCCCCCCCCCCCCCA1
CCCCCCCCCCCCCCA=F0
BBBBBBBBBBBBBB@2
666666666666664 7
1
6
5
1
2
6
2
0
63
7777777777777751
CCCCCCCCCCCCCCA
=F0
BBBBBBBBBBBBBB@62
666666666666664 1
0
1
0
0
0
0
0
0
03
777777777777775+ ( 1)2
666666666666664 1
1
0
1
1
0
0
0
0
03
777777777777775+ ( 6)2
6666666666666641
0
0
1
0
2
1
0
0
03
777777777777775+ 22
666666666666664 1
1
0
0
0
1
0
1
0
03
777777777777775+ ( 6)2
666666666666664 1
0
0
2
0
2
0
0
0
13
7777777777777751
CCCCCCCCCCCCCCA=2
666646
1
6
2
63
77775
So we have the matrix representation of the restriction of T(again recycling and redening the matrix M)
M=MTjGT( 1)
F;F=2
666640 5 1 2 6
0 1 0 1 1
2 4 1 1 6
0 0 0 1 2
1 3 1 2 63
77775
By Theorem RGEN [716] we can obtain a nilpotent matrix from this matrix representation by subtracting
the eigenvalue from the diagonal elements, and then we can apply Theorem CFNLT [694] to M ( 1)I5.
First check that ( M ( 1)I5)3=O, so we know that the index of M ( 1)I5as a nilpotent matrix, and
that therefore = 1 is an eigenvalue of Twith index 3, T( 1) = 3. To determine a basis of C5that
convertsM ( 1)I5to canonical form, we need the null spaces of the powers of M ( 1)I5. Again, for
convenience, set N=M ( 1)I5.
N
N1
=*8
>>>><
>>>>:2
666641
0
1
0
03
77775;2
66664 3
1
0
2
23
777759
>>>>=
>>>>;+
Version 2.30
Subsection JCF.JCF Jordan Canonical Form 739
N
N2
=*8
>>>><
>>>>:2
666643
1
0
0
03
77775;2
666641
0
1
0
03
77775;2
666640
0
0
1
03
77775;2
66664 3
0
0
0
13
777759
>>>>=
>>>>;+
N
N3
=*8
>>>><
>>>>:2
666641
0
0
0
03
77775;2
666640
1
0
0
03
77775;2
666640
0
1
0
03
77775;2
666640
0
0
1
03
77775;2
666640
0
0
0
13
777759
>>>>=
>>>>;+
=C5
Then we choose a vector from N
N3
that is not an element of N
N2
. The sum of the four basis vectors
forN
N2
sum to a vector with all ve entries equal to 1. We will mess with the rst entry to create a
vector not inN
N2
,
z3;1=2
666640
1
1
1
13
77775
where we are employing the notation in Theorem CFNLT [694]. The next step is to multiply this vector
byNto get a portion of the basis for N
N2
,
z2;1=Nz3;1=2
666641 5 1 2 6
0 0 0 1 1
2 4 2 1 6
0 0 0 2 2
1 3 1 2 53
777752
666640
1
1
1
13
77775=2
666642
2
1
4
33
77775
We have a basis for the two-dimensional subspace N
N1
and we can add to that the vector z2;1and we
have three of four basis vectors for N
N2
. These three vectors span the subspace we call Q2. We need a
fourth vector outside of Q2to complete a basis of the four-dimensional subspace N
N2
. Check that the
vector
z2;2=2
666643
1
3
1
13
77775
is an element ofN
N2
that lies outside of the subspace Q2. This vector was constructed by getting a
nice basis for Q2and forming a linear combination of this basis that species three of the ve entries of
the result. Of the remaining two entries, one was changed to move the vector outside of Q2and this was
followed by a change to the remaining entry to place the vector into N
N2
. The vector z2;2is the lone
basis vector for the subspace we call R2.
The remaining two basis vectors are easy to come by. They are the result of applying Nto each of the
two most recently determined basis vectors,
z1;1=Nz2;1=2
666643
1
0
2
23
77775z1;2=Nz2;2=2
666643
2
3
4
43
77775
Version 2.30
740 Section JCF Jordan Canonical Form
Now we reorder these basis vectors, to arrive at the basis
C=fz1;1;z2;1;z3;1;z1;2;z2;2g=8
>>>><
>>>>:2
666643
1
0
2
23
77775;2
666642
2
1
4
33
77775;2
666640
1
1
1
13
77775;2
666643
2
3
4
43
77775;2
666643
1
3
1
13
777759
>>>>=
>>>>;
A matrix representation of Nrelative toCis
2
666640 1 0 0 0
0 0 1 0 0
0 0 0 0 0
0 0 0 0 1
0 0 0 0 03
77775=J3(0)O
OJ2(0)
To obtain a matrix representation of M, we add back in the matrix ( 1)I5, placing the eigenvalue back
along the diagonal, and slightly modifying the Jordan blocks,
2
66664 1 1 0 0 0
0 1 1 0 0
0 0 1 0 0
0 0 0 1 1
0 0 0 0 13
77775=J3( 1)O
OJ2( 1)
The basisCyields a pleasant matrix representation for the restriction of the linear transformation T
( 1)Ito the generalized eigenspace GT( 1). However, we must remember that these vectors in C5are
representations of vectors in C10relative to the basis F. Each needs to be \un-coordinatized" before joining
our nal basis. Here we go,
v6= 1
F0
BBBB@2
666643
1
0
2
23
777751
CCCCA= 32
666666666666664 1
0
1
0
0
0
0
0
0
03
777777777777775+ ( 1)2
666666666666664 1
1
0
1
1
0
0
0
0
03
777777777777775+ 02
6666666666666641
0
0
1
0
2
1
0
0
03
777777777777775+ 22
666666666666664 1
1
0
0
0
1
0
1
0
03
777777777777775+ ( 2)2
666666666666664 1
0
0
2
0
2
0
0
0
13
777777777777775=2
666666666666664 2
1
3
3
1
2
0
2
0
23
777777777777775
v7= 1
F0
BBBB@2
666642
2
1
4
33
777751
CCCCA= 22
666666666666664 1
0
1
0
0
0
0
0
0
03
777777777777775+ ( 2)2
666666666666664 1
1
0
1
1
0
0
0
0
03
777777777777775+ ( 1)2
6666666666666641
0
0
1
0
2
1
0
0
03
777777777777775+ 42
666666666666664 1
1
0
0
0
1
0
1
0
03
777777777777775+ ( 3)2
666666666666664 1
0
0
2
0
2
0
0
0
13
777777777777775=2
666666666666664 2
2
2
3
2
0
1
4
0
33
777777777777775
Version 2.30
Subsection JCF.JCF Jordan Canonical Form 741
v8= 1
F0
BBBB@2
666640
1
1
1
13
777751
CCCCA= 02
666666666666664 1
0
1
0
0
0
0
0
0
03
777777777777775+ 12
666666666666664 1
1
0
1
1
0
0
0
0
03
777777777777775+ 12
6666666666666641
0
0
1
0
2
1
0
0
03
777777777777775+ 12
666666666666664 1
1
0
0
0
1
0
1
0
03
777777777777775+ 12
666666666666664 1
0
0
2
0
2
0
0
0
13
777777777777775=2
666666666666664 2
2
0
0
1
1
1
1
0
13
777777777777775
v9= 1
F0
BBBB@2
666643
2
3
4
43
777751
CCCCA= 32
666666666666664 1
0
1
0
0
0
0
0
0
03
777777777777775+ ( 2)2
666666666666664 1
1
0
1
1
0
0
0
0
03
777777777777775+ ( 3)2
6666666666666641
0
0
1
0
2
1
0
0
03
777777777777775+ 42
666666666666664 1
1
0
0
0
1
0
1
0
03
777777777777775+ ( 4)2
666666666666664 1
0
0
2
0
2
0
0
0
13
777777777777775=2
666666666666664 4
2
3
3
2
2
3
4
0
43
777777777777775
v10= 1
F0
BBBB@2
666643
1
3
1
13
777751
CCCCA= 32
666666666666664 1
0
1
0
0
0
0
0
0
03
777777777777775+ 12
666666666666664 1
1
0
1
1
0
0
0
0
03
777777777777775+ 32
6666666666666641
0
0
1
0
2
1
0
0
03
777777777777775+ 12
666666666666664 1
1
0
0
0
1
0
1
0
03
777777777777775+ 12
666666666666664 1
0
0
2
0
2
0
0
0
13
777777777777775=2
666666666666664 3
2
3
2
1
3
3
1
0
13
777777777777775
To summarize, we list the entire basis B=fv1;v2;v3; :::; v10g,
v1=2
6666666666666642
1
1
1
1
2
1
0
1
03
777777777777775v2=2
6666666666666641
1
2
2
0
1
0
1
0
13
777777777777775v3=2
666666666666664 1
0
2
1
0
2
0
0
0
13
777777777777775v4=2
6666666666666640
1
0
0
0
1
1
0
0
03
777777777777775v5=2
6666666666666641
2
1
2
1
2
1
1
1
03
777777777777775
Version 2.30
742 Section JCF Jordan Canonical Form
v6=2
666666666666664 2
1
3
3
1
2
0
2
0
23
777777777777775v7=2
666666666666664 2
2
2
3
2
0
1
4
0
33
777777777777775v8=2
666666666666664 2
2
0
0
1
1
1
1
0
13
777777777777775v9=2
666666666666664 4
2
3
3
2
2
3
4
0
43
777777777777775v10=2
666666666666664 3
2
3
2
1
3
3
1
0
13
777777777777775
The resulting matrix representation is
MT
B;B=2
6666666666666642 0 0 0 0 0 0 0 0 0
0 2 0 0 0 0 0 0 0 0
0 0 0 1 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 1 1 0 0 0
0 0 0 0 0 0 1 1 0 0
0 0 0 0 0 0 0 1 0 0
0 0 0 0 0 0 0 0 1 1
0 0 0 0 0 0 0 0 0 13
777777777777775
If you are not inclined to check all of these computations, here are a few that should convince you of the
amazing properties of the basis B. Compute the matrix-vector products Avi, 1i10. In each case the
result will be a vector of the form vi+vi 1, whereis one of the eigenvalues (you should be able to
predict ahead of time which one) and2f0;1g.
Alternatively, if we can write inputs to the linear transformation Tas linear combinations of the vectors
inB(which we can do uniquely since Bis a basis, Theorem VRRB [360]), then the \action" of Tis reduced
to a matrix-vector product with the exceedingly simple matrix that is the Jordan canonical form. Wow!
Subsection CHT
Cayley-Hamilton Theorem
Jordan was a French mathematician who was active in the late 1800's. Cayley and Hamilton were 19th-
century contemporaries of Jordan from Britain. The theorem that bears their names is perhaps one of
the most celebrated in basic linear algebra. While our result applies only to vector spaces and linear
transformations with scalars from the set of complex numbers, C, the result is equally true if we restrict
our scalars to the real numbers, R. It says that every matrix satises its own characteristic polynomial.
Theorem CHT
Cayley-Hamilton Theorem
SupposeAis a square matrix with characteristic polynomial pA(x). ThenpA(A) =O.
Proof SupposeBandCare similar matrices via the matrix S,B=S 1CS, andq(x) is any polynomial.
Thenq(B) is similar to q(C) viaS,q(B) =S 1q(C)S. (See Example HPDM [502] for hints on how to
convince yourself of this.)
By Theorem JCFLT [728] and Theorem SCB [656] we know Ais similar to a matrix, J, in Jordan
canonical form. Suppose 1; 2; 3; :::; mare the distinct eigenvalues of A(and are therefore the eigen-
values and diagonal entries of J). Then by Theorem EMRCP [461] and Denition AME [463], we can
Version 2.30
Subsection JCF.CHT Cayley-Hamilton Theorem 743
factor the characteristic polynomial as
pA(x) = (x 1)A(1)(x 2)A(2)(x 3)A(3)(x m)A(m)
On substituting the matrix Jwe have
pA(J) = (J 1I)A(1)(J 2I)A(2)(J 3I)A(3)(J mI)A(m)
The matrix J kIwill be block diagonal, and the block arising from the generalized eigenspace for k
will have zeros along the diagonal. Suitably adjusted for matrices (rather than linear transformations),
Theorem RGEN [716] tells us this matrix is nilpotent. Since the size of this nilpotent matrix is equal to
the algebraic multiplicity of k, the power ( J kI)A(k)will be a zero matrix (Theorem KPNLT [692])
in the location of this block.
Repeating this argument for each of the meigenvalues will place a zero block in some term of the product
at every location on the diagonal. The entire product will then be zero blocks on the diagonal, and zero o
the diagonal. In other words, it will be the zero matrix. Since AandJare similar, pA(A) =pA(J) =O.
Version 2.30
744 Section JCF Jordan Canonical Form
Version 2.30
Annotated Acronyms JCF.R Representations 745
Annotated Acronyms R
Representations
Denition VR [603]
Matrix representations build on vector representations, so this is the denition that gets us started. A
representation depends on the choice of a single basis for the vector space. Theorem VRRB [360] is what
tells us this idea might be useful.
Theorem VRILT [608]
As an invertible linear transformation, vector representation allows us to translate, back and forth, between
abstract vector spaces ( V) and concrete vector spaces ( Cn). This is key to all our notions of representations
in this chapter.
Theorem CFDVS [608]
Every vector space with nite dimension \looks like" a vector space of column vectors. Vector representa-
tion is the isomorphism that establishes that these vector spaces are isomorphic.
Denition MR [615]
Building on the denition of a vector representation, we dene a representation of a linear transformation,
determined by a choice of two bases, one for the domain and one for the codomain. Notice that vectors
are represented by columnar lists of scalars, while linear transformations are represented by rectangular
tables of scalars. Building a matrix representation is as important a skill as row-reducing a matrix.
Theorem FTMR [617]
Denition MR [615] is not really very interesting until we have this theorem. The second form tells us that
we can compute outputs of linear transformations via matrix multiplication, along with some bookkeeping
for vector representations. Searching forward through the text on \FTMR" is an interesting exercise. You
will nd reference to this result buried inside many key proofs at critical points, and it also appears in
numerous examples and solutions to exercises.
Theorem MRCLT [622]
Turns out that matrix multiplication is really a very natural operation, it is just the chaining together
(composition) of functions (linear transformations). Beautiful. Even if you don't try to work the problem,
study Solution MR.T80 [645] for more insight.
Theorem KNSI [625]
Kernels \are" null spaces. For this reason you'll see these terms used interchangeably.
Theorem RCSI [628]
Ranges \are" column spaces. For this reason you'll see these terms used interchangeably.
Theorem IMR [630]
Invertible linear transformations are represented by invertible (nonsingular) matrices.
Theorem NME9 [633]
The NMEx series has always been important, but we've held o saying so until now. This is the end of
the line for this one, so it is a good time to contemplate all that it means.
Version 2.30
746 Section JCF Jordan Canonical Form
Theorem SCB [656]
Diagonalization back in Section SD [493] was really a change of basis to achieve a diagonal matrix repe-
sentation. Maybe we should be highlighting the more general Theorem MRCB [654] here, but its overly
technical description just isn't as appealing. However, it will be important in some of the matrix decom-
postions in Chapter MD [903].
Theorem EER [659]
This theorem, with the companion denition, Denition EELT [647], tells us that eigenvalues, and eigen-
vectors, are fundamentally a characteristic of linear transformations (not matrices). If you study matrix
decompositions in Chapter MD [903] you will come to appreciate that almost all of a matrix's secrets can
be unlocked with knowledge of the eigenvalues and eigenvectors.
Theorem OD [681]
Can you imagine anything nicer than an orthonormal diagonalization? A basis of pairwise orthogonal, unit
norm, eigenvectors that provide a diagonal representation for a matrix? Here we learn just when this can
happen | precisely when a matrix is normal, which is a disarmingly simple property to dene.
Theorem CFNLT [694]
Nilpotent linear transformations are the fundamental obstacle to a matrix (or linear transformation) being
diagonalizable. This specialized representation theorem is the fundamental expression of just how close we
can come to surmounting the obstacle, i.e. how close we can come to a diagonal representation.
Theorem DGES [727]
This theorem is a long time in coming, but perhaps it best explains our interest in generalized eigenspaces.
When the dimension of a \regular" eigenspace (the geometic multiplicity) does not meet the algebraic
multiplicity of the corresponding eigenvalue, then a matrix is not diagonalizable (Theorem DMFE [499]).
However, if we generalize the idea of an eigenspace (Denition GES [707]), then we arrive at invariant
subspaces that together give a complete decomposition of the domain as a direct sum. And these subspaces
have dimensions equal to the corresponding algebraic multiplicities.
Theorem JCFLT [728]
If you can't diagonalize, just how close can you come? This is an answer (there are others, like rational
canonical form). \Canonicalism" is in the eye of the beholder. But this is a good place to conclude our
study of a widely accepted canonical form that is possible for every matrix or linear transformation.
Version 2.30
Appendix CN
Computation Notes
Section MMA
Mathematica
Computation Note ME.MMA
Matrix Entry
Matrices are input as lists of lists, since a list is a basic data structure in Mathematica . A matrix is a list
of rows, with each row entered as a list. Mathematica uses braces ((f,g)) to delimit lists. So the input
a=ff1;2;3;4g;f5;6;7;8g;f9;10;11;12gg
would create a 34 matrix named athat is equal to
2
41 2 3 4
5 6 7 8
9 10 11 123
5
To display a matrix named a\nicely" in Mathematica , type MatrixForm[a] , and the output will be
displayed with rows and columns. If you just type a, then you will get a list of lists, like how you input
the matrix in the rst place.
Computation Note RR.MMA
Row Reduce
Ifais the name of a matrix in Mathematica, then the command RowReduce[a] will output the reduced
row-echelon form of the matrix.
747
748 Section MMA Mathematica
Computation Note LS.MMA
Linear Solve
Mathematica will solve a linear system of equations using the LinearSolve[ ] command. The inputs are
a matrix with the coecients of the variables (but not the column of constants), and a list containing the
constant terms of each equation. This will look a bit odd, since the lists in the matrix are rows, but the
column of constants is also input as a list and so looks like a row rather than a column. The result will
be a single solution (even if there are innitely many), reported as a list, or the statement that there is
no solution. When there are innitely many, the single solution reported is exactly that solution used in
the proof of Theorem RCLS [58], where the free variables are all set to zero, and the dependent variables
come along with values from the nal column of the row-reduced matrix.
As an example, Archetype A [781] is
x1 x2+ 2x3= 1
2x1+x2+x3= 8
x1+x2= 5
To ask Mathematica for a solution, enter
LinearSolve [ff1; 1;2g;f2;1;1g;f1;1;0gg;f1;8;5g]
and you will get back the single solution
f3;2;0g
We will see later how to coax Mathematica into giving us innitely many solutions for this system (Com-
putation VFSS.MMA [747]).
Computation Note VLC.MMA
Vector Linear Combinations
Contributed by Robert Beezer
Vectors in Mathematica are represented as lists, written and displayed horizontally. For example, the
vector
v=2
6641
2
3
43
775
would be entered and named via the command
v=f1;2;3;4g
Vector addition and scalar multiplication are then very natural. If uand vare two lists of equal length,
then
2u+ ( 3)v
will compute the correct vector and return it as a list. If uand vhave dierent sizes, then Mathematica
will complain about \objects of unequal length."
Version 2.30
Computation Note MMA.NS.MMA Null Space 749
Computation Note NS.MMA
Null Space
Given a matrix A, Mathematica will compute a set of column vectors whose span is the null space of the ma-
trix with the NullSpace[ ] command. Perhaps not coincidentally, this set is exactly fzjj1jn rg.
However, Mathematica prefers to output the vectors in the opposite order than one we have chosen. Here's
a small example.
Begin with the 3 4 matrixA, and its row-reduced version B,
A=2
41 2 1 0
3 4 1 2
1 1 5 33
5RREF ! B=2
410 3 2
01 2 1
0 0 0 03
5
We could extract entries from Bto build the vectors z1andz2according to Theorem SSNS [137] and
describeN(A) as a span of the set fz1;z2g. Instead, if ahas been set to A, then executing the command
NullSpace[a] yields the list of lists (column vectors),
ff2; 1;0;1g;f 3;2;1;0gg
Notice how our z1is second in the list. To \correct" this we can use a list-processing command from
Mathematica, Reverse[ ] , as follows,
Reverse[NullSpace[a]]
and receive the output in our preferred order. Give it a try yourself.
Computation Note VFSS.MMA
Vector Form of Solution Set
Suppose that Ais anmnmatrix and b2Cmis a column vector. We might wish to nd all of the
solutions to the linear system LS(A;b). Mathematica's LinearSolve[A, b] will return at most one
solution (Computation LS.MMA [746]). However, when the system is consistent, then this one solution
reported is exactly the vector c, described in the statement of Theorem VFSLS [118].
The vectors uj, 1jn rof Theorem VFSLS [118] are exactly the output of Mathematica's
NullSpace[ ] command, though Mathematica lists them in the opposite order from the order we have
chosen. These are the same vectors listed as zj, 1jn rin Theorem SSNS [137]. With cproduced
from the LinearSolve[ ] command, and the ujcoming from the NullSpace[ ] command we can use
Mathematica's symbolic manipulation commands to create an expression that describes all of the solutions.
Begin with the system LS(A;b). Row-reduce A(Computation RR.MMA [745]) and identify the free
variables by determining the non-pivot columns. Suppose, for the sake of argument, that we have the
three free variables x3,x7andx8. Then the following command will build an expression for an arbitrary
solution:
LinearSolve[A, b]+ fx8, x7, x3g.NullSpace[A]
Be sure to include the \dot" right before the NullSpace[ ] command | it has the eect of creating
a linear combination of the vectors in the null space, using scalars that are symbols reminiscent of the
variables.
Version 2.30
750 Section MMA Mathematica
A concrete example should help here. Suppose we want a solution set for the linear system with
coecient matrix Aand vector of constants b,
A=2
41 2 3 5 1 1 2
2 4 0 8 4 1 8
3 6 4 0 2 5 73
5 b=2
48
1
53
5
If we were to apply Theorem VFSLS [118], we would extract the components of candujfrom the
row-reduced version of the augmented matrix of the system (obtained with Mathematica, Computation
RR.MMA [745]),
2
412 0 4 2 0 5 2
0 0 1 3 1 0 3 1
0 0 0 0 0 1 2 33
5
Instead, we will use this augmented matrix in reduced row-echelon form only to identify the free variables.
In this example, we locate the non-pivot columns and see that x2,x4,x5andx7are free. If we have set a
to the coecient matrix and bto the vector of constants, then we execute the Mathematica command,
LinearSolve[a, b]+ fx7, x5, x4, x2g.NullSpace[a]
As output we obtain the column vector (list),
2
6666666642 2x2 4x4+ 2x5+ 5x7
x2
1 + 3 x4 x5 3x7
x4
x5
3 2x7
x73
777777775
Computation Note GSP.MMA
Gram-Schmidt Procedure
Mathematica has a built-in routine that will do the Gram-Schmidt procedure (Theorem GSP [199]).
The input is a set of vectors, which must be linearly independent. This is written as a list, contain-
ing lists that are the vectors. Let abe such a list of lists, containing the vectors vi, 1ipfrom
the statement of the theorem. You will need to rst load the right Mathematica package | execute
<<LinearAlgebra`Orthogonalization` to make this happen. Then execute GramSchmidt[a] . The
output will be another list of lists containing the vectors ui, 1ipfrom the statement of the theorem.
Mathematica will complain if you do not provide a linearly independent set as input (try it!).
An example. Suppose our linearly independent set (check this!) is
S=8
>>>><
>>>>:2
66664 1
4
1
0
33
77775;2
666640
3
0
3
33
77775;2
66664 1
2
0
1
23
77775;2
66664 1
2
3
1
43
77775;2
666641
6
1
4
63
777759
>>>>=
>>>>;
Version 2.30
Computation Note MMA.TM.MMA Transpose of a Matrix 751
The output of the GramSchmidt[ ] command will be the set,
T=8
>>>>>>>><
>>>>>>>>:2
666664 1
3p
34
3p
31
3p
3
0
1p
33
777775;2
6666666641
12p
1523
12p
15
1
12p
15
3q
3
5
4
q
5
3
23
777777775;2
66666664 37
4p
68529
4p
685
3
4p
685
79
4p
685
5q
5
137
23
77777775;2
6666664 337
2p
120423
37
6p
120423
1763
6p
120423337
6p
12042350p
1204233
7777775;2
666666423p
87926
3p
879
44
3p
879
23
3p
8791p
8793
77777759
>>>>>>>>=
>>>>>>>>;
Ugly, but true. At this stage, you might just as well be encouraged to think of the Gram-Schmidt procedure
as a computational black box, linearly independent set in, orthogonal span-preserving set out.
To check that the output set is orthogonal, we can easily check the orthogonality of individual pairs
of vectors. Suppose the output was set equal to b(say via b=GramSchmidt[a] ). We can extract the
individual vectors of cas \parts" with syntax like c[[3]] , which would return the third vector in the
set. When our vectors have only real number entries, we can accomplish an innerproduct with a \dot."
So, for example, you should discover that c[[3]].c[[5]] will return zero. Try it yourself with another
pair of vectors.
Computation Note TM.MMA
Transpose of a Matrix
Contributed by Robert Beezer
Suppose ais the name of a matrix stored in Mathematica . Then Transpose[a] will create the transpose
ofa.
Computation Note MM.MMA
Matrix Multiplication
IfAandBare matrices dened in Mathematica , then A.B will return the product of the two matrices
(notice the dot between the matrices). If Ais a matrix and vis a vector, then A.v will return the vector
that is the matrix-vector product of Aandv. In every case the sizes of the matrices and vectors need to
be correct.
Some examples:
ff1;2g;f3;4gg:ff5;6;7g;f8;9;10gg=ff21;24;27g;f47;54;61gg
ff1;2g;f3;4gg:ff5g;f6gg=ff17g;f39gg
ff1;2g;f3;4gg:f5;6g=f17;39g
Understanding the dierence between the last two examples will go a long way to explaining how some
Mathematica constructs work.
Computation Note MI.MMA
Matrix Inverse
IfAis a matrix dened in Mathematica , then Inverse[A] will return the inverse of A, should it exist. In
the case where Adoes not have an inverse Mathematica will tell you the matrix is singular (see Theorem
NI [261]).
Version 2.30
752 Section TI86 Texas Instruments 86
Section TI86
Texas Instruments 86
Computation Note ME.TI86
Matrix Entry
On the TI-86, press the MATRX key (Yellow-7) . Press the second menu key over, F2, to bring up
the EDIT screen. Give your matrix a name, one letter or many, then press ENTER . You can then change
the size of the matrix (rows, then columns) and begin editing individual entries (which are initially zero).
ENTER will move you from entry to entry, or the down arrow key will move you to the next row. A menu
gives you extra options for editing.
Matrices may also be entered on the home screen as follows. Use brackets ([ , ]) to enclose rows with
elements separated by commas. Group rows, in order, into a nal set of brackets (with no commas between
rows). This can then be stored in a name with the STO key. So, for example,
[[1;2;3;4] [5;6;7;8] [9;10;11;12]]!A
will create a matrix named Athat is equal to
2
41 2 3 4
5 6 7 8
9 10 11 123
5
Computation Note RR.TI86
Row Reduce
IfAis the name of a matrix stored in the TI-86, then the command rref A will return the reduced
row-echelon form of the matrix. This command can also be found by pressing the MATRX key, then F4
forOPS, and nally, F5forrref .
Note that this command will not work for a matrix with more rows than columns. (Ed. Not sure just
why this is!) A work-around is to pad the matrix with extra columns of zeros until the matrix is square.
Computation Note VLC.TI86
Vector Linear Combinations
Contributed by Robert Beezer
Vector operations on the TI-86 can be accessed via the VECTR key, which is Yellow-8 . The EDIT tool
appears when the F2key is pressed. After providing a name and giving a \dimension" (the size) then
you can enter the individual entries, one at a time. Vectors can also be entered on the home screen using
brackets ( [,]). To create the vector
v=2
6641
2
3
43
775
Version 2.30
Computation Note TI86.TM.TI86 Transpose of a Matrix 753
use brackets and the store key ( STO),
[1;2;3;4]!v
Vector addition and scalar multiplication are then very natural. If uand vare two vectors of equal size,
then
2u+ ( 3)v
will compute the correct vector and display the result as a vector.
Computation Note TM.TI86
Transpose of a Matrix
Contributed by Eric Fickenscher
Suppose Ais the name of a matrix stored in the TI-86. Use the command ATto transpose A. This
command can be found by pressing the MATRX key, then F3forMATH , then F2forT.
Section TI83
Texas Instruments 83
Computation Note ME.TI83
Matrix Entry
Contributed by Douglas Phelps
On the TI-83, press the MATRX key. Press the right arrow key twice so that EDIT is highlighted. Move
the cursor down so that it is over the desired letter of the matrix and press ENTER . For example, let's call
our matrix B, so press the down arrow once and press ENTER . To enter a 23 matrix, press 2 ENTER
3 ENTER . To create the matrix1 2 3
4 5 6
press 1 ENTER 2 ENTER 3 ENTER 4 ENTER 5 ENTER 6 ENTER .
Computation Note RR.TI83
Row Reduce
Contributed by Douglas Phelps
Suppose Bis the name of a matrix stored in the TI-83. Press the MATRX key. Press the right arrow
key once so that MATH is highlighted. Press the down arrow eleven times so that rref ( is highlighted,
then press ENTER . to choose the matrix B, press MATRX , then the down arrow once followed by ENTER
. Supply a right parenthesis ( )) and press ENTER .
Note that this command will not work for a matrix with more rows than columns. (Ed. Not sure just
why this is!) A work-around is to pad the matrix with extra columns of zeros until the matrix is square.
Version 2.30
754 Section SAGE SAGE: Open Source Mathematics Software
Computation Note VLC.TI83
Vector Linear Combinations
Contributed by Douglas Phelps
Entering a vector on the TI-83 is the same process as entering a matrix. You press 4 ENTER 3 ENTER for
a 43 matrix. Likewise, you press 4 ENTER 1 ENTER for a vector of size 4. To multiply a vector by 8,
press the number 8, then press the MATRX key, then scroll down to the letter you named your vector (A,
B, C, etc) and press ENTER .
To add vectors Aand Bfor example, press the MATRX key, then ENTER . Then press the +key.
Then press the MATRX key, then the down arrow once, then ENTER .[A] + [B] will appear on the
screen. Press ENTER .
Section SAGE
SAGE: Open Source Mathematics Software
Computation Note R.SAGE
Rings
Contributed by Steve Caneld
SAGE uses dierent rings to denote the type of an object. The rings are as follows:
ZZ: The set of integers
QQ: The set of rational numbers
RR: The real numbers
CC: The complex numbers
Most objects in SAGE will tell you which they are using with the base ring() command. Keep this in
mind, especially when row reducing or factoring. Here's a quick example of where you might go wrong.
m=matrix ([[2;3];[4;7]])
m:base ring ()
IntegerRing
m:echelon form ()
2 0
0 1
As you can clearly see, misn't even in reduced row-echelon form. This is because mis dened over the
ZZ. You have to create matrices with the correct ring or you will get this type of odd result. This problem
comes up in more places than just calculating the reduced row-echelon form, so unless you are specically
working with integers take note.
Version 2.30
Computation Note SAGE.ME.SAGE Matrix Entry 755
Computation Note ME.SAGE
Matrix Entry
Contributed by Steve Caneld
A matrix in SAGE can be made a few ways. The rst is simply to dene the matrix as an array of rows.
SAGE uses brackets ([ , ]) to delimit arrays. So the input
a=matrix ([[1;2;3;4];[5;6;7;8];[9;10;11;12]])
would create a 34 matrix named athat is equal to
2
41 2 3 4
5 6 7 8
9 10 11 123
5
SAGE will guess what type of matrix you are working with based on the inputs. If all the entries are
integers, you will get back an integer matrix. If your matrix contains an entry in the RorCspace, the
matrix will be of those types. This can cause problems as integers cannot become fractions, which is an
issue when calculating reduced row-echelon form. We therefore recommend using the following construction
to make your matrices,
a=matrix (QQ;[[1;2;3;4];[5;6;7;8];[9;10;11;12]])
This gives you a matrix over the rational numbers which will be sucient for most of the course. If your
matrix has entries that are complex numbers you would replace the QQwith CC.
To display a matrix named a, type a, and the output will be displayed with rows and columns. If you
type latex(a) you will get L ATEX code to display the matrix. Very handy.
Computation Note RR.SAGE
Row Reduce
Contributed by Steve Caneld and Robert Beezer
Row-reducing a matrix is a simple operation in SAGE. However, because of Sage's
exibility with dierent
types of numbers (integers, rationals, reals, complexes), we need to be a bit more careful.
Ifais a matrix entered in in SAGE (see Computation ME.SAGE [753]) then a.echelon form() will
return a new matrix that is the reduced row-echelon form of a(Denition RREF [33]).
If your matrix has only integer entries (as is the case with many examples and exercises in this book),
then row operations might introduce rational numbers (\fractions"). So when you enter your matrix, you
need to tell SAGE that rational numbers are allowable in its calculations. This is the advice in Computation
R.SAGE [752] to use the ring QQ. As an illustration create
a=matrix (QQ;[[1;2;3;4];[5;6;7;8];[9;10;11;12]])
and issue the command a.echelon form() . The result is
2
41 0 1 2
0 1 2 3
0 0 0 03
5
However, if we adjust the entry by neglecting to specify QQ, then SAGE assumes that we only want to
work with integers, since every entry of the matrix is an integer. So as an experiment, enter
b=matrix ([[1;2;3;4];[5;6;7;8];[9;10;11;12]])
Version 2.30
756 Section SAGE SAGE: Open Source Mathematics Software
and issue the command b.echelon form() . The result is
2
41 2 3 4
0 4 8 12
0 0 0 03
5
You can now clearly see Sage's reluctance to multiply row 2 by1
4.
The ring QQwill of course suce if your matrix has rational numbers for entries. Decimal entries are
another place to be careful. If an entry of your matrix is the real number 2 :17, you are free to enter it as
the rational number217
100and keep the ring QQin the specication of your matrix. If you want to consider
your entries as real numbers, then you might as well just specify your ring as the complex numbers CC.
This advice also applies if you have complex numbers as entries.
If you allow SAGE to work with real or complex numbers, then the problem of round-o error becomes
relevant. Computer arithmetic with real numbers is, of necessity, subject to minor inaccuracies and errors.
This becomes problematic when row-reducing a matrix. If a zero entry is computed instead as an extremely
small number, such as 1 :28710 18, then an incorrect sequence of row operations will follow (with further
incorrect results). So if you use CCbe on the lookout for these kinds of potential pitfalls.
So, in summary, remember to always specify the ring you will be using for your matrices, and most
matrices can be handled with a choice of QQorCC.
When you need to do signicant scientic computing with SAGE, there are extra facilities that will
help you work with these subtleties.
Finally, you can also use a command of the form a.echelonize() to replace awith its reduced
row-echelon form.
Computation Note LS.SAGE
Linear Solve
SAGE can solve a variety of systems of equations with the solve( ) command, even when the equations
are not linear (see Exercise SSLE.M70 [22]). But we can aord to specialize here to just linear systems.
First, you must specify your variables in advance, so for example, var('x1,x2,x3') might precede a
system with three equations. Equations are then written just as you might expect, except that equality is
written as ==, since computer programs have traditionally reserved =to assign values to variables. And
remember to use a *to indicate that a coecient multiplies a variable.
The example below illustrates the use of the command and the possibilities for results. Each system
would be preceded by establishing the variables with the command var('x,y') . In the case of an innite
solution set, free variables are denoted as rxwhere xis an integer that increases throughout a session.
The style of this description of a solution set is reminiscent of the style we used in Chapter SLE [3] before
we were accustomed to using linear combinations of vectors (Theorem VFSLS [118]).
System Solution Set Result
solve([2*x+y==5, 3*x+2*y==15], x, y) Unique [[x == -5, y == 15]]
solve([2*x+y==5, 6*x+3*y==15], x, y) Innite [[x == (5 - r1)/2, y == r1]]
solve([2*x+y==5, 6*x+3*y==10], x, y) Empty ValueError: Unable to solve...
Notice how the output contains equations written a format that might be suitable as input for further use
within SAGE .
Version 2.30
Computation Note SAGE.VLC.SAGE Vector Linear Combinations 757
Computation Note VLC.SAGE
Vector Linear Combinations
Contributed by Robert Beezer
Vectors in SAGE are constructed from lists, and are displayed horizontally. For example, the vector
v=2
6641
2
3
43
775
would be entered and named via the command
v=vector (QQ;[1;2;3;4])
See the notes about rings (Computation R.SAGE [752]) and matrix entry (Computation ME.SAGE [753])
for reminders about specifying the relevant ring.
Vector addition and scalar multiplication are then very natural. If uand vare two vectors of the
same size, then
2u+ ( 3)v
will compute the correct vector. The result can be assigned to a variable (which will then contain a vector),
or be printed. If printed, it will be written horizontally with parentheses for grouping. If uand vhave
dierent sizes, then SAGE will complain about \unsupported operand(s)."
Computation Note MI.SAGE
Matrix Inverse
Contributed by Steve Caneld
Ifais a matrix dened in SAGE , then a.inverse() will return the inverse of a, should it exist. In the
case where adoes not have an inverse SAGE will tell you the matrix must be nonsingular (see Theorem
NI [261]).
Computation Note TM.SAGE
Transpose of a Matrix
Suppose ais the name of a matrix stored in SAGE . Then a.transpose() will return the transpose of
a.
Computation Note E.SAGE
Eigenspaces
Contributed by Steve Caneld
SAGE can compute eigenspaces and eigenvalues for you. If you have a matrix named aand you type
a:eigenspaces ()
Version 2.30
758 Section SAGE SAGE: Open Source Mathematics Software
you will get a listing of the eigenvalues and the eigenspace for each. Let's do an example. Your output
may be formatted slightly dierent from what we have here.
m=matrix (QQ;[[ 13; 8; 4];[12;7;4];[24;16;7]])
m:eigenspaces ()
[(3;[(1;2=3;1=3)]);( 1;[(1;0;1=2);(0;1; 1=2)])]
Whew, that looks like a mess. At the top level, eigenspaces() returns a dictionary whose keys are the
eigenvalues. So in this case we have eigenvalues 3 and -1. Each eigenvalue has an array after it that forms
the basis of the eigenspace. In our example, there is 1 vector for = 3 and 2 vectors for = 1. Finally,
the vectors SAGE spits out may not be the nicest ones to work with. In particular, we might want to scale
the vectors to get rid of fractions.
Version 2.30
Appendix P
Preliminaries
This appendix contains important ideas about complex numbers, sets, and the logic and techniques of
forming proofs. It is not meant to be read straight through, but you should head here when you need to
review these ideas.
We choose to expand the set of scalars from the real numbers, R, to the set of complex numbers, C. So
basic operations with complex numbers (like addition and division) will be necessary. This can be safely
postponed until your arrival in Section O [191], and a refresher before Chapter E [453] would be a good
idea as well.
Sets are extremely important in all of mathematics, but maybe you have not had much exposure to
the basic operations. Check out Section SET [761]. The text will send you here frequently as well. Visit
often.
This book is as much about doing mathematics as it is about linear algebra. The \Proof Techniques"
are vignettes about logic, types of theorems, structure of proofs, or just plain old-fashioned advice about
how to doadvanced mathematics. The text will frequently point to one of these techniques in advance
of their rst use, and for specic instructions there will be additional references. If you nd constructing
proofs dicult (we all did once), then head back here and browse through the advice for second or third
readings.
Section CNO
Complex Number Operations
In this section we review of the basics of working with complex numbers.
Subsection CNA
Arithmetic with complex numbers
A complex number is a linear combination of 1 and i=p 1, typically written in the form a+bi.
Complex numbers can be added, subtracted, multiplied and divided, just like we are used to doing with
real numbers, including the restriction on division by zero. We will not dene these operations carefully,
but instead illustrate with examples.
Example ACN
Arithmetic of complex numbers
(2 + 5i) + (6 4i) = (2 + 6) + (5 + ( 4))i= 8 +i
759
760 Section CNO Complex Number Operations
(2 + 5i) (6 4i) = (2 6) + (5 ( 4))i= 4 + 9i
(2 + 5i)(6 4i) = (2)(6) + (5 i)(6) + (2)( 4i) + (5i)( 4i) = 12 + 30 i 8i 20i2
= 12 + 22i 20( 1) = 32 + 22 i
Division takes just a bit more care. We multiply the denominator by a complex number chosen to produce
a real number and then we can produce a complex number as a result.
2 + 5i
6 4i=2 + 5i
6 4i6 + 4i
6 + 4i= 8 + 38i
52= 8
52+38
52i= 2
13+19
26i
In this example, we used 6 + 4 ito convert the denominator in the fraction to a real number. This
number is known as the conjugate, which we dene in the next section. We will often exploit the basic
properties of complex number addition, subtraction, multiplication and division, so we will carefully dene
the two basic operations, together with a denition of equality, and then collect nine basic properties in a
theorem.
Denition CNE
Complex Number Equality
The complex numbers =a+biand=c+diareequal , denoted=, ifa=candb=d.
(This denition contains Notation CNE.) 4
Denition CNA
Complex Number Addition
Thesum of the complex numbers =a+biand=c+di, denoted+, is (a+c) + (b+d)i.
(This denition contains Notation CNA.) 4
Denition CNM
Complex Number Multiplication
Theproduct of the complex numbers =a+biand=c+di, denoted, is (ac bd) + (ad+bc)i.
(This denition contains Notation CNM.) 4
Theorem PCNA
Properties of Complex Number Arithmetic
The operations of addition and multiplication of complex numbers have the following properties.
ACCN Additive Closure, Complex Numbers
If;2C, then+2C.
MCCN Multiplicative Closure, Complex Numbers
If;2C, then2C.
CACN Commutativity of Addition, Complex Numbers
For any; 2C,+=+.
CMCN Commutativity of Multiplication, Complex Numbers
For any; 2C,=.
AACN Additive Associativity, Complex Numbers
For any; ;
2C,+ (+
) = (+) +
.
MACN Multiplicative Associativity, Complex Numbers
For any; ;
2C,(
) = ()
.
Version 2.30
Subsection CNO.CCN Conjugates of Complex Numbers 761
DCN Distributivity, Complex Numbers
For any; ;
2C,(+
) =+
.
ZCN Zero, Complex Numbers
There is a complex number 0 = 0 + 0 iso that for any 2C, 0 +=.
OCN One, Complex Numbers
There is a complex number 1 = 1 + 0 iso that for any 2C, 1=.
AICN Additive Inverse, Complex Numbers
For every2Cthere exists 2Cso that+ ( ) = 0.
MICN Multiplicative Inverse, Complex Numbers
For every2C,6= 0 there exists1
2Cso that 1
= 1.
Proof We could derive each of these properties of complex numbers with a proof that builds on the
identical properties of the real numbers. The only proof that might be at all interesting would be to show
Property MICN [759] since we would need to trot out a conjugate. For this property, and especially for
the others, we might be tempted to construct proofs of the identical properties for the reals. This would
take us way too far aeld, so we will draw a line in the sand right here and just agree that these nine
fundamental behaviors are true. OK?
Mostly we have stated these nine properties carefully so that we can make reference to them later in
other proofs. So we will be linking back here often.
Subsection CCN
Conjugates of Complex Numbers
Denition CCN
Conjugate of a Complex Number
Theconjugate of the complex number =a+bi2Cis the complex number =a bi.
(This denition contains Notation CCN.) 4
Example CSCN
Conjugate of some complex numbers
2 + 3i= 2 3i 5 4i= 5 + 4i 3 + 0i= 3 + 0i 0 + 0i= 0 + 0i
Notice how the conjugate of a real number leaves the number unchanged. The conjugate enjoys some
basic properties that are useful when we work with linear expressions involving addition and multiplication.
Theorem CCRA
Complex Conjugation Respects Addition
Suppose that andare complex numbers. Then +=+.
Proof Let=a+biand=r+si. Then
+=(a+r) + (b+s)i= (a+r) (b+s)i= (a bi) + (r si) =+
Version 2.30
762 Section CNO Complex Number Operations
Theorem CCRM
Complex Conjugation Respects Multiplication
Suppose that andare complex numbers. Then =.
Proof Let=a+biand=r+si. Then
=(ar bs) + (as+br)i= (ar bs) (as+br)i
= (ar ( b)( s)) + (a( s) + ( b)r)i= (a bi)(r si) =
Theorem CCT
Complex Conjugation Twice
Suppose that is a complex number. Then =.
Proof Let=a+bi. Then
=a bi=a ( bi) =a+bi=
Subsection MCN
Modulus of a Complex Number
We dene one more operation with complex numbers that may be new to you.
Denition MCN
Modulus of a Complex Number
Themodulus of the complex number =a+bi2C, is the nonnegative real number
jj=p
=p
a2+b2:
4
Example MSCN
Modulus of some complex numbers
j2 + 3ij=p
13j5 4ij=p
41j 3 + 0ij= 3j0 + 0ij= 0
The modulus can be interpreted as a version of the absolute value for complex numbers, as is suggested
by the notation employed. You can see this in how j 3j=j 3 + 0ij= 3. Notice too how the modulus of
the complex zero, 0 + 0 i, has value 0.
Version 2.30
Section SET Sets 763
Section SET
Sets
Denition SET
Set
Asetis an unordered collection of objects. If Sis a set and xis an object that is in the set S, we write
x2S. Ifxis not inS, then we write x62S. We refer to the objects in a set as its elements .
(This denition contains Notation SETM.) 4
Hard to get much more basic than that. Notice that the objects in a set can be anything , and there is
no notion of order among the elements of the set. A set can be nite as well as innite. A set can contain
other sets as its objects. At a primitive level, a set is just a way to break up some class of objects into two
groupings: those objects in the set, and those objects not in the set.
Example SETM
Set membership
From the set of all possible symbols, construct the following set of three symbols,
S=f;;Fg
Then the statement 2Sis true, while the statement N2Sis false. However, then the statement N62S
is true.
A portion of a set is known as a subset. Notice how the following denition uses an implication (if
whenever. . . then. . . ). Note too how the denition of a subset relies on the denition of a set through the
idea of set membership.
Denition SSET
Subset
IfSandTare two sets, then Sis a subset of T, writtenSTif whenever x2Sthenx2T.
(This denition contains Notation SSET.) 4
If we want to disallow the possibility that Sis the same as T, we use the notation STand we say
thatSis aproper subset ofT. We'll do an example, but rst we'll dene a special set.
Denition ES
Empty Set
The empty set is the set with no elements. Its is denoted by ;.
(This denition contains Notation ES.) 4
Example SSET
Subset
IfS=f;;Fg,T=fF;g,R=fN;Fg, then
TS R 6T ;S
TS S S S 6S
What does it mean for two sets to be equal? They must be the same. Well, that explanation is not
really too helpful, is it? How about: If ABandBA, thenAequalsB. This gives us something to
work with, if Ais a subset of B, and vice versa , then they must really be the same set. We will now make
Version 2.30
764 Section SET Sets
the symbol \=" do double-duty and extend its use to statements like A=B, whereAandBare sets.
Here's the denition, which we will reference often.
Denition SE
Set Equality
Two sets,SandT, are equal, if STandTS. In this case, we write S=T.
(This denition contains Notation SE.) 4
Sets are typically written inside of braces, as fg, as we have seen above. However, when sets have
more than a few elements, a description will typically have two components. The rst is a description of
the general type of objects contained in a set, while the second is some sort of restriction on the properties
the objects have. Every object in the set must be of the type described in the rst part and it must satisfy
the restrictions in the second part. Conversely, any object of the proper type for the rst part, that also
meets the conditions of the second part, will be in the set. These two parts are set o from each other
somehow, often with a vertical bar ( j) or a colon (:).
I like to think of sets as clubs. The rst part is some description of the type of people who might
belong to the club, the basic objects. For example, a bicycle club would describe its members as being
people who like to ride bicycles. The second part is like a membership committee, it restricts the people
who are allowed in the club. Continuing with our bicycle club analogy, we might decide to limit ourselves
to \serious" riders and only have members who can document having ridden 100 kilometers or more in a
single day at least one time.
The restrictions on membership can migrate around some between the rst and second part, and there
may be several ways to describe the same set of objects. Here's a more mathematical example, employing
the set of all integers, Z, to describe the set of even integers.
E=fx2Zjxis an even number g
=fx2Zj2 dividesxevenlyg
=f2kjk2Zg
Notice how this set tells us that its objects are integer numbers (not, say, matrices or functions, for example)
and just those that are even. So we can write that 10 2E, while 1762Eonce we check the membership
criteria. We also recognize the question
1 3 5
2 0 3
2E?
as being simply ridiculous.
Subsection SC
Set Cardinality
On occasion, we will be interested in the number of elements in a nite set. Here's the denition and the
associated notation.
Denition C
Cardinality
SupposeSis a nite set. Then the number of elements in Sis called the cardinality orsize ofS, and is
denotedjSj.
(This denition contains Notation C.) 4
Example CS
Cardinality and Size
IfS=f;F;g, thenjSj= 3.
Version 2.30
Subsection SET.SO Set Operations 765
Subsection SO
Set Operations
In this subsection we dene and illustrate the three most common basic ways to manipulate sets to create
other sets. Since much of linear algebra is about sets, we will use these often.
Denition SU
Set Union
SupposeSandTare sets. Then the union ofSandT, denotedS[T, is the set whose elements are those
that are elements of Sor ofT, or both. More formally,
x2S[Tif and only if x2Sorx2T
(This denition contains Notation SU.) 4
Notice that the use of the word \or" in this denition is meant to be non-exclusive. That is, it allows
forxto be an element of both SandTand still qualify for membership in S[T.
Example SU
Set union
IfS=f;F;gandT=f;F;NgthenS[T=f;F;;Ng.
Denition SI
Set Intersection
SupposeSandTare sets. Then the intersection ofSandT, denotedS\T, is the set whose elements
are only those that are elements of Sand ofT. More formally,
x2S\Tif and only if x2Sandx2T
(This denition contains Notation SI.) 4
Example SI
Set intersection
IfS=f;F;gandT=f;F;NgthenS\T=f;Fg.
The union and intersection of sets are operations that begin with two sets and produce a third, new,
set. Our nal operation is the set complement, which we usually think of as an operation that takes a
single set and creates a second, new, set. However, if you study the denition carefully, you will see that
it needs to be computed relative to some \universal" set.
Denition SC
Set Complement
SupposeSis a set that is a subset of a universal set U. Then the complement ofS, denotedS, is the
set whose elements are those that are elements of Uand not elements of S. More formally,
x2Sif and only if x2Uandx62S
(This denition contains Notation SC.) 4
Notice that there is nothing at all special about the universal set. This is simply a term that suggests
thatUcontains all of the possible objects we are considering. Often this set will be clear from the context,
and we won't think much about it, nor reference it in our notation. In other cases (rarely in our work in
Version 2.30
766 Section SET Sets
this course) the exact nature of the universal set must be made explicit, and reference to it will possibly
be carried through in our choice of notation.
Example SC
Set complement
IfU=f;F;;NgandS=f;F;gthenS=fNg.
There are many more natural operations that can be performed on sets, such as an exclusive-or and the
symmetric dierence. Many of these can be dened in terms of the union, intersection and complement.
We will not have much need of them in this course, and so we will not give precise descriptions here in this
preliminary section.
There is also an interesting variety of basic results that describe the interplay of these operations with
each other. We mention just two as an example, these are known as DeMorgan's Laws.
(S[T) =S\T
(S\T) =S[T
Besides having an appealing symmetry, we mention these two facts, since constructing the proofs of each
is a useful exercise that will require a solid understanding of all but one of the denitions presented in this
section. Give it a try.
Version 2.30
Section PT Proof Techniques 767
Section PT
Proof Techniques
In this section we collect many short essays designed to help you understand how to read, understand and
construct proofs. Some are very factual, while others consist of advice. They appear in the order that
they are rst needed (or advisable) in the text, and are meant to be self-contained. So you should not
think of reading through this section in one sitting as you begin this course. But be sure to head back here
for a rst reading whenever the text suggests it. Also think about returning to browse at various points
during the course, and especially as you struggle with becoming an accomplished mathematician who is
comfortable with the dicult process of designing new proofs.
Proof Technique D
Denitions
A denition is a made-up term, used as a kind of shortcut for some typically more complicated idea. For
example, we say a whole number is even as a shortcut for saying that when we divide the number by two
we get a remainder of zero. With a precise denition, we can answer certain questions unambiguously. For
example, did you ever wonder if zero was an even number? Now the answer should be clear since we have
a precise denition of what we mean by the term even.
A single term might have several possible denitions. For example, we could say that the whole number
nis even if there is another whole number ksuch thatn= 2k. We say this is an equivalent denition since
it categorizes even numbers the same way our rst denition does.
Denitions are like two-way streets | we can use a denition to replace something rather complicated
by its denition (if it ts) andwe can replace a denition by its more complicated description. A denition
is usually written as some form of an implication, such as \If something-nice-happens, then blatzo ."
However, this also means that \If blatzo, then something-nice-happens," even though this may not be
formally stated. This is what we mean when we say a denition is a two-way street | it is really two
implications, going in opposite \directions."
Anybody (including you) can make up a denition, so long as it is unambiguous, but the real test of a
denition's utility is whether or not it is useful for describing interesting or frequent situations.
We will talk about theorems later (and especially equivalences). For now, be sure not to confuse the
notion of a denition with that of a theorem.
In this book, we will display every new denition carefully set-o from the text, and the term being
dened will be written thus: denition . Additionally, there is a full list of all the denitions, in order of
their appearance located at the front of the book (Denitions [xi]). Finally, the acronym for each denition
can be found in the index (Index [ ??]). Denitions are critical to doing mathematics and proving theorems,
so we've given you lots of ways to locate a denition should you forget its. . . uh, well, . . . denition.
Can you formulate a precise denition for what it means for a number to be odd? (Don't just say
it is the opposite of even. Act as if you don't have a denition for even yet.) Can you formulate your
denition a second, equivalent, way? Can you employ your denition to test an odd and an even number
for \odd-ness"?
Version 2.30
768 Section PT Proof Techniques
Proof Technique T
Theorems
Higher mathematics is about understanding theorems. Reading them, understanding them, applying them,
proving them. Every theorem is a shortcut | we prove something in general, and then whenever we nd
a specic instance covered by the theorem we can immediately say that we know something else about
the situation by applying the theorem. In many cases, this new information can be gained with much less
eort than if we did not know the theorem.
The rst step in understanding a theorem is to realize that the statement of every theorem can be rewrit-
ten using statements of the form \If something-happens, then something-else-happens." The \something-
happens" part is the hypothesis and the \something-else-happens" is the conclusion . To understand
a theorem, it helps to rewrite its statement using this construction. To apply a theorem, we verify that
\something-happens" in a particular instance and immediately conclude that \something-else-happens."
To prove a theorem, we must argue based on the assumption that the hypothesis is true, and arrive through
the process of logic that the conclusion must then also be true.
Proof Technique L
Language
Like any science, the language of math must be understood before further study can continue.
Erin Wilson, Student
September, 2004
Mathematics is a language. It is a way to express complicated ideas clearly, precisely, and unambiguously.
Because of this, it can be dicult to read. Read slowly, and have pencil and paper at hand. It will usually
be necessary to read something several times. While reading can be dicult, it is even harder to speak
mathematics, and so that is the topic of this technique.
\Natural" language, in the present case English, is fraught with ambiguity. Consider the possible
meanings of the sentence: The sh is ready to eat. One sh, or two sh? Are the sh hungry, or will the
sh be eaten? (See Exercise SSLE.M10 [21], Exercise SSLE.M11 [21], Exercise SSLE.M12 [22], Exercise
SSLE.M13 [22].) In your daily interactions with others, give some thought to how many mis-understandings
arise from the ambiguity of pronouns, modiers and objects.
I am going to suggest a simple modication to the way you use language that will make it much, much
easier to become procient at speaking mathematics and eventually it will become second nature. Think
of it as a training aid or practice drill you might use when learning to become skilled at a sport.
First, eliminate pronouns from your vocabulary when discussing linear algebra, in class or with your
colleagues. Do not use: it, that, those, their or similar sources of confusion. This is the single easiest step
you can take to make your oral expression of mathematics clearer to others, and in turn, it will greatly
help your own understanding.
Now rid yourself of the word \thing" (or variants like \something"). When you are tempted to use this
word realize that there is some object you want to discuss, and we likely have a denition for that object
(see the discussion at Technique D [765]). Always \think about your objects" and many aspects of the
study of mathematics will get easier. Ask yourself: \Am I working with a set, a number, a function, an
operation, a dierential equation, or what?" Knowing what an object iswill allow you to narrow down the
procedures you may apply to it. If you have studied an object-oriented computer programming language,
then you will already have experience identifying objects and thinking carefully about what procedures are
allowed to be applied to them.
Version 2.30
Proof Technique PT.GS Getting Started 769
Third, eliminate the verb \works" (as in \the equation works") from your vocabulary. This term is
used as a substitute when we are not sure just what we are trying to accomplish. Usually we are trying to
say that some object fullls some condition. The condition might even have a denition associated with
it, making it even easier to describe.
Last, speak slooooowly and thoughtfully as you try to get by without all these lazy words. It is hard
at rst, but you will get better with practice. Especially in class, when the pressure is on and all eyes are
on you, don't succumb to the temptation to use these weak words. Slow down, we'd all rather wait for a
slow, well-formed question or answer than a fast, sloppy, incomprehensible one.
You will nd the improvement in your ability to speak clearly about complicated ideas will greatly
improve your ability to think clearly about complicated ideas. And I believe that you cannot think clearly
about complicated ideas if you cannot formulate questions or answers clearly in the correct language. This
is as applicable to the study of law, economics or philosophy as it is to the study of science or mathematics.
In this spirit, Dupont Hubert has contributed the following quotation, which is widely used in French
mathematics courses (and which might be construed as the contrapositive of Technique CP [769])
Ce que l'on concoit bien s'enonce clairement,
Et les mots pour le dire arrivent aisement.
| Nicolas Boileau, L'art po etique, Chant I, 1674
which translates as
Whatever is well conceived is clearly said,
And the words to say it
ow with ease.
So when you come to class, check your pronouns at the door, along with other weak words. And
when studying with friends, you might make a game of catching one another using pronouns, \thing," or
\works." I know I'll be calling you on it!
Proof Technique GS
Getting Started
\I don't know how to get started!" is often the lament of the novice proof-builder. Here are a few pieces
of advice.
1. As mentioned in Technique T [766], rewrite the statement of the theorem in an \if-then" form. This
will simplify identifying the hypothesis and conclusion, which are referenced in the next few items.
2. Ask yourself what kind of statement you are trying to prove. This is always part of your conclusion.
Are you being asked to conclude that two numbers are equal, that a function is dierentiable or a set
is a subset of another? You cannot bring other techniques to bear if you do not know what type of
conclusion you have.
3. Write down reformulations of your hypotheses. Interpret and translate each denition properly.
4. Write your hypothesis at the top of a sheet of paper and your conclusion at the bottom. See if you
can formulate a statement that precedes the conclusion and also implies it. Work down from your
hypothesis, and up from your conclusion, and see if you can meet in the middle. When you are
nished, rewrite the proof nicely, from hypothesis to conclusion, with veriable implications giving
each subsequent statement.
Version 2.30
770 Section PT Proof Techniques
5. As you work through your proof, think about what kinds of objects your symbols represent. For
example, suppose Ais a set and f(x) is a real-valued function. Then the expression A+fmight
make no sense if we have not dened what it means to \add" a set to a function, so we can stop at
that point and adjust accordingly. On the other hand we might understand 2 fto be the function
whose rule is described by (2 f)(x) = 2f(x). \Think about your objects" means to always verify that
your objects and operations are compatible.
Proof Technique C
Constructive Proofs
Conclusions of proofs come in a variety of types. Often a theorem will simply assert that something exists.
The best way, but not the only way, to show something exists is to actually build it. Such a proof is
called constructive . The thing to realize about constructive proofs is that the proof itself will contain a
procedure that might be used computationally to construct the desired object. If the procedure is not too
cumbersome, then the proof itself is as useful as the statement of the theorem.
Proof Technique E
Equivalences
When a theorem uses the phrase \if and only if" (or the abbreviation \i") it is a shorthand way of saying
that two if-then statements are true. So if a theorem says \P if and only if Q," then it is true that \if P,
then Q" while it is also true that \if Q, then P." For example, it may be a theorem that \I wear bright
yellow knee-high plastic boots if and only if it is raining." This means that I never forget to wear my
super-duper yellow boots when it is raining andI wouldn't be seen in such silly boots unless it was raining.
You never have one without the other. I've got my boots on and it is raining orI don't have my boots on
and it is dry.
The upshot for proving such theorems is that it is like a 2-for-1 sale, we get to do twoproofs. Assume
Pand conclude Q, then start over and assume Qand conclude P. For this reason, \if and only if" is
sometimes abbreviated by () , while proofs indicate which of the two implications is being proved by
prefacing each with )or(. A carefully written proof will remind the reader which statement is being
used as the hypothesis, a quicker version will let the reader deduce it from the direction of the arrow.
Tradition dictates we do the \easy" half rst, but that's hard for a student to know until you've nished
doing both halves! Oh well, if you rewrite your proofs (a good habit), you can then choose to put the easy
half rst.
Theorems of this type are called \equivalences" or \characterizations," and they are some of the most
pleasing results in mathematics. They say that two objects, or two situations, are really the same. You
don't have one without the other, like rain and my yellow boots. The more dierent PandQseem to be,
the more pleasing it is to discover they are really equivalent. And if Pdescribes a very mysterious solution
or involves a tough computation, while Qis transparent or involves easy computations, then we've found
a great shortcut for better understanding or faster computation. Remember that every theorem really is a
shortcut in some form. You will also discover that if proving P)Qis very easy, then proving Q)Pis
likely to be proportionately harder. Sometimes the two halves are about equally hard. And in rare cases,
you can string together a whole sequence of other equivalences to form the one you're after and you don't
even need to do two halves. In this case, the argument of one half is just the argument of the other half,
but in reverse.
One last thing about equivalences. If you see a statement of a theorem that says two things are
\equivalent," translate it rst into an \if and only if" statement.
Version 2.30
Proof Technique PT.N Negation 771
Proof Technique N
Negation
When we construct the contrapositive of a theorem (Technique CP [769]), we need to negate the two
statements in the implication. And when we construct a proof by contradiction (Technique CD [770]),
we need to negate the conclusion of the theorem. One way to construct a converse (Technique CV [769])
is to simultaneously negate the hypothesis and conclusion of an implication (but remember that this is
not guaranteed to be a true statement). So we often have the need to negate statements, and in some
situations it can be tricky.
If a statement says that a set is empty, then its negation is the statement that the set is nonempty.
That's straightforward. Suppose a statement says \something-happens" for all i, or everyi, or anyi. Then
the negation is that \something-doesn't-happen" for at least one value of i. If a statement says that there
exists at least one \thing," then the negation is the statement that there is no \thing." If a statement says
that a \thing" is unique, then the negation is that there is zero, or more than one, of the \thing."
We are not covering all of the possibilities, but we wish to make the point that logical qualiers like
\there exists" or \for every" must be handled with care when negating statements. Studying the proofs
which employ contradiction (as listed in Technique CD [770]) is a good rst step towards understanding
the range of possibilities.
Proof Technique CP
Contrapositives
Thecontrapositive of an implication P)Qis the implication not( Q))not(P), where \not" means the
logical negation, or opposite. An implication is true if and only if its contrapositive is true. In symbols,
(P)Q)() (not(Q))not(P)) is a theorem. Such statements about logic, that are always true, are
known as tautologies .
For example, it is a theorem that \if a vehicle is a re truck, then it has big tires and has a siren."
(Yes, I'm sure you can conjure up a counterexample, but play along with me anyway.) The contrapositive
is \if a vehicle does not have big tires or does not have a siren, then it is not a re truck." Notice how the
\and" became an \or" when we negated the conclusion of the original theorem.
It will frequently happen that it is easier to construct a proof of the contrapositive than of the original
implication. If you are having diculty formulating a proof of some implication, see if the contrapositive
is easier for you. The trick is to construct the negation of complicated statements accurately. More on
that later.
Proof Technique CV
Converses
Theconverse of the implication P)Qis the implication Q)P. There is no guarantee that the truth
of these two statements are related. In particular, if an implication has been proven to be a theorem, then
do not try to use its converse too, as if it were a theorem. Sometimes the converse is true (and we have an
equivalence, see Technique E [768]). But more likely the converse is false, especially if it wasn't included
in the statement of the original theorem.
For example, we have the theorem, \if a vehicle is a re truck, then it is has big tires and has a siren."
The converse is false. The statement that \if a vehicle has big tires and a siren, then it is a re truck" is
false. A police vehicle for use on a sandy public beach would have big tires and a siren, yet is not equipped
to ght res.
Version 2.30
772 Section PT Proof Techniques
We bring this up now, because Theorem CSRN [59] has a tempting converse. Does this theorem say
that ifr < n , then the system is consistent? Denitely not, as Archetype E [799] has r= 3<4 =n,
yet is inconsistent. This example is then said to be a counterexample to the converse. Whenever you
think a theorem that is an implication might actually be an equivalence, it is good to hunt around for
a counterexample that shows the converse to be false (the archetypes, Appendix A [777], can be a good
hunting ground).
Proof Technique CD
Contradiction
Another proof technique is known as \proof by contradiction" and it can be a powerful (and satisfying)
approach. Simply put, suppose you wish to prove the implication, \If A, thenB." As usual, we assume
thatAis true, but we also make the additional assumption that Bis false. If our original implication
is true, then these twin assumptions should lead us to a logical inconsistency. In practice we assume the
negation of Bto be true (see Technique N [769]). So we argue from the assumptions Aand not(B) looking
for some obviously false conclusion such as 1 = 6, or a set is simultaneously empty and nonempty, or a
matrix is both nonsingular and singular.
You should be careful about formulating proofs that look like proofs by contradiction, but really aren't.
This happens when you assume Aand not(B) and proceed to give a \normal" and direct proof that B
is true by only using the assumption that Ais true. Your last step is to then claim that Bis true and
you then appeal to the assumption that not( B) is true, thus getting the desired contradiction. Instead,
you could have avoided the overhead of a proof by contradiction and just run with the direct proof. This
stylistic
aw is known, quite graphically, as \setting up the strawman to knock him down."
Here is a simple example of a proof by contradiction. There are direct proofs that are just about as
easy, but this will demonstrate the point, while narrowly avoiding knocking down the straw man.
Theorem : Ifaandbare odd integers, then their product, ab, is odd.
Proof : To begin a proof by contradiction, assume the hypothesis, that aandbare odd. Also assume
the negation of the conclusion, in this case, that abis even. Then there are integers, j,k,`so that
a= 2j+ 1,b= 2k+ 1,ab= 2`. Then
0 =ab ab
= (2j+ 1)(2k+ 1) (2`)
= 4jk+ 2j+ 2k 2`+ 1
= 2 (2jk+j+k `) + 1
Notice how we used both our hypothesis and the negation of the conclusion in the second line. Now divide
the integer on each end of this string of equalities by 2. On the left we get a remainder of 0, while on
the right we see that the remainder will be 1. Both remainders cannot be correct, so this is our desired
contradiction. Thus, the conclusion (that abis odd) is true.
Again, we do not oer this example as the bestproof of this fact about even and odd numbers, but rather
it is a simple illustration of a proof by contradiction. You can nd examples of proofs by contradiction in
Theorem RREFU [35], Theorem NMUS [86], Theorem NPNT [259], Theorem TTMI [246], Theorem GSP
[199], Theorem ELIS [407], Theorem EDYES [410], Theorem EMHE [457], Theorem EDELI [479], and
Theorem DMFE [499], in addition to several examples and solutions to exercises.
Version 2.30
Proof Technique PT.U Uniqueness 773
Proof Technique U
Uniqueness
A theorem will sometimes claim that some object, having some desirable property, is unique. In other
words, there should be only one such object. To prove this, a standard technique is to assume there
are two such objects and proceed to analyze the consequences. The end result may be a contradiction
(Technique CD [770]), or the conclusion that the two allegedly dierent objects really are equal.
Proof Technique ME
Multiple Equivalences
A very specialized form of a theorem begins with the statement \The following are equivalent. . . ," which
is then followed by a list of statements. Informally, this lead-in sometimes gets abbreviated by \TFAE."
This formulation means that any two of the statements on the list can be connected with an \if and only
if" to form a theorem. So if the list has nstatements then, there aren(n 1)
2possible equivalences that can
be constructed (and are claimed to be true).
Suppose a theorem of this form has statements denoted as A,B,C,. . .Z. To prove the entire theorem,
we can prove A)B,B)C,C)D,. . . ,Y)Zand nally, Z)A. This circular chain of nequivalences
would allow us, logically, if not practically, to form any one of then(n 1)
2possible equivalences by chasing
the equivalences around the circle as far as required.
Proof Technique PI
Proving Identities
Many theorems have conclusions that say two objects are equal. Perhaps one object is hard to compute or
understand, while the other is easy to compute or understand. This would make for a pleasing theorem.
Whether the result is pleasing or not, we take the same approach to formulate a proof. Sometimes we need
to employ specialized notions of equality, such as Denition SE [762] or Denition CVE [98], but in other
cases we can string together a list of equalities.
The wrong way to prove an identity is to begin by writing it down and then beating on it until it
reduces to an obvious identity. The rst
aw is that you would be writing down the statement you wish
to prove, as if you already believed it to be true. But more dangerous is the possibility that some of your
maneuvers are not reversible. Here's an example. Let's prove that 3 = 3.
3 = 3 (This is a bad start)
32= ( 3)2Square both sides
9 = 9
0 = 0 Subtract 9 from both sides
So because 0 = 0 is a true statement, does it follow that 3 = 3 is a true statement? Nope. Of course,
we didn't really expect a legitimate proof of 3 = 3, but this attempt should illustrate the dangers of this
(incorrect) approach.
What you have just seen in the proof of Theorem VSPCV [100], and what you will see consistently
throughout this text, is proofs of the following form. To prove that A=Dwe write
A=B Theorem, Denition or Hypothesis justifying A=B
=C Theorem, Denition or Hypothesis justifying B=C
Version 2.30
774 Section PT Proof Techniques
=D Theorem, Denition or Hypothesis justifying C=D
In your scratch work exploring possible approaches to proving a theorem you may massage a variety of
expressions, sometimes making connections to various bits and pieces, while some parts get abandoned.
Once you see a line of attack, rewrite your proof carefully mimicking this style.
Proof Technique DC
Decompositions
Much of your mathematical upbringing, especially once you began a study of algebra, revolved around
simplifying expressions | combining like terms, obtaining common denominators so as to add fractions,
factoring in order to solve polynomial equations. However, as often as not, we will do the opposite.
Many theorems and techniques will revolve around taking some object and \decomposing" it into some
combination of other objects, ostensibly in a more complicated fashion. When we say something can \be
written as" something else, we mean that the one object can be decomposed into some combination of
other objects. This may seem unnatural at rst, but results of this type will give us insight into the
structure of the original object by exposing its inner workings. An appropriate analogy might be stripping
the wallboards away from the interior of a building to expose the structural members supporting the whole
building.
Perhaps you have studied integral calculus, or a pre-calculus course, where you learned about partial
fractions. This is a technique where a fraction of two polynomials is decomposed (written as, expressed
as) a sum of simpler fractions. The purpose in calculus is to make nding an antiderivative simpler. For
example, you can verify the truth of the expression
12x5+ 2x4 20x3+ 66x2 294x+ 308
x6+x5 3x4+ 21x3 52x2+ 20x 48=5x+ 2
x2 x+ 6+3x 7
x2+ 1+3
x+ 4+1
x 2
In an early course in algebra, you might be expected to combine the four terms on the right over a
common denominator to create the \simpler" expression on the left. Going the other way, the partial
fraction technique would allow you to systematically decompose the fraction of polynomials on the left into
the sum of the four (arguably) simpler fractions of polynomials on the right.
This is a major shift in thinking, so come back here often, especially when we say \can be written as",
or \can be expressed as," or \can be decomposed as."
Proof Technique I
Induction
\Induction" or \mathematical induction" is a framework for proving statements that are indexed by in-
tegers. In other words, suppose you have a statement to prove that is really multiple statements, one for
n= 1, another for n= 2, a third for n= 3, and so on. If there is enough similarity between the statements,
then you can use a script (the framework) to prove them all at once.
For example, consider the theorem
Theorem 1 + 2 + 3 ++n=n(n+ 1)
2forn1.
This is shorthand for the many statements 1 =1(1+1)
2, 1 + 2 =2(2+1)
2, 1 + 2 + 3 =3(3+1)
2, 1 + 2 + 3 + 4 =
4(4+1)
2, and so on. Forever. You can do the calculations in each of these statements and verify that all four
are true. We might not be surprised to learn that the fth statement is true as well (go ahead and check).
However, do we think the theorem is true for n= 872? Or n= 1;234;529?
Version 2.30
Proof Technique PT.I Induction 775
To see that these questions are not so ridiculous, consider the following example from Rotman's Journey
into Mathematics . The statement \ n2 n+ 41 is prime" is true for integers 1 n40 (check a few).
However, when we check n= 41 we nd 412 41 + 41 = 412, which is not prime.
So how do we prove innitely many statements all at once? More formally, lets denote our statements
asP(n). Then, if we can prove the two assertions
1.P(1) is true.
2. IfP(k) is true, then P(k+ 1) is true.
then it follows that P(n) is true for all n1. To understand this, I liken the process to climbing an
innitely long ladder with equally spaced rungs. Confronted with such a ladder, suppose I tell you that
you are able to step up onto the rst rung, and if you are on any particular rung, then you are capable of
stepping up to the next rung. It follows that you can climb the ladder as far up as you wish. The rst
formal assertion above is akin to stepping onto the rst rung, and the second formal assertion is akin to
assuming that if you are on any one rung then you can always reach the next rung.
In practice, establishing that P(1) is true is called the \base case" and in most cases is straightforward.
Establishing that P(k))P(k+ 1) is referred to as the \induction step," or in this book (and elsewhere)
we will typically refer to the assumption of P(k) as the \induction hypothesis." This is perhaps the most
mysterious part of a proof by induction, since it looks like you are assuming ( P(k)) what you are trying
to prove (P(n)). Sometimes it is even worse, since as you get more comfortable with induction, we often
don't bother to use a dierent letter ( k) for the index ( n) in the induction step. Notice that the second
formal assertion never says that P(k) is true, it simply says that ifP(k) were true, what might logically
follow. We can establish statements like \If I lived on the moon, then I could pole-vault over a bar 12
meters high." This may be a true statement, but it does not say we live on the moon, and indeed we may
never live there.
Enough generalities. Let's work an example and prove the theorem above about sums of integers.
Formally, our statement is P(n) : 1 + 2 + 3 ++n=n(n+ 1)
2.
Proof : Base Case. P(1) is the statement 1 =1(1+1)
2, which we see simplies to the true statement
1 = 1.
Induction Step: We will assume P(k) is true, and will try to prove P(k+ 1). Given what we want to
accomplish, it is natural to begin by examining the sum of the rst k+ 1 integers.
1 + 2 + 3 ++ (k+ 1) = (1 + 2 + 3 + +k) + (k+ 1)
=k(k+ 1)
2+ (k+ 1) Induction Hypothesis
=k2+k
2+2k+ 2
2
=k2+ 3k+ 2
2
=(k+ 1)(k+ 2)
2
=(k+ 1)((k+ 1) + 1)
2
We then recognize the two ends of this chain of equalities as P(k+ 1). So, by mathematical induction, the
theorem is true for all n.
How do you recognize when to use induction? The rst clue is a statement that is really many state-
ments, one for each integer. The second clue would be that you begin a more standard proof and you nd
yourself using words like \and so on" (as above!) or lots of ellipses (dots) to establish patterns that you are
Version 2.30
776 Section PT Proof Techniques
convinced continue on and on forever. However, there are many minor instances where induction might be
warranted but we don't bother.
Induction is important enough, and used often enough, that it appears in various variations. The base
case sometimes begins with n= 0, or perhaps an integer greater than n. Some formulate the induction
step asP(k 1))P(k). There is also a \strong form" of induction where we assume all of P(1),P(2),
P(3), . . .P(k) as a hypothesis for showing the conclusion P(k+ 1).
You can nd examples of induction in the proofs of Theorem GSP [199], Theorem DER [429], Theorem DT
[430], Theorem DIM [443], Theorem EOMP [481], Theorem DCP [484], and Theorem KPLT [691].
Proof Technique P
Practice
Here is a technique used by many practicing mathematicians when they are teaching themselves new
mathematics. As they read a textbook, monograph or research article, they attempt to prove each new
theorem themselves, before reading the proof. Often the proofs can be very dicult, so it is wise not to
spend too much time on each. Maybe limit your losses and try each proof for 10 or 15 minutes. Even if
the proof is not found, it is time well-spent. You become more familiar with the denitions involved, and
the hypothesis and conclusion of the theorem. When you do work through the proof, it might make more
sense, and you will gain added insight about just how to construct a proof.
Proof Technique LC
Lemmas and Corollaries
Theorems often go by dierent titles. Two of the most popular being \lemma" and \corollary." Before we
describe the ne distinctions, be aware that lemmas, corollaries, propositions, claims and facts are all just
theorems. And every theorem can be rephrased as an \if-then" statement, or perhaps a pair of \if-then"
statements expressed as an equivalence (Technique E [768]).
A lemma is a theorem that is not too interesting in its own right, but is important for proving other
theorems. It might be a generalization or abstraction of a key step of several dierent proofs. For this
reason you often hear the phrase \technical lemma" though some might argue that the adjective \technical"
is redundant.
A corollary is a theorem that follows very easily from another theorem. For this reason, corollaries
frequently do not have proofs. You are expected to easily and quickly see how a previous theorem implies
the corollary.
A proposition or fact is really just a codeword for a theorem. A claim might be similar, but some
authors like to use claims within a proof to organize key steps. In a similar manner, some long proofs are
organized as a series of lemmas.
In order to not confuse the novice, we have just called all our theorems theorems. It is also an
organizational convenience. With only theorems and denitions, the theoretical backbone of the course is
laid bare in the two lists of Denitions [xi] and Theorems [xiii].
Version 2.30
Proof Technique PT.LC Lemmas and Corollaries 777
Version 2.30
778 Section PT Proof Techniques
Version 2.30
Appendix A
Archetypes
WordNet (an open-source lexical database) gives the following denition of \archetype": something that
serves as a model or a basis for making copies.
Our archetypes are typical examples of systems of equations, matrices and linear transformations. They
have been designed to demonstrate the range of possibilities, allowing you to compare and contrast them.
Several are of a size and complexity that is usually not presented in a textbook, but should do a better
job of being \typical."
We have made frequent reference to many of these throughout the text, such as the frequent comparisons
between Archetype A [781] and Archetype B [786]. Some we have left for you to investigate, such as
Archetype J [820], which parallels Archetype I [816].
How should you use the archetypes? First, consult the description of each one as it is mentioned in the
text. See how other facts about the example might illuminate whatever property or construction is being
described in the example. Second, each property has a short description that usually includes references to
the relevant theorems. Perform the computations and understand the connections to the listed theorems.
Third, each property has a small checkbox in front of it. Use the archetypes like a workbook and chart
your progress by \checking-o" those properties that you understand.
The next page has a chart that summarizes some (but not all) of the properties described for each
archetype. Notice that while there are several types of objects, there are fundamental connections between
them. That some lines of the table do double-duty is meant to convey some of these connections. Consult
this table when you wish to quickly nd an example of a certain phenomenon.
779
780 Appendix A Archetypes
Version 2.30
Appendix A Archetypes 781ABCDE F GHIJK LMNOPQRSTUVW X
Type SSSSS S SSSSMM LLLLLLLLLLLL
Vars, Cols, Domain 33444 4 227955553355356434
Eqns, Rows, CoDom 33333 4 554655335555464434
Solution Set IUIIN U UNII
Rank 23322 4 223453232345254433
Nullity 10122 0 004502321010102001
Injective XXNYNYNYXYYN
Surjective NYXXNYXXYYYN
Full Rank NYYNN Y YYNNYN
Nonsingular NY Y YN
Invertible NY Y YN NY YYN
Determinant 0-2 -18 16 0 -2-3 0
Diagonalizable NY Y YY YY
Archetype Facts
S=System of Equations, M=Matrix, L=Linear Transformation
U=Unique solution, I=Innitely many solutions, N=No solutions
Y=Yes, N=No, X=Impossible, blank=Not Applicable
Version 2.30
782 Appendix A Archetypes
Version 2.30
Archetype A 783
Archetype A
Summary Linear system of three equations, three unknowns. Singular coecient matrix with dimension
1 null space. Integer eigenvalues and a degenerate eigenspace for coecient matrix.
A system of linear equations (Denition SLE [11]):
x1 x2+ 2x3= 1
2x1+x2+x3= 8
x1+x2= 5
Some solutions to the system of linear equations (not necessarily exhaustive):
x1= 2; x 2= 3; x 3= 1
x1= 3; x 2= 2; x 3= 0
Augmented matrix of the linear system of equations (Denition AM [30]):
2
41 1 2 1
2 1 1 8
1 1 0 53
5
Matrix in reduced row-echelon form, row-equivalent to augmented matrix:
2
410 1 3
01 1 2
0 0 0 03
5
Analysis of the augmented matrix (Notation RREFA [33]):
r= 2 D=f1;2g F=f3;4g
Vector form of the solution set to the system of equations (Theorem VFSLS [118]). Notice the rela-
tionship between the free variables and the set Fabove. Also, notice the pattern of 0's and 1's in the
entries of the vectors corresponding to elements of the set Ffor the larger examples.
Version 2.30
784 Archetype A
2
4x1
x2
x33
5=2
43
2
03
5+x32
4 1
1
13
5
Given a system of equations we can always build a new, related, homogeneous system (Denition HS
[71]) by converting the constant terms to zeros and retaining the coecients of the variables. Properties
of this new system will have precise relationships with various properties of the original system.
x1 x2+ 2x3= 0
2x1+x2+x3= 0
x1+x2 = 0
Some solutions to the associated homogenous system of linear equations (not necessarily exhaustive):
x1= 0; x 2= 0; x 3= 0
x1= 1; x 2= 1; x 3= 1
x1= 5; x 2= 5; x 3= 5
Form the augmented matrix of the homogenous linear system, and use row operations to convert to
reduced row-echelon form. Notice how the entries of the nal column remain zeros:
2
410 1 0
01 1 0
0 0 0 03
5
Analysis of the augmented matrix for the homogenous system (Notation RREFA [33]). Notice the
slight variation for the same analysis of the original system only when the original system was consistent:
r= 2 D=f1;2g F=f3;4g
Coecient matrix of original system of equations, and of associated homogenous system. This matrix
will be the subject of further analysis, rather than the systems of equations.
2
41 1 2
2 1 1
1 1 03
5
Matrix brought to reduced row-echelon form:
2
410 1
01 1
0 0 03
5
Version 2.30
Archetype A 785
Analysis of the row-reduced matrix (Notation RREFA [33]):
r= 2 D=f1;2g F=f3g
Matrix (coecient matrix) is nonsingular or singular? (Theorem NMRRI [84]) at the same time,
examine the size of the set Fabove.Notice that this property does not apply to matrices that are not
square.
Singular.
This is the null space of the matrix. The set of vectors used in the span construction is a linearly
independent set of column vectors that spans the null space of the matrix (Theorem SSNS [137], Theorem
BNS [160]). Solve the homogenous system with this matrix as the coecient matrix and write the solutions
in vector form (Theorem VFSLS [118]) to see these vectors arise.
*8
<
:2
4 1
1
13
59
=
;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors that are
also columns of the matrix. These columns have indices that form the set Dabove. (Theorem BCS [274])
*8
<
:2
41
2
13
5;2
4 1
1
13
59
=
;+
The column space of the matrix, as it arises from the extended echelon form of the matrix. The matrix
Lis computed as described in Denition EEF [297]. This is followed by the column space described by a
set of linearly independent vectors that span the null space of L, computed as according to Theorem FS
[299] and Theorem BNS [160]. When r=m, the matrix Lhas no rows and the column space is all of Cm.
L=
1 2 3
*8
<
:2
4 3
0
13
5;2
42
1
03
59
=
;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors. These
vectors are computed by row-reducing the transpose of the matrix into reduced row-echelon form, tossing
out the zero rows, and writing the remaining nonzero rows as column vectors. By Theorem CSRST [282]
and Theorem BRS [280], and in the style of Example CSROI [282], this yields a linearly independent set
of vectors that span the column space.
Version 2.30
786 Archetype A
*8
<
:2
41
0
1
33
5;2
40
1
2
33
59
=
;+
Row space of the matrix, expressed as a span of a set of linearly independent vectors, obtained from
the nonzero rows of the equivalent matrix in reduced row-echelon form. (Theorem BRS [280])
*8
<
:2
41
0
13
5;2
40
1
13
59
=
;+
Inverse matrix, if it exists. The inverse is not dened for matrices that are not square, and if the matrix
is square, then the matrix must be nonsingular. (Denition MI [244], Theorem NI [261])
Subspace dimensions associated with the matrix. (Denition NOM [397], Denition ROM [397]) Verify
Theorem RPNC [398]
Matrix columns: 3 Rank: 2 Nullity: 1
Determinant of the matrix, which is only dened for square matrices. The matrix is nonsingular if and
only if the determinant is nonzero (Theorem SMZD [445]). (Product of all eigenvalues?)
Determinant = 0
Eigenvalues, and bases for eigenspaces. (Denition EEM [453],Denition EM [461])
= 0 EA(0) =*8
<
:2
4 1
1
13
59
=
;+
= 2 EA(2) =*8
<
:2
41
5
33
59
=
;+
Geometric and algebraic multiplicities. (Denition GME [463]Denition AME [463])
A(0) = 1 A(0) = 2
A(2) = 1 A(2) = 1
Diagonalizable? (Denition DZM [496])
Version 2.30
Archetype A 787
No,
A(0)6=B(0), Theorem DMFE [499].
Version 2.30
788 Archetype B
Archetype B
Summary System with three equations, three unknowns. Nonsingular coecient matrix. Distinct
integer eigenvalues for coecient matrix.
A system of linear equations (Denition SLE [11]):
7x1 6x2 12x3= 33
5x1+ 5x2+ 7x3= 24
x1+ 4x3= 5
Some solutions to the system of linear equations (not necessarily exhaustive):
x1= 3; x 2= 5; x 3= 2
Augmented matrix of the linear system of equations (Denition AM [30]):
2
4 7 6 12 33
5 5 7 24
1 0 4 53
5
Matrix in reduced row-echelon form, row-equivalent to augmented matrix:
2
410 0 3
010 5
0 0 1 23
5
Analysis of the augmented matrix (Notation RREFA [33]):
r= 3 D=f1;2;3g F=f4g
Vector form of the solution set to the system of equations (Theorem VFSLS [118]). Notice the rela-
tionship between the free variables and the set Fabove. Also, notice the pattern of 0's and 1's in the
entries of the vectors corresponding to elements of the set Ffor the larger examples.
2
4x1
x2
x33
5=2
4 3
5
23
5
Given a system of equations we can always build a new, related, homogeneous system (Denition HS
[71]) by converting the constant terms to zeros and retaining the coecients of the variables. Properties
Version 2.30
Archetype B 789
of this new system will have precise relationships with various properties of the original system.
11x1+ 2x2 14x3= 0
23x1 6x2+ 33x3= 0
14x1 2x2+ 17x3= 0
Some solutions to the associated homogenous system of linear equations (not necessarily exhaustive):
x1= 0; x 2= 0; x 3= 0
Form the augmented matrix of the homogenous linear system, and use row operations to convert to
reduced row-echelon form. Notice how the entries of the nal column remain zeros:
2
410 0 0
010 0
0 0 103
5
Analysis of the augmented matrix for the homogenous system (Notation RREFA [33]). Notice the
slight variation for the same analysis of the original system only when the original system was consistent:
r= 3 D=f1;2;3g F=f4g
Coecient matrix of original system of equations, and of associated homogenous system. This matrix
will be the subject of further analysis, rather than the systems of equations.
2
4 7 6 12
5 5 7
1 0 43
5
Matrix brought to reduced row-echelon form:
2
410 0
010
0 0 13
5
Analysis of the row-reduced matrix (Notation RREFA [33]):
r= 3 D=f1;2;3g F=fg
Matrix (coecient matrix) is nonsingular or singular? (Theorem NMRRI [84]) at the same time,
examine the size of the set Fabove.Notice that this property does not apply to matrices that are not
Version 2.30
790 Archetype B
square.
Nonsingular.
This is the null space of the matrix. The set of vectors used in the span construction is a linearly
independent set of column vectors that spans the null space of the matrix (Theorem SSNS [137], Theorem
BNS [160]). Solve the homogenous system with this matrix as the coecient matrix and write the solutions
in vector form (Theorem VFSLS [118]) to see these vectors arise.
hfgi
Column space of the matrix, expressed as the span of a set of linearly independent vectors that are
also columns of the matrix. These columns have indices that form the set Dabove. (Theorem BCS [274])
*8
<
:2
4 7
5
13
5;2
4 6
5
03
5;2
4 12
7
43
59
=
;+
The column space of the matrix, as it arises from the extended echelon form of the matrix. The matrix
Lis computed as described in Denition EEF [297]. This is followed by the column space described by a
set of linearly independent vectors that span the null space of L, computed as according to Theorem FS
[299] and Theorem BNS [160]. When r=m, the matrix Lhas no rows and the column space is all of Cm.
L=
*8
<
:2
41
0
03
5;2
40
1
03
5;2
40
0
13
59
=
;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors. These
vectors are computed by row-reducing the transpose of the matrix into reduced row-echelon form, tossing
out the zero rows, and writing the remaining nonzero rows as column vectors. By Theorem CSRST [282]
and Theorem BRS [280], and in the style of Example CSROI [282], this yields a linearly independent set
of vectors that span the column space.
*8
<
:2
41
0
03
5;2
40
1
03
5;2
40
0
13
59
=
;+
Row space of the matrix, expressed as a span of a set of linearly independent vectors, obtained from
the nonzero rows of the equivalent matrix in reduced row-echelon form. (Theorem BRS [280])
Version 2.30
Archetype B 791
*8
<
:2
41
0
03
5;2
40
1
03
5;2
40
0
13
59
=
;+
Inverse matrix, if it exists. The inverse is not dened for matrices that are not square, and if the matrix
is square, then the matrix must be nonsingular. (Denition MI [244], Theorem NI [261])
2
4 10 12 9
13
2811
25
235
23
5
Subspace dimensions associated with the matrix. (Denition NOM [397], Denition ROM [397]) Verify
Theorem RPNC [398]
Matrix columns: 3 Rank: 3 Nullity: 0
Determinant of the matrix, which is only dened for square matrices. The matrix is nonsingular if and
only if the determinant is nonzero (Theorem SMZD [445]). (Product of all eigenvalues?)
Determinant = 2
Eigenvalues, and bases for eigenspaces. (Denition EEM [453],Denition EM [461])
= 1 EB( 1) =*8
<
:2
4 5
3
13
59
=
;+
= 1 EB(1) =*8
<
:2
4 3
2
13
59
=
;+
= 2 EB(2) =*8
<
:2
4 2
1
13
59
=
;+
Geometric and algebraic multiplicities. (Denition GME [463]Denition AME [463])
B( 1) = 1 B( 1) = 1
B(1) = 1 B(1) = 1
B(2) = 1 B(2) = 1
Diagonalizable? (Denition DZM [496])
Version 2.30
792 Archetype B
Yes, distinct eigenvalues, Theorem DED [501].
The diagonalization. (Theorem DC [497])
2
4 1 1 1
2 3 1
1 2 13
52
4 7 6 12
5 5 7
1 0 43
52
4 5 3 2
3 2 1
1 1 13
5
=2
4 1 0 0
0 1 0
0 0 23
5
Version 2.30
Archetype C 793
Archetype C
Summary System with three equations, four variables. Consistent. Null space of coecient matrix has
dimension 1.
A system of linear equations (Denition SLE [11]):
2x1 3x2+x3 6x4= 7
4x1+x2+ 2x3+ 9x4= 7
3x1+x2+x3+ 8x4= 8
Some solutions to the system of linear equations (not necessarily exhaustive):
x1= 7; x 2= 2; x 3= 7; x 4= 1
x1= 1; x 2= 7; x 3= 4; x 4= 2
Augmented matrix of the linear system of equations (Denition AM [30]):
2
42 3 1 6 7
4 1 2 9 7
3 1 1 8 83
5
Matrix in reduced row-echelon form, row-equivalent to augmented matrix:
2
410 0 2 5
010 3 1
0 0 1 1 63
5
Analysis of the augmented matrix (Notation RREFA [33]):
r= 3 D=f1;2;3g F=f4;5g
Vector form of the solution set to the system of equations (Theorem VFSLS [118]). Notice the rela-
tionship between the free variables and the set Fabove. Also, notice the pattern of 0's and 1's in the
entries of the vectors corresponding to elements of the set Ffor the larger examples.
Version 2.30
794 Archetype C
2
664x1
x2
x3
x43
775=2
664 5
1
6
03
775+x42
664 2
3
1
13
775
Given a system of equations we can always build a new, related, homogeneous system (Denition HS
[71]) by converting the constant terms to zeros and retaining the coecients of the variables. Properties
of this new system will have precise relationships with various properties of the original system.
2x1 3x2+x3 6x4= 0
4x1+x2+ 2x3+ 9x4= 0
3x1+x2+x3+ 8x4= 0
Some solutions to the associated homogenous system of linear equations (not necessarily exhaustive):
x1= 0; x 2= 0; x 3= 0; x 4= 0
x1= 2; x 2= 3; x 3= 1; x 4= 1
x1= 4; x 2= 6; x 3= 2; x 4= 2
Form the augmented matrix of the homogenous linear system, and use row operations to convert to
reduced row-echelon form. Notice how the entries of the nal column remain zeros:
2
410 0 2 0
010 3 0
0 0 1 1 03
5
Analysis of the augmented matrix for the homogenous system (Notation RREFA [33]). Notice the
slight variation for the same analysis of the original system only when the original system was consistent:
r= 3 D=f1;2;3g F=f4;5g
Coecient matrix of original system of equations, and of associated homogenous system. This matrix
will be the subject of further analysis, rather than the systems of equations.
2
42 3 1 6
4 1 2 9
3 1 1 83
5
Matrix brought to reduced row-echelon form:
2
410 0 2
010 3
0 0 1 13
5
Version 2.30
Archetype C 795
Analysis of the row-reduced matrix (Notation RREFA [33]):
r= 3 D=f1;2;3g F=f4g
This is the null space of the matrix. The set of vectors used in the span construction is a linearly
independent set of column vectors that spans the null space of the matrix (Theorem SSNS [137], Theorem
BNS [160]). Solve the homogenous system with this matrix as the coecient matrix and write the solutions
in vector form (Theorem VFSLS [118]) to see these vectors arise.
*8
>><
>>:2
664 2
3
1
13
7759
>>=
>>;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors that are
also columns of the matrix. These columns have indices that form the set Dabove. (Theorem BCS [274])
*8
<
:2
42
4
33
5;2
4 3
1
13
5;2
41
2
13
59
=
;+
The column space of the matrix, as it arises from the extended echelon form of the matrix. The matrix
Lis computed as described in Denition EEF [297]. This is followed by the column space described by a
set of linearly independent vectors that span the null space of L, computed as according to Theorem FS
[299] and Theorem BNS [160]. When r=m, the matrix Lhas no rows and the column space is all of Cm.
L=
*8
<
:2
41
0
03
5;2
40
1
03
5;2
40
0
13
59
=
;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors. These
vectors are computed by row-reducing the transpose of the matrix into reduced row-echelon form, tossing
out the zero rows, and writing the remaining nonzero rows as column vectors. By Theorem CSRST [282]
and Theorem BRS [280], and in the style of Example CSROI [282], this yields a linearly independent set
of vectors that span the column space.
*8
<
:2
41
0
03
5;2
40
1
03
5;2
40
0
13
59
=
;+
Row space of the matrix, expressed as a span of a set of linearly independent vectors, obtained from
the nonzero rows of the equivalent matrix in reduced row-echelon form. (Theorem BRS [280])
Version 2.30
796 Archetype C
*8
>><
>>:2
6641
0
0
23
775;2
6640
1
0
33
775;2
6640
0
1
13
7759
>>=
>>;+
Subspace dimensions associated with the matrix. (Denition NOM [397], Denition ROM [397]) Verify
Theorem RPNC [398]
Matrix columns: 4 Rank: 3 Nullity: 1
Version 2.30
Archetype D 797
Archetype D
Summary System with three equations, four variables. Consistent. Null space of coecient matrix has
dimension 2. Coecient matrix identical to that of Archetype E, vector of constants is dierent.
A system of linear equations (Denition SLE [11]):
2x1+x2+ 7x3 7x4= 8
3x1+ 4x2 5x3 6x4= 12
x1+x2+ 4x3 5x4= 4
Some solutions to the system of linear equations (not necessarily exhaustive):
x1= 0; x 2= 1; x 3= 2; x 4= 1
x1= 4; x 2= 0; x 3= 0; x 4= 0
x1= 7; x 2= 8; x 3= 1; x 4= 3
Augmented matrix of the linear system of equations (Denition AM [30]):
2
42 1 7 7 8
3 4 5 6 12
1 1 4 5 43
5
Matrix in reduced row-echelon form, row-equivalent to augmented matrix:
2
410 3 2 4
011 3 0
0 0 0 0 03
5
Analysis of the augmented matrix (Notation RREFA [33]):
r= 2 D=f1;2g F=f3;4;5g
Vector form of the solution set to the system of equations (Theorem VFSLS [118]). Notice the rela-
tionship between the free variables and the set Fabove. Also, notice the pattern of 0's and 1's in the
entries of the vectors corresponding to elements of the set Ffor the larger examples.
Version 2.30
798 Archetype D
2
664x1
x2
x3
x43
775=2
6644
0
0
03
775+x32
664 3
1
1
03
775+x42
6642
3
0
13
775
Given a system of equations we can always build a new, related, homogeneous system (Denition HS
[71]) by converting the constant terms to zeros and retaining the coecients of the variables. Properties
of this new system will have precise relationships with various properties of the original system.
2x1+x2+ 7x3 7x4= 0
3x1+ 4x2 5x3 6x4= 0
x1+x2+ 4x3 5x4= 0
Some solutions to the associated homogenous system of linear equations (not necessarily exhaustive):
x1= 0; x 2= 0; x 3= 0; x 4= 0
x1= 3; x 2= 1; x 3= 1; x 4= 0
x1= 2; x 2= 3; x 3= 0; x 4= 1
x1= 1; x 2= 2; x 3= 1; x 4= 1
Form the augmented matrix of the homogenous linear system, and use row operations to convert to
reduced row-echelon form. Notice how the entries of the nal column remain zeros:
2
410 3 2 0
011 3 0
0 0 0 0 03
5
Analysis of the augmented matrix for the homogenous system (Notation RREFA [33]). Notice the
slight variation for the same analysis of the original system only when the original system was consistent:
r= 2 D=f1;2g F=f3;4;5g
Coecient matrix of original system of equations, and of associated homogenous system. This matrix
will be the subject of further analysis, rather than the systems of equations.
2
42 1 7 7
3 4 5 6
1 1 4 53
5
Matrix brought to reduced row-echelon form:
2
410 3 2
011 3
0 0 0 03
5
Version 2.30
Archetype D 799
Analysis of the row-reduced matrix (Notation RREFA [33]):
r= 2 D=f1;2g F=f3;4g
This is the null space of the matrix. The set of vectors used in the span construction is a linearly
independent set of column vectors that spans the null space of the matrix (Theorem SSNS [137], Theorem
BNS [160]). Solve the homogenous system with this matrix as the coecient matrix and write the solutions
in vector form (Theorem VFSLS [118]) to see these vectors arise.
*8
>><
>>:2
664 3
1
1
03
775;2
6642
3
0
13
7759
>>=
>>;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors that are
also columns of the matrix. These columns have indices that form the set Dabove. (Theorem BCS [274])
*8
<
:2
42
3
13
5;2
41
4
13
59
=
;+
The column space of the matrix, as it arises from the extended echelon form of the matrix. The matrix
Lis computed as described in Denition EEF [297]. This is followed by the column space described by a
set of linearly independent vectors that span the null space of L, computed as according to Theorem FS
[299] and Theorem BNS [160]. When r=m, the matrix Lhas no rows and the column space is all of Cm.
L=
11
7 11
7
*8
<
:2
411
7
0
13
5;2
4 1
7
1
03
59
=
;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors. These
vectors are computed by row-reducing the transpose of the matrix into reduced row-echelon form, tossing
out the zero rows, and writing the remaining nonzero rows as column vectors. By Theorem CSRST [282]
and Theorem BRS [280], and in the style of Example CSROI [282], this yields a linearly independent set
of vectors that span the column space.
*8
<
:2
41
0
7
113
5;2
40
1
1
113
59
=
;+
Row space of the matrix, expressed as a span of a set of linearly independent vectors, obtained from
the nonzero rows of the equivalent matrix in reduced row-echelon form. (Theorem BRS [280])
Version 2.30
800 Archetype D
*8
>><
>>:2
6641
0
3
23
775;2
6640
1
1
33
7759
>>=
>>;+
Subspace dimensions associated with the matrix. (Denition NOM [397], Denition ROM [397]) Verify
Theorem RPNC [398]
Matrix columns: 4 Rank: 2 Nullity: 2
Version 2.30
Archetype E 801
Archetype E
Summary System with three equations, four variables. Inconsistent. Null space of coecient matrix
has dimension 2. Coecient matrix identical to that of Archetype D, constant vector is dierent.
A system of linear equations (Denition SLE [11]):
2x1+x2+ 7x3 7x4= 2
3x1+ 4x2 5x3 6x4= 3
x1+x2+ 4x3 5x4= 2
Some solutions to the system of linear equations (not necessarily exhaustive):
None. (Why?)
Augmented matrix of the linear system of equations (Denition AM [30]):
2
42 1 7 7 2
3 4 5 6 3
1 1 4 5 23
5
Matrix in reduced row-echelon form, row-equivalent to augmented matrix:
2
410 3 2 0
011 3 0
0 0 0 0 13
5
Analysis of the augmented matrix (Notation RREFA [33]):
r= 3 D=f1;2;5g F=f3;4g
Vector form of the solution set to the system of equations (Theorem VFSLS [118]). Notice the rela-
tionship between the free variables and the set Fabove. Also, notice the pattern of 0's and 1's in the
entries of the vectors corresponding to elements of the set Ffor the larger examples.
Inconsistent system, no solutions exist.
Given a system of equations we can always build a new, related, homogeneous system (Denition HS
[71]) by converting the constant terms to zeros and retaining the coecients of the variables. Properties
Version 2.30
802 Archetype E
of this new system will have precise relationships with various properties of the original system.
2x1+x2+ 7x3 7x4= 0
3x1+ 4x2 5x3 6x4= 0
x1+x2+ 4x3 5x4= 0
Some solutions to the associated homogenous system of linear equations (not necessarily exhaustive):
x1= 0; x 2= 0; x 3= 0; x 4= 0
x1= 4; x 2= 13; x 3= 2; x 4= 5
Form the augmented matrix of the homogenous linear system, and use row operations to convert to
reduced row-echelon form. Notice how the entries of the nal column remain zeros:
2
410 3 2 0
011 3 0
0 0 0 0 03
5
Analysis of the augmented matrix for the homogenous system (Notation RREFA [33]). Notice the
slight variation for the same analysis of the original system only when the original system was consistent:
r= 2 D=f1;2g F=f3;4;5g
Coecient matrix of original system of equations, and of associated homogenous system. This matrix
will be the subject of further analysis, rather than the systems of equations.
2
42 1 7 7
3 4 5 6
1 1 4 53
5
Matrix brought to reduced row-echelon form:
2
410 3 2
011 3
0 0 0 03
5
Analysis of the row-reduced matrix (Notation RREFA [33]):
r= 2 D=f1;2g F=f3;4g
This is the null space of the matrix. The set of vectors used in the span construction is a linearly
independent set of column vectors that spans the null space of the matrix (Theorem SSNS [137], Theorem
Version 2.30
Archetype E 803
BNS [160]). Solve the homogenous system with this matrix as the coecient matrix and write the solutions
in vector form (Theorem VFSLS [118]) to see these vectors arise.
*8
>><
>>:2
664 3
1
1
03
775;2
6642
3
0
13
7759
>>=
>>;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors that are
also columns of the matrix. These columns have indices that form the set Dabove. (Theorem BCS [274])
*8
<
:2
42
3
13
5;2
41
4
13
59
=
;+
The column space of the matrix, as it arises from the extended echelon form of the matrix. The matrix
Lis computed as described in Denition EEF [297]. This is followed by the column space described by a
set of linearly independent vectors that span the null space of L, computed as according to Theorem FS
[299] and Theorem BNS [160]. When r=m, the matrix Lhas no rows and the column space is all of Cm.
L=
11
7 11
7
*8
<
:2
411
7
0
13
5;2
4 1
7
1
03
59
=
;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors. These
vectors are computed by row-reducing the transpose of the matrix into reduced row-echelon form, tossing
out the zero rows, and writing the remaining nonzero rows as column vectors. By Theorem CSRST [282]
and Theorem BRS [280], and in the style of Example CSROI [282], this yields a linearly independent set
of vectors that span the column space.
*8
<
:2
41
0
7
113
5;2
40
1
1
113
59
=
;+
Row space of the matrix, expressed as a span of a set of linearly independent vectors, obtained from
the nonzero rows of the equivalent matrix in reduced row-echelon form. (Theorem BRS [280])
*8
>><
>>:2
6641
0
3
23
775;2
6640
1
1
33
7759
>>=
>>;+
Subspace dimensions associated with the matrix. (Denition NOM [397], Denition ROM [397]) Verify
Version 2.30
804 Archetype E
Theorem RPNC [398]
Matrix columns: 4 Rank: 2 Nullity: 2
Version 2.30
Archetype F 805
Archetype F
Summary System with four equations, four variables. Nonsingular coecient matrix. Integer eigenval-
ues, one has \high" multiplicity.
A system of linear equations (Denition SLE [11]):
33x1 16x2+ 10x3 2x4= 27
99x1 47x2+ 27x3 7x4= 77
78x1 36x2+ 17x3 6x4= 52
9x1+ 2x2+ 3x3+ 4x4= 5
Some solutions to the system of linear equations (not necessarily exhaustive):
x1= 1; x 2= 2; x 3= 2; x 4= 4
Augmented matrix of the linear system of equations (Denition AM [30]):
2
66433 16 10 2 27
99 47 27 7 77
78 36 17 6 52
9 2 3 4 53
775
Matrix in reduced row-echelon form, row-equivalent to augmented matrix:
2
666410 0 0 1
010 0 2
0 0 10 2
0 0 0 1 43
7775
Analysis of the augmented matrix (Notation RREFA [33]):
r= 4 D=f1;2;3;4g F=f5g
Vector form of the solution set to the system of equations (Theorem VFSLS [118]). Notice the rela-
tionship between the free variables and the set Fabove. Also, notice the pattern of 0's and 1's in the
entries of the vectors corresponding to elements of the set Ffor the larger examples.
Version 2.30
806 Archetype F
2
664x1
x2
x3
x43
775=2
6641
2
2
43
775
Given a system of equations we can always build a new, related, homogeneous system (Denition HS
[71]) by converting the constant terms to zeros and retaining the coecients of the variables. Properties
of this new system will have precise relationships with various properties of the original system.
33x1 16x2+ 10x3 2x4= 0
99x1 47x2+ 27x3 7x4= 0
78x1 36x2+ 17x3 6x4= 0
9x1+ 2x2+ 3x3+ 4x4= 0
Some solutions to the associated homogenous system of linear equations (not necessarily exhaustive):
x1= 0; x 2= 0; x 3= 0; x 4= 0
Form the augmented matrix of the homogenous linear system, and use row operations to convert to
reduced row-echelon form. Notice how the entries of the nal column remain zeros:
2
666410 0 0 0
010 0 0
0 0 10 0
0 0 0 103
7775
Analysis of the augmented matrix for the homogenous system (Notation RREFA [33]). Notice the
slight variation for the same analysis of the original system only when the original system was consistent:
r= 4 D=f1;2;3;4g F=f5g
Coecient matrix of original system of equations, and of associated homogenous system. This matrix
will be the subject of further analysis, rather than the systems of equations.
2
66433 16 10 2
99 47 27 7
78 36 17 6
9 2 3 43
775
Matrix brought to reduced row-echelon form:
2
666410 0 0
010 0
0 0 10
0 0 0 13
7775
Version 2.30
Archetype F 807
Analysis of the row-reduced matrix (Notation RREFA [33]):
r= 4 D=f1;2;3;4g F=fg
Matrix (coecient matrix) is nonsingular or singular? (Theorem NMRRI [84]) at the same time,
examine the size of the set Fabove.Notice that this property does not apply to matrices that are not
square.
Nonsingular.
This is the null space of the matrix. The set of vectors used in the span construction is a linearly
independent set of column vectors that spans the null space of the matrix (Theorem SSNS [137], Theorem
BNS [160]). Solve the homogenous system with this matrix as the coecient matrix and write the solutions
in vector form (Theorem VFSLS [118]) to see these vectors arise.
hfgi
Column space of the matrix, expressed as the span of a set of linearly independent vectors that are
also columns of the matrix. These columns have indices that form the set Dabove. (Theorem BCS [274])
*8
>><
>>:2
66433
99
78
93
775;2
664 16
47
36
23
775;2
66410
27
17
33
775;2
664 2
7
6
43
7759
>>=
>>;+
The column space of the matrix, as it arises from the extended echelon form of the matrix. The matrix
Lis computed as described in Denition EEF [297]. This is followed by the column space described by a
set of linearly independent vectors that span the null space of L, computed as according to Theorem FS
[299] and Theorem BNS [160]. When r=m, the matrix Lhas no rows and the column space is all of Cm.
L=
*8
>><
>>:2
6641
0
0
03
775;2
6640
1
0
03
775;2
6640
0
1
03
775;2
6640
0
0
13
7759
>>=
>>;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors. These
vectors are computed by row-reducing the transpose of the matrix into reduced row-echelon form, tossing
out the zero rows, and writing the remaining nonzero rows as column vectors. By Theorem CSRST [282]
and Theorem BRS [280], and in the style of Example CSROI [282], this yields a linearly independent set
of vectors that span the column space.
Version 2.30
808 Archetype F
*8
>><
>>:2
6641
0
0
03
775;2
6640
1
0
03
775;2
6640
0
1
03
775;2
6640
0
0
13
7759
>>=
>>;+
Row space of the matrix, expressed as a span of a set of linearly independent vectors, obtained from
the nonzero rows of the equivalent matrix in reduced row-echelon form. (Theorem BRS [280])
*8
>><
>>:2
6641
0
0
03
775;2
6640
1
0
03
775;2
6640
0
1
03
775;2
6640
0
0
13
7759
>>=
>>;+
Inverse matrix, if it exists. The inverse is not dened for matrices that are not square, and if the matrix
is square, then the matrix must be nonsingular. (Denition MI [244], Theorem NI [261])
2
664 86
338
3 11
37
3
129
286
3 17
231
6
13 6 2 1
45
229
3 5
213
63
775
Subspace dimensions associated with the matrix. (Denition NOM [397], Denition ROM [397]) Verify
Theorem RPNC [398]
Matrix columns: 4 Rank: 4 Nullity: 0
Determinant of the matrix, which is only dened for square matrices. The matrix is nonsingular if and
only if the determinant is nonzero (Theorem SMZD [445]). (Product of all eigenvalues?)
Determinant = 18
Eigenvalues, and bases for eigenspaces. (Denition EEM [453],Denition EM [461])
= 1 EF( 1) =*8
>><
>>:2
6641
2
0
13
7759
>>=
>>;+
= 2 EF(2) =*8
>><
>>:2
6642
5
2
13
7759
>>=
>>;+
= 3 EF(3) =*8
>><
>>:2
6641
1
0
73
775;2
66417
45
21
03
7759
>>=
>>;+
Version 2.30
Archetype F 809
Geometric and algebraic multiplicities. (Denition GME [463]Denition AME [463])
F( 1) = 1 F( 1) = 1
F(2) = 1 F(2) = 1
F(3) = 2 F(3) = 2
Diagonalizable? (Denition DZM [496])
Yes, full eigenspaces, Theorem DMFE [499].
The diagonalization. (Theorem DC [497])
2
66412 5 1 1
39 18 7 3
27
7 13
76
7 1
726
7 12
75
7 2
73
7752
66433 16 10 2
99 47 27 7
78 36 17 6
9 2 3 43
7752
6641 2 1 17
2 5 1 45
0 2 0 21
1 1 7 03
775
=2
664 1 0 0 0
0 2 0 0
0 0 3 0
0 0 0 33
775
Version 2.30
810 Archetype G
Archetype G
Summary System with ve equations, two variables. Consistent. Null space of coecient matrix has
dimension 0. Coecient matrix identical to that of Archetype H, constant vector is dierent.
A system of linear equations (Denition SLE [11]):
2x1+ 3x2= 6
x1+ 4x2= 14
3x1+ 10x2= 2
3x1 x2= 20
6x1+ 9x2= 18
Some solutions to the system of linear equations (not necessarily exhaustive):
x1= 6; x 2= 2
Augmented matrix of the linear system of equations (Denition AM [30]):
2
666642 3 6
1 4 14
3 10 2
3 1 20
6 9 183
77775
Matrix in reduced row-echelon form, row-equivalent to augmented matrix:
2
6666410 6
01 2
0 0 0
0 0 0
0 0 03
77775
Analysis of the augmented matrix (Notation RREFA [33]):
r= 2 D=f1;2g F=f3g
Vector form of the solution set to the system of equations (Theorem VFSLS [118]). Notice the rela-
tionship between the free variables and the set Fabove. Also, notice the pattern of 0's and 1's in the
Version 2.30
Archetype G 811
entries of the vectors corresponding to elements of the set Ffor the larger examples.
x1
x2
=6
2
Given a system of equations we can always build a new, related, homogeneous system (Denition HS
[71]) by converting the constant terms to zeros and retaining the coecients of the variables. Properties
of this new system will have precise relationships with various properties of the original system.
2x1+ 3x2= 0
x1+ 4x2= 0
3x1+ 10x2= 0
3x1 x2= 0
6x1+ 9x2= 0
Some solutions to the associated homogenous system of linear equations (not necessarily exhaustive):
x1= 0; x 2= 0
Form the augmented matrix of the homogenous linear system, and use row operations to convert to
reduced row-echelon form. Notice how the entries of the nal column remain zeros:
2
6666410 0
010
0 0 0
0 0 0
0 0 03
77775
Analysis of the augmented matrix for the homogenous system (Notation RREFA [33]). Notice the
slight variation for the same analysis of the original system only when the original system was consistent:
r= 2 D=f1;2g F=f3g
Coecient matrix of original system of equations, and of associated homogenous system. This matrix
will be the subject of further analysis, rather than the systems of equations.
2
666642 3
1 4
3 10
3 1
6 93
77775
Version 2.30
812 Archetype G
Matrix brought to reduced row-echelon form:
2
6666410
01
0 0
0 0
0 03
77775
Analysis of the row-reduced matrix (Notation RREFA [33]):
r= 2 D=f1;2g F=fg
This is the null space of the matrix. The set of vectors used in the span construction is a linearly
independent set of column vectors that spans the null space of the matrix (Theorem SSNS [137], Theorem
BNS [160]). Solve the homogenous system with this matrix as the coecient matrix and write the solutions
in vector form (Theorem VFSLS [118]) to see these vectors arise.
hfgi
Column space of the matrix, expressed as the span of a set of linearly independent vectors that are
also columns of the matrix. These columns have indices that form the set Dabove. (Theorem BCS [274])
*8
>>>><
>>>>:2
666642
1
3
3
63
77775;2
666643
4
10
1
93
777759
>>>>=
>>>>;+
The column space of the matrix, as it arises from the extended echelon form of the matrix. The matrix
Lis computed as described in Denition EEF [297]. This is followed by the column space described by a
set of linearly independent vectors that span the null space of L, computed as according to Theorem FS
[299] and Theorem BNS [160]. When r=m, the matrix Lhas no rows and the column space is all of Cm.
L=2
41 0 0 0 1
3
0 1 0 1 1
3
0 0 1 1 13
5
*8
>>>><
>>>>:2
666641
31
3
1
0
13
77775;2
666640
1
1
1
03
777759
>>>>=
>>>>;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors. These
Version 2.30
Archetype G 813
vectors are computed by row-reducing the transpose of the matrix into reduced row-echelon form, tossing
out the zero rows, and writing the remaining nonzero rows as column vectors. By Theorem CSRST [282]
and Theorem BRS [280], and in the style of Example CSROI [282], this yields a linearly independent set
of vectors that span the column space.
*8
>>>><
>>>>:2
666641
0
2
1
33
77775;2
666640
1
1
1
03
777759
>>>>=
>>>>;+
Row space of the matrix, expressed as a span of a set of linearly independent vectors, obtained from
the nonzero rows of the equivalent matrix in reduced row-echelon form. (Theorem BRS [280])
1
0
;0
1
Subspace dimensions associated with the matrix. (Denition NOM [397], Denition ROM [397]) Verify
Theorem RPNC [398]
Matrix columns: 2 Rank: 2 Nullity: 0
Version 2.30
814 Archetype H
Archetype H
Summary System with ve equations, two variables. Inconsistent, overdetermined. Null space of
coecient matrix has dimension 0. Coecient matrix identical to that of Archetype G, constant vector is
dierent.
A system of linear equations (Denition SLE [11]):
2x1+ 3x2= 5
x1+ 4x2= 6
3x1+ 10x2= 2
3x1 x2= 1
6x1+ 9x2= 3
Some solutions to the system of linear equations (not necessarily exhaustive):
None. (Why?)
Augmented matrix of the linear system of equations (Denition AM [30]):
2
666642 3 5
1 4 6
3 10 2
3 1 1
6 9 33
77775
Matrix in reduced row-echelon form, row-equivalent to augmented matrix:
2
66666410 0
010
0 0 1
0 0 0
0 0 03
777775
Analysis of the augmented matrix (Notation RREFA [33]):
r= 3 D=f1;2;3g F=fg
Vector form of the solution set to the system of equations (Theorem VFSLS [118]). Notice the rela-
tionship between the free variables and the set Fabove. Also, notice the pattern of 0's and 1's in the
Version 2.30
Archetype H 815
entries of the vectors corresponding to elements of the set Ffor the larger examples.
Inconsistent system, no solutions exist.
Given a system of equations we can always build a new, related, homogeneous system (Denition HS
[71]) by converting the constant terms to zeros and retaining the coecients of the variables. Properties
of this new system will have precise relationships with various properties of the original system.
2x1+ 3x2= 0
x1+ 4x2= 0
3x1+ 10x2= 0
3x1 x2= 0
6x1+ 9x2= 0
Some solutions to the associated homogenous system of linear equations (not necessarily exhaustive):
x1= 0; x 2= 0
Form the augmented matrix of the homogenous linear system, and use row operations to convert to
reduced row-echelon form. Notice how the entries of the nal column remain zeros:
2
6666410 0
010
0 0 0
0 0 0
0 0 03
77775
Analysis of the augmented matrix for the homogenous system (Notation RREFA [33]). Notice the
slight variation for the same analysis of the original system only when the original system was consistent:
r= 2 D=f1;2g F=f3g
Coecient matrix of original system of equations, and of associated homogenous system. This matrix
will be the subject of further analysis, rather than the systems of equations.
2
666642 3
1 4
3 10
3 1
6 93
77775
Version 2.30
816 Archetype H
Matrix brought to reduced row-echelon form:
2
6666410
01
0 0
0 0
0 03
77775
Analysis of the row-reduced matrix (Notation RREFA [33]):
r= 2 D=f1;2g F=fg
This is the null space of the matrix. The set of vectors used in the span construction is a linearly
independent set of column vectors that spans the null space of the matrix (Theorem SSNS [137], Theorem
BNS [160]). Solve the homogenous system with this matrix as the coecient matrix and write the solutions
in vector form (Theorem VFSLS [118]) to see these vectors arise.
hfgi
Column space of the matrix, expressed as the span of a set of linearly independent vectors that are
also columns of the matrix. These columns have indices that form the set Dabove. (Theorem BCS [274])
*8
>>>><
>>>>:2
666642
1
3
3
63
77775;2
666643
4
10
1
93
777759
>>>>=
>>>>;+
The column space of the matrix, as it arises from the extended echelon form of the matrix. The matrix
Lis computed as described in Denition EEF [297]. This is followed by the column space described by a
set of linearly independent vectors that span the null space of L, computed as according to Theorem FS
[299] and Theorem BNS [160]. When r=m, the matrix Lhas no rows and the column space is all of Cm.
L=
*8
>>>><
>>>>:2
666641
31
3
1
0
13
77775;2
666640
1
1
1
03
777759
>>>>=
>>>>;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors. These
vectors are computed by row-reducing the transpose of the matrix into reduced row-echelon form, tossing
Version 2.30
Archetype H 817
out the zero rows, and writing the remaining nonzero rows as column vectors. By Theorem CSRST [282]
and Theorem BRS [280], and in the style of Example CSROI [282], this yields a linearly independent set
of vectors that span the column space.
*8
>>>><
>>>>:2
666641
0
2
1
33
77775;2
666640
1
1
1
03
777759
>>>>=
>>>>;+
The column space of the matrix, as it arises from the extended echelon form of the matrix. The matrix
Lis computed as described in Denition EEF [297]. This is followed by the column space described by a
set of linearly independent vectors that span the null space of L, computed as according to Theorem FS
[299] and Theorem BNS [160]. When r=m, the matrix Lhas no rows and the column space is all of Cm.
L=2
41 0 0 0 1
3
0 1 0 1 1
3
0 0 1 1 13
5
*8
>>>><
>>>>:2
666641
31
3
1
0
13
77775;2
666640
1
1
1
03
777759
>>>>=
>>>>;+
Row space of the matrix, expressed as a span of a set of linearly independent vectors, obtained from
the nonzero rows of the equivalent matrix in reduced row-echelon form. (Theorem BRS [280])
1
0
;0
1
Subspace dimensions associated with the matrix. (Denition NOM [397], Denition ROM [397]) Verify
Theorem RPNC [398]
Matrix columns: 2 Rank: 2 Nullity: 0
Version 2.30
818 Archetype I
Archetype I
Summary System with four equations, seven variables. Consistent. Null space of coecient matrix has
dimension 4.
A system of linear equations (Denition SLE [11]):
x1+ 4x2 x4+ 7x6 9x7= 3
2x1+ 8x2 x3+ 3x4+ 9x5 13x6+ 7x7= 9
2x3 3x4 4x5+ 12x6 8x7= 1
x1 4x2+ 2x3+ 4x4+ 8x5 31x6+ 37x7= 4
Some solutions to the system of linear equations (not necessarily exhaustive):
x1= 25,x2= 4,x3= 22,x4= 29,x5= 1,x6= 2,x7= 3
x1= 7,x2= 5,x3= 7,x4= 15,x5= 4,x6= 2,x7= 1
x1= 4,x2= 0,x3= 2,x4= 1,x5= 0,x6= 0,x7= 0
Augmented matrix of the linear system of equations (Denition AM [30]):
2
6641 4 0 1 0 7 9 3
2 8 1 3 9 13 7 9
0 0 2 3 4 12 8 1
1 4 2 4 8 31 37 43
775
Matrix in reduced row-echelon form, row-equivalent to augmented matrix:
2
66414 0 0 2 1 3 4
0 0 10 1 3 5 2
0 0 0 12 6 6 1
0 0 0 0 0 0 0 03
775
Analysis of the augmented matrix (Notation RREFA [33]):
r= 3 D=f1;3;4g F=f2;5;6;7;8g
Vector form of the solution set to the system of equations (Theorem VFSLS [118]). Notice the rela-
tionship between the free variables and the set Fabove. Also, notice the pattern of 0's and 1's in the
Version 2.30
Archetype I 819
entries of the vectors corresponding to elements of the set Ffor the larger examples.
2
666666664x1
x2
x3
x4
x5
x6
x73
777777775=2
6666666644
0
2
1
0
0
03
777777775+x22
666666664 4
1
0
0
0
0
03
777777775+x52
666666664 2
0
1
2
1
0
03
777777775+x62
666666664 1
0
3
6
0
1
03
777777775+x72
6666666643
0
5
6
0
0
13
777777775
Given a system of equations we can always build a new, related, homogeneous system (Denition HS
[71]) by converting the constant terms to zeros and retaining the coecients of the variables. Properties
of this new system will have precise relationships with various properties of the original system.
x1+ 4x2 x4+ 7x6 9x7= 0
2x1+ 8x2 x3+ 3x4+ 9x5 13x6+ 7x7= 0
2x3 3x4 4x5+ 12x6 8x7= 0
x1 4x2+ 2x3+ 4x4+ 8x5 31x6+ 37x7= 0
Some solutions to the associated homogenous system of linear equations (not necessarily exhaustive):
x1= 0,x2= 0,x3= 0,x4= 0,x5= 0,x6= 0,x7= 0
x1= 3,x2= 0,x3= 5,x4= 6,x5= 0,x6= 0,x7= 1
x1= 1,x2= 0,x3= 3,x4= 6,x5= 0,x6= 1,x7= 0
x1= 2,x2= 0,x3= 1,x4= 2,x5= 1,x6= 0,x7= 0
x1= 4,x2= 1,x3= 0,x4= 0,x5= 0,x6= 0,x7= 0
x1= 4,x2= 1,x3= 3,x4= 2,x5= 1,x6= 1,x7= 1
Form the augmented matrix of the homogenous linear system, and use row operations to convert to
reduced row-echelon form. Notice how the entries of the nal column remain zeros:
2
66414 0 0 2 1 3 0
0 0 10 1 3 5 0
0 0 0 12 6 6 0
0 0 0 0 0 0 0 03
775
Analysis of the augmented matrix for the homogenous system (Notation RREFA [33]). Notice the
slight variation for the same analysis of the original system only when the original system was consistent:
r= 3 D=f1;3;4g F=f2;5;6;7;8g
Version 2.30
820 Archetype I
Coecient matrix of original system of equations, and of associated homogenous system. This matrix
will be the subject of further analysis, rather than the systems of equations.
2
6641 4 0 1 0 7 9
2 8 1 3 9 13 7
0 0 2 3 4 12 8
1 4 2 4 8 31 373
775
Matrix brought to reduced row-echelon form:
2
66414 0 0 2 1 3
0 0 10 1 3 5
0 0 0 12 6 6
0 0 0 0 0 0 03
775
Analysis of the row-reduced matrix (Notation RREFA [33]):
r= 3 D=f1;3;4g F=f2;5;6;7g
This is the null space of the matrix. The set of vectors used in the span construction is a linearly
independent set of column vectors that spans the null space of the matrix (Theorem SSNS [137], Theorem
BNS [160]). Solve the homogenous system with this matrix as the coecient matrix and write the solutions
in vector form (Theorem VFSLS [118]) to see these vectors arise.
*8
>>>>>>>><
>>>>>>>>:2
666666664 4
1
0
0
0
0
03
777777775;2
666666664 2
0
1
2
1
0
03
777777775;2
666666664 1
0
3
6
0
1
03
777777775;2
6666666643
0
5
6
0
0
13
7777777759
>>>>>>>>=
>>>>>>>>;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors that are
also columns of the matrix. These columns have indices that form the set Dabove. (Theorem BCS [274])
*8
>><
>>:2
6641
2
0
13
775;2
6640
1
2
23
775;2
664 1
3
3
43
7759
>>=
>>;+
The column space of the matrix, as it arises from the extended echelon form of the matrix. The matrix
Lis computed as described in Denition EEF [297]. This is followed by the column space described by a
Version 2.30
Archetype I 821
set of linearly independent vectors that span the null space of L, computed as according to Theorem FS
[299] and Theorem BNS [160]. When r=m, the matrix Lhas no rows and the column space is all of Cm.
L=
1 12
31 13
317
31
*8
>><
>>:2
664 7
31
0
0
13
775;2
66413
31
0
1
03
775;2
66412
31
1
0
03
7759
>>=
>>;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors. These
vectors are computed by row-reducing the transpose of the matrix into reduced row-echelon form, tossing
out the zero rows, and writing the remaining nonzero rows as column vectors. By Theorem CSRST [282]
and Theorem BRS [280], and in the style of Example CSROI [282], this yields a linearly independent set
of vectors that span the column space.
*8
>><
>>:2
6641
0
0
31
73
775;2
6640
1
0
12
73
775;2
6640
0
1
13
73
7759
>>=
>>;+
Row space of the matrix, expressed as a span of a set of linearly independent vectors, obtained from
the nonzero rows of the equivalent matrix in reduced row-echelon form. (Theorem BRS [280])
*8
>>>>>>>><
>>>>>>>>:2
6666666641
4
0
0
2
1
33
777777775;2
6666666640
0
1
0
1
3
53
777777775;2
6666666640
0
0
1
2
6
63
7777777759
>>>>>>>>=
>>>>>>>>;+
Subspace dimensions associated with the matrix. (Denition NOM [397], Denition ROM [397]) Verify
Theorem RPNC [398]
Matrix columns: 7 Rank: 3 Nullity: 4
Version 2.30
822 Archetype J
Archetype J
Summary System with six equations, nine variables. Consistent. Null space of coecient matrix has
dimension 5.
A system of linear equations (Denition SLE [11]):
x1+ 2x2 2x3+ 9x4+ 3x5 5x6 2x7+x8+ 27x9= 5
2x1+ 4x2+ 3x3+ 4x4 x5+ 4x6+ 10x7+ 2x8 23x9= 18
x1+ 2x2+x3+ 3x4+x5+x6+ 5x7+ 2x8 7x9= 6
2x1+ 4x2+ 3x3+ 4x4 7x5+ 2x6+ 4x7 11x9= 20
x1+ 2x2+ 5x4+ 2x5 4x6+ 3x7+ 8x8+ 13x9= 4
3x1 6x2 x3 13x4+ 2x5 5x6 4x7+ 13x8+ 10x9= 29
Some solutions to the system of linear equations (not necessarily exhaustive):
x1= 6,x2= 0,x3= 1,x4= 0,x5= 1,x6= 2,x7= 0,x8= 0,x9= 0
x1= 4,x2= 1,x3= 1,x4= 0,x5= 1,x6= 2,x7= 0,x8= 0,x9= 0
x1= 17,x2= 7,x3= 3,x4= 2,x5= 1,x6= 14,x7= 1,x8= 3,x9= 2
x1= 11,x2= 6,x3= 1,x4= 5,x5= 4,x6= 7,x7= 3,x8= 1,x9= 1
Augmented matrix of the linear system of equations (Denition AM [30]):
2
66666641 2 2 9 3 5 2 1 27 5
2 4 3 4 1 4 10 2 23 18
1 2 1 3 1 1 5 2 7 6
2 4 3 4 7 2 4 0 11 20
1 2 0 5 2 4 3 8 13 4
3 6 1 13 2 5 4 13 10 293
7777775
Matrix in reduced row-echelon form, row-equivalent to augmented matrix:
2
6666666412 0 5 0 0 1 2 3 6
0 0 1 2 0 0 3 5 6 1
0 0 0 0 10 1 1 1 1
0 0 0 0 0 10 2 3 2
0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 03
77777775
Version 2.30
Archetype J 823
Analysis of the augmented matrix (Notation RREFA [33]):
r= 4 D=f1;3;5;6g F=f2;4;7;8;9;10g
Vector form of the solution set to the system of equations (Theorem VFSLS [118]). Notice the rela-
tionship between the free variables and the set Fabove. Also, notice the pattern of 0's and 1's in the
entries of the vectors corresponding to elements of the set Ffor the larger examples.
2
6666666666664x1
x2
x3
x4
x5
x6
x7
x8
x93
7777777777775=2
66666666666646
0
1
0
1
2
0
0
03
7777777777775+x22
6666666666664 2
1
0
0
0
0
0
0
03
7777777777775+x42
6666666666664 5
0
2
1
0
0
0
0
03
7777777777775+x72
6666666666664 1
0
3
0
1
0
1
0
03
7777777777775+x82
66666666666642
0
5
0
1
2
0
1
03
7777777777775+x92
6666666666664 3
0
6
0
1
3
0
0
13
7777777777775
Given a system of equations we can always build a new, related, homogeneous system (Denition HS
[71]) by converting the constant terms to zeros and retaining the coecients of the variables. Properties
of this new system will have precise relationships with various properties of the original system.
x1+ 2x2 2x3+ 9x4+ 3x5 5x6 2x7+x8+ 27x9= 0
2x1+ 4x2+ 3x3+ 4x4 x5+ 4x6+ 10x7+ 2x8 23x9= 0
x1+ 2x2+x3+ 3x4+x5+x6+ 5x7+ 2x8 7x9= 0
2x1+ 4x2+ 3x3+ 4x4 7x5+ 2x6+ 4x7 11x9= 0
x1+ 2x2+ +5x4+ 2x5 4x6+ 3x7+ 8x8+ 13x9= 0
3x1 6x2 x3 13x4+ 2x5 5x6 4x7+ 13x8+ 10x9= 0
Some solutions to the associated homogenous system of linear equations (not necessarily exhaustive):
x1= 0,x2= 0,x3= 0,x4= 0,x5= 0,x6= 0,x7= 0,x8= 0,x9= 0
x1= 2,x2= 1,x3= 0,x4= 0,x5= 0,x6= 0,x7= 0,x8= 0,x9= 0
x1= 23,x2= 7,x3= 4,x4= 2,x5= 0,x6= 12,x7= 1,x8= 3,x9= 2
x1= 17,x2= 6,x3= 2,x4= 5,x5= 3,x6= 5,x7= 3,x8= 1,x9= 1
Form the augmented matrix of the homogenous linear system, and use row operations to convert to
Version 2.30
824 Archetype J
reduced row-echelon form. Notice how the entries of the nal column remain zeros:
2
6666666412 0 5 0 0 1 2 3 0
0 0 1 2 0 0 3 5 6 0
0 0 0 0 10 1 1 1 0
0 0 0 0 0 10 2 3 0
0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 03
77777775
Analysis of the augmented matrix for the homogenous system (Notation RREFA [33]). Notice the
slight variation for the same analysis of the original system only when the original system was consistent:
r= 4 D=f1;3;5;6g F=f2;4;7;8;9;10g
Coecient matrix of original system of equations, and of associated homogenous system. This matrix
will be the subject of further analysis, rather than the systems of equations.
2
66666641 2 2 9 3 5 2 1 27
2 4 3 4 1 4 10 2 23
1 2 1 3 1 1 5 2 7
2 4 3 4 7 2 4 0 11
1 2 0 5 2 4 3 8 13
3 6 1 13 2 5 4 13 103
7777775
Matrix brought to reduced row-echelon form:
2
6666666412 0 5 0 0 1 2 3
0 0 1 2 0 0 3 5 6
0 0 0 0 10 1 1 1
0 0 0 0 0 10 2 3
0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 03
77777775
Analysis of the row-reduced matrix (Notation RREFA [33]):
r= 4 D=f1;3;5;6g F=f2;4;7;8;9g
This is the null space of the matrix. The set of vectors used in the span construction is a linearly
independent set of column vectors that spans the null space of the matrix (Theorem SSNS [137], Theorem
BNS [160]). Solve the homogenous system with this matrix as the coecient matrix and write the solutions
in vector form (Theorem VFSLS [118]) to see these vectors arise.
Version 2.30
Archetype J 825
*8
>>>>>>>>>>>><
>>>>>>>>>>>>:2
6666666666664 2
1
0
0
0
0
0
0
03
7777777777775;2
6666666666664 5
0
2
1
0
0
0
0
03
7777777777775;2
6666666666664 1
0
3
0
1
0
1
0
03
7777777777775;2
66666666666642
0
5
0
1
2
0
1
03
7777777777775;2
6666666666664 3
0
6
0
1
3
0
0
13
77777777777759
>>>>>>>>>>>>=
>>>>>>>>>>>>;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors that are
also columns of the matrix. These columns have indices that form the set Dabove. (Theorem BCS [274])
*8
>>>>>><
>>>>>>:2
66666641
2
1
2
1
33
7777775;2
6666664 2
3
1
3
0
13
7777775;2
66666643
1
1
7
2
23
7777775;2
6666664 5
4
1
2
4
53
77777759
>>>>>>=
>>>>>>;+
The column space of the matrix, as it arises from the extended echelon form of the matrix. The matrix
Lis computed as described in Denition EEF [297]. This is followed by the column space described by a
set of linearly independent vectors that span the null space of L, computed as according to Theorem FS
[299] and Theorem BNS [160]. When r=m, the matrix Lhas no rows and the column space is all of Cm.
L=1 0186
13151
131 188
13177
131
0 1 272
131 45
13158
131 14
131
*8
>>>>>><
>>>>>>:2
6666664 77
13114
131
0
0
0
13
7777775;2
6666664188
131
58
131
0
0
1
03
7777775;2
6666664 51
13145
131
0
1
0
03
7777775;2
6666664 186
131272
131
1
0
0
03
77777759
>>>>>>=
>>>>>>;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors. These
vectors are computed by row-reducing the transpose of the matrix into reduced row-echelon form, tossing
out the zero rows, and writing the remaining nonzero rows as column vectors. By Theorem CSRST [282]
and Theorem BRS [280], and in the style of Example CSROI [282], this yields a linearly independent set
of vectors that span the column space.
Version 2.30
826 Archetype J
*8
>>>>>><
>>>>>>:2
66666641
0
0
0
1
29
73
7777775;2
66666640
1
0
0
11
2
94
73
7777775;2
66666640
0
1
0
10
223
7777775;2
66666640
0
0
1
3
2
33
77777759
>>>>>>=
>>>>>>;+
Row space of the matrix, expressed as a span of a set of linearly independent vectors, obtained from
the nonzero rows of the equivalent matrix in reduced row-echelon form. (Theorem BRS [280])
*8
>>>>>>>>>>>><
>>>>>>>>>>>>:2
66666666666641
2
0
5
0
0
1
2
33
7777777777775;2
66666666666640
0
1
2
0
0
3
5
63
7777777777775;2
66666666666640
0
0
0
1
0
1
1
13
7777777777775;2
66666666666640
0
0
0
0
1
0
2
33
77777777777759
>>>>>>>>>>>>=
>>>>>>>>>>>>;+
Subspace dimensions associated with the matrix. (Denition NOM [397], Denition ROM [397]) Verify
Theorem RPNC [398]
Matrix columns: 9 Rank: 4 Nullity: 5
Version 2.30
Archetype K 827
Archetype K
Summary Square matrix of size 5. Nonsingular. 3 distinct eigenvalues, 2 of multiplicity 2.
A matrix:
2
6666410 18 24 24 12
12 2 6 0 18
30 21 23 30 39
27 30 36 37 30
18 24 30 30 203
77775
Matrix brought to reduced row-echelon form:
2
66666410 0 0 0
010 0 0
0 0 10 0
0 0 0 10
0 0 0 0 13
777775
Analysis of the row-reduced matrix (Notation RREFA [33]):
r= 5 D=f1;2;3;4;5g F=fg
Matrix (coecient matrix) is nonsingular or singular? (Theorem NMRRI [84]) at the same time,
examine the size of the set Fabove.Notice that this property does not apply to matrices that are not
square.
Nonsingular.
This is the null space of the matrix. The set of vectors used in the span construction is a linearly
independent set of column vectors that spans the null space of the matrix (Theorem SSNS [137], Theorem
BNS [160]). Solve the homogenous system with this matrix as the coecient matrix and write the solutions
in vector form (Theorem VFSLS [118]) to see these vectors arise.
hfgi
Column space of the matrix, expressed as the span of a set of linearly independent vectors that are
also columns of the matrix. These columns have indices that form the set Dabove. (Theorem BCS [274])
Version 2.30
828 Archetype K
*8
>>>><
>>>>:2
6666410
12
30
27
183
77775;2
6666418
2
21
30
243
77775;2
6666424
6
23
36
303
77775;2
6666424
0
30
37
303
77775;2
66664 12
18
39
30
203
777759
>>>>=
>>>>;+
The column space of the matrix, as it arises from the extended echelon form of the matrix. The matrix
Lis computed as described in Denition EEF [297]. This is followed by the column space described by a
set of linearly independent vectors that span the null space of L, computed as according to Theorem FS
[299] and Theorem BNS [160]. When r=m, the matrix Lhas no rows and the column space is all of Cm.
L=
*8
>>>><
>>>>:2
666641
0
0
0
03
77775;2
666640
1
0
0
03
77775;2
666640
0
1
0
03
77775;2
666640
0
0
1
03
77775;2
666640
0
0
0
13
777759
>>>>=
>>>>;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors. These
vectors are computed by row-reducing the transpose of the matrix into reduced row-echelon form, tossing
out the zero rows, and writing the remaining nonzero rows as column vectors. By Theorem CSRST [282]
and Theorem BRS [280], and in the style of Example CSROI [282], this yields a linearly independent set
of vectors that span the column space.
*8
>>>><
>>>>:2
666641
0
0
0
03
77775;2
666640
1
0
0
03
77775;2
666640
0
1
0
03
77775;2
666640
0
0
1
03
77775;2
666640
0
0
0
13
777759
>>>>=
>>>>;+
Row space of the matrix, expressed as a span of a set of linearly independent vectors, obtained from
the nonzero rows of the equivalent matrix in reduced row-echelon form. (Theorem BRS [280])
*8
>>>><
>>>>:2
666641
0
0
0
03
77775;2
666640
1
0
0
03
77775;2
666640
0
1
0
03
77775;2
666640
0
0
1
03
77775;2
666640
0
0
0
13
777759
>>>>=
>>>>;+
Inverse matrix, if it exists. The inverse is not dened for matrices that are not square, and if the matrix
is square, then the matrix must be nonsingular. (Denition MI [244], Theorem NI [261])
Version 2.30
Archetype K 829
2
666641 9
4
3
2
3 6
21
243
421
29 9
15 21
2
11 1539
2
915
49
210 15
9
23
43
26 19
23
77775
Subspace dimensions associated with the matrix. (Denition NOM [397], Denition ROM [397]) Verify
Theorem RPNC [398]
Matrix columns: 5 Rank: 5 Nullity: 0
Determinant of the matrix, which is only dened for square matrices. The matrix is nonsingular if and
only if the determinant is nonzero (Theorem SMZD [445]). (Product of all eigenvalues?)
Determinant = 16
Eigenvalues, and bases for eigenspaces. (Denition EEM [453],Denition EM [461])
= 2 EK( 2) =*8
>>>><
>>>>:2
666642
2
1
0
13
77775;2
66664 1
2
2
1
03
777759
>>>>=
>>>>;+
= 1 EK(1) =*8
>>>><
>>>>:2
666644
10
7
0
23
77775;2
66664 4
18
17
5
03
777759
>>>>=
>>>>;+
= 4 EK(4) =*8
>>>><
>>>>:2
666641
1
0
1
13
777759
>>>>=
>>>>;+
Geometric and algebraic multiplicities. (Denition GME [463]Denition AME [463])
K( 2) = 2 K( 2) = 2
K(1) = 2 K(1) = 2
K(4) = 1 K(4) = 1
Diagonalizable? (Denition DZM [496])
Version 2.30
830 Archetype K
Yes, full eigenspaces, Theorem DMFE [499].
The diagonalization. (Theorem DC [497])
2
66664 4 3 4 6 7
7 5 6 8 10
1 1 1 1 3
1 0 0 1 2
2 5 6 4 03
777752
6666410 18 24 24 12
12 2 6 0 18
30 21 23 30 39
27 30 36 37 30
18 24 30 30 203
777752
666642 1 4 4 1
2 2 10 18 1
1 2 7 17 0
0 1 0 5 1
1 0 2 0 13
77775
=2
66664 2 0 0 0 0
0 2 0 0 0
0 0 1 0 0
0 0 0 1 0
0 0 0 0 43
77775
Version 2.30
Archetype L 831
Archetype L
Summary Square matrix of size 5. Singular, nullity 2. 2 distinct eigenvalues, each of \high" multiplicity.
A matrix:
2
66664 2 1 2 4 4
6 5 4 4 6
10 7 7 10 13
7 5 6 9 10
4 3 4 6 63
77775
Matrix brought to reduced row-echelon form:
2
66666410 0 1 2
010 2 2
0 0 1 2 1
0 0 0 0 0
0 0 0 0 03
777775
Analysis of the row-reduced matrix (Notation RREFA [33]):
r= 5 D=f1;2;3g F=f4;5g
Matrix (coecient matrix) is nonsingular or singular? (Theorem NMRRI [84]) at the same time,
examine the size of the set Fabove.Notice that this property does not apply to matrices that are not
square.
Singular.
This is the null space of the matrix. The set of vectors used in the span construction is a linearly
independent set of column vectors that spans the null space of the matrix (Theorem SSNS [137], Theorem
BNS [160]). Solve the homogenous system with this matrix as the coecient matrix and write the solutions
in vector form (Theorem VFSLS [118]) to see these vectors arise.
*8
>>>><
>>>>:2
66664 1
2
2
1
03
77775;2
666642
2
1
0
13
777759
>>>>=
>>>>;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors that are
Version 2.30
832 Archetype L
also columns of the matrix. These columns have indices that form the set Dabove. (Theorem BCS [274])
*8
>>>><
>>>>:2
66664 2
6
10
7
43
77775;2
66664 1
5
7
5
33
77775;2
66664 2
4
7
6
43
777759
>>>>=
>>>>;+
The column space of the matrix, as it arises from the extended echelon form of the matrix. The matrix
Lis computed as described in Denition EEF [297]. This is followed by the column space described by a
set of linearly independent vectors that span the null space of L, computed as according to Theorem FS
[299] and Theorem BNS [160]. When r=m, the matrix Lhas no rows and the column space is all of Cm.
L=1 0 2 6 5
0 1 4 10 9
*8
>>>><
>>>>:2
66664 5
9
0
0
13
77775;2
666646
10
0
1
03
77775;2
666642
4
1
0
03
777759
>>>>=
>>>>;+
Column space of the matrix, expressed as the span of a set of linearly independent vectors. These
vectors are computed by row-reducing the transpose of the matrix into reduced row-echelon form, tossing
out the zero rows, and writing the remaining nonzero rows as column vectors. By Theorem CSRST [282]
and Theorem BRS [280], and in the style of Example CSROI [282], this yields a linearly independent set
of vectors that span the column space.
*8
>>>><
>>>>:2
666641
0
0
9
45
23
77775;2
666640
1
0
5
43
23
77775;2
666640
0
1
1
2
13
777759
>>>>=
>>>>;+
Row space of the matrix, expressed as a span of a set of linearly independent vectors, obtained from
the nonzero rows of the equivalent matrix in reduced row-echelon form. (Theorem BRS [280])
*8
>>>><
>>>>:2
666641
0
0
1
23
77775;2
666640
1
0
2
23
77775;2
666640
0
1
2
13
777759
>>>>=
>>>>;+
Inverse matrix, if it exists. The inverse is not dened for matrices that are not square, and if the matrix
is square, then the matrix must be nonsingular. (Denition MI [244], Theorem NI [261])
Version 2.30
Archetype L 833
Subspace dimensions associated with the matrix. (Denition NOM [397], Denition ROM [397]) Verify
Theorem RPNC [398]
Matrix columns: 5 Rank: 3 Nullity: 2
Determinant of the matrix, which is only dened for square matrices. The matrix is nonsingular if and
only if the determinant is nonzero (Theorem SMZD [445]). (Product of all eigenvalues?)
Determinant = 0
Eigenvalues, and bases for eigenspaces. (Denition EEM [453],Denition EM [461])
= 1 EL( 1) =*8
>>>><
>>>>:2
66664 5
9
0
0
13
77775;2
666646
10
0
1
03
77775;2
666642
4
1
0
03
777759
>>>>=
>>>>;+
= 0 EL(0) =*8
>>>><
>>>>:2
666642
2
1
0
13
77775;2
66664 1
2
2
1
03
777759
>>>>=
>>>>;+
Geometric and algebraic multiplicities. (Denition GME [463]Denition AME [463])
L( 1) = 3 L( 1) = 3
L(0) = 2 L(0) = 2
Diagonalizable? (Denition DZM [496])
Yes, full eigenspaces, Theorem DMFE [499].
The diagonalization. (Theorem DC [497])
2
666644 3 4 6 6
7 5 6 9 10
10 7 7 10 13
4 3 4 6 7
7 5 6 8 103
777752
66664 2 1 2 4 4
6 5 4 4 6
10 7 7 10 13
7 5 6 9 10
4 3 4 6 63
777752
66664 5 6 2 2 1
9 10 4 2 2
0 0 1 1 2
0 1 0 0 1
1 0 0 1 03
77775
Version 2.30
834 Archetype L
=2
66664 1 0 0 0 0
0 1 0 0 0
0 0 1 0 0
0 0 0 0 0
0 0 0 0 03
77775
Version 2.30
Archetype M 835
Archetype M
Summary Linear transformation with bigger domain than codomain, so it is guaranteed to not be
injective. Happens to not be surjective.
A linear transformation: (Denition LT [515])
T:C5!C3; T0
BBBB@2
66664x1
x2
x3
x4
x53
777751
CCCCA=2
4x1+ 2x2+ 3x3+ 4x4+ 4x5
3x1+x2+ 4x3 3x4+ 7x5
x1 x2 5x4+x53
5
A basis for the null space of the linear transformation: (Denition KLT [545])
8
>>>><
>>>>:2
66664 2
1
0
0
13
77775;2
666642
3
0
1
03
77775;2
66664 1
1
1
0
03
777759
>>>>=
>>>>;
Injective: No. (Denition ILT [541])
Since the kernel is nontrivial Theorem KILT [548] tells us that the linear transformation is not injective.
Also, since the rank can not exceed 3, we are guaranteed to have a nullity of at least 2, just from checking
dimensions of the domain and the codomain. In particular, verify that
T0
BBBB@2
666641
2
1
4
53
777751
CCCCA=2
438
24
163
5 T0
BBBB@2
666640
3
0
5
63
777751
CCCCA=2
438
24
163
5
This demonstration that Tis not injective is constructed with the observation that
2
666640
3
0
5
63
77775=2
666641
2
1
4
53
77775+2
66664 1
5
1
1
13
77775
Version 2.30
836 Archetype M
and
z=2
66664 1
5
1
1
13
777752K(T)
so the vector zeectively \does nothing" in the evaluation of T.
A basis for the range of the linear transformation: (Denition RLT [563])
Evaluate the linear transformation on a standard basis to get a spanning set for the range (Theorem SSRLT
[567]):
8
<
:2
41
3
13
5;2
42
1
13
5;2
43
4
03
5;2
44
3
53
5;2
44
7
13
59
=
;
If the linear transformation is injective, then the set above is guaranteed to be linearly independent (The-
orem ILTLI [549]). This spanning set may be converted to a \nice" basis, by making the vectors the rows
of a matrix (perhaps after using a vector reperesentation), row-reducing, and retaining the nonzero rows
(Theorem BRS [280]), and perhaps un-coordinatizing. A basis for the range is:
8
<
:2
41
0
4
53
5;2
40
1
3
53
59
=
;
Surjective: No. (Denition SLT [559])
Notice that the range is not all of C3since its dimension 2, not 3. In particular, verify that2
43
4
53
562R(T),
by setting the output equal to this vector and seeing that the resulting system of linear equations has no
solution, i.e. is inconsistent. So the preimage, T 10
@2
43
4
53
51
A, is empty. This alone is sucient to see that
the linear transformation is not onto.
Subspace dimensions associated with the linear transformation. Examine parallels with earlier results
for matrices. Verify Theorem RPNDD [588].
Domain dimension: 5 Rank: 2 Nullity: 3
Invertible: No.
Version 2.30
Archetype M 837
Not injective or surjective.
Matrix representation (Theorem MLTCV [523]):
T:C5!C3; T (x) =Ax; A =2
41 2 3 4 4
3 1 4 3 7
1 1 0 5 13
5
Version 2.30
838 Archetype N
Archetype N
Summary Linear transformation with domain larger than its codomain, so it is guaranteed to not be
injective. Happens to be onto.
A linear transformation: (Denition LT [515])
T:C5!C3; T0
BBBB@2
66664x1
x2
x3
x4
x53
777751
CCCCA=2
42x1+x2+ 3x3 4x4+ 5x5
x1 2x2+ 3x3 9x4+ 3x5
3x1+ 4x3 6x4+ 5x53
5
A basis for the null space of the linear transformation: (Denition KLT [545])
8
>>>><
>>>>:2
666641
1
2
0
13
77775;2
66664 2
1
3
1
03
777759
>>>>=
>>>>;
Injective: No. (Denition ILT [541])
Since the kernel is nontrivial Theorem KILT [548] tells us that the linear transformation is not injective.
Also, since the rank can not exceed 3, we are guaranteed to have a nullity of at least 2, just from checking
dimensions of the domain and the codomain. In particular, verify that
T0
BBBB@2
66664 3
1
2
3
13
777751
CCCCA=2
46
19
63
5 T0
BBBB@2
66664 4
4
2
1
43
777751
CCCCA=2
46
19
63
5
This demonstration that Tis not injective is constructed with the observation that
2
66664 4
4
2
1
43
77775=2
66664 3
1
2
3
13
77775+2
66664 1
5
0
2
33
77775
Version 2.30
Archetype N 839
and
z=2
66664 1
5
0
2
33
777752K(T)
so the vector zeectively \does nothing" in the evaluation of T.
A basis for the range of the linear transformation: (Denition RLT [563])
Evaluate the linear transformation on a standard basis to get a spanning set for the range (Theorem SSRLT
[567]):
8
<
:2
42
1
33
5;2
41
2
03
5;2
43
3
43
5;2
4 4
9
63
5;2
45
3
53
59
=
;
If the linear transformation is injective, then the set above is guaranteed to be linearly independent (The-
orem ILTLI [549]). This spanning set may be converted to a \nice" basis, by making the vectors the rows
of a matrix (perhaps after using a vector reperesentation), row-reducing, and retaining the nonzero rows
(Theorem BRS [280]), and perhaps un-coordinatizing. A basis for the range is:
8
<
:2
41
0
03
5;2
40
1
03
5;2
40
0
13
59
=
;
Surjective: Yes. (Denition SLT [559])
Notice that the basis for the range above is the standard basis for C3. So the range is all of C3and thus
the linear transformation is surjective.
Subspace dimensions associated with the linear transformation. Examine parallels with earlier results
for matrices. Verify Theorem RPNDD [588].
Domain dimension: 5 Rank: 3 Nullity: 2
Invertible: No.
Not surjective, and the relative sizes of the domain and codomain mean the linear transformation cannot
be injective. (Theorem ILTIS [582])
Matrix representation (Theorem MLTCV [523]):
T:C5!C3; T (x) =Ax; A =2
42 1 3 4 5
1 2 3 9 3
3 0 4 6 53
5
Version 2.30
840 Archetype N
Version 2.30
Archetype O 841
Archetype O
Summary Linear transformation with a domain smaller than the codomain, so it is guaranteed to not
be onto. Happens to not be one-to-one.
A linear transformation: (Denition LT [515])
T:C3!C5; T0
@2
4x1
x2
x33
51
A=2
66664 x1+x2 3x3
x1+ 2x2 4x3
x1+x2+x3
2x1+ 3x2+x3
x1+ 2x33
77775
A basis for the null space of the linear transformation: (Denition KLT [545])
8
<
:2
4 2
1
13
59
=
;
Injective: No. (Denition ILT [541])
Since the kernel is nontrivial Theorem KILT [548] tells us that the linear transformation is not injective.
Also, since the rank can not exceed 3, we are guaranteed to have a nullity of at least 2, just from checking
dimensions of the domain and the codomain. In particular, verify that
T0
@2
45
1
33
51
A=2
66664 15
19
7
10
113
77775T0
@2
41
1
53
51
A=2
66664 15
19
7
10
113
77775
This demonstration that Tis not injective is constructed with the observation that
2
41
1
53
5=2
45
1
33
5+2
4 4
2
23
5
and
z=2
4 4
2
23
52K(T)
so the vector zeectively \does nothing" in the evaluation of T.
A basis for the range of the linear transformation: (Denition RLT [563])
Version 2.30
842 Archetype O
Evaluate the linear transformation on a standard basis to get a spanning set for the range (Theorem SSRLT
[567]):
8
>>>><
>>>>:2
66664 1
1
1
2
13
77775;2
666641
2
1
3
03
77775;2
66664 3
4
1
1
23
777759
>>>>=
>>>>;
If the linear transformation is injective, then the set above is guaranteed to be linearly independent (The-
orem ILTLI [549]). This spanning set may be converted to a \nice" basis, by making the vectors the rows
of a matrix (perhaps after using a vector reperesentation), row-reducing, and retaining the nonzero rows
(Theorem BRS [280]), and perhaps un-coordinatizing. A basis for the range is:
8
>>>><
>>>>:2
666641
0
3
7
23
77775;2
666640
1
2
5
13
777759
>>>>=
>>>>;
Subspace dimensions associated with the linear transformation. Examine parallels with earlier results
for matrices. Verify Theorem RPNDD [588].
Domain dimension: 3 Rank: 2 Nullity: 1
Surjective: No. (Denition SLT [559])
The dimension of the range is 2, and the codomain ( C5) has dimension 5. So the transformation is not onto.
Notice too that since the domain C3has dimension 3, it is impossible for the range to have a dimension
greater than 3, and no matter what the actual denition of the function, it cannot possibly be onto.
To be more precise, verify that2
666642
3
1
1
13
7777562R(T), by setting the output equal to this vector and seeing that
the resulting system of linear equations has no solution, i.e. is inconsistent. So the preimage, T 10
BBBB@2
666642
3
1
1
13
777751
CCCCA,
is empty. This alone is sucient to see that the linear transformation is not onto.
Invertible: No.
Not injective, and the relative dimensions of the domain and codomain prohibit any possibility of being
surjective.
Matrix representation (Theorem MLTCV [523]):
Version 2.30
Archetype O 843
T:C3!C5; T (x) =Ax; A =2
66664 1 1 3
1 2 4
1 1 1
2 3 1
1 0 23
77775
Version 2.30
844 Archetype P
Archetype P
Summary Linear transformation with a domain smaller that its codomain, so it is guaranteed to not
be surjective. Happens to be injective.
A linear transformation: (Denition LT [515])
T:C3!C5; T0
@2
4x1
x2
x33
51
A=2
66664 x1+x2+x3
x1+ 2x2+ 2x3
x1+x2+ 3x3
2x1+ 3x2+x3
2x1+x2+ 3x33
77775
A basis for the null space of the linear transformation: (Denition KLT [545])
fg
Injective: Yes. (Denition ILT [541])
SinceK(T) =f0g, Theorem KILT [548] tells us that Tis injective.
A basis for the range of the linear transformation: (Denition RLT [563])
Evaluate the linear transformation on a standard basis to get a spanning set for the range (Theorem SSRLT
[567]):
8
>>>><
>>>>:2
66664 1
1
1
2
23
77775;2
666641
2
1
3
13
77775;2
666641
2
3
1
33
777759
>>>>=
>>>>;
If the linear transformation is injective, then the set above is guaranteed to be linearly independent (The-
orem ILTLI [549]). This spanning set may be converted to a \nice" basis, by making the vectors the rows
of a matrix (perhaps after using a vector reperesentation), row-reducing, and retaining the nonzero rows
(Theorem BRS [280]), and perhaps un-coordinatizing. A basis for the range is:
8
>>>><
>>>>:2
666641
0
0
10
63
77775;2
666640
1
0
7
33
77775;2
666640
0
1
1
13
777759
>>>>=
>>>>;
Version 2.30
Archetype P 845
Surjective: No. (Denition SLT [559])
The dimension of the range is 3, and the codomain ( C5) has dimension 5. So the transformation is not
surjective. Notice too that since the domain C3has dimension 3, it is impossible for the range to have a
dimension greater than 3, and no matter what the actual denition of the function, it cannot possibly be
surjective in this situation.
To be more precise, verify that2
666642
1
3
2
63
7777562R(T), by setting the output equal to this vector and seeing that
the resulting system of linear equations has no solution, i.e. is inconsistent. So the preimage, T 10
BBBB@2
666642
1
3
2
63
777751
CCCCA,
is empty. This alone is sucient to see that the linear transformation is not onto.
Subspace dimensions associated with the linear transformation. Examine parallels with earlier results
for matrices. Verify Theorem RPNDD [588].
Domain dimension: 3 Rank: 3 Nullity: 0
Invertible: No.
The relative dimensions of the domain and codomain prohibit any possibility of being surjective, so apply
Theorem ILTIS [582].
Matrix representation (Theorem MLTCV [523]):
T:C3!C5; T (x) =Ax; A =2
66664 1 1 1
1 2 2
1 1 3
2 3 1
2 1 33
77775
Version 2.30
846 Archetype Q
Archetype Q
Summary Linear transformation with equal-sized domain and codomain, so it has the potential to be
invertible, but in this case is not. Neither injective nor surjective. Diagonalizable, though.
A linear transformation: (Denition LT [515])
T:C5!C5; T0
BBBB@2
66664x1
x2
x3
x4
x53
777751
CCCCA=2
66664 2x1+ 3x2+ 3x3 6x4+ 3x5
16x1+ 9x2+ 12x3 28x4+ 28x5
19x1+ 7x2+ 14x3 32x4+ 37x5
21x1+ 9x2+ 15x3 35x4+ 39x5
9x1+ 5x2+ 7x3 16x4+ 16x53
77775
A basis for the null space of the linear transformation: (Denition KLT [545])
8
>>>><
>>>>:2
666643
4
1
3
33
777759
>>>>=
>>>>;
Injective: No. (Denition ILT [541])
Since the kernel is nontrivial Theorem KILT [548] tells us that the linear transformation is not injective.
Also, since the rank can not exceed 3, we are guaranteed to have a nullity of at least 2, just from checking
dimensions of the domain and the codomain. In particular, verify that
T0
BBBB@2
666641
3
1
2
43
777751
CCCCA=2
666644
55
72
77
313
77775T0
BBBB@2
666644
7
0
5
73
777751
CCCCA=2
666644
55
72
77
313
77775
This demonstration that Tis not injective is constructed with the observation that
2
666644
7
0
5
73
77775=2
666641
3
1
2
43
77775+2
666643
4
1
3
33
77775
Version 2.30
Archetype Q 847
and
z=2
666643
4
1
3
33
777752K(T)
so the vector zeectively \does nothing" in the evaluation of T.
A basis for the range of the linear transformation: (Denition RLT [563])
Evaluate the linear transformation on a standard basis to get a spanning set for the range (Theorem SSRLT
[567]):
8
>>>><
>>>>:2
66664 2
16
19
21
93
77775;2
666643
9
7
9
53
77775;2
666643
12
14
15
73
77775;2
66664 6
28
32
35
163
77775;2
666643
28
37
39
163
777759
>>>>=
>>>>;
If the linear transformation is injective, then the set above is guaranteed to be linearly independent (The-
orem ILTLI [549]). This spanning set may be converted to a \nice" basis, by making the vectors the rows
of a matrix (perhaps after using a vector reperesentation), row-reducing, and retaining the nonzero rows
(Theorem BRS [280]), and perhaps un-coordinatizing. A basis for the range is:
8
>>>><
>>>>:2
666641
0
0
0
13
77775;2
666640
1
0
0
13
77775;2
666640
0
1
0
13
77775;2
666640
0
0
1
23
777759
>>>>=
>>>>;
Surjective: No. (Denition SLT [559])
The dimension of the range is 4, and the codomain ( C5) has dimension 5. So R(T)6=C5and by Theorem
RSLT [565] the transformation is not surjective.
To be more precise, verify that2
66664 1
2
3
1
43
7777562R(T), by setting the output equal to this vector and seeing that
the resulting system of linear equations has no solution, i.e. is inconsistent. So the preimage, T 10
BBBB@2
66664 1
2
3
1
43
777751
CCCCA,
is empty. This alone is sucient to see that the linear transformation is not onto.
Subspace dimensions associated with the linear transformation. Examine parallels with earlier results
Version 2.30
848 Archetype Q
for matrices. Verify Theorem RPNDD [588].
Domain dimension: 5 Rank: 4 Nullity: 1
Invertible: No.
Neither injective nor surjective. Notice that since the domain and codomain have the same dimension,
either the transformation is both onto and one-to-one (making it invertible) or else it is both not onto and
not one-to-one (as in this case) by Theorem RPNDD [588].
Matrix representation (Theorem MLTCV [523]):
T:C5!C5; T (x) =Ax; A =2
66664 2 3 3 6 3
16 9 12 28 28
19 7 14 32 37
21 9 15 35 39
9 5 7 16 163
77775
Eigenvalues and eigenvectors (Denition EELT [647], Theorem EER [659]):
= 1 ET( 1) =*8
>>>><
>>>>:2
666640
2
3
3
13
777759
>>>>=
>>>>;+
= 0 ET(0) =*8
>>>><
>>>>:2
666643
4
1
3
33
777759
>>>>=
>>>>;+
= 1 ET(1) =*8
>>>><
>>>>:2
666645
3
0
0
23
77775;2
66664 3
1
0
2
03
77775;2
666641
1
2
0
03
777759
>>>>=
>>>>;+
Evaluate the linear transformation with each of these eigenvectors as an interesting check.
A diagonal matrix representation relative to a basis of eigenvectors, B.
B=8
>>>><
>>>>:2
666640
2
3
3
13
77775;2
666643
4
1
3
33
77775;2
666645
3
0
0
23
77775;2
66664 3
1
0
2
03
77775;2
666641
1
2
0
03
777759
>>>>=
>>>>;
Version 2.30
Archetype Q 849
MT
B;B=2
66664 1 0 0 0 0
0 0 0 0 0
0 0 1 0 0
0 0 0 1 0
0 0 0 0 13
77775
Version 2.30
850 Archetype R
Archetype R
Summary Linear transformation with equal-sized domain and codomain. Injective, surjective, invert-
ible, diagonalizable, the works.
A linear transformation: (Denition LT [515])
T:C5!C5; T0
BBBB@2
66664x1
x2
x3
x4
x53
777751
CCCCA=2
66664 65x1+ 128x2+ 10x3 262x4+ 40x5
36x1 73x2 x3+ 151x4 16x5
44x1+ 88x2+ 5x3 180x4+ 24x5
34x1 68x2 3x3+ 140x4 18x5
12x1 24x2 x3+ 49x4 5x53
77775
A basis for the null space of the linear transformation: (Denition KLT [545])
fg
Injective: Yes. (Denition ILT [541])
Since the kernel is trivial Theorem KILT [548] tells us that the linear transformation is injective.
A basis for the range of the linear transformation: (Denition RLT [563])
Evaluate the linear transformation on a standard basis to get a spanning set for the range (Theorem SSRLT
[567]):
8
>>>><
>>>>:2
66664 65
36
44
34
123
77775;2
66664128
73
88
68
243
77775;2
6666410
1
5
3
13
77775;2
66664 262
151
180
140
493
77775;2
6666440
16
24
18
53
777759
>>>>=
>>>>;
If the linear transformation is injective, then the set above is guaranteed to be linearly independent (The-
orem ILTLI [549]). This spanning set may be converted to a \nice" basis, by making the vectors the rows
of a matrix (perhaps after using a vector reperesentation), row-reducing, and retaining the nonzero rows
(Theorem BRS [280]), and perhaps un-coordinatizing. A basis for the range is:
8
>>>><
>>>>:2
666641
0
0
0
03
77775;2
666640
1
0
0
03
77775;2
666640
0
1
0
03
77775;2
666640
0
0
1
03
77775;2
666640
0
0
0
13
777759
>>>>=
>>>>;
Version 2.30
Archetype R 851
Surjective: Yes. (Denition SLT [559])
A basis for the range is the standard basis of C5, soR(T) =C5and Theorem RSLT [565] tells us T
is surjective. Or, the dimension of the range is 5, and the codomain ( C5) has dimension 5. So the
transformation is surjective.
Subspace dimensions associated with the linear transformation. Examine parallels with earlier results
for matrices. Verify Theorem RPNDD [588].
Domain dimension: 5 Rank: 5 Nullity: 0
Invertible: Yes.
Both injective and surjective (Theorem ILTIS [582]). Notice that since the domain and codomain have the
same dimension, either the transformation is both injective and surjective (making it invertible, as in this
case) or else it is both not injective and not surjective.
Matrix representation (Theorem MLTCV [523]):
T:C5!C5; T (x) =Ax; A =2
66664 65 128 10 262 40
36 73 1 151 16
44 88 5 180 24
34 68 3 140 18
12 24 1 49 53
77775
The inverse linear transformation (Denition IVLT [579]):
T 1:C5!C5; T 10
BBBB@2
66664x1
x2
x3
x4
x53
777751
CCCCA=2
66664 47x1+ 92x2+x3 181x4 14x5
27x1 55x2+7
2x3+221
2x4+ 11x5
32x1+ 64x2 x3 126x4 12x5
25x1 50x2+3
2x3+199
2x4+ 9x5
9x1 18x2+1
2x3+71
2x4+ 4x53
77775
Verify that T
T 1(x)
=xandT
T 1(x)
=x, and notice that the representations of the transformation
and its inverse are matrix inverses (Theorem IMR [630], Denition MI [244]).
Eigenvalues and eigenvectors (Denition EELT [647], Theorem EER [659]):
= 1 ET( 1) =*8
>>>><
>>>>:2
66664 57
0
18
14
53
77775;2
666642
1
0
0
03
777759
>>>>=
>>>>;+
Version 2.30
852 Archetype R
= 1 ET(1) =*8
>>>><
>>>>:2
66664 10
5
6
0
13
77775;2
666642
3
1
1
03
777759
>>>>=
>>>>;+
= 2 ET(2) =*8
>>>><
>>>>:2
66664 6
3
4
3
13
777759
>>>>=
>>>>;+
Evaluate the linear transformation with each of these eigenvectors as an interesting check.
A diagonal matrix representation relative to a basis of eigenvectors, B.
B=8
>>>><
>>>>:2
66664 57
0
18
14
53
77775;2
666642
1
0
0
03
77775;2
66664 10
5
6
0
13
77775;2
666642
3
1
1
03
77775;2
66664 6
3
4
3
13
777759
>>>>=
>>>>;
MT
B;B=2
66664 1 0 0 0 0
0 1 0 0 0
0 0 1 0 0
0 0 0 1 0
0 0 0 0 23
77775
Version 2.30
Archetype S 853
Archetype S
Summary Domain is column vectors, codomain is matrices. Domain is dimension 3 and codomain is
dimension 4. Not injective, not surjective.
A linear transformation: (Denition LT [515])
T:C3!M22; T0
@2
4a
b
c3
51
A=a b 2a+ 2b+c
3a+b+c 2a 6b 2c
A basis for the null space of the linear transformation: (Denition KLT [545])
8
<
:2
4 1
1
43
59
=
;
Injective: No. (Denition ILT [541])
Since the kernel is nontrivial Theorem KILT [548] tells us that the linear transformation is not injective.
Also, since the rank can not exceed 3, we are guaranteed to have a nullity of at least 1, just from checking
dimensions of the domain and the codomain. In particular, verify that
T0
@2
42
1
33
51
A=1 9
10 16
T0
@2
40
1
113
51
A=1 9
10 16
This demonstration that Tis not injective is constructed with the observation that
2
40
1
113
5=2
42
1
33
5+2
4 2
2
83
5
and
z=2
4 2
2
83
52K(T)
so the vector zeectively \does nothing" in the evaluation of T.
A basis for the range of the linear transformation: (Denition RLT [563])
Evaluate the linear transformation on a standard basis to get a spanning set for the range (Theorem SSRLT
Version 2.30
854 Archetype S
[567]):
1 2
3 2
; 1 2
1 6
;0 1
1 2
If the linear transformation is injective, then the set above is guaranteed to be linearly independent (The-
orem ILTLI [549]). This spanning set may be converted to a \nice" basis, by making the vectors the rows
of a matrix (perhaps after using a vector reperesentation), row-reducing, and retaining the nonzero rows
(Theorem BRS [280]), and perhaps un-coordinatizing. A basis for the range is:
1 0
1 2
;0 1
1 2
Surjective: No. (Denition SLT [559])
The dimension of the range is 2, and the codomain ( M22) has dimension 4. So the transformation is not
surjective. Notice too that since the domain C3has dimension 3, it is impossible for the range to have a
dimension greater than 3, and no matter what the actual denition of the function, it cannot possibly be
surjective in this situation.
To be more precise, verify that2 1
1 3
62R(T), by setting the output of Tequal to this matrix and
seeing that the resulting system of linear equations has no solution, i.e. is inconsistent. So the preimage,
T 12 1
1 3
, is empty. This alone is sucient to see that the linear transformation is not onto.
Subspace dimensions associated with the linear transformation. Examine parallels with earlier results
for matrices. Verify Theorem RPNDD [588].
Domain dimension: 3 Rank: 2 Nullity: 1
Invertible: No.
Not injective (Theorem ILTIS [582]), and the relative dimensions of the domain and codomain prohibit
any possibility of being surjective.
Matrix representation (Denition MR [615]):
B=8
<
:2
41
0
03
5;2
40
1
03
5;2
40
0
13
59
=
;
C=1 0
0 0
;0 1
0 0
;0 0
1 0
;0 0
0 1
MT
B;C=2
6641 1 0
2 2 1
3 1 1
2 6 23
775
Version 2.30
Archetype S 855
Version 2.30
856 Archetype T
Archetype T
Summary Domain and codomain are polynomials. Domain has dimension 5, while codomain has
dimension 6. Is injective, can't be surjective.
A linear transformation: (Denition LT [515])
T:P4!P5; T (p(x)) = (x 2)p(x)
A basis for the null space of the linear transformation: (Denition KLT [545])
fg
Injective: Yes. (Denition ILT [541])
Since the kernel is trivial Theorem KILT [548] tells us that the linear transformation is injective.
A basis for the range of the linear transformation: (Denition RLT [563])
Evaluate the linear transformation on a standard basis to get a spanning set for the range (Theorem SSRLT
[567]):
x 2; x2 2x; x3 2x2; x4 2x3;x5 2x4;x6 2x5
If the linear transformation is injective, then the set above is guaranteed to be linearly independent (The-
orem ILTLI [549]). This spanning set may be converted to a \nice" basis, by making the vectors the rows
of a matrix (perhaps after using a vector reperesentation), row-reducing, and retaining the nonzero rows
(Theorem BRS [280]), and perhaps un-coordinatizing. A basis for the range is:
1
32x5+ 1; 1
16x5+x; 1
8x5+x2; 1
4x5+x3; 1
2x5+x4
Surjective: No. (Denition SLT [559])
The dimension of the range is 5, and the codomain ( P5) has dimension 6. So the transformation is not
surjective. Notice too that since the domain P4has dimension 5, it is impossible for the range to have a
dimension greater than 5, and no matter what the actual denition of the function, it cannot possibly be
surjective in this situation.
To be more precise, verify that 1+ x+x2+x3+x462R(T), by setting the output equal to this vector and
seeing that the resulting system of linear equations has no solution, i.e. is inconsistent. So the preimage,
Version 2.30
Archetype T 857
T 1
1 +x+x2+x3+x4
, is nonempty. This alone is sucient to see that the linear transformation is
not onto.
Subspace dimensions associated with the linear transformation. Examine parallels with earlier results
for matrices. Verify Theorem RPNDD [588].
Domain dimension: 5 Rank: 5 Nullity: 0
Invertible: No.
The relative dimensions of the domain and codomain prohibit any possibility of being surjective, so apply
Theorem ILTIS [582].
Matrix representation (Denition MR [615]):
B=
1; x; x2; x3; x4
C=
1; x; x2; x3; x4; x5
MT
B;C=2
6666664 2 0 0 0 0
1 2 0 0 0
0 1 2 0 0
0 0 1 2 0
0 0 0 1 2
0 0 0 0 13
7777775
Version 2.30
858 Archetype U
Archetype U
Summary Domain is matrices, codomain is column vectors. Domain has dimension 6, while codomain
has dimension 4. Can't be injective, is surjective.
A linear transformation: (Denition LT [515])
T:M23!C4; Ta b c
d e f
=2
664a+ 2b+ 12c 3d+e+ 6f
2a b c+d 11f
a+b+ 7c+ 2d+e 3f
a+ 2b+ 12c+ 5e 5f3
775
A basis for the null space of the linear transformation: (Denition KLT [545])
3 4 0
1 2 1
; 2 5 1
0 0 0
Injective: No. (Denition ILT [541])
Since the kernel is nontrivial Theorem KILT [548] tells us that the linear transformation is not injective.
Also, since the rank can not exceed 4, we are guaranteed to have a nullity of at least 2, just from checking
dimensions of the domain and the codomain. In particular, verify that
T1 10 2
3 1 1
=2
664 7
14
1
133
775T5 3 1
5 3 3
=2
664 7
14
1
133
775
This demonstration that Tis not injective is constructed with the observation that
5 3 1
5 3 3
=1 10 2
3 1 1
+4 13 1
2 4 2
and
z=4 13 1
2 4 2
2K(T)
so the vector zeectively \does nothing" in the evaluation of T.
A basis for the range of the linear transformation: (Denition RLT [563])
Evaluate the linear transformation on a standard basis to get a spanning set for the range (Theorem SSRLT
[567]):
Version 2.30
Archetype U 859
2
6641
2
1
13
775;2
6642
1
1
23
775;2
66412
1
7
123
775;2
664 3
1
2
03
775;2
6641
0
1
53
775;2
6646
11
3
53
775
If the linear transformation is injective, then the set above is guaranteed to be linearly independent (The-
orem ILTLI [549]). This spanning set may be converted to a \nice" basis, by making the vectors the rows
of a matrix (perhaps after using a vector reperesentation), row-reducing, and retaining the nonzero rows
(Theorem BRS [280]), and perhaps un-coordinatizing. A basis for the range is:
8
>><
>>:2
6641
0
0
03
775;2
6640
1
0
03
775;2
6640
0
1
03
775;2
6640
0
0
13
7759
>>=
>>;
Surjective: Yes. (Denition SLT [559])
A basis for the range is the standard basis of C4, soR(T) =C4and Theorem RSLT [565] tells us T
is surjective. Or, the dimension of the range is 4, and the codomain ( C4) has dimension 4. So the
transformation is surjective.
Subspace dimensions associated with the linear transformation. Examine parallels with earlier results
for matrices. Verify Theorem RPNDD [588].
Domain dimension: 6 Rank: 4 Nullity: 2
Invertible: No.
The relative sizes of the domain and codomain mean the linear transformation cannot be injective. (The-
orem ILTIS [582])
Matrix representation (Denition MR [615]):
B=1 0 0
0 0 0
;0 1 0
0 0 0
;0 0 1
0 0 0
;0 0 0
1 0 0
;0 0 0
0 1 0
;0 0 0
0 0 1
C=8
>><
>>:2
6641
0
0
03
775;2
6640
1
0
03
775;2
6640
0
1
03
775;2
6640
0
0
13
7759
>>=
>>;
MT
B;C=2
6641 2 12 3 1 6
2 1 1 1 0 11
1 1 7 2 1 3
1 2 12 0 5 53
775
Version 2.30
860 Archetype V
Archetype V
Summary Domain is polynomials, codomain is matrices. Domain and codomain both have dimension
4. Injective, surjective, invertible. Square matrix representation, but domain and codomain are unequal,
so no eigenvalue information.
A linear transformation: (Denition LT [515])
T:P3!M22; T
a+bx+cx2+dx3
=a+b a 2c
d b d
A basis for the null space of the linear transformation: (Denition KLT [545])
fg
Injective: Yes. (Denition ILT [541])
Since the kernel is trivial Theorem KILT [548] tells us that the linear transformation is injective.
A basis for the range of the linear transformation: (Denition RLT [563])
Evaluate the linear transformation on a standard basis to get a spanning set for the range (Theorem SSRLT
[567]):
1 1
0 0
;1 0
0 1
;0 2
0 0
;0 0
1 1
If the linear transformation is injective, then the set above is guaranteed to be linearly independent (The-
orem ILTLI [549]). This spanning set may be converted to a \nice" basis, by making the vectors the rows
of a matrix (perhaps after using a vector reperesentation), row-reducing, and retaining the nonzero rows
(Theorem BRS [280]), and perhaps un-coordinatizing. A basis for the range is:
1 0
0 0
;0 1
0 0
;0 0
1 0
;0 0
0 1
Surjective: Yes. (Denition SLT [559])
A basis for the range is the standard basis of M22, soR(T) =M22and Theorem RSLT [565] tells us
Version 2.30
Archetype V 861
Tis surjective. Or, the dimension of the range is 4, and the codomain ( M22) has dimension 4. So the
transformation is surjective.
Subspace dimensions associated with the linear transformation. Examine parallels with earlier results
for matrices. Verify Theorem RPNDD [588].
Domain dimension: 4 Rank: 4 Nullity: 0
Invertible: Yes.
Both injective and surjective (Theorem ILTIS [582]). Notice that since the domain and codomain have the
same dimension, either the transformation is both injective and surjective (making it invertible, as in this
case) or else it is both not injective and not surjective.
Matrix representation (Denition MR [615]):
B=
1; x; x2; x3
C=1 0
0 0
;0 1
0 0
;0 0
1 0
;0 0
0 1
MT
B;C=2
6641 1 0 0
1 0 2 0
0 0 0 1
0 1 0 13
775
Since invertible, the inverse linear transformation. (Denition IVLT [579])
T 1:M22!P3; T 1a b
c d
= (a c d) + (c+d)x+1
2(a b c d)x2+cx3
Version 2.30
862 Archetype W
Archetype W
Summary Domain is polynomials, codomain is polynomials. Domain and codomain both have dimen-
sion 3. Injective, surjective, invertible, 3 distinct eigenvalues, diagonalizable.
A linear transformation: (Denition LT [515])
T:P2!P2; T
a+bx+cx2
= (19a+ 6b 4c) + ( 24a 7b+ 4c) + (36a+ 12b 9c)
A basis for the null space of the linear transformation: (Denition KLT [545])
fg
Injective: Yes. (Denition ILT [541])
Since the kernel is trivial Theorem KILT [548] tells us that the linear transformation is injective.
A basis for the range of the linear transformation: (Denition RLT [563])
Evaluate the linear transformation on a standard basis to get a spanning set for the range (Theorem SSRLT
[567]):
19 24x+ 36x2;6 7x+ 12x2; 4 + 4x 9x2
If the linear transformation is injective, then the set above is guaranteed to be linearly independent (The-
orem ILTLI [549]). This spanning set may be converted to a \nice" basis, by making the vectors the rows
of a matrix (perhaps after using a vector reperesentation), row-reducing, and retaining the nonzero rows
(Theorem BRS [280]), and perhaps un-coordinatizing. A basis for the range is:
1; x; x2
Surjective: Yes. (Denition SLT [559])
A basis for the range is the standard basis of C5, soR(T) =C5and Theorem RSLT [565] tells us T
is surjective. Or, the dimension of the range is 5, and the codomain ( C5) has dimension 5. So the
transformation is surjective.
Subspace dimensions associated with the linear transformation. Examine parallels with earlier results
for matrices. Verify Theorem RPNDD [588].
Domain dimension: 3 Rank: 3 Nullity: 0
Version 2.30
Archetype W 863
Invertible: Yes.
Both injective and surjective (Theorem ILTIS [582]). Notice that since the domain and codomain have the
same dimension, either the transformation is both injective and surjective (making it invertible, as in this
case) or else it is both not injective and not surjective.
Matrix representation (Denition MR [615]):
B=
1; x; x2
C=
1; x; x2
MT
B;C=2
419 6 4
24 7 4
36 12 93
5
Since invertible, the inverse linear transformation. (Denition IVLT [579])
T 1:P2!P2; T 1
a+bx+cx2
= ( 5a 2b+4
3c) + (24a+ 9b 20
3c)x+ (12a+ 4b 11
3c)x2
Eigenvalues and eigenvectors (Denition EELT [647], Theorem EER [659]):
= 1 ET( 1) =
2x+ 3x2
= 1 ET(1) =hf 1 + 3xgi
= 3 ET(3) =
1 2x+x2
Evaluate the linear transformation with each of these eigenvectors as an interesting check.
A diagonal matrix representation relative to a basis of eigenvectors, B.
B=
2x+ 3x2; 1 + 3x;1 2x+x2
MT
B;B=2
4 1 0 0
0 1 0
0 0 33
5
Version 2.30
864 Archetype X
Archetype X
Summary Domain and codomain are square matrices. Domain and codomain both have dimension 4.
Not injective, not surjective, not invertible, 3 distinct eigenvalues, diagonalizable.
A linear transformation: (Denition LT [515])
T:M22!M22; Ta b
c d
= 2a+ 15b+ 3c+ 27d 10b+ 6c+ 18d
a 5b 9d a 4b 5c 8d
A basis for the null space of the linear transformation: (Denition KLT [545])
6 3
2 1
Injective: No. (Denition ILT [541])
Since the kernel is nontrivial Theorem KILT [548] tells us that the linear transformation is not injective.
In particular, verify that
T 2 0
1 4
=115 78
38 35
T4 3
1 3
=115 78
38 35
This demonstration that Tis not injective is constructed with the observation that
4 3
1 3
= 2 0
1 4
+6 3
2 1
and
z=6 3
2 1
2K(T)
so the vector zeectively \does nothing" in the evaluation of T.
A basis for the range of the linear transformation: (Denition RLT [563])
Evaluate the linear transformation on a standard basis to get a spanning set for the range (Theorem SSRLT
[567]):
2 0
1 1
;15 10
5 4
;3 6
0 5
;27 18
9 8
If the linear transformation is injective, then the set above is guaranteed to be linearly independent (The-
orem ILTLI [549]). This spanning set may be converted to a \nice" basis, by making the vectors the rows
Version 2.30
Archetype X 865
of a matrix (perhaps after using a vector reperesentation), row-reducing, and retaining the nonzero rows
(Theorem BRS [280]), and perhaps un-coordinatizing. A basis for the range is:
1 0
1
20
;0 1
1
40
;0 0
0 1
Surjective: No. (Denition SLT [559])
The dimension of the range is 3, and the codomain ( M22) has dimension 5. So R(T)6=M22and by
Theorem RSLT [565] the transformation is not surjective.
To be more precise, verify that2 4
3 1
62R(T), by setting the output of Tequal to this matrix and
seeing that the resulting system of linear equations has no solution, i.e. is inconsistent. So the preimage,
T 12 4
3 1
, is empty. This alone is sucient to see that the linear transformation is not onto.
Subspace dimensions associated with the linear transformation. Examine parallels with earlier results
for matrices. Verify Theorem RPNDD [588].
Domain dimension: 4 Rank: 3 Nullity: 1
Invertible: No.
Neither injective nor surjective (Theorem ILTIS [582]). Notice that since the domain and codomain have
the same dimension, either the transformation is both injective and surjective or else it is both not injective
and not surjective (making it not invertible, as in this case).
Matrix representation (Denition MR [615]):
B=1 0
0 0
;0 1
0 0
;0 0
1 0
;0 0
0 1
C=1 0
0 0
;0 1
0 0
;0 0
1 0
;0 0
0 1
MT
B;C=2
664 2 15 3 27
0 10 6 18
1 5 0 9
1 4 5 83
775
Eigenvalues and eigenvectors (Denition EELT [647], Theorem EER [659]):
= 0 ET(0) = 6 3
2 1
Version 2.30
866 Archetype X
= 1 ET(1) = 7 2
3 0
; 1 2
0 1
= 3 ET(3) = 3 2
1 1
Evaluate the linear transformation with each of these eigenvectors as an interesting check.
A diagonal matrix representation relative to a basis of eigenvectors, B.
B= 6 3
2 1
; 7 2
3 0
; 1 2
0 1
; 3 2
1 1
MT
B;B=2
6640 0 0 0
0 1 0 0
0 0 3 0
0 0 0 33
775
Version 2.30
Appendix GFDL
GNU Free Documentation License
Version 1.2, November 2002
Copyright c
2000,2001,2002 Free Software Foundation, Inc.
59 Temple Place, Suite 330, Boston, MA 02111-1307 USA
Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is
not allowed.
Preamble
The purpose of this License is to make a manual, textbook, or other functional and useful document
\free" in the sense of freedom: to assure everyone the eective freedom to copy and redistribute it, with
or without modifying it, either commercially or noncommercially. Secondarily, this License preserves for
the author and publisher a way to get credit for their work, while not being considered responsible for
modications made by others.
This License is a kind of \copyleft", which means that derivative works of the document must themselves
be free in the same sense. It complements the GNU General Public License, which is a copyleft license
designed for free software.
We have designed this License in order to use it for manuals for free software, because free software
needs free documentation: a free program should come with manuals providing the same freedoms that
the software does. But this License is not limited to software manuals; it can be used for any textual work,
regardless of subject matter or whether it is published as a printed book. We recommend this License
principally for works whose purpose is instruction or reference.
1. APPLICABILITY AND DEFINITIONS
This License applies to any manual or other work, in any medium, that contains a notice placed by
the copyright holder saying it can be distributed under the terms of this License. Such a notice grants a
world-wide, royalty-free license, unlimited in duration, to use that work under the conditions stated herein.
The\Document" , below, refers to any such manual or work. Any member of the public is a licensee,
and is addressed as \you" . You accept the license if you copy, modify or distribute the work in a way
requiring permission under copyright law.
A\Modied Version" of the Document means any work containing the Document or a portion of
it, either copied verbatim, or with modications and/or translated into another language.
A\Secondary Section" is a named appendix or a front-matter section of the Document that deals
exclusively with the relationship of the publishers or authors of the Document to the Document's overall
subject (or to related matters) and contains nothing that could fall directly within that overall subject.
(Thus, if the Document is in part a textbook of mathematics, a Secondary Section may not explain any
867
868 Appendix GFDL GNU Free Documentation License
mathematics.) The relationship could be a matter of historical connection with the subject or with related
matters, or of legal, commercial, philosophical, ethical or political position regarding them.
The\Invariant Sections" are certain Secondary Sections whose titles are designated, as being those
of Invariant Sections, in the notice that says that the Document is released under this License. If a section
does not t the above denition of Secondary then it is not allowed to be designated as Invariant. The
Document may contain zero Invariant Sections. If the Document does not identify any Invariant Sections
then there are none.
The\Cover Texts" are certain short passages of text that are listed, as Front-Cover Texts or Back-
Cover Texts, in the notice that says that the Document is released under this License. A Front-Cover Text
may be at most 5 words, and a Back-Cover Text may be at most 25 words.
A\Transparent" copy of the Document means a machine-readable copy, represented in a format
whose specication is available to the general public, that is suitable for revising the document straightfor-
wardly with generic text editors or (for images composed of pixels) generic paint programs or (for drawings)
some widely available drawing editor, and that is suitable for input to text formatters or for automatic
translation to a variety of formats suitable for input to text formatters. A copy made in an otherwise
Transparent le format whose markup, or absence of markup, has been arranged to thwart or discourage
subsequent modication by readers is not Transparent. An image format is not Transparent if used for
any substantial amount of text. A copy that is not \Transparent" is called \Opaque" .
Examples of suitable formats for Transparent copies include plain ASCII without markup, Texinfo input
format, LaTeX input format, SGML or XML using a publicly available DTD, and standard-conforming
simple HTML, PostScript or PDF designed for human modication. Examples of transparent image
formats include PNG, XCF and JPG. Opaque formats include proprietary formats that can be read and
edited only by proprietary word processors, SGML or XML for which the DTD and/or processing tools
are not generally available, and the machine-generated HTML, PostScript or PDF produced by some word
processors for output purposes only.
The\Title Page" means, for a printed book, the title page itself, plus such following pages as are
needed to hold, legibly, the material this License requires to appear in the title page. For works in formats
which do not have any title page as such, \Title Page" means the text near the most prominent appearance
of the work's title, preceding the beginning of the body of the text.
A section \Entitled XYZ" means a named subunit of the Document whose title either is precisely
XYZ or contains XYZ in parentheses following text that translates XYZ in another language. (Here XYZ
stands for a specic section name mentioned below, such as \Acknowledgements" ,\Dedications" ,
\Endorsements" , or\History" .) To \Preserve the Title" of such a section when you modify the
Document means that it remains a section \Entitled XYZ" according to this denition.
The Document may include Warranty Disclaimers next to the notice which states that this License
applies to the Document. These Warranty Disclaimers are considered to be included by reference in this
License, but only as regards disclaiming warranties: any other implication that these Warranty Disclaimers
may have is void and has no eect on the meaning of this License.
2. VERBATIM COPYING
You may copy and distribute the Document in any medium, either commercially or noncommercially,
provided that this License, the copyright notices, and the license notice saying this License applies to the
Document are reproduced in all copies, and that you add no other conditions whatsoever to those of this
License. You may not use technical measures to obstruct or control the reading or further copying of the
copies you make or distribute. However, you may accept compensation in exchange for copies. If you
distribute a large enough number of copies you must also follow the conditions in section 3.
You may also lend copies, under the same conditions stated above, and you may publicly display copies.
3. COPYING IN QUANTITY
Version 2.30
Appendix GFDL GNU Free Documentation License 869
If you publish printed copies (or copies in media that commonly have printed covers) of the Document,
numbering more than 100, and the Document's license notice requires Cover Texts, you must enclose the
copies in covers that carry, clearly and legibly, all these Cover Texts: Front-Cover Texts on the front cover,
and Back-Cover Texts on the back cover. Both covers must also clearly and legibly identify you as the
publisher of these copies. The front cover must present the full title with all words of the title equally
prominent and visible. You may add other material on the covers in addition. Copying with changes
limited to the covers, as long as they preserve the title of the Document and satisfy these conditions, can
be treated as verbatim copying in other respects.
If the required texts for either cover are too voluminous to t legibly, you should put the rst ones
listed (as many as t reasonably) on the actual cover, and continue the rest onto adjacent pages.
If you publish or distribute Opaque copies of the Document numbering more than 100, you must
either include a machine-readable Transparent copy along with each Opaque copy, or state in or with
each Opaque copy a computer-network location from which the general network-using public has access
to download using public-standard network protocols a complete Transparent copy of the Document, free
of added material. If you use the latter option, you must take reasonably prudent steps, when you begin
distribution of Opaque copies in quantity, to ensure that this Transparent copy will remain thus accessible
at the stated location until at least one year after the last time you distribute an Opaque copy (directly
or through your agents or retailers) of that edition to the public.
It is requested, but not required, that you contact the authors of the Document well before redistributing
any large number of copies, to give them a chance to provide you with an updated version of the Document.
4. MODIFICATIONS
You may copy and distribute a Modied Version of the Document under the conditions of sections 2
and 3 above, provided that you release the Modied Version under precisely this License, with the Modied
Version lling the role of the Document, thus licensing distribution and modication of the Modied Version
to whoever possesses a copy of it. In addition, you must do these things in the Modied Version:
A. Use in the Title Page (and on the covers, if any) a title distinct from that of the Document, and from
those of previous versions (which should, if there were any, be listed in the History section of the
Document). You may use the same title as a previous version if the original publisher of that version
gives permission.
B. List on the Title Page, as authors, one or more persons or entities responsible for authorship of
the modications in the Modied Version, together with at least ve of the principal authors of the
Document (all of its principal authors, if it has fewer than ve), unless they release you from this
requirement.
C. State on the Title page the name of the publisher of the Modied Version, as the publisher.
D. Preserve all the copyright notices of the Document.
E. Add an appropriate copyright notice for your modications adjacent to the other copyright notices.
F. Include, immediately after the copyright notices, a license notice giving the public permission to use
the Modied Version under the terms of this License, in the form shown in the Addendum below.
G. Preserve in that license notice the full lists of Invariant Sections and required Cover Texts given in
the Document's license notice.
H. Include an unaltered copy of this License.
Version 2.30
870 Appendix GFDL GNU Free Documentation License
I. Preserve the section Entitled \History", Preserve its Title, and add to it an item stating at least
the title, year, new authors, and publisher of the Modied Version as given on the Title Page. If
there is no section Entitled \History" in the Document, create one stating the title, year, authors,
and publisher of the Document as given on its Title Page, then add an item describing the Modied
Version as stated in the previous sentence.
J. Preserve the network location, if any, given in the Document for public access to a Transparent copy
of the Document, and likewise the network locations given in the Document for previous versions it
was based on. These may be placed in the \History" section. You may omit a network location for a
work that was published at least four years before the Document itself, or if the original publisher of
the version it refers to gives permission.
K. For any section Entitled \Acknowledgements" or \Dedications", Preserve the Title of the section,
and preserve in the section all the substance and tone of each of the contributor acknowledgements
and/or dedications given therein.
L. Preserve all the Invariant Sections of the Document, unaltered in their text and in their titles. Section
numbers or the equivalent are not considered part of the section titles.
M. Delete any section Entitled \Endorsements". Such a section may not be included in the Modied
Version.
N. Do not retitle any existing section to be Entitled \Endorsements" or to con
ict in title with any
Invariant Section.
O. Preserve any Warranty Disclaimers.
If the Modied Version includes new front-matter sections or appendices that qualify as Secondary
Sections and contain no material copied from the Document, you may at your option designate some or all
of these sections as invariant. To do this, add their titles to the list of Invariant Sections in the Modied
Version's license notice. These titles must be distinct from any other section titles.
You may add a section Entitled \Endorsements", provided it contains nothing but endorsements of
your Modied Version by various parties{for example, statements of peer review or that the text has been
approved by an organization as the authoritative denition of a standard.
You may add a passage of up to ve words as a Front-Cover Text, and a passage of up to 25 words
as a Back-Cover Text, to the end of the list of Cover Texts in the Modied Version. Only one passage of
Front-Cover Text and one of Back-Cover Text may be added by (or through arrangements made by) any
one entity. If the Document already includes a cover text for the same cover, previously added by you or
by arrangement made by the same entity you are acting on behalf of, you may not add another; but you
may replace the old one, on explicit permission from the previous publisher that added the old one.
The author(s) and publisher(s) of the Document do not by this License give permission to use their
names for publicity for or to assert or imply endorsement of any Modied Version.
5. COMBINING DOCUMENTS
You may combine the Document with other documents released under this License, under the terms
dened in section 4 above for modied versions, provided that you include in the combination all of the
Invariant Sections of all of the original documents, unmodied, and list them all as Invariant Sections of
your combined work in its license notice, and that you preserve all their Warranty Disclaimers.
The combined work need only contain one copy of this License, and multiple identical Invariant Sections
may be replaced with a single copy. If there are multiple Invariant Sections with the same name but dierent
contents, make the title of each such section unique by adding at the end of it, in parentheses, the name
Version 2.30
Appendix GFDL GNU Free Documentation License 871
of the original author or publisher of that section if known, or else a unique number. Make the same
adjustment to the section titles in the list of Invariant Sections in the license notice of the combined work.
In the combination, you must combine any sections Entitled \History" in the various original documents,
forming one section Entitled \History"; likewise combine any sections Entitled \Acknowledgements", and
any sections Entitled \Dedications". You must delete all sections Entitled \Endorsements".
6. COLLECTIONS OF DOCUMENTS
You may make a collection consisting of the Document and other documents released under this License,
and replace the individual copies of this License in the various documents with a single copy that is included
in the collection, provided that you follow the rules of this License for verbatim copying of each of the
documents in all other respects.
You may extract a single document from such a collection, and distribute it individually under this
License, provided you insert a copy of this License into the extracted document, and follow this License in
all other respects regarding verbatim copying of that document.
7. AGGREGATION WITH INDEPENDENT WORKS
A compilation of the Document or its derivatives with other separate and independent documents or
works, in or on a volume of a storage or distribution medium, is called an \aggregate" if the copyright
resulting from the compilation is not used to limit the legal rights of the compilation's users beyond what
the individual works permit. When the Document is included in an aggregate, this License does not apply
to the other works in the aggregate which are not themselves derivative works of the Document.
If the Cover Text requirement of section 3 is applicable to these copies of the Document, then if the
Document is less than one half of the entire aggregate, the Document's Cover Texts may be placed on covers
that bracket the Document within the aggregate, or the electronic equivalent of covers if the Document is
in electronic form. Otherwise they must appear on printed covers that bracket the whole aggregate.
8. TRANSLATION
Translation is considered a kind of modication, so you may distribute translations of the Document
under the terms of section 4. Replacing Invariant Sections with translations requires special permission
from their copyright holders, but you may include translations of some or all Invariant Sections in addition
to the original versions of these Invariant Sections. You may include a translation of this License, and all
the license notices in the Document, and any Warranty Disclaimers, provided that you also include the
original English version of this License and the original versions of those notices and disclaimers. In case
of a disagreement between the translation and the original version of this License or a notice or disclaimer,
the original version will prevail.
If a section in the Document is Entitled \Acknowledgements", \Dedications", or \History", the re-
quirement (section 4) to Preserve its Title (section 1) will typically require changing the actual title.
9. TERMINATION
You may not copy, modify, sublicense, or distribute the Document except as expressly provided for
under this License. Any other attempt to copy, modify, sublicense or distribute the Document is void, and
will automatically terminate your rights under this License. However, parties who have received copies, or
rights, from you under this License will not have their licenses terminated so long as such parties remain
in full compliance.
10. FUTURE REVISIONS OF THIS LICENSE
Version 2.30
872 Appendix GFDL GNU Free Documentation License
The Free Software Foundation may publish new, revised versions of the GNU Free Documentation
License from time to time. Such new versions will be similar in spirit to the present version, but may dier
in detail to address new problems or concerns. See http://www.gnu.org/copyleft/.
Each version of the License is given a distinguishing version number. If the Document species that
a particular numbered version of this License \or any later version" applies to it, you have the option of
following the terms and conditions either of that specied version or of any later version that has been
published (not as a draft) by the Free Software Foundation. If the Document does not specify a version
number of this License, you may choose any version ever published (not as a draft) by the Free Software
Foundation.
ADDENDUM: How to use this License for your documents
To use this License in a document you have written, include a copy of the License in the document and
put the following copyright and license notices just after the title page:
Copyright c
YEAR YOUR NAME. Permission is granted to copy, distribute and/or modify this
document under the terms of the GNU Free Documentation License, Version 1.2 or any later
version published by the Free Software Foundation; with no Invariant Sections, no Front-Cover
Texts, and no Back-Cover Texts. A copy of the license is included in the section entitled \GNU
Free Documentation License".
If you have Invariant Sections, Front-Cover Texts and Back-Cover Texts, replace the \with...Texts."
line with this:
with the Invariant Sections being LIST THEIR TITLES, with the Front-Cover Texts being
LIST, and with the Back-Cover Texts being LIST.
If you have Invariant Sections without Cover Texts, or some other combination of the three, merge
those two alternatives to suit the situation.
If your document contains nontrivial examples of program code, we recommend releasing these examples
in parallel under your choice of free software license, such as the GNU General Public License, to permit
their use in free software.
Version 2.30
Part T
Topics
Section F
Fields
Draft: This Section Complete, But Subject To Change
We have chosen to present introductory linear algebra in the Core (Part C [3]) using scalars from the
set of complex numbers, C. We could have instead chosen to use scalars from the set of real numbers,
R. This would have presented certain diculties when we encountered characteristic polynomials with
complex roots (Denition CP [460]) or when we needed to be sure every matrix had at least one eigenvalue
(Theorem EMHE [457]). However, much of the basics would be unchanged. The denition of a vector space
would not change, nor would the ideas of linear independence, spanning, or bases. Linear transformations
would still behave the same and we would still obtain matrix representations, though our ideas about
canonical forms would have to be adjusted slightly.
The real numbers and the complex numbers are both examples of what are called elds, and we can \do"
linear algebra in just a bit more generality by letting our scalars take values from some unspecied eld.
So in this section we will describe exactly what constitutes a eld, give some nite examples, and discuss
another connection between elds and vector spaces. Vector spaces over nite elds are very important
in certain applications, so this is partially background for other topics. As such, we will not prove every
claim we make.
Subsection F
Fields
Like a vector space, a eld is a set along with two binary operations. The distinction is that both operations
accept two elements of the set, and then produce a new element of the set. In a vector space we have two
sets | the vectors and the scalars, and scalar multiplication mixes one of each to produce a vector. Here
is the careful denition of a eld.
Denition F
Field
Suppose that Fis a set upon which we have dened two operations: (1) addition , which combines two
elements of Fand is denoted by \+", and (2) multiplication , which combines two elements of Fand
is denoted by juxtaposition. Then F, along with the two operations, is a eld if the following properties
hold.
ACF Additive Closure, Field
If;2F, then+2F.
MCF Multiplicative Closure, Field
If;2F, then2F.
CAF Commutativity of Addition, Field
If;2F, then+=+.
CMF Commutativity of Multiplication, Field
If;2F, then=.
AAF Additive Associativity, Field
If; ;
2F, then+ (+
) = (+) +
.
876 Section F Fields
MAF Multiplicative Associativity, Field
If; ;
2F, then(
) = ()
.
DF Distributivity, Field
If; ;
2F, then(+
) =+
.
ZF Zero, Field
There is an element, 0 2F, called zero , such that + 0 =for all2F.
OF One, Field
There is an element, 1 2F, called one, such that (1) =for all2F.
AIF Additive Inverse, Field
If2F, then there exists 2Fso that+ ( ) = 0.
MIF Multiplicative Inverse, Field
If2F,6= 0, then there exists1
2Fso that 1
= 1.
4
Mostly this denition says that all the good things you might expect, really do happen in a eld. The
one technicality is that the special element, 0, the additive identity element, does not have a multiplicative
inverse. In other words, no dividing by zero.
This denition should remind you of Theorem PCNA [758], and indeed, Theorem PCNA [758] provides
the justication for the statement that the complex numbers form a eld. Another example of eld is the
set of rational numbers
Q=p
qp; qare integers, q6= 0
Of course, the real numbers, R, also form a eld. It is this eld that you probably studied for many years.
You began studying the integers (\counting"), then the rationals (\fractions"), then the reals (\algebra"),
along with some excursions in the complex numbers (\imaginary numbers"). So you should have seen
three elds already in your previous studies.
Our rst observation about elds is that we can go back to our denition of a vector space (Denition
VS [317]) and replace every occurrence of Cby some general, unspecied eld, F, and all our subsequent
denitions and theorems are still true, so long as we avoid roots of polynomials (or equivalently, factoring
polynomials). So if you consult more advanced texts on linear algebra, you will see this sort of approach.
You might study some of the rst theorems we proved about vector spaces in Subsection VS.VSP [323] and
work through their proofs in the more general setting of an arbitrary eld. This exercise should convince
you that very little changes when we move from Cto an arbitrary eld F. (See Exercise F.T10 [880].)
Subsection FF
Finite Fields
It may sound odd at rst, but there exist nite elds, and even nite vector spaces. We will nd certain
of these important in subsequent applications, so we collect some ideas and properties here.
Denition IMP
Integers Modulo a Prime
Suppose that pis a prime number. Let Zp=f0;1;2; :::; p 1g. Add and multiply elements of Zpas
integers, but whenever a result lies outside of the set Zp, nd its remainder after division by pand replace
the result by this remainder. 4
We have dened a set, and two binary operations. The result is a eld.
Version 2.30
Subsection F.FF Finite Fields 877
Theorem FIMP
Field of Integers Modulo a Prime
The set of integers modulo a prime p,Zp, is a eld.
Example IM11
Integers mod 11
Z11is a eld by Theorem FIMP [875]. Here we provide some sample calculations.
8 + 5 = 2 8 = 3 5 9 = 7
5(7) = 21
7= 86
5= 10
25= 10 1 = 101
0= ?
We can now \do" linear algebra using scalars from a nite eld.
Example VSIM5
Vector space over integers mod 5
Let (Z5)3be the set of all column vectors of length 3 with entries from Z5. Use Z5as the set of scalars.
Dene addition and multiplication the usual way. We exhibit a few sample calculations.
2
42
3
43
5+2
44
1
33
5=2
41
4
23
5 32
42
0
43
5=2
41
0
23
5
We can, of course, build linear combinations, such as
22
41
3
03
5 42
42
1
13
5+2
41
2
43
5=2
40
4
03
5
which almost looks like a relation of linear dependence. The set
8
<
:2
41
3
13
5;2
42
2
03
59
=
;
is linearly independent, while the set
8
<
:2
41
3
13
5;2
42
2
03
5;2
44
3
23
59
=
;
is linearly dependent, as can be seen from the relation of linear dependence formed by the scalars a1= 2,
a2= 1 anda3= 4. To nd these scalars, one would take the same approach as Example LDS [153], but
in performing row operations to solve a homogeneous system, you would need to take care that all scalar
(eld) operations are performed over Z5, especially when multiplying a row by a scalar to make a leading
entry equal to 1. One more observation about this example | the set
8
<
:2
41
0
03
5;2
41
1
03
5;2
41
1
13
59
=
;
Version 2.30
878 Section F Fields
is a basis for ( Z5)3, since it is both linearly independent and spans ( Z5)3.
In applications to computer science or electrical engineering, Z2is the most important eld, since it can
be used to describe the binary nature of logic, circuitry, communications and their intertwined relationships.
The vector space of column vectors with entries from Z2, (Z2)n, with scalars taken from Z2is the natural
extension of this idea. Notice that Z2has the minimum number of elements to be a eld, since any eld
must contain a zero and a one (Property ZF [874], Property OF [874]).
Example SM2Z7
Symmetric matrices of size 2 over Z7
We can employ the eld of integers modulo a prime to build other examples of vector spaces with novel
elds of scalars. Dene
S22(Z7) =a b
b ca; b; c2Z7
which is the set of all 2 2 symmetric matrices with entries from Z7. Use the eld Z7as the set of scalars,
and dene vector addition and scalar multiplication in the natural way. The result will be a vector space.
Notice that the eld of scalars is nite, as is the vector space, since there are 73= 343 matrices in
S22(Z7). The set
1 0
0 0
;0 1
1 0
;0 0
0 1
is a basis, so dim ( S22(Z7)) = 3.
In a more advanced algebra course it is possible to prove that the number of elements in a nite eld
must be of the form pn, wherepis a prime. We can't go so far aeld as to prove this here, but we can
demonstrate an example.
Example FF8
Finite eld of size 8
Dene the set FasF=
a+bt+ct2a; b; c2Z2
. Add and multiply these quantities as polynomials in
the variable t, but replace any occurrence of t3byt+ 1.
This denes a set, and the two operations on elements of that set. Do not be concerned with what t
\is," because it isn't. tis just a handy device that makes the example a eld. We'll say a bit more about
twhen we nish. But rst, some examples. Remember that 1 + 1 = 0 in Z2. Addition is quite simple, for
example,
1 +t+t2
+
1 +t2
= (1 + 1) + (1 + 0) t+ (1 + 1)t2=t
Multiplication gets more involved, for example,
1 +t+t2
1 +t2
= 1 +t2+t+t3+t2+t4
= 1 +t+ (1 + 1)t2+t3(1 +t)
= 1 +t+ (1 +t) (1 +t)
= 1 +t+ 1 +t+t+t2
= (1 + 1) + (1 + 1 + 1) t+t2
=t+t2
Every element has a multiplicative inverse (Property MIF [874]). What is the inverse of t+t2? Check that
t+t2
(1 +t) =t+t2+t2+t3
Version 2.30
Subsection F.FF Finite Fields 879
=t+ (1 + 1)t2+ (1 +t)
=t+ 1 +t
= 1 + (1 + 1) t
= 1
So we can write1
t+t2= 1 +t. So that you may experiment, we give you the complete addition and
multiplication tables for this eld. Addition is simple, while multiplication is more interesting, so verify
a few entries of each table. Because of the commutativity of addition and multiplication (Property CAF
[873], Property CMF [873]), we have just listed half of each table.
+ 0 1t t2t+ 1t2+t t2+t+ 1t2+ 1
0 0 1t t2t+ 1t2+t t2+t+ 1t2+ 1
1 0t+ 1t2+ 1t t2+t+ 1t2+t t2
t 0t2+t1 t2t2+ 1t2+t+ 1
t20t2+t+ 1t t + 1 1
t+ 1 0 t2+ 1t2t2+t
t2+t 0 1 t+ 1
t2+t+ 1 0 t
t2+ 1 0
0 1t t2t+ 1t2+t t2+t+ 1t2+ 1
0 0 0 0 0 0 0 0 0
1 1t t2t+ 1t2+t t2+t+ 1t2+ 1
t t2t+ 1t2+t t2+t+ 1t2+ 1 1
t2t2+t t2+t+ 1t2+ 1 1 t
t+ 1 t2+ 1 1 t t2
t2+t t t2t+ 1
t2+t+ 1 1 +t t2+t
t2+ 1 t2+t+ 1
Note that every element of Fis a linear combination (with scalars from Z2) of the polynomials 1, t,t2.
SoB=
1; t; t2
is a spanning set for F. Further,Bis linearly independent since there is no nontrivial
relation of linear dependence, and Bis a basis. So dim ( F) = 3. Of course, this paragraph presumes that
Fis also a vector space over Z2(which it is).
The dening relation for t(t3=t+ 1) in Example FF8 [876] arises from the polynomial t3+t+ 1,
which has no factorization with coecients from Z2. This is an example of an irreducible polynomial ,
which involves considerable theory to fully understand. In the exercises, we provide you with a few more
irreducible polynomials to experiment with. See the suggested readings if you would like to learn more.
Trivially, every eld (nite or otherwise) is a vector space. Suppose we begin with a eld F. From this
we knowFhas two binary operations dened on it. We need to somehow create a vector space from F,
in a general way. First we need a set of vectors. That'll be F. We also need a set of scalars. That'll be F
as well. How do we dene the addition of two vectors? By the same rule that we use to add them when
they are in the eld. How do we dene scalar multiplication? Since a scalar is an element of F, and a
vector is an element of F, we can dene scalar multiplication to be the same rule that we use to multiply
the two elements as members of the eld. With these denitions, Fwill be a vector space (Exercise F.T20
[880]). This is something of a trivial situation, since the set of vectors and the set of scalars are identical.
In particular, do not confuse this with Example FF8 [876] where the set of vectors has eight elements, and
the set of scalars has just two elements.
Further Reading
Robert J. McEliece, Finite Fields for Scientists and Engineers. Kluwer Academic Publishers, 1987.
Version 2.30
880 Section F Fields
Rudolpf Lidl, Harald Niederreiter, Introduction to Finite Fields and Their Applications, Revised Edi-
tion. Cambridge University Press, 1994.
Version 2.30
Subsection F.EXC Exercises 881
Subsection EXC
Exercises
C60 Consider the vector space ( Z5)4composed of column vectors of size 4 with entries from Z5. The
matrixAis a square matrix composed of four such column vectors.
A=2
6643 3 0 3
1 2 3 0
1 1 0 2
4 2 2 13
775
Find the inverse of A. Use this to nd a solution to LS(A;b) when
b=2
6643
3
2
03
775
Contributed by Robert Beezer Solution [881]
M10 Suppose we relax the restriction in Denition IMP [874] to allow pto not be a prime. Will the
construction given still be a eld? Is Z6a eld? Can you generalize?
Contributed by Robert Beezer
M40 Construct a nite eld with 9 elements using the set
F=fa+btja; b2Z3g
wheret2is consistently replaced by 2 t+1 in any intermediate results obtained with polynomial multiplica-
tion. Compute the rst nine powers of t(t0throught8). Use this information to aid you in the construction
of the multiplication table for this eld. What is the multiplicative inverse of 2 t?
Contributed by Robert Beezer
M45 Construct a nite eld with 25 elements using the set
F=fa+btja; b2Z5g
wheret2is consistently replaced by t+3 in any intermediate results obtained with polynomial multiplication.
Compute the rst 25 powers of t(t0throught24). Use this information to aid you in computing in this
eld. What is the multiplicative inverse of 2 t? What is the multiplicative inverse of 4? What is the
multiplicative inverse of 1 + 4 t?
Find a basis for Fas a vector space with Z5used as the set of scalars.
Contributed by Robert Beezer
M50 Construct a nite eld with 16 elements using the set
F=
a+bt+ct2+dt3a; b; c; d2Z2
wheret4is consistently replaced by t+1 in any intermediate results obtained with polynomial multiplication.
Compute the rst 16 powers of t(t0throught15). Consider the set G=
0;1; t5; t10
. ThenGwill also
be a nite eld, a subeld of F. Construct the addition and multiplication tables for G. Notice that since
bothGandFare vector spaces over Z2, andGF, by Denition S [333], Gis a subspace of F.
Contributed by Robert Beezer
Version 2.30
882 Section F Fields
T10 Give a new proof of Theorem ZVSM [325] for a vector space whose scalars come from an arbitrary
eldF.
Contributed by Robert Beezer
T20 By applying Denition VS [317], prove that every eld is also a vector space. (See the construction
at the end of this section.)
Contributed by Robert Beezer
Version 2.30
Subsection F.SOL Solutions 883
Subsection SOL
Solutions
C60 Contributed by Robert Beezer Statement [879]
Remember that every computation must be done with arithmetic in the eld, reducing any intermediate
number outside of f0;1;2;3;4gto its remainder after division by 5.
The matrix inverse can be found with Theorem CINM [248] (and we discover along the way that Ais
nonsingular). The inverse is
A 1=2
6641 1 3 1
3 4 1 4
1 4 0 2
3 0 1 03
775
Then by an application of Theorem SNCM [261] the (unique) solution to the system will be
A 1b=2
6641 1 3 1
3 4 1 4
1 4 0 2
3 0 1 03
7752
6643
3
2
03
775=2
6642
3
0
13
775
Version 2.30
884 Section F Fields
Version 2.30
Section T Trace 885
Section T
Trace
This section contributed by Andy Zimmer.
The matrix trace is a function that sends square matrices to scalars. In some ways it is reminiscent of
the determinant. And like the determinant, it has many useful and surprising properties.
Denition T
Trace
SupposeAis a square matrix of size n. Then the trace ofA,t(A), is the sum of the diagonal entries of
A. Symbolically,
t(A) =nX
i=1[A]ii
(This denition contains Notation T.) 4
The next three proofs make for excellent practice. In some books they would be left as exercises for the
reader as they are all \trivial" in the sense they do not rely on anything but the denition of the matrix
trace.
Theorem TL
Trace is Linear
SupposeAandBare square matrices of size n. Thent(A+B) =t(A) +t(B). Furthermore, if 2C,
thent(A) =t(A).
Proof These properties are exactly those required for a linear transformation. To prove these results we
just manipulate sums,
t(A+B) =nX
k=1[A+B]ii Denition T [883]
=nX
i=1[A]ii+ [B]ii Denition MA [207]
=nX
i=1[A]ii+nX
i=1[B]ii Property CACN [758]
=t(A) +t(B) Denition T [883]
The second part is as straightforward as the rst,
t(A) =nX
i=1[A]ii Denition T [883]
=nX
i=1[A]ii Denition MSM [208]
=nX
i=1[A]ii Property DCN [759]
=t(A) Denition T [883]
Version 2.30
886 Section T Trace
Theorem TSRM
Trace is Symmetric with Respect to Multiplication
SupposeAandBare square matrices of size n. Thent(AB) =t(BA).
Proof
t(AB) =nX
k=1[AB]kk Denition T [883]
=nX
k=1nX
`=1[A]k`[B]`k Theorem EMP [227]
=nX
`=1nX
k=1[A]k`[B]`k Property CACN [758]
=nX
`=1nX
k=1[B]`k[A]k` Property CMCN [758]
=nX
`=1[BA]`` Theorem EMP [227]
=t(BA) Denition T [883]
Theorem TIST
Trace is Invariant Under Similarity Transformations
SupposeAandSare square matrices of size nandSis invertible. Then t
S 1AS
=t(A).
Proof Invariant means constant under some operation. In this case the operation is a similarity trans-
formation. A lengthy exercise (but possibly a educational one) would be to prove this result without
referencing Theorem TSRM [884]. But here we will,
t
S 1AS
=t
S 1A
S
Theorem MMA [231]
=t
S
S 1A
Theorem TSRM [884]
=t
SS 1
A
Theorem MMA [231]
=t(A) Denition MI [244]
Now we could dene the trace of a linear transformation as the trace of any matrix representation of
the transformation. Would this denition be well-dened? That is, will two dierent representations of
the same linear transformation always have the same trace? Why? (Think Theorem SCB [656].) We will
now prove one of the most interesting and surprising results about the trace.
Theorem TSE
Trace is the Sum of the Eigenvalues
Suppose that Ais a square matrix of size nwith distinct eigenvalues 1; 2; 3; :::; k. Then
t(A) =kX
i=1A(i)i
Version 2.30
Section T Trace 887
Proof It is amazing that the eigenvalues would have anything to do with the sum of the diagonal entries.
Our proof will rely on double counting. We will demonstrate two dierent ways of counting the same thing
therefore proving equality. Our object of interest is the coecient of xn 1in the characteristic polynomial
ofA(Denition CP [460]), which will be denoted n 1. From the proof of Theorem NEM [485] we have,
pA(x) = ( 1)n(x 1)A(1)(x 2)A(2)(x 3)A(3)(x k)A(k)
First we want to prove that n 1is equal to ( 1)n+1Pk
i=1A(i)iand to do this we will use a straight
forward counting argument. Induction can be used here as well (try it), but the intuitive approach is a
much stronger technique. Let's imagine creating each term one by one from the extended product. How
do we do this? From each ( x i) we pick either a xor ai. But we are only interested in the terms
that result in xto the power n 1. AsPk
i=1A(i) =n, we havenfactors of the form ( x i). Then
to get terms with xn 1we need to pick x's in every ( x i), except one. Since we have nlinear factors
there arenways to do this, namely each eigenvalue represented as many times as it's algebraic multiplicity.
Now we have to take into account the sign of each term. As we pick n 1x's and one i(which has a
negative sign in the linear factor) we get a factor of 1. Then we have to take into account the ( 1)nin
the characteristic polynomial. Thus n 1is the sum of these terms,
n 1= ( 1)n+1kX
i=1A(i)i
Now we will now show that n 1is also equal to ( 1)n 1t(A). For this we will proceed by induction on the
size ofA. IfAis a 11 square matrix then pA(x) = det (A xIn) = ([A]11 x) and ( 1)1 1t(A) = [A]11.
With our base case in hand let's assume Ais a square matrix of size n. By Denition CP [460]
pA(x) = det (A xIn)
= [A xIn]11det ((A xIn) (1j1)) [A xIn]12det ((A xIn) (1j2)) +
[A xIn]13det ((A xInn) (1j3)) + ( 1)n+1[A xIn]1ndet ((A xIn) (1jn))
First let's consider the maximum degree of [ A xIn]1idet ((A xIn) (1ji)) wheni6= 1. For polynomials,
the degree of f, denotedd(f), is the highest power of xin the expression f(x). A well known result of this
denition is: if f(x) =g(x)h(x) thend(f) =d(g) +d(h) (can you prove this?). Now [ A xIn]1ihas degree
zero wheni6= 1. Furthermore ( A xIn) (1ji) hasn 1 rows, one of which has all of its entries of degree
zero, since column iis removed. The other n 2 rows have one entry with degree one and the remainder
of degree zero. Then by Exercise T.T30 [887], the maximum degree of [ A xIn]1idet ((A xIn) (1ji)) is
n 2. So these terms will not aect the coecient of xn 1. Now we are free to focus all of our attention on
the term [A xIn]11det ((A xIn) (1j1)). AsA(1j1) is a (n 1)(n 1) matrix the induction hypothesis
tells us that det (( A xIn) (1j1)) has a coecient of ( 1)n 2t(A(1j1)) forxn 2. We also note that the
proof of Theorem NEM [485] tells us that the leading coecient of det (( A xIn) (1j1)) is ( 1)n 1. Then,
[A xIn]11det ((A xIn) (1j1)) = ([ A]11 x)
( 1)n 1xn 1+ ( 1)n 2t(A(1j1))xn 2+:::
Expanding the product shows n 1(the coecient of xn 1) to be
n 1= ( 1)n 1[A]11+ ( 1)n 1t(A(1j1))
= ( 1)n 1[A]11+ ( 1)n 1n 1X
k=1[A(1j1)]kk Denition T [883]
= ( 1)n 1
[A]11+n 1X
k=1[A(1j1)]kk
Property DCN [759]
Version 2.30
888 Section T Trace
= ( 1)n 1
[A]11+nX
k=2[A]kk
Denition SM [428]
= ( 1)n 1t(A) Denition T [883]
With two expressions for n 1, we have our result,
t(A) = ( 1)n+1( 1)n 1t(A)
= ( 1)n+1n 1
= ( 1)n+1( 1)n+1kX
i=1A(i)i
=kX
i=1A(i)i
Version 2.30
Subsection T.EXC Exercises 889
Subsection EXC
Exercises
T10 Prove there are no square matrices AandBsuch thatAB BA=In.
Contributed by Andy Zimmer
T12 AssumeAis a square matrix of size nmatrix. Prove t(A) =t
At
.
Contributed by Andy Zimmer
T20 IfTn=fM2Mnnjt(M) = 0gthen prove Tnis a subspace of Mnnand determine it's dimension.
Contributed by Andy Zimmer
T30 AssumeAis annmatrix with polynomial entries. Dene md(A;i) to be the maximum degree of
the entries in row i. Thend(det (A))md(A;1)+md(A;2)+:::+md(A;n). (Hint: If f(x) =h(x)+g(x),
thend(f)maxfd(h);d(g)g.)
Contributed by Andy Zimmer Solution [888]
T40 IfAis a square matrix, the matrix exponential is dened as
eA=1X
i=0Ai
i!
Prove that det
eA
=et(A). (You might want to give some thought to the convergence of the innite sum
as well.)
Contributed by Andy Zimmer
Version 2.30
890 Section T Trace
Subsection SOL
Solutions
T30 Contributed by Andy Zimmer Statement [887]
We will proceed by induction. If Ais a square matrix of size 1, then clearly d(det (A))md(A;1). Now
assumeAis a square matrix of size nthen by Theorem DER [429],
det (A) = ( 1)2[A]1;1det (A(1j1)) + ( 1)3[A]1;2det (A(1j2))
+ ( 1)4[A]1;3det (A(1j3)) ++ ( 1)n+1[A]1;ndet (A(1jn))
Let's consider the degree of term j, ( 1)1+j[A]1;jdet (A(1jj)). By denition of the function md,d([A]1;j)
md(A;j). We use our induction hypothesis to examine the other part of the product which tells us that
d(det (A(1jj)))md(A(1jj);1) +md(A(1jj);2) ++md(A(1jj);n 1)
Furthermore by denition of A(1jj) (Denition SM [428]) row iof matrixAcontains all the entries of the
corresponding row in A(1jj) then,
md(A(1jj);1)md(A;1)
md(A(1jj);2)md(A;2)
...
md(A(1jj);j 1)md(A;j 1)
md(A(1jj);j)md(A;j+ 1)
...
md(A(1jj);n 1)md(A;n)
So,
d(det (A(1jj)))md(A(1jj);1) +md(A(1jj);2) ++md(A(1jj);n 1)
md(A;1) +md(A;2) ++md(A;j 1) +md(A;j+ 1) ++md(A;n 1)
Then using the property that if f(x) =g(x)h(x) thend(f) =d(g) +d(h),
d
( 1)1+j[A]1;jdet (A(1jj))
=d
[A]1;j
+d(det (A(1jj)))
md(A;j) +md(A;1) +md(A;2) ++
md(A;j 1) +md(A;j+ 1) ++md(A;n)
=md(A;1) +md(A;2) ++md(A;n)
Asjis arbitrary the degree of all terms in the determinant are so bounded. Finally using the fact that if
f(x) =g(x) +h(x) thend(f)maxfd(h);d(g)gwe have
d(det (A))md(A;1) +md(A;2) ++md(A;n)
Version 2.30
Section HP Hadamard Product 891
Section HP
Hadamard Product
This section is contributed by Elizabeth Million.
You may have once thought that the natural denition for matrix multiplication would be entrywise
multiplication, much in the same way that a young child might say, \I writed my name." The mistake is
understandable, but it still makes us cringe. Unlike poor grammar, however, entrywise matrix multiplica-
tion has reason to be studied; it has nice properties in matrix analysis and additionally plays a role with
relative gain arrays in chemical engineering, covariance matrices in probability and serves as an inertia
preserver for Hermitian matrices in physics. Here we will only explore the properties of the Hadamard
product in matrix analysis.
Denition HP
Hadamard Product
LetAandBbemnmatrices. The Hadamard Product ofAandBis dened by [ AB]ij= [A]ij[B]ij
for all 1im, 1jn.
(This denition contains Notation HP.) 4
As we can see, the Hadamard product is simply \entrywise multiplication". Because of this, the
Hadamard product inherits the same benets (and restrictions) of multiplication in C. Note also that
bothAandBneed to be the same size, but not necessarily square. To avoid confusion, juxtaposition of
matrices will imply the \usual" matrix multiplication, and we will use \ " for the Hadamard product.
Example HP
Hadamard Product
Consider
A=1 0 6
35
B=3 13i
1
32 4
Then
AB=(1)(3) (0)(13) (6)( i)
3(1
3) ()(2) (5)(4)
=3 0 6i
1 220
:
Now we will explore some basics properties of the Hadamard Product.
Theorem HPC
Hadamard Product is Commutative
IfAandBaremnmatrices then AB=BA.
Proof The proof follows directly from the fact that multiplication in Cis commutative. Let AandBbe
mnmatrices. Then
[AB]ij= [A]ij[B]ij Denition HP [889]
= [B]ij[A]ij Property CMCN [758]
= [BA]ij Denition HP [889]
Version 2.30
892 Section HP Hadamard Product
With equality of each entry of the matrices being equal we know by Denition ME [207] that the two
matrices are equal.
Denition HID
Hadamard Identity
TheHadamard identity is themnmatrixJmndened by [ Jmn]ij= 1 for all 1im, 1jn.
(This denition contains Notation HID.) 4
Theorem HPHID
Hadamard Product with the Hadamard Identity
SupposeAis anmnmatrix. Then AJmn=JmnA=A.
Proof
[AJmn]ij= [JmnA]ij Theorem HPC [889]
= [Jmn]ij[A]ij Denition HP [889]
= (1) [A]ij Denition HID [890]
= [A]ij Property OCN [759]
With equality of each entry of the matrices being equal we know by Denition ME [207] that the two
matrices are equal.
Denition HI
Hadamard Inverse
LetAbe anmnmatrix and suppose [ A]ij6= 0 for all 1im, 1jn. Then the Hadamard
Inverse ,bA, is given byh
bAi
ij= ([A]ij) 1for all 1im, 1jn.
(This denition contains Notation HI.) 4
Theorem HPHI
Hadamard Product with Hadamard Inverses
LetAbe anmnmatrix such that [ A]ij6= 0 for all 1im, 1jn. ThenAbA=bAA=Jmn.
Proof
h
AbAi
ij=h
bAAi
ijTheorem HPC [889]
=h
bAi
ij[A]ij Denition HP [889]
= ([A]ij) 1[A]ij Denition HI [890], [ A]ij6= 0
= 1 Property MICN [759]
= [Jmn]ij Denition HID [890]
With equality of each entry of the matrices being equal we know by Denition ME [207] that the two
matrices are equal.
Since matrices have a dierent inverse and identity under the Hadamard product, we have used special
notation to distinguish them from what we have been using with \normal" matrix multiplication. That
is, compare \usual" matrix inverse, A 1, with the Hadamard inverse bA, and the \usual" matrix identity,
In, with the Hadamard identity, Jmn. The Hadamard identity matrix and the Hadamard inverse are both
more limiting than helpful, so we will not explore their use further. One last fun fact for those of you
who may be familiar with group theory: the set of mnmatrices with nonzero entries form an abelian
(commutative) group under the Hadamard product (prove this!).
Version 2.30
Subsection HP.DMHP Diagonal Matrices and the Hadamard Product 893
Theorem HPDAA
Hadamard Product Distributes Across Addition
SupposeA,BandCaremnmatrices. Then C(A+B) =CA+CB.
Proof
[C(A+B)]ij= [C]ij[A+B]ij Denition HP [889]
= [C]ij([A]ij+ [B]ij) Denition MA [207]
= [C]ij[A]ij+ [C]ij[B]ij Property DCN [759]
= [CA]ij+ [CB]ij Denition HP [889]
= [CA+CB]ij Denition MA [207]
With equality of each entry of the matrices being equal we know by Denition ME [207] that the two
matrices are equal.
Theorem HPSMM
Hadamard Product and Scalar Matrix Multiplication
Suppose2C, andAandBaremnmatrices. Then (AB) = (A)B=A(B).
Proof
[AB]ij=[AB]ij Denition MSM [208]
=[A]ij[B]ij Denition HP [889]
= [A]ij[B]ij Denition MSM [208]
= [(A)B]ij Denition HP [889]
=[A]ij[B]ij Denition MSM [208]
= [A]ij[B]ij Property CMCN [758]
= [A]ij[B]ij Denition MSM [208]
= [A(B)]ij Denition HP [889]
With equality of each entry of the matrices being equal we know by Denition ME [207] that the two
matrices are equal.
Subsection DMHP
Diagonal Matrices and the Hadamard Product
We can relate the Hadamard product with matrix multiplication by considering diagonal matrices, since
AB=ABif and only if both AandBare diagonal (Citation!!!). For example, a simple calculation reveals
that the Hadamard product relates the diagonal values of a diagonalizable matrix Awith its eigenvalues:
Theorem DMHP
Diagonalizable Matrices and the Hadamard Product
LetAbe a diagonalizable matrix of size nwith eigenvalues 1; 2; 3; :::; n. LetDbe a diagonal matrix
from the diagonalization of A,A=SDS 1, and dbe a vector such that [ D]ii=[d]i=ifor all 1in.
Then
[A]ii=
S(S 1)td
ifor all 1in:
Version 2.30
894 Section HP Hadamard Product
That is,2
666664[A]11
[A]22
[A]33...
[A]nn3
777775=S(S 1)t2
6666641
2
3
...
n3
777775
Proof
S(S 1)td
i=nX
k=1
S(S 1)t
ik[d]k Denition MVP [223]
=nX
k=1
S(S 1)t
ikk Denition of d
=nX
k=1[S]ik
(S 1)t
ikk Denition HP [889]
=nX
k=1[S]ik
S 1
kik Denition TM [210]
=nX
k=1[S]ikk
S 1
kiProperty CMCN [758]
=nX
k=1[S]ik[D]kk
S 1
kiDenition of D
=nX
j=1nX
k=1[S]ik[D]kj
S 1
ji[D]kj= 0 for allk6=j
=nX
j=1[SD]ij
S 1
jiTheorem EMP [227]
=
SDS 1
iiTheorem EMP [227]
= [A]ii Denition ME [207]
With equality of each entry of the matrices being equal we know by Denition ME [207] that the two
matrices are equal.
We obtain a similar result when we look at the singular value decomposition of square matrices (see
exercises).
Theorem DMMP
Diagonal Matrices and Matrix Products
SupposeA,Baremnmatrices, and DandEare diagonal matrices of size mandn, respectively. Then,
D(AB)E= (DAE )B= (DA)(BE)
Proof
[D(AB)E]ij=mX
k=1[D]ik[(AB)E]kj Theorem EMP [227]
Version 2.30
Subsection HP.DMHP Diagonal Matrices and the Hadamard Product 895
=mX
k=1nX
l=1[D]ik[AB]kl[E]lj Theorem EMP [227]
=mX
k=1nX
l=1[D]ik[A]kl[B]kl[E]lj Denition HP [889]
=mX
k=1[D]ik[A]kj[B]kj[E]jj [E]lj= 0 for alll6=j
= [D]ii[A]ij[B]ij[E]jj [D]ik= 0 for alli6=k
= [D]ii[A]ij[E]jj[B]ij Property CMCN [758]
= [D]ii(nX
l=1[A]il[E]lj) [B]ij [E]lj= 0 for alll6=j
= [D]ii[AE]ij[B]ij Theorem EMP [227]
= (mX
k=1[D]ik[AE]kj) [B]ij [D]ik= 0 for alli6=k
= [DAE ]ij[B]ij Theorem EMP [227]
= [(DAE )B]ij Denition HP [889]
With equality of each entry of the matrices being equal we know by Denition ME [207] that the two
matrices are equal.
Also,
[(DAE )B]ij= [DAE ]ij[B]ij Denition HP [889]
= (nX
k=1[DA]ik[E]kj) [B]ij Theorem EMP [227]
= [DA]ij[E]jj[B]ij [E]kj= 0 for allk6=j
= [DA]ij[B]ij[E]jj Property CMCN [758]
= [DA]ij(nX
k=1[B]ik[E]kj) [ E]kj= 0 for allk6=j
= [DA]ij[BE]ij Theorem EMP [227]
= [(DA)(BE)]ij Denition HP [889]
With equality of each entry of the matrices being equal we know by Denition ME [207] that the two
matrices are equal.
Version 2.30
896 Section HP Hadamard Product
Subsection EXC
Exercises
T10 Prove that AB=ABif and only if both AandBare diagonal matrices.
Contributed by Elizabeth Million
T20 SupposeA,Baremnmatrices, and DandEare diagonal matrices of size mandn, respectively.
Prove both parts of the following equality hold:
D(AB)E= (AE)(DB) =A(DBE )
Contributed by Elizabeth Million
T30 LetAbe a square matrix of size nwith singular values 1; 2; 3; :::; n. LetDbe a diagonal
matrix from the singular value decomposition of A,A=UDV(Theorem SVD [921]). Dene the vector
dby [d]i= [D]ii=i, 1in. Prove the following equality,
[A]ii=
(UV)d
i
Contributed by Elizabeth Million
T40 SupposeA,BandCaremnmatrices. Prove that for all 1 im,
(AB)Ct
ii=
(AC)Bt
ii
Contributed by Elizabeth Million
T50 Dene the diagonal matrix Dof sizenwith entries from a vector x2Cnby
[D]ij=(
[x]iifi=j
0 otherwise
Furthermore, suppose A,Baremnmatrices. Prove that
ADBt
ii= [(AB)x]ifor all 1im.
Contributed by Elizabeth Million
Version 2.30
Section VM Vandermonde Matrix 897
Section VM
Vandermonde Matrix
This Section is a Draft, Subject to Changes
Alexandre-Th eophile Vandermonde was a French mathematician in the 1700's who was among the rst
to write about basic properties of the determinant (such as the eect of swapping two rows). However,
the determinant that bears his name (Theorem DVM [895]) does not appear in any of his four published
mathematical papers.
Denition VM
Vandermonde Matrix
An square matrix of size n,A, is a Vandermonde matrix if there are scalars, x1; x2; x3; :::; xnsuch
that [A]ij=xj 1
i, 1in, 1jn. 4
Example VM4
Vandermonde matrix of size 4
A=2
6641 2 4 8
1 3 9 27
1 1 1 1
1 4 16 643
775
is a Vandermonde matrix since it meets the denition with x1= 2,x2= 3,x3= 1,x4= 4.
Vandermonde matrices are not very interesting as numerical matrices, but instead appear more often
in proofs and applications where the scalars xiare carried as symbols. Two such applications are in the
sections on secret-sharing (Section SAS [937]) and curve-tting (Section CF [931]). Principally, we would
like to know when Vandermonde matrices are nonsingular, and the most convenient way to check this is
by determining when the determinant is nonzero (Theorem SMZD [445]). As a bonus, the determinant of
a Vandermonde matrix has an especially pleasing formula.
Theorem DVM
Determinant of a Vandermonde Matrix
Suppose that Ais a Vandermonde matrix of size nbuilt with the scalars x1; x2; x3; :::; xn. Then
det (A) =Y
1i<jn(xj xi)
Proof The proof is by induction (Technique I [772]) on n, the size of the matrix. An empty product for
a 11 matrix might make a good base case, but we'll start at n= 2 instead. For a 2 2 Vandermonde
matrix, we have
det (A) =1x1
1x2=x2 x1=Y
1i<j2(xj xi)
For the induction step we will perform row operations on Ato obtain the determinant of Aas multiple of
the determinant of an ( n 1)(n 1) Vandermonde matrix. the notation in this theorem tens to obscure
your intuition about the changes eected by various row and column manipulations. Construct a 4 4
Version 2.30
898 Section VM Vandermonde Matrix
Vandermonde matrix with four symbols as the scalars ( x1,x2,x2,x4, or perhaps a,b,c,d) and play along
with the example as you study the proof.
First we convert most of the rst column to zeros. Subtract row nfrom each of the other n 1 rows
to form a matrix B. By Theorem DRCMA [441], Bhas the same determinant as A. The entries of B, in
the rstn 1 rows, i.e. for 1in 1, 1jn 1, are
[B]ij=xj 1
i xj 1
n= (xi xn)j 2X
k=0xj 2 k
ixk
n
As the elements of row i, 1in 1, have the common factor ( xi xn), we form the new matrix
Cthat diers from Bby the removal of this factor from each of the rst n 1 rows. This will change
the determinant, as we will track carefully in a moment. We also have a rst column with zeros in each
location, except row n, so we can use it for a column expansion computation of the determinant. We now
know,
det (A) = det (B) Theorem DRCMA [441]
= (x1 xn)(x2 xn)(xn 1 xn) det (C) Theorem DRCM [440]
= (x1 xn)(x2 xn)(xn 1 xn)(1)( 1)n+1det (C(n 1j1)) Theorem DEC [431]
= (x1 xn)(x2 xn)(xn 1 xn)( 1)n 1det (C(n 1j1))
= (xn x1)(xn x2)(xn xn 1) det (C(n 1j1))
For convenience, denote D=C(n 1j1). Entries of this matrix are similar to those of B, but the factors
used to build Care gone, and since the rst column is gone, there is a slight re-indexing relative to the
columns. For 1in 1, 1jn 1,
[D]ij=j 1X
k=0xj 1 k
ixk
n
We will perform many column operations on the matrix D, always of the type where we multiply a
column by a scalar and add the result to another column. As such, Theorem DRCM [440] insures that
the determinant will remain constant. We will work column by column, left to right, to convert Dinto a
Vandermonde matrix with scalars x1; x2; x3; :::; xn 1. More precisely, we will build a sequence of matrices
D=D1,D2, . . . ,Dn 1, where each obtainable from the previous by a sequence of determinant-preserving
column operations and the rst `columns of D`are the rst `columns of a Vandermonde matrix with
scalarsx1; x2; x3; :::; xn 1. We could establish this claim by induction (Technique I [772]) on `if we were
to expand the claim to specify the exact values of the nal n 1 `columns as well. Since the claim is
that matrices with certain properties exist, we will instead establish the claim by constructing the desired
matrices one-by-one procedurally. The extension to an inductive proof should be clear, but not especially
illuminating.
SetD1=Dto begin, and note that the entries of the rst column of D1are, for 1in 1,
[D1]i1=1 1X
k=0x1 1 k
ixk
n= 1 =x1 1
i
So the rst column of D1has the properties we desire. We will use this column of all 1's to remove the
highest power of xnfrom each of the remaining columns and so build D2. Precisely, perform the n 2
column operations where column 1 is multiplied by xj 1
nand subtracted from column j, for 2jn 1.
Call the result D2, and examine its entries in columns 2 through n 1. For 1in 1, 2jn 1,
[D2]ij= xj 1
n[D1]i1+ [D1]ij
Version 2.30
Section VM Vandermonde Matrix 899
= xj 1
n(1) +j 1X
k=0xj 1 k
ixk
n
= xj 1
n+xj 1 (j 1)
ixj 1
n+j 2X
k=0xj 1 k
ixk
n
=j 2X
k=0xj 1 k
ixk
n
In particular, we examine column 2 of D2. For 1in 1,
[D2]i2=2 2X
k=0x2 1 k
ixk
n=x1
i=x2 1
i
Now, form D3. Perform the n 3 column operations where column 2 of D2is multiplied by xj 2
nand
subtracted from column j, for 3jn 1. The result is D3, whose entries we now compute. For
1in 1,
[D3]ij= xj 2
n[D2]i2+ [D2]ij
= xj 2
nx1
i+j 2X
k=0xj 1 k
ixk
n
= xj 2
nx1
i+xj 1 (j 2)
ixj 2
n+j 3X
k=0xj 1 k
ixk
n
=j 3X
k=0xj 1 k
ixk
n
Specically, we examine column 3 of D3. For 1in 1,
[D3]i3=3 3X
k=0x3 1 k
ixk
n=x2
i=x3 1
i
We could continue this procedure n 4 more times, eventually totaling1
2
n2 3n+ 2
column operations,
and arriving at Dn 1, the Vandermonde matrix of size n 1 built from the scalars x1; x2; x3; :::; xn 1.
Informally, we chop o the last term of every sum, until a single term is left in a column, and it is of the
right form for the Vandermonde matrix. This desired column is then used in the next iteration to chop
o some more nal terms for columns to the right. Now we can apply our induction hypothesis to the
determinant of Dn 1and arrive at an expression for det A,
det (A) = det (C)
=n 1Y
k=1(xn xk) det (D)
=n 1Y
k=1(xn xk) det (Dn 1)
=n 1Y
k=1(xn xk)Y
1i<jn 1(xj xi)
=Y
1i<jn(xj xi)
Version 2.30
900 Section VM Vandermonde Matrix
which is the desired result.
Before we had Theorem DVM [895] we could see that if two of the scalar values were equal, then the
Vandermonde matrix would have two equal rows and hence be singular (Theorem DERC [441], Theorem
SMZD [445]). But with this expression for the determinant, we can establish the converse.
Theorem NVM
Nonsingular Vandermonde Matrix
A Vandermonde matrix of size nwith scalars x1; x2; x3; :::; xnis nonsingular if and only if the scalars
are all dierent.
Proof LetAdenote the Vandermonde matrix with scalars x1; x2; x3; :::; xn. By Theorem SMZD [445],
Ais nonsingular if and only if the determinant of Ais nonzero. The determinant is given by Theorem
DVM [895], and this product is nonzero if and only if each term of the product is nonzero. This condition
translates to xi xj6= 0 whenever i6=j. In other words, the matrix is nonsingular if and only if the scalars
are all dierent.
Version 2.30
Section PSM Positive Semi-denite Matrices 901
Section PSM
Positive Semi-denite Matrices
This Section is a Draft, Subject to Changes
Needs Numerical Examples
Positive semi-denite matrices (and their cousins, positive denite matrices) are square matrices which
in many ways behave like non-negative (respectively, positive) real numbers. Results given here are em-
ployed in the decompositions of Section SVD [917], Section SR [923] and Section PD [407].
Subsection PSM
Positive Semi-Denite Matrices
Denition PSM
Positive Semi-Denite Matrix
A square matrix Aof sizenispositive semi-denite ifAis Hermitian and for all x2Cn,hAx;xi0.
4
For a denition of positive denite replace the inequality in the denition with a strict inequality,
and exclude the zero vector from the vectors xrequired to meet the condition. Similar variations allow
denitions of negative denite andnegative semi-denite . Our rst theorem in this section gives us
an easy way to build positive semi-denite matrices.
Theorem CPSM
Creating Positive Semi-Denite Matrices
Suppose that Ais anymnmatrix. Then the matrices AAandAAare positive semi-denite matrices.
Proof We will give the proof for the rst matrix, the proof for the second is entirely similar. First we
check thatAAis Hermitian,
(AA)=A(A)Theorem MMAD [233]
=AA Theorem AA [215]
so by Denition HM [234], the matrix AAis Hermitian. Second, for any x2Cn,
hAAx;xi=hAx;(A)xi Theorem AIP [233]
=hAx; Axi Theorem AA [215]
0 Theorem PIP [196]
which is the second criteria in the denition of a positive semi-denite matrix (Denition PSM [899]).
A statement very similar to the converse of this theorem is also true. Any positive semi-denite matrix
can be realized as the product of a square matrix, B, with its adjoint, B. (See Exercise PSM.T20 [902]
after studying this entire section.) The matrices AAandAAwill be important later when we dene
singular values (Section SVD [917]).
Positive semi-denite matrices can also be characterized by their eigenvalues, without any mention of
inner products. This next result further reinforces the notion that positive semi-denite matrices behave
like non-negative real numbers.
Version 2.30
902 Section PSM Positive Semi-denite Matrices
Theorem EPSM
Eigenvalues of Positive Semi-denite Matrices
Suppose that Ais a Hermitian matrix. Then Ais positive semi-denite matrix if and only if whenever
is an eigenvalue of A, then0.
Proof Notice rst that since we are considering only Hermitian matrices in this theorem, it is always
possible to compare eigenvalues with the real number zero, since eigenvalues of Hermitian matrices are all
real numbers (Theorem HMRE [487]). Let ndenote the size of A.
()) Let x6= 0 be an eigenvector of Afor. Then by Theorem PIP [196] we know hx;xi6= 0. So
=1
hx;xihx;xi Property MICN [759]
=1
hx;xihx;xi Theorem IPSM [194]
=1
hx;xihAx;xi Denition EEM [453]
By Theorem PIP [196], hx;xi>0 and by Denition PSM [899] we have hAx;xi0. Withexpressed
as the product of these two quantities, we have 0.
(() Suppose now that 1; 2; 3; :::; nare the (not necessarily distinct) eigenvalues of the Her-
mitian matrix A, each of which is non-negative. Let B=fx1;x2;x3; :::; xngbe a set of associated
eigenvectors for these eigenvalues. Since a Hermitian matrix is normal (Denition HM [234], Denition
NM [83]), Theorem OBNM [683] allows us to choose this set of eigenvectors to also be an orthonormal basis
ofCn. Choose any x2Cnand leta1; a2; a3; :::; anbe the scalars guaranteed by the spanning property
of the basis Bsuch that
x=a1x1+a2x2+a3x3++anxn=nX
i=1aixi
Since we have presumed Ais Hermitian, we need only check the other dening property,
hAx;xi=*
AnX
i=1aixi;nX
j=1ajxj+
Denition TSVS [356]
=*nX
i=1Aaixi;nX
j=1ajxj+
Theorem MMDAA [230]
=*nX
i=1aiAxi;nX
j=1ajxj+
Theorem MMSMM [230]
=*nX
i=1aiixi;nX
j=1ajxj+
Denition EEM [453]
=nX
i=1nX
j=1haiixi; ajxji Theorem IPVA [193]
=nX
i=1nX
j=1aiiajhxi;xji Theorem IPSM [194]
=nX
i=1aiiaihxi;xii+nX
i=1nX
j=1
j6=iaiiajhxi;xji Property CACN [758]
Version 2.30
Subsection PSM.PSM Positive Semi-Denite Matrices 903
=nX
i=1aiiai(1) +nX
i=1nX
j=1
j6=iaiiaj(0) Denition ONS [201]
=nX
i=1aiiai
=nX
i=1ijaij2Denition MCN [760]
With non-negative values for each eigenvalue i, 1in, and each modulus squared, it should be clear
that this sum is non-negative. Which is exactly what is required by Denition PSM [899] to establish that
Ais positive semi-denite.
As positive semi-denite matrices are dened to be Hermitian, they are then normal and subject to
orthonormal diagonalization (Theorem OD [681]). Now consider the interpretation of orthonormal diago-
nalization as a rotation to principal axes, a stretch by a diagonal matrix and a rotation back (Subsection
OD.OD [681]). For a positive semi-denite matrix, the diagonal matrix has diagonal entries that are the
non-negative eigenvalues of the original positive semi-denite matrix. So the \stretching" along each axis
is never a re
ection.
Version 2.30
904 Section PSM Positive Semi-denite Matrices
Subsection EXC
Exercises
T20 Suppose that Ais a positive semi-denite matrix of size n. Prove that there is a square matix Bof
sizensuch thatA=BB.
Contributed by Robert Beezer
Version 2.30
Chapter MD
Matrix Decompositions
This chapter is about breaking up a matrix Ainto pieces that somehow combine to recreate A. Usually
the pieces are again matrices, and usually they are then combined via matrix multiplication (Denition
MM [226]). In some cases, the decomposition will be valid for any matrix, but often we might need
extra conditions on A, such as being square (Denition SQM [83]), nonsingular (Denition NM [83]) or
diagonalizable (Denition DZM [496]) before we can guarantee the decomposition. If you are comfortable
with topics like decomposing a solution vector into linear combinations (Subsection LC.VFSS [113]) or
decomposing vector spaces into direct sums (Subsection PD.DS [413]), then we will be doing similar things
in this chapter. If not, review these ideas and take another look at Technique DC [772] on decompositions.
We have studied one matrix decomposition already, so we will review that here in this introduction,
both as a way of previewing the topic in a familiar setting, but also since it does not deserve another
section all of its own.
A diagonalizable matrix (Denition DZM [496]) is dened to be a square matrix Asuch that there is
an invertible matrix Sand a diagonal matrix DwhereS 1AS=D. We can re-write this as A=SDS 1.
Here we have a decomposition of Ainto three matrices, S,DandS 1, which recombine through matrix
multiplication to recreate A. We also know that the diagonal entries of Dare the eigenvalues of A. We
cannot form this decomposition for just any matrix | Amust be square and we know from Theorem
DC [497] that a matrix of size nis diagonalizable if and only if there is a basis for Cncomposed entirely
of eigenvectors of A, or by Theorem DMFE [499] we know that Ais diagonalizable if and only if each
eigenvalue of Ahas a geometric multiplicity equal to its algebraic multiplicity. Some authors prefer to call
this an eigen decomposition ofArather than a matrix diagonalization .
Another decomposition, which is similar in
avor to matrix diagonalization, is orthonormal diagonal-
ization (Theorem OD [681]). Here we require the matrix Ato be normal and we get the decomposition
A=UDU, whereDis a diagonal matrix with the eigenvalues of Aon the diagonal, and Uis unitary.
The hypothesis that Ais normal guarantees the decomposition and we get the extra information that U
is unitary.
Each section of this chapter features a dierent matrix decomposition, with the exception of Section
PSM [899], which presents background information on positive semi-denite matrices required for singular
value decompositions, square roots and polar decompositions.
Section ROD
Rank One Decomposition
This Section is a Draft, Subject to Changes
Our rst decomposition applies only to diagonalizable (Denition DZM [496]) matrices, and yields a
905
906 Section ROD Rank One Decomposition
decomposition into a sum of very simple matrices.
Theorem ROD
Rank One Decomposition
Suppose that Ais a diagonalizable matrix of size nand rankr. Then there are rsquare matrices
A1; A2; A3; :::; Ar, each of size nand rank 1 such that
A=A1+A2+A3++Ar
Furthermore, if 1; 2; 3; :::; rare the nonzero eigenvalues of A, then there are two sets of rlinearly
independent vectors from Cn,
X=fx1;x2;x3; :::; xrg Y=fy1;y2;y3; :::; yrg
such thatAk=kxkyt
k, 1kr.
Proof The proof is constructive. Generally, we will diagonalize A, creating a nonsingular matrix Sand
a diagonal matrix D. Then we split up the diagonal matrix into a sum of matrices with a single nonzero
entry (on the diagonal). This fundamentally creates the decomposition in the statement of the theorem,
the remainder is just bookkeeping. The vectors in XandYwill result from the columns of Sand the rows
ofS 1.
Let1; 2; 3; :::; nbe the eigenvalues of A(repeated according to their algebraic multiplicity). If
Ahas rankr, then dim (N(A)) =n r(Theorem RPNC [398]). The null space of Ais the eigenspace of
the eigenvalue = 0 (Theorem EMNS [462]), so it follows that the algebraic multiplicity of = 0 isn r,
A(0) =n r. Presume that the complete list of eigenvalues is ordered so that k= 0 forr+ 1kn.
SinceAis hypothesized to be diagonalizable, there exists a diagonal matrix Dand an invertible matrix
S, such that D=S 1AS. We can rearrange tis equation to read, A=SDS 1. Also, the proof of Theorem
DC [497] says that the diagonal elements of Dare the eigenvalues of Aand we have the
exibility to assume
they lie on the diagonal in the same order as we have specied above. Now, let X=fx1;x2;x3; :::; xng
be the columns of S, and letY=fy1;y2;y3; :::; yngbe the rows of S 1converted to column vectors.
With little motivation other than the statement of the theorem, dene size nmatricesAk, 1knby
Ak=kxkyt
k. Finally, let Dkbe the size nmatrix that is totally zero, other than having kin rowkand
columnk.
With everything in place, we compute entry-by-entry,
[A]ij=
SDS 1
ijDenition DZM [496]
="
S nX
k=1Dk!
S 1#
ijDenition MA [207]
="
S nX
k=1DkS 1!#
ijTheorem MMDAA [230]
="nX
k=1SDkS 1#
ijTheorem MMDAA [230]
=nX
k=1
SDkS 1
ijDenition MA [207]
=nX
k=1nX
`=1[SDk]i`
S 1
`jTheorem EMP [227]
=nX
k=1nX
`=1nX
p=1[S]ip[Dk]p`
S 1
`jTheorem EMP [227]
Version 2.30
Section ROD Rank One Decomposition 907
=nX
k=1[S]ik[Dk]kk
S 1
kj[Dk]p`= 0 ifp6=k, or`6=k
=nX
k=1[S]ikk
S 1
kj[Dk]kk=k
=nX
k=1k[S]ik
S 1
kjProperty CMCN [758]
=nX
k=1k[xk]i1
yt
k
1jDenition of X,Y
=nX
k=1k1X
q=1[xk]iq
yt
k
qj
=nX
k=1k
xkyt
k
ijTheorem EMP [227]
=nX
k=1
kxkyt
k
ijDenition MSM [208]
=nX
k=1[Ak]ij Denition of Ak
="nX
k=1Ak#
ijDenition MA [207]
So by Denition ME [207] we have the desired equality of matrices. The careful reader will have noted that
Ak=O,r+ 1kn, sincek= 0 in these instances. To get the sets XandYfromXandY, simply
discard the last n rvectors. We can safely ignore (or remove) Ar+1; Ar+2; :::; Anfrom the summation
just derived.
One last assertion to check. What is the rank of Ak, 1kr? Every row of Akis a scalar multiple
ofyt
k, rowkof the nonsingular matrix S 1(Theorem MIMI [251]). As a row of a nonsingular matrix, yt
k
cannot be all zeros. In particular, row iofAkis obtained as a scalar multiple of yt
kby the scalar k[xk]i.
We have restricted ourselves to the nonzero eigenvalues of A, and asSis nonsingular, some entry of xk
is nonzero. This all implies that some row of Akwill be nonzero. Now consider row-reducing Ak. Swap
the nonzero row up into row 1. Use scalar multiples of this row to zero out every other row. This leaves a
single nonzero row in the reduced row-echelon form, so Akhas rank one.
We record two observations that was not stated in our theorem above. First, the vectors in X, chosen
as columns of S, are eigenvectors of A. Second, the product of two vectors from XandYin the opposite
order, by which we mean yt
ixj, is the entry in row iand column jof the matrix product S 1S=In
(Theorem EMP [227]). In particular,
yt
ixj=(
1 ifi=j
0 ifi6=j
We give two computational examples. One small, one a bit bigger.
Example ROD2
Rank one decomposition, size 2
Version 2.30
908 Section ROD Rank One Decomposition
Consider the 22 matrix,
A= 16 6
45 17
By the techniques of Chapter E [453] we nd the eigenvalues and eigenspaces,
1= 2EA(2) = 1
3
2= 1EA( 1) = 2
5
Withn= 2 distinct eigenvalues, Theorem DED [501] tells us that Ais diagonalizable, and with no zero
eigenvalues we see that Ahas full rank. Theorem DC [497] says we can construct the nonsingular matrix
Swith eigenvectors of Aas columns, so we have
S= 1 2
3 5
S 1=5 2
3 1
From these matrices we obtain the sets of vectors
X= 1
3
; 2
5
Y=5
2
; 3
1
And we have the matrices,
A1= 2 1
35
2t
= 2 5 2
15 6
= 10 4
30 12
A2= ( 1) 2
5 3
1t
= ( 1)6 2
15 5
= 6 2
15 5
And you can easily verify that A=A1+A2.
Here's a slightly larger example, and the matrix does not have full rank.
Example ROD4
Rank one decomposition, size 4
Consider the 44 matrix,
B=2
66434 18 1 6
44 24 1 9
36 18 3 6
36 18 6 33
775
By the techniques of Chapter E [453] we nd the eigenvalues and eigenvectors,
1= 3 EB(3) =*8
>><
>>:2
6641
2
1
13
775;2
6641
1
1
23
7759
>>=
>>;+
2= 2 EB( 2) =*8
>><
>>:2
664 1
2
0
03
7759
>>=
>>;+
3= 0 EA(0) =*8
>><
>>:2
6642
3
2
23
7759
>>=
>>;+
Version 2.30
Section ROD Rank One Decomposition 909
The algebraic and geometric multiplicities of each eigenvalue are equal, so Theorem DMFE [499] tells us
thatAis diagonalizable. With a single zero eigenvalue we see that Ahas rank 4 1 = 3. Theorem DC
[497] says we can construct the nonsingular matrix Swith eigenvectors of Aas columns, so we have
S=2
6641 1 1 2
2 1 2 3
1 1 0 2
1 2 0 23
775S 1=2
6644 2 0 1
8 4 1 1
1 0 1 0
6 3 1 13
775
Sincer= 3, we need only collect three vectors from each of these matrices,
X=8
>><
>>:2
6641
2
1
13
775;2
6641
1
1
23
775;2
664 1
2
0
03
7759
>>=
>>;Y=8
>><
>>:2
6644
2
0
13
775;2
6648
4
1
13
775;2
664 1
0
1
03
7759
>>=
>>;
And we obtain the matrices,
B1= 32
6641
2
1
13
7752
6644
2
0
13
775t
= 32
6644 2 0 1
8 4 0 2
4 2 0 1
4 2 0 13
775=2
66412 6 0 3
24 12 0 6
12 6 0 3
12 6 0 33
775
B2= 32
6641
1
1
23
7752
6648
4
1
13
775t
= 32
6648 4 1 1
8 4 1 1
8 4 1 1
16 8 2 23
775=2
66424 12 3 3
24 12 3 3
24 12 3 3
48 24 6 63
775
B3= ( 2)2
664 1
2
0
03
7752
664 1
0
1
03
775t
= ( 2)2
6641 0 1 0
2 0 2 0
0 0 0 0
0 0 0 03
775=2
664 2 0 2 0
4 0 4 0
0 0 0 0
0 0 0 03
775
Then we verify that
B=B1+B2+B3
=2
66412 6 0 3
24 12 0 6
12 6 0 3
12 6 0 33
775+2
66424 12 3 3
24 12 3 3
24 12 3 3
48 24 6 63
775+2
664 2 0 2 0
4 0 4 0
0 0 0 0
0 0 0 03
775
=2
66434 18 1 6
44 24 1 9
36 18 3 6
36 18 6 33
775
Version 2.30
910 Section ROD Rank One Decomposition
Version 2.30
Section TD Triangular Decomposition 911
Section TD
Triangular Decomposition
This Section is a Draft, Subject to Changes
Our next decomposition will break a square matrix into a product of two matrices, one lower triangular
and the other upper triangular. So we will write A=LU, and hence many refer to this as LU decompo-
sition . We will see that this decomposition is very easy to compute and that it has a direct application
to solving systems of equations. Since this section is about triangular matrices you might want to review
the denitions and a couple of basic theorems back in Subsection OD.TM [675].
Subsection TD
Triangular Decomposition
With a slight condition on the nonsingularity of certain submatrices, we can split a matrix into a product
of two triangular matrices.
Theorem TD
Triangular Decomposition
SupposeAis a square matrix of size n. LetAkbe thekkmatrix formed from Aby taking the rst k
rows and the rst kcolumns. Suppose that Akis nonsingular for all 1 kn. Then there is a lower
triangular matrix Lwith all of its diagonal entries equal to 1 and an upper triangular matrix Usuch that
A=LU. Furthermore, this decomposition is unique.
Proof We will row reduce Ato a row-equivalent upper triangular matrix through a series of row operations,
forming intermediate matrices A0
j, 1jn, that denote the state of the conversion after working on
columnj. First, the lone entry of A1is [A]11and this scalar must be nonzero if A1is nonsingular (Theorem
SMZD [445]). We can use row operations Denition RO [31] of the form R1+Rk, 2kn, where
= [A]1k=[A]11to place zeros in the rst column below the diagonal. The rst two rows and columns
ofA0
1are a 22 upper triangular matrix whose determinant is equal to the determinant of A2, since
the matrices are row-equivalent through a sequence of row operations strictly of the third type (Theorem
DRCMA [441]). As such the diagonal entries of this 2 2 submatrix of A0
1are nonzero. We can employ
this nonzero diagonal element with row operations of the form R2+Rk, 3knto place zeros below
the diagonal in the second column. We can continue this process, column by column. The key observations
are that our hypothesis on the nonsingularity of the Akwill guarantee a nonzero diagonal entry for each
column when we need it, that the row operations employed are always of the third type using a multiple
of a row to transform another row with a greater row index , and that the nal result will be a nonsingular
upper triangular matrix. This is the desired matrix U.
Each row operation described in the previous paragraph can be accomplished with matrix multiplication
by the appropriate elementary matrix (Theorem EMDRO [425]). Since every row operation employed is
adding a multiple of a row to a subsequent row these elementary matrices are of the form Ej;k() with
j <k . By Denition ELEM [423], these matrices are lower triangular with every diagonal entry equal to 1.
We know that the product of two such matrices will again be lower triangular (Theorem PTMT [675]), but
also, as you can also easily check using a proof with a style similar to one above, that the product maintains
all 1's on the diagonal. Let E1; E2; E3; :::; Emdenote the elementary matrices for this sequence of row
operations. Then
U=EmEm 1:::E 3E2E1A=L0A
Version 2.30
912 Section TD Triangular Decomposition
whereL0is the product of the elementary matrices, and we know L0is lower triangular with all 1's on the
diagonal. Our desired matrix Lis thenL= (L0) 1. By Theorem ITMT [676], Lis lower triangular with
all 1's on the diagonal and A=LU, as desired.
The process just described is deterministic. That is, the proof is constructive, with no freedom for each
of us to walk through it dierently. But could there be other matrices with the same properties as Land
Uthat give such a decomposition of A. In other words, is the decomposition unique (Technique U [771])?
Suppose that we have two triangular decompositions, A=L1U1andA=L2U2. SinceAis nonsingular,
two applications of Theorem NPNT [259] imply that L1; L2; U1; U2are all nonsingular. We have
L 1
2L1=L 1
2InL1 Theorem MMIM [229]
=L 1
2AA 1L1 Denition MI [244]
=L 1
2L2U2(L1U1) 1L1
=L 1
2L2U2U 1
1L 1
1L1 Theorem SS [250]
=InU2U 1
1In Denition MI [244]
=U2U 1
1 Theorem MMIM [229]
Theorem ITMT [676] tells us that L 1
2is lower triangular and has 1's as the diagonal entries. By Theorem
PTMT [675], the product L 1
2L1is again lower triangular, and it is simple to check (as before) that the
diagonal entries of the product are again all 1's. By the entirely similar process we can conclude that the
productU2U 1
1is upper triangular. Because these two products are equal, their common value is a matrix
that is both lower triangular andupper triangular, with all 1's on the diagonal. The only matrix meeting
these three requirements is the identity matrix (Denition IM [84]). So, we have,
In=L 1
2L1)L2=L1 In=U2U 1
1)U1=U2
which establishes the uniqueness of the decomposition.
Studying the proofs of some previous theorems will perhaps give you an idea for an approach to
computing a triangular decomposition. In the proof of Theorem CINM [248] we augmented a nonsingular
matrix with an identity matrix of the same size, and row-reduced until the original matrix became the
identity matrix (as we knew in advance would happen, since we knew Theorem NMRRI [84]). Theorem
PEEF [298] tells us about properties of extended echelon form, and in particular, that B=JA, whereAis
the matrix that begins on the left, and Bis the reduced row-echelon form of A. The matrix Jis the result
on the right side of the augmented matrix, which is the result of applying the same row operations to the
identity matrix. We should recognize now that Jis just the product of the elementary matrices (Subsection
DM.EM [423]) that perform these row operations. Theorem ITMT [676] used the extended echelon form
to discern properties of the inverse of a triangular matrix. Theorem TD [909] proves the existence of a
triangular decomposition by applying specic row operations, and tracking the relevant elementary row
operations. It is not a great leap to combine these observations into a computational procedure.
To nd the triangular decomposition of A, augmentAwith the identity matrix of the same size and
call this new 2 nnmatrix,M. Perform row operations on Mthat convert the rst ncolumns to an upper
triangular matrix. Do this using only row operations that add a scalar multiple of one row to another row
with higher index (i.e. lower down). In this way, the last ncolumns of Mwill be converted into a lower
triangular matrix with 1's on the diagonal (since Mhas 1's in these locations initially). We could think of
this process as doing about half of the work required to compute the inverse of A. Take the rst ncolumns
of the row-equivalent version of Mand call this matrix U. Take the nal ncolumns of the row-equivalent
version ofMand call this matrix L0. Then by a proof employing elementary matrices, or a proof similar in
spirit to the one used to prove Theorem PEEF [298], we arrive at a result similar to the second assertion
of Theorem PEEF [298]. Namely, U=L0A. Multiplication on the left, by the inverse of L0, will give us a
decomposition of A(which we know to be unique). Ready? Lets try it.
Version 2.30
Subsection TD.TD Triangular Decomposition 913
Example TD4
Triangular decomposition, size 4
In this example, we will illustrate the process for computing a triangular decomposition, as described in
the previous paragraphs. Consider the nonsingular square matrix Aof size 4,
A=2
664 2 6 8 7
4 16 14 15
6 22 23 26
6 26 18 173
775
We formMby augmenting Awith the size 4 identity matrix I4. We will perform the allowed operations,
column by column, only reporting intermediate results as we nish converting each column. It is easy to
determine exactly which row operations we perform, since the nal four columns contain a record of each
such operation. We will not verify our hypotheses about the nonsingularity of the Ak, since if we do not
have these conditions, we will reach a stage where a diagonal entry is zero and we cannot create the row
operations we need to zero out the bottom portion of the associated column. In other words, we can boldly
proceed and the necessity of our hypotheses will become apparent.
M=2
664 2 6 8 7 1 0 0 0
4 16 14 15 0 1 0 0
6 22 23 26 0 0 1 0
6 26 18 17 0 0 0 13
775
!2
664 2 6 8 7 1 0 0 0
0 4 2 1 2 1 0 0
0 4 1 5 3 0 1 0
0 8 6 4 3 0 0 13
775
!2
664 2 6 8 7 1 0 0 0
0 4 2 1 2 1 0 0
0 0 1 4 1 1 1 0
0 0 2 6 1 2 0 13
775
!2
664 2 6 8 7 1 0 0 0
0 4 2 1 2 1 0 0
0 0 1 4 1 1 1 0
0 0 0 2 1 4 2 13
775
So at this point, we have UandL0,
U=2
664 2 6 8 7
0 4 2 1
0 0 1 4
0 0 0 23
775L0=2
6641 0 0 0
2 1 0 0
1 1 1 0
1 4 2 13
775
Then by whatever procedure we like (such as Theorem CINM [248]), we nd
L=
L0 1=2
6641 0 0 0
2 1 0 0
3 1 1 0
3 2 2 13
775
It is instructive to verify that indeed LU=A.
Version 2.30
914 Section TD Triangular Decomposition
Subsection TDSSE
Triangular Decomposition and Solving Systems of Equations
In this section we give an explanation of why you might be interested in a triangular decomposition for
a matrix. Many of the computational problems in linear algebra revolve around solving large systems
of equations, or nearly equivalently, nding inverses of large matrices. Suppose we have a system of
equations with coecient matrix Aand vector of constants b, and suppose further that Ahas the triangular
decomposition A=LU.
Letybe the solution to the linear system LS(L;b), so that by Theorem SLEMM [224], we have
Ly=b. Notice that since Lis nonsingular, this solution is unique, and the form of Lmakes it trivial
to solve the system. The rst component of yis determined easily, and we can continue on through
determining the components of y, without even ever dividing. Now, with yin hand, consider the linear
system,LS(U;y). Let xbe the unique solution to this system, so by Theorem SLEMM [224] we have
Ux=y. Notice that a system of equations with Uas a coecient matrix is also straightforward to solve,
though we will compute the bottom entries of xrst, and we will need to divide. The upshot of all this is
thatxis a solution toLS(A;b), as we now show,
Ax=LUx=L(Ux) =Ly=b
An application of Theorem SLEMM [224] demonstrates that xis a solution toLS(A;b).
Example TDSSE
Triangular decomposition solves a system of equations
Here we illustrate the previous discussion, recycling the decomposition found previously in Example TD4
[911]. Consider the linear system LS(A;b) with
A=2
664 2 6 8 7
4 16 14 15
6 22 23 26
6 26 18 173
775b=2
664 10
2
1
83
775
First we solve the system LS(L;b) (see Example TD4 [911] for L),
y1= 10
2y1+y2= 2
3y1+y2+y3= 1
3y1+ 2y2 2y3+y4= 8
Then
y1= 10
y2= 2 2y1= 2 2( 10) = 18
y3= 1 3y1 y2= 1 3( 10) 18 = 11
y4= 8 3y1 2y2+ 2y3= 8 3( 10) 2(18) + 2(11) = 8
so
y=2
664 10
18
11
83
775
Version 2.30
Subsection TD.CTD Computing Triangular Decompositions 915
Then we solve the system LS(U;y) (see Example TD4 [911] for U),
2x1+ 6x2 8x3+ 7x4= 10
4x2+ 2x3+x4= 18
x3+ 4x4= 11
2x4= 8
Then
x4= 8=2 = 4
x3= (11 4x4)=( 1) = (11 4(4))=( 1) = 5
x2= (18 2x3 x4)=4 = (18 2(5) 4)=4 = 1
x1= ( 10 6x2+ 8x3 7x4)=( 2) = ( 10 6(1) + 8(5) 7(4))=( 2) = 2
And so
x=2
6644
5
1
23
775
is the solution to LS(U;y) and consequently is the unique solution to LS(A;b), as you can easily verify.
Subsection CTD
Computing Triangular Decompositions
It would be a simple matter to adjust the algorithm for converting a matrix to reduced row-echelon form
and obtain an algorithm to compute the triangular decomposition of the matrix, along the lines of Example
TD4 [911] and the discussion preceding this example. However, it is possible to obtain relatively simple
formulas for the entries of the decomposition, and if computed in the proper order, an implementation will
be straightforward. We will state the result as a theorem and then give an example of its use.
Theorem TDEE
Triangular Decomposition, Entry by Entry
Suppose that Ais a squarematrix of size nwith a triangular decomposition A=LU, whereLis lower
triangular with diagonal entries all equal to 1, and Uis upper triangular. Then
[U]ij= [A]ij i 1X
k=1[L]ik[U]kj 1ijn
[L]ij=1
[U]jj
[A]ij j 1X
k=1[L]ik[U]kj!
1j <in
Proof Consider a single scalar product of an entry of Lwith an entry of Uof the form [ L]ik[U]kj. By
Denition LTM [675], if k>i then [L]ik= 0, while Denition UTM [675], says that if k>j then [U]kj= 0.
So we can combine these two facts to assert that if k>min(i; j), [L]ik[U]kj= 0 since at least one term of
the product will be zero. Employing this observation,
[A]ij=nX
k=1[L]ik[U]kj Theorem EMP [227]
Version 2.30
916 Section TD Triangular Decomposition
=min(i;j)X
k=1[L]ik[U]kj
Now, assume that 1 ijn,
[U]ij= [A]ij [A]ij+ [U]ij
= [A]ij min(i;j)X
k=1[L]ik[U]kj+ [U]ij
= [A]ij iX
k=1[L]ik[U]kj+ [U]ij
= [A]ij i 1X
k=1[L]ik[U]kj [L]ii[U]ij+ [U]ij
= [A]ij i 1X
k=1[L]ik[U]kj [U]ij+ [U]ij
= [A]ij i 1X
k=1[L]ik[U]kj
And for 1j <in,
[L]ij=1
[U]jj
[L]ij[U]jj
=1
[U]jj
[A]ij [A]ij+ [L]ij[U]jj
=1
[U]jj0
@[A]ij min(i;j)X
k=1[L]ik[U]kj+ [L]ij[U]jj1
A
=1
[U]jj
[A]ij jX
k=1[L]ik[U]kj+ [L]ij[U]jj!
=1
[U]jj
[A]ij j 1X
k=1[L]ik[U]kj [L]ij[U]jj+ [L]ij[U]jj!
=1
[U]jj
[A]ij j 1X
k=1[L]ik[U]kj!
At rst glance, these formulas may look exceedingly complex. Upon closer examination, it looks even
worse. We have expressions for entries of Uthat depend on other entries of Uand also on entries of L.
But then the formula for entries of Ldepend on entries from Land entries from U. Do these formula have
circular dependencies? Or perhaps equivalently, how do we get started? The key is to be organized about
the computations and employ these two (similar) formulas in a specic order. First compute the rst row
ofL, followed by the rst column of U. Then the second row of L, followed by the second column of U.
And so on. In this way, all of the values required for each new entry will have already been computed
previously.
Of course, the formula for entries of Lrequire division by diagonal entries of U. These entries might
be zero, but in this case Ais nonsingular and does not have a triangular decomposition. So we need not
Version 2.30
Subsection TD.CTD Computing Triangular Decompositions 917
check the hypothesis carefully and can launch into the arithmetic dictated by the formulas, condent that
we will be reminded when a decomposition is not possible. Note that these formula give us all of the values
that we need for the decomposition, since we require that Lhas 1's on the diagonal. If we replace the 1's
on the diagonal of Lby zeros, and add the matrix U, we get an nnmatrix containing all the information
we need to resurrect the triangular decomposition. This is mostly a notational convenience, but it is a
frequent way of presenting the information. We'll employ it in the next example.
Example TDEE6
Triangular decomposition, entry by entry, size 6
We illustrate the application of the formulas in Theorem TDEE [913] for the 6 6 matrixA.
A=2
66666643 3 3 2 1 0
6 4 5 2 4 2
9 9 7 7 0 1
6 10 8 10 1 7
6 4 9 2 10 1
9 3 12 3 21 23
7777775
Using the notational convenience of packaging the two triangular matrices into one matrix, and using the
ordering of the computations mentioned above, we display the results after computing a single row and
column of each of the two triangular matrices.
2
66666643 3 3 2 1 0
2
3
2
2
33
77777752
66666643 3 3 2 1 0
2 2 1 2 2 2
3 0
2 2
2 1
3 33
7777775
2
66666643 3 3 2 1 0
2 2 1 2 2 2
3 0 2 1 3 1
2 2 0
2 1 2
3 3 33
77777752
66666643 3 3 2 1 0
2 2 1 2 2 2
3 0 2 1 3 1
2 2 0 2 1 3
2 1 2 1
3 3 3 33
7777775
2
66666643 3 3 2 1 0
2 2 1 2 2 2
3 0 2 1 3 1
2 2 0 2 1 3
2 1 2 1 1 2
3 3 3 3 03
77777752
66666643 3 3 2 1 0
2 2 1 2 2 2
3 0 2 1 3 1
2 2 0 2 1 3
2 1 2 1 1 2
3 3 3 3 0 23
7777775
Splitting out the pieces of this matrix, we have the decomposition,
L=2
66666641 0 0 0 0 0
2 1 0 0 0 0
3 0 1 0 0 0
2 2 0 1 0 0
2 1 2 1 1 0
3 3 3 3 0 13
7777775U=2
66666643 3 3 2 1 0
0 2 1 2 2 2
0 0 2 1 3 1
0 0 0 2 1 3
0 0 0 0 1 2
0 0 0 0 0 23
7777775
The hypotheses of Theorem TD [909] can be weakened slightly to include matrices where not every
Akis nonsingular. The introduces a rearrangement of the rows and columns of Ato force as many as
Version 2.30
918 Section TD Triangular Decomposition
possible of the smaller submatrices to be nonsingular. Then permutation matrices also enter into the
decomposition. We will not present the details here, but instead suggest consulting a more advanced text
on matrix analysis.
Version 2.30
Section SVD Singular Value Decomposition 919
Section SVD
Singular Value Decomposition
This Section is a Draft, Subject to Changes
Needs Numerical Examples
The singular value decomposition is one of the more useful ways to represent any matrix, even rectan-
gular ones. We can also view the singular values of a (rectangular) matrix as analogues of the eigenvalues
of a square matrix. Our denitions and theorems in this section rely heavily on the properties of the
matrix-adjoint products ( AAandAA), which we rst met in Theorem CPSM [899]. We start by exam-
ining some of the basic properties of these two matrices. Now would be a good time to review the basic
facts about positive semi-denite matrices in Section PSM [899].
Subsection MAP
Matrix-Adjoint Product
Theorem EEMAP
Eigenvalues and Eigenvectors of Matrix-Adjoint Product
Suppose that Ais anmnmatrix and AAhas rankr. Let1; 2; 3; :::; pbe the nonzero distinct
eigenvalues of AAand let1; 2; 3; :::; qbe the nonzero distinct eigenvalues of AA. Then,
1.p=q.
2. The distinct nonzero eigenvalues can be ordered such that i=i, 1ip.
3. Properly ordered, AA(i) =AA(i), 1ip.
4. The rank of AAis equal to the rank of AA.
5. There is an orthonormal basis, fx1;x2;x3; :::; xngofCncomposed of eigenvectors of AAand an
orthonormal basis, fy1;y2;y3; :::; ymgofCmcomposed of eigenvectors of AAwith the following
properties. Order the eigenvectors so that xi,r+ 1inare the eigenvectors of AAfor the zero
eigenvalue. Let i, 1irdenote the nonzero eigenvalues of AA. ThenAxi=piyi, 1ir
andAxi=0,r+1in. Finally, yi,r+1im, are eigenvectors of AAfor the zero eigenvalue.
Proof Suppose that x2Cnis any eigenvector of AAfor a nonzero eigenvalue . We will show that Ax
is an eigenvector of AAfor the same eigenvalue, . First, we ascertain that Axis not the zero vector.
hAx; Axi=hAx;(A)xi Theorem AA [215]
=hAAx;xi Theorem AIP [233]
=hx;xi Denition EEM [453]
=hx;xi Theorem IPSM [194]
Since xis an eigenvector, x6=0, and by Theorem PIP [196], hx;xi6= 0. Aswas assumed to be nonzero,
we see thathAx; Axi6= 0. Again, Theorem PIP [196] tells us that Ax6=0.
Much of the sequel turns on the following simple computation. If you ever wonder what all the fuss is
about adjoints, Hermitian matrices, square roots, and singular values, return to this brief computation, as
Version 2.30
920 Section SVD Singular Value Decomposition
it holds the key. There is much more to do in this proof, but after this it is mostly bookkeeping. Here we
go. We check that Axfunctions as an eigenvector of AAfor the eigenvalue ,
(AA)Ax=A(AA)x Theorem MMA [231]
=Ax Denition EEM [453]
=(Ax) Theorem MMSMM [230]
That's it. If xis an eigenvector of AA(for a nonzero eigenvalue), then Axis an eigenvector for AAfor
the same eigenvalue. Let's see what this buys us.
AAandAAare Hermitian matrices (Denition HM [234]), and hence are normal (Denition NRML
[680]). This provides the existence of orthonormal bases of eigenvectors for each matrix by Theorem
OBNM [683]. Also, since each matrix is diagonalizable (Denition DZM [496]) by Theorem OD [681] we
can interchange algebraic and geometric multiplicities by Theorem DMFE [499].
Our rst step is to establish that an eigenvalue has the same geometric multiplicity for both AA
andAA. Supposefx1;x2;x3; :::; xsgis an orthonormal basis of eigenvectors of AAfor the eigenspace
EAA(). Then for 1i<js, note
hAxi; Axji=hAxi;(A)xji Theorem AA [215]
=hAAxi;xji Theorem AIP [233]
=hxi;xji Denition EEM [453]
=hxi;xji Theorem IPSM [194]
=(0) Denition ONS [201]
= 0 Property ZCN [759]
Then the set E=fAx1; Ax2; Ax3; :::; A xsgis an orthogonal set of nonzero eigenvectors of AAfor the
eigenvalue. By Theorem OSLI [198], the set Eis linearly independent and so the geometric multiplicity
ofas an eigenvalue of AAissor greater. We have
AA() =
AA()
AA() =AA()
This inequality applies to any matrix, so long as the eigenvalue is nonzero. We now apply it to the matrix
A,
AA() =(A)A()A(A)() =AA()
So for a nonzero eigenvalue, its algebraic multiplicities as an eigenvalue of AAandAAare equal. This
is enough to establish that p=qand the eigenvalues can be ordered such that i=ifor 1ip.
For any matrix B, the null space is identical to the eigenspace of the zero eigenvalue, N(B) =EB(0),
and thus the nullity of the matrix is equal to the geometric multiplicity of the zero eigenvalue. With this,
we can examine the ranks of AAandAA.
r(AA) =n n(AA) Theorem RPNC [398]
=
AA(0) +pX
i=1AA(i)!
n(AA) Theorem NEM [485]
=
AA(0) +pX
i=1AA(i)!
AA(0) Denition GME [463]
=
AA(0) +pX
i=1AA(i)!
AA(0) Theorem DMFE [499]
Version 2.30
Subsection SVD.MAP Matrix-Adjoint Product 921
=pX
i=1AA(i)
=pX
i=1AA(i)
=
AA(0) +pX
i=1AA(i)!
AA(0)
=
AA(0) +pX
i=1AA(i)!
AA(0) Theorem DMFE [499]
=
AA(0) +pX
i=1AA(i)!
n(AA) Denition GME [463]
=m n(AA) Theorem NEM [485]
=r(AA) Theorem RPNC [398]
WhenAis rectangular, the square matrices AAandAAhave dierent sizes. With equal algebraic and
geometric multiplicities for their common nonzero eigenvalues, the dierence in their sizes is manifest in
dierent algebraic multiplicities for the zero eigenvalue and dierent nullities. Specically,
n(AA) =n r n (AA) =m r
Suppose that x1;x2;x3; :::; xnis an orthonormal basis of Cncomposed of eigenvectors of AAand ordered
so that xi,r+ 1inare eigenvectors of AAfor the zero eigenvalue. Denote the associated nonzero
eigenvalues of AAfor these eigenvectors by i, 1ir. Then dene
yi=1piAxi 1ir
Letyr+1;yr+2;yr+2; :::; ymbe an orthonormal basis for the eigenspace EAA(0), whose existence is
guaranteed by Theorem GSP [199]. As scalar multiples of demonstrated eigenvectors of AA,yi, 1ir
are also eigenvectors of AA, and yi,r+ 1inhave been chosen as eigenvectors of AA. These
eigenvectors also have norm 1, as we now show. For 1 ir,
kyik=
1piAxi
=s1piAxi;1piAxi
Theorem IPN [195]
=s
1pi1pihAxi; Axii Theorem IPSM [194]
=s
1pi1pihAxi; Axii Theorem HMRE [487]
=1pip
hAxi; Axii
=1piq
hAxi;(A)xii Theorem AA [215]
=1pip
hAAxi;xii Theorem AIP [233]
=1pip
hixi;xii Denition EEM [453]
Version 2.30
922 Section SVD Singular Value Decomposition
=1pip
ihxi;xii Theorem IPSM [194]
=1pip
i(1) Denition ONS [201]
= 1
Forr+ 1in, theyihave been chosen to have norm 1.
Finally we check orthogonality. Consider two eigenvectors yiandyjwith 1i<jm. If these two
vectors have dierent eigenvalues, then Theorem HMOE [488] establishes that the two eigenvectors are
orthogonal. If the two eigenvectors have a zero eigenvalue, then they are orthogonal by the choice of the
orthonormal basis of EAA(0). If the two eigenvectors have identical, nonzero, eigenvalues, then
hyi;yji=*
1piAxi;1p
jAxj+
=1pi1p
jhAxi; Axji Theorem IPSM [194]
=1p
ijhAxi; Axji Theorem HMRE [487]
=1p
ijhAxi;(A)xji Theorem AA [215]
=1p
ijhAAxi;xji Theorem AIP [233]
=1p
ijhixi;xji Denition EEM [453]
=ip
ijhxi;xji Theorem IPSM [194]
=ip
ij(0) Denition ONS [201]
= 0
Sofy1;y2;y3; :::; ymgis an orthonormal set of eigenvectors for AA. The critical relationship between
these two orthonormal bases is present by design. For 1 ir,
Axi=p
i1piAxi=p
iyi
Forr+ 1inwe have
hAxi; Axii=hAxi;(A)xii Theorem AA [215]
=hAAxi;xii Theorem AIP [233]
=h0;xii Denition EEM [453]
= 0 Denition IP [192]
So by Theorem PIP [196], Axi=0.
Subsection SVD
Singular Value Decomposition
The square roots of the eigenvalues of AA(or almost equivalently, AA!) are known as the singular values
ofA. Here is the denition.
Version 2.30
Subsection SVD.SVD Singular Value Decomposition 923
Denition SV
Singular Values
SupposeAis anmnmatrix. If the eigenvalues of AAare1; 2; 3; :::; n, then the singular values
ofAarep1;p2;p3; :::;pn. 4
Theorem EEMAP [917] is a total setup for the singular value decomposition. This remarkable theorem
says that anymatrix can be broken into a product of three matrices. Two are square, and unitary. In
light of Theorem UMPIP [264], we can view these matrices as transforming vectors or coordinates in a
rotational fashion. The middle matrix of this decomposition is rectangular, but is as close to being diagonal
as a rectangular matrix can be. Viewed as a transformation, this matrix eects, re
ections, contractions
or expansions along axes | it stretches vectors. So any matrix, viewed as a transformation is the product
of a rotation, a stretch and a rotation.
The singular value theorem can also be viewed as an application of our most general statement about
matrix representations of linear transformations relative to dierent bases. Theorem MRCB [654] concerns
linear transformations T:U!VwhereUandVare possibly dierent vector spaces. When UandV
have dierent dimensions, the resulting matrix representation will be rectangular. In Section CB [647] we
quickly specialized to the case where U=Vand the matrix representations are square with one of our
most central results, Theorem SCB [656]. Theorem SVD [921] is an application of the full generality of
Theorem MRCB [654] where the relevant bases are now orthonormal sets.
Theorem SVD
Singular Value Decomposition
SupposeAis anmnmatrix of rank rwith nonzero singular values s1; s2; s3; :::; sr. ThenA=UDV
whereUis a unitary matrix of size m,Vis a unitary matrix of size nandDis anmnmatrix given by
[D]ij=(
siif 1i=jr
0 otherwise
Proof Letx1;x2;x3; :::; xnandy1;y2;y3; :::; ymbe the orthonormal bases described by the conclu-
sion of Theorem EEMAP [917]. Dene Uto be themmmatrix whose columns are yi, 1im, and
deneVto be thennmatrix whose columns are xi, 1in. With orthonormal sets of columns, by
Theorem CUMOS [263] both UandVare unitary matrices.
Then for 1im, 1jn,
[AV]ij= [Axj]iDenition MM [226]
=hp
jyji
iTheorem EEMAP [917]
= [sjyj]iDenition SV [921]
= [yj]isj Denition CVSM [99]
= [U]ij[D]jj
=mX
k=1[U]ik[D]kj
= [UD]ij Theorem EMP [227]
So by Theorem ME [485], AV=UDand thus
A=AIn=AVV=UDV
Version 2.30
924 Section SVD Singular Value Decomposition
Version 2.30
Section SR Square Roots 925
Section SR
Square Roots
This Section is a Draft, Subject to Changes
Needs Numerical Examples
With all our results about Hermitian matrices, their eigenvalues and their diagonalizations, it will be a
nearly trivial matter to now construct a \square root" of a positive semi-denite matrix. We will describe
the square root of a matrix Aas a matrix Ssuch thatA=S2. In general, a matrix Amight have many
such square roots. But with a few results in hand we will be able to impose an extra condition on Sthat
will make a unique Ssuch thatA=S2. At that point we can dene thesquare root of Aformally.
Subsection SRM
Square Root of a Matrix
Theorem PSMSR
Positive Semi-Denite Matrices and Square Roots
SupposeAis a square matrix. There is a positive semi-denite matrix Ssuch thatA=S2if and only if
Ais positive semi-denite.
Proof Letndenote the size of A.
(() Suppose that Ais positive semi-denite. Since Ais Hermitian (Denition PSM [899]) we know A
is normal (Denition NRML [680]) and so by Theorem OD [681] there is a unitary matrix Uand a diagonal
matrixD, whose diagonal entries are the eigenvalues of A, such that D=UAU. The eigenvalues of A
are all non-negative (Theorem EPSM [900]), which allows us to dene a diagonal matrix Ewhose diagonal
entries are the positive square roots of the eigenvalues of A, in the same order as they appear in D. More
precisely, dene Eto be the diagonal matrix with non-negative diagonal entries such that E2=D. Set
S=UEU, and compute
S2=UEUUEU
=UEInEUDenition UM [262]
=UEEUTheorem MMIM [229]
=UDU
=UUAUUTheorem OD [681]
=InAIn Denition UM [262]
=A Theorem MMIM [229]
We need to rst verify that Sis Hermitian.
S= (UEU)
= (UEU)
= (U)EUTheorem MMAD [233]
=UEUTheorem AA [215]
=U
EtUDenition A [214]
=UEtUTheorem HMRE [487]
Version 2.30
926 Section SR Square Roots
=UEUDiagonal matrix
=S
And nally, we want to check the use of Sin an inner product. Notice that Eis Hermitian since it is a
diagonal matrix with real entries. Furthermore, as a diagonal matrix, the eigenvalues of Eare precisely
the diagonal entries, and since these were chosen to be positive, an application of Theorem EPSM [900]
tells us that Eis positive semi-denite. Now, for any x2Cn,
hSx;xi=hUEUx;xi
=hEUx; Uxi Theorem AIP [233]
=hE(Ux); Uxi
0 Denition PSM [899]
So, according to Denition PSM [899], Sis positive semi-denite.
()) Assume that A=S2, withSpositive semi-denite. Then Sis Hermitian, and we check that Ais
Hermitian.
A= (SS)
=SSTheorem MMAD [233]
=SS Denition HM [234]
=A
Now for the use of Ain an inner product. For any x2Cn,
hAx;xi=
S2x;x
=hSx; Sxi Theorem AIP [233]
=hSx; Sxi Denition HM [234]
0 Theorem PIP [196]
So by Denition PSM [899], Ais positive semi-denite.
There is a very close relationship between the eigenvalues and eigenspaces of a positive semi-denite
matrix and its positive semi-denite square root. The next theorem is interesting in its own right, but is
also an important technical step in some other important results, such as the upcoming uniqueness of the
square root (Theorem USR [926]).
Theorem EESR
Eigenvalues and Eigenspaces of a Square Root
Suppose that Ais a positive semi-denite matrix and Sis a positive semi-denite matrix such that
A=S2. If1; 2; 3; :::; pare the distinct eigenvalues of A, then the distinct eigenvalues of Sarep1;p2;p3; :::;p
p, andES pi
=EA(i) for 1ip.
Proof Letxbe an eigenvector of Sfor an eigenvalue . Then, in the style of Theorem EPM [481],
Ax=S2x=S(Sx) =S(x) =Sx=2x
so2is an eigenvalue of Aand must equal some i. Furthermore, because Sis positive semi-denite,
Theorem EPSM [900] tells us that 0. The impact for us here is that we cannot have two dierent
eigenvalues of Swhose squares equal the same eigenvalue of A, so we can pair each eigenvalue of Swith a
dierent eigenvalue of A, equal to its square. (A good exercise is to track through the rest of this proof in
the situation where Sis not assumed to be positive semi-denite and we do not have this condition on the
eigenvalues. Where does the proof then break down?) Let i, 1iqdenote theqdistinct eigenvalues of
Version 2.30
Subsection SR.SRM Square Root of a Matrix 927
S. The discussion above implies that we can order the eigenvalues of AandSso thati=2
ifor 1iq.
Notice that at this point we know that qp, though we will be showing that q=p.
Additionally, the equation above tells us that every eigenvector of Sforiis again an eigenvector of A
for2
i. So for 1iq, the relevant eigenspaces are related by
ESp
i
=ES(i)EA
2
i
=EA(i)
So the eigenspaces of Sare subsets of the eigenspaces of A, for the related eigenvalues. However, we will
be showing that these sets are indeed equal to each other.
BothAandSare positive semi-denite, hence Hermitian and therefore normal. Theorem OD [681]
then tells us that each is diagonalizable (Denition DZM [496]). Then Theorem DMFE [499] says that the
algebraic multiplicity and geometric multiplicity of each eigenvalue are equal. Then, if we let ndenote the
size ofA,
n=qX
i=1Sp
i
Theorem NEM [485]
=qX
i=1
Sp
i
Theorem DMFE [499]
=qX
i=1dim
ESp
i
Denition GME [463]
qX
i=1dim (EA(i)) Theorem PSSD [410]
pX
i=1dim (EA(i)) Denition D [391]
=pX
i=1
A(i) Denition GME [463]
=pX
i=1A(i) Theorem DMFE [499]
=n Theorem NEM [485]
With equal values at the two ends of this chain of equalities and inequalities, we know that the two
inequalities are forced to actually be equalities. In particular, the second inequality implies that p=qand
the rst, in conjunction with Theorem EDYES [410], implies that ES pi
=EA(i) for 1ip.
Notice that we dened the singular values of a matrix Aas the square roots of the eigenvalues of AA
(Denition SV [921]). With Theorem EESR [924] in hand we recognize the singular values of Aas simply
the eigenvalues of AA1=2. Indeed, many authors take this as the denition of singular values, since it is
equivalent to our denition. We have chosen not to wait for a discussion of square roots before making a
denition of singular values, allowing us to present the singular value decomposition (Theorem SVD [921])
all the sooner.
In the rst half of the proof of Theorem PSMSR [923] we could have chosen the matrix E(which was
the essential component of the desired matrix S) in a variety of ways. Any collection of diagonal entries
ofEcould be replaced by their negatives and we would maintain the property that E2=D. However,
if we decide to enforce the entries of Eas non-negative quantities then Eis positive semi-denite, and
thenSfollows along as a positive semi-denite matrix. We now show that of all the possible square roots
of a positive semi-denite matrix, only one is itself again positive semi-denite. In other words, the Sof
Theorem PSMSR [923] is unique.
Version 2.30
928 Section SR Square Roots
Theorem USR
Unique Square Root
SupposeAis a positive semi-denite matrix. Then there is a unique positive semi-denite matrix Ssuch
thatA=S2.
Proof Theorem PSMSR [923] gives us the existence of at least one positive semi-denite matrix Ssuch
thatA=S2. As usual, we will assume that S1andS2are positive semi-denite matrices such that
A=S2
1=S2
2(Technique U [771]).
AsAis diagonalizable, there is a basis of Cncomposed entirely of eigenvectors of A(Theorem DC [497]),
sayB=fx1;x2;x3; :::; xng. Let1; 2; 3; :::; ndenote the associated eigenvalues. Theorem EESR
[924] allows to conclude that EA(i) =ES1 pi
=ES2 pi
. SoS1xi=pixi=S2xifor 1in.
Choose any x2Cn. The spanning property of Ballows us to conclude the existence of a set of scalars,
a1; a2; a3; :::; an, yielding xas a linear combination of the vectors in B. So,
S1x=S1nX
i=1aixi=nX
i=1aiS1xi=nX
i=1aip
ixi=nX
i=1aiS2xi=S2nX
i=1aixi=S2x
SinceS1andS2have the same action on every vector, Theorem EMMVP [225] yields the conclusion that
S1=S2.
With a criteria that distinguishes one square root from all the rest (positive semi-deniteness) we can
now dene thesquare root of a positive semi-denite matrix.
Denition SRM
Square Root of a Matrix
SupposeAis a positive semi-denite matrix and Sis the positive semi-denite matrix such that S2=
SS=A. ThenSis the square root ofAand we write S=A1=2.
(This denition contains Notation SRM.) 4
Version 2.30
Section POD Polar Decomposition 929
Section POD
Polar Decomposition
This Section is a Draft, Subject to Changes
Needs Numerical Examples
The polar decomposition of a matrix writes any matrix as the product of a unitary matrix (Denition
UM [262])and a positive semi-denite matrix (Denition PSM [899]). It takes its name from a special way
to write complex numbers. If you've had a basic course in complex analysis, the next paragraph will help
explain the name. If the next paragraph makes no sense to you, there's no harm in skipping it.
Any complex number z2Ccan be written as z=reiwhereris a positive number (computed as a
square root of a function of the real amd imaginary parts of z) andis an angle of rotation that converts 1
to the complex number ei= cos() +isin(). The polar form of a square matrix is a product of a positive
semi-denite matrix that is a square root of a function of the matrix together with a unitary matrix, which
can be viewed as achieving a rotation (Theorem UMPIP [264]).
OK, enough preliminaries. We have all the tools in place to jump straight to our main theorem.
Theorem PDM
Polar Decomposition of a Matrix
Suppose that Ais a square matrix. Then there is a unitary matrix Usuch thatA= (AA)1=2U.
Proof This theorem only claims the existence of a unitary matrix Uthat does a certain job. We will
manufacture Uand check that it meets the requirements.
SupposeAhas sizenand rankr. We begin by applying Theorem EEMAP [917] to A. LetB=
fx1;x2;x3; :::; xngbe the orthonormal basis of Cncomposed of eigenvectors for AA, and letC=
fy1;y2;y3; :::; yngbe the orthonormal basis of Cncomposed of eigenvectors for AA. We have Axi=pixi, 1ir, andAxi=0,r+ 1in, wherei, 1irare the distinct nonzero eigenvalues of
AA.
DeneT:Cn!Cnto be the unique linear transformation such that T(xi) =yi, 1in, as
guaranteed by Theorem LTDB [525]. Let Ebe the basis of standard unit vectors for Cn(Denition SUV
[197]), and dene Uto be the matrix representation (Denition MR [615]) of Twith respect to E, more
carefullyU=MT
E;E. This is the matrix we are after. Notice that
Uxi=MT
E;EE(xi) Denition VR [603]
=E(T(xi)) Theorem FTMR [617]
=E(yi) Theorem FTMR [617]
=yi Denition VR [603]
SinceBandCare orthonormal bases, and Cis the result of multiplying the vectors of BbyU, we conclude
thatUis unitary by Theorem UMCOB [380]. So once again, Theorem EEMAP [917] is a big part of the
setup for a decomposition.
Letx2Cnbe any vector. Since Bis a basis of Cn, there are scalars a1; a2; a3; :::; anexpressing xas
a linear combination of the vectors in B. then
(AA)1=2Ux= (AA)1=2UnX
i=1aixi Denition B [371]
=nX
i=1(AA)1=2Uaixi Theorem MMDAA [230]
Version 2.30
930 Section POD Polar Decomposition
=nX
i=1ai(AA)1=2Uxi Theorem MMSMM [230]
=nX
i=1ai(AA)1=2yi
=rX
i=1ai(AA)1=2yi+nX
i=r+1ai(AA)1=2yi Property AAC [100]
=rX
i=1aip
iyi+nX
i=r+1ai(0)yi Theorem EESR [924]
=rX
i=1aip
iyi+nX
i=r+1ai0 Theorem ZSSM [324]
=rX
i=1aiAxi+nX
i=r+1aiAxi Theorem EEMAP [917]
=nX
i=1aiAxi Property AAC [100]
=nX
i=1Aaixi Theorem MMSMM [230]
=AnX
i=1aixi Theorem MMDAA [230]
=Ax
So by Theorem EMMVP [225] we have the matrix equality ( AA)1=2U=A.
Version 2.30
Part A
Applications
931
Section CF
Curve Fitting
This Section is Incomplete
Given two points in the plane, there is a unique line through them. Given three points in the plane,
and not in a line, there is a unique parabola through them. Given four points in the plane, there is a
unique polynomial, of degree 3 or less, passing through them. And so on. We can prove this result, and
give a procedure for nding the polynomial with the help of Vandermonde matrices (Section VM [895]).
Theorem IP
Interpolating Polynomial
Supposef(xi; yi)j1in+ 1gis a set of n+ 1 points in the plane where the x-coordinates are all
dierent. Then there is a unique polynomial of degree nor less,p(x), such that p(xi) =yi, 1in+ 1.
Proof Writep(x) =a0+a1x+a2x2++anxn. To meet the conclusion of the theorem, we desire,
yi=p(xi) =a0+a1xi+a2x2
i++anxn
i 1in+ 1
This is a system of n+ 1 linear equations in the n+ 1 variables a0; a1; a2; :::; an. The vector of constants
in this system is the vector containing the y-coordinates of the points. More importantly, the coecient
matrix is a Vandermonde matrix (Denition VM [895]) built from the x-coordinates x1; x2; x3; :::; xn+1.
Since we have required that these scalars all be dierent, Theorem NVM [898] tells us that the coecient
matrix is nonsingular and Theorem NMUS [86] says the solution for the coecients of the polynomial
exists, and is unique. As a practical matter, Theorem SNCM [261] provides an expression for the solution.
Example PTFP
Polynomial through ve points
Suppose we have the following 5 points in the plane and we wish to pass a degree 4 polynomial through
them.
i 1 2 3 4 5
xi-3 -1 2 3 6
yi276 16 31 144 2319
The required system of equations has a coecient matrix that is the Vandermonde matrix where row iis
successive powers of xi
A=2
666641 3 9 27 81
1 1 1 1 1
1 2 4 8 16
1 3 9 27 81
1 6 36 216 12963
77775
Theorem NMUS [86] provides a solution as
2
66664a0
a1
a2
a3
a43
77775=A 12
66664276
16
31
144
23193
77775=2
66664 1
159
149
10 1
21
42
0 3
73
4 1
31
845
108 1
56 1
417
72 11
756
1
541
21 1
121
18 1
7561
540 1
1681
60 1
721
7563
777752
66664276
16
31
144
23193
77775=2
666643
4
5
2
23
77775
934 Section CF Curve Fitting
So the polynomial is p(x) = 3 4x+ 5x2 2x3+ 2x4.
The unique polynomial passing through a set of points is known as the interpolating polynomial
and it has many uses. Unfortunately, when confronted with data from an experiment the situation may
not be so simple or clear cut. Read on.
Subsection DF
Data Fitting
Suppose that we have nreal variables, x1; x2; x3; :::; xn, that we can measure in an experiment. We
believe that these variables combine, in a linear fashion, to equal another real variable, y. In other words,
we have reason to believe from our understanding of the experiment, that
y=a1x1+a2x2+a3x3++anxn
where the scalars a1; a2; a3; :::; anare not known to us, but are instead desirable. We would call this our
model of the situation. Then we run the experiment mtimes, collecting sets of values for the variables
of the experiment. For run number kwe might denote these values as yk,xk1,xk2,xk3, . . . ,xkn. If we
substitute these values into the model equation, we get mlinear equations in the unknown coecients
a1; a2; a3; :::; an. Ifm=n, then we have a square coecient matrix of the system which might happen
to be nonsingular and there would be a unique solution.
However, more likely m>n (the more data we collect, the greater our condence in the results) and
the resulting system is inconsistent. It may be that our model is only an approximate understanding of
the relationship between the xiandy, or our measurements are not completely accurate. Still we would
like to understand the situation we are studying, and would like some best answer for a1; a2; a3; :::; an.
Letydenote the vector with [ y]i=yi, 1im, letadenote the vector with [ a]j=aj, 1jn,
and letXdenote the mnmatrix with [ X]ij=xij, 1im, 1jn. Then the model equation,
evaluated with each run of the experiment, translates to Xa=y. With the presumption that this system
has no solution, we can try to minimize the dierence between the two side of the equation y Xa. As a
vector, it is hard to imagine what the minimum might be, so we instead minimize the square of its norm
S= (y Xa)t(y Xa)
To keep the logical
ow accurate, we will dene the minimizing value and then give the proof that it
behaves as desired.
Denition LSS
Least Squares Solution
Given the equation Xa=y, whereXis anmnmatrix of rank n, theleast squares solution forais
XtX 1Xty. 4
Theorem LSMR
Least Squares Minimizes Residuals
Suppose that Xis anmnmatrix of rank n. The least squares solution of Xa=y,a0=
XtX 1Xty,
minimizes the expression
S= (y Xa)t(y Xa)
Proof We begin by nding the critical points of S. In preparation, let Xjdenote column jofX, for
1jnand compute partial derivatives with respect to aj, 1jn. A matrix product of the form
Version 2.30
Subsection CF.DF Data Fitting 935
xtyis a sum of products, so a derivative is a sum of applications of the product rule,
@
@ajS=@
@aj
(y Xa)t(y Xa)
=mX
i=1@
@aj([y Xa]i) [y Xa]i+ [y Xa]i@
@aj([y Xa]i)
= 2mX
i=1@
@aj([y Xa]i) [y Xa]i
= 2mX
i=1@
@aj
[y]i nX
k=1[X]ik[a]k!
[y Xa]i
= 2mX
i=1 [X]ij[y Xa]i
= 2 (Xj)t(y Xa)
The rst partial derivatives will allow us to nd critical points, while second partial derivatives will be
needed to conrm that a critical point will yield a minimum. Return to the next-to-last expression for the
rst partial derivative of S,
@
@a`ajS=@
@a`2mX
i=1 [X]ij[y Xa]i
= 2mX
i=1@
@a`[X]ij[y Xa]i
= 2mX
i=1[X]ij@
@a`
[y]i nX
k=1[X]ik[a]k!
= 2mX
i=1[X]ij( [X]i`)
= 2mX
i=1[X]ij[X]i`
= 2mX
i=1
Xt
ji[X]i`
= 2
XtX
j`
For 1jn, set@
@ajS= 0. This results in the nscalar equations
(Xj)tXa= (Xj)ty 1jn
Thesenvector equations can be summarized in the single vector equation,
XtXa=Xty
XtXis annnmatrix and since we have assumed that Xhas rankn,XtXwill also have rank n. Since
XtXis invertible, we have a critical point at
a0=
XtX 1Xty
Version 2.30
936 Section CF Curve Fitting
Is this lone critical point really a minimum? The matrix of second partial derivatives is constant, and a
positive multiple of XtX. Theorem CPSM [899] tells us that this matrix is positive semi-denite. In an
advanced course on multivariable calculus, it is shown that a minimum occurs exactly where the matrix of
second partial derivatives is positive semi-denite. You may have seen this in the two-variable case, where
a check on the positive semi-deniteness is disguised with a determinant of the 2 2 matrix of second
partial derivatives.
Version 2.30
Subsection CF.EXC Exercises 937
Subsection EXC
Exercises
T20 Theorem IP [931] constructs a unique polynomial through a set of n+ 1 points in the plane,
f(xi; yi)j1in+ 1g, where the x-coordinates are all dierent. Prove that the expression below is
the same polynomial and include an explanation of the necessity of the hypothesis that the x-coordinates
are all dierent.
p(x) =n+1X
i=1yin+1Y
j=1
j6=ix xj
xi xj
This is known as the Lagrange form of the interpolating polynomial.
Contributed by Robert Beezer
Version 2.30
938 Section CF Curve Fitting
Version 2.30
Section SAS Sharing A Secret 939
Section SAS
Sharing A Secret
This Section is a Draft, Subject to Changes
In this section we will see how to use solutions to systems of equations to share a secret among a group
of people. We will be able to break a secret up into, say 10 pieces, so as to distribute the secret among 10
people. But rather than requiring all 10 people to collaborate on restoring the secret, we can design the
split so that any smaller group, of say just 4 of these people, can collaborate and restore the secret. The
numbers 10 and 4 here are arbitrary, we can choose them to be anything.
Suppose we have a secret, S. This could be the combination to a lock, a password on an account, or
a recipe for chocolate chip cookies. If the secret is text, we will assume that the characters have been
translated into integers (say with the ASCII code), and these numbers have been rolled up into one grand
positive integer (perhaps by concatenating binary strings for the ASCII code numbers, and interpreting
the longer string as one big base 2 integer). So we will assume Sis some positive integer.
Suppose you wish to give parts of your secret to npeople, and you wish to require that any group of
m(or more) of these people should be able to combine their parts and recover the secret. Perhaps you are
President and CEO of a small company and only you know the password that authorizes large transfers of
money among the company's bank accounts. If you were to die or become incapacitated, it would perhaps
hamper the company's ability to function if they couldn't quickly rearrange their assets, especially since
they are also without a CEO. So you might wish to give this secret to six of your trusted Vice-Presidents.
But you don't trust them that much and you certainly don't want any one of these people to be able to
access the company's accounts all by themselves without anybody else in the company knowing about it.
Simultaneously, you know that in an emergency, it might not be possible to get all six Vice-Presidents
together and maybe even one or two of them have met the same unfortunate fate you did. So you would
like any group of three Vice-Presidents to be able to combine their parts and recover S. So you would
choosen= 6 andm= 3.
We will describe the split, with no motivation. The explanation of how the secret recovery is handled
will explain our choices here. Choose a large prime number, p, bigger than any possible secret. For a single
number in a combination lock, pcould be small. For a one-page recipe, pwould need to be huge. All of
our subsequent arithmetic will be modulo p, so consult Subsection F.FF [874] for a brief description of
how we do linear algebra when our eld is Zp. Build a polynomial, r(x), of degree m 1 as follows. Set
the constant term to S, and choose the other m 1 coecients at random from Zp. The quality of your
random generator will ultimately aect the quality of how hidden your secret remains.
Compute the pairs ( i; r(i)), 1in. To person i, of thenpersons you will give a part of your secret,
present the pair ( i; r(i)), and instruct them to keep this secret, for all 1 in. They could perhaps
encrypt their pairs with AES (Advanced Encryption Standard) using a password known only to them
individually. Or you could do this for each of them in advance and tell them the chose password orally, in
private. At any rate, each person gets a pair of integers, an input to the polynomial, and the output of
evaluating the polynomial, and they keep this information secret. They do not know the polynomial itself,
and certainly not the constant term S, so the secret is still safe.
Now suppose that mof these people get together, in the event you are unable to act, or perhaps without
your permission. Suppose they pool all of their pairs, or even just turn them over to one member of the
group. What do they now know collectively? Suppose that
r(x) =a0+a1x+a2x2++am 1xm 1
where, of course, a0=Sis the secret. A single pair, ( i; r(i)), results in a linear equation whose unknowns
are themcoecients of r(x). Withmpairs revealed, we now have mequations in mvariables. Furthermore,
Version 2.30
940 Section SAS Sharing A Secret
the coecient matrix of this system is a Vandermonde matrix (Denition VM [895]). With our inputs to
the polynomial all dierent (we used 1 ;2;3; :::; n ), the Vandermonde matrix is nonsingular (Theorem
NVM [898]). Thus by Theorem NMUS [86] there is a unique solution for the coecients of r(x). We only
desire the constant term | the other coecients (the randomly chosen ones) are of no interest, they were
used to mask the secret as it was split into parts.
A few practical considerations. If certain individuals in your group are more important, or more
trustworthy, you can give them more than one part. You could split a secret into 30 parts, giving 5 Vice-
Presidents each 4 parts and give 10 department heads each 1 part. Then you might require 12 parts to
be present. This way three Vice-Presidents could recover the secret, or 4 department heads could stand-in
for a Vice-President. Furthermore, the 10 department heads could not recover the secret without having
at least one Vice-President present.
The inputs do not have to be consecutive integers, starting at 1. Any set of dierent integers will
suce. Why make it any easier for an attacker? Mix it up and choose the inputs randomly as well, just
keep them dierent.
Why do all this arithmetic over Zp? If we worked with polynomials having real number coecients,
properties of polynomials as continuous functions might give an attacker the ability to compute the secret
with a reasonable amount of computing time. For example, the magnitude of the output is going to
dominated by the term of r(x) having degree m 1. Suppose an attacker had a few of the pairs, but not
a full set of mof them. Or even worse, suppose some group of fewer than mof your trusted acquaintances
were to conspire against you. It might be possible to guess a limited range of values for the coecient
of the largest term. With a limited range of values here, the next term might fall to a similar analysis.
And so on. However, modular arithmetic is in some ways very unpredictable looking and as high powers
\wrap-around" this sort of analysis will be frustrated. And we know it is no harder to do linear algebra in
Zpthan in C.
OK, here's a non-trivial example.
Example SS6W
Sharing a secret 6 ways
Let's return to the CEO and his six Vice-Presidents. Suppose the password for the company's accounts
is a sequence of 5 two-digit numbers, which we will concatenate into a 10-digit number, in this case S=
0603725962. For a prime pwe choose the 11-digit prime number p= 22801761379. From the requirement
thatm= 3 Vice-Presidents are needed to recover the secret, we need a second-degree polynomial and so
need two more coecients, which we will construct at random between 1 and p. The resulting polynomial
is
r(x) = 603725962 + 22561982919 x+ 8844088338 x2
We will now build six pairs of inputs and outputs, where we will choose the inputs at random (not allowing
duplicates) and we do all our arithmetic modulo p,
VP x r (x)
Finance 20220406046 7205699654
Human Resources 8862377358 17357568951
Marketing 13747127957 18503158079
Legal 15835120319 14060705999
Research 6530855859 5628836054
Manufacturing 9222703664 2608052019
The two numbers of each row of the table are then given to the indicated Vice-President. Done. The secret
has been split six ways, and any three VP's can jointly recover the secret.
Let's test the recovery process, especially since it contains the relevant linear algebra. Suppose we write
the unknown polynomial as r(x) =a0+a1x+a2x2and the VP's for Finance, Marketing and Legal all get
Version 2.30
Section SAS Sharing A Secret 941
together to recover the secret. The equations we arrive at are,
Finance 7205699654 = r(20220406046)
=a0+a1(20220406046) + a2(20220406046)2
=a0+ 20220406046 a1+ 7793596215 a2
Marketing 18503158079 = r(13747127957)
=a0+a1(13747127957) + a2(13747127957)2
=a0+ 13747127957 a1+ 18840301370 a2
Legal 14060705999 = r(15835120319)
=a0+a1(15835120319) + a2(15835120319)2
=a0+ 15835120319 a1+ 8874412999 a2
So they have a linear system, LS(A;b) with
A=2
41 20220406046 7793596215
1 13747127957 18840301370
1 15835120319 88744129993
5 b=2
47205699654
18503158079
140607059993
5
With a Vandermonde matrix as the coecient matrix, they know there is a solution, and it is unique. By
Theorem SNCM [261] (or through row-reducing the augmented matrix) they arrive at the solution,
A 1b=2
45716900879 9234437646 7850422855
20952200747 16452595922 8198726089
17286943796 18018241597 102983373653
52
47205699654
18503158079
140607059993
5=2
4603725962
22561982919
88440883383
5
So the CEO's password is the secret S=a0= 603725962 = 0603725962 (as expected).
Version 2.30
942 Section SAS Sharing A Secret
Version 2.30
Index
A (appendix), 777
A (archetype), 781
A (denition), 214
A (notation), 214
A (part), 931
AA (Property), 317
AA (subsection, section WILA), 4
AA (theorem), 215
AAC (Property), 100
AACN (Property), 758
AAF (Property), 873
AALC (example), 111
AAM (Property), 209
ABLC (example), 110
ABS (example), 131
AC (Property), 317
ACC (Property), 100
ACCN (Property), 758
ACF (Property), 873
ACM (Property), 209
ACN (example), 757
additive associativity
column vectors
Property AAC, 100
complex numbers
Property AACN, 758
matrices
Property AAM, 209
vectors
Property AA, 317
additive closure
column vectors
Property ACC, 100
complex numbers
Property ACCN, 758
eld
Property ACF, 873
matrices
Property ACM, 209
vectors
Property AC, 317
additive commutativitycomplex numbers
Property CACN, 758
additive inverse
complex numbers
Property AICN, 759
from scalar multiplication
theorem AISM, 325
additive inverses
column vectors
Property AIC, 100
matrices
Property AIM, 209
unique
theorem AIU, 324
vectors
Property AI, 318
adjoint
denition A, 214
inner product
theorem AIP, 233
notation, 214
of a matrix sum
theorem AMA, 214
of an adjoint
theorem AA, 215
of matrix scalar multiplication
theorem AMSM, 214
AHSAC (example), 71
AI (Property), 318
AIC (Property), 100
AICN (Property), 759
AIF (Property), 874
AIM (Property), 209
AIP (theorem), 233
AISM (theorem), 325
AIU (theorem), 324
AIVLT (example), 579
ALT (example), 516
ALTMM (example), 619
AM (denition), 30
AM (example), 27
AM (notation), 30
943
944 INDEX
AM (subsection, section MO), 214
AMA (theorem), 214
AMAA (example), 30
AME (denition), 463
AME (notation), 463
AMSM (theorem), 214
ANILT (example), 580
ANM (example), 680
AOS (example), 197
Archetype A
column space, 276
linearly dependent columns, 158
singular matrix, 83
solving homogeneous system, 72
system as linear combination, 111
archetype A
augmented matrix
example AMAA, 30
Archetype B
column space, 276
inverse
example CMIAB, 249
linearly independent columns, 158
nonsingular matrix, 84
not invertible
example MWIAA, 244
solutions via inverse
example SABMI, 243
solving homogeneous system, 72
system as linear combination, 110
vector equality, 98
archetype B
solutions
example SAB, 39
Archetype C
homogeneous system, 71
Archetype D
column space, original columns, 275
solving homogeneous system, 72
vector form of solutions, 114
Archetype I
column space from row operations, 282
null space, 74
row space, 278
vector form of solutions, 121
Archetype I:casting out vectors, 177
Archetype L
null space span, linearly independent, 161
vector form of solutions, 122
ASC (example), 609augmented matrix
notation, 30
AVR (example), 359
B (archetype), 786
B (denition), 371
B (section), 371
B (subsection, section B), 371
basis
columns nonsingular matrix
example CABAK, 376
common size
theorem BIS, 394
crazy vector apace
example BC, 374
denition B, 371
matrices
example BM, 372
example BSM22, 373
polynomials
example BP, 372
example BPR, 408
example BSP4, 372
example SVP4, 409
subspace of matrices
example BDM22, 409
BC (example), 374
BCS (theorem), 274
BDE (example), 482
BDM22 (example), 409
best cities
money magazine
example MBC, 224
BIS (theorem), 394
BM (example), 372
BNM (subsection, section B), 376
BNS (theorem), 160
BP (example), 372
BPR (example), 408
BRLT (example), 568
BRS (theorem), 280
BS (theorem), 180
BSCV (subsection, section B), 374
BSM22 (example), 373
BSP4 (example), 372
C (archetype), 791
C (denition), 762
C (notation), 762
C (part), 3
C (Property), 317
Version 2.30
INDEX 945
C (technique, section PT), 768
CABAK (example), 376
CACN (Property), 758
CAEHW (example), 458
CAF (Property), 873
canonical form
nilpotent linear transformation
example CFNLT, 698
theorem CFNLT, 694
CAV (subsection, section O), 191
Cayley-Hamilton
theorem CHT, 740
CB (section), 647
CB (theorem), 649
CBCV (example), 652
CBM (denition), 648
CBM (subsection, section CB), 648
CBP (example), 649
CC (Property), 100
CCCV (denition), 191
CCCV (notation), 191
CCM (denition), 212
CCM (example), 212
CCM (notation), 212
CCM (theorem), 213
CCN (denition), 759
CCN (notation), 759
CCN (subsection, section CNO), 759
CCRA (theorem), 759
CCRM (theorem), 760
CCT (theorem), 760
CD (subsection, section DM), 429
CD (technique, section PT), 770
CEE (subsection, section EE), 460
CELT (example), 665
CELT (subsection, section CB), 660
CEMS6 (example), 466
CF (section), 931
CFDVS (theorem), 608
CFNLT (example), 698
CFNLT (subsection, section NLT), 694
CFNLT (theorem), 694
CFV (example), 60
change of basis
between polynomials
example CBP, 649
change-of-basis
between column vectors
example CBCV, 652
matrix representationtheorem MRCB, 654
similarity
theorem SCB, 656
theorem CB, 649
change-of-basis matrix
denition CBM, 648
inverse
theorem ICBM, 649
characteristic polynomial
denition CP, 460
degree
theorem DCP, 484
size 3 matrix
example CPMS3, 460
CHT (subsection, section JCF), 740
CHT (theorem), 740
CILT (subsection, section ILT), 551
CILTI (theorem), 551
CIM (subsection, section MISLE), 245
CINM (theorem), 248
CIVLT (example), 583
CIVLT (theorem), 585
CLI (theorem), 609
CLTLT (theorem), 533
CM (denition), 28
CM (Property), 209
CM32 (example), 611
CMCN (Property), 758
CMF (Property), 873
CMI (example), 247
CMIAB (example), 249
CMVEI (theorem), 61
CN (appendix), 745
CNA (denition), 758
CNA (notation), 758
CNA (subsection, section CNO), 757
CNE (denition), 758
CNE (notation), 758
CNM (denition), 758
CNM (notation), 758
CNMB (theorem), 376
CNO (section), 757
CNS1 (example), 74
CNS2 (example), 75
CNSV (example), 195
COB (theorem), 378
coecient matrix
denition CM, 28
nonsingular
theorem SNCM, 261
Version 2.30
946 INDEX
column space
as null space
theorem FS, 299
Archetype A
example CSAA, 276
Archetype B
example CSAB, 276
as null space
example CSANS, 294
as null space, Archetype G
example FSAG, 305
as row space
theorem CSRST, 282
basis
theorem BCS, 274
consistent system
theorem CSCS, 272
consistent systems
example CSMCS, 271
isomorphic to range, 628
matrix, 271
nonsingular matrix
theorem CSNM, 277
notation, 271
original columns, Archetype D
example CSOCD, 275
row operations, Archetype I
example CSROI, 282
subspace
theorem CSMS, 343
testing membership
example MCSM, 272
two computations
example CSTW, 274
column vector addition
notation, 99
column vector scalar multiplication
notation, 99
commutativity
column vectors
Property CC, 100
matrices
Property CM, 209
vectors
Property C, 317
complexm-space
example VSCV, 319
complex arithmetic
example ACN, 757
complex numberconjugate
example CSCN, 759
modulus
example MSCN, 760
complex number
conjugate
denition CCN, 759
modulus
denition MCN, 760
complex numbers
addition
denition CNA, 758
notation, 758
arithmetic properties
theorem PCNA, 758
equality
denition CNE, 758
notation, 758
multiplication
denition CNM, 758
notation, 758
complex vector space
dimension
theorem DCM, 395
composition
injective linear transformations
theorem CILTI, 551
surjective linear transformations
theorem CSLTS, 570
conjugate
addition
theorem CCRA, 759
column vector
denition CCCV, 191
matrix
denition CCM, 212
notation, 212
multiplication
theorem CCRM, 760
notation, 759
of conjugate of a matrix
theorem CCM, 213
scalar multiplication
theorem CRSM, 191
twice
theorem CCT, 760
vector addition
theorem CRVA, 191
conjugate of a vector
notation, 191
Version 2.30
INDEX 947
conjugation
matrix addition
theorem CRMA, 213
matrix scalar multiplication
theorem CRMSM, 213
matrix transpose
theorem MCT, 214
consistent linear system, 58
consistent linear systems
theorem CSRN, 59
consistent system
denition CS, 55
constructive proofs
technique C, 768
contradiction
technique CD, 770
contrapositive
technique CP, 769
converse
technique CV, 769
coordinates
orthonormal basis
theorem COB, 378
coordinatization
linear combination of matrices
example CM32, 611
linear independence
theorem CLI, 609
orthonormal basis
example CROB3, 379
example CROB4, 378
spanning sets
theorem CSS, 610
coordinatization principle, 611
coordinatizing
polynomials
example CP2, 610
COV (example), 177
COV (subsection, section LDS), 177
CP (denition), 460
CP (subsection, section VR), 609
CP (technique, section PT), 769
CP2 (example), 610
CPMS3 (example), 460
CPSM (theorem), 899
crazy vector space
example CVSR, 609
properties
example PCVS, 326
CRMA (theorem), 213CRMSM (theorem), 213
CRN (theorem), 397
CROB3 (example), 379
CROB4 (example), 378
CRS (section), 271
CRS (subsection, section FS), 294
CRSM (theorem), 191
CRVA (theorem), 191
CS (denition), 55
CS (example), 762
CS (subsection, section TSS), 55
CSAA (example), 276
CSAB (example), 276
CSANS (example), 294
CSCN (example), 759
CSCS (theorem), 272
CSIP (example), 192
CSLT (subsection, section SLT), 570
CSLTS (theorem), 570
CSM (denition), 271
CSM (notation), 271
CSMCS (example), 271
CSMS (theorem), 343
CSNM (subsection, section CRS), 276
CSNM (theorem), 277
CSOCD (example), 275
CSRN (theorem), 59
CSROI (example), 282
CSRST (diagram), 307
CSRST (theorem), 282
CSS (theorem), 610
CSSE (subsection, section CRS), 271
CSSOC (subsection, section CRS), 274
CSTW (example), 274
CTD (subsection, section TD), 913
CTLT (example), 533
CUMOS (theorem), 263
curve tting
polynomial through 5 points
example PTFP, 931
CV (denition), 27
CV (notation), 28
CV (technique, section PT), 769
CVA (denition), 98
CVA (notation), 99
CVC (notation), 28
CVE (denition), 98
CVE (notation), 98
CVS (example), 322
CVS (subsection, section VR), 608
Version 2.30
948 INDEX
CVSM (denition), 99
CVSM (example), 100
CVSM (notation), 99
CVSR (example), 609
D (acronyms, section PDM), 451
D (archetype), 795
D (chapter), 423
D (denition), 391
D (notation), 391
D (section), 391
D (subsection, section D), 391
D (subsection, section SD), 496
D (technique, section PT), 765
D33M (example), 428
DAB (example), 496
DC (example), 396
DC (technique, section PT), 772
DC (theorem), 497
DCM (theorem), 395
DCN (Property), 759
DCP (theorem), 484
DD (subsection, section DM), 427
DEC (theorem), 431
decomposition
technique DC, 772
DED (theorem), 501
denition
A, 214
AM, 30
AME, 463
B, 371
C, 762
CBM, 648
CCCV, 191
CCM, 212
CCN, 759
CM, 28
CNA, 758
CNE, 758
CNM, 758
CP, 460
CS, 55
CSM, 271
CV, 27
CVA, 98
CVE, 98
CVSM, 99
D, 391
DIM, 496
DM, 428DS, 413
DZM, 496
EEF, 297
EELT, 647
EEM, 453
ELEM, 423
EM, 461
EO, 14
ES, 761
ESYS, 14
F, 873
GES, 707
GEV, 707
GME, 463
HI, 890
HID, 890
HM, 234
HP, 889
HS, 71
IDLT, 579
IDV, 57
IE, 717
ILT, 541
IM, 84
IMP, 874
IP, 192
IS, 703
IVLT, 579
IVS, 586
JB, 687
JCF, 727
KLT, 545
LC, 338
LCCV, 109
LI, 351
LICV, 153
LNS, 293
LSS, 932
LT, 515
LTA, 530
LTC, 532
LTM, 675
LTR, 711
LTSM, 531
M, 27
MA, 207
MCN, 760
ME, 207
MI, 244
MM, 226
Version 2.30
INDEX 949
MR, 615
MRLS, 29
MSM, 208
MVP, 223
NLT, 685
NM, 83
NOLT, 588
NOM, 397
NRML, 680
NSM, 73
NV, 195
ONS, 201
OSV, 197
OV, 196
PI, 528
PSM, 899
REM, 31
RLD, 351
RLDCV, 153
RLT, 563
RO, 31
ROLT, 588
ROM, 397
RR, 42
RREF, 33
RSM, 278
S, 333
SC, 763
SE, 762
SET, 761
SI, 763
SIM, 493
SLE, 11
SLT, 559
SM, 428
SOLV, 29
SQM, 83
SRM, 926
SS, 339
SSCV, 131
SSET, 761
SSLE, 12
SSSLE, 12
SU, 763
SUV, 197
SV, 921
SYM, 211
T, 883
technique D, 765
TM, 210TS, 337
TSHSE, 71
TSVS, 356
UM, 262
UTM, 675
VM, 895
VOC, 28
VR, 603
VS, 317
VSCV, 97
VSM, 207
ZCV, 28
ZM, 210
DEHD (example), 501
DEM (theorem), 444
DEMMM (theorem), 445
DEMS5 (example), 468
DER (theorem), 429
DERC (theorem), 441
determinant
computed two ways
example TCSD, 432
denition DM, 428
equal rows or columns
theorem DERC, 441
expansion, columns
theorem DEC, 431
expansion, rows
theorem DER, 429
identity matrix
theorem DIM, 443
matrix multiplication
theorem DRMM, 447
nonsingular matrix, 445
notation, 428
row or column multiple
theorem DRCM, 440
row or column swap
theorem DRCS, 439
size 2 matrix
theorem DMST, 429
size 3 matrix
example D33M, 428
transpose
theorem DT, 430
via row operations
example DRO, 442
zero
theorem SMZD, 445
zero row or column
Version 2.30
950 INDEX
theorem DZRC, 439
zero versus nonzero
example ZNDAB, 446
determinant, upper triangular matrix
example DUTM, 432
determinants
elementary matrices
theorem DEMMM, 445
DF (Property), 874
DF (subsection, section CF), 932
DFS (subsection, section PD), 412
DFS (theorem), 412
DGES (theorem), 727
diagonal matrix
denition DIM, 496
diagonalizable
denition DZM, 496
distinct eigenvalues
example DEHD, 501
theorem DED, 501
full eigenspaces
theorem DMFE, 499
not
example NDMS4, 501
diagonalizable matrix
high power
example HPDM, 502
diagonalization
Archetype B
example DAB, 496
criteria
theorem DC, 497
example DMS3, 498
diagram
CSRST, 307
DLTA, 516
DLTM, 516
DTSLS, 61
FTMR, 618
FTMRA, 619
GLT, 519
ILT, 543
MRCLT, 625
NILT, 542
DIM (denition), 496
DIM (theorem), 443
dimension
crazy vector space
example DC, 396
denition D, 391notation, 391
polynomial subspace
example DSP4, 396
proper subspaces
theorem PSSD, 410
subspace
example DSM22, 395
direct sum
decomposing zero vector
theorem DSZV, 414
denition DS, 413
dimension
theorem DSD, 416
example SDS, 413
from a basis
theorem DSFB, 413
from one subspace
theorem DSFOS, 414
notation, 413
zero intersection
theorem DSZI, 415
direct sums
linear independence
theorem DSLI, 416
repeated
theorem RDS, 417
distributivity
complex numbers
Property DCN, 759
eld
Property DF, 874
distributivity, matrix addition
matrices
Property DMAM, 209
distributivity, scalar addition
column vectors
Property DSAC, 101
matrices
Property DSAM, 209
vectors
Property DSA, 318
distributivity, vector addition
column vectors
Property DVAC, 101
vectors
Property DVA, 318
DLDS (theorem), 175
DLTA (diagram), 516
DLTM (diagram), 516
DM (denition), 428
Version 2.30
INDEX 951
DM (notation), 428
DM (section), 423
DM (theorem), 395
DMAM (Property), 209
DMFE (theorem), 499
DMHP (subsection, section HP), 891
DMHP (theorem), 891
DMMP (theorem), 892
DMS3 (example), 498
DMST (theorem), 429
DNLT (theorem), 691
DNMMM (subsection, section PDM), 445
DP (theorem), 395
DRCM (theorem), 440
DRCMA (theorem), 441
DRCS (theorem), 439
DRMM (theorem), 447
DRO (example), 442
DRO (subsection, section PDM), 439
DROEM (subsection, section PDM), 443
DS (denition), 413
DS (notation), 413
DS (subsection, section PD), 413
DSA (Property), 318
DSAC (Property), 101
DSAM (Property), 209
DSD (theorem), 416
DSFB (theorem), 413
DSFOS (theorem), 414
DSLI (theorem), 416
DSM22 (example), 395
DSP4 (example), 396
DSZI (theorem), 415
DSZV (theorem), 414
DT (theorem), 430
DTSLS (diagram), 61
DUTM (example), 432
DVA (Property), 318
DVAC (Property), 101
DVM (theorem), 895
DVS (subsection, section D), 395
DZM (denition), 496
DZRC (theorem), 439
E (acronyms, section SD), 513
E (archetype), 799
E (chapter), 453
E (technique, section PT), 768
E.SAGE (computation, section SAGE), 755
ECEE (subsection, section EE), 463
EDELI (theorem), 479EDYES (theorem), 410
EE (section), 453
EEE (subsection, section EE), 456
EEF (denition), 297
EEF (subsection, section FS), 297
EELT (denition), 647
EELT (subsection, section CB), 647
EEM (denition), 453
EEM (subsection, section EE), 453
EEMAP (theorem), 917
EENS (example), 496
EER (theorem), 659
EESR (theorem), 924
EHM (subsection, section PEE), 487
eigenspace
as null space
theorem EMNS, 462
denition EM, 461
invariant subspace
theorem EIS, 705
subspace
theorem EMS, 461
eigenspaces
sage, 755
eigenvalue
algebraic multiplicity
denition AME, 463
notation, 463
complex
example CEMS6, 466
denition EEM, 453
existence
example CAEHW, 458
theorem EMHE, 457
geometric multiplicity
denition GME, 463
notation, 463
index, 717
linear transformation
denition EELT, 647
multiplicities
example EMMS4, 463
power
theorem EOMP, 481
root of characteristic polynomial
theorem EMRCP, 461
scalar multiple
theorem ESMM, 481
symmetric matrix
example ESMS4, 464
Version 2.30
952 INDEX
zero
theorem SMZE, 480
eigenvalues
building desired
example BDE, 482
complex, of a linear transformation
example CELT, 665
conjugate pairs
theorem ERMCP, 483
distinct
example DEMS5, 468
example SEE, 453
Hermitian matrices
theorem HMRE, 487
inverse
theorem EIM, 482
maximum number
theorem MNEM, 487
multiplicities
example HMEM5, 465
theorem ME, 485
number
theorem NEM, 485
of a polynomial
theorem EPM, 481
size 3 matrix
example EMS3, 461
example ESMS3, 462
transpose
theorem ETM, 483
eigenvalues, eigenvectors
vector, matrix representations
theorem EER, 659
eigenvector, 453
linear transformation, 647
eigenvectors, 454
conjugate pairs, 483
Hermitian matrices
theorem HMOE, 488
linear transformation
example ELTBM, 647
example ELTBP, 648
linearly independent
theorem EDELI, 479
of a linear transformation
example ELTT, 660
EILT (subsection, section ILT), 541
EIM (theorem), 482
EIS (example), 705
EIS (theorem), 705ELEM (denition), 423
ELEM (notation), 424
elementary matrices
denition ELEM, 423
determinants
theorem DEM, 444
nonsingular
theorem EMN, 427
notation, 424
row operations
example EMRO, 424
theorem EMDRO, 425
ELIS (theorem), 407
ELTBM (example), 647
ELTBP (example), 648
ELTT (example), 660
EM (denition), 461
EM (subsection, section DM), 423
EMDRO (theorem), 425
EMHE (theorem), 457
EMMS4 (example), 463
EMMVP (theorem), 225
EMN (theorem), 427
EMNS (theorem), 462
EMP (theorem), 227
empty set, 761
notation, 761
EMRCP (theorem), 461
EMRO (example), 424
EMS (theorem), 461
EMS3 (example), 461
ENLT (theorem), 690
EO (denition), 14
EOMP (theorem), 481
EOPSS (theorem), 14
EPM (theorem), 481
EPSM (theorem), 900
equal matrices
via equal matrix-vector products
theorem EMMVP, 225
equation operations
denition EO, 14
theorem EOPSS, 14
equivalence statements
technique E, 768
equivalences
technique ME, 771
equivalent systems
denition ESYS, 14
ERMCP (theorem), 483
Version 2.30
INDEX 953
ES (denition), 761
ES (notation), 761
ESEO (subsection, section SSLE), 13
ESLT (subsection, section SLT), 559
ESMM (theorem), 481
ESMS3 (example), 462
ESMS4 (example), 464
ESYS (denition), 14
ETM (theorem), 483
EVS (subsection, section VS), 319
example
AALC, 111
ABLC, 110
ABS, 131
ACN, 757
AHSAC, 71
AIVLT, 579
ALT, 516
ALTMM, 619
AM, 27
AMAA, 30
ANILT, 580
ANM, 680
AOS, 197
ASC, 609
AVR, 359
BC, 374
BDE, 482
BDM22, 409
BM, 372
BP, 372
BPR, 408
BRLT, 568
BSM22, 373
BSP4, 372
CABAK, 376
CAEHW, 458
CBCV, 652
CBP, 649
CCM, 212
CELT, 665
CEMS6, 466
CFNLT, 698
CFV, 60
CIVLT, 583
CM32, 611
CMI, 247
CMIAB, 249
CNS1, 74
CNS2, 75CNSV, 195
COV, 177
CP2, 610
CPMS3, 460
CROB3, 379
CROB4, 378
CS, 762
CSAA, 276
CSAB, 276
CSANS, 294
CSCN, 759
CSIP, 192
CSMCS, 271
CSOCD, 275
CSROI, 282
CSTW, 274
CTLT, 533
CVS, 322
CVSM, 100
CVSR, 609
D33M, 428
DAB, 496
DC, 396
DEHD, 501
DEMS5, 468
DMS3, 498
DRO, 442
DSM22, 395
DSP4, 396
DUTM, 432
EENS, 496
EIS, 705
ELTBM, 647
ELTBP, 648
ELTT, 660
EMMS4, 463
EMRO, 424
EMS3, 461
ESMS3, 462
ESMS4, 464
FDV, 57
FF8, 876
FRAN, 564
FS1, 303
FS2, 304
FSAG, 305
FSCF, 503
GE4, 708
GE6, 709
GENR6, 717
Version 2.30
954 INDEX
GSTV, 200
HISAA, 72
HISAD, 72
HMEM5, 465
HP, 889
HPDM, 502
HUSAB, 72
IAP, 549
IAR, 542
IAS, 281
IAV, 544
ILTVR, 632
IM, 84
IM11, 875
IS, 17
ISJB, 706
ISMR4, 714
ISMR6, 715
ISSI, 56
IVSAV, 586
JB4, 687
JCF10, 729
KPNLT, 693
KVMR, 626
LCM, 338
LDCAA, 158
LDHS, 156
LDP4, 394
LDRN, 157
LDS, 153
LIC, 355
LICAB, 158
LIHS, 155
LIM32, 353
LINSB, 159
LIP4, 351
LIS, 154
LLDS, 157
LNS, 293
LTDB1, 526
LTDB2, 527
LTDB3, 527
LTM, 520
LTPM, 518
LTPP, 518
LTRGE, 711
MA, 208
MBC, 224
MCSM, 272
MFLT, 522MI, 245
MIVS, 609
MMNC, 227
MNSLE, 224
MOLT, 524
MPMR, 622
MRBE, 657
MRCM, 654
MSCN, 760
MSM, 208
MTV, 223
MWIAA, 244
NDMS4, 501
NIAO, 549
NIAQ, 541
NIAQR, 548
NIDAU, 550
NJB5, 688
NKAO, 545
NLT, 517
NM, 84
NM62, 686
NM64, 685
NM83, 689
NRREF, 33
NSAO, 566
NSAQ, 559
NSAQR, 566
NSC2A, 336
NSC2S, 337
NSC2Z, 336
NSDAT, 569
NSDS, 138
NSE, 12
NSEAI, 74
NSLE, 29
NSLIL, 161
NSNM, 86
NSR, 85
NSS, 85
OLTTR, 615
ONFV, 202
ONTV, 201
OSGMD, 61
OSMC, 263
PCVS, 326
PM, 455
PSHS, 125
PTFP, 931
PTM, 226
Version 2.30
INDEX 955
PTMEE, 228
RAO, 563
RES, 182
RNM, 397
RNSM, 398
ROD2, 905
ROD4, 906
RREF, 33
RREFN, 55
RRTI, 411
RS, 375
RSAI, 278
RSB, 374
RSC4, 182
RSC5, 176
RSNS, 338
RSREM, 280
RVMR, 629
S, 83
SAA, 40
SAB, 39
SABMI, 243
SAE, 41
SAN, 567
SAR, 560
SAV, 561
SC, 764
SC3, 333
SCAA, 133
SCAB, 135
SCAD, 139
SDS, 413
SEE, 453
SEEF, 297
SETM, 761
SI, 763
SM2Z7, 876
SM32, 341
SMLT, 532
SMS3, 494
SMS5, 493
SP4, 335
SPIAS, 528
SRR, 85
SS, 428
SS6W, 938
SSC, 358
SSET, 761
SSM22, 357
SSNS, 137SSP, 340
SSP4, 356
STLT, 531
STNE, 11
SU, 763
SUVOS, 197
SVP4, 409
SYM, 211
TCSD, 432
TD4, 911
TDEE6, 915
TDSSE, 912
TIS, 703
TIVS, 609
TKAP, 546
TLC, 109
TM, 210
TMP, 4
TOV, 196
TREM, 31
TTS, 13
UM3, 262
UPM, 262
US, 16
USR, 32
VA, 99
VESE, 98
VFS, 115
VFSAD, 114
VFSAI, 121
VFSAL, 122
VM4, 895
VRC4, 604
VRP2, 606
VSCV, 319
VSF, 321
VSIM5, 875
VSIS, 320
VSM, 319
VSP, 319
VSPUD, 396
VSS, 321
ZNDAB, 446
EXC (subsection, section B), 383
EXC (subsection, section CB), 669
EXC (subsection, section CF), 935
EXC (subsection, section CRS), 284
EXC (subsection, section D), 401
EXC (subsection, section DM), 434
EXC (subsection, section EE), 471
Version 2.30
956 INDEX
EXC (subsection, section F), 879
EXC (subsection, section FS), 308
EXC (subsection, section HP), 894
EXC (subsection, section HSE), 76
EXC (subsection, section ILT), 552
EXC (subsection, section IS), 720
EXC (subsection, section IVLT), 593
EXC (subsection, section LC), 127
EXC (subsection, section LDS), 185
EXC (subsection, section LI), 163
EXC (subsection, section LISS), 362
EXC (subsection, section LT), 535
EXC (subsection, section MINM), 266
EXC (subsection, section MISLE), 253
EXC (subsection, section MM), 236
EXC (subsection, section MO), 216
EXC (subsection, section MR), 635
EXC (subsection, section NM), 88
EXC (subsection, section O), 203
EXC (subsection, section PD), 418
EXC (subsection, section PDM), 449
EXC (subsection, section PEE), 489
EXC (subsection, section PSM), 902
EXC (subsection, section RREF), 44
EXC (subsection, section S), 345
EXC (subsection, section SD), 507
EXC (subsection, section SLT), 571
EXC (subsection, section SS), 142
EXC (subsection, section SSLE), 20
EXC (subsection, section T), 887
EXC (subsection, section TSS), 63
EXC (subsection, section VO), 103
EXC (subsection, section VR), 613
EXC (subsection, section VS), 328
EXC (subsection, section WILA), 8
extended echelon form
submatrices
example SEEF, 297
extended reduced row-echelon form
properties
theorem PEEF, 298
F (archetype), 803
F (denition), 873
F (section), 873
F (subsection, section F), 873
FDV (example), 57
FF (subsection, section F), 874
FF8 (example), 876
Fibonacci sequence
example FSCF, 503eld
denition F, 873
FIMP (theorem), 875
nite eld
size 8
example FF8, 876
four subsets
example FS1, 303
example FS2, 304
four subspaces
dimension
theorem DFS, 412
FRAN (example), 564
free variables
example CFV, 60
free variables, number
theorem FVCS, 60
free, independent variables
example FDV, 57
FS (section), 293
FS (subsection, section FS), 299
FS (subsection, section SD), 503
FS (theorem), 299
FS1 (example), 303
FS2 (example), 304
FSAG (example), 305
FSCF (example), 503
FTMR (diagram), 618
FTMR (theorem), 617
FTMRA (diagram), 619
FV (subsection, section TSS), 60
FVCS (theorem), 60
G (archetype), 808
G (theorem), 407
GE4 (example), 708
GE6 (example), 709
GEE (subsection, section IS), 706
GEK (theorem), 708
generalized eigenspace
as kernel
theorem GEK, 708
denition GES, 707
dimension
theorem DGES, 727
dimension 4 domain
example GE4, 708
dimension 6 domain
example GE6, 709
invariant subspace
theorem GESIS, 707
Version 2.30
INDEX 957
nilpotent restriction
theorem RGEN, 716
nilpotent restrictions, dimension 6 domain
example GENR6, 717
notation, 707
generalized eigenspace decomposition
theorem GESD, 721
generalized eigenvector
denition GEV, 707
GENR6 (example), 717
GES (denition), 707
GES (notation), 707
GESD (subsection, section JCF), 721
GESD (theorem), 721
GESIS (theorem), 707
GEV (denition), 707
GFDL (appendix), 865
GLT (diagram), 519
GME (denition), 463
GME (notation), 463
goldilocks
theorem G, 407
Gram-Schmidt
column vectors
theorem GSP, 199
three vectors
example GSTV, 200
gram-schmidt
mathematica, 748
GS (technique, section PT), 767
GSP (subsection, section O), 199
GSP (theorem), 199
GSP.MMA (computation, section MMA), 748
GSTV (example), 200
GT (subsection, section PD), 407
H (archetype), 812
Hadamard Identity
notation, 890
Hadamard identity
denition HID, 890
Hadamard Inverse
notation, 890
Hadamard inverse
denition HI, 890
Hadamard Product
Diagonalizable Matrices
theorem DMHP, 891
notation, 889
Hadamard product
commutativitytheorem HPC, 889
denition HP, 889
diagonal matrices
theorem DMMP, 892
distributivity
theorem HPDAA, 891
example HP, 889
identity
theorem HPHID, 890
inverse
theorem HPHI, 890
scalar matrix multiplication
theorem HPSMM, 891
hermitian
denition HM, 234
Hermitian matrix
inner product
theorem HMIP, 234
HI (denition), 890
HI (notation), 890
HID (denition), 890
HID (notation), 890
HISAA (example), 72
HISAD (example), 72
HM (denition), 234
HM (subsection, section MM), 233
HMEM5 (example), 465
HMIP (theorem), 234
HMOE (theorem), 488
HMRE (theorem), 487
HMVEI (theorem), 73
homogeneous system
Archetype C
example AHSAC, 71
consistent
theorem HSC, 71
denition HS, 71
innitely many solutions
theorem HMVEI, 73
homogeneous systems
linear independence, 155
HP (denition), 889
HP (example), 889
HP (notation), 889
HP (section), 889
HPC (theorem), 889
HPDAA (theorem), 891
HPDM (example), 502
HPHI (theorem), 890
HPHID (theorem), 890
Version 2.30
958 INDEX
HPSMM (theorem), 891
HS (denition), 71
HSC (theorem), 71
HSE (section), 71
HUSAB (example), 72
I (archetype), 816
I (technique, section PT), 772
IAP (example), 549
IAR (example), 542
IAS (example), 281
IAV (example), 544
ICBM (theorem), 649
ICLT (theorem), 585
identities
technique PI, 771
identity matrix
determinant, 444
example IM, 84
notation, 84
IDLT (denition), 579
IDV (denition), 57
IE (denition), 717
IE (notation), 717
IFDVS (theorem), 609
IILT (theorem), 582
ILT (denition), 541
ILT (diagram), 543
ILT (section), 541
ILTB (theorem), 550
ILTD (subsection, section ILT), 550
ILTD (theorem), 550
ILTIS (theorem), 582
ILTLI (subsection, section ILT), 549
ILTLI (theorem), 549
ILTLT (theorem), 582
ILTVR (example), 632
IM (denition), 84
IM (example), 84
IM (notation), 84
IM (subsection, section MISLE), 244
IM11 (example), 875
IMILT (theorem), 633
IMP (denition), 874
IMR (theorem), 630
inconsistent linear systems
theorem ISRN, 59
independent, dependent variables
denition IDV, 57
indesxstring
example SM2Z7, 876example SSET, 761
index
eigenvalue
denition IE, 717
notation, 717
indexstring
theorem DRCMA, 441
theorem OBUTR, 679
theorem UMCOB, 380
induction
technique I, 772
innite solution set
example ISSI, 56
innite solutions, 3 4
example IS, 17
injective
example IAP, 549
example IAR, 542
not
example NIAO, 549
example NIAQ, 541
example NIAQR, 548
not, by dimension
example NIDAU, 550
polynomials to matrices
example IAV, 544
injective linear transformation
bases
theorem ILTB, 550
injective linear transformations
dimension
theorem ILTD, 550
inner product
anti-commutative
theorem IPAC, 194
example CSIP, 192
norm
theorem IPN, 195
notation, 192
positive
theorem PIP, 196
scalar multiplication
theorem IPSM, 194
vector addition
theorem IPVA, 193
integers
modp
denition IMP, 874
modp, eld
theorem FIMP, 875
Version 2.30
INDEX 959
mod 11
example IM11, 875
interpolating polynomial
theorem IP, 931
invariant subspace
denition IS, 703
eigenspace, 705
eigenspaces
example EIS, 705
example TIS, 703
Jordan block
example ISJB, 706
kernels of powers
theorem KPIS, 705
inverse
composition of linear transformations
theorem ICLT, 585
example CMI, 247
example MI, 245
notation, 244
of a matrix, 244
invertible linear transformation
dened by invertible matrix
theorem IMILT, 633
invertible linear transformations
composition
theorem CIVLT, 585
computing
example CIVLT, 583
IP (denition), 192
IP (notation), 192
IP (subsection, section O), 192
IP (theorem), 931
IPAC (theorem), 194
IPN (theorem), 195
IPSM (theorem), 194
IPVA (theorem), 193
IS (denition), 703
IS (example), 17
IS (section), 703
IS (subsection, section IS), 703
ISJB (example), 706
ISMR4 (example), 714
ISMR6 (example), 715
isomorphic
multiple vector spaces
example MIVS, 609
vector spaces
example IVSAV, 586
isomorphic vector spacesdimension
theorem IVSED, 587
example TIVS, 609
ISRN (theorem), 59
ISSI (example), 56
ITMT (theorem), 676
IV (subsection, section IVLT), 582
IVLT (denition), 579
IVLT (section), 579
IVLT (subsection, section IVLT), 579
IVLT (subsection, section MR), 630
IVS (denition), 586
IVSAV (example), 586
IVSED (theorem), 587
J (archetype), 820
JB (denition), 687
JB (notation), 687
JB4 (example), 687
JCF (denition), 727
JCF (section), 721
JCF (subsection, section JCF), 727
JCF10 (example), 729
JCFLT (theorem), 728
Jordan block
denition JB, 687
nilpotent
theorem NJB, 689
notation, 687
size 4
example JB4, 687
Jordan canonical form
denition JCF, 727
size 10
example JCF10, 729
K (archetype), 825
kernel
injective linear transformation
theorem KILT, 548
isomorphic to null space
theorem KNSI, 625
linear transformation
example NKAO, 545
notation, 545
of a linear transformation
denition KLT, 545
pre-image, 547
subspace
theorem KLTS, 546
trivial
Version 2.30
960 INDEX
example TKAP, 546
via matrix representation
example KVMR, 626
KILT (theorem), 548
KLT (denition), 545
KLT (notation), 545
KLT (subsection, section ILT), 545
KLTS (theorem), 546
KNSI (theorem), 625
KPI (theorem), 547
KPIS (theorem), 705
KPLT (theorem), 691
KPNLT (example), 693
KPNLT (theorem), 692
KVMR (example), 626
L (archetype), 829
L (technique, section PT), 766
LA (subsection, section WILA), 3
LC (denition), 338
LC (section), 109
LC (subsection, section LC), 109
LC (technique, section PT), 774
LCCV (denition), 109
LCM (example), 338
LDCAA (example), 158
LDHS (example), 156
LDP4 (example), 394
LDRN (example), 157
LDS (example), 153
LDS (section), 175
LDSS (subsection, section LDS), 175
least squares
minimizes residuals
theorem LSMR, 932
least squares solution
denition LSS, 932
left null space
as row space, 299
denition LNS, 293
example LNS, 293
notation, 293
subspace
theorem LNSMS, 344
lemma
technique LC, 774
LI (denition), 351
LI (section), 153
LI (subsection, section LISS), 351
LIC (example), 355
LICAB (example), 158LICV (denition), 153
LIHS (example), 155
LIM32 (example), 353
linear combination
system of equations
example ABLC, 110
denition LC, 338
denition LCCV, 109
example TLC, 109
linear transformation, 525
matrices
example LCM, 338
system of equations
example AALC, 111
linear combinations
solutions to linear systems
theorem SLSLC, 112
linear dependence
more vectors than size
theorem MVSLD, 158
linear independence
denition LI, 351
denition LICV, 153
homogeneous systems
theorem LIVHS, 155
injective linear transformation
theorem ILTLI, 549
matrices
example LIM32, 353
orthogonal, 198
r and n
theorem LIVRN, 156
linear solve
mathematica, 746
sage, 754
linear system
consistent
theorem RCLS, 58
matrix representation
denition MRLS, 29
notation, 29
linear systems
notation
example MNSLE, 224
example NSLE, 29
linear transformation
polynomials to polynomials
example LTPP, 518
addition
denition LTA, 530
Version 2.30
INDEX 961
theorem MLTLT, 531
theorem SLTLT, 530
as matrix multiplication
example ALTMM, 619
basis of range
example BRLT, 568
checking
example ALT, 516
composition
denition LTC, 532
theorem CLTLT, 533
dened by a matrix
example LTM, 520
dened on a basis
example LTDB1, 526
example LTDB2, 527
example LTDB3, 527
theorem LTDB, 525
denition LT, 515
identity
denition IDLT, 579
injection
denition ILT, 541
inverse
theorem ILTLT, 582
inverse of inverse
theorem IILT, 582
invertible
denition IVLT, 579
example AIVLT, 579
invertible, injective and surjective
theorem ILTIS, 582
Jordan canonical form
theorem JCFLT, 728
kernels of powers
theorem KPLT, 691
linear combination
theorem LTLC, 525
matrix of, 523
example MFLT, 522
example MOLT, 524
not
example NLT, 517
not invertible
example ANILT, 580
notation, 515
polynomials to matrices
example LTPM, 518
rank plus nullity
theorem RPNDD, 588restriction
denition LTR, 711
notation, 711
scalar multiple
example SMLT, 532
scalar multiplication
denition LTSM, 531
spanning range
theorem SSRLT, 567
sum
example STLT, 531
surjection
denition SLT, 559
vector space of, 532
zero vector
theorem LTTZZ, 519
linear transformation inverse
via matrix representation
example ILTVR, 632
linear transformation restriction
on generalized eigenspace
example LTRGE, 711
linear transformations
compositions
example CTLT, 533
from matrices
theorem MBLT, 522
linearly dependent
r<n
example LDRN, 157
via homogeneous system
example LDHS, 156
linearly dependent columns
Archetype A
example LDCAA, 158
linearly dependent set
example LDS, 153
linear combinations within
theorem DLDS, 175
polynomials
example LDP4, 394
linearly independent
crazy vector space
example LIC, 355
extending sets
theorem ELIS, 407
polynomials
example LIP4, 351
via homogeneous system
example LIHS, 155
Version 2.30
962 INDEX
linearly independent columns
Archetype B
example LICAB, 158
linearly independent set
example LIS, 154
example LLDS, 157
LINM (subsection, section LI), 158
LINSB (example), 159
LIP4 (example), 351
LIS (example), 154
LISS (section), 351
LISV (subsection, section LI), 153
LIVHS (theorem), 155
LIVRN (theorem), 156
LLDS (example), 157
LNS (denition), 293
LNS (example), 293
LNS (notation), 293
LNS (subsection, section FS), 293
LNSMS (theorem), 344
lower triangular matrix
denition LTM, 675
LS.MMA (computation, section MMA), 746
LS.SAGE (computation, section SAGE), 754
LSMR (theorem), 932
LSS (denition), 932
LT (acronyms, section IVLT), 601
LT (chapter), 515
LT (denition), 515
LT (notation), 515
LT (section), 515
LT (subsection, section LT), 515
LTA (denition), 530
LTC (denition), 532
LTC (subsection, section LT), 519
LTDB (theorem), 525
LTDB1 (example), 526
LTDB2 (example), 527
LTDB3 (example), 527
LTLC (subsection, section LT), 524
LTLC (theorem), 525
LTM (denition), 675
LTM (example), 520
LTPM (example), 518
LTPP (example), 518
LTR (denition), 711
LTR (notation), 711
LTRGE (example), 711
LTSM (denition), 531
LTTZZ (theorem), 519M (acronyms, section FS), 315
M (archetype), 833
M (chapter), 207
M (denition), 27
M (notation), 27
MA (denition), 207
MA (example), 208
MA (notation), 208
MACN (Property), 758
MAF (Property), 874
MAP (subsection, section SVD), 917
mathematica
gram-schmidt (computation), 748
linear solve (computation), 746
matrix entry (computation), 745
matrix inverse (computation), 749
matrix multiplication (computation), 749
null space (computation), 747
row reduce (computation), 745
transpose of a matrix (computation), 749
vector form of solutions (computation), 747
vector linear combinations (computation), 746
mathematical language
technique L, 766
matrix
addition
denition MA, 207
notation, 208
augmented
denition AM, 30
column space
denition CSM, 271
complex conjugate
example CCM, 212
denition M, 27
equality
denition ME, 207
notation, 207
example AM, 27
identity
denition IM, 84
inverse
denition MI, 244
nonsingular
denition NM, 83
notation, 27
of a linear transformation
theorem MLTCV, 523
product
example PTM, 226
Version 2.30
INDEX 963
example PTMEE, 228
product with vector
denition MVP, 223
rectangular, 83
row space
denition RSM, 278
scalar multiplication
denition MSM, 208
notation, 208
singular, 83
square
denition SQM, 83
submatrices
example SS, 428
submatrix
denition SM, 428
symmetric
denition SYM, 211
transpose
denition TM, 210
unitary
denition UM, 262
unitary is invertible
theorem UMI, 263
zero
denition ZM, 210
matrix addition
example MA, 208
matrix components
notation, 27
matrix entry
mathematica, 745
sage, 753
ti83, 751
ti86, 750
matrix inverse
Archetype B, 249
computation
theorem CINM, 248
mathematica, 749
nonsingular matrix
theorem NI, 261
of a matrix inverse
theorem MIMI, 251
one-sided
theorem OSIS, 260
product
theorem SS, 250
sage, 755
scalar multipletheorem MISM, 252
size 2 matrices
theorem TTMI, 246
transpose
theorem MIT, 251
uniqueness
theorem MIU, 250
matrix multiplication
adjoints
theorem MMAD, 233
associativity
theorem MMA, 231
complex conjugation
theorem MMCC, 232
denition MM, 226
distributivity
theorem MMDAA, 230
entry-by-entry
theorem EMP, 227
identity matrix
theorem MMIM, 229
inner product
theorem MMIP, 231
mathematica, 749
noncommutative
example MMNC, 227
scalar matrix multiplication
theorem MMSMM, 230
systems of linear equations
theorem SLEMM, 224
transposes
theorem MMT, 232
zero matrix
theorem MMZM, 229
matrix product
as composition of linear transformations
example MPMR, 622
matrix representation
basis of eigenvectors
example MRBE, 657
composition of linear transformations
theorem MRCLT, 622
denition MR, 615
invertible
theorem IMR, 630
multiple of a linear transformation
theorem MRMLT, 621
notation, 615
restriction to generalized eigenspace
theorem MRRGE, 719
Version 2.30
964 INDEX
sum of linear transformations
theorem MRSLT, 621
theorem FTMR, 617
upper triangular
theorem UTMR, 676
matrix representations
converting with change-of-basis
example MRCM, 654
example OLTTR, 615
matrix scalar multiplication
example MSM, 208
matrix vector space
dimension
theorem DM, 395
matrix-adjoint product
eigenvalues, eigenvectors
theorem EEMAP, 917
matrix-vector product
example MTV, 223
notation, 223
MBC (example), 224
MBLT (theorem), 522
MC (notation), 27
MCC (subsection, section MO), 212
MCCN (Property), 758
MCF (Property), 873
MCN (denition), 760
MCN (subsection, section CNO), 760
MCSM (example), 272
MCT (theorem), 214
MD (chapter), 903
ME (denition), 207
ME (notation), 207
ME (subsection, section PEE), 484
ME (technique, section PT), 771
ME (theorem), 485
ME.MMA (computation, section MMA), 745
ME.SAGE (computation, section SAGE), 753
ME.TI83 (computation, section TI83), 751
ME.TI86 (computation, section TI86), 750
MEASM (subsection, section MO), 207
MFLT (example), 522
MI (denition), 244
MI (example), 245
MI (notation), 244
MI.MMA (computation, section MMA), 749
MI.SAGE (computation, section SAGE), 755
MICN (Property), 759
MIF (Property), 874
MIMI (theorem), 251MINM (section), 259
MISLE (section), 243
MISM (theorem), 252
MIT (theorem), 251
MIU (theorem), 250
MIVS (example), 609
MLT (subsection, section LT), 520
MLTCV (theorem), 523
MLTLT (theorem), 531
MM (denition), 226
MM (section), 223
MM (subsection, section MM), 226
MM.MMA (computation, section MMA), 749
MMA (section), 745
MMA (theorem), 231
MMAD (theorem), 233
MMCC (theorem), 232
MMDAA (theorem), 230
MMEE (subsection, section MM), 227
MMIM (theorem), 229
MMIP (theorem), 231
MMNC (example), 227
MMSMM (theorem), 230
MMT (theorem), 232
MMZM (theorem), 229
MNEM (theorem), 487
MNSLE (example), 224
MO (section), 207
MOLT (example), 524
more variables than equations
example OSGMD, 61
theorem CMVEI, 61
MPMR (example), 622
MR (denition), 615
MR (notation), 615
MR (section), 615
MRBE (example), 657
MRCB (theorem), 654
MRCLT (diagram), 625
MRCLT (theorem), 622
MRCM (example), 654
MRLS (denition), 29
MRLS (notation), 29
MRMLT (theorem), 621
MRRGE (theorem), 719
MRS (subsection, section CB), 654
MRSLT (theorem), 621
MSCN (example), 760
MSM (denition), 208
MSM (example), 208
Version 2.30
INDEX 965
MSM (notation), 208
MTV (example), 223
multiplicative associativity
complex numbers
Property MACN, 758
multiplicative closure
complex numbers
Property MCCN, 758
eld
Property MCF, 873
multiplicative commutativity
complex numbers
Property CMCN, 758
multiplicative inverse
complex numbers
Property MICN, 759
MVNSE (subsection, section RREF), 27
MVP (denition), 223
MVP (notation), 223
MVP (subsection, section MM), 223
MVSLD (theorem), 158
MWIAA (example), 244
N (archetype), 836
N (subsection, section O), 195
N (technique, section PT), 769
NDMS4 (example), 501
negation of statements
technique N, 769
NEM (theorem), 485
NI (theorem), 261
NIAO (example), 549
NIAQ (example), 541
NIAQR (example), 548
NIDAU (example), 550
nilpotent
linear transformation
denition NLT, 685
NILT (diagram), 542
NJB (theorem), 689
NJB5 (example), 688
NKAO (example), 545
NLT (denition), 685
NLT (example), 517
NLT (section), 685
NLT (subsection, section NLT), 685
NLTFO (subsection, section LT), 530
NM (denition), 83
NM (example), 84
NM (section), 83
NM (subsection, section NM), 83NM (subsection, section OD), 680
NM62 (example), 686
NM64 (example), 685
NM83 (example), 689
NME1 (theorem), 87
NME2 (theorem), 159
NME3 (theorem), 261
NME4 (theorem), 277
NME5 (theorem), 377
NME6 (theorem), 399
NME7 (theorem), 446
NME8 (theorem), 480
NME9 (theorem), 633
NMI (subsection, section MINM), 259
NMLIC (theorem), 159
NMPEM (theorem), 427
NMRRI (theorem), 84
NMTNS (theorem), 86
NMUS (theorem), 86
NOILT (theorem), 588
NOLT (denition), 588
NOLT (notation), 588
NOM (denition), 397
NOM (notation), 397
nonsingular
columns as basis
theorem CNMB, 376
nonsingular matrices
linearly independent columns
theorem NMLIC, 159
nonsingular matrix
Archetype B
example NM, 84
column space, 277
elementary matrices
theorem NMPEM, 427
equivalences
theorem NME1, 87
theorem NME2, 159
theorem NME3, 261
theorem NME4, 277
theorem NME5, 377
theorem NME6, 399
theorem NME7, 446
theorem NME8, 480
theorem NME9, 633
matrix inverse, 261
null space
example NSNM, 86
nullity, 399
Version 2.30
966 INDEX
product of nonsingular matrices
theorem NPNT, 259
rank
theorem RNNM, 399
row-reduced
theorem NMRRI, 84
trivial null space
theorem NMTNS, 86
unique solutions
theorem NMUS, 86
nonsingular matrix, row-reduced
example NSR, 85
norm
example CNSV, 195
inner product, 195
notation, 195
normal matrix
denition NRML, 680
example ANM, 680
orthonormal basis, 683
notation
A, 214
AM, 30
AME, 463
C, 762
CCCV, 191
CCM, 212
CCN, 759
CNA, 758
CNE, 758
CNM, 758
CSM, 271
CV, 28
CVA, 99
CVC, 28
CVE, 98
CVSM, 99
D, 391
DM, 428
DS, 413
ELEM, 424
ES, 761
GES, 707
GME, 463
HI, 890
HID, 890
HP, 889
IE, 717
IM, 84
IP, 192JB, 687
KLT, 545
LNS, 293
LT, 515
LTR, 711
M, 27
MA, 208
MC, 27
ME, 207
MI, 244
MR, 615
MRLS, 29
MSM, 208
MVP, 223
NOLT, 588
NOM, 397
NSM, 73
NV, 195
RLT, 563
RO, 31
ROLT, 588
ROM, 397
RREFA, 33
RSM, 278
SC, 763
SE, 762
SETM, 761
SI, 763
SM, 428
SRM, 926
SSET, 761
SSV, 131
SU, 763
SUV, 197
T, 883
TM, 210
VR, 603
VSCV, 97
VSM, 207
ZCV, 28
ZM, 210
notation for a linear system
example NSE, 12
NPNT (theorem), 259
NRFO (subsection, section MR), 621
NRML (denition), 680
NRREF (example), 33
NS.MMA (computation, section MMA), 747
NSAO (example), 566
NSAQ (example), 559
Version 2.30
INDEX 967
NSAQR (example), 566
NSC2A (example), 336
NSC2S (example), 337
NSC2Z (example), 336
NSDAT (example), 569
NSDS (example), 138
NSE (example), 12
NSEAI (example), 74
NSLE (example), 29
NSLIL (example), 161
NSM (denition), 73
NSM (notation), 73
NSM (subsection, section HSE), 73
NSMS (theorem), 337
NSNM (example), 86
NSNM (subsection, section NM), 85
NSR (example), 85
NSS (example), 85
NSSLI (subsection, section LI), 159
Null space
as a span
example NSDS, 138
null space
Archetype I
example NSEAI, 74
basis
theorem BNS, 160
computation
example CNS1, 74
example CNS2, 75
isomorphic to kernel, 625
linearly independent basis
example LINSB, 159
mathematica, 747
matrix
denition NSM, 73
nonsingular matrix, 86
notation, 73
singular matrix, 85
spanning set
example SSNS, 137
theorem SSNS, 137
subspace
theorem NSMS, 337
null space span, linearly independent
Archetype L
example NSLIL, 161
nullity
computing, 397
injective linear transformationtheorem NOILT, 588
linear transformation
denition NOLT, 588
matrix, 397
denition NOM, 397
notation, 397, 588
square matrix, 398
NV (denition), 195
NV (notation), 195
NVM (theorem), 898
O (archetype), 839
O (Property), 318
O (section), 191
OBC (subsection, section B), 377
OBNM (theorem), 683
OBUTR (theorem), 679
OC (Property), 101
OCN (Property), 759
OD (section), 675
OD (subsection, section OD), 681
OD (theorem), 681
OF (Property), 874
OLTTR (example), 615
OM (Property), 209
one
column vectors
Property OC, 101
complex numbers
Property OCN, 759
eld
Property OF, 874
matrices
Property OM, 209
vectors
Property O, 318
ONFV (example), 202
ONS (denition), 201
ONTV (example), 201
orthogonal
linear independence
theorem OSLI, 198
set
example AOS, 197
set of vectors
denition OSV, 197
vector pairs
denition OV, 196
orthogonal vectors
example TOV, 196
orthonormal
Version 2.30
968 INDEX
denition ONS, 201
matrix columns
example OSMC, 263
orthonormal basis
normal matrix
theorem OBNM, 683
orthonormal diagonalization
theorem OD, 681
orthonormal set
four vectors
example ONFV, 202
three vectors
example ONTV, 201
OSGMD (example), 61
OSIS (theorem), 260
OSLI (theorem), 198
OSMC (example), 263
OSV (denition), 197
OV (denition), 196
OV (subsection, section O), 196
P (appendix), 757
P (archetype), 842
P (technique, section PT), 774
particular solutions
example PSHS, 125
PCNA (theorem), 758
PCVS (example), 326
PD (section), 407
PDM (section), 439
PDM (theorem), 927
PEE (section), 479
PEEF (theorem), 298
PI (denition), 528
PI (subsection, section LT), 528
PI (technique, section PT), 771
PIP (theorem), 196
PM (example), 455
PM (subsection, section EE), 455
PMI (subsection, section MISLE), 250
PMM (subsection, section MM), 229
PMR (subsection, section MR), 625
PNLT (subsection, section NLT), 690
POD (section), 927
polar decomposition
theorem PDM, 927
polynomial
of a matrix
example PM, 455
polynomial vector space
dimensiontheorem DP, 395
positive semi-denite
creating
theorem CPSM, 899
positive semi-denite matrix
denition PSM, 899
eigenvalues
theorem EPSM, 900
practice
technique P, 774
pre-image
denition PI, 528
kernel
theorem KPI, 547
pre-images
example SPIAS, 528
principal axis theorem, 683
product of triangular matrices
theorem PTMT, 675
Property
AA, 317
AAC, 100
AACN, 758
AAF, 873
AAM, 209
AC, 317
ACC, 100
ACCN, 758
ACF, 873
ACM, 209
AI, 318
AIC, 100
AICN, 759
AIF, 874
AIM, 209
C, 317
CACN, 758
CAF, 873
CC, 100
CM, 209
CMCN, 758
CMF, 873
DCN, 759
DF, 874
DMAM, 209
DSA, 318
DSAC, 101
DSAM, 209
DVA, 318
DVAC, 101
Version 2.30
INDEX 969
MACN, 758
MAF, 874
MCCN, 758
MCF, 873
MICN, 759
MIF, 874
O, 318
OC, 101
OCN, 759
OF, 874
OM, 209
SC, 317
SCC, 100
SCM, 209
SMA, 318
SMAC, 100
SMAM, 209
Z, 318
ZC, 100
ZCN, 759
ZF, 874
ZM, 209
PSHS (example), 125
PSHS (subsection, section LC), 124
PSM (denition), 899
PSM (section), 899
PSM (subsection, section PSM), 899
PSM (subsection, section SD), 494
PSMSR (theorem), 923
PSPHS (theorem), 124
PSS (subsection, section SSLE), 13
PSSD (theorem), 410
PSSLS (theorem), 60
PT (section), 765
PTFP (example), 931
PTM (example), 226
PTMEE (example), 228
PTMT (theorem), 675
Q (archetype), 844
R (acronyms, section JCF), 743
R (archetype), 848
R (chapter), 603
R.SAGE (computation, section SAGE), 752
range
full
example FRAN, 564
isomorphic to column space
theorem RCSI, 628
linear transformationexample RAO, 563
notation, 563
of a linear transformation
denition RLT, 563
pre-image
theorem RPI, 568
subspace
theorem RLTS, 564
surjective linear transformation
theorem RSLT, 565
via matrix representation
example RVMR, 629
rank
computing
theorem CRN, 397
linear transformation
denition ROLT, 588
matrix
denition ROM, 397
example RNM, 397
notation, 397, 588
of transpose
example RRTI, 411
square matrix
example RNSM, 398
surjective linear transformation
theorem ROSLT, 588
transpose
theorem RMRT, 411
rank one decomposition
size 2
example ROD2, 905
size 4
example ROD4, 906
theorem ROD, 904
rank+nullity
theorem RPNC, 398
RAO (example), 563
RCLS (theorem), 58
RCSI (theorem), 628
RD (subsection, section VS), 326
RDS (theorem), 417
READ (subsection, section B), 382
READ (subsection, section CB), 668
READ (subsection, section CRS), 283
READ (subsection, section D), 400
READ (subsection, section DM), 433
READ (subsection, section EE), 470
READ (subsection, section FS), 307
READ (subsection, section HSE), 75
Version 2.30
970 INDEX
READ (subsection, section ILT), 551
READ (subsection, section IVLT), 592
READ (subsection, section LC), 126
READ (subsection, section LDS), 184
READ (subsection, section LI), 162
READ (subsection, section LISS), 361
READ (subsection, section LT), 534
READ (subsection, section MINM), 265
READ (subsection, section MISLE), 252
READ (subsection, section MM), 235
READ (subsection, section MO), 215
READ (subsection, section MR), 634
READ (subsection, section NM), 87
READ (subsection, section O), 202
READ (subsection, section PD), 417
READ (subsection, section PDM), 448
READ (subsection, section PEE), 488
READ (subsection, section RREF), 42
READ (subsection, section S), 344
READ (subsection, section SD), 506
READ (subsection, section SLT), 570
READ (subsection, section SS), 141
READ (subsection, section SSLE), 19
READ (subsection, section TSS), 62
READ (subsection, section VO), 101
READ (subsection, section VR), 612
READ (subsection, section VS), 327
READ (subsection, section WILA), 7
reduced row-echelon form
analysis
notation, 33
denition RREF, 33
example NRREF, 33
example RREF, 33
extended
denition EEF, 297
notation
example RREFN, 55
unique
theorem RREFU, 35
reducing a span
example RSC5, 176
relation of linear dependence
denition RLD, 351
denition RLDCV, 153
REM (denition), 31
REMEF (theorem), 34
REMES (theorem), 31
REMRS (theorem), 279
RES (example), 182RGEN (theorem), 716
rings
sage, 752
RLD (denition), 351
RLDCV (denition), 153
RLT (denition), 563
RLT (notation), 563
RLT (subsection, section IS), 711
RLT (subsection, section SLT), 563
RLTS (theorem), 564
RMRT (theorem), 411
RNLT (subsection, section IVLT), 588
RNM (example), 397
RNM (subsection, section D), 397
RNNM (subsection, section D), 398
RNNM (theorem), 399
RNSM (example), 398
RO (denition), 31
RO (notation), 31
RO (subsection, section RREF), 30
ROD (section), 903
ROD (theorem), 904
ROD2 (example), 905
ROD4 (example), 906
ROLT (denition), 588
ROLT (notation), 588
ROM (denition), 397
ROM (notation), 397
ROSLT (theorem), 588
row operations
denition RO, 31
elementary matrices, 424, 425
notation, 31
row reduce
mathematica, 745
sage, 753
ti83, 751
ti86, 750
row space
Archetype I
example RSAI, 278
as column space, 282
basis
example RSB, 374
theorem BRS, 280
matrix, 278
notation, 278
row-equivalent matrices
theorem REMRS, 279
subspace
Version 2.30
INDEX 971
theorem RSMS, 344
row-equivalent matrices
denition REM, 31
example TREM, 31
row space, 279
row spaces
example RSREM, 280
theorem REMES, 31
row-reduce
the verb
denition RR, 42
row-reduced matrices
theorem REMEF, 34
RPI (theorem), 568
RPNC (theorem), 398
RPNDD (theorem), 588
RR (denition), 42
RR.MMA (computation, section MMA), 745
RR.SAGE (computation, section SAGE), 753
RR.TI83 (computation, section TI83), 751
RR.TI86 (computation, section TI86), 750
RREF (denition), 33
RREF (example), 33
RREF (section), 27
RREF (subsection, section RREF), 32
RREFA (notation), 33
RREFN (example), 55
RREFU (theorem), 35
RRTI (example), 411
RS (example), 375
RSAI (example), 278
RSB (example), 374
RSC4 (example), 182
RSC5 (example), 176
RSLT (theorem), 565
RSM (denition), 278
RSM (notation), 278
RSM (subsection, section CRS), 278
RSMS (theorem), 344
RSNS (example), 338
RSREM (example), 280
RT (subsection, section PD), 410
RVMR (example), 629
S (archetype), 851
S (denition), 333
S (example), 83
S (section), 333
SAA (example), 40
SAB (example), 39
SABMI (example), 243SAE (example), 41
sage
eigenspaces (computation), 755
linear solve (computation), 754
matrix entry (computation), 753
matrix inverse (computation), 755
rings (computation), 752
row reduce (computation), 753
transpose of a matrix (computation), 755
vector linear combinations (computation), 755
SAGE (section), 752
SAN (example), 567
SAR (example), 560
SAS (section), 937
SAV (example), 561
SC (denition), 763
SC (example), 764
SC (notation), 763
SC (Property), 317
SC (subsection, section S), 343
SC (subsection, section SET), 762
SC3 (example), 333
SCAA (example), 133
SCAB (example), 135
SCAD (example), 139
scalar closure
column vectors
Property SCC, 100
matrices
Property SCM, 209
vectors
Property SC, 317
scalar multiple
matrix inverse, 252
scalar multiplication
zero scalar
theorem ZSSM, 324
zero vector
theorem ZVSM, 325
zero vector result
theorem SMEZV, 326
scalar multiplication associativity
column vectors
Property SMAC, 100
matrices
Property SMAM, 209
vectors
Property SMA, 318
SCB (theorem), 656
SCC (Property), 100
Version 2.30
972 INDEX
SCM (Property), 209
SD (section), 493
SDS (example), 413
SE (denition), 762
SE (notation), 762
secret sharing
6 ways
example SS6W, 938
SEE (example), 453
SEEF (example), 297
SER (theorem), 494
set
cardinality
denition C, 762
example CS, 762
notation, 762
complement
denition SC, 763
example SC, 764
notation, 763
denition SET, 761
empty
denition ES, 761
equality
denition SE, 762
notation, 762
intersection
denition SI, 763
example SI, 763
notation, 763
membership
example SETM, 761
notation, 761
size, 762
subset, 761
union
denition SU, 763
example SU, 763
notation, 763
SET (denition), 761
SET (section), 761
SETM (example), 761
SETM (notation), 761
shoes, 250
SHS (subsection, section HSE), 71
SI (denition), 763
SI (example), 763
SI (notation), 763
SI (subsection, section IVLT), 586
SIM (denition), 493similar matrices
equal eigenvalues
example EENS, 496
eual eigenvalues
theorem SMEE, 495
example SMS3, 494
example SMS5, 493
similarity
denition SIM, 493
equivalence relation
theorem SER, 494
singular matrix
Archetype A
example S, 83
null space
example NSS, 85
singular matrix, row-reduced
example SRR, 85
singular value decomposition
theorem SVD, 921
singular values
denition SV, 921
SLE (acronyms, section NM), 95
SLE (chapter), 3
SLE (denition), 11
SLE (subsection, section SSLE), 11
SLELT (subsection, section IVLT), 591
SLEMM (theorem), 224
SLSLC (theorem), 112
SLT (denition), 559
SLT (section), 559
SLTB (theorem), 568
SLTD (subsection, section SLT), 569
SLTD (theorem), 569
SLTLT (theorem), 530
SM (denition), 428
SM (notation), 428
SM (subsection, section SD), 493
SM2Z7 (example), 876
SM32 (example), 341
SMA (Property), 318
SMAC (Property), 100
SMAM (Property), 209
SMEE (theorem), 495
SMEZV (theorem), 326
SMLT (example), 532
SMS (theorem), 211
SMS3 (example), 494
SMS5 (example), 493
SMZD (theorem), 445
Version 2.30
INDEX 973
SMZE (theorem), 480
SNCM (theorem), 261
SO (subsection, section SET), 763
socks, 250
SOL (subsection, section B), 385
SOL (subsection, section CB), 670
SOL (subsection, section CRS), 288
SOL (subsection, section D), 403
SOL (subsection, section DM), 436
SOL (subsection, section EE), 473
SOL (subsection, section F), 881
SOL (subsection, section FS), 310
SOL (subsection, section HSE), 79
SOL (subsection, section ILT), 555
SOL (subsection, section IVLT), 596
SOL (subsection, section LC), 129
SOL (subsection, section LDS), 187
SOL (subsection, section LI), 167
SOL (subsection, section LISS), 364
SOL (subsection, section LT), 537
SOL (subsection, section MINM), 268
SOL (subsection, section MISLE), 256
SOL (subsection, section MM), 239
SOL (subsection, section MO), 219
SOL (subsection, section MR), 638
SOL (subsection, section NM), 90
SOL (subsection, section O), 204
SOL (subsection, section PD), 419
SOL (subsection, section PDM), 450
SOL (subsection, section PEE), 490
SOL (subsection, section RREF), 48
SOL (subsection, section S), 347
SOL (subsection, section SD), 508
SOL (subsection, section SLT), 574
SOL (subsection, section SS), 145
SOL (subsection, section SSLE), 23
SOL (subsection, section T), 888
SOL (subsection, section TSS), 67
SOL (subsection, section VO), 106
SOL (subsection, section VR), 614
SOL (subsection, section VS), 330
SOL (subsection, section WILA), 9
solution set
Archetype A
example SAA, 40
archetype E
example SAE, 41
theorem PSPHS, 124
solution set of a linear system
denition SSSLE, 12solution sets
possibilities
theorem PSSLS, 60
solution to a linear system
denition SSLE, 12
solution vector
denition SOLV, 29
SOLV (denition), 29
solving homogeneous system
Archetype A
example HISAA, 72
Archetype B
example HUSAB, 72
Archetype D
example HISAD, 72
solving nonlinear equations
example STNE, 11
SP4 (example), 335
span
basic
example ABS, 131
basis
theorem BS, 180
denition SS, 339
denition SSCV, 131
improved
example IAS, 281
notation, 131
reducing
example RSC4, 182
reduction
example RS, 375
removing vectors
example COV, 177
reworking elements
example RES, 182
set of polynomials
example SSP, 340
subspace
theorem SSS, 339
span of columns
Archetype A
example SCAA, 133
Archetype B
example SCAB, 135
Archetype D
example SCAD, 139
spanning set
crazy vector space
example SSC, 358
Version 2.30
974 INDEX
denition TSVS, 356
matrices
example SSM22, 357
more vectors
theorem SSLD, 391
polynomials
example SSP4, 356
SPIAS (example), 528
SQM (denition), 83
square root
eigenvalues, eigenspaces
theorem EESR, 924
matrix
denition SRM, 926
notation, 926
positive semi-denite matrix
theorem PSMSR, 923
unique
theorem USR, 926
SR (section), 923
SRM (denition), 926
SRM (notation), 926
SRM (subsection, section SR), 923
SRR (example), 85
SS (denition), 339
SS (example), 428
SS (section), 131
SS (subsection, section LISS), 355
SS (theorem), 250
SS6W (example), 938
SSC (example), 358
SSCV (denition), 131
SSET (denition), 761
SSET (example), 761
SSET (notation), 761
SSLD (theorem), 391
SSLE (denition), 12
SSLE (section), 11
SSM22 (example), 357
SSNS (example), 137
SSNS (subsection, section SS), 136
SSNS (theorem), 137
SSP (example), 340
SSP4 (example), 356
SSRLT (theorem), 567
SSS (theorem), 339
SSSLE (denition), 12
SSSLT (subsection, section SLT), 567
SSV (notation), 131
SSV (subsection, section SS), 131standard unit vector
notation, 197
starting proofs
technique GS, 767
STLT (example), 531
STNE (example), 11
SU (denition), 763
SU (example), 763
SU (notation), 763
submatrix
notation, 428
subset
denition SSET, 761
notation, 761
subspace
as null space
example RSNS, 338
characterized
example ASC, 609
denition S, 333
inP4
example SP4, 335
not, additive closure
example NSC2A, 336
not, scalar closure
example NSC2S, 337
not, zero vector
example NSC2Z, 336
testing
theorem TSS, 334
trivial
denition TS, 337
verication
example SC3, 333
example SM32, 341
subspaces
equal dimension
theorem EDYES, 410
surjective
Archetype N
example SAN, 567
example SAR, 560
not
example NSAQ, 559
example NSAQR, 566
not, Archetype O
example NSAO, 566
not, by dimension
example NSDAT, 569
polynomials to matrices
Version 2.30
INDEX 975
example SAV, 561
surjective linear transformation
bases
theorem SLTB, 568
surjective linear transformations
dimension
theorem SLTD, 569
SUV (denition), 197
SUV (notation), 197
SUVB (theorem), 371
SUVOS (example), 197
SV (denition), 921
SVD (section), 917
SVD (subsection, section SVD), 920
SVD (theorem), 921
SVP4 (example), 409
SYM (denition), 211
SYM (example), 211
symmetric matrices
theorem SMS, 211
symmetric matrix
example SYM, 211
system of equations
vector equality
example VESE, 98
system of linear equations
denition SLE, 11
T (archetype), 854
T (denition), 883
T (notation), 883
T (part), 873
T (section), 883
T (technique, section PT), 766
TCSD (example), 432
TD (section), 909
TD (subsection, section TD), 909
TD (theorem), 909
TD4 (example), 911
TDEE (theorem), 913
TDEE6 (example), 915
TDSSE (example), 912
TDSSE (subsection, section TD), 912
technique
C, 768
CD, 770
CP, 769
CV, 769
D, 765
DC, 772
E, 768GS, 767
I, 772
L, 766
LC, 774
ME, 771
N, 769
P, 774
PI, 771
T, 766
U, 771
theorem
AA, 215
AIP, 233
AISM, 325
AIU, 324
AMA, 214
AMSM, 214
BCS, 274
BIS, 394
BNS, 160
BRS, 280
BS, 180
CB, 649
CCM, 213
CCRA, 759
CCRM, 760
CCT, 760
CFDVS, 608
CFNLT, 694
CHT, 740
CILTI, 551
CINM, 248
CIVLT, 585
CLI, 609
CLTLT, 533
CMVEI, 61
CNMB, 376
COB, 378
CPSM, 899
CRMA, 213
CRMSM, 213
CRN, 397
CRSM, 191
CRVA, 191
CSCS, 272
CSLTS, 570
CSMS, 343
CSNM, 277
CSRN, 59
CSRST, 282
Version 2.30
976 INDEX
CSS, 610
CUMOS, 263
DC, 497
DCM, 395
DCP, 484
DEC, 431
DED, 501
DEM, 444
DEMMM, 445
DER, 429
DERC, 441
DFS, 412
DGES, 727
DIM, 443
DLDS, 175
DM, 395
DMFE, 499
DMHP, 891
DMMP, 892
DMST, 429
DNLT, 691
DP, 395
DRCM, 440
DRCMA, 441
DRCS, 439
DRMM, 447
DSD, 416
DSFB, 413
DSFOS, 414
DSLI, 416
DSZI, 415
DSZV, 414
DT, 430
DVM, 895
DZRC, 439
EDELI, 479
EDYES, 410
EEMAP, 917
EER, 659
EESR, 924
EIM, 482
EIS, 705
ELIS, 407
EMDRO, 425
EMHE, 457
EMMVP, 225
EMN, 427
EMNS, 462
EMP, 227
EMRCP, 461EMS, 461
ENLT, 690
EOMP, 481
EOPSS, 14
EPM, 481
EPSM, 900
ERMCP, 483
ESMM, 481
ETM, 483
FIMP, 875
FS, 299
FTMR, 617
FVCS, 60
G, 407
GEK, 708
GESD, 721
GESIS, 707
GSP, 199
HMIP, 234
HMOE, 488
HMRE, 487
HMVEI, 73
HPC, 889
HPDAA, 891
HPHI, 890
HPHID, 890
HPSMM, 891
HSC, 71
ICBM, 649
ICLT, 585
IFDVS, 609
IILT, 582
ILTB, 550
ILTD, 550
ILTIS, 582
ILTLI, 549
ILTLT, 582
IMILT, 633
IMR, 630
IP, 931
IPAC, 194
IPN, 195
IPSM, 194
IPVA, 193
ISRN, 59
ITMT, 676
IVSED, 587
JCFLT, 728
KILT, 548
KLTS, 546
Version 2.30
INDEX 977
KNSI, 625
KPI, 547
KPIS, 705
KPLT, 691
KPNLT, 692
LIVHS, 155
LIVRN, 156
LNSMS, 344
LSMR, 932
LTDB, 525
LTLC, 525
LTTZZ, 519
MBLT, 522
MCT, 214
ME, 485
MIMI, 251
MISM, 252
MIT, 251
MIU, 250
MLTCV, 523
MLTLT, 531
MMA, 231
MMAD, 233
MMCC, 232
MMDAA, 230
MMIM, 229
MMIP, 231
MMSMM, 230
MMT, 232
MMZM, 229
MNEM, 487
MRCB, 654
MRCLT, 622
MRMLT, 621
MRRGE, 719
MRSLT, 621
MVSLD, 158
NEM, 485
NI, 261
NJB, 689
NME1, 87
NME2, 159
NME3, 261
NME4, 277
NME5, 377
NME6, 399
NME7, 446
NME8, 480
NME9, 633
NMLIC, 159NMPEM, 427
NMRRI, 84
NMTNS, 86
NMUS, 86
NOILT, 588
NPNT, 259
NSMS, 337
NVM, 898
OBNM, 683
OBUTR, 679
OD, 681
OSIS, 260
OSLI, 198
PCNA, 758
PDM, 927
PEEF, 298
PIP, 196
PSMSR, 923
PSPHS, 124
PSSD, 410
PSSLS, 60
PTMT, 675
RCLS, 58
RCSI, 628
RDS, 417
REMEF, 34
REMES, 31
REMRS, 279
RGEN, 716
RLTS, 564
RMRT, 411
RNNM, 399
ROD, 904
ROSLT, 588
RPI, 568
RPNC, 398
RPNDD, 588
RREFU, 35
RSLT, 565
RSMS, 344
SCB, 656
SER, 494
SLEMM, 224
SLSLC, 112
SLTB, 568
SLTD, 569
SLTLT, 530
SMEE, 495
SMEZV, 326
SMS, 211
Version 2.30
978 INDEX
SMZD, 445
SMZE, 480
SNCM, 261
SS, 250
SSLD, 391
SSNS, 137
SSRLT, 567
SSS, 339
SUVB, 371
SVD, 921
TD, 909
TDEE, 913
technique T, 766
TIST, 884
TL, 883
TMA, 211
TMSM, 212
TSE, 884
TSRM, 884
TSS, 334
TT, 212
TTMI, 246
UMCOB, 380
UMI, 263
UMPIP, 264
USR, 926
UTMR, 676
VFSLS, 118
VRI, 607
VRILT, 608
VRLT, 603
VRRB, 360
VRS, 608
VSLT, 532
VSPCV, 100
VSPM, 209
ZSSM, 324
ZVSM, 325
ZVU, 324
ti83
matrix entry (computation), 751
row reduce (computation), 751
vector linear combinations (computation), 752
TI83 (section), 751
ti86
matrix entry (computation), 750
row reduce (computation), 750
transpose of a matrix (computation), 751
vector linear combinations (computation), 750
TI86 (section), 750TIS (example), 703
TIST (theorem), 884
TIVS (example), 609
TKAP (example), 546
TL (theorem), 883
TLC (example), 109
TM (denition), 210
TM (example), 210
TM (notation), 210
TM (subsection, section OD), 675
TM.MMA (computation, section MMA), 749
TM.SAGE (computation, section SAGE), 755
TM.TI86 (computation, section TI86), 751
TMA (theorem), 211
TMP (example), 4
TMSM (theorem), 212
TOV (example), 196
trace
denition T, 883
linearity
theorem TL, 883
matrix multiplication
theorem TSRM, 884
notation, 883
similarity
theorem TIST, 884
sum of eigenvalues
theorem TSE, 884
trail mix
example TMP, 4
transpose
matrix scalar multiplication
theorem TMSM, 212
example TM, 210
matrix addition
theorem TMA, 211
matrix inverse, 251
notation, 210
scalar multiplication, 212
transpose of a matrix
mathematica, 749
sage, 755
ti86, 751
transpose of a transpose
theorem TT, 212
TREM (example), 31
triangular decomposition
entry by entry, size 6
example TDEE6, 915
entry by entry
Version 2.30
INDEX 979
theorem TDEE, 913
size 4
example TD4, 911
solving systems of equations
example TDSSE, 912
theorem TD, 909
triangular matrix
inverse
theorem ITMT, 676
trivial solution
system of equations
denition TSHSE, 71
TS (denition), 337
TS (subsection, section S), 334
TSE (theorem), 884
TSHSE (denition), 71
TSM (subsection, section MO), 210
TSRM (theorem), 884
TSS (section), 55
TSS (subsection, section S), 338
TSS (theorem), 334
TSVS (denition), 356
TT (theorem), 212
TTMI (theorem), 246
TTS (example), 13
typical systems, 2 2
example TTS, 13
U (archetype), 856
U (technique, section PT), 771
UM (denition), 262
UM (subsection, section MINM), 262
UM3 (example), 262
UMCOB (theorem), 380
UMI (theorem), 263
UMPIP (theorem), 264
unique solution, 3 3
example US, 16
example USR, 32
uniqueness
technique U, 771
unit vectors
basis
theorem SUVB, 371
denition SUV, 197
orthogonal
example SUVOS, 197
unitary
permutation matrix
example UPM, 262
size 3example UM3, 262
unitary matrices
columns
theorem CUMOS, 263
unitary matrix
inner product
theorem UMPIP, 264
UPM (example), 262
upper triangular matrix
denition UTM, 675
US (example), 16
USR (example), 32
USR (theorem), 926
UTM (denition), 675
UTMR (subsection, section OD), 676
UTMR (theorem), 676
V (acronyms, section O), 205
V (archetype), 858
V (chapter), 97
VA (example), 99
Vandermonde matrix
denition VM, 895
vandermonde matrix
determinant
theorem DVM, 895
nonsingular
theorem NVM, 898
size 4
example VM4, 895
VEASM (subsection, section VO), 98
vector
addition
denition CVA, 98
column
denition CV, 27
equality
denition CVE, 98
notation, 98
inner product
denition IP, 192
norm
denition NV, 195
notation, 28
of constants
denition VOC, 28
product with matrix, 223, 226
scalar multiplication
denition CVSM, 99
vector addition
example VA, 99
Version 2.30
980 INDEX
vector component
notation, 28
vector form of solutions
Archetype D
example VFSAD, 114
Archetype I
example VFSAI, 121
Archetype L
example VFSAL, 122
example VFS, 115
mathematica, 747
theorem VFSLS, 118
vector linear combinations
mathematica, 746
sage, 755
ti83, 752
ti86, 750
vector representation
example AVR, 359
example VRC4, 604
injective
theorem VRI, 607
invertible
theorem VRILT, 608
linear transformation
denition VR, 603
notation, 603
theorem VRLT, 603
surjective
theorem VRS, 608
theorem VRRB, 360
vector representations
polynomials
example VRP2, 606
vector scalar multiplication
example CVSM, 100
vector space
characterization
theorem CFDVS, 608
column vectors
denition VSCV, 97
denition VS, 317
innite dimension
example VSPUD, 396
linear transformations
theorem VSLT, 532
over integers mod 5
example VSIM5, 875
vector space of column vectors
notation, 97vector space of functions
example VSF, 321
vector space of innite sequences
example VSIS, 320
vector space of matrices
denition VSM, 207
example VSM, 319
notation, 207
vector space of polynomials
example VSP, 319
vector space properties
column vectors
theorem VSPCV, 100
matrices
theorem VSPM, 209
vector space, crazy
example CVS, 322
vector space, singleton
example VSS, 321
vector spaces
isomorphic
denition IVS, 586
theorem IFDVS, 609
VESE (example), 98
VFS (example), 115
VFSAD (example), 114
VFSAI (example), 121
VFSAL (example), 122
VFSLS (theorem), 118
VFSS (subsection, section LC), 113
VFSS.MMA (computation, section MMA), 747
VLC.MMA (computation, section MMA), 746
VLC.SAGE (computation, section SAGE), 755
VLC.TI83 (computation, section TI83), 752
VLC.TI86 (computation, section TI86), 750
VM (denition), 895
VM (section), 895
VM4 (example), 895
VO (section), 97
VOC (denition), 28
VR (denition), 603
VR (notation), 603
VR (section), 603
VR (subsection, section LISS), 359
VRC4 (example), 604
VRI (theorem), 607
VRILT (theorem), 608
VRLT (theorem), 603
VRP2 (example), 606
VRRB (theorem), 360
Version 2.30
INDEX 981
VRS (theorem), 608
VS (acronyms, section PD), 421
VS (chapter), 317
VS (denition), 317
VS (section), 317
VS (subsection, section VS), 317
VSCV (denition), 97
VSCV (example), 319
VSCV (notation), 97
VSF (example), 321
VSIM5 (example), 875
VSIS (example), 320
VSLT (theorem), 532
VSM (denition), 207
VSM (example), 319
VSM (notation), 207
VSP (example), 319
VSP (subsection, section MO), 209
VSP (subsection, section VO), 100
VSP (subsection, section VS), 323
VSPCV (theorem), 100
VSPM (theorem), 209
VSPUD (example), 396
VSS (example), 321
W (archetype), 860
WILA (section), 3
X (archetype), 862
Z (Property), 318
ZC (Property), 100
ZCN (Property), 759
ZCV (denition), 28
ZCV (notation), 28
zero
complex numbers
Property ZCN, 759
eld
Property ZF, 874
zero column vector
denition ZCV, 28
notation, 28
zero matrix
notation, 210
zero vector
column vectors
Property ZC, 100
matrices
Property ZM, 209
uniquetheorem ZVU, 324
vectors
Property Z, 318
ZF (Property), 874
ZM (denition), 210
ZM (notation), 210
ZM (Property), 209
ZNDAB (example), 446
ZSSM (theorem), 324
ZVSM (theorem), 325
ZVU (theorem), 324
Version 2.30